Icosa Computing

Local AI that learns the work.

Everything you love about your AI assistant, without usage limits or bills that grow out of control. Icosa runs on your computer, not in the cloud.

  • No token costs. Use it as much as you want.
  • Super fast. Answers instantly, right on your device.
  • Private. Your files never leave your computer.

Trusted by

FujitsuNECToshibaNASANational Science FoundationUSRAAlgoquantUNLU & CoBener Law OfficeDualityCreative Destruction LabFujitsuNECToshibaNASANational Science FoundationUSRAAlgoquantUNLU & CoBener Law OfficeDualityCreative Destruction LabFujitsuNECToshibaNASANational Science FoundationUSRAAlgoquantUNLU & CoBener Law OfficeDualityCreative Destruction LabFujitsuNECToshibaNASANational Science FoundationUSRAAlgoquantUNLU & CoBener Law OfficeDualityCreative Destruction Lab
$0
Token costs
100%
Compute on your hardware
U.S. National Science Foundation
$1.4M
NSF research funding →

Products

Zeno

macOS

The local AI workspace that learns your work. Zeno adapts to your documents, code, and terminology, entirely on your machine, with no model selection or setup.

Apply for early access

Launching August 19

--days
--hrs
--min
--sec

LM Shop

Web

The no-code platform for training your own AI models. Upload your documents and get a custom model out, fast. Download it and run it locally with llama.cpp or any app you like, or directly with Zeno.

Use LM Shop

How it works

  1. 1. Upload your docs
  2. 2. Train with no code, in minutes
  3. 3. Download & run locally

Your hardware does the compute.
Usage never adds a bill.

01

Zero token costs

Agentic work turns one request into many model calls. On rented AI, that compounds a bill. On Icosa, it costs nothing.

02

Private by design

Files, prompts, and model weights never leave machines you control. No cloud dependency and no vendor shutdowns.

03

Expert accuracy

Combinatorial Reasoning scores many answer paths and selects the strongest, closing the gap with hosted frontier models.

Research

Many reasoning paths. One better answer.

Combinatorial Reasoning is a physics-based optimization method developed with NASA and funded by the National Science Foundation. It samples many reasoning traces, scores them, and selects the strongest set before answering, so small local models perform like much larger ones on the work they know.

Combinatorial Reasoning pipeline from the Icosa paper: sampled reasons are mapped to a QUBO problem, optimized, and the selected reasons feed the final prompt

Esencan et al., arXiv:2407.00071

Team

Deep-tech researchers, shipping product.

Mert Esencan

Founder & CEO

Stanford BS/MS, Oxford quantum PhD. Ex-Fidelity researcher.

Can Unlu

Software

UVA Math & CS. Founding engineer of the optimization stack.

Tarun Advaith

Research

Waterloo PhD. Drives Icosa research; author of ML and quantum papers.

Baran Melik

Go-to-Market

Duke MBA. Drives go-to-market and revenue operations.

Alan Ho

Product

Ex-Google Head of Quantum AI Product. Co-founder of Qolab.

Mark Gurvis

Revenue

Ex-Google Head of Strategic Platforms. Scales revenue.

FAQ

Frequently asked questions

Local AI, private and secure AI, on-prem deployment, data sovereignty, small language models, and how Combinatorial Reasoning works.

Combinatorial Reasoning is Icosa's physics-based optimization method, developed with NASA and USRA and supported by $1.4 million in funding from the U.S. National Science Foundation. Instead of accepting a model's first chain of thought, it samples many candidate reasoning paths, scores them, and selects the strongest subset before the model answers. This is how small local models reach reasoning quality normally associated with much larger hosted systems.

When the model runs entirely on your own hardware, yes. There is no API call, so nothing is transmitted anywhere. With Icosa, your prompts, documents, and the model weights themselves stay on machines you control, which means no third party ever receives your data in the first place.

Anything you type or upload passes through servers you do not control. That data can be retained, logged, reviewed by staff, or used to train future models, depending on the provider's terms. It can also be exposed in a breach, and shared conversation links have been indexed by search engines. Beyond privacy, the vendor can change pricing, restrict access, or shut a model down without notice.

It depends entirely on where inference happens. With cloud AI, your data sits inside someone else's system under their retention and access policies. With local AI, you keep custody: your compliance team can audit exactly where data lives, there is no third-party processor to vet, and there is no cloud breach that can expose your files.

Data sovereignty is the principle that data is subject to the laws of the country where it is stored and processed. Cloud AI can move your data across borders and into jurisdictions with different rules, which creates real problems under regimes like GDPR. Running models on your own infrastructure keeps data physically inside your jurisdiction and under your own governance.

On-prem (on-premises) AI is AI deployed inside your own infrastructure, such as a workstation, an office server, or your own data center, rather than rented from a cloud provider. Icosa specializes in on-prem AI for regulated firms: small, domain-expert models that run on a few GPUs or even a MacBook, instead of frontier-scale models that need a 1.5 TB cluster.

Open-weight models are models whose trained parameters are published for anyone to download, so you can run them, fine-tune them, and inspect them on your own hardware. They now catch frontier releases within weeks. Icosa turns open-weight models into domain experts for your specific work, so you get frontier-level results without depending on any single vendor.

Icosa removes the cloud from the loop at every stage. LM Shop trains a custom model from your own documents, and you download that model and own it outright. Zeno then runs it on your Mac, entirely offline. Combinatorial Reasoning keeps the quality high enough that going private costs you nothing in capability.

An SLM is a language model compact enough to run on local hardware, typically billions of parameters instead of trillions. A general frontier model spends most of its capacity on knowledge your team never uses. An SLM fine-tuned on your domain matches frontier output on the work that matters, while being faster, cheaper, and fully private.

The inference is. Cloud AI charges per token, so heavy use means growing bills. With local AI, your computer does the work, so usage costs nothing beyond electricity. An agent can run thousands of steps overnight for pennies, with no subscription scaling against you, no token meter, and no rate limits.

Zeno is Icosa's local AI workspace for macOS, launching August 19, 2026. It learns your documents, code, and terminology on your machine, with nothing to configure and no model selection to manage. All you need is an Apple Silicon Mac, because Icosa's models are optimized for local inference.

AI shouldn't be rented by the token.

Bring owned, private AI to your team, on the hardware you already have.