Vertical Inference

Vision

Today most AI inference is sold by the “token”. A token is a unit of effort that may or may not have contributed to an advancement. We at Vertical Inference are striving for a future where inference is by the task. We feel a task completion should be the unit of value not token, token count should be irrelevant. In labor markets, things move the same way, a persons effort is valued for their outcome and not by amount of “thought” put into a task, why should AGI systems be treated differently?

We’re rethinking the entire inference stack from how agents run → down to the datacenters and silicon. In order for the task to be the unit of value and token consumption to be inconsequential, inference should be extremely efficient 50-60x compared to current. Long-running agents need inference that adapts to the task, lossless memory, verification of task completion, and automated research that develops custom kernels for each workload. Delivering these capabilities at scale requires portable datacenters built for distributed inference and sovereign AI, alongside purpose-built chips - all designed to maximize task success within a budget of time, compute, and power. We envision a world where inference is accountable for completing tasks, with pricing tied to results rather than tokens consumed. Making that possible requires rethinking the entire stack - from agents to datacenters to silicon. That’s what we’re building – We are building the new inference stack for the AGI era.

Our mission: build an inference stack tailored to each task that operates at the Pareto frontier of quality, latency and cost.

Why AI inference is hard?

Nothing about Inference is deterministic, that makes AI infrastructure harder to build. A one-word question can take a thousand tokens to answer. Everything is probabilistic - from the question complexity, to the models behavior to that question, repeated invocations for same input only reinforce the probabilistic nature. In classical systems design optimisation assume the opposite. The work is known, system can be deterministically profiled, bottlenecks found, removed and the output stays correct. AI inference makes both the workload and the consequences of optimization uncertain. The amount of work is discovered while doing it. And almost every change that makes a model faster can also make it wrong, so speed and correctness stop being independent.

Our goal is a future where inference adapts to the task as it unfolds, allocating the right computation at each step to deliver a reliable result.

Why now?

Token volume is exploding. So what? In our task-level thinking, the implications are different. The interesting question is what happens when the price per token falls far enough. At some point it becomes cheaper to attempt a run twenty times and keep the best result than to attempt it once and accept what you got. Every engineering design could compete against a million alternatives. Every testable scientific hypothesis could trigger a 1000 automated searches. Entire businesses could continuously experiment with better ways to operate. Systems like AlphaEvolve already generate, evaluate and select among candidates. Making this possible at production scale requires an inference stack built for sustained search, verification and task completion.

When that becomes economical, the posit that unit of demand changes: people can ask for a problem to be solved and give the system room to find a solution, as token cost is not the consideration. Vertical Inference is building the infrastructure to make that practical - from inference that adapts to the task-at-hand to portable and distributed datacenters and chips designed to sustain it.

Our Approach

Our bet is that inference should be optimized around the task. We start with an inference stack tailored to the task, the user’s SLOs and explicit success criteria. Our system learns the workload’s execution characteristics, then develops custom kernels, adapts KV-cache and memory reuse strategies and schedules compute across attempts and verification. After deployment, it continually identifies and addresses efficiency gaps, using accumulated workload knowledge to push inference toward the Pareto frontier of quality, latency and cost. Our stack is designed to run across NVIDIA, AMD and Intel GPUs, from large datacenters to smaller sovereign deployments. This is our immediate focus.

For long-running agents, compute, memory and communication must be designed together around how tasks execute. We’re taking a top-down approach: deploying our inference stack first, then extending into portable, distributed datacenters and purpose-built silicon. The objective at every layer is the same: complete more tasks successfully within the required quality, latency and resource budgets.

Why us?

Our founding team brings AI researchers and hardware scientists with experience working on complex AI problems at Amazon, Meta and Microsoft. We’ve built and deployed AI products impacting hundreds of millions, and our expertise spans the software and hardware this vision requires.

Contact

Tell us what you are working on.

Operators and researchers welcome. We read everything that comes in.