Industry Solutions Banking & Finance Healthcare Manufacturing Legal Government & Defense How It Works Cost Savings Knowledge Blog About Request Demo
6 min read

Two Thirds of AI Work Now Runs on Smaller Models

OpenRouter's August charts show the money and the work pulling apart. Three labs take two thirds of the spend while smaller models do two thirds of the work, and that work is the kind you can own.

AI spending flowing to three large labs while the token volume flows to a row of smaller models
Two and a half years of OpenRouter traffic: the money went one way, the work went the other.

In January 2024, Anthropic, Google, and OpenAI processed 61% of the tokens moving through OpenRouter. This month their share is 35%. The money moved less. The same three labs still collect 69% of what gets spent there.

Those two lines come from charts published in August 2026 by Peter Walker, who runs insights at OpenRouter. Put together, they say the work is moving to smaller models, and it is moving much faster than the money.

The quick version

AI usage through OpenRouter grew 600x in two years. Most of the money now goes to agents and code, work that runs all day on a meter. The three biggest labs still collect 69% of the spend, but smaller labs handle 65% of the tokens. The heavy work fits smaller models, and the best smaller models are open weight, which means you can own them. We made the case for owning in July. These numbers show the market working through the same choice.

What does the OpenRouter data show?

OpenRouter is a switchboard for AI. A developer sends a request, picks a model, and the platform routes the call. Hundreds of models compete there on the same terms, so its records show which ones get chosen for real work.

The scale has changed fast. Token volume through the platform grew 600x in 24 months. In the first 19 days of August alone, users pushed 246 trillion tokens through it, and weekly volume is closing in on 100 trillion.

It is one platform, and a particular slice of the market. OpenRouter serves developers who pick their own models, so a company buying straight from one lab under contract is invisible here. Keep that in mind with every number below. A quarter of a quadrillion tokens in 19 days is still a decent sample.

What are companies paying AI to do?

Walker's August breakdown sorts the spending by task. Code work took 31.6% of the money: writing new code at 10.9%, debugging at 5.9%, and a tail of smaller jobs. Agent work took 28.7%, and one line inside it towers over the rest. Workflow execution, agents carrying out multi-step jobs from start to finish, was 19% of all spending on its own. Add code and agents together and they take 60 cents of every dollar.

The rest is everyday office work. Sorting content into buckets took 9.8%. Pulling data out of documents took 5.7%. Writing took 4.2%, and answering questions took 4%.

Read that list as a budget owner. This is production work. It runs on schedules and queues, all day, with no one waiting for a person to type a prompt. A bill that meters by the token grows with every job the system picks up, which is a big part of why AI costs are so hard to forecast.

Who is doing the work?

Now set the spending chart against the token chart. On spending, the big three went from 81% in January 2024 to 69% today. On tokens, they went from 61% to 35%. Both lines fall, but the token line falls off a cliff. Every other lab combined now handles 65% of the work on 31% of the money.

The gap has a plain cause. Frontier models charge far more per token, so the big labs can lose most of the volume and keep most of the revenue. Each token they serve earns a premium price. Each token that moves to a smaller lab costs the customer a fraction of that.

Why did the work move to smaller models?

Look at what the work is. Sorting records. Pulling fields out of documents. Running one defined step in a workflow. Reviewing code against a rule set. These are bounded tasks with checkable answers, and a smaller model tuned to a bounded task handles it well, at a fraction of the frontier price. Open-weight models closed most of the quality gap for this kind of work, and the cost gap kept widening. We covered that shift in our State of Enterprise AI review, where open-weight serving had reached up to 60x cheaper inference.

OpenRouter makes the choice visible. A developer there can put any model on any task and watch the results. The token share shows what they keep choosing.

What does the split mean for a regulated business?

The tasks eating these budgets are the same ones eating yours. Document work. Sorting. Agents running steps against internal systems. The OpenRouter charts say that work fits smaller models, and that matters in a bank or a hospital for a second reason. Many of the best smaller models are open weight. You can download the weights, run them on hardware inside your own network, and the work never leaves.

Those 65% of tokens still flow through the cloud, though. Renting a smaller model through an API is cheaper renting, and the data still travels. The same open-weight models can run inside your own walls, where the meter goes away and the files stay put. That is the difference between a cheaper meter and no meter at all.

We laid out the own-versus-rent decision in July, when Palantir's CEO was calling token pricing a wealth tax on live TV. That piece gives you the framework: who holds the weights, whether you can trace an answer, and what happens to the bill as use grows. Walker's charts read like the market answering the first of those questions, one task at a time.

What should you do with the numbers?

Build your own version of Walker's chart. Your AI tools keep usage logs, so sort last month's spend by task the way OpenRouter does. Then draw one line through the list. Which tasks need a frontier model? Which are bounded jobs with checkable answers?

For the bounded list, price an owned, open-weight model doing the work inside your network at a fixed platform cost. Our cost savings breakdown shows the math. Then check the agent line in your bill, because that is the line that grows on its own.

On OpenRouter, smaller models already handle 65% of the tokens. Find out how much of your work is bounded.

Go deeper

The decision framework this data feeds: Should You Own or Rent Your AI Model? The wider 2026 evidence on control and returns: The State of Enterprise AI in 2026.

Price your heavy work both ways

A short assessment sorts your AI workload by task, the way this data does, and prices the bounded work on an owned model running inside your network.

Book a Free AI Strategy Assessment
Keith Kennedy

Keith Kennedy, CISSP

Founder & CEO, Cognetryx

Keith is an IT thought leader with nearly 20 years of experience architecting secure technology solutions for regulated industries. He holds a CISSP certification and advises institutions on secure AI architecture, access control, and keeping sensitive data inside the network. About Keith

The AI Spend and Token Split, Answered

On OpenRouter they do. As of August 2026, labs other than Anthropic, Google, and OpenAI handled 65% of the tokens moving through the platform, up from 39% in January 2024. OpenRouter is one platform, and it serves developers who pick their own models, so the wider market may split differently. The trend line has pointed the same way since January 2024.

OpenRouter's August 2026 numbers sort spending by task. Code work took 31.6% of the money and agent work took 28.7%, so those two make up about 60% of all spend. The single biggest task was workflow execution, agents carrying out multi-step jobs, at 19%. The rest went to everyday work such as sorting content, writing, answering questions, and pulling data out of documents.

Price. Frontier models charge far more per token than smaller ones, so the three biggest labs kept 69% of the spending on OpenRouter while handling 35% of the tokens. Some work does need a frontier model. Much of the rest gets routed to smaller models that do the job at a fraction of the price.

Sort last month's AI spend by task, the way OpenRouter does. Mark which tasks are bounded jobs with checkable answers, such as sorting, extraction, and workflow steps. Those fit smaller open-weight models, and an open-weight model can run on your own hardware at a fixed cost, with your data staying inside your network. Price that list both ways and compare.

Sources: OpenRouter model and usage rankings (openrouter.ai/rankings), and figures published by Peter Walker, Head of Insights at OpenRouter, on LinkedIn in August 2026: share of tokens and share of spending across OpenRouter from January 1, 2024 to August 20, 2026; spend by task type and total token volume for August 1 to 19, 2026; weekly token volume from August 2025 to August 2026; and growth in total token volume over the 24 months to August 2026. OpenRouter figures describe traffic routed through that platform. Companies that buy directly from a single AI lab do not appear in them, so shares across the whole market may differ.