In January 2024, Anthropic, Google, and OpenAI processed 61% of the tokens moving through OpenRouter. This month their share is 35%. The money moved less. The same three labs still collect 69% of what gets spent there.
Those two lines come from charts published in August 2026 by Peter Walker, who runs insights at OpenRouter. Put together, they say the work is moving to smaller models, and it is moving much faster than the money.
AI usage through OpenRouter grew 600x in two years. Most of the money now goes to agents and code, work that runs all day on a meter. The three biggest labs still collect 69% of the spend, but smaller labs handle 65% of the tokens. The heavy work fits smaller models, and the best smaller models are open weight, which means you can own them. We made the case for owning in July. These numbers show the market working through the same choice.
What does the OpenRouter data show?
OpenRouter is a switchboard for AI. A developer sends a request, picks a model, and the platform routes the call. Hundreds of models compete there on the same terms, so its records show which ones get chosen for real work.
The scale has changed fast. Token volume through the platform grew 600x in 24 months. In the first 19 days of August alone, users pushed 246 trillion tokens through it, and weekly volume is closing in on 100 trillion.
It is one platform, and a particular slice of the market. OpenRouter serves developers who pick their own models, so a company buying straight from one lab under contract is invisible here. Keep that in mind with every number below. A quarter of a quadrillion tokens in 19 days is still a decent sample.
What are companies paying AI to do?
Walker's August breakdown sorts the spending by task. Code work took 31.6% of the money: writing new code at 10.9%, debugging at 5.9%, and a tail of smaller jobs. Agent work took 28.7%, and one line inside it towers over the rest. Workflow execution, agents carrying out multi-step jobs from start to finish, was 19% of all spending on its own. Add code and agents together and they take 60 cents of every dollar.
The rest is everyday office work. Sorting content into buckets took 9.8%. Pulling data out of documents took 5.7%. Writing took 4.2%, and answering questions took 4%.
Read that list as a budget owner. This is production work. It runs on schedules and queues, all day, with no one waiting for a person to type a prompt. A bill that meters by the token grows with every job the system picks up, which is a big part of why AI costs are so hard to forecast.
Who is doing the work?
Now set the spending chart against the token chart. On spending, the big three went from 81% in January 2024 to 69% today. On tokens, they went from 61% to 35%. Both lines fall, but the token line falls off a cliff. Every other lab combined now handles 65% of the work on 31% of the money.
The gap has a plain cause. Frontier models charge far more per token, so the big labs can lose most of the volume and keep most of the revenue. Each token they serve earns a premium price. Each token that moves to a smaller lab costs the customer a fraction of that.
Why did the work move to smaller models?
Look at what the work is. Sorting records. Pulling fields out of documents. Running one defined step in a workflow. Reviewing code against a rule set. These are bounded tasks with checkable answers, and a smaller model tuned to a bounded task handles it well, at a fraction of the frontier price. Open-weight models closed most of the quality gap for this kind of work, and the cost gap kept widening. We covered that shift in our State of Enterprise AI review, where open-weight serving had reached up to 60x cheaper inference.
OpenRouter makes the choice visible. A developer there can put any model on any task and watch the results. The token share shows what they keep choosing.
What does the split mean for a regulated business?
The tasks eating these budgets are the same ones eating yours. Document work. Sorting. Agents running steps against internal systems. The OpenRouter charts say that work fits smaller models, and that matters in a bank or a hospital for a second reason. Many of the best smaller models are open weight. You can download the weights, run them on hardware inside your own network, and the work never leaves.
Those 65% of tokens still flow through the cloud, though. Renting a smaller model through an API is cheaper renting, and the data still travels. The same open-weight models can run inside your own walls, where the meter goes away and the files stay put. That is the difference between a cheaper meter and no meter at all.
We laid out the own-versus-rent decision in July, when Palantir's CEO was calling token pricing a wealth tax on live TV. That piece gives you the framework: who holds the weights, whether you can trace an answer, and what happens to the bill as use grows. Walker's charts read like the market answering the first of those questions, one task at a time.
What should you do with the numbers?
Build your own version of Walker's chart. Your AI tools keep usage logs, so sort last month's spend by task the way OpenRouter does. Then draw one line through the list. Which tasks need a frontier model? Which are bounded jobs with checkable answers?
For the bounded list, price an owned, open-weight model doing the work inside your network at a fixed platform cost. Our cost savings breakdown shows the math. Then check the agent line in your bill, because that is the line that grows on its own.
On OpenRouter, smaller models already handle 65% of the tokens. Find out how much of your work is bounded.
The decision framework this data feeds: Should You Own or Rent Your AI Model? The wider 2026 evidence on control and returns: The State of Enterprise AI in 2026.
Price your heavy work both ways
A short assessment sorts your AI workload by task, the way this data does, and prices the bounded work on an owned model running inside your network.
Book a Free AI Strategy Assessment