Look at the jobs a bank or a fund would hand to AI in a normal week. Tag each document in a loan file. Pull balances and dates off statements. Sum up the exceptions in a vendor's SOC 2 report. Answer a teller's question about the callback rule for a large wire, from the current procedures manual.
These are bounded jobs. Each one starts from a document you supply, asks for a narrow result, and ends in an answer someone can check against the page. Smaller open-weight models handle jobs like these well, and the research behind that has grown fast.
Open weights also mean your firm can hold the model. It can run on your own servers, next to the documents, at a cost that stays the same however many files it reads.
What makes an AI job bounded?
A bounded job has a clear input, a narrow task, and a result you can check. Tag this document. Pull these fields. Sum up this file. Answer this question from this policy. The model works from text you hand it, so it needs to read well and follow instructions more than it needs broad knowledge of the world.
At a bank, credit union, or fund, bounded jobs include:
- Sorting. Tagging each document in a loan file, or routing complaints by product and issue.
- Pulling fields. Account numbers, balances, dates, and signers from statements and forms.
- Summaries. A long credit memo, a vendor's SOC 2 report, or a board packet, cut down to what the reader needs.
- Policy answers. A question from a teller or a loan officer, answered from the current manual with the passage cited.
- First drafts. DDQ answers built from past responses and current policies, for a person to edit.
Some work does need a frontier model. Open-ended research across many sources, hard reasoning with no single right answer, and tricky code can earn the higher price. The test is whether you can check the result against something. If you can, give a smaller model the first trial.
What does the research say about smaller models?
It says the size needed for a given skill has shrunk fast. Stanford's AI Index found that the smallest model to score above 60% on MMLU, a broad test of knowledge, went from 540 billion parameters in 2022 to 3.8 billion in 2024.[1] That's a 142-fold drop in just over two years.
The cost of a given level of skill fell even more. The same report found that the cost to run a model at the level of GPT-3.5 dropped more than 280-fold between November 2022 and October 2024.[1]
Narrow jobs make the case stronger. In 2024, engineers at Predibase, a company that sells model tuning and hosting, trained small open models on 31 narrow tasks, from pulling names out of text to sorting content. The models had 2 to 7 billion parameters. After tuning, 224 of the 310 models beat GPT-4 on their task, and the team served 25 of the tuned models from a single GPU.[2] Tuning takes extra work, so read this as a sign of what a small model can reach on a narrow job.
The same logic runs through agents. Researchers at NVIDIA argued in 2025 that many agent systems use a model for a small number of narrow tasks, repeated with little change. Small models, they wrote, are strong enough and cheaper for many of those calls.[3] It's a position paper from a company that sells GPUs, so read it as an argument. It fits what the tuning results show.
At the top of the charts, open-weight models trail by a little. In March 2026, the best closed model led the best open-weight model by 3.3%, up from 0.5% in August 2024.[4] For a bounded job, a test on your own documents will tell you more than that gap.
What are companies paying AI to do?
OpenRouter, a marketplace developers use to call models from many labs, posts charts of its own traffic. In August 2026 its data page ranked the top three things people paid for: agents carrying out workflows, writing code, and what it called "classifying lots of things."[5] Sorting, one of the plainest bounded jobs, is a top-three expense there.
The top category has more in common with sorting than it seems. An agent that works a loan file runs a chain of bounded steps. It opens the file, tags each document, pulls the key fields, compares them with the credit policy, and writes a note for the processor. Each step can be checked, so each step is a candidate for a smaller model.
The same page shows which models carry the volume. In the 30 days to September 8, 2026, one closed model made OpenRouter's top ten by token volume.[5] The other nine were open-weight.
Open weights tell you who can hold a model, and size is a separate question. Open-weight models range from a few billion parameters to hundreds of billions, and Meta's Llama 3.1 alone came in 8, 70, and 405 billion versions.[6] So the ranking shows developers picking open weights for volume work. It says little about size.
OpenRouter is one platform, and it serves developers who pick their own models. A bank that buys from one AI vendor under contract won't show up in these numbers, so treat them as a view of one market.
Why does owning the model matter for a bank or a fund?
Open weights mean the model can run on your own servers. Your documents stay inside, the model changes only when your team installs an update, and the cost stops rising with every call. For bounded jobs that run all day, that last point decides how much you can automate.
Renting a smaller model through an API is cheaper renting. The meter still runs, and your files still travel to someone else's servers. Running the same kind of model inside your own walls ends both.
Lumen, the private AI platform from Cognetryx, runs the model your firm chooses, on your own hardware. The platform has one fixed price, so an agent that tags every loan file every night costs the same as one that runs once. The fixed-cost section of Why Cognetryx explains how the price is set.
Owning the model helps in an exam, too. Updates arrive as signed packages that your IT team installs on its own schedule, so the model behind last month's answers is the one you chose. Each answer cites its source, and a click opens the document with the passage highlighted.
How do you test a smaller model on your own work?
Start from your own list of jobs. Take last month's AI requests, or the jobs you plan to automate, and sort them by task. Mark each one bounded or open-ended. Then run a sample of the bounded jobs through a smaller open-weight model and score the results against your staff's work.
- Build a test set from real files. Use documents your team has already handled, so you know the right answer for each one. Put the hard cases in on purpose: very long files, poor scans, tables, and questions that need facts from several documents at once. They show where a smaller model's limits are for your work.
- Score what the job needs. The right tag, the right field values, the right passage cited. For summaries, have the person who would have written one grade it.
- Check every citation. A policy answer should link to the passage it came from. Our guide on checking an AI answer before it reaches an examiner covers how.
- Price it both ways. Compare the per-token cost of the bounded list with a fixed cost for owned capacity. Our guide to fixed-cost AI covers when each one wins.
- Keep the big model where it earns its price. Jobs that fail the test stay with a frontier model.
Start with one queue, such as tagging incoming loan documents. A week of results on your own files will tell you more than any leaderboard.
Watch bounded work run on real documents
See Lumen answer from real financial documents, with each answer linked to the passage it came from.
See it in actionSources
- Stanford HAI, The 2025 AI Index Report. Source of the drop from 540 billion to 3.8 billion parameters for a score above 60% on MMLU, and the more than 280-fold fall in the cost of GPT-3.5-level results. hai.stanford.edu
- Justin Zhao and others, Predibase, LoRA Land: 310 Fine-tuned LLMs that Rival GPT-4, A Technical Report, April 29, 2024. Ten base models of 2 and 7 billion parameters, tuned on 31 tasks. arxiv.org/abs/2405.00732
- Peter Belcak and others, NVIDIA, Small Language Models are the Future of Agentic AI, position paper, first posted June 2, 2025. arxiv.org/abs/2506.02153
- Stanford HAI, The 2026 AI Index Report, technical performance chapter. hai.stanford.edu
- OpenRouter data page: "Spend by task category," August 20, 2026, and "Top models by token volume," September 8, 2026, covering the prior 30 days. The figures cover OpenRouter traffic only. openrouter.ai/data
- Meta, Introducing Llama 3.1, July 23, 2024. ai.meta.com