01
Reading large files
Agents pull whole files into a frontier model's context just to find the ten lines that matter.
Cheap worker model
Token routing for AI coding tools
Coding agents send file reads, repo scans, and boilerplate to the same frontier model that makes your hard decisions. We route that work to cheap models and keep Claude, Cursor, and Copilot for the decisions that need them.
Where the bill comes from
AI coding tools at scale blow budgets because agents use frontier models for work that needs none of their intelligence. Three of these four jobs don't need a frontier model.
01
Agents pull whole files into a frontier model's context just to find the ten lines that matter.
Cheap worker model
02
Search, grep, and “where is this used?” loops re-read the same code again and again, turn after turn.
Cheap worker model
03
Tests, types, config, and predictable code carry no judgment — only volume.
Cheap worker model
04
Architecture, debugging, and novel problems are what you pay a frontier model for. That stays put.
Claude / Cursor / Copilot
The pattern works
Spotify Engineering routed bulk file reading and predictable code generation to cheaper worker models and let Claude handle only the novel reasoning. They reported Claude Code token usage falling by about 90% in their tests.
Source: Spotify Engineering, September 2026. Tested on a Java monorepo. Spotify is not a TokenShunt customer; your savings depend on your task mix, which is why we measure first.
Reported by Spotify
~90%fewer Claude Code tokens with two-model routing
How it works
No rip-and-replace. Your engineers keep the tools they already use; the routine work just stops being billed at frontier rates.
Step 01
We baseline your agents’ real sessions: which tools you pay for, where the tokens go, and how much of the bill is reading and boilerplate rather than reasoning.
Output
Token map of your current spend
Step 02
Hard rules at the routing layer — not prompt suggestions — send large reads, repo scans, and predictable code to cheap models. Decisions stay on the frontier model your engineers already use.
Output
Routing policy for your stack
Step 03
Routed and unrouted runs on your own tasks, side by side. Anything that costs quality goes back to the frontier model before it ever reaches your engineers.
Output
Before/after cost and quality report
Who it's for
Claude Code, Cursor, or usage-based Copilot across an engineering org. If your AI coding bill is a rounding error, you don't need us — and we'll tell you so in the first call.
Claude Code
Cursor
GitHub Copilot (usage-based)
Founding partner program
We are partnering with a select group of enterprises for our first engagements. Founding partners work directly with our principals, influence what we build next, and receive founding-partner terms.
20 minutes
Tell us which coding tools your engineers use and roughly what you spend. We'll show you where the tokens go and what routing would change — or tell you it won't.
We respond within 1 business day. Mutual NDA available before any data discussion.