Skip to content
TokenShunt

Token routing for AI coding tools

Your engineers are paying frontier prices to read files.

Coding agents send file reads, repo scans, and boilerplate to the same frontier model that makes your hard decisions. We route that work to cheap models and keep Claude, Cursor, and Copilot for the decisions that need them.

Where the bill comes from

Most of the token bill isn't reasoning.

AI coding tools at scale blow budgets because agents use frontier models for work that needs none of their intelligence. Three of these four jobs don't need a frontier model.

01

Reading large files

Agents pull whole files into a frontier model's context just to find the ten lines that matter.

Cheap worker model

02

Scanning the repo

Search, grep, and “where is this used?” loops re-read the same code again and again, turn after turn.

Cheap worker model

03

Writing boilerplate

Tests, types, config, and predictable code carry no judgment — only volume.

Cheap worker model

04

Making decisions

Architecture, debugging, and novel problems are what you pay a frontier model for. That stays put.

Claude / Cursor / Copilot

The pattern works

Same output. Far fewer frontier tokens.

Spotify Engineering routed bulk file reading and predictable code generation to cheaper worker models and let Claude handle only the novel reasoning. They reported Claude Code token usage falling by about 90% in their tests.

Source: Spotify Engineering, September 2026. Tested on a Java monorepo. Spotify is not a TokenShunt customer; your savings depend on your task mix, which is why we measure first.

Reported by Spotify

~90%fewer Claude Code tokens with two-model routing

  • Hard blocks at the routing layer, not prompt guidance
  • Large files never reach the expensive model
  • Frontier model reserved for novel reasoning

How it works

Measure, route, prove — on your own code.

No rip-and-replace. Your engineers keep the tools they already use; the routine work just stops being billed at frontier rates.

  1. Step 01

    Measure

    We baseline your agents’ real sessions: which tools you pay for, where the tokens go, and how much of the bill is reading and boilerplate rather than reasoning.

    Output

    Token map of your current spend

  2. Step 02

    Route

    Hard rules at the routing layer — not prompt suggestions — send large reads, repo scans, and predictable code to cheap models. Decisions stay on the frontier model your engineers already use.

    Output

    Routing policy for your stack

  3. Step 03

    Prove it

    Routed and unrouted runs on your own tasks, side by side. Anything that costs quality goes back to the frontier model before it ever reaches your engineers.

    Output

    Before/after cost and quality report

Who it's for

If you're spending real money on AI coding tools, this applies.

Claude Code, Cursor, or usage-based Copilot across an engineering org. If your AI coding bill is a rounding error, you don't need us — and we'll tell you so in the first call.

Claude Code

Cursor

GitHub Copilot (usage-based)

Founding partner program

Shape the practice with us.

We are partnering with a select group of enterprises for our first engagements. Founding partners work directly with our principals, influence what we build next, and receive founding-partner terms.

  • Principal-led engagement from day one
  • Early access to our evaluation and routing tooling
  • A published case study only if and when you choose
Apply to the program

20 minutes

See whether it fits your stack.

Tell us which coding tools your engineers use and roughly what you spend. We'll show you where the tokens go and what routing would change — or tell you it won't.

We respond within 1 business day. Mutual NDA available before any data discussion.