All posts
2 min readby Romiel Inolino

LangChain cut its coding agent's median cost per task by 64% with a model router. Most requests never needed the top model

AI agentsLLM costsmodel routingAI engineering

Most AI agent bills are high for one reason: every request goes to the most expensive model.

What happened

LangChain built a model router for Open SWE, its internal coding agent. It labeled a week of real requests by task type, picked three models along the cost and intelligence curve (fast, balanced and performance tiers), and wrote plain language criteria for which work each tier should take. A classifier reads the first message in a thread and picks the cheapest tier likely to succeed.

In an A/B test across 973 threads against always using GPT-6 Astra, merged PR rates were statistically flat (29.2% routed vs 27.3% control). Median cost per thread fell from $2.61 to $0.94, a 64% drop, with the mean down 42% and the 90th percentile down 37%. Only 10% of routed threads needed the performance tier. A second test that sent everything to the fast tier was stopped within a day because engineers complained about quality.

The classifier now runs on Jev, a decision model, which LangChain says made classification almost 50 times faster. Separately, TechCrunch reports AWS released Strands Decider 2B, an open source decision model small enough to run locally that picks among predefined options and returns a confidence score. Its creator, Marc Brooker, described it as a natural decider for a workflow step.

My take

This maps directly onto n8n and Make builds. A router is just a classification step at the top of the workflow, followed by a switch.

How I would apply it:

  1. Pull a week of real inputs first. You cannot design tiers without knowing your task mix.
  2. Define success before routing. LangChain used merged PRs. For a support flow, it might be tickets resolved without escalation.
  3. Start with two tiers, not five.
  4. Do not route everything cheap. Their fast only test failed in a day.
  5. Log the tier on every run, so you can see when the router is wrong.

The savings come from measurement, not from the model.

More posts