All posts
2 min readby Romiel Inolino

Nvidia cut coding agent token use by up to 49% without changing the model. The harness is where your costs hide

AI agentsLLM costsNvidiaagent engineeringautomation

Before you switch to a cheaper model, look at what your agent sends it on every loop.

What happened

Nvidia researchers published SoL-Pi, a system that optimizes the harness of a coding agent: the control layer that decides what the model sees, how actions run and how feedback comes back. A research agent reads another agent's traces, proposes harness changes and keeps only the ones that hold performance while cutting cost.

The Decoder reports the search produced four mechanisms. Action Fusion merges two consecutive steps, such as an edit followed by a test run, removing a model call. Online Context Compact trims accumulated context after planning steps. ObservationPack archives long tool outputs and sends a short summary on later steps. An Evidence Preserving Reducer routes large logs to a cheaper model for summarizing, with a verification step for key clues.

On EdgeBench, token usage dropped 44.7 to 49 percent. The leanest variant reached 93.7 percent of the original Pi harness score; one test run fell from $1,339 to $894. Results were mixed elsewhere: on 63 Terminal-Bench tasks it solved 15 versus 18 for Codex and Pi, though at about a quarter lower cost. The article also cites an August Composio test where the same model varied nearly 3x in cost per solved task across four frameworks.

My take

This is a coding agent paper, but the lesson is general. Most of the bill in an agent workflow is not the model price, it is the loop: full tool outputs pasted back every turn, whole CRM records when three fields would do, a separate model call for every tiny step.

Things I apply in n8n and custom agent builds:

  • Send summaries of big API responses, not the raw JSON, on later steps.
  • Combine steps that always run together.
  • Push log reading and cleanup to a small, cheap model.
  • Measure cost per completed task, not cost per token.

Note the tradeoffs the researchers flagged: compact context can reduce prompt cache reuse, and the gains were smaller on unfamiliar tasks. Test on your real workload before and after.

I design and build AI automations that do real jobs, from lead routing and reporting to data entry, content pipelines and agent workflows, and I tune them so the monthly bill stays sane. See the work at romielwillautomate.dev.

More posts