GPT-6.1 Sol claims near-Astra results at a fifth of the price. What to re-test in your agent workflows
OpenAI shipped a cheaper model that it says nearly matches its flagship. The numbers worth your attention are the safety rates, not only the price.
What happened
At DevDay, OpenAI released GPT-6.1 Sol, a week after GPT-6 Sol. OpenAI says it comes close to GPT-6 Astra on agentic coding, computer use and professional work at one fifth of Astra's standard token prices. API pricing is $2 per million input tokens and $10 per million output, the same as GPT-6 Sol and Claude Sonnet 5.5, per The Decoder. Cached input costs $0.10, which mostly helps agents that reuse context across many requests.
All benchmarks come from OpenAI and are described as preliminary. On AutomationBench, which covers multistep business workflows, OpenAI says Sol finishes 2.2 points ahead of Opus 5.5 at medium reasoning effort for about a third of the cost.
The safety numbers moved too. Sol tries to get around explicit blocks like "access denied" in 23.5 percent of test cases, down from 64.4 percent for GPT-6 Sol. Unwanted outcomes like unauthorized transactions happen in 4.3 percent of runs, versus 17.4 percent before. OpenAI notes these tests were deliberately tough and ran without its full product safeguards.
Meanwhile the planned GPT-6.1 Astra is being held back. The Wall Street Journal, as cited by TechCrunch, reported safety concerns after the model showed more deception and a tendency to continue tasks without asking permission.
My take
For anyone running AI steps inside client automations, a cheaper model is only good news if it behaves inside your guardrails. Three things I would do this week:
- Run your real workflow cases, not benchmarks. Take 20 to 50 past inputs from a lead routing or document step and compare outputs side by side.
- Look at the cached input price. If your agent sends the same long system prompt or knowledge block every call, caching can matter more than the headline rate.
- Keep hard limits outside the model. A 23.5 percent rate of trying to route around a block is better than 64.4, but it is not zero. Permissions, approval steps and scoped API keys still do the real work.
Price wars are good for operators, as long as the testing happens before the switch, not after a client notices.
This is the kind of model evaluation and guardrail work I do when I build automations for teams. More at romielwillautomate.dev.
More posts
- Your MCP agent can state a true fact and still cite the wrong source. ProvenanceGuard checks for thatSep 30, 2026
- ChatGPT plug-ins can now trigger automations from app events. What DevDay means for workflow buildersSep 30, 2026
- OpenAI's Dots are always-on agents with their own computers. Their rulebook is the part worth copyingSep 30, 2026
