Microsoft's Decision-1 returns a confidence score with every label. Use it to decide when a human should check
Another decision model launched. The useful part is not the speed, it is the probability that comes back with each answer.
What happened
Microsoft released Microsoft-Decision-1 on October 9, a model built for routing, classification, prioritization, verification and workflow control. It is a post-trained version of Qwen3.5-9B, available in Microsoft Foundry and through OpenRouter. Input costs $0.042 per million tokens and output is free.
Given a fixed set of options, it returns a calibrated probability for each one. It handles yes or no, multiple choice, ratings and rubric-based grading of AI responses and agent actions.
Microsoft's own numbers:
- Highest accuracy in its 36-benchmark comparison, nearly 150,000 questions.
- 2.5 times faster than the runner-up, H2O-Lightning-4B.
- When the same request was reworded in eight ways, the decision changed 1.3% of the time on average, with zero flips when options were paraphrased, reversed or shuffled.
- Xbox Research sorted more than 10,000 pieces of feedback into fixed themes, with quality competitive with GPT-6 Sol at over 14 times the speed and 200 times less cost.
The Decoder notes these are Microsoft's own tests, and that Cloudflare's open Clef models were not in the comparison.
My take
Microsoft states the key idea plainly: applications use confidence to decide when to act, defer, or ask for review, and a 90% prediction should be right about nine times out of ten.
That gives you a clean design for lead and ticket routing:
- Above your threshold: route automatically and log it.
- In the middle: send it to a review queue with the top two options shown.
- Below: fall back to a default owner.
Do not guess the thresholds. Pull 100 to 200 past leads or tickets with known correct answers, run them through, and see where accuracy drops. That is an afternoon of work, and it tells you how much you can safely automate.
The reorder result matters too. Pipeline stages and team names change often in a CRM, and a router that changes its answer when you reorder a dropdown is a support ticket waiting to happen.
Treat vendor benchmarks as a shortlist, not a verdict. Your own labeled records decide.
More posts
- A signed BAA does not make your clinic client's CRM automations HIPAA safe. Map where the data goes nextOct 11, 2026
- Claude can now split one job across up to 1,000 agents. Set a budget before you try itOct 11, 2026
- Satya Nadella says to assume every AI model is compromised and give it an emergency brake. Your automations need one tooOct 11, 2026
