Microsoft released Microsoft-Decision-1, a decision-scoring model post-trained from Alibaba's Qwen3.5-9B that returns calibrated probabilities for fixed answer options instead of text. Available via Microsoft Foundry and OpenRouter, it led Microsoft's 36-benchmark comparison at 83.5% average accuracy with 85 ms p50 latency, ahead of Quyet-1.0-Large at 81.9%. Pricing is $0.042 per million input tokens with free output. Latency comparisons are disputed, and all benchmarks are vendor-run.
No score is assigned. Sources and their independence are shown in the citation chain below.