Microsoft has introduced a new AI model called Microsoft-Decision-1, designed for routing, classification, prioritization, data validation, and workflow management. This model is built on Alibaba's open-source Qwen3.5-9B.
Unlike traditional large language models that generate text and reason, decision-making models provide structured responses that can be acted upon immediately.
Microsoft-Decision-1 evaluates a fixed set of options and assigns calibrated probabilities to each.
The neural network supports binary answers, multiple-choice selections, and ratings. It also assesses AI responses and agent actions based on specified criteria.
Developers have fine-tuned Qwen3.5-9B to enable rapid single-pass evaluations.
In the near future, Microsoft plans to adapt this solution to other foundational models, including its own MAI and those developed by OpenAI.
Testing Results
According to Microsoft, the new model achieved the highest accuracy across 36 benchmarks, featuring nearly 150,000 questions that were not included in its training.
Average accuracy of decision-making models across 36 benchmarks. Source: Microsoft.Additionally, this model was the fastest among those tested, outperforming its closest competitor, H2O-Lightning-4B v1.1, by 2.5 times, and GPT-6 Sol by 35 times.
Median latency per request in the JevBench test set. Source: Microsoft.Microsoft also assessed the model's robustness by altering the same query in eight different ways, including rephrasing instructions, rearranging answer options, and introducing formatting noise. On average, the solution varied in only 1.3% of cases, with rephrasing and rearranging options having no impact on the outcome.
In safety tests, developers used 5,250 queries from 11 benchmarks that included malicious attempts, jailbreak efforts, and prompt injections. Microsoft claims the model effectively rejected harmful actions while maintaining utility.
Internal Applications
The Xbox Research division utilized Microsoft-Decision-1 to analyze over 10,000 reviews from surveys, Steam, and X. In terms of quality, the model matched GPT-6 Sol but operated over 14 times faster and was 200 times cheaper.
The Copilot team employed the neural network to evaluate the quality of chatbot and agent responses, achieving results comparable to GPT-5.6 Luna with a hundredfold increase in speed.
In Microsoft Discovery, the model assigned ratings to experiments for adaptive rescheduling, resulting in 46 times more stability than LLM while being three times faster. The rescheduling process itself was nearly four times quicker.
Other applications mentioned by the company include selecting suitable models for requests, content moderation, code verification, interface management, and robotics.
The cost of usage is $0.042 per million input tokens, with output tokens being free of charge.
Notably, at the end of September, Microsoft revealed a new version of Copilot that will integrate chat features, application creation tools, automation, and autonomous AI agents.
Stay updated by following ForkLog on social media
Telegram (main channel) Facebook X