Skip to content

Microsoft launches Decision-1, a model that returns only yes/no and multiple-choice answers

· by Pondero Newsdesk

The short version

Microsoft's Decision-1 skips text generation for calibrated yes/no and multiple-choice scoring, arriving weeks after TypeSafe's similar Jev model drew a $7.5 billion valuation.

Microsoft launches Decision-1, a model that returns only yes/no and multiple-choice answers

Microsoft released a model on October 9 that refuses to write a single sentence of prose. Decision-1 answers only yes/no questions, picks from a multiple-choice list, assigns a rating, or grades another system's output against a rubric, and it does so weeks after a similarly narrow model from startup TypeSafe AI, Jev, drew a $7.5 billion valuation in an $870 million Series A, per TechCrunch.

What Microsoft built

Per Microsoft's announcement, Decision-1 was made by post-training the open-weight Qwen3.5-9B model to output a calibrated probability score for a fixed set of answer options instead of generating free text, aimed at routing, classification, prioritization, verification, and grading AI agent actions. Microsoft said the model achieved the highest accuracy across a 36-benchmark comparison spanning nearly 150,000 questions held out from training, and that it ran 2.5 times faster than the runner-up, H2O-Lightning-4B v1.1, and 35 times faster than GPT-6 Sol. In a robustness test that paraphrased, reordered, or reformatted the same question eight different ways, Microsoft said Decision-1 changed its answer on 1.3% of those variations on average, with zero flips when option descriptions were paraphrased or shuffled. The model is live now in Microsoft Foundry and on OpenRouter, priced at $0.042 per million input tokens with output tokens free, and Microsoft said it plans further versions rebased on its own MAI models and on OpenAI technology.

A contested category, not just a product launch

Decision-1 is Microsoft's entry into a model class TypeSafe effectively named: its Jev model spawned a public leaderboard, JevBench, that Microsoft says it used to pull in rival models for its own benchmark run. That detail matters more than the product itself. A model category barely a month old already has a Big Tech incumbent and a venture-backed startup racing on the same axis (speed and cost on structured decisions, not open-ended generation), which raises the stakes for any team currently routing classification or moderation tasks through a general-purpose LLM. Operators building agent pipelines now have a second, larger-company option alongside TypeSafe's, and Microsoft's Foundry and OpenRouter distribution means the model reaches existing Azure customers without a new vendor relationship.

Numbers to treat as Microsoft's own

Every benchmark figure above, accuracy, the 2.5x and 35x speed multiples, and the 1.3% perturbation rate, comes from Microsoft's own blog post and benchmark run, not an independent test. Microsoft's post itself notes it was updated after publication to add benchmarks against Jev on accuracy and calibration, which suggests the comparison to TypeSafe's model was a late addition rather than part of the original testing plan.

What to watch next

Whether any third party runs Decision-1 against Jev or other decision models on a neutral benchmark, and whether Microsoft's promised MAI- and OpenAI-rebased versions ship with different performance tradeoffs than the Qwen3.5-9B base.

Sources