Published
Updated
Microsoft-Decision-1 lists $0.042 per million input tokens and calls itself 4.5 times quicker than Quyet-1.0-Large
On October 9, 2026 Microsoft posted Microsoft-Decision-1 at $0.042 per million input tokens, with output free. The post says the model was 4.5 times quicker than Quyet-1.0-Large and 35 times quicker than GPT-6 Sol, and it does not print an accuracy. The Foundry catalog, opened October 10, lists a 32,768-token window and says the weights are not distributed. A search of that day's JevBench API page found no Microsoft-Decision-1. We did not call the API.
Achint Srivastava posted Microsoft-Decision-1 on Command Line on October 9, 2026. The byline lists him as VP of Software Engineering in Microsoft’s Office of the CTO. Input tokens are $0.042 per million. Output tokens are free. The post says the model is available in Microsoft Foundry and on OpenRouter.
Satya Nadella posted the same announcement at 18:37 UTC. The attachment is a 10.2 second video. We did not transcribe a score from it. OpenRouter posted at 22:43 UTC that the model was live, at $0.042 per million input tokens, with free output and a 32K context.
The Foundry catalog, opened October 10, 2026, labels the lifecycle Generally available and the version 1. The context window is 32,768 tokens. OpenRouter’s header the same day prints 33K. The model id there is microsoft/microsoft-decision-1. A caller sends text and a closed list of options. The model returns a probability for each option. The catalog says the call is text only, and that the model does not write an explanation.
The Learn page, updated October 9, names three question types. Noul is a probability from 0 through 1. Choice returns the selected option and a probability for every option. Score is the probability-weighted average of the level indexes, so a four-level scale runs from 0 through 3 and the result can fall between levels. That page says to prefer scores for relative ordering and thresholds rather than as absolute, calibrated ratings. The Command Line post says a 90% prediction should be right about nine times out of ten on representative cases. The catalog also calls the scores calibrated.
The Foundry path is {resource}/providers/microsoft/v1/systemone. The request’s model field is the deployment name. The response’s model field can read microsoft-decision-1. Learn names DataZoneStandard in selected regions and GlobalStandard. The section we read does not list the regions. We did not deploy the model or send a request.
Microsoft post-trained Qwen3.5-9B. The technical specs call the result a dense decoder-only Transformer in the 5B to 15B range, and they link Qwen/Qwen3.5-9B. We did not open that card. The catalog says the weights are not distributed. Training time is September 2026 to October 2026. The cutoff date is listed as not supplied. Post-training text is between 1 billion and 1 trillion tokens. The disclosure says no Microsoft customer data was used, and that collection is ongoing. The post says a later version will rebase onto Microsoft AI and OpenAI models. OpenRouter’s description says the weights are updated continually and the API shape stays the same.
The post says the model had the highest accuracy in a 36-benchmark comparison of nearly 150,000 questions, on benchmarks kept blind from training. It does not print that accuracy. It says this was the fastest model measured: 4.5 times quicker than Quyet-1.0-Large, and 35 times quicker than GPT-6 Sol. A later line says P50 latency is about 35 times faster than GPT-6 Sol. The appendix names Quyet-1.0-Large, Surogate Rune 26B-A4B, GPT-6 Luna Decisions, deck-31B, H2O-Lightning-4B, and Strands-Decider 2B. It does not name Jev. The post says the comparison also took several of the top public models on JevBench and tested them on 36 further public and private benchmarks.
The catalog’s benchmarks tab, opened October 10, prints no accuracy percentage and no JevBench string. It says the model performs on par with leading decision models and ahead of other open decision models under that methodology. It says the model is strongest on reasoning, rule application, and robustness to prompt formatting, competitive on classification, retrieval, fairness, tool use, and most multilingual tasks, and weaker on specialized domain knowledge.
Opened the same day, the JevBench API board still said release v1.6.1, and the header said board v1.7.29. A search of that page’s text for Microsoft-Decision-1 returned no match. We did not rerun the suite.
The robustness test perturbs one request in eight ways. The decision changed on 1.3% of perturbations on average. The post says there were zero flips when option descriptions were paraphrased, or when options were reversed or shuffled. The safety test is 5,250 requests across 11 benchmarks, covering harmful content, jailbreaks, and prompt injection. The post says the model refused harmful behavior and kept its utility. It does not print a refusal rate.
XBOX Research, in the post, sorted more than 10,000 pieces of open-ended feedback from surveys, STEAM, and Twitter/X into themes the researchers had already named. The post says quality was competitive with GPT-6 Sol, at over 14 times the speed and 200 times lower cost. The Copilot team, it says, found the model competitive with GPT5.6 Luna and 100 times faster. On-call engineers used it to retrieve knowledge during live incidents, and the post says it performed better and faster than an LLM. That line has no number. Microsoft Discovery, it says, found the model 46 times more consistent than the LLM-based score, at three times the speed, and nearly four times the speed on adaptive replanning.
OpenRouter, opened October 10, lists one provider, Azure, with the privacy label Private. The listed price is $0.042 per million input tokens and free output. That page’s P50 latency was 0.18 seconds, and uptime was 98.73%. The weighted average input price shown there was $0.0415 per million tokens. The Azure Foundry pricing page we opened the same day did not include the name Decision-1 in its visible text. The playground asked for a sign-in. We did not call the API.
The catalog says the model is optimized for English and was evaluated in Chinese, Japanese, Korean, Spanish, French, Portuguese including Brazilian Portuguese, German, Italian, Russian, Arabic, Indonesian, Vietnamese, Thai, Turkish, Hindi, Bengali, Swahili, Dutch, Polish, Hebrew, Persian, and Ukrainian. It says the Qwen3.5-9B base supports more than 200 languages, and that quality and calibration can vary by language. It also says the model is not for open-ended generation, conversation, translation, or summarization, and that it should not be the only basis for a decision about credit, employment, housing, insurance, education, healthcare, or legal rights.
Other hosted calls of this shape are on decision models like Jev.
This site's reading
Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.
Verify
Primary post is Achint Srivastava on Command Line, dated 2026.10.09, https://commandline.microsoft.com/microsoft-decision-1-model-foundry/. The byline is VP of Software Engineering, Office of the CTO, Microsoft. Input tokens are $0.042 per million. Output tokens are free. The post says the model is in Microsoft Foundry and on OpenRouter. It says Microsoft post-trained Qwen3.5-9B and will rebase onto Microsoft AI and OpenAI models. A fixed option list gets a calibrated probability per option. The formats named are yes/no, multiple-choice, rating, and rubric grading.
The post says Microsoft-Decision-1 had the highest accuracy in a 36-benchmark comparison of nearly 150,000 questions, kept blind from training, and does not print that accuracy. It says the model was the fastest measured: 4.5 times quicker than Quyet-1.0-Large, and 35 times quicker than GPT-6 Sol. The speed section says P50 latency is about 35 times faster than GPT-6 Sol. Eight perturbations changed the decision on 1.3% of cases on average, with zero flips when option descriptions were paraphrased or when options were reversed or shuffled. Safety is 5,250 requests across 11 benchmarks. The post does not print a refusal rate.
Internal lines, as printed. XBOX Research: more than 10,000 pieces from surveys, STEAM, and Twitter/X, competitive with GPT-6 Sol, over 14 times faster and 200 times less expensive. Copilot: competitive with GPT5.6 Luna and 100 times faster. Incident knowledge retrieval: better and faster than an LLM, no number. Microsoft Discovery: 46 times more consistent than the LLM-based score, at three times the speed, and nearly four times the speed on adaptive replanning. The post embeds a classification demo and a backpack-buying demo. We did not transcribe a time or a score from either. The appendix names Quyet-1.0-Large, Surogate Rune 26B-A4B, GPT-6 Luna Decisions, deck-31B, H2O-Lightning-4B, and Strands-Decider 2B. It does not name Jev.
Satya Nadella, October 9, 2026, 18:37 UTC, status 2108627923888754862. The attachment is a 10.2 second video. We did not transcribe a score from it. OpenRouter, the same day, 22:43 UTC, status 2108689756372775370, says the model is live, quotes the 36 benchmarks, the 4.5 times and 35 times lines, the 1.3% flips, the Qwen3.5-9B base, $0.042 per million input tokens, free output, and a 32K context.
https://ai.azure.com/catalog/models/Microsoft-Decision-1, opened October 10, 2026. Version 1. Lifecycle label Generally available. Context window 32768. Direct from Azure. Built on Qwen3.5-9B, post-trained by Microsoft. Training data is public datasets under Microsoft's Open Data process plus synthetic data. Weights are not distributed. No Microsoft customer data. Post-training text is between 1 billion and 1 trillion tokens. The training dataset was first used in September 2026, and collection is ongoing. Technical specs: training time September 2026 to October 2026, cutoff date not supplied, dense decoder-only, 5B to 15B parameters, up to 32,768 input tokens, text in, numerical probabilities out, no explanatory text. Optimized for English and evaluated in the languages named in the body. The benchmarks tab prints no accuracy percentage and no JevBench string. It says the model is on par with leading decision models and ahead of other open decision models on that methodology.
Learn, https://learn.microsoft.com/en-us/azure/foundry/foundry-models/how-to/use-foundry-models-microsoft-decision, updated_at 2026-10-09. The path is {resource}/providers/microsoft/v1/systemone. The request model is the deployment name. The response model can read microsoft-decision-1. Types are noul, choice, and score. A score is the probability-weighted average of the level indexes. The page says to prefer scores for relative ordering and thresholds rather than as absolute, calibrated ratings. Deployment types named there are DataZoneStandard in selected regions and GlobalStandard. That section does not list the regions. We did not deploy the model or send a request.
OpenRouter's model page, opened October 10, 2026. Model id microsoft/microsoft-decision-1. Header price $0.042 / $0 per 1M. Header context 33K. Released Oct 9, 2026. One provider, Azure, privacy label Private. P50 latency 0.18 s. Uptime 98.73%. Weighted average input price $0.0415 per million tokens, output $0. The description says weights are updated continually and the API shape stays the same. The playground asked for a sign-in. We did not run it.
https://azure.microsoft.com/en-us/pricing/details/microsoft-foundry/, opened the same day. A search of the visible text for Decision-1 returned no match. https://benchmarkheaven.com/jev-models/api, opened the same day, still said release v1.6.1, and the header said board v1.7.29. A search of that page's text for Microsoft, Decision-1, and Quyet returned no match. We did not rerun the board.
Compare
The 4.5 times and 35 times lines are Microsoft's own speed comparison. The post does not print milliseconds for those rows, and it does not print the accuracy it calls the highest. The appendix does not name Jev. JevBench's API page, opened October 10, 2026, had no Microsoft-Decision-1 string. That search is not a score.
OpenRouter's 0.18 second P50 is the Azure provider row on October 10. It is not the bake-off against Quyet-1.0-Large or GPT-6 Sol. The $0.042 price matches the Command Line post. The $0.0415 figure is that page's weighted average input price the same day. TypeSafe's models page, in the October 7 reading on this desk, also lists $0.042 per million input tokens. The two prices are separate rate cards.
The Learn page says score results are for ordering and thresholds, and that they are not absolute calibrated ratings. The Command Line post says a 90% prediction should be right about nine times out of ten on representative cases. Those two sentences are both Microsoft's. This desk did not measure either one.
Terms
- $0.042 per million input tokens
- The input price on the October 9 Command Line post and on the OpenRouter page opened October 10. Output tokens are free on both. The Azure Foundry pricing page opened that day did not show the name Decision-1.
- 4.5 times quicker
- Microsoft's October 9 speed line against Quyet-1.0-Large. The same post says 35 times quicker than GPT-6 Sol, and that P50 latency is about 35 times faster than GPT-6 Sol. It does not print the milliseconds.
- 1.3% of perturbations
- The average rate at which eight perturbations of one request changed the decision, in the October 9 post. Paraphrased option descriptions, and reversed or shuffled options, had zero flips.
- 32,768 tokens
- The context window on the Foundry catalog opened October 10, 2026. OpenRouter's header that day printed 33K. OpenRouter's October 9 post said 32K.