Published
Updated
Open d1's d1-3B card prints 48.57 on a self-run of Decision Index 0.2.1
On October 7, 2026 Liquid AI posted Open d1, and the d1-3B card prints 48.57 on a self-run of Decision Index 0.2.1. Winnow-12B on that table is 50.02. d1-omni-600M is 15.95. The LICENSE file is LFM Open License v1.0, with commercial use under a $10,000,000 revenue line. We did not run the weights.
Liquid AI posted Open d1 on October 7, 2026, at 17:01 UTC. The thread names two open checkpoints, d1-3B and d1-omni-600M. d1-3B reads text and images. d1-omni-600M reads text with an image, or text with audio. A call sends a state and named questions and returns Noul, Choice, or Score from one forward pass. Both cards print output_tokens at 0. The hosted model is a separate page, Liquid d1.
The d1-3B card prints 48.57 on Decision Index 0.2.1. It says that figure used the official scorer and is not a leaderboard submission. The blog chart caption says the same 120,226 requests, from the public leaderboard of October 6, and marks the run not submitted. The area cells are Knowledge 23.8, Language 56.4, Retrieval 52.8, Tools 74.5, and Arts 36.3. The same table prints Winnow-12B at 50.02 and Decider 35B-A3B at 47.11. The card calls d1-3B the best decision model under 10B on this index. Winnow-12B is the row above it. The under-10B rows the table lists sit below 48.57. We did not open the full leaderboard. The omni card prints 15.95 on the same index, with Knowledge 8.3, Language 12.9, Retrieval 35.0, Tools 15.1, and Arts 6.8.
48.57 is this self-run. The September 29 hosted chart is 58.9 against Jev at 57.9. The JevBench API board opened October 7 ranks the hosted name d1 at composite 73.0. Decision Index 0.3 on this desk is a later edition. The blog says edition 0.3 has a private vision split, and that this release does not report it.
The blog’s seven-benchmark table leaves HelpSteer2 out. Its means are 82.9 for d1-3B, 78.4 for d1-omni-600M, 81.1 for Decider 4B, and 77.1 for Decider 2B. On that table d1-omni-600M leads Civil Comments at 95.8 and PAWS-X at 79.5. The d1-3B card adds HelpSteer2 at 36.7, against Decider 4B at 42.0 and Decider 2B at 32.0, and prints a mean of 77.1 against 76.2 and 71.5. The omni card says HelpSteer2 may overlap that model’s training data, which is why the seven-benchmark table drops it. 82.9 and 77.1 are two means. The 77.1 on the seven-benchmark row is Decider 2B.
The d1-3B card also prints 71.8 on DecisionBench, eng v1, all 23,900 rows, and 69.3 on Fast Decisions, dev split. The omni card prints 76.9 on that Fast Decisions dev split. Eleven image benchmarks, read as decisions, average 74.1 for d1-3B and 73.9 for LFM2.5-VL-3B. With the images removed, the same questions score 45.1. ImajevBench, dev and calibration, 253 rows, is 64.0 against the base at 66.8.
d1-3B is 3.12B parameters, context 32,768, vocabulary 128,000, with a SigLIP2 NaFlex 400M vision encoder. The index table rounds the size to 3B. The docs page’s specs table also prints 3.12B. The overview card on the docs index says 3.1B. d1-omni-600M is 587M in the details table, and the title says 600M. The card splits that into a 381M trunk and decision head, a 94M vision encoder, and a 112M audio encoder. Context is 16,384 tokens. With images, the state and question text is cut to 896 tokens. Both cards name license lfm1.0. The LICENSE file on d1-3B is LFM Open License v1.0. Section 5 conditions commercial use on annual revenue under $10,000,000. A legal entity above that line is not licensed for commercial use. A qualified nonprofit’s non-commercial or research use sits outside the threshold. The blog’s closing list says “Download, fine-tune, and deploy without restrictions.”
The blog says d1-3B starts from LFM2.5-VL-3B, after averaging LFM2.5-2.6B with that model’s text backbone. d1-omni-600M starts from LFM2.5-Encoder-350M, a bidirectional encoder, then adds a 17-layer FastConformer and the vision encoder from LFM2.5-VL-450M. Audio training clips are cut at 30 seconds, on English speaker-and-assistant requests. The card says a request carries images or audio, and that passing both raises ValueError. It also says float16 on GPU kept the float32 top answer on the text, image, and audio rows they checked, and that bfloat16 changed the top answer on 0.8% of text rows and 1.7% of audio rows. Text answers use per-type temperatures in config.json. Image and audio answers are the softmax as trained. The d1-3B example on its card loads bfloat16.
The d1-3B latency table is one request at a time. On an RTX 4090 the one-question cell is 8 ms, and the card says that row uses model.compile as CUDA graphs. Without it, one question takes 16 ms. The GPU rows are bf16, median of 20 runs. Jetson AGX Thor is 16 ms, Jetson AGX Orin 64 GB is 26 ms, Jetson Orin Nano is 50 ms, and an Apple M5 Pro is 30 ms. Three questions on the Thor go from 16 ms to 20 ms. On the 4090 they go from 8 ms to 21 ms. A 3.4K-token state is 102 ms on the 4090 and 1,640 ms on the Orin Nano. The omni card says it does not report inference numbers. The GGUF card prints llama-server -hf LiquidAI/d1-3B-GGUF:Q8_0 and then POST /v1/systemone on port 8080. We did not run it.
This site's reading
Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.
Verify
Primary post is @liquidai, October 7, 2026, 17:01 UTC, status 2107878924831379676. It names Open d1, d1-3B for text and vision, and d1-omni-600M for text plus an image or text plus audio. The attached chart is titled "Decision Index against model size." It marks d1-3b and d1-omni-600m and draws lines labelled d1 (hosted API) and Jev 1.13 (hosted API). We quote the card's printed scores, not a reading off that plot.
Reply 2107878926953701478 says d1-3B ranks first among models under 10B on Decision Index v0.2.1 and is built from LFM2.5-VL-3B. Reply 2107878929256304819 calls d1-omni-600M an experimental 600M-parameter model and says it leads their text comparison on toxicity detection and paraphrase identification. Reply 2107878931433152755 prints one-question latency, one request at a time: RTX 4090 8 ms, Jetson AGX Thor 16 ms, Jetson AGX Orin 64 GB 26 ms, Jetson Orin Nano 50 ms, and says both models have day-one llama.cpp support. Reply 2107878932892832065 says they built 10 live-camera demos and, with NVIDIA Robotics, showed d1-3B in Isaac Sim with the model on a Jetson. Reply 2107878935270965420 points at the blog and the two Hugging Face repos. We did not open the arcade space or the Jetson AI Lab pages the blog links.
The blog https://www.liquid.ai/blog/d1-open is dated OCT 7, 2026. It prints d1-3B at 48.57 on Decision Index v0.2.1, public split, and d1-omni-600M at 15.95. The chart caption says public leaderboard from Oct 6, 2026, scored with the official scorer on the same 120,226 requests, not submitted. The blog's BibTeX note prints www.liquid.ai/blog/open-d1. The page we opened is /blog/d1-open. We did not open the other path.
Blog table, seven benchmarks, columns d1-omni-600M / d1-3B / Decider 2B / Decider 4B. SQuAD 2.0: 74.0 / 85.3 / 67.7 / 76.0. Civil Comments: 95.8 / 93.0 / 93.6 / 92.8. MASSIVE intent: 86.1 / 87.3 / 81.1 / 88.3. PubMedQA: 61.3 / 66.0 / 65.7 / 63.3. BoolQ: 77.7 / 86.7 / 87.3 / 89.0. XNLI: 74.7 / 85.0 / 85.0 / 88.6. PAWS-X: 79.5 / 76.9 / 59.5 / 69.8. Mean: 78.4 / 82.9 / 77.1 / 81.1. The blog says edition 0.3 has a private vision split and that this release does not report it.
Blog edge table, one question / 3 questions / 3.4K-token state / 384px image / 64 states packed. Apple M5 Pro: 30 ms / 41 ms / 640 ms / 62 ms / 78 per second. Jetson AGX Thor: 16 / 20 / 220 / 35 / 262. Jetson AGX Orin 64 GB: 26 / 35 / 560 / 83 / 110. Jetson Orin Nano: 50 / 73 / 1,640 / 202 / 38. GPU table. RTX 4090: 8 / 21 / 102 / 17 / 475. AMD MI325X: 9 / 14 / 44 / 18 / 1,106. The blog says inference numbers are for d1-3B.
https://huggingface.co/LiquidAI/d1-3B, opened October 8, 2026. Total parameters 3.12B. The index table rounds the size to 3B. Context 32,768. Vocabulary 128,000. Vision encoder SigLIP2 NaFlex 400M. Base LiquidAI/LFM2.5-VL-3B. license_name lfm1.0. transformers>=5.14 and trust_remote_code. Question types noul, choice, and score. Score is 2 to 10 ordered levels. The return line includes output_tokens 0. Decision Index 0.2.1, official scorer, not a leaderboard submission. Other rows are from the public leaderboard v0.2.1. d1-3B 48.57, Knowledge 23.8, Language 56.4, Retrieval 52.8, Tools 74.5, Arts 36.3. Winnow-12B 50.02. Decider 35B-A3B 47.11, size column 36B. JPT-9B 46.89. Decision 1.0 Lux 43.49. JPT-4B 43.04. Jet v6.2 42.60. Decider 4B 40.70. Winnow-E4B 39.89. Decider 2B 28.97. The card's table has no Jev row.
The same card's eight-benchmark table adds HelpSteer2 36.7 against Decider 4B 42.0 and Decider 2B 32.0, and prints means 77.1 / 76.2 / 71.5. DecisionBench eng v1, all 23,900 rows, is 71.8. Fast Decisions dev split is 69.3. Eleven image benchmarks average 74.1 against LFM2.5-VL-3B at 73.9. Cells, d1-3B / base: AI2D 79.9 / 80.9, BLINK 59.2 / 58.7, CV-Bench 82.1 / 87.6, HallusionBench 65.3 / 65.0, MMBench 84.9 / 84.3, MME 82.1 / 82.4, MMStar 59.9 / 61.2, MMVP 77.0 / 73.7, POPE 88.5 / 90.1, VisualWebBench 71.4 / 78.3, VL-RewardBench 65.0 / 50.9. ImajevBench, dev and calibration, 253 rows, is 64.0 against 66.8. With the images removed, the same questions score 45.1.
The card's GPU rows are bf16, median of 20 runs. It says the RTX 4090 one-question cell uses model.compile as CUDA graphs, and that without it one question takes 16 ms. The first call with a new shape pays for compilation.
https://huggingface.co/LiquidAI/d1-omni-600M, opened the same day. Title says 600M. Details say 587M: 381M trunk and decision head, 94M vision encoder, 112M audio encoder. Base LFM2.5-Encoder-350M. Vision tower from LFM2.5-VL-450M. Audio encoder is a 17-layer FastConformer. Context 16,384. With images, state and question text is cut to 896 tokens. Vocabulary 65,536. license_name lfm1.0. transformers>=5.15. Images or one 16 kHz mono clip, not both. Training clips are cut at 30 seconds, on English speaker-and-assistant requests. The card says float16 matched the float32 top answer on 243 text, 214 image, and 416 audio rows, and that bfloat16 changed the top answer on 0.8% of text rows and 1.7% of audio rows. Text answers use per-type temperatures in config.json. Image and audio answers are the softmax as trained. The index row is 15.95, Knowledge 8.3, Language 12.9, Retrieval 35.0, Tools 15.1, Arts 6.8. Fast Decisions dev split is 76.9. The card says it does not report inference numbers, and that HelpSteer2 may overlap training data so that row is left out of its seven-benchmark table.
The LICENSE file on LiquidAI/d1-3B is titled LFM Open License v1.0. Section 5 says commercial use is conditioned on the legal entity not exceeding $10,000,000 annual revenue, and that commercial use above that threshold is not licensed. A qualified nonprofit's use for non-commercial or research purposes is outside the threshold. We did not compare the omni LICENSE file word for word. Its card names the same license_name.
https://huggingface.co/LiquidAI/d1-3B-GGUF prints llama-server -hf LiquidAI/d1-3B-GGUF:Q8_0 and then POST /v1/systemone on 127.0.0.1:8080. The d1-3B docs page prints the same command and the same path. https://docs.liquid.ai/lfm/models/decision-models lists d1-3B at 3.1B, d1-omni-600M at 587M, and hosted d1 as a separate API card. The d1-3B specs table on the docs page prints 3.12B and 32K tokens. A Hugging Face model list for author LiquidAI, fetched October 8, also names LiquidAI/d1-omni-600M-GGUF and LiquidAI/d1-3B-w8a8. We did not open those two READMEs. We did not download the weights, and we did not run llama-server.
Compare
48.57 is the d1-3B card's self-run of Decision Index 0.2.1. 15.95 is the omni card on that same index. The September 29 hosted chart, on Liquid d1, is 58.9 against Jev 1.13 at 57.9. The JevBench API board opened October 7 ranks the hosted name d1 at composite 73.0. Decision Index 0.3 on this desk ranks Perplexity Decider v1.1 at 62.8, with Jev 1.13.0 as a reference at 60.1. The blog says this release does not report the 0.3 vision split.
The card's index table has no Jev row. Winnow-12B at 50.02 on that table is a different figure from JevBench v1.5.0, where Winnow-12B Q8 is 73.2. Tools at 74.5 on this self-run is a different cell from the hosted chart's tools cell of 74.1.
The seven-benchmark mean is 82.9 for d1-3B. The eight-benchmark mean on the d1-3B card, which adds HelpSteer2 at 36.7, is 77.1. Decider 2B's 77.1 is the seven-benchmark mean. Those three 77.1 figures are three rows.
Fast Decisions dev split is 69.3 for d1-3B and 76.9 for d1-omni-600M. DecisionBench 71.8 is the d1-3B card only. ImajevBench 64.0 on 253 rows is the card's comparison with LFM2.5-VL-3B, not a rank on the imajev-4b page.
The local call on the GGUF card is POST /v1/systemone on port 8080. Hosted d1 on the September 29 page is POST https://api.liquid.ai/decisions/v1/systemone, model d1:free. Access on this desk does not list either host as a door for calling Jev. The 8 ms cell is the compiled 4090 row. The card also prints 16 ms without that compile.
Terms
- Open d1
- Liquid AI's October 7, 2026 name for the open checkpoints d1-3B and d1-omni-600M. The hosted model d1:free stays on the Liquid d1 page.
- 48.57
- d1-3B on the card's self-run of Decision Index 0.2.1, official scorer, not a leaderboard submission. d1-omni-600M on the omni card is 15.95. Winnow-12B on the 3B table is 50.02.
- lfm1.0
- The license_name on both cards. The d1-3B LICENSE file is LFM Open License v1.0. Section 5 conditions commercial use on annual revenue under $10,000,000.