Published
Updated
DiffusionGemma as Jev reads one denoising step, and a hand chart sums to 198 of 201
Google's Gemma account pointed at vLLM pull 57250 on September 29: seed a canvas, read yes/no, choice, and score probabilities in one denoising step. Matt Mastracci's September 17 chart adds to 198 of 201 against Jev at 191 of 201. A public 231-task replay is 196 against a published 200. On a 100-point routing file, djev-spark's joint score is 32.2% against Jev at 61.4%.
Google’s Gemma account posted on September 29 that vLLM can turn DiffusionGemma into a Jev-like model. The server seeds a canvas with a response template and, in one denoising step, reads confidence and a probability distribution for a yes/no question, a multiple-choice question, or a scored question. The next post thanks the vLLM project and Matt Mastracci and links pull request 57250. Lucas Wilkinson merged it on September 22.
The name is older than that post. On September 18 the same account called the setup DiffusionGemma as Jev, quoted Matt’s September 17 comparison, and put a single pass at about 0.2 seconds on a DGX Spark. That post says bidirectional attention yields well-calibrated decision distributions. Matt’s September 16 thread says the probabilities are the model’s own marginals, and that they are not calibrated against labelled data. He averages four reads with different noise and reports the standard error.
The model card lists 25.2 billion parameters, 3.8 billion active, a canvas of 256 tokens, a 262K vocabulary, and Apache 2.0. The specification table says text and image. The introduction also names video. Normal generation denoises a canvas of up to 48 steps. The Jev-style path asks for the first step only.
The pull request is the machinery. A discrete diffusion model denoises the whole canvas on each forward pass. If the canvas already holds the answer’s fixed text, and only the answer slots are noise, that one step is a distribution over each slot. The example server is not a standard vLLM route. It speaks POST /v1/systemone, with noul, choice, and score. A label has to be a single token, so a long option name is mapped to a short one. The write-up says that with 26 options only 1 to 8 of the label ids ranked in the top-k, which is why the sampler now asks for those ids.
On the pull request’s DGX Spark test, one-way traffic is 8.7 requests a second at 0.12 seconds. At 32-way it is 54.0 requests a second at 0.58 seconds, about 162 decisions a second. The setup is read-only single reads, a 32-row canvas, and three decisions per request. A programming-language battery in the same note is 10/10. Human language is 9/10: Portuguese was called French, at 0.47. Unit comparison is 10/12.
mmastrac/djev is Apache 2.0. The README says this repository is the home of that example server, and that the copy inside the pull request is a snapshot. The vLLM recipe, updated September 23, pins a nightly image, serves the model as dgemma, and starts structured_server.py on port 8011 with a canvas of 64. The curl example in that recipe sends "model": "jev-latest". The djev README says a router that discovers models from GET /v1/models has to send the served name, dgemma.
Matt’s September 17 chart is eight hand sets. The fractions on the bars add to 198 of 201 for DiffusionGemma and 191 of 201 for Jev. Both are 12/12 on tickets, 10/10 on code language, 10/10 on a 10-way language set, 10/10 on a 26-way language set, 10/12 on unit comparison, and 38/38 on entity type. Reviews are 20/20 against 19/20. PII per word, five sentences, is 88/89 against 82/89. The entity-type row is labelled “42 words” and scored 38/38. He called the pair roughly tied, and said DiffusionGemma wins a bit on the PII test. The disagreement table shows the unit-comparison misses landing on different items: each model is wrong on two of the twelve, and they are not the same two. The items, the code, and the per-item outputs are not in the thread.
The latency chart on that thread times a warm second pass on the DGX Spark, one read per request, against Jev on one kept-alive connection. It subtracts a measured 46 ms round trip from Jev and says that trip is still inside the request. One read is faster on seven of the eight sets. Entity type is the exception, 215 ms against 166 ms. With the default policy that re-reads an uncertain question up to four times, DiffusionGemma is slower than Jev on all eight. On September 25 he posted that a lightly tuned Vertex A100 80 GB endpoint, about $3,000 a month, runs 250 decisions a second at a latency on par with the official API. That post has no table.
Chirag’s sieve README, MIT, says reversing the options flipped 24% of DiffusionGemma-Jev’s answers. Averaging three reads in shuffled order halved that, and added 17 to 26 points on sets with 60 to 151 options. The chart on the September 26 post prints 23.6% on one read, 14.3% on two, and 11.8% on three, against GLiNER at 3.0% on one read. The point gains on that chart are +17 on CLINC150 (151 options), +26 on MASSIVE Hindi (60), +24 on MASSIVE English (60), and +18 on Banking77 (77).
Josh Brown posted a Perch Bug Bench chart the same day as Google’s vLLM note. It is average precision on 2,220 public before-and-after pairs: DiffusionGemma Jev 46.8%, Jev 46.2%, SemIf 44.6%, Span-01 43.8%, CLM-v0.1-8B 42.1%. His parent post says the rule-based checks gave the strongest results, so 46.8% is the standard-detection figure on the chart. The source line points at perchscan.com/benchmark-results. That URL returned 404 when we opened it.
A separate replay, Hangzhi/diffusion-jev-sglang, measured September 21 on one A100 80 GB. On the 231 public JevBench tasks at commit c6004e0, local DiffusionGemma is 196/231 (84.8%) against the published Jev 1.13.0 artifact of 200/231 (86.6%). Easy is 48/48 for both. Original is 70/72 against 71/72. Hard is 78/111 against 81/111. The client-observed median is 183.3 ms against the published 665.0 ms. The report says the local clock is loopback on the A100 and the Jev clock includes the production path from Germany, and that no fresh hosted run completed because the gateway worker exited before any hosted item was measured. Ten-bin expected calibration error on the text tasks is 0.1222. Twenty of 35 wrong text answers had a top probability of at least 90%. A 100-image flower test is 98/100, both misses sunflowers called daisies, with no hosted Jev image column. That service reads the last active denoising step. Text requests used 2 to 6 steps, and 212 of 231 used two, beside the one-step server in the vLLM example.
ywchiu’s summary is 100 decision points, five repeats, teacher-forced, on an 11-question design. djev-spark, named there as DiffusionGemma 26B-A4B, scores 32.2% on the four-field joint against Jev 1.13.0 at 61.4%. The six-field joint is 13.4% against 48.8%. Median latency is 1,230 ms against 749 ms. The file says djev varies from vLLM nondeterminism at a fixed seed, and that Jev is the system that varies on its own.
Maisa’s hosted djev is a different service, and djev-run is the Cloud Run recipe. The hosted JevBench row, and the 35 to 60 ms samples=1 note, stay on those pages. We did not serve the model.
This site's reading
Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.
Verify
Google's Gemma account, September 29, 2026, 17:42 UTC, says vLLM seeds a canvas with a response template and reads confidence and probability distributions for yes/no, multiple-choice, and scored questions in one denoising step. The next post links vLLM pull 57250 and thanks @vllm_project and @mmastrac. Lucas Wilkinson merged that pull request on September 22.
The September 18 post is the one that uses the name DiffusionGemma as Jev. It quotes about 0.2 seconds on a DGX Spark and says the distributions are well calibrated. Matt's September 16 thread says the probabilities are the model's own marginals and are not calibrated against labelled data.
The model card lists 25.2 billion parameters, 3.8 billion active, a canvas of 256, vocabulary 262K, and Apache 2.0. Its specification table says text and image. The introduction also names video.
Pull 57250's DGX Spark note: one-way 8.7 requests a second at 0.12 seconds, 32-way 54.0 requests a second at 0.58 seconds, about 162 decisions a second. Read-only single reads, a 32-row canvas, three decisions per request. The same write-up prints programming language 10/10, human language 9/10 (Portuguese called French), and unit comparison 10/12. It also says that with 26 options only 1 to 8 label ids ranked in the top-k before the sampler asked for those ids.
mmastrac/djev is Apache 2.0. The README says this repository is the home of the example server and the copy in the pull request is a snapshot. The vLLM recipe curl sends model jev-latest. The djev README says a router that discovers models from GET /v1/models has to send the served name, dgemma.
The September 17 chart prints eight fractions. They add to 198 of 201 for DiffusionGemma and 191 of 201 for Jev: 12/12, 10/10, 10/10, 10/10, 10/12, 20/20 against 19/20, 38/38, and 88/89 against 82/89. The entity-type bar is labelled "42 words" and scored 38/38. The items are not in the thread. The latency chart is a warm second pass on a DGX Spark against Jev on a kept-alive connection, with a measured 46 ms round trip subtracted from Jev. One read is faster on seven of eight sets. The entity-type set is 215 ms against 166 ms. Auto re-reads, up to four, are slower than Jev on all eight. The September 25 post says a Vertex A100 80 GB endpoint, about $3,000 a month, runs 250 decisions a second. That post has no table.
NakliTechie/sieve, MIT, says reversing the options flipped 24% of DiffusionGemma-Jev's answers, and that three shuffled reads halved that and added 17 to 26 points on sets with 60 to 151 options. The chart prints 23.6%, 14.3%, and 11.8%, against GLiNER at 3.0%, and +17, +26, +24, and +18 on CLINC150, MASSIVE Hindi, MASSIVE English, and Banking77.
Josh Brown's September 29 chart is average precision on 2,220 public before-and-after pairs: DiffusionGemma Jev 46.8%, Jev 46.2%, SemIf 44.6%, Span-01 43.8%, CLM-v0.1-8B 42.1%. The parent post says the rule-based checks gave the strongest results. perchscan.com/benchmark-results returned 404.
Hangzhi's report, September 21, one A100 80 GB: 196/231 against the published Jev 1.13.0 artifact of 200/231 at commit c6004e0. Easy 48/48 and 48/48, original 70/72 and 71/72, hard 78/111 and 81/111. Client median 183.3 ms against the published 665.0 ms. The report says the local clock is loopback and the Jev clock includes the path from Germany, and that no fresh hosted run completed. Text ECE is 0.1222. Twenty of 35 wrong text answers had top probability at least 90%. Flowers are 98/100, with no Jev image column. That run reads the last active denoising step, often two steps.
ywchiu's summary: 100 decision points, five repeats, teacher-forced. djev-spark's four-field joint is 32.2% against Jev 61.4%. Six-field joint is 13.4% against 48.8%. p50 is 1,230 ms against 749 ms.
We did not serve the model. TypeSafe's Master Customer Agreement section 2.3(f) forbids publishing benchmarks of the Services. The Jev cells above are reported as published.
Compare
Maisa's hosted djev is a different service. On the v1.2.8 geometric mean this desk already recorded, that API is 74.3, with hard-tier 69.5% and Calibration 65.4. djev-run is the Cloud Run recipe, 35 to 60 ms at samples=1, and it prints no accuracy table.
Open-Jev's 27B v1.1 card on the same 231 public tasks is 197/231 against the published Jev row of 200/231. Hangzhi's DiffusionGemma read on that slice is 196/231. The hand chart is a different set, and the author built the reader he is scoring.
ywchiu's routing file is the place the same family of read falls away: 32.2% joint against 61.4%, and 13.4% on six fields against 48.8%. Chirag's option-order chart is a separate failure. One read flips 23.6% of answers when the options are reversed. GLiNER on that chart flips 3.0%.
Terms
- DiffusionGemma as Jev
- A one-step read of google/diffusiongemma-26B-A4B-it. The canvas is seeded with the answer template, and the label slots stay noisy. mmastrac/djev, Apache 2.0, is the example server in front of vLLM pull 57250.
- Seeded canvas
- The fixed text of the answer is written into the diffusion canvas before the forward pass. One denoise step then returns a distribution at each noisy slot.
Sources
- Google Gemma, September 29
- Google Gemma, pull 57250
- Google Gemma, DiffusionGemma as Jev
- Matt Mastracci, September 17 chart
- Matt Mastracci, September 16 note
- Matt Mastracci, 250 decisions a second
- vLLM pull 57250
- mmastrac/djev
- vLLM DiffusionGemma recipe
- DiffusionGemma model card
- Chirag, option order
- Josh Brown, Perch Bug Bench chart
- Hangzhi, public 231
- ywchiu, routing summary