A Cloud Run recipe serves DiffusionGemma-Jev at 35 to 60 ms

This topic was created 10 days ago, and the information it contains may have evolved or changed since then.

Google's Gemma account posted a one-command Cloud Run deploy for DiffusionGemma-Jev. The README puts single-request latency at 35 to 60 ms, batch throughput at 100 to 123 requests a second, cold start at about 47.5 seconds, and the GPU at $3.19 an hour while it is up. The repository lists no license. It does not print an accuracy table.

Google’s Gemma account posted on September 22 that DiffusionGemma-Jev can be deployed to Cloud Run with one command. The post puts single-step latency at about 35 to 60 ms, batch throughput at 32-wide concurrency at about 100 to 123 requests a second, and the bill at roughly $3 an hour, dropping to $0 when the service is idle. The next post credits Matt Mastracci and Daniel Lee, and quotes Lee’s September 20 note. That note links taeold/djev-run. GitHub lists no license on the repository.

The README serves ghcr.io/taeold/djev-run:latest, which it says is built from mmastrac/djev-spark. The deploy command asks for one nvidia-rtx-pro-6000, 20 vCPU, 80 GiB of memory, concurrency 32, and a maximum of one instance. Minimum instances is 0. The regions it names for that GPU are us-central1, europe-west4, asia-southeast1, and asia-south2. Weights are nvidia/diffusiongemma-26B-A4B-it-NVFP4. The page describes them as 17.5 GB, copied into /dev/shm while Python imports torch, at about 1.05 GiB a second from Cloud Storage. Cold start from zero instances is about 47.5 seconds. The page says an earlier baseline was 4 minutes 5 seconds, and that baking the weights into the image took 80 to 115 seconds.

The latency line matches the post, with one extra setting: about 35 to 60 ms when steps=1 and samples=1, and about 160 ms when samples is auto. Batch throughput is about 100 to 123 requests a second at concurrency 32. The price on the page is $3.19 an hour while the GPU is active. The post’s figure is “roughly $3.”

The server implements POST /v1/systemone. A sample uses @ai-sdk/typesafe-ai and points baseURL at the Cloud Run /v1 path, then calls evaluationModel('jev-latest'). snake.html is a single file that calls the same route from the browser, at about 15 moves a second. The README prints no accuracy table and no JevBench row.

Maisa’s hosted djev is a different service. JevBench lists that API at 74.3, with hard-tier accuracy 69.5% and a p50 of 0.24 seconds. The local stack at Davipar/djev-dev speaks POST /v1/request. This recipe speaks /v1/systemone on a GPU billed by the hour. The 35 to 60 ms cell is the samples=1 note on an RTX PRO 6000, not the hosted bench timing. We did not deploy it.

This site's reading

Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.

Verify

Primary post is @googlegemma, September 22, 2026, 23:13 UTC. Text: one command, single-step latency about 35 to 60 ms, batch at 32 about 100 to 123 requests a second, roughly $3 an hour and $0 when idle.

The next post in the thread credits @mmastrac and @dylayed and quotes Daniel Lee's September 20 post, 16:56 UTC, which links https://github.com/taeold/djev-run. GitHub says the repository was created 2026-09-20 and lists no license. README: image ghcr.io/taeold/djev-run:latest, built from mmastrac/djev-spark.

GPU nvidia-rtx-pro-6000, 20 vCPU, 80 GiB, concurrency 32, min instances 0, max instances 1. Regions named for that GPU: us-central1, europe-west4, asia-southeast1, asia-south2. Weights are nvidia/diffusiongemma-26B-A4B-it-NVFP4, described as 17.5 GB, staged into /dev/shm.

Cold start about 47.5 seconds, down from a 4 minute 5 second baseline. Single-request latency about 35 to 60 ms at steps=1, samples=1, and about 160 ms with samples set to auto. Batch throughput about 100 to 123 requests a second at concurrency 32.

Price on the page: $3.19 per hour while active, $0 at min-instances 0. The server implements POST /v1/systemone. snake.html calls that route from the browser at about 15 moves a second.

The sample client uses @ai-sdk/typesafe-ai and evaluationModel('jev-latest') pointed at the Cloud Run URL. No accuracy, Brier, or JevBench row is printed. We did not deploy it.

Compare

Maisa's hosted djev API is a different service. JevBench lists that row at 74.3, hard-tier 69.5%, p50 0.24 s, and the announced price is $0.035 per million input tokens. Davipar/djev-dev speaks POST /v1/request.

This recipe speaks /v1/systemone on a GPU you pay for by the hour. The 35 to 60 ms figure is the README's samples=1 note, not that JevBench timing.

Terms

djev-run
Daniel Lee's Cloud Run recipe, GitHub taeold/djev-run. It serves mmastrac/djev-spark on an NVIDIA RTX PRO 6000 and exposes POST /v1/systemone. The repository page lists no license.
samples=1
The README's fast latency setting, about 35 to 60 ms. The same page says samples set to auto is about 160 ms. Neither cell is a JevBench score.

Sources

  1. Google Gemma, Cloud Run deploy
  2. Google Gemma, credits
  3. Daniel Lee, original deploy post
  4. taeold/djev-run