Published
Updated
Cloudflare posts Clef and Clef-flash, Apache 2.0 decision models on Workers AI
Cloudflare posted Clef and Clef-flash on October 1, 2026, two Apache 2.0 decision models hosted on Workers AI and released on Hugging Face. The model cards list input at $0.24 and $0.09 per million tokens. The blog's median latency row is 209.3 ms and 38.8 ms against a printed Jev cell of 524.1 ms. On October 3 a Fastino chart prints Clef at 61.2 and Clef-flash at 57.1, and captions those Cloudflare scores as self-reported. llama.cpp v0.6.0, published October 5, serves Clef on /v1/systemone, including images. The October 6 note leaves Clef's speed cell blank.
On October 1, 2026 at 19:51 UTC, @Cloudflare posted Clef and Clef-flash. The post calls them open-source decision models on Workers AI for classification and agent workflows, and it links the Birthday Week write-up. Michelle Chen is the byline on that post. The page says the weights are Apache 2.0 on Hugging Face, and that the hosted API accepts a Jev-shaped request.
Clef is the 27B model, post-trained from Qwen3.8-27B. Clef-flash is the 9B model, post-trained from Qwen3.5-9B. Both cards say the checkpoint keeps the base model’s vision encoder, adds a joint schema head, and returns one logit per allowed option. A softmax over each question turns those logits into probabilities. The cards say there is no free-form text to parse.
What you send
The hosted ids are @cf/cloudflare/clef and @cf/cloudflare/clef-flash. The docs show a Workers binding, env.AI.run, and a REST call:
POST https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run/@cf/cloudflare/clef
The JSON body sets model to clef or clef-flash, plus state and questions. Question types are noul, choice, and score, the same three names Jev uses. The clef input schema allows 1 to 64 questions. A choice map is described as 2 to 255 options. A score list is 2 to 10 levels, lowest first. instructions is required on that schema. The local encode_record helper on the model card says instructions can be omitted, and that the question id is used instead.
The docs print a context window of 65,536 tokens for both models. The blog rounds that to 64k and sets it next to “Jev’s 32k”. The Workers AI card for typesafe/jev, already on Access, prints 32,000 tokens. TypeSafe’s own models page, as recorded on Access, says a Jev request is 64k tokens in total, and that state plus the longest question is 32k. The local helper’s default max_length is 16,384. The hosted window on the docs is 65,536.
Images on the hosted schema are an optional list of at most four embedded PNG, JPEG, or WebP items. Remote URLs are rejected. Each image is capped at 4 MiB and 16 megapixels, the decoded set at 8 MiB, and the whole body at 13 MiB. The model card also accepts video as frame arrays in a local record. The hosted schema we opened has no videos field. The docs intro still says the model reads video.
The clef card lists $0.24 per million input tokens. The clef-flash card lists $0.09. Neither page prints an output price. Jev’s published price on this desk is $0.042 per million input tokens, with output free. The blog says Cloudflare does not read, store, or train on requests or responses, unless you use the fine-tuning product described later on the same page.
The table they printed
The blog says Clef leads the Jev Decision Index on a live demo. We did not open that demo. The table under the sentence is ten rows. On seven of them the higher cell is Clef or Clef-flash. Jev is higher on When2Call and on BRIGHT. DiffusionGemma Jev is higher on PhishNChips.
| Benchmark | Clef | Clef-flash | Jev | DiffusionGemma Jev | Kev 9B | Laya |
|---|---|---|---|---|---|---|
| BFCL, case exact | 98.47 | 98.76 | 95.75 | 96.52 | 94.51 | 38.13 |
| ToolRet, nDCG@10 | 69.19 | 66.43 | 65.28 | 61.21 | 64.26 | 12.69 |
| API-Bank, accuracy | 91.93 | 93.11 | 88.19 | 83.66 | 56.30 | 11.41 |
| Home appliances, case exact | 82.95 | 97.73 | 52.27 | 42.05 | 25.00 | 0.00 |
| When2Call, accuracy | 72.37 | 65.58 | 80.97 | 75.44 | 49.62 | 11.94 |
| BANKING77, macro-F1 | 94.20 | 90.93 | 79.74 | 74.28 | 84.83 | 14.29 |
| CLINC150+OOS, macro-F1 | 97.43 | 66.77 | 89.27 | 83.49 | 79.03 | 3.19 |
| BRIGHT, nDCG@10 | 45.91 | 39.26 | 47.52 | 42.94 | 38.53 | 19.90 |
| Amazon ESCI, macro-F1 | 57.48 | 57.39 | 55.21 | 53.37 | 49.22 | 24.40 |
| PhishNChips, accuracy | 79.60 | 75.05 | 62.55 | 85.35 | 50.75 | 50.15 |
Home appliances is the wide gap on that grid: Clef-flash 97.73, Clef 82.95, Jev 52.27. CLINC150+OOS is the wide miss for the small model: Clef-flash 66.77, Jev 89.27, Clef 97.43. BANKING77 here is macro-F1. Laya’s author table and Charly Poly’s encoder bake-off use other Banking77 setups. This page does not treat the three as one comparison.
A second blog table is four workflows, scored against TypeSafe’s eval suite. Invoice processing, exact actions: Clef 64.7, Clef-flash 57.1, Jev 61.8. Customer service: 76.3, 77, 76.0. Security incidents: 62.9, 61.7, 61.7. Agent trace observability: 68.5, 69.8, 71.6. The blog says the Clef models beat Jev in 3 of 4 areas. The invoice row is one of those areas for Clef, and Clef-flash is behind Jev there. Agent trace is the row Jev leads. The model card adds an invoice “primary action” line the blog table omits: Clef 86.2, Clef-flash 73.3, Jev 83.1.
Latency is one median and one p95, in milliseconds, which the blog says was taken across 43 evals. Median: Clef 209.3, Clef-flash 38.8, Jev 524.1, DiffusionGemma Jev 84.4, Kev 9B 51.4, Laya 5.8. p95: 238.6, 122.4, 536.0, 211.2, 187.9, 222.5. The blog says the Clef models are faster than the other decision models in that set except Laya. That exception matches the median row. On p95, Clef-flash’s 122.4 ms is under Laya’s 222.5 ms. 524.1 divided by 38.8 is about 13.5. 524.1 divided by 209.3 is about 2.5. Those ratios are arithmetic on the printed cells. The blog does not print them.
Where the longer card still has Jev ahead
The model card prints the same lab as an internal run of Decision Index 0.2.1, with scores rounded to one decimal. It marks a best cell in bold. Jev has that mark on eleven rows. Five of them, as Clef, Clef-flash, then Jev: GPQA Diamond 48.0, 51.0, 78.3. BBH 73.7, 68.9, 92.9. MMLU-Pro 65.9, 65.3, 82.7. When2Call 72.4, 65.6, 81.0. BRIGHT 45.9, 39.3, 47.5. The other six are ANLI, POP909-CL, VAST, NLI4CT, HoVer, and New Yorker.
CLINC150+OOS on that card is 97.4, 66.8, 89.3. RAGTruth hallucination F1 is 79.4, 35.6, 76.5. ForecastBench is a Brier score, lower better: 13.9, 10.6, 17.4. Clef-flash has the low cell on ForecastBench, and the low cell on RAGTruth.
This is not the composite on the Decision Index page. That page’s kit test locks Jev at 51.67, and the September 28 board file lists Jev 1.13.0 at 57.91. The Clef card does not print a number on that scale. Drex 1.5 and Liquid d1 argue over that composite. They are a different chart from these per-benchmark cells.
How they say it was trained
The blog says the week Jev shipped, the same team posted a DiffusionGemma experiment that read token probabilities. Michelle Chen’s September 18 post is that demo, at kev.workers-ai-mle.workers.dev, on Workers AI. The October post says Clef keeps the idea of scoring schema choices from a backbone, and that the backbone is Qwen. It points at Matt Mastracci’s vLLM pull 57250, which is already written up on DiffusionGemma as Jev.
The training note says Qwen3.8-27B and Qwen3.5-9B stay frozen. A routing head is trained together with rank-256 adapters. The loss they name is label-smoothed cross-entropy on valid schema outputs, plus a Brier term. The data they name is internal and synthetic, with field order, prompts, and schemas permuted. At inference the model runs a prefill, then scores the legal choices in parallel, without generating the answer token by token.
The same section says they developed Reinforcement Learning for Calibrated Decisions, RLCD, as a second target. The mechanics they list are partial credit for a neighboring ordinal choice, a reward for a fully correct record, and a penalty that keeps the distribution from drifting. This desk already uses RLCD as TypeSafe’s name for its own training, on the System One explainer. The Cloudflare post does not say the two procedures are the same one.
One internal example is on the blog, from Cloudflare’s threat-intelligence workflow. A domain plus a browser render is classified in 2.2 seconds. The post’s example distribution is about 95% fashion, 85% ecommerce, and under 1% phishing. The same workflow on gpt-oss-120b took 4.7 seconds and returned two labels. That is one workflow on their side. It is not a row in the table above.
Fine-tuning is a service, not a button on the model card. The blog says a forward-deployed team will tune Clef with a customer first, and that a self-serve trainer comes after that. The pieces they name are AI Gateway for capturing traffic, Workers AI for rollouts, Containers as a scoring sandbox, a new trainer, and bring-your-own-model for redeploying the result. The interest form is cloudflare.com/resource/clef-rl-interest. We did not submit it, and we did not call @cf/cloudflare/clef.
On October 3 Sahibzada Allahyar’s chart prints Clef at 61.2 and Clef-flash at 57.1 on Decision Index 0.2.1, with Jev at the public-board 57.91. The caption says the Cloudflare scores are self-reported, and the axis starts at 56. The model card we read still does not print a composite on that scale. 61.2 is not a replacement for the per-benchmark cells above, and it is not the blog’s latency row.
llama.cpp release v0.6.0, published October 5, serves these weights on POST /v1/systemone. Pull 29831 merged October 3 with commands for ggml-org/Clef-GGUF and ggml-org/Clef-Flash-GGUF, and said images were still missing. Pull 29969 merged October 5 and adds images for the clef type. The server README at that tag says a Clef process serves only that endpoint, and that Clef reads every question in one prompt. A choice can have 255 options. Opened October 6, the Hugging Face note lists Clef at 27B, images, Apache 2.0, and leaves the speed cell blank. Clef-flash is not a row there. The server write-up is llamacpp-jev. We did not run it.
This site's reading
Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.
Verify
@Cloudflare, October 1, 2026, 19:51 UTC, status 2105747536510099540. The text introduces Clef and Clef-flash as open-source decision models hosted on Workers AI, links the blog, and tags Birthday Week. The post carries a link card for the blog. It has no attached photo. We did not call the endpoint.
The blog is https://blog.cloudflare.com/clef-decision-models/. Its schema.org block says datePublished 2026-10-01T15:34:02.111Z. The byline is Michelle Chen. We opened the page on October 2, 2026. The ten-row benchmark table, the four workflow rows, and the latency row on this page are the figures printed there. The page links a live demo at https://clef-evals.workers-ai-mle.workers.dev. We did not open that demo, so no figure here comes from it.
Workers AI docs, opened the same day. @cf/cloudflare/clef is described as a 27B multimodal model with a 65,536-token window, vision marked yes, and unit pricing of $0.24 per million input tokens. @cf/cloudflare/clef-flash is described as a 9B model with the same window and vision flag, at $0.09 per million input tokens. Neither card prints an output price.
The clef input schema JSON, opened the same day. model matches clef or clef-flash. questions allows 1 to 64 ids. A choice criteria map is described as 2 to 255 options. A score criteria list is 2 to 10 levels. images allows at most 4 embedded PNG, JPEG, or WebP items, and says remote URLs are not accepted. The schema has no videos field.
Hugging Face Cloudflare/clef and Cloudflare/clef-flash, opened the same day. Both say Apache-2.0. Clef is 27B, post-trained from Qwen/Qwen3.8-27B. Clef-flash is 9B, post-trained from Qwen/Qwen3.5-9B. The results section says the table is an internal run of Decision Index suite 0.2.1. Scores there are printed to one decimal. encode_record documents a default max_length of 16,384 tokens. The usage note says the snippet was tested with torch 2.11 and transformers 5.10.2 on one H200. We did not download the weights and we did not run that snippet.
Michelle Chen, @michellechen, September 18, 2026, 23:28 UTC, status 2101091012559151480. The post points at a DiffusionGemma demo on Workers AI, https://kev.workers-ai-mle.workers.dev/. The October 1 blog says Clef uses a Qwen backbone after that experiment, and it names Matt Mastracci's vLLM pull 57250. That pull is already on the DiffusionGemma page. We did not open the September demo.
We did not send a Workers AI request. TypeSafe's Master Customer Agreement section 2.3(f) forbids customers from publishing benchmarks of the Services. The tables here are Cloudflare's published tables.
Sahibzada Allahyar, October 3, 2026, 02:42 UTC, status 2106213366813319230, posts a Decision Index 0.2.1 chart. Clef is 61.2 and Clef-flash is 57.1. The caption says those Cloudflare scores are self-reported. Jev on the same chart is 57.91, marked as the public leaderboard. The axis starts at 56. We did not find 61.2 or 57.1 as composites on the model card.
llama.cpp pull 29831 merged October 3 at 00:50 UTC. The commands are llama serve -hf ggml-org/Clef-GGUF and ggml-org/Clef-Flash-GGUF. The pull said images were still missing. Pull 29969 merged October 5 at 12:02 UTC and adds images for Clef. Release v0.6.0 was published October 5 at 16:56 UTC. The server README at that tag says a Clef server serves only /v1/systemone, and that Clef reads every question in one prompt. A choice on Clef can have 255 options. Opened October 6, the Hugging Face note lists Clef at 27B, images, Apache 2.0, with a blank speed cell. Clef-flash is not a row. We did not run llama-server.
Compare
TypeSafe's models page, as recorded on Access on October 1, lists jev-1.13.0 at $0.042 per million input tokens and free output. The Clef cards list $0.24 and $0.09 per million input tokens and do not print an output price. Those are two price lists.
The blog's median row is Clef 209.3 ms, Clef-flash 38.8 ms, Jev 524.1 ms, DiffusionGemma Jev 84.4 ms, Kev 9B 51.4 ms, and Laya 5.8 ms. Dividing 524.1 by 38.8 is about 13.5. Dividing 524.1 by 209.3 is about 2.5. The blog's p95 row puts Clef-flash at 122.4 ms and Laya at 222.5 ms. JevBench medians on this desk are a different task mix.
Cloudflare's BANKING77 row is macro-F1 94.20, 90.93, and 79.74. Laya's author table is 0.425 against a published Jev 0.870, with 77 labels against 72. Charly Poly's Banking77 figure is a third setup. The three pages are not one run.
Decision Index on this desk prints a composite: 51.67 in the kit test, and 57.91 for Jev 1.13.0 in the September 28 board file. Cloudflare's card prints per-benchmark scores from its own 0.2.1 run. We did not find a composite of that kind on the card.
Access already lists Workers AI typesafe/jev with a 32,000-token card. The Clef docs print 65,536. The blog's phrase "Jev's 32k" matches that Workers AI card. Access still records TypeSafe's own cut as a 64k whole request and 32k for state plus the longest question.
OpenAI's Decisions API is a separate host. On October 2 a POST to https://api.openai.com/v1/decisions from this desk returned HTTP 403.
Fastino's October 3 chart prints Clef at 61.2 and Clef-flash at 57.1 on Decision Index 0.2.1, and says the Cloudflare numbers are self-reported. This card still does not print that composite. The blog's median row, 209.3 ms and 38.8 ms against 524.1 ms, is a different measurement.
The llama.cpp note's blank Clef speed cell is not the blog's 209.3 ms median. The 3 ms to 43 ms row on that note is Julia-1, Laya, Kev-4B, lev, and OpenJev, measured October 2 on one RTX PRO 6000.
Terms
- Clef
- Cloudflare's 27B decision model, Apache 2.0, hosted as @cf/cloudflare/clef. The model card says it is post-trained from Qwen3.8-27B.
- Clef-flash
- The 9B model in the same release, hosted as @cf/cloudflare/clef-flash. The model card says it is post-trained from Qwen3.5-9B.
- @cf/cloudflare/clef
- The Workers AI run id for Clef. clef-flash uses @cf/cloudflare/clef-flash. typesafe/jev on the same host is the Jev listing, a different card.
Sources
- Cloudflare, October 1
- Cloudflare blog, Introducing Clef
- Workers AI, clef
- Workers AI, clef-flash
- Clef input schema
- Cloudflare/clef on Hugging Face
- Cloudflare/clef-flash on Hugging Face
- Michelle Chen, September 18
- Sahibzada Allahyar, October 3 chart
- llama.cpp release v0.6.0
- llama.cpp pull 29831, Clef text
- llama.cpp pull 29969, Clef images
- ggml-org decision models note