DSPy 3.4.0 calls Jev from Predict, with a gate at 0.2 and 0.6

This topic was created 8 days ago, and the information it contains may have evolved or changed since then.

DSPy 3.4.0 adds experimental Noul, Choice, and Score types plus a TypeSafe client, and a patterns repo pins jev-1.13.0.

DSPy 3.4.0, released September 25, 2026, adds a TypeSafe client to stanfordnlp/dspy. The repository is MIT.

The notes, written by Isaac Miller with Drew Breunig on the Jev section, install the extra as dspy[typesafe] and configure dspy.experimental.TypeSafe("jev-latest"). Predict turns a signature into Jev questions when every output is a decision.

Noul returns .value, .probability, and .confidence. Choice returns a value plus option probabilities. Score returns a continuous value, a level, and rubric probabilities. The notes say Noul’s confidence is the distance from the configured threshold. They do not call that distance a calibration statistic.

A signature that needs generated text has no fallback to a language model. Decision streaming, and generative optimizers with this client, are listed as unsupported. ReAnchor fits Boolean thresholds, Score cuts, and Choice weights against the program’s metric. The notes do not print an accuracy for that fit.

A six-file agent on jev-1.13.0

Drew Breunig posted cmpnd-ai/dspy-system-one-agent-patterns on October 1. GitHub’s license field for that repository is empty.

The README pins dspy[typesafe]==3.4.0. config.py sets Jev to TypeSafe("jev-1.13.0"), the model that writes code to openai/gpt-6-luna, and the cheaper model to openai/gpt-6-sol.

A comment in that file says Claude Sonnet 5.5 and Opus 5.5 refused the RLM action prompt in the author’s testing, which is why the series uses the OpenAI model. The sandbox is Monty. It has no filesystem, network, or environment access unless a tool the series registers provides it.

The permission gate sits inside run_shell_command. DENY_BELOW is 0.2 and ASK_BELOW is 0.6. Below 0.2 the command is refused. From 0.2 up to 0.6 the user is asked. At 0.6 or above it runs.

fit_thresholds.py fits the ask line with ReAnchor on 32 labeled commands, and it scores a false approval as ten times a false denial. The README does not print the fitted number. agent_v5_final.py is written to print token counts and whether a patch fixes one failing test. Those figures are not in the README either.

Slides from the office hours

Nathan LeClaire posted highlights from the DSPy office hours the same night. One slide matches the gate: pytest -q at 0.91 runs, rm -rf / at 0.01 is denied, and a curl piped to sh at 0.02 is denied. A second slide is a dspy.Flex and GEPA sketch and has no score.

The third slide is a cost chart, “Jev to Luna,” on 100 test pairs. The caption says the marked policy was frozen on a validation set of 30.

Standalone Luna is printed at 92.5 percent and $0.746 per 1,000. A live threshold near 0.78 is printed at 88.0 percent and $0.583 per 1,000. The left end of the curve is labeled “Jev only” and has no numeric callout.

One sentence in the thread quotes a participant: Jev agreed with Opus 5, used as a DSPy judge, 90 percent of the time and was 240 times cheaper. The thread does not attach the set or the prices behind that sentence.

TypeSafe called the recording a useful set of examples. That post does not add a figure. The September 30 invite links a StreamYard watch page. We did not play it.

This site's reading

Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.

Verify

DSPy 3.4.0 was released September 25, 2026 by Isaac Miller. The notes add dspy.experimental.TypeSafe and the types Noul, Choice, and Score. stanfordnlp/dspy is MIT. The notes print no accuracy table. ReAnchor is described there and also prints no score.

Drew Breunig, @dbreunig, October 1, 2026, 00:10 UTC, status 2105450254791512371, links cmpnd-ai/dspy-system-one-agent-patterns. GitHub's license field for that repo is empty. config.py pins TypeSafe("jev-1.13.0"). agent_v1 sets DENY_BELOW to 0.2 and ASK_BELOW to 0.6.

The README says fit_thresholds.py uses 32 labeled commands and does not print a fitted threshold. Nathan LeClaire's thread, status 2105417902543528327, attaches the chart saved for this page.

Printed callouts are Luna at 92.5 percent and $0.746 per 1,000, and a live threshold near 0.78 at 88.0 percent and $0.583 per 1,000, on 100 test pairs. We did not run the demos or watch the recording.

Compare

The release notes and the patterns README print no JevBench row. Strands Decider's 167 of 231 is a different repository and a public-task count. This chart's 100 pairs are not that 231.

The permission gate's 0.2 and 0.6 are hand-set constants in one file. toolgate and the Vercel form router use other cutoffs, 0.15 in a stored script note and 0.95 on a form destination. Those are other programs.

Terms

ReAnchor
The experimental DSPy 3.4 optimizer that fits Boolean thresholds, Score cuts, and Choice weights to a program metric. The release notes print no accuracy, and the patterns README does not print the fitted ask threshold.
0.6
ASK_BELOW in agent_v1_permission_gate.py. At or above 0.6 the shell command runs. From 0.2 up to 0.6 the user is asked. Below 0.2 it is refused.

Sources

  1. DSPy 3.4.0 release notes
  2. stanfordnlp/dspy
  3. Drew Breunig, October 1
  4. cmpnd-ai/dspy-system-one-agent-patterns
  5. Nathan LeClaire, September 30
  6. Drew Breunig, September 30
  7. TypeSafe AI, October 1