postgres-search can sink a bad hit, and the README has not measured that step

This topic was created 10 days ago, and the information it contains may have evolved or changed since then.

Alex posted an Algolia-style search in Postgres with an optional Jev pass. On 50 handwritten queries the SQL ranking puts a correct product first 82% of the time. The Jev step asks one Noul per candidate in the top 10 and sinks anything under 0.30. The docs say that pass was not measured, because no API key was available.

Alex posted on September 23 that postgres-search is an Algolia alternative built with Jev. The repository is alexforman1/postgres-search, MIT. The search itself is SQL. Jev is an optional second pass, and the docs say that pass has not been scored.

The demo index is USDA FoodData Central branded foods, the 2025-12-18 release, public domain. The full file is 440,302 products. Warm queries on that file: about 2 ms for cheerios, 22 ms for milk, 53 ms for the misspelling cheerois. Facet counts for milk take 128 to 135 ms. On 50 handwritten queries against the full release, a correct product is first for 82% of them and inside the top 10 for 88%. Those rates are the SQL order. A misspelling that some product already uses can hide the correctly spelled rows: the README’s example is one PARMESEAN product, so parmesean never surfaces the 2,734 parmesan rows.

docs/jev.md describes the model step. After the SQL query returns, one POST /v1/systemone sends the top 10 candidates and asks one Noul per candidate: is this the product the query names? The default model is jev-latest. The default threshold is 0.30. A candidate at or above that line keeps its place. A candidate under it moves to the bottom of the 10, still in the original relative order. Nothing past the 10th result is touched, and the step does not sort by the probability. The page’s reason is that a correct product and a close variant can both score near 1, so a sort would reshuffle good rows. The call is skipped when fewer than two results come back. Any failure leaves the SQL order in place: a missing key, a network error, a non-2xx response, the 1.5 second timeout, or an answer that is missing a number. There is no retry.

The same page says the worked examples (pepper, cream, crackers, apple) were not run through Jev for this release, because no API key was available. scripts/eval.ts would add jev hit@1 and jev hit@3 once a key is set. Hit@10 would match the SQL figure, because Jev can only reorder the top 10. The 0.3 cut was not calibrated on this data. The docs say to pin JEV_MODEL to a versioned id after a threshold is chosen, since jev-latest moves.

jevsearch did run its re-ranker: Hit@1 83% against 41% for keywords, on 41 labelled docs queries. This README prints the SQL rates and leaves the Jev columns empty. We did not load the database or call the API.

This site's reading

Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.

Verify

Primary post is @alex_forman1, September 23, 2026, 00:06 UTC. Display name Alex. Text: an Algolia alternative using Jev, linking the repository. README is MIT. Demo data is USDA FoodData Central branded foods, release 2025-12-18, public domain, 440,302 products in the full file.

Warm timings: cheerios about 2 ms, milk 22 ms, misspelled cheerois 53 ms. Facet counts for milk 128 to 135 ms. On 50 handwritten queries against the full release, a correct product is first for 82% and in the top 10 for 88%.

docs/jev.md: one POST /v1/systemone per search, model default jev-latest, one Noul per candidate in the top 10. JEV_THRESHOLD default 0.3. Candidates at or above the threshold keep their order.

Candidates below it move to the bottom of the 10, still in their original order. Results past the 10th are not touched. The step never sorts by the probability. Fail open: missing key, network error, non-2xx, 1.5 second timeout, or a non-numeric score.

No retry. The call is skipped when fewer than two results come back. The page says what Jev does on the worked examples was not measured for this release, because no key was available.

With a key, scripts/eval.ts would add jev hit@1 and jev hit@3. Jev hit@10 would equal the SQL hit@10, because only the top 10 can move. The 0.3 default was not calibrated on this data.

We did not run the database or call Jev.

Compare

jevsearch re-ranks 41 labelled TypeSafe-docs queries and prints Hit@1 83% against keyword search at 41%, at $0.26 per 1,000 queries. This repository prints the SQL hit rates and says the Jev columns were not filled in. The 82% figure is the Postgres order on 50 handwritten food queries.

Terms

sink, don't sort
postgres-search keeps a top-10 hit in place when its Noul is at least 0.30, and moves a lower score to the bottom of those 10 without reordering either group. The docs say a correct product scores near 1 for close variants too, so sorting by the probability would shuffle good rows.
JEV_THRESHOLD 0.3
The default cut in docs/jev.md. The page says it was not calibrated on the USDA queries, and that the Jev hit rates for this release were not measured.

Sources

  1. Alex, postgres-search
  2. alexforman1/postgres-search
  3. The Jev step