Published
Updated
imajev-4b: Hugging Face weights and JevBench 70.4
Opened October 8, 2026, the JevBench v1.6.1 open-weights list puts imajev-4b at capability rank 31 and 61.9. Image JevBench v0.3.0, opened the same day, lists capability 68.0 at rank 1 and official composite 63.7 at rank 2 of 52. On October 7, 2026 the open-weights list put imajev-4b at capability rank 17 and 61.9. imajev-4b is Mohit Garg's Apache-2.0 LoRA on Qwen3.5-4B, with Hugging Face weights, JevBench v1.5.6 rank 10 at 70.4, and Image JevBench rank 2 at 76.4 on October 4. The September 27 changelog scores the same checkpoint 67.37 at rank 1, and Jev 1.13.0 stays 63.29 on that board. A September 26 adapter moved ImajevBench from 82.4% to 83.9%. We did not rerun the boards.
What is imajev-4b?
imajev-4b is the 4B checkpoint. Weights, the local server, and the October 4 boards are on imajev-4b.
mohit67890/imajev is Mohit Garg’s open typed-decision model. The LICENSE file is Apache-2.0. The 4B checkpoint is a LoRA on Qwen3.5-4B, with a 256-code readout and a frozen vision tower. Weights are at mohit67890/imajev-4b. A request can carry up to two photos. Text-only calls use the same Choice, Score, and Noul shape as Jev, plus an unknown_probability and an abstained flag. We did not run the server.
What did imajev-4b score on JevBench v1.4.2.2?
The independent rank is in the JevBench changelog. v1.4.2.2, dated September 27, adds Imajev-4B at 67.37 and rank 1, using the unchanged v1.4.2 scorer. Prior measurement fields stay. A September 28 note says the cost basis was rewritten to a single pinned server pass, and that no score, axis, rank, eligibility, or numeric value changed. The one-rotation run used the author’s pinned adapter, calibration.json, and pinned kernels.
The capability view on benchmarkheaven.com/jev-models, which we opened September 28, puts Imajev-4B first among Jev-class systems at 66.3. That figure is the average of Intelligence 52.2 and Calibration 80.4. Estimated cost is $0.022 per 1,000 decisions. Median latency is 0.23 s. The same view lists official rank 1. Jev 1.13.0 is capability 64.7, Intelligence 53.1, Calibration 76.3, $0.040, median 0.65 s, and official rank 4. Garg’s screenshot of the composite, dated the same day, prints 67.4, then Plumb-4B 65.8, decider-4b v2 64.1, and Jev 63.3. Those four match the changelog’s 67.37, the Plumb row at 65.84, and the v1.4.2 cells 64.13 and 63.29, rounded.
What did the September 26 adapter change?
His own September 26 note is a different protocol. The phase-3 adapter, trained on decisions the previous release got wrong, moves ImajevBench from 82.4% to 83.9% and a hidden split from 84.2% to 85.6%. JevBench public hard, as shipped with four option orders and the calibration file, moves from 70.3% to 72.1%. DecisionBench’s full suite in that note moves from 77.5% to 79.7%. Abstention on the 21 ImajevBench items whose answer is unknown goes from 14 to 18. Abstention on 258 answerable items goes from 3 to 9. Calibration on DecisionBench gets worse: ECE 0.024 to 0.069. The 2B and 9B checkpoints are unchanged. The 4B leads the 9B on ImajevBench, 83.9% against 82.1%, and the README’s paired test is p = 0.57.
ImajevBench has 279 test questions. At a 90% confidence bar the 4B automates 58% of them and is right on 97.5% of the ones it automates. At 99% it automates 40% and the README prints 100% right. A playground of five small apps, scored without the calibration file, passes 129 of 145 checks. With the file the top answer stays put and confidence drops, so the same 80% bar passes 113 of 145. One miss the README names: a blank payment note answered “not paid” instead of unknown.
The README also shows a screenshot of the Hanno Labs DecisionBench board, captured September 28: imajev-4b at 79.65, third of 56, with jev-1.13 at 71.90 and gpt-5.6-luna at 69.04. A reasoning tab in that screenshot puts imajev-4b at 80.58 and jev-1.13 at 74.50. We did not load the live Gradio app, so those two boards are the screenshot’s figures. The two models above imajev-4b on the main tab are the benchmark team’s own Bosun checkpoints, at 87.29 and 83.20.
Where does imajev-4b sit on later boards?
JevBench v1.5.0, which we opened September 29, is a new board. Imajev-4B is a roster addendum there, not an official rank. The page’s placement against the frozen scores is sixth, at 70.4, with an interval of 67.8 to 71.6. The capability view on that addendum prints 70.8, from Intelligence 53.5 and Calibration 88.1, at about $0.017 per 1,000 decisions. The 67.37 above is the v1.4.2.2 composite. We did not rerun v1.5.
The page we opened September 30 is v1.5.4. It numbers that same 70.4 as official rank 9 and still tags the row roster addendum A2. Decision 4B v1.1 is also 70.4, at rank 10. The September 29 “sixth, outside the ranked order” sentence describes the page opened that day. We did not rerun v1.5.4.
The same URL on October 4 says v1.5.6 and numbers 110 ranked systems of 116. imajev-4b is official rank 10 at 70.4, still tagged roster addendum A2. The axis cells are 53.5, 88.1, 91.1, and 63.3. Estimated cost is $0.017 per 1,000 decisions. Cygnet, Winnow-12B Q8, and Jev 1.13.0 stay 73.7, 73.2, and 72.1, and the page still calls Cygnet and Winnow joint leaders. The September 30 rank for this row was 9. We did not rerun v1.5.6.
On October 7, 2026 the same URL says headline release v1.6.1 and board revision v1.7.12. The open-weights capability list puts imajev-4b at rank 17 and 61.9, Intelligence 34.6, Calibration 89.2, estimated $0.017 per 1,000 decisions, p50 0.29 s. The rank line says official composite 26. 70.4 stays the v1.5.6 point estimate. We did not rerun v1.6.1.
Image JevBench v0.1.5, opened the same day, lists imajev-4b at 76.4, official rank 2 of 50. The compare table on that page prints 76.39, with Intelligence 73.8, Calibration 90.5, Speed 87.6, and Cost 61.2, about $0.020 per 1,000. Wity-1 is rank 1 at 80.2, and that row says the server build ID was not recorded and the score may change after verification. Jev-Omni is 73.1 at rank 3. The setting line for imajev-4b says PyTorch, one rotation, --fast, --merge-lora, and calibration.json. The author’s September 28 screenshot of v0.1.3 prints 76.39 as rank 1 of 49. We did not rerun the image board.
Opened October 8, 2026, the open-weights URL still says headline release v1.6.1. The board revision is v1.7.21. The capability card for imajev-4b still prints 61.9, Intelligence 34.6, Calibration 89.2, estimated $0.017 per 1,000 decisions, and p50 0.29 s. The rank line says capability rank 31 and official composite 47. The v1.7.21 note says that revision left headline scores, Capability, Composite, and every rank unchanged. The rank move against the October 7 reading comes from rows added after v1.7.12. We did not rerun the suite.
Opened the same day, Image JevBench headlines JevImageBench v0.3.0. Imajev-4B leads the Jev-class systems at capability 68.0, rank 1, official composite rank 2. The composite table lists 52 ranked systems, and Imajev-4B is 63.7, with Intelligence 50, Calibration 86, Speed 86, Cost 52, and $0.038 per 1,000. The capability card prints Intelligence 49.9, Calibration 86.1, cost score 51.8, $0.038 per 1,000, and p50 0.37 s. Wity-1 reasoning auto is composite rank 1 at 71.2. The page still prints a 50-system table with imajev-4b at 76.4 and rank 2. A note on the page says scores from v0.1.5 are not ranked on the v0.3 scale. We did not rerun the image board.
This site's reading
Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.
Verify
mohit67890/imajev, Apache-2.0. The 4B is a LoRA on Qwen3.5-4B with a 256-code readout. The vision tower is frozen. A note dated 2026-09-26 says the phase-3 adapter moved ImajevBench from 82.4% to 83.9%, the hidden split from 84.2% to 85.6%, JevBench public hard as shipped from 70.3% to 72.1%, and DecisionBench from 77.5% to 79.7%.
Unknown abstentions on ImajevBench went from 14 of 21 to 18 of 21. Abstentions on 258 answerable items went from 3 to 9. DecisionBench ECE went from 0.024 to 0.069. The JevBench changelog dated 2026-09-27 adds Imajev-4B at rank 1, score 67.37, on the unchanged v1.4.2 scorer.
A 28 September note says the cost basis was corrected to one pinned server pass and that no score, axis, rank, or numeric value changed. The live capability view we opened the same day lists Imajev-4B at capability 66.3, Intelligence 52.2, Calibration 80.4, an estimated $0.022 per 1,000, p50 0.23 s, and official rank 1.
Jev 1.13.0 on that view is capability 64.7, Intelligence 53.1, Calibration 76.3, $0.040, p50 0.65 s, official rank 4. The README's screenshot of the composite, captured 28 September, prints 67.4, Plumb-4B 65.8, decider-4b v2 64.1, and Jev 63.3.
Its DecisionBench screenshot prints imajev-4b 79.65, third of 56, jev-1.13 71.90, and gpt-5.6-luna 69.04. We did not load the live Gradio board. ImajevBench is 279 questions, 21 of them unknown.
At a 90% bar the 4B automates 58% and is right on 97.5% of those. Scenario checks without the calibration file are 129 of 145. With the file, 113 of 145. The README says the top answer does not change.
We did not rerun it.
The page we opened September 30 is v1.5.4. It numbers the same 70.4 as official rank 9 and still tags the row roster addendum A2.
The same URL on October 4 says v1.5.6 and numbers 110 ranked systems of 116. imajev-4b is official rank 10 at 70.4, still tagged roster addendum A2. Axes are 53.5, 88.1, 91.1, and 63.3. Estimated cost is $0.017 per 1,000 decisions. Cygnet, Winnow-12B Q8, and Jev stay 73.7, 73.2, and 72.1.
Image JevBench v0.1.5, opened the same day, lists imajev-4b at 76.4, official rank 2 of 50. The compare table prints 76.39. Wity-1 is rank 1 at 80.2, and that row says the score may change after verification. The author's September 28 screenshot of v0.1.3 prints 76.39 as rank 1 of 49. We did not rerun either board.
On October 7, 2026 the same ranking URL says headline release v1.6.1 and board revision v1.7.12. The open-weights capability list puts imajev-4b at rank 17 and 61.9, Intelligence 34.6, Calibration 89.2, estimated $0.017 per 1,000 decisions, p50 0.29 s. The rank line says official composite 26. We did not rerun v1.6.1.
Compare
Plumb-4B's official composite in the same changelog is 65.84, and its own public-hard runner prints 89 of 111. decider-4b v2 remains 64.13. Jev's 63.29 is the same cell as v1.4.2. Imajev's 72.1% public-hard line is the author's as-shipped protocol, four option orders plus a calibration file.
It is not the 74.1% hard cell on the older Jev row. Jared Palmer's Kev and Sandipan Kundu's 400M encoder are different models that also use the name Kev.
On v1.5.4 the same 70.4 is official rank 9, tagged roster addendum A2. Decision 4B v1.1 is also 70.4, at rank 10.
On the v1.5.6 page opened October 4, imajev-4b is official rank 10 at 70.4, still tagged roster addendum A2. Image JevBench v0.1.5 lists it at 76.4, rank 2 of 50, behind Wity-1 at 80.2.
On the v1.6.1 open-weights list opened October 7, the same checkpoint is capability rank 17 at 61.9, official composite rank 26. 70.4 stays the v1.5.6 point estimate.
Opened October 8, revision v1.7.21, that card is still 61.9, capability rank 31, official composite 47. Image JevBench v0.3.0 lists capability 68.0 at rank 1 and composite 63.7 at rank 2 of 52. 76.4 stays the v0.1.5 figure.
Terms
- 67.37
- Imajev-4B on the JevBench v1.4.2.2 composite, changelog date September 27. Garg's September 28 screenshot rounds the neighboring rows to 65.8, 64.1, and 63.3. Jev's unrounded cell remains 63.29.
- 83.9%
- imajev-4b on ImajevBench after the September 26 phase-3 adapter. The previous release was 82.4%. The set is 279 questions, including 21 whose labeled answer is unknown.
- 18 of 21
- ImajevBench unknown items the 4B abstains on after phase 3. The previous release abstained on 14 of 21. It also abstains on 9 of 258 answerable items, up from 3.
- 70.4
- imajev-4b on JevBench v1.5.6, the page opened October 4. Official rank 10, roster addendum A2. The same point estimate was rank 9 on the v1.5.4 page opened September 30, and would-place sixth on the v1.5.0 page opened September 29.
- 61.9
- imajev-4b capability on the JevBench v1.6.1 open-weights list. Opened October 8, 2026, revision v1.7.21, the rank line says capability rank 31 and official composite 47. Intelligence 34.6, Calibration 89.2, estimated $0.017 per 1,000 decisions, p50 0.29 s. Opened October 7, revision v1.7.12, the same score was capability rank 17 and official composite 26.
- 76.4
- imajev-4b on Image JevBench v0.1.5, the page opened October 4. Official rank 2 of 50. The compare table prints 76.39. The author's September 28 screenshot of v0.1.3 prints 76.39 as rank 1 of 49.