imajev/ case study
Case study · Mohit Garg · September 2026

Decisions for real-world cases: an open 4B model, #1 on JevBench for text and images

imajev-4b · JevBench v1.4.2.2, scored 27 Sep 2026 · Image JevBench v0.1.3, released 28 Sep 2026 · run by the benchmark's maintainer

I built a system that makes the small, repeated decisions a team makes all day: does this photo match the listing, is this the item we shipped, which queue owns this ticket. It answers in the options you set, acts when it is sure, and hands the rest to a person.

#1 of 91
JevBench Score, ahead of Jev 1.13.0, the closed original it followsv1.4.2.2 · scored 27 Sep 2026 · board
#1 of 49
Image JevBench: decisions from photos, ahead of Jev-Omni (12B)v0.1.3 · released 28 Sep 2026 · board
#3 of 56
DecisionBench: ahead of GPT-5.6 Luna, DeepSeek V4.1 Flash (552B) and GLM-5.3 Flash (320B)eng, v1 · 28 Sep 2026 · leaderboard
01 · The problem

Thousands of small checks, each one a person's time

Operations, support, QC and back-office teams answer the same narrow questions all day. Each takes a person seconds to a minute; added up, they fill whole shifts.

  • Does the photo match the listing?marketplace and D2C catalogue
  • Is the returned item the one we shipped?returns and refunds
  • Is this part chipped?production-line QC
  • Which queue does this ticket go to, and is it urgent?customer support
  • Does this refund meet the policy?support and finance
  • Does this email contradict the CRM record?back office and records

A chat model can answer these, but not in a form a workflow can use. It replies in prose that has to be parsed before it can go in an if. The frontier vision APIs we measured took 5 to 8 seconds per answer, every call is billed, and customer photos leave the building. And unless it is prompted to, a chat model does not say "I can't tell", so there is no clean point at which to hand a case to a person.

A general chat model

  • Free-text answer to parse
  • Seconds per call, a bill per call
  • Photos go to a hosted API
  • Guesses when the evidence is thin

What a team actually needs

  • One of the answers you allowed, with a probability
  • Fast enough to sit inside a queue
  • Runs where the data already is
  • Says can't tell, so a person can take it
02 · What I built

A decision system that knows when to hand over

imajev is a family of small open models (2B, 4B and 9B parameters) trained for one job: read the evidence a business already has and answer a closed question with a calibrated probability.

  1. In: the evidence and a closed question

    A photo, or two (what we shipped and what came back), the record you already have, and a question with the answers you allow: which field is wrong? same item, yes or no? which department?

  2. Out: a probability for each answer

    Every allowed answer gets a probability, plus one for can't tell. No prose to parse; the result drops straight into an if.

  3. Act when sure, route the rest

    The business sets a confidence bar per decision, based on what a mistake costs. Above it the system acts; below it, or when it says can't tell, the case goes to a person.

Runs on your own hardware: a Mac or one GPUNo per-call API billPhotos and customer data stay in-houseOpen weights, Apache-2.0

Marketplace listing check: the listing says red, the photo shows beige suede boat shoes; imajev-4b names listing.color at 0.960 and the app holds the listing
Listing check. The listing says red; the photo shows beige suede. listing.color 0.960 → hold the listing until the seller fixes it.
Return check: pink and black knit sneakers were shipped, beige boat shoes came back; same product? no 0.986; the app holds the refund
Return check. Pink knit sneakers shipped, beige boat shoes returned. same item: no 0.986 → hold the refund.
Production-line QC: a known-good part next to the part under inspection; chipped or broken edge 0.981; the app rejects the part
Production-line QC. A known-good reference against the part on the line. chipped or broken edge 0.981 → reject the part.
Support ticket triage, text only: a duplicate-billing complaint; billing 0.941; the app routes it to billing
Ticket triage, text only. A duplicate-billing complaint. billing 0.941 → route to billing.

These are real outputs of imajev-4b with its shipped calibration, from the checked example apps on the project site. Every clickable combination in those apps was run and compared with the right answer; the misses are published too. One worth knowing: with a blank payment note, the model answered "not paid" instead of can't tell. That is why a pilot on your own cases comes before any automation.

03 · Measured results

Independent boards, run by their maintainers

The three rankings below were measured by the benchmarks' own maintainers, not by me. Each links to its source. Rankings move as new systems are added, so each is quoted with its version and date.

#1 of 91JevBench Score
v1.4.2.2 · scored 27 Sep 2026

JevBench ranks decision models of the Jev class. Its score weighs four things equally: accuracy (Intelligence), whether the stated confidence can be trusted (Calibration), speed and cost. imajev-4b leads Jev 1.13.0, the closed commercial model whose interface it follows, by 4.08 points. It is #3 on accuracy alone; the lead comes from balance, and it has the best calibration of the top 8 (80.4).

#SystemScoreIntell.Calib.SpeedCost
1imajev-4bopen weights67.3752.280.490.659.7
2Plumb-4B65.8453.075.593.555.8
3decider-4b v264.1349.475.092.960.9
4Jev 1.13.0TypeSafe AI, closed63.2953.176.383.352.0
5JevK5 v0.2.062.0448.974.591.159.5
JevBench v1.4.2.2 board screenshot: 1 Imajev-4B 67.4

Top 5 of 91 (Speed and Cost columns on wider screens). Board: benchmarkheaven.com/jev-models (maintainer Florian Standhartinger) · data: jevbench-v1.4.2.2-results.json. Measured with one option order and the shipped calibration file. The board estimates cost at $0.022 per 1,000 decisions (Jev $0.040), from the base model's public list price.

#1 of 49Image JevBench v0.1.3
released 28 Sep 2026

The same board's image benchmark: 684 decisions made from photos, screenshots and documents. imajev-4b scores 76.39, ahead of Jev-Omni, a 12B model (73.10), and NeoHorse Jev 4B (71.94). Its accuracy on the sealed items is 83.3%, second of the 44 systems that run on your own hardware.

Image JevBench v0.1.3 board screenshot: 1 Imajev-4B 76.39

Board: benchmarkheaven.com/image-jev-bench (screenshot 29 Sep 2026). Measured on the fast server setting; the board notes that this version's fresh sealed items are its own synthetic images.

#3 of 56DecisionBench (eng, v1), Hanno-Labs
leaderboard, 28 Sep 2026

79.65 across 23 tasks and 23,900 rows, every row answered. It is ahead of much larger general models: GLM-5.3 Flash (320B, 73.41), DeepSeek V4.1 Flash (552B, 70.68) and GPT-5.6 Luna (69.04), and of Jev 1.13 (71.90). The only two models ahead are the benchmark team's own, so imajev-4b is the highest-scoring model not built by the benchmark's team. On the separate Reasoning track it is #3 too (80.58), behind GLM-5.3 Flash and GPT-5.6 Luna.

DecisionBench (eng, v1) leaderboard screenshot: 3 imajev-4b 79.65

Leaderboard: Hanno-Labs/decision-bench-leaderboard (screenshot 28 Sep 2026) · registry: decision-bench-results (results PR #68).

58%decided automatically at a 90% bar,
97.5% of them right

The confidence bar is the dial a business controls. Raise it and fewer cases are automated, with fewer mistakes; everything below it goes to a person, including every can't tell. imajev-4b on the 279 test questions of ImajevBench:

Act when at leastBarDecided automaticallyAutoOf those, rightRightTo a personPerson
80% sure63%94.9%37%
90% sure58%97.5%42%
99% sure40%100%60%

ImajevBench is the project's own benchmark (photos, records and text; 21 questions whose honest answer is can't tell), built to be hard. Treat these as a starting point: the right bar for your team comes from a few hundred of your own cases. Source: README, "Automate what is clear, route the rest".

~$676rented GPU time,
the whole project

About a million training decisions in four stages, on open base models, every run included. Small models and a careful data pipeline, not a large compute budget, produced these results, and the same pipeline works on your own data.

Source: README, "Technical specification" (Compute).

04 · How I work

What I can build for your team

Each step stands on its own and ends with something you can use, so you can stop after any of them.

  1. Decision audit

    Pick one workflow. We list the decisions in it, how often each happens, what a mistake costs, and what evidence (photos, records, text) is already on hand.

    the decisions worth automating, written as closed questions, with a target bar for each

  2. Pilot on your own data

    A few hundred of your real cases, labelled with your team, including the ones nobody can decide. The model runs on them on your hardware.

    a measured accuracy-vs-automation curve on your cases and a recommended bar

  3. Tuning, if the pilot needs it

    Where accuracy falls short, I train a small adapter on your decisions and fit its calibration, with the same recipe used for imajev.

    a model tuned to your cases, re-measured on held-out ones

  4. Production hand-off

    Served on your own machines and wired into your queue or tools: the routing rule, a review queue for hand-offs, and tracking of the automated share and spot-checked accuracy.

    a running system, documentation, and a team that knows how to adjust the bar

Good fit

High-volume, well-defined decisions with a clear set of answers: listings, returns, QC, ticket routing, refunds against a policy, records against documents. E-commerce and D2C operations, support, QC and back-office teams.

Not a fit

One-off strategic calls, open-ended writing, or anything that needs a safety, medical, legal or hiring certificate. The system answers the questions you set; it does not replace the judgement behind them.

05 · Next step

Have a workflow like this?

Tell me which decision your team makes most often and roughly how many times a week. I will tell you whether it is a good candidate and what a pilot would measure.