A 4B-parameter model (about 4 billion parameters), built for $1,200 in 15 days, has produced a decision AI that beats GPT-5.6 Luna. What we find most worth paying attention to is what lies beyond the scores: it got right something general LLMs have never quite managed — knowing when to say "I don't know."
What this is
Developer Mohit is a process consultant who has served Fortune 500 companies. He fine-tuned ImaJev-4b on Qwen3.5-4B (an open-source lightweight model from the Qwen family): it takes text records or up to two photos as input, outputs decision probabilities such as "refund" or "return," and is allowed to answer "I don't know." It ranks first among 91 models on JevBench and third on DecisionBench, surpassing GPT-5.6 Luna and DeepSeek V4.1. Cost: about $1,200 in GPU rental. Weights are open-sourced under Apache-2.0.
The author acknowledges that pure accuracy only ranks third; the find real strength is "calibration" (confidence matches actual accuracy): when the model says it's 90% sure, it really is 90% right. That makes it suited for finance, refunds, and customer service — decisions where error tolerance is low.
Industry view
The optimistic camp sees this as confirmation of an existing judgment: general LLMs are not good at vertical business decisions, because a good decision means "knowing what you don't know." Vertical small models — cheaper, with more deterministic answers — are the tools that are actually usable for enterprises.
The cautious camp raises three points. Both JevBench and DecisionBench are relatively new benchmarks with limited sample sizes, so a first-place finish doesn't necessarily translate to broad advantage. The author trained on items where "two AI answers agreed," effectively narrowing the training set — this could cause performance to regress on more complex scenarios. In fact, his first fine-tuning attempt dropped the model from 64.9 to 42.3 points. And "being able to say I don't know" is a double-edged sword for process automation: too many "unknown" responses push the work back to humans.
One often-overlooked context: the base model is Qwen3.5-4B, and Chinese open-source models are becoming the default starting point for overseas independent developers.
Impact on regular people
- For enterprise IT: claims review, returns and exchanges, and customer-service triage now have "good-enough and cheap" options. There's no longer a need to pay for general LLM APIs.
- For individual professionals: process-oriented and judgment-oriented roles (review, operations, junior analysis) will feel the change first — not by being replaced, but by being required to "work with an AI assistant."
- For the consumer market: little impact for now. These products require enterprise-grade integration and won't show up directly on consumer phones.