What this is
A 2B (2 billion parameter) small model scored 83.1% across five classification benchmarks — 0.1 percentage point above commercial product Jev (launched by TypeSafe) at 83.0%. That tells us open-source small models have caught up to closed commercial products on "judgment" tasks.
The model is called Jeff, published this week by Reddit user firelex on r/LocalLLaMA. It's a fine-tune of Qwen3.5 (Alibaba's lightweight Tongyi Qianwen variant) and Gemma (Google's open-source small model), purpose-built for classification: feed it a context and a list of options, and it outputs a probability for each option without generating text. The 0.8B version runs at roughly 28ms per inference on an M4 Max chip — effectively real-time.
The training cost is the real story here. The 0.8B took 2 hours to train; the 2B took 3.5 hours. Both ran on a single RTX PRO 6000 (96GB VRAM workstation GPU). Synthetic training data was generated locally on two DGX Sparks. Released under Apache 2.0 — commercial use permitted.
Industry view
This is another concrete data point in the "small models catching up" narrative. Over the past year, Mistral, Alibaba, and Zhipu have all placed bets on the small-parameter route; the open-source community keeps demonstrating that sub-7B (sub-70 billion parameter) models can hit the floor that big models set on vertical tasks. firelex himself sits squarely in the current — he used Alibaba's updated Qwen3.8-Flash-Next to generate synthetic training data, essentially distilling a bigger model's capabilities into targeted optimization for the small one.
Two cautions, though. First, the multi-step reasoning gap is stark: Jeff scores only 64–68% on BBH (a complex reasoning benchmark), versus Jev's 94%. Small models excel at "judgment," not "thinking." Second, synthetic data has a ceiling. If the upstream generator carries bias, the downstream small model amplifies it rather than dilutes it. A company that ships Jeff as a "big-model replacement" will likely see document summarization and long-chain QA collapse into the low 60% range — directly degrading user experience.
Impact on regular people
For enterprise IT: it's worth re-evaluating the cost structure of AI use cases. One 96GB workstation GPU plus a few hours of training yields a judgment model at 80%+ accuracy — an order of magnitude cheaper than a commercial API subscription. But distinguish "judgment tasks" from "reasoning tasks" — the former can be cut, the latter still needs a big model.
For working professionals: the cost of standardized judgment tools (contract clause classification, email priority sorting, ticket routing) is dropping. Watch whether your company's IT department is piloting local small models; over the next one to two years, this kind of automation may become more pervasive.
For the consumer market: 28ms response latency puts on-device AI (running locally, without uploading to the cloud) within closer reach. Future phones or laptops may run these judgment models directly — privacy and latency improve in tandem.