This week, the InternLM (Shusheng) team open-sourced Intern-Decision 0.8B and 4B—two models that don't write text, only scoring "multiple-choice questions." We've noticed this kind of "do one thing" small model quietly becoming the workhorse of enterprise AI deployment.
What this is
It takes three inputs: a shared state, a set of questions with options, and optional images. In a single pass, it returns an answer probability distribution for each question, output as structured JSON. The entire process generates no free-form text; at its core, it's a scorer, not a chatbot.
The technical trick: each option is mapped to a single character (A, B…Z, a…z, 0…9), so the model only predicts probabilities over those characters. The 0.8B version runs on a single consumer-grade GPU—or even a laptop; the 4B version targets higher-accuracy needs.
Industry view
Supporters argue this is exactly what enterprise AI deployment needs—interpretable (outputs are probabilities, not text), auditable, and cost-controlled. Approval workflows, credit pre-screening, customer-service triage—these "decision-type" scenarios rarely need ChatGPT-style free-form generation, and often fear it.
But skepticism exists. First, the model is constrained to a "multiple-choice" framework—real business problems often lack preset options and demand open-ended reasoning. Second, the base is still Qwen3.5-4B; the ceiling on decision capability depends on fine-tuning, not the base itself. Third, open-source small models typically need accompanying engineering work to land in enterprises, and the InternLM team's commercialization path remains unclear.
There's an even sharper critique: forcing decision problems into a multiple-choice wrapper may simply be masking the model's underlying reasoning inadequacies.
Impact on regular people
For enterprise IT: Structured-judgment scenarios like credit pre-screening, compliance approval, and survey analysis can now run on 1-2 consumer-grade GPUs, significantly lowering the upfront investment bar.
For individual careers: In roles like customer-service dispatch, junior review, and rules-based analyst work, the "mechanical judgment" components will be replaced first—but work requiring trade-offs, negotiation, and complex reasoning is safe for now.
For consumer markets: Local deployment means sensitive data never leaves the building—clear good news for heavily regulated industries like healthcare, finance, and law.