返回首页

对比阅读

对比阅读:Larger LLMs Confidently Hallucinate: 120B Test Breaks 'Bigger Is Better' Myth 与 大模型越大越靠谱?120B 测试给出反常识结论:越大越会'自信地胡说'

AEN
OllamaLocal LLMHallucination Detection·

Larger LLMs Confidently Hallucinate: 120B Test Breaks 'Bigger Is Better' Myth

What This Is

A team from Reddit's r/LocalLLaMA community tested open-source models ranging from 1.5B to 120B parameters (B stands for billion parameters — more parameters theoretically means more capable). They found that 120B models output exactly the same wrong answer when "confidently hallucinating" — breaking the conventional wisdom that "bigger models are more reliable."

The problem they tackled is highly practical: when running LLMs locally with tools like Ollama, how do you know when the AI is confidently making things up (the industry term is "hallucination")? The traditional academic method is "Semantic Entropy" — sample the model's answer 5 times, use an extra small neural network like DeBERTa to judge whether the answers mean the same thing, then calculate the disorder. But running this on local hardware is costly: it eats VRAM, adds over 100 milliseconds of latency, and tanks throughput.

Their alternative is disarmingly simple: just use string matching plus Shannon entropy (a mathematical way of measuring "uncertainty"), running purely on CPU, completing in 1.3 microseconds without touching the GPU.

Industry View

They ran controlled experiments on Qwen and Mistral families across GSM8K (math problems) and TriviaQA (general knowledge Q&A). The results split into four tiers:

  • 1.5B small models: AUROC (a classifier effectiveness metric, 1.0 is perfect) around 0.58. The models are too small — output formats aren't even consistent, so matching fails.
  • 7B mid-size: AUROC around 0.71, unexpectedly tying the heavyweight DeBERTa model (0.706 vs 0.705).
  • 27B large: AUROC around 0.89, best in class.
  • 120B flagship: AUROC crashes to 0.09, effectively broken.

The last finding is the study's most explosive — they call it the "Frontier Trap": when a 120B model hallucinates, five independent samples produce identical wrong answers. "Confident hallucinations" stay highly consistent, and even traditional detection methods can't save you.

Pushback is worth noting: tests only ran on TriviaQA and GSM8K, so coverage is narrow. Some in the community worry the finding will be misread as "smaller models are safer" — but in reality, smaller models are just easier to catch hallucinating, not less prone to hallucinating. These are two completely different things.

Impact on Regular People

For enterprise IT: when deploying open-source models locally for structured tasks (math, code, SQL extraction, JSON parsing), 27B is the cost-effectiveness sweet spot. Going beyond that requires building a dedicated "hallucination consistency" defense layer — we can no longer assume bigger means more stable.

For working professionals: when we use AI for weekly reports, data lookups, or research, beware — if the AI gives the same answer multiple times in a row, that doesn't mean it's true. It might be a "confident hallucination." For critical decisions, switch models or rephrase the question and verify again.

For the consumer market: local AI tools are getting cheaper (open-source models plus Ollama-style tooling are going mainstream), but "bigger and pricier" no longer equals "bigger and more reliable." Future AI procurement will shift from a "parameter arms race" toward "task matching plus hallucination defense mechanisms."

BZH
Ollama本地大模型幻觉检测·

大模型越大越靠谱?120B 测试给出反常识结论:越大越会'自信地胡说'

这是什么

Reddit r/LocalLLaMA 社区一个团队测试了 1.5B 到 120B 的开源模型(B 指十亿参数,参数越多模型理论上能力越强),发现 120B 模型在'自信胡说'时输出完全相同的错误答案 —— 这打破了'模型越大越靠谱'的常识。

他们研究的问题很实际:本地用 Ollama 等工具跑大模型时,怎么知道 AI 在一本正经胡说八道(行业术语叫'幻觉' hallucination)?传统学术方法是'语义熵'(Semantic Entropy)—— 让模型回答 5 次,用 DeBERTa 这种额外的小神经网络判断答案意思是否一致,再算混乱度。但本地硬件跑这个代价高:占显存、慢 100 毫秒以上、吞吐量崩。

他们的替代方案很朴素:直接用字符串匹配 + Shannon 信息熵(一种衡量'不确定性'的数学方法),纯 CPU 跑,1.3 微秒完成,不吃 GPU。

行业怎么看

他们拿 Qwen、Mistral 系列在 GSM8K(数学题)和 TriviaQA(百科问答)上跑了对照实验,结果分四档:

  • 1.5B 小模型:AUROC(分类器效果指标,1.0 满分)约 0.58。模型太小,格式都不统一,匹配不上。
  • 7B 中型:AUROC 约 0.71,意外追平重型 DeBERTa 模型(0.706 vs 0.705)。
  • 27B 大模型:AUROC 约 0.89,效果最好。
  • 120B 顶级模型:AUROC 跌到 0.09,几乎完全失灵。

最后一条是研究最炸的发现 —— 他们叫它'前沿陷阱'(Frontier Trap):120B 模型幻觉时,5 次独立采样给出完全相同的错误答案,'自信地胡说'还高度一致,连传统检测方法也救不了。

反对意见值得说:测试只在 TriviaQA 和 GSM8K 两个数据集上跑,覆盖场景有限;社区也有人担心这个结论会被误读成'小模型更安全'——实际上小模型只是更容易被发现胡说,不是更不容易胡说。这是两个完全不同的事。

对普通人的影响

对企业 IT:本地部署开源模型做结构化任务(数学、代码、SQL 抽取、JSON 解析),27B 是性价比甜点;超过这个规模,需要专门搭'幻觉一致性'防线,不能默认越大越稳。

对个人职场:用 AI 写周报、做数据查询、查资料时注意 —— 如果 AI 连续几次给出相同答案,不一定是真的,可能是'自信的胡说'。关键决策换个模型或换个问法再验一次。

对消费市场:本地 AI 工具在变便宜(开源模型 + Ollama 类工具普及),但'越大越贵'不等于'越大越靠谱'——未来 AI 选型会从'参数竞赛'转向'任务匹配 + 防幻觉机制'。