Back to home

Compare

Comparing: 50 Records In, 20 Seen by Model: Agent Silent Truncation Is the Real Killer & 数据 50 条进 AI,模型只看到 20 条:Agent 不报错的截断才致命

AEN
AgentMCPLLM·

50 Records In, 20 Seen by Model: Agent Silent Truncation Is the Real Killer

What This Is

A source database held 50 records; only 20 actually reached the large language model. Developers recently debugged an Agent (an AI assistant that autonomously invokes tools to complete tasks) case that looked successful but had failed: the remaining 30 records vanished without any error, and the model produced a logically coherent but factually incomplete summary based on truncated input.

The trouble is that traditional ops only watches for errors, but Agent data truncation happens outside the error-reporting layer. Text formatting is normal, the model shows no hallucinations, system logs are green—but factual coverage is already incomplete.

The author's fix: before sending data to the model, the system explicitly logs several fields—how many records were requested, how many were actually loaded, whether truncation occurred from row caps or byte limits, and how many were filtered out. Once inputComplete=false, downstream answers can no longer pretend to be based on the full set.

Further, "completeness" splits into three layers: whether the source set is defined completely, whether the target data fully enters the model, and whether the final answer correctly covers the data—these three cannot be conflated. Counting records alone isn't enough—if the source is A B C D E and the middle layer receives A B C D D, the count is still 5, but E is already missing; only ID-set reconciliation or hash comparison can catch this.

Industry View

The mainstream tech reaction is "this isn't new"—anyone who has built data pipelines has seen silent data loss; it's essentially the same as a paginated API returning only the first 100 records. Applied to Agents, it's just a new shell.

But another voice deserves more of our attention: Agents are shifting from "toys" to "business systems." When AI handles customer lists, order ledgers, and compliance documents for enterprises, "the process finished" does not equal "the job was done right." Traditional ops watches error rates and latency; the Agent era needs a new metric—"input completeness rate."

The risk: most enterprises currently don't have this monitoring dimension at all. When a model gives wrong answers, it's easy to spot. When a model sounds "plausible but is based on incomplete data," the problem often surfaces only after business outcomes go wrong.

Impact on Regular People

For enterprise IT: shift monitoring from "system errors" to "data landing points." Critical paths should at minimum log three numbers—how many records at source, how many received by the model, how many covered in the answer.

For individual professionals: when using AI assistants for tasks like client lists, financial summaries, or meeting notes, be wary of conclusions that "sound smooth." When necessary, ask the AI to list original record counts for cross-verification.

For consumer markets: AI customer service, AI investment advisors, AI health-check reports, and similar products may be making recommendations based on incomplete data. Maintaining a healthy suspicion—"it hasn't necessarily seen everything"—is safer than trusting a polished answer.

Source: juejin.cn
BZH
AgentMCP大模型·

数据 50 条进 AI,模型只看到 20 条:Agent 不报错的截断才致命

这是什么

源头数据库有 50 条记录,真正进入大模型的只有 20 条——开发者最近排查了一个「看着成功实际失败」的 Agent(能自主调用工具完成任务的 AI 助手)案例:剩下 30 条没有任何报错地消失了,模型却基于残缺输入给出了逻辑通顺但事实不全的总结。 麻烦在于,传统运维只看「有没有报错」,而 Agent 的数据截断发生在报错体系之外。文本格式正常,模型也没出现幻觉,系统日志一片绿色——但「事实覆盖已经不完整」。 作者的修复思路是:在送入模型前,由系统显式记录几个字段——本次请求多少条、实际加载多少条、是否被条目上限或字节限制截断、过滤掉了多少。一旦 inputComplete=false,后面的回答就不能继续假装基于全集。 更进一步,「完整」被拆成三层:源集合是否定义完整、目标数据是否完整进入模型、最终回答是否正确覆盖数据,三者不能混为一谈。光看数量还不够——如果源端是 A B C D E,中间层拿到 A B C D D,数量仍是 5,但 E 已经丢了,需要用 ID 集合或哈希对账才能发现。

行业怎么看

技术社区的主流反应是「这不新鲜」——任何做过数据管道的人都见过静默丢数据,本质上和分页接口只返回前 100 条是一回事。套到 Agent 上,只是换了个壳。 但另一种声音更值得我们注意:Agent 正在从「玩具」变成「业务系统」。当 AI 替企业处理客户名单、订单流水、合规文档时,「流程跑完了」不等于「事情做对了」。传统运维盯的是错误率和延迟,Agent 时代还需要一个新的指标——「输入完整率」。 风险在于,多数企业目前根本没有这个监控维度。模型答错了容易被发现,模型答得「看起来对但基于残缺数据」,往往要等到业务结果出问题才暴露。

对普通人的影响

对企业 IT:需要从「系统报错」转向「数据落点」监控,关键链路至少要记录「源头几条、模型收到几条、回答覆盖几条」三个数。 对个人职场:用 AI 助手处理客户清单、财报摘要、会议纪要这类任务时,遇到「看着很顺」的结论要警惕,必要时要求 AI 列出原始条数做交叉验证。 对消费市场:AI 客服、AI 投顾、AI 体检报告等产品可能正在基于不完整数据做推荐,保持一份「它不一定看全了」的怀疑,比相信一个漂亮答案更安全。
Source: juejin.cn