返回首页

对比阅读

对比阅读:AI Agents Run Like Blind Boxes — LangSmith Wants to Be Their ECG 与 AI Agent 跑起来像开盲盒 — LangSmith 想给它装心电监护仪

AEN
LangSmithLangChainAgent·

AI Agents Run Like Blind Boxes — LangSmith Wants to Be Their ECG

Developers running AI Agents have an old metaphor: like opening a blind box — you can't see which tool got called, how long each step took, or how much money the tokens burned (Agent: an AI program that autonomously calls tools to complete multi-step tasks). LangChain's LangSmith wants to solve this with "full-chain observability" (source: a recent deep-dive article on Juejin). What we're watching is the signal behind it: as AI moves from demo to enterprise deployment, "invisibility" is becoming the biggest cost black hole.

What this is

LangSmith is LangChain's AI observability platform, purpose-built for RAG (Retrieval-Augmented Generation — AI looks up information before answering) and Agents. Developers use it to "illuminate" every AI run: which tool each step called, how long each step took, where the tokens went. The article uses a sharp medical metaphor — real-time monitoring (Tracing/Monitoring) is the ECG, telling you "that last heartbeat was off"; batch evaluation (Datasets/Evaluators) is the annual physical, telling you "what's the overall score." You need both.

Industry view

Supporters argue: as AI moves from experiment to production, observability tools will become infrastructure like New Relic was. Andrew Ng has been repeating lately that "90% of Agent projects don't stall on technical issues at deployment" — and one root cause is exactly "invisibility."

But the dissent is sharp. First, observability ≠ accuracy — no matter how pretty the monitoring dashboard, if AI answers wrong, it answers wrong. Second, LangSmith is LangChain's own child — enterprises that adopt it risk vendor lock-in, and open-source alternatives Langfuse and Arize are taking market share. Third, for many traditional enterprises, the real problem isn't "how do I observe AI," but "do I even have qualified data for AI to answer from."

Impact on regular people

For enterprise IT: In the next 1–2 years, enterprise IT budgets will grow a new line item — "AI Observability," similar to the rise of APM (Application Performance Monitoring) back in the day.

For careers: Engineers who know LangSmith and Agent Ops tools will be in high demand; non-technical roles need to get used to the idea that "when AI gets it wrong, it can be traced and audited."

For consumers: Next time AI gives you a wrong or off-target answer, customer support will at least be able to tell you "which step went wrong" — giving you real grounds to push back.

来源: juejin.cn
BZH
LangSmithLangChainAgent·

AI Agent 跑起来像开盲盒 — LangSmith 想给它装心电监护仪

跑 AI Agent 的开发者有个老比喻:像开盲盒 — 看不见调了哪个工具、每步耗时多少、token 烧了多少钱(Agent:能自主调用工具完成多步任务的 AI 程序)。LangChain 的 LangSmith 想用"全链路观测"解决这件事(来源:掘金最近一篇深度文章)。我们关心的是背后信号:当 AI 从 Demo 走向企业部署,"看不见"这件事,正在变成最大的成本黑洞。

这是什么

LangSmith 是 LangChain 推出的 AI 观测平台,专为 RAG(检索增强生成,即 AI 先查资料再回答)和 Agent 设计。开发者用它把每一次 AI 运行"照亮":哪一步调了哪个工具、每步耗时、token 花在哪。文章用了一个精妙的医疗类比 — 实时监控(Tracing/Monitoring)是心电监护仪,告诉你"刚才那下心跳不对";批量评估(Datasets/Evaluators)是年度体检,告诉你"整体能打几分"。两者缺一不可。

行业怎么看

支持方认为:当 AI 从实验走向生产,观测工具会变成像 New Relic 那样的基础设施。Andrew Ng 近期反复提"90% 的 Agent 项目卡在落地不是技术",根因之一正是"看不见"。

但反对意见也很尖锐:第一,观测 ≠ 准确率,监控得再漂亮,AI 答错还是答错;第二,LangSmith 是 LangChain 亲儿子,企业用了容易绑定单一供应商,开源替代品 Langfuse、Arize 在抢市场;第三,对很多传统企业来说,真正的问题不是"怎么观测 AI",而是"我有没有合格数据让 AI 回答"。

对普通人的影响

对企业 IT:未来 1-2 年,企业 IT 预算里会多出"AI 可观测性"这一栏,类似当年 APM(应用性能监控)的崛起。

对个人职场:懂 LangSmith、Agent Ops 这类工具的工程师会很抢手;非技术岗要习惯"AI 出错是可以被追溯和审计的"这件事。

对消费市场:未来遇到 AI 答错、答偏,客服至少能告诉你"是哪一步出了问题",维权有依据。

来源: juejin.cn