返回首页

对比阅读

对比阅读:LangSmith Lets AI Apps Self-Check — LLM Deployment Enters the Observability Era 与 LangSmith 让 AI 应用学会自检 — 大模型落地进入「可观测」时代

AEN
LangSmithLangChainRAG·

LangSmith Lets AI Apps Self-Check — LLM Deployment Enters the Observability Era

We've noticed a shift: more and more enterprise AI Agent (AI programs that autonomously complete multi-step tasks) and RAG (technology that lets AI retrieve internal knowledge bases to answer questions) projects are past the demo stage — not stuck on model performance, but stuck on "we can't clearly tell whether it's actually running correctly." LangSmith being frequently discussed in Chinese tech communities this week is a signal of exactly this shift.

What this is

LangSmith is LangChain's AI application observability and evaluation platform. It does two things: first, it makes visible what happens inside a single AI request — which step called the model, which documents were retrieved, how long it took, how many tokens were spent, and where errors occurred; second, it judges whether the AI application is actually "good" — instead of relying on subjective human spot-checks, it runs batch regression tests against standard datasets and scoring metrics.

Simply put, it's the equivalent of "New Relic + an automated testing platform" for AI applications. Setup cost is low — three lines of code in an environment variable and you're connected.

Industry view

Optimistic voices see this as a sign of AI engineering maturity — moving from "it runs" to "observable, regressable, measurable," a path the industry must take.

But there are plenty of counterarguments. On one hand, LangSmith is tightly coupled to the LangChain ecosystem — enterprises not building on the LangChain framework see limited benefit. On the other, observability doesn't equal problem-solving — even after you see clearly that the AI is hallucinating, someone still has to fix it. Add cost concerns: every AI call generates a Trace (tracking record), and fees under high concurrency aren't trivial — enterprises need to run the numbers before scaling.

Impact on regular people

For enterprise IT: CTOs are starting to bring AI projects under formal operations and quality management systems, leaving behind the "ship it after the demo" state.

For individual careers: people who understand AI application observability and quality evaluation are forming a new specialization, similar to the SRE (Site Reliability Engineer) and QA roles in traditional software.

For consumer markets: users will eventually benefit indirectly — more stable AI customer service, fewer glitchy AI assistants — all fundamentally dependent on this kind of tooling maturing.

来源: juejin.cn
BZH
LangSmithLangChainRAG·

LangSmith 让 AI 应用学会自检 — 大模型落地进入「可观测」时代

我们注意到一个转变:越来越多企业的 AI Agent(能自主完成多步任务的 AI 程序)和 RAG(让 AI 检索内部知识库回答问题的技术)项目过了 demo 阶段,不是卡在模型效果上,而是卡在「说不清它到底跑得对不对」上。LangSmith 这周在中文技术社区被频繁讨论,正是这个转变的信号。

这是什么

LangSmith 是 LangChain 推出的 AI 应用观测与评估平台。它解决两件事:第一,看清一次 AI 请求内部发生了什么——哪一步调用了模型、检索了哪些文档、耗时多久、Token 花了多少、哪里报错;第二,判断这套 AI 应用到底「好不好」——不再靠人随便问几个问题主观判断,而是用标准数据集和评分指标做批量回归测试。

简单说,它相当于 AI 应用的「New Relic + 自动化测试平台」。配置成本很低,在环境变量里加三行代码就能接入。

行业怎么看

乐观的声音认为,这标志着 AI 工程的成熟——从「能跑就行」走向「可观测、可回归、可量化」,是行业必经之路。

但也有不少反对意见。一方面,LangSmith 与 LangChain 生态深度绑定,企业如果不基于 LangChain 框架开发,收益有限;另一方面,可观测不等于能解决问题——看清了 AI 在胡说八道,还是得有人能改。此外还有成本顾虑:每一次 AI 调用都会产生 Trace(追踪记录),高并发场景下的费用并不低,企业上规模前需要算清楚账。

对普通人的影响

对企业 IT:CTO 们开始要把 AI 项目纳入正式的运维和质量管理体系,告别「demo 完就放一边」的状态。

对个人职场:懂 AI 应用可观测和质量评估的人,正在形成一个新岗位细分需求,类似传统软件里的 SRE(站点可靠性工程师)和 QA。

对消费市场:用户最终会间接受益——更稳定的 AI 客服、更少抽风的智能助手,本质上都依赖这类工具的成熟。

来源: juejin.cn