What This Is
We noticed a public benchmark: LoCoMo's 300 turns across 35 sessions—the hard threshold for measuring AI memory, and the reason most commercial agents still have goldfish brains. The solution in Chapter 3 of "Deep Understanding AI Agents," open-sourced by Chinese engineers on GitHub, is this: AI memory isn't a recording. When a conversation ends, you call the LLM again to extract what's worth keeping. It splits into two scales—"user memory" remembers who you are, "knowledge base" remembers what the world is. Memory capability is stacked across three floors: basic recall (verbatim regurgitation), cross-session retrieval (piecing info across sessions), and proactive service (flagging issues without instructions).
Industry View
Supporters argue that treating memory systems as an engineering problem to decompose is a sign the industry is maturing.
But the dissent is just as sharp: calling an LLM again to "distill memory" creates new problems—high cost, slow speed, and distilled outputs may contradict the original context. More aggressive researchers argue memory should be built into the base model, not bolted on later—a direction with no consensus yet. We believe another overlooked risk is privacy: long-term memory means every user has a permanent file, and if the company gets sold or hacked, that file is you.
Impact on Regular People
For enterprise IT: when upgrading customer service and assistant systems, the real bottleneck is no longer model parameters but the engineering architecture of "retrieval → processing → re-invocation."
For individual professionals: when using AI to take meeting notes or write minutes, the "it forgot again" issue often lies in the engineering layer—switching tools may not solve it.
For the consumer market: smart assistants will increasingly "get you," but the price is that your preferences, weaknesses, and relationship networks get permanently archived—convenience and risk grow in the same direction.