An AI response typically takes a dozen-plus seconds to produce, but the "typewriter-style" output users see hides a frequently overlooked problem: text scrolling past on screen doesn't mean the answer is complete. This week, a developer tutorial re-dissected it — streaming (SSE, server-sent events) makes responses "look fast," but the cost is that the frontend must distinguish "displayed" from "generation-complete."
What This Is
A normal HTTP request is like mailing a letter: the server finishes the entire letter before sending it back, and users stare at a blank screen for the dozen-plus seconds in between. Streaming turns "waiting" into "in progress" — and that's the core reason ChatGPT, Doubao, and Wenxin Yiyan feel "fine enough" to use. This week's tutorial on Juejin re-demonstrates the underlying mechanism using Python's standard library: loop over the network stream, print each chunk as it arrives, and only treat the response as done when the "finished" event lands.
But the author deliberately separates two things — "content has appeared on screen" and "the answer is reliably complete." If the network drops, the user refreshes, or the model errors mid-stream, the first still holds; the second does not. It's a small engineering detail that we keep seeing ignored at the product level.
Industry View
The case for streaming is hard to argue with: if responses had to wait a dozen-plus seconds to come out as a whole block, users seeing no cursor movement would click again — and the extra requests would crowd out inference resources, making everyone slower. Streaming word-by-word is, in essence, disguising "waiting" as "in progress."
But there is pushback. Streaming magnifies the illusion of authority: a half-finished sentence scrolls across the screen and users tend to treat it as a complete answer. Once the network drops, the page refreshes, or the model errors mid-generation, "displayed" content becomes "already trusted" content. The article's author specifically wrote handling for disconnections and error events — to remind us that engineering teams are responsible for distinguishing the two, rather than assuming a finished frontend render equals a complete answer.
Impact on Regular People
For enterprise IT: if your company is building an internal AI assistant or customer-service bot, don't let the frontend just show "generating." Streaming output must be paired with a notice that the answer may be incomplete — otherwise frontline employees will treat half-sentences as official conclusions.
For individual careers: before pasting AI output into weekly reports, contracts, or client emails, wait for full generation. Words scrolling past on screen aren't the same as words you've actually verified.
For the consumer market: nearly all upcoming Chinese AI products — phone assistants, smart speakers, in-car systems — will almost certainly move to streaming. We all need to build a new habit: seeing AI typing doesn't mean it's "thinking." It might just mean the network is still alive.