We noticed a brand-new account on Xiaohongshu recently published six Chinese-style healing MVs made entirely with AI, from Nian Wei Xie to You Feng Wu Feng Jie Ziyou. After six videos, the workflow was fully validated—all through "learning by doing." The angle worth caring about: once AI has pushed the technical floor low enough, what decides output quality is no longer whether you can use the tool, but taste and judgment.
What this is
The creator had no formal training and didn't start by grinding through tutorials. His workflow: Doubao breaks down lyrics into shot-list prompts → Jimeng generates images → Jimeng's image-to-video → CapCut assembles and edits. The toolset is unglamorous—domestic Chinese AI alone carries the full pipeline.
The key iteration hit on the fourth video. For the first three, he went straight to "text-to-video" (generating video directly from a text description). Abstract text made the visuals nearly impossible to predict, and the repeated "gacha pulls" (regenerating over and over) burned both money and time. On video four, he switched to a two-step "text-to-image + image-to-video" approach—lock down the keyframes first, fix anything off at the image stage, then feed the stills into video generation. Efficiency jumped sharply; the later videos mostly landed in one pass.
Another lesson was style consistency. He had Doubao lock in a Morandi low-saturation palette and a consistent character setup of a Q-version (chibi-style) little girl plus a kitten, then produced one character sheet (a three-view turnaround: front, side, and back reference) as the standard. From the fifth video on, image-to-video nailed the look on the first try.
Industry view
Feedback from the AIGC (AI-Generated Content) creator community consistently validates "learning by doing." Multiple experienced practitioners note that AI tools iterate so fast that any tutorial you finish is likely already obsolete—adjusting on the fly while building is actually faster.
But there are pushbacks worth heeding. Some in the industry caution: this "build first, learn later" approach is fine for personal projects but a different story for enterprise applications. Production environments demand deliverable stability, compliance, and clean copyright lineage (a clear, traceable, licensable chain of source assets). Pure trial-and-error is a minefield. An MV is a personal work—you can fix mistakes. Commercial content delivery has completely different rework costs. Separately, "text-to-image + image-to-video" is more controllable than pure text-to-video, but current AI video still struggles with camera-movement logic and character consistency, especially on longer clips. The creator himself acknowledged this in his own retrospective.
Impact on regular people
For enterprise IT: A solo operator plus AI tools can now produce content that once required a professional team. Enterprise budgets for content production and headcount-productivity assumptions both need to be reworked.
For individual careers: People in non-creative roles can now produce decent visual work with AI—skill boundaries are widening. But taste and judgment are getting more valuable, not less.
For the consumer market: AI-generated polished content will keep flooding Xiaohongshu and Douyin. The cost for users to tell what's "original" or "authentic" is rising. Platforms and creators alike face a new trust problem.