What this is
This week on Reddit's LocalLLaMA subreddit, an independent developer posted progress on post-training Yandex's AliceAI-80B-A3B model. He used 3 V100 GPUs—32GB VRAM each, 96GB total—to instruction fine-tune an 80-billion-parameter base model (i.e., tuning a general-purpose model into one that understands human instructions and follows commands).
His method: deploy Qwen 3.8 27B locally as a 'teacher' to generate training data, then distill (i.e., have the target model learn the larger model's outputs) into the 80B model. After the first run with ~5 million tokens, his own testing conclusion: 'technically it runs, but it's basically useless.' Reason: the data was concentrated on coding; the model outputs garbled text on ambiguous queries.
He decided not to release this checkpoint, planning to generate another 5 million tokens of broader instruction data and retrain at a lower learning rate. He also open-sourced his self-developed data generation tool, sftmill.
Industry view
We see two sides to this story.
The optimistic side: post-training 80B-scale models used to require a full data center's compute; today, one person with 3 second-hand GPUs can run it. We see the open-source ecosystem rapidly closing the 'can-build' gap with big tech. This is a typical signal of last-generation big-tech-exclusive capabilities starting to spill into the community.
The cool—even pessimistic—side: he himself called his first run 'basically useless.' This tells us that between 'pipeline executes' and 'ship a production-grade model' there are still multiple gates of data quality, training strategy, and evaluation methodology. One developer's effort and goodwill doesn't guarantee outcomes. We see this disparity sting harder in business scenarios—many companies adopt open-source models only to find 'looks like it can chat, but goes off the rails on actual business,' which is fundamentally the same problem: treating demo as capability.
Impact on regular people
For enterprise IT: the cost structure of post-training open-source models is changing, but don't be misled by 'it runs.' Before deployment, you must have an evaluation set tailored to your own business scenarios; external demos don't prove it can handle your real data.
For working professionals: AI tools' real capability boundary is narrower than vendors advertise. They fail most easily on ambiguous, colloquial, or rare-in-training-data queries—exactly the same class of problem that the developer's first version exposed through insufficient training data.
For the consumer market: this story won't affect what you buy short-term, but it signals the barrier to customized AI is falling; in the future, the cost for SMBs to build 'dedicated AI assistants' will be much lower than today—we recommend keeping an eye on this trend.