A Russian developer admitted this week on r/LocalLLaMA: his attempt to fine-tune Yandex's open-source 80B model (continuing training on his own data) failed on the very first run. The cause was a self-written low-level module for V100 GPUs that produced NaN (numerical anomaly), wiping gradients (training update signals) on every layer except two. The loss curve (training health metric) "looked fine" — but it was an illusion.

What this is

Yandex Alice 80B uses a MoE architecture (Mixture of Experts — only a subset of parameters fires per token, with 80B total and 3B activated, hence the A3B designation). On paper it stands shoulder-to-shoulder with mainstream open-source models. The developer ran QLoRA (a low-VRAM fine-tuning method) plus 4 V100 GPUs through post-training from scratch, and got stuck on a hardware-level compatibility issue. We want to flag this: open-source weights ≠ a usable model. Between the two sit three gates — toolchain, quantization (compressing into smaller VRAM), and community maintenance.

Industry view

Supporters argue the loss curve is more stable this time and worth pushing on. That dissent matters more: one failed training run means days of compute down the drain. The cost-benefit of post-training an 80B model from scratch is questionable on its face. The deeper risk — every month now sees a few more 80B-tier "open-source releases" appear, but fewer than one in ten actually gets fine-tuned and shipped as something usable. Most end up as benchmark fodder.

Impact on regular people

For enterprise IT: when adopting open-source models, engineering capability and toolchain matter more than the weights.
For working professionals: "AI is strong" often really means "the foundation is strong, the tuning is weak." The last mile is where actual work gets done.
For the consumer market: the open-source map is fragmenting. Russia and Europe are entering the field — it is no longer just the US and China.