What This Is

On September 24, Amazon's team published the Rufus-Air training manual: it takes GLM-4.5-Air-Base (a 106-billion-parameter open-source base model) and runs it through 8 sequential stages—supervised fine-tuning, reasoning, code, instruction following, general agent, coding agent, search agent, and RLHF (reinforcement learning from human feedback)—while disclosing full details on data filtering, reward design, stage ordering, and infrastructure.

What deserves attention is not the number "8," but a more fundamental principle—we summarize it as "hard-first, soft-later, gated progression": early stages rely only on hard rewards with clear right/wrong answers (math answers, code test pass/fail), and only later introduce soft rewards tied to subjective judgment (whether a response is pleasant). Each stage has a gate: if the target capability doesn't improve by at least 2%, no promotion to the next stage; if general capability regresses by more than 1%, immediate rollback.

Industry View

Supporters argue this manual puts engineering details back on the table that the industry often glosses over—data filtering, reward reliability, stage gating, rollback mechanisms—delivering more long-term value than yet another benchmark leaderboard. Mid-sized and small teams don't need to replicate 106-billion-parameter training, but they can copy the engineering skeleton: reward ordering, difficulty filtering, stage gating, and rollback checkpoints.

Criticisms are equally clear. "Open" doesn't mean low-cost—reproduction still demands massive compute and strict data version management. The absence of new human annotation and internal teacher distillation silently caps certain capability ceilings. "Competitive results" doesn't mean optimal for every model and language. We also flag a hidden concern: chasing "every stage must improve" metrics may train models to please the judges rather than genuinely master the tasks themselves.

Impact on Regular People

  • For enterprise IT: when procuring AI services, you can now ask vendors "are your rewards hard metrics or soft judges?"—a concrete question that separates engineering maturity from marketing.
  • For working professionals: when using AI for code or document review, the verifiable parts (does the code run, are the numbers right) will be more reliable, while tone and style still need human oversight.
  • For consumer markets: differences in consumer AI product experience fundamentally come from how the training stages are designed—not all from "parameter size."