What this is
This week Reddit user jjusko20 open-sourced SFTMill—a tool that drops the barrier to model distillation (distillation: using a large model to generate training data that teaches a smaller model the same capability) to the level of a YAML (a clean configuration file). Users first define a "training curriculum" in YAML—say, teaching an AI to call tools (operate external software), fix bugs, or trace errors—and the system has a large model generate questions and the target model answer them, auto-producing a dataset ready for fine-tuning (continued training on proprietary data).
Industry view
Supporters argue: the barrier to model customization is collapsing. A year ago fine-tuning a model required an engineering team and a compute budget; now any developer who can write a config file can do it. This is happening in lockstep with the maturation of open-source foundations like Llama, Qwen, and Mistral.
But there are cooler heads: the tool itself isn't a moat. The real bottleneck is curriculum design and data quality assessment. Reddit personal projects have short lifespans and unstable maintenance—a major vendor updating an API spec could break the whole thing. We noticed a more direct signal: the author himself publicly job-hunted at the end of his post—"hope someone sees the project and wants to hire me." The independent developer's tooling ecosystem remains fragile.
Impact on regular people
For enterprise IT: Deploying AI no longer means wrestling with big-vendor APIs. Combining open-source tools with open-source foundations for internal customization could cut costs by an order of magnitude.
For individual careers: "Translating business requirements into training objectives" becomes a new skill—effectively one level above writing prompts.
For the consumer market: Vertical small models will become more common—lightweight AIs specialized for contract review or customer-service scripts, no longer dependent on expensive large models.