This week, a notable post surfaced on Reddit's r/LocalLLaMA: an independent developer trained a 3.87B-parameter MoE model called Apex-2 from scratch on just 86.5B tokens — less than 1/200 of the 18T tokens used by Qwen2.5-1.5B — and it nearly matched the latter on HumanEval+ code benchmarks. The takeaway for us: the barrier to training large AI models is sliding from "big-company exclusive" toward "something a serious individual researcher can pull off."
What this is
Apex-2 is a small AI model trained from scratch by independent developer YOON1v — not fine-tuned from any existing large model. It was pre-trained on 86.5B tokens (leading models typically use 10T–20T), then supervised fine-tuned (SFT — training the model on human-written Q&A pairs to follow instructions) on 2.5B additional tokens. Training ran on NVIDIA GH200 GPUs (a high-end AI chip) using DiLoCo (a distributed training method that coordinates across machines), scaling from 1 to 2 GPUs.
MoE (Mixture of Experts) is an architecture that activates only a subset of sub-networks on demand. Of Apex-2's 3.87B total parameters, just 1.45B activate per inference — effectively delivering more compute per dollar.
Key numbers: HumanEval (AI writes a function that passes tests) 43.9, MBPP (similar programming problems) 56.3; HumanEval+ 41.5, MBPP+ 48.9. For comparison, Qwen2.5-1.5B needed 200x more training data just to tie on HumanEval+.
Industry view
Supporters see this as evidence that data efficiency is leaping forward — architectural innovation (MoE + DiLoCo) now lets small teams produce domain-usable small models. If the trend holds, the cost of training proprietary AI in-house could drop from "starting in the millions" to "tens of thousands or less." We see this efficiency dividend landing most directly on mid-size software companies and vertical-industry service providers.
Skeptics push back hard. Code is just one slice of AI capability — Apex-2's MMLU (a comprehensive knowledge test, essentially an AI's "general exam") sits at just 28.6%, far below the 70%+ of mainstream models. GSM8K (grade-school math word problems) 32.4 and MATH-500 21.0 show math and reasoning remain weak; the 4k-token context window also makes long-document handling a struggle. In short: writing code is not the same as serving as a general-purpose assistant, and an order-of-magnitude gap still separates Apex-2 from production-ready models.
There's another concern: the original post's author admits that DPO (a fine-tuning method that aligns AI outputs more closely with human preferences) made responses longer while tanking every benchmark, so he dropped that training stage entirely. Results that "run but aren't stable" are still a long way from productization.
Impact on regular people
For enterprise IT: the cost of small models for vertical tasks (code generation, customer support replies, document summarization) will keep falling. Training proprietary AI for a single use case may no longer be a fantasy for SMBs.
For working professionals: AI tools will increasingly split by scenario — coding AI gets cheaper and more local, while general-purpose AI for reports and decision-making stays in big-company territory.
For consumers: on-device AI in phones, smart home devices, and car infotainment systems will get stronger thanks to these efficient small models. Next time you upgrade your phone, "on-device large model" may shift from a selling point to a baseline feature.