A Reddit user this week pulled off something notable: ran Qwen 35B on a 16GB GPU at 60 token/s, with scores 2.5 points higher than the original. The target of the rework was the open-source inference engine llama.cpp (the core program that lets large models run locally); the new version is called AgrillaMoE, optimized specifically for Qwen-series MoE models (Mixture of Experts architecture—models containing multiple "expert" sub-networks, only a portion of which activate per query).

What this is

The core data points are these: a 35B-parameter model (total model size) running on a rented V100 16GB GPU, hitting 57-60 token/s generation speed under extreme 2-bit quantization, with compatibility for both OpenAI and Anthropic API interfaces—meaning coding tools like Claude Code can directly interface with the local model.

More critically, the developer implemented a runtime technique called "MoE expert expansion": the model originally activates only 8 expert sub-networks per call, but a parameter (--moe-experts 20) dynamically expands activation to 20 experts, with the threshold being that any new expert only needs to reach 80% of the confidence of the original top-8 to be included—no retraining required. On GPQA-Diamond (a graduate-level science reasoning benchmark), scores climbed from 81.82% to 84.34%, a 2.5-point gain.

Industry view

Supporters see this as a sign that the open-source ecosystem is "changing tactics"—no longer obsessing over model size, but mining gold from engineering layers like inference efficiency and routing strategy (the algorithm deciding which experts to activate each turn), creating cost pressure on cloud API vendors.

But we also flag three reservations: first, this is a single developer's Reddit post, and reproducibility remains questionable; second, "expert expansion" is fundamentally an engineering hack, and stability under long-context, complex-instruction scenarios hasn't been thoroughly tested by the community; third, whether the rental cost of a 16GB V100 actually beats direct API calls—the math hasn't been done. In other words, this approach works at the hobbyist level, but it's still some distance from enterprise production environments.

Impact on regular people

For enterprise IT: The hardware bar for self-hosted AI inference has dropped another notch, but whether to replace cloud APIs depends on the actual latency and stability requirements of the business.

For individual professionals: If you're using tools like Claude Code to handle sensitive code, it's worth considering the local route—at least there's now an additional option for "data never leaves the company server."

For consumer markets: Cloud API price wars will be pressured by community innovation like this, but in the short term, consumer experience still favors the cloud—don't let tech blog hype set your expectations.