This week, r/LocalLLaMA user Excellent-Issue-5956 published a set of test data: after swapping their local AI assistant from Qwen3 27B to Ornith 1.5 35B-A3B, inference speed jumped from roughly 60 chars/sec to 180 chars/sec—about 3x—and both self-built tests passed. Our judgment: the hardware bar for on-premise "data never leaves the building" AI may be lower than we think.
What this is
The user ran a routine "AI assistant" benchmark: a local machine powering an Agent (an AI program that can autonomously complete multi-step tasks) capable of tool use, file I/O, and decision-making. The original setup used two RTX 5070 Ti GPUs running a Qwen3 27B quantized version (a model with weights compressed for lower VRAM usage), hitting roughly 55-70 chars/sec. After swapping to Ornith 1.5 35B-A3B, output climbed to 176-183 chars/sec; two self-built tests—a 9-step tool-calling plus memory suite (9/9) and a 10-question coding suite (10/10)—both scored full marks.
The speedup hinges on a Mixture of Experts architecture (MoE): 35B total parameters, but only ~3B activated per token (8 out of 256 expert subnetworks). Out of 41 layers, only 10 use full attention, while the remaining 31 use linear attention (where VRAM usage barely scales with context length). The result: doubling context from 128K to 256K added only ~2GB of VRAM and dropped speed from 180 to 139 chars/sec—a far gentler curve than traditional architectures.
Industry view
The bullish read: MoE plus linear attention is becoming the "efficiency default" for open-source large models. Qwen, DeepSeek, and Ornith are all on this path. If the curve holds, by 2025-2026 "two consumer-grade GPUs running a 30B-class model" will be routine configuration, not a hobbyist toy.
The sober judgment has to be on the table too. First, this is a single user's test, not a benchmark. Artificial Analysis has not yet scored Ornith; the vendor's self-reported SWE-bench Verified 79 and Terminal-Bench 2.1 68.5 have not been third-party replicated. Third—but not least—the author concedes: a perfect test score only proves "not bad," not "smarter"—it is not uncommon for MoE models to fly on easy tasks and lose to smaller dense ones on tough reasoning. Second, at 256K context speed drops from 180 to 139 chars/sec: long-text slowdown is postponed, not eliminated.
Impact on regular people
For enterprise IT: on-prem AI assistants used to require tens of thousands of dollars in server-grade GPUs. If this curve holds, a $2000-3000 workstation setup could hit "good enough" within 12-18 months. Scenarios touching customer contracts, source code, and internal data are worth tracking.
For individual professionals: freelancers, consultants, independent developers—anyone sensitive to "customer data cannot go to the cloud"—could within 12 months run a usable local assistant on a high-end workstation. Today it is a hobbyist toy; tomorrow it may be standard tooling.
For the consumer market: do not rush out to buy an "AI laptop" because of this test. Built-in AI silicon in laptops—Apple Neural Engine, Qualcomm Hexagon, Intel NPU—still trails discrete graphics cards by 5-10x, but the price gap is narrowing.