What This Is
A post on r/LocalLLaMA this week caught our attention: a Polish legal AI startup is stuck on architecture choice — 50B+ MoE or 27B Dense — and can't get the cost math to close.
The poster is building a B2C legal AI product: document drafting, legal Q&A, paired with automation scripts. Their current 27B Dense model is already good enough after fine-tuning. They want to step up to a 50B+ total-parameter MoE (Mixture of Experts) model that activates only a small slice of its weights per inference.
Their dilemma isn't capability — it's production cost math. They want to rent inference hardware that scales with concurrency. Fine-tuning via LoRA (Low-Rank Adaptation, a lightweight fine-tuning method) is capped at 4 consumer-grade GPUs. And every tool call inside an Agent (an AI that autonomously executes multi-step tasks) loop has a latency budget. This isn't a benchmark race — it's accounting.
Industry View
Proponents will point to the Mixtral, DeepSeek-V3, and Qwen3-MoE track record: MoE delivers near-Dense capability with 5%–20% active parameters, theoretically faster and cheaper. This is a path the big labs have validated.
But the counterpoints are equally sharp. A few replies in the thread deserve flagging. First, "small active parameters ≠ small VRAM" — the full MoE weight set still has to sit in GPU memory, and fine-tuning a 70B MoE on 4 consumer GPUs is essentially a non-starter. Second, stability feedback on MoE under long documents and multi-turn Agent calls is inconsistent; some users report that certain MoE models are less reliable on tool calling (letting the AI invoke external tools) than same-size Dense models — a red flag for low-tolerance domains like legal.
Impact on Regular People
For enterprise IT: mid-size companies building vertical AI products (legal, medical, document processing) have entered a phase where "wrong architecture choice, doubled bill." Dense or MoE directly determines inference cost and concurrency ceiling — this is fundamentally a finance problem.
For individual professionals: white-collar workers using AI for drafting and retrieval don't need to understand the architecture, but should know that the "fast and cheap" AI services on the market are often MoE models routing between experts — which can drop critical details in specific scenarios.
For the consumer market: expect a flood of "vertical AI" products over the next year. Most won't be new models — they'll be tight fine-tunes on top of large models, packaged with toolchains. The cost structure will show up directly in subscription pricing.