Back to home

Compare

Comparing: Three Old PCs Form an AI Team — Multi-Model Routing Heads to the Garage & 三台旧电脑拼出AI团队 — 多模型调度正在从大厂走进车库

AEN
Claude CodeAnthropicLiteLLM·

Three Old PCs Form an AI Team — Multi-Model Routing Heads to the Garage

This week on Reddit's r/LocalLLaMA (the open-source community for running large models locally), a user posed a question: how do you string together three machines at home — a PC, a Mac mini, and a Raspberry Pi — into an AI team that splits work across different models? On the surface it looks like a geek toy, but it shows that "multi-model routing" — a concept once exclusive to big tech — is entering everyday view.

His hardware list: the main desktop runs a small Qwen 3.8B-class model (13–15 tokens per second), the 16GB Mac mini runs Orninth 9B or Gemma4 12B, and the Raspberry Pi 5 runs a 3B model. What he wants is a "router" — a dispatcher that breaks tasks into pieces and sends each to the most suitable machine and model. What got him thinking about this was the response speed of Claude Code (Anthropic's code agent) — simple conversations go to fast small models, complex reasoning stays with the big ones, and the whole thing speeds up.

What this is

In essence, this is an AI orchestration problem: deciding "who should handle this piece of work." It's the same logic as load balancing and microservices scheduling in traditional software — except the thing being routed is "models" instead of "services."

For non-technical readers, here's the analogy: in the past we let one AI model handle everything from start to finish — like asking a general practitioner to diagnose, operate, and prescribe. Multi-model routing turns AI into a clinic: reception does triage, specialists perform surgery, pharmacists dispense the drugs.

Industry view

We've noticed this kind of demand spiking noticeably in the technical community. Multi-model routing used to be a big-company game — techniques like "ensemble learning" (multiple models vote for the best answer) and "Mixture of Experts / MoE" (one large model split into sub-models that each handle part of the work) have been quietly running inside OpenAI and Anthropic, out of reach for ordinary developers.

The open-source ecosystem is catching up. Projects like LiteLLM (a unified interface layer for calling multiple model providers) and OpenRouter (a routing service aggregating multiple cloud model providers) already run in the cloud — but a mature "home router" for local setups doesn't exist yet. That's exactly where this user is stuck.

That said, we want to add two sober notes. First, for most people, "one machine running one strong model" is still the simplest and good-enough setup, and the complexity from routing often isn't worth it. Second, routing itself carries overhead — splitting tasks, passing data, aggregating results — and if poorly designed, it eats the speed gains. Third, multi-machine deployment means more failure points: one Raspberry Pi going offline brings the whole pipeline down.

In short, multi-model routing is moving from "internal dark art" to "civilian tool," but it hasn't reached "out-of-the-box ready" yet.

Impact on regular people

  • For enterprise IT: Future AI procurement may shift from "buy one giant model" to "buy a combination package," and cost structures and ops thinking will have to adapt accordingly.
  • For working professionals: You don't need to worry about this term just yet, but people who understand a bit of "AI architecture" are seeing growing wage premiums — especially those who can clearly explain "when to use a big model, when to use a small one."
  • For the consumer market: A new hardware category may emerge — small AI servers, home inference boxes — and it also means old PCs and laptops at home could get a "second career."
BZH
Claude CodeAnthropicLiteLLM·

三台旧电脑拼出AI团队 — 多模型调度正在从大厂走进车库

本周 Reddit 的 r/LocalLLaMA(本地部署大模型的开源社区)上一位用户抛出个问题:怎么把家里的三台机器——一台 PC、一台 MacMini、一个树莓派——串成一支 AI 团队,分工跑不同模型。这事看似极客玩具,但说明「多模型调度」这个原本只属于大厂的概念,正在走进普通人的视野。

他的设备清单是:主力台式机跑 Qwen 3.8B 类小模型(每秒 13-15 tokens)、MacMini 16GB 跑 Orninth 9B 或 Gemma4 12B、树莓派 5 跑 3B 模型。他想要的,是一个「调度器」——把任务拆碎,按需派给最合适的那台机器和那个模型。让他动这个念头的,是 Claude Code(Anthropic 出品的代码 Agent)的响应速度——简单对话派给快速小模型、复杂推理留给大模型,整体就快了。

这是什么

本质上,这是一个「AI 编排」(orchestration)问题:决定「这一段活该交给谁」。它和传统软件里的负载均衡、微服务调度是同一种思路,被调度的对象从「服务」变成了「模型」。

对非技术读者来说,可以这样理解:过去我们让一个 AI 模型从头干到尾,像让一个全科医生既看病又动刀又开药;多模型调度则是把 AI 变成一个诊所,前台分诊、专科医生手术、药剂师配药。

行业怎么看

我们注意到,这类需求最近在技术社区出现频率明显升高。多模型调度过去是大公司的玩法——「集成学习」(多个模型投票取最优)、「混合专家 MoE」(一个大模型切成多个子模型各自处理一部分)这些技术,过去都在 OpenAI、Anthropic 内部默默运转,普通开发者碰不到。

开源生态正在跟上。LiteLLM(统一调用多家模型的接口层)、OpenRouter(聚合多家云端模型的路由服务)这类项目已经在云端跑起来了;但本地版的「家用路由器」还没有成熟方案——这正是这位用户卡住的原因。

不过,我们也要提醒两个冷静的声音。第一,对绝大多数人来说,「一台机器跑一个强模型」仍然是最简单也够用的方案,调度带来的复杂度常常得不偿失。第二,调度本身有 overhead(额外开销)——拆任务、传数据、汇总结果,如果设计不好,速度优势会被吃掉。第三,多机器部署意味着更多故障点,一台树莓派掉线,整条流水线就停了。

换句话说,多模型调度正在从「内部黑科技」走向「民用工具」,但还没到「开箱即用」的阶段。

对普通人的影响

  • 对企业 IT:未来采购 AI 可能从「买一个超大模型」变成「买一套组合套餐」,成本结构、运维思路都得跟着调整。
  • 对个人职场:暂时不用担心被这个词砸到,但懂一点「AI 架构」的人溢价空间在变大,尤其是能讲清楚「什么时候用大模型、什么时候用小模型」的。
  • 对消费市场:可能催生一类新硬件——小型 AI 服务器、家用推理盒子;同时也意味着家里的旧电脑、旧笔记本有望「再就业」。