A Reddit post circulated in the hardware community this week: a user who upgraded to an RTX 5070 Ti stuffed their idle RTX 4070 Ti into a homemade OCuLink (a chassis-external GPU interface standard, a simplified Thunderbolt alternative) enclosure, connected it to their main rig, and ran the Qwen 3.8 27B model with "perfectly usable" results. Small story, but worth our noting—hardware barriers for running usable mid-size LLMs locally are being chipped away by this kind of DIY tinkering.

What this is

The poster originally upgraded their GPU to run local models in the 27B tier. Two cards in one chassis overheat, and they didn't want the 4070 Ti collecting dust, so they put it in an external enclosure, connected to the host via a PCIe 4.0 OCuLink expansion card on the motherboard. The combo ran Alibaba's open-source Qwen 3.8 27B well enough, prompting the post asking for experiences and warning signs.

OCuLink isn't new—it delivers higher bandwidth than Thunderbolt at lower cost, but sacrifices portability; cabling and power delivery both require manual work. It was previously more common in industrial and crypto-mining setups. Using it for local AI inference is a niche play.

Industry view

The r/LocalLLaMA community is treating this post as a new case study in "low-cost dual-GPU inference." The bull case: consumer GPUs plus external enclosures can already cover part of the inference demand for sub-70B models, and personal AI workstations will become more common.

The bearish and risk-skeptical voices are equally sharp. Commenters warn that OCuLink cables degrade noticeably past 1 meter, and sustained high loads can burn the connector; others point out the same budget goes further renting cloud GPUs by the hour. The real bottleneck isn't hardware—it's software. Local workflows that smoothly deliver production-grade results at the 27B tier remain a niche tinkerer toy, far from "plug-and-play for normal users."

Impact on regular people

For enterprise IT: No need to consider local 27B deployment in the short term—procurement, compliance, and ops costs far outweigh simply calling cloud APIs.

For working professionals: Daily AI use for emails and edits requires zero understanding of eGPUs; but product managers and engineers who need to demo local AI should start paying attention to this "second-hand hardware + external enclosure" playbook.

For the consumer market: Second-hand GPU prices will stay under pressure, but the 5070 Ti and similar new products still target gaming and creators—local AI is just a bonus scenario.