When a client said, “Don’t upload the files to the cloud,” I honestly panicked a little
I got stuck on this too. When I first heard “run AI locally,” my instinct was to buy the strongest GPU I could afford. Later I realized that with a lot of models, the issue isn’t that they run slowly—it’s that they don’t fit at all.
What this really is: it’s not just horsepower, it’s trunk space too
Put simply, running AI locally comes down to two things: whether you have enough RAM/VRAM, and whether the read speed is fast enough. The first is like trunk space; the second is like loading speed. Last Tuesday at 8 p.m., at a Starbucks in Binjiang, Hangzhou, Ajie—a freelance designer—unscrewed his iced Americano and showed me the mini PC he had just ordered. He didn’t buy the most expensive GPU. He picked a machine with more memory, because he needed to fit a larger model first so he could “load in” the client’s documents. I got this wrong before too. I used to think bigger compute numbers automatically meant better value. Later I learned: if it can’t fit, fast doesn’t help.
What it costs to copy this today
Money: trying the basics can cost 0 yuan. If we actually buy a machine, a high-memory mini PC is about 20,000 yuan, while a top-tier consumer GPU setup is about 30,000 yuan. Time: 15–30 minutes is enough to understand the key idea first. Technical barrier: no coding needed—we just need to grasp that capacity decides whether it fits, and speed decides how quickly it responds. First step: open the page for whatever local AI tool you’re looking at, click “model details” or “parameter info,” check how large the model is, and only then decide what hardware to buy. This tool and way of thinking won’t be necessary for everyone. If privacy-sensitive files, long documents, or local processing still aren’t part of the work, it’s fine not to bother with it yet.
If I were choosing at different stages, here’s how I’d do it
If I were just getting started, I wouldn’t buy hardware first. I’d test smaller models on the computer I already have and confirm that I really do have a local-use case.
If I already had one or two clients and only occasionally needed to handle contracts, interviews, or proposals, I’d look at the “more memory” option first. Even if it answers a bit more slowly, I’d rather make sure it runs at all.
If I were scaling up and starting to process lots of documents steadily, or share the setup across multiple people, I’d budget for “fits the model” and “responds fast” separately: capacity for big models, speed for small models and high-frequency Q&A. We don’t have to do it all in one shot. Buying the right thing matters more than buying the expensive thing.