6 min read

Local AI Got Good Enough. Should You Actually Run It?

A couple of years ago, running a language model on your own machine was mostly a party trick. It was slow, it said dumb things, and you'd go back to a cloud API by lunchtime. That has changed. Open-weight models are genuinely useful now, the tooling is close to painless, and a decent laptop can do…

A couple of years ago, running a language model on your own machine was mostly a party trick. It was slow, it said dumb things, and you'd go back to a cloud API by lunchtime. That has changed. Open-weight models are genuinely useful now, the tooling is close to painless, and a decent laptop can do real work.

That doesn't mean you should cancel your API keys. "Good enough" depends heavily on what you're doing. Here's how I'd think it through: four reasons to go local, and four things you give up when you do.

Reason one: your data stays put

This is the clearest win. When inference runs on hardware you control, your prompts never leave your network. If you handle medical records, legal documents or financial data, that matters a lot. MindStudio and Ertas both point to these as the cases where sending data to a cloud API creates compliance risk, with regimes like GDPR and GLBA in the background.

One honest caveat, which MindStudio makes too: local isn't automatically secure. You've moved the risk from someone else's data center to your own machines, so now you're the one who has to lock them down. A model server that's open to the office Wi-Fi isn't a privacy upgrade.

Reason two: cost, but only at volume

This is where people tend to get carried away. Local inference isn't free. You pay up front for hardware and then keep paying for electricity. Whether it saves money comes down to how many tokens you push through.

MindStudio's rule of thumb is that cloud is usually cheaper below about 1M tokens a day, and local often pays off above about 5M a day. Their worked example: 10M tokens a day at $5 per million comes to roughly $18k a year, against a $15–20k server that should last three or more years. That sounds like an easy call until you add electricity. A 300–500W GPU box running all the time costs roughly $150–300 a month. And some cheap cloud models now cost under $0.50 per million tokens, which wrecks the math entirely.

The other estimates don't agree with each other. Ertas, which sells local deployment tooling (so factor that in), puts the crossover at 10–50M tokens a month. PromptQuorum says local breaks even in 6–12 months above about 10M tokens a month. None of these numbers come with sources, so treat them as rough guidance and not as gospel.

The practical advice I'd take from MindStudio: once your cloud bill is steadily above $500–700 a month and your volume is predictable, sit down and model the hardware option. Below that, you're probably buying a hobby.

Reason three: latency, specifically for agents

Raw speed is mixed, and we'll get to that. But there's one latency argument that's strong. MindStudio puts a cloud round trip at 200–800ms. That's nothing for a single chat message. An agent that makes 10–30 calls to finish a task, though, picks up 5–15 seconds of pure network and queueing overhead. Run inference locally and most of that goes away.

Throughput is a different story, and it depends on your hardware:

  • CPU only: PromptQuorum cites 10–25 tokens/sec for a 7B model, which is 4–10x slower than cloud. Usable, but you'll feel it.
  • A proper GPU: 130–160 tokens/sec on an RTX 4090. The catch is that the card is discontinued and goes for $2,000–2,600 used.
  • Apple Silicon: In March, Ollama shipped an MLX-powered preview (v0.19) and reported decode going from 58 to 112 tokens/sec and prefill from 1154 to 1810 tokens/sec on Qwen3.5-35B-A3B. Read the fine print, though. It's a preview, it's one model, and the before-and-after runs used different quantizations (NVFP4 vs Q4_K_M), so it isn't a clean comparison. Ollama also recommends a Mac with more than 32GB of memory.

If you're wondering whether Ollama's convenience costs you speed compared with raw llama.cpp, it mostly doesn't. AgenticWire measured them as basically tied on an M1 (48.1 vs 47.2 tokens/sec on a 1.5B model), with Ollama using about 54MB more memory. That's one machine and one small model, but it matches what most people find in practice.

Reason four: it works on a plane

This one's short. No network, no problem. If you work offline, in an air-gapped environment or anywhere with flaky connectivity, a local model is the only kind you can count on.

What you give up

Now for the part that local AI enthusiasts tend to rush past.

Frontier capability. PromptQuorum puts local 7B models 10–20 points below frontier models, for example roughly 45–55% on HumanEval against about 90%. The gap narrows a lot at 70B, but that model size needs 40–48GB of RAM. MindStudio puts open-weight models about 3–6 months behind the frontier overall, with the smallest gap on extraction, summarization and classification. That's a useful hint about which jobs to hand them. Ertas argues that fine-tuned local models can match frontier models on narrow tasks, which is plausible, but they don't publish a benchmark to back it up.

Long context. Local models typically handle somewhere between 4K and 128K tokens of context. Cloud models offer 128K to over 1M. If your workflow involves feeding in whole codebases or long document stacks, that difference is real.

Elasticity. A cloud API scales up for a traffic spike and scales back down when it's over. Your GPU box doesn't. You size it for peak load or you queue requests.

Someone else running the uptime. There's no SLA when the server is under your desk. Drivers break, disks fill up with 2–50GB model files, and you're on call. Local models also don't come with built-in web access or the other conveniences cloud providers bundle in.

So what should you actually do?

Honestly, probably both. MindStudio and Ertas land on the same pattern, and I think it's right: use a frontier cloud model for planning and hard reasoning, and send the high-volume or sensitive execution work to local models. Let the big model write the plan, then let the local one classify ten thousand support tickets or summarize the patient notes that shouldn't leave the building.

If you're a solo developer with a modest API bill, local is mostly worth it for privacy, offline work and tinkering, and that's a perfectly good reason. If you're running agents at scale with steady volume and sensitive data, run the numbers this quarter. The models are finally good enough that the answer could surprise you.