← All posts

Engineering · AI

Local LLMs on Apple Silicon: An Honest Verdict

Mihajlo Petrović5 min read

How much RAM you actually need, what quantisation costs you, which tools are worth installing - and a clear answer on where a laptop model beats an API and where it is not remotely close.

Running a model on your own laptop is genuinely appealing: no per-token cost, no data leaving the machine, works on a plane. Apple Silicon is unusually good at it thanks to unified memory — the GPU can address the same pool as the CPU, so a MacBook can load models that would need an expensive discrete GPU on other hardware.

I've spent enough time on this to have a clear opinion. It's useful. It is not a replacement for a frontier model, and most of the disappointment comes from expecting it to be one.


The Only Spec That Matters at First: RAM

Model size is the binding constraint, and it's simple arithmetic. A quantised model needs roughly its parameter count in bytes at 4-bit precision, plus overhead for context:

Model size ~4-bit footprint Comfortable on
7-8B ~5 GB 16 GB Mac
13-14B ~9 GB 16 GB (tight) / 24 GB
30-32B ~20 GB 32 GB+
70B ~40 GB 64 GB+

"Comfortable" matters: the OS and your actual work need memory too. Filling RAM to the brim means swapping, and a swapping model is unusably slow.

The second spec is memory bandwidth, which is what governs generation speed. This is why a Pro/Max chip generates noticeably faster than a base chip with the same amount of RAM — and why more cores don't help nearly as much as people expect.


Quantisation, Briefly

Quantisation shrinks weights from 16-bit to 8, 4, or fewer. The rough consensus, which matches my experience:

  • 8-bit — nearly indistinguishable from full precision, twice the footprint of 4-bit
  • 4-bit — the sweet spot; small quality loss, half the memory
  • Below 4-bit — degrades quickly, noticeable on reasoning and instruction-following

The more useful rule: a bigger model at 4-bit beats a smaller model at 8-bit at the same memory budget, almost always.


The Tools

Ollama — the easiest start. ollama run <model>, an HTTP API on localhost, model management handled. This is where to begin.

LM Studio — a GUI, easy model browsing, good for trying things without touching a terminal. Also useful for non-developers on your team.

MLX — Apple's own array framework, tuned for Apple Silicon. Noticeably faster than generic runtimes on Mac for some workloads, and the right choice if you're building something rather than just chatting.

llama.cpp — what several of the above build on. Reach for it directly when you want control over quantisation and runtime flags.


Where Local Genuinely Wins

Sensitive data. The strongest case by far. Documents that can't leave the building, personal notes, client material under NDA. A locally-run 8B model that never touches a network beats a better model you're not permitted to send the data to — often the deciding factor in regulated environments.

Bulk, simple, repetitive work. Classifying 50,000 support tickets, extracting fields from documents, tagging a dataset. Per-item quality requirements are low, volume is high, and the marginal cost is electricity. Run it overnight.

Offline and latency-sensitive. Planes, trains, air-gapped networks. And for short completions, a small local model on a fast chip can start responding sooner than a network round trip.

Learning how any of this works. Running a model locally teaches you more about context windows, quantisation, sampling and prompt formats in a weekend than months of API usage.


Where It Loses, Clearly

Agentic coding. This is the honest headline. Multi-step tool use requires reliably producing well-formed calls, recovering from errors, and holding a plan across many turns. Local models in the 7-30B range are dramatically less reliable at this than frontier models — not "slightly worse", a different category of experience. If you've been frustrated trying to drive a local model as a coding agent, the model is the problem, not your setup.

Long context. Advertised context lengths and useful context lengths diverge sharply, and memory use grows with the context you actually fill. Frontier models measure context in the hundreds of thousands to a million tokens; a local setup is usually working with far less in practice.

Hard reasoning. Non-trivial refactors, subtle debugging, ambiguous requirements. The gap is real and it's the whole reason frontier models cost what they do.

Battery and heat. Sustained generation runs the machine hot and drains it fast. Fine at a desk, less fine on that plane you bought it for.


The Verdict

Local for private, bulk, or simple. Frontier API for anything agentic or hard.

The setup I'd actually recommend to a developer with a 32-64 GB Mac: a mid-size local model for the private and repetitive work, a frontier model through an API for coding agents and anything requiring real reasoning. They're complementary, not competing — and the money you save by routing bulk work locally pays for the calls where quality matters.

What I'd avoid is the trap of measuring local models against the frontier and concluding they're useless. That's the wrong comparison. The right one is against not being able to do the task at all because the data can't leave your machine — and by that standard, a laptop that runs a capable model offline is remarkable.

  • #local llm
  • #ollama
  • #apple silicon
  • #mlx
  • #macbook
  • #ai

Written by

Mihajlo Petrović

Software engineer in Belgrade. Builds his own products and the AI automations that keep them running.

Have a task that repeats every week?

Tell me about it. If it can be automated well, I will show you how. If it cannot, I will say that too.

Tell me what to automate