AI

Fix: Ollama slow on Mac M-series — the real fixes

Ollama running at 2 tokens/sec on your M1/M2/M3/M4 Mac when it should be 30+. Here's why, and the actual fixes that recover full speed.

Ollama on an M-series MacBook can serve 30-60 tokens per second on a 7B model. When it drops to 2-5 tokens per second, something is wrong. The Mac isn’t underpowered — it’s misconfigured, thermally throttled, or being starved of memory. Here’s each cause and the specific fix.

Baseline — what “fast” looks like

Rough expected throughput on Apple Silicon at Q4_K_M quantization:

  • M1 8GB — 20-25 tok/s on 7B, unusable on 13B
  • M1/M2 16GB — 30-40 tok/s on 7B, 12-18 tok/s on 13B
  • M2 Pro/Max 32GB — 40-55 tok/s on 13B, 15-20 tok/s on 34B
  • M3 Max 64GB — 50+ tok/s on 34B, viable on 70B
  • M4 / M4 Pro / M4 Max — improved bandwidth over M3; expect 10-25% headroom over comparable M3 tiers on the same models

If numbers are 5x lower than the row above, one of the following causes applies.

Diagnose first

Get the actual tokens per second before guessing:

ollama run llama3 --verbose "Write a haiku about coffee."

At the end Ollama prints:

total duration:       4.234s
load duration:        1.891s
prompt eval count:    18 token(s)
prompt eval rate:     42.31 tokens/s
eval count:           47 token(s)
eval rate:            22.45 tokens/s

eval rate is the number that matters for generation speed. prompt eval rate matters for long contexts.

Also check memory pressure while it runs:

memory_pressure

If output shows System-wide memory free percentage: 5% while Ollama runs, the model is spilling to swap — jump to Cause 2.

Cause 1 — Metal GPU not being used

Ollama on Mac uses Metal (Apple’s GPU API). If it fell back to CPU, throughput collapses.

Check the server logs:

tail -f ~/.ollama/logs/server.log

Look for a line like:

llama_new_context_with_model: using Metal GPU with 10 layers

If it says using CPU only or the GPU line is missing, Metal isn’t active.

Fix: update Ollama to the latest version. Metal detection improved significantly across 2024-2026 releases.

brew upgrade ollama
# or download the DMG from ollama.com
ollama --version

Then restart the server:

# Simplest: quit and reopen the Ollama menu bar app (click icon → Quit, then relaunch)
# From terminal:
pkill Ollama && open -a Ollama

Force GPU layers on all-Metal:

OLLAMA_NUM_GPU=999 ollama run llama3

999 tells llama.cpp to offload every layer to GPU. Any number higher than the model’s layer count acts the same.

Cause 2 — Model too big, spilling to swap

macOS quietly pushes RAM to swap when pressure rises. Once the model straddles RAM and swap, generation stalls per token.

Check swap usage:

sysctl vm.swapusage

If used = 8000M or higher while Ollama is running, the model doesn’t fit in physical RAM.

Fix — use a smaller quant of the same model:

ModelQ8_0Q5_K_MQ4_K_MQ3_K_M
Llama 3 8B~9 GB~6 GB~5 GB~4 GB
Llama 3 13B~14 GB~9 GB~7.5 GB~6 GB
Llama 3 34B~36 GB~24 GB~20 GB~16 GB

For an 8GB Mac stick to 7B/8B at Q4 or lower. For 16GB, 13B at Q4 fits comfortably.

Pull a smaller quant:

ollama pull llama3:8b-instruct-q4_K_M

Cause 3 — Other GPU apps stealing memory

Ollama shares GPU memory with everything else running. Chrome, Slack, Notion, Zoom, Figma, and Xcode all use Metal.

Fix — close them before benchmarking:

# Quick free-RAM check
memory_pressure

# Kill big GPU users
killall "Google Chrome"
killall Slack
killall zoom.us

Real-world impact: closing Chrome alone can raise Ollama throughput by 30-50% on 16GB machines.

For long inference sessions, prefer a browser like Safari that shares Metal cheaply, or route the LLM to a headless remote if you need Chrome open.

Cause 4 — Thermal throttling

Long generation runs heat the SoC. macOS throttles the GPU when the chassis exceeds thermal limits — plastic MacBook Air chassis (no fan) hit this fast; MacBook Pros with fans last longer.

Check throttle state:

sudo powermetrics --samplers smc -i1 -n1 | grep -i "cpu die\|gpu die"

Temperatures above 95°C on GPU die = throttling.

Fixes in order:

  1. Raise the laptop on a stand for airflow. This alone recovers ~15-20% on Air chassis.
  2. Use an active cooling pad if generating for extended sessions.
  3. Reduce the model size — smaller models produce less heat per token.
  4. Set a lower generation temperature or shorter max tokens if you were testing long outputs.

Cause 5 — Long context sizes

Large context windows (32k, 128k) allocate large KV caches in GPU memory. That memory competes with weights.

Check current context size:

ollama show llama3 --parameters | grep num_ctx

Fix — reduce context size to what you actually need:

Create a Modelfile:

FROM llama3
PARAMETER num_ctx 2048

Then:

ollama create llama3-short -f Modelfile
ollama run llama3-short

Dropping from 32k to 4k context can double throughput on a memory-constrained Mac.

Cause 6 — Prompt too long, prefill dominates

Generation is measured in eval rate. But if your prompt is 8000 tokens long, prompt eval rate (prefill) can take longer than the actual generation.

Check both rates:

ollama run llama3 --verbose "..."

If prompt eval count is huge and prefill takes 20 seconds before the first output token appears, the fix is prompt engineering, not Ollama config:

  • Summarize retrieved context before prompting
  • Cache system prompts (Ollama reuses KV state for identical prefixes)
  • Chunk long documents; iterate rather than dumping all at once

Cause 7 — Wrong model architecture for the task

Some models are inherently slow on Metal. Mixture-of-Experts (MoE) architectures like Mixtral 8x7B run poorly on Apple Silicon compared to dense models of similar quality.

Fix — pick dense models for Mac:

  • Llama 3 / 3.1 / 3.2 / 3.3 series (and small Llama 4 dense variants where available)
  • Qwen 2 / 2.5 / 3 series
  • Mistral 7B / Nemo / Small variants
  • Gemma 2 / Gemma 4 small variants

Avoid on Mac unless you have 64GB+ RAM:

  • Mixtral 8x7B, 8x22B (MoE)
  • Command R+ 104B (too large)
  • DeepSeek V3 / V4 flagship MoE (too large)
  • Llama 4 Maverick (400B MoE — needs enterprise hardware)
  • Mistral Large 3 (675B MoE)

The universal Ollama speed audit

Run this checklist any time throughput drops:

# 1. Version — must be recent
ollama --version

# 2. Metal active in logs
tail -20 ~/.ollama/logs/server.log | grep -i metal

# 3. Memory pressure clean
memory_pressure

# 4. No swap in use
sysctl vm.swapusage

# 5. Actual tok/s
ollama run llama3 --verbose "test"

# 6. Context size sane
ollama show llama3 --parameters | grep num_ctx

Any red flag in that list is the fix target.

Prevention — sensible defaults per machine class

Save these Modelfiles once and use them going forward.

8GB Mac (Air, base M1/M2):

FROM llama3:8b-instruct-q4_0
PARAMETER num_ctx 2048
PARAMETER num_gpu 999

16GB Mac:

FROM llama3:8b-instruct-q5_K_M
PARAMETER num_ctx 4096
PARAMETER num_gpu 999

32GB+ Mac Pro/Max:

FROM qwen2.5:14b-instruct-q5_K_M
PARAMETER num_ctx 8192
PARAMETER num_gpu 999

Bottom line

Slow Ollama on Apple Silicon is almost always one of: outdated version missing Metal fixes, model too big for RAM (swap thrash), other GPU apps hogging memory, thermal throttling, or oversized context window. Check ollama run --verbose, memory_pressure, and sysctl vm.swapusage — the true cause is visible in under a minute. Right-size the model to the machine and Ollama recovers full native speed.