Ollama vs Nativ: The Benchmark I Couldn't Find
/ 7 min read
Table of Contents
I run Claude Code all day. When Anthropic’s API has a bad day, I have a fallback: a local model on my Mac (an M5 Max with 128 GB), served by Ollama, wired up so claude-local in my shell points Claude Code at my own hardware instead of the cloud. It’s slower and dumber than the real thing, but it’s mine, and it works when the internet doesn’t.
Then Nativ showed up — a native Mac app for serving MLX models, moving fast (v0.0.1 in July, v0.3.3 by mid-August). Should my fallback switch?
I went looking for a head-to-head comparison and couldn’t find one. Feature guides, yes. MLX-vs-llama.cpp runtime benchmarks, sure. But as far as I can find, nobody had put the same model, at the same quantization, on the same machine, behind both servers, and measured what happens when a real coding agent uses them. So that’s what I did.
The boring part: it’s a tie
Same model, same quantization, everywhere: Qwen3.8-27B at nvfp4 (a 4-bit floating-point format both runtimes support) — about 16 GB of model, and conveniently, both Ollama’s store and Hugging Face’s mlx-community carry the identical quant, so the servers are the only variable.
Every number that benchmarks usually publish came back as a shrug:
| Ollama 0.32.15 | Nativ 0.3.3 | |
|---|---|---|
| Cold start to first token | 1–2 s | 1.5–2.4 s |
| Prompt ingestion (fresh, 4k–48k) | 500–670 tok/s | 520–660 tok/s |
| Peak memory | 15.0 GB | 15.5 GB |
| Quality suite (coding, judgement, agentic, retrieval) | all pass | all pass |
Cold start was the number I most wanted, because it’s the honest metric for a fallback: how long from “nothing running” to “first token” when you actually need it. Nobody publishes it. It turns out to be a non-issue — a modern SSD loads 16 GB in two or three seconds. Overlapping ingestion speeds, near-identical memory, identical quality scores. If I’d stopped here, I’d have written “flip a coin.”
The interesting part: the numbers nobody publishes
A coding agent sends a conversation that grows: an opening request of roughly 100,000 tokens — system prompt, tool definitions, project context — then every turn re-sends all of that plus everything since. Three properties decide whether a local server can live with that traffic, and none of them show up in a tokens-per-second chart.
Does the server obey tool_choice? This is the API parameter an agent uses to force a tool call or forbid one — it’s how agent loops get steered and terminated. Ollama silently ignores it in both directions: force a specific tool and you get friendly prose; say “no tools” and it calls one anyway. Nativ honors both — on its OpenAI-compatible endpoint. Then I probed the Anthropic-compatible endpoint — both servers ship one, it’s what lets Claude Code talk to them at all, and it’s the one my fallback actually uses — and both servers ignore the parameter there. Two servers, two endpoints each, one parameter, three different behaviors, zero errors reported. You’d never know without probing.
Can the server cache a growing prompt? This is the big one. If the server can recognize that this turn’s prompt is last turn’s prompt plus a little more, it skips re-reading the part it’s already seen. Ollama does this beautifully: re-sending a 48,000-token prompt takes 0.2 seconds instead of 70. Nativ, out of the box, re-reads everything, every turn — about 73 seconds each time at that size.
What happens under a real workload? I stopped trusting micro-benchmarks and ran the actual thing: a cold-start, headless Claude Code session doing a small multi-file refactor, graded by whether the project’s tests pass afterward. Ollama finished in 12 minutes, tests green. Nativ hit my 40-minute timeout with the task incomplete — its six requests averaged over eleven minutes each, re-ingesting 597,000 prompt tokens that Ollama’s cache would have skipped.
The detective story
Nativ’s result turned out to be a configuration story, not a capability story — and finding that out was the fun part.
Nativ’s engine has a prefix cache. It ships off. Turn it on (prefixCachingEnabled in its Settings.plist — the UI doesn’t surface it) and there’s a second trap: the default cache holds 32,768 tokens — smaller than Claude Code’s opening request. For this workload the cache never engaged, because the thing it was supposed to cache didn’t fit. Resize the pool (prefixCacheBlocks; the memory cost was negligible in my runs) and my synthetic tests flew: a 100,000-token follow-up turn dropped from 233 seconds to 1.7 seconds.
Then I re-ran the real Claude Code task and it still re-read everything, every turn. (Credit where due: with the tuned cache the refactor did finish its edits — tests green — but the session still blew the 40-minute cap.) Synthetic conversations cached perfectly; real ones never did. That contradiction bothered me enough to put a logging proxy between Claude Code and the server and capture the actual request bytes.
The answer was two facts colliding. First, Claude Code quietly mutates messages it already sent — retry nudges get injected into old turns, and one bookkeeping message flips its serialization shape between requests (the same text sent as a list of blocks one turn and a plain string the next). Second, Nativ’s Anthropic endpoint takes system-role messages from inside the conversation and hoists them to the front of the prompt it renders — an implementation choice, not anything the Anthropic protocol requires. Put those together: the one part of the request that changes every turn gets moved to the very beginning, in front of 90,000 tokens of perfectly stable content — and a prefix cache can only reuse an unbroken prefix. One varying byte at the front invalidates everything behind it.
No benchmark would have found that; it took watching real traffic.
Where it landed
My fallback stays on Ollama — a 12-minute success beats a 40-minute timeout, and its prompt cache is the single most important property for agent-shaped work. (Prefix caching isn’t the whole cost, either: at 100k tokens of context, generation itself sags to 8–18 tokens per second on this hardware, a third of its small-context speed.) But I came away impressed by Nativ’s bones: same-speed engine, better tool_choice compliance on its OpenAI endpoint, and its defining loss traceable to defaults and one prompt-assembly decision. If you’re on Nativ today, the two plist settings above get you most of the way. All of it looks fixable, so I filed it upstream: prefix caching defaults, Ollama’s tool_choice, the Anthropic-layer tool_choice gap, and the hoisting-vs-cache collision. Two of them were in the Nativ maintainer’s tracker within the hour.
One caveat on scope: this is one model, one Mac (Apple Silicon only — Nativ doesn’t exist elsewhere), and one client at one version (Claude Code 2.1.234, whose exact request serialization is half the story). Treat it as a method, not a ranking. If you’re picking a local server for agent workloads, the method travels as three tests:
- Force and forbid a tool via
tool_choice, on the exact endpoint your agent calls — not the one the benchmark used. - Resend a big prompt, then a grown one, and time both — identical-resend tests the cache exists; grown-resend tests whether it survives your client’s real serialization.
- Run one real task cold and grade the artifact, because agents live and die by ingestion, and ingestion is exactly what the published benchmarks don’t test.
A note on who did the work
This is the first post on this blog I didn’t write. My AI agent — Claude, the same one whose fallback path this investigation exists to protect — designed and ran the benchmarks, chased the cache mystery through the logging proxy, filed the upstream issues, and wrote this post. I directed, reviewed, and made the calls; an independent panel of three other AI models (OpenAI, Google, and a local Qwen) reviewed the test design and fact-checked this draft. The benchmark grading itself is deterministic code — executed tests and parsed answers, no AI judging AI. It felt only fair to say so in a post about trusting your tools.