MLX vs Ollama on Apple Silicon (2026)
The question changed this year. Ollama now runs MLX itself, and its v0.40 release candidate makes MLX the default on Apple Silicon. So the real questions are which engine a given model runs on, and whether running mlx_lm yourself still buys anything.
If you run local AI on a Mac, you eventually hit the question: serve models through Ollama, or through MLX (Apple's own machine-learning framework)? Until this year those were two separate choices. Now Ollama can run MLX itself, so the comparison has three sides: Ollama on its llama.cpp engine, Ollama on its MLX engine, and mlx_lm.server run directly. We run Ollama and mlx_lm side by side on the same hardware, so here is the honest version.
Short answer
- Ollama and MLX are no longer either/or. Since v0.19 (March 2026), Ollama on Apple Silicon can run models on MLX. Which engine you get depends on the model's format, not a setting: MLX (safetensors) models run on MLX, GGUF models run on llama.cpp.
- v0.40 makes MLX the Mac default. The v0.40.0 release candidate (September 25, 2026) runs a model on MLX automatically wherever the MLX runtime supports its architecture. It is a pre-release: the latest stable Ollama is v0.34.4.
- Running
mlx_lm.serveryourself is now about coverage and control (any MLX conversion on Hugging Face, your own KV-cache and draft-model flags), not raw speed. - Not on a Mac? None of this applies. Ollama runs llama.cpp on Linux and Windows, with or without an NVIDIA GPU, and Ollama Herd routes Macs and other machines behind one endpoint.
How Ollama's MLX engine works
| Ollama version | What it does on Apple Silicon |
|---|---|
| 0.19 (March 2026) | First MLX engine, as a preview. Ollama recommended a Mac with more than 32 GB of unified memory, and the preview accelerated one model (Qwen3.5-35B-A3B). |
| 0.30 (May 2026) | llama.cpp runs GGUF models alongside the MLX engine, widening model and hardware support. |
| 0.34.4 (latest stable, September 23, 2026) | MLX runs when you pull an MLX version of a model. GGUF versions still run on llama.cpp. |
| 0.40.0-rc0 (release candidate, September 25, 2026) | One model name can carry both an MLX and a GGUF variant. On Apple Silicon, MLX is picked automatically for architectures it supports; more models are being enabled during the pre-release. |
There is no OLLAMA_MLX or OLLAMA_BACKEND environment variable, whatever some guides say, and the 32 GB figure was a recommendation in Ollama's announcement rather than a limit it enforces. Sources: Ollama's MLX announcement and Ollama's GitHub releases, checked September 29, 2026.
Which engine did I get?
On stable Ollama (0.34.4 and earlier), the tag decides: an MLX or safetensors tag runs on MLX, and anything GGUF runs on llama.cpp. On the v0.40 release candidate, Ollama tells you:
ollama ps # new RUNNER column: mlx, llamacpp or ggml
ollama show MODEL # shows the selected runner and the available ones
ollama run --runner llamacpp MODEL # force an engine for one run (also on ollama pull)
This is the check most comparisons skip. A benchmark of "Ollama" that ran a GGUF model measured llama.cpp, not MLX, and ours did exactly that (below).
Our benchmark, and what it actually measured
We simulated a realistic 25-turn Claude Code session (each turn appends about 500 tokens of tool output plus a new question) and measured time-to-first-token on identical hardware, an M3 Ultra, with the same 4-bit coding model at the same context length.
What "Ollama" means here: this April 2026 run used Ollama 0.20.4 on its llama.cpp engine, because Ollama's MLX library did not load on that install. So it compares mlx_lm.server with a tuned llama.cpp, not with Ollama's MLX engine. Updated September 29, 2026.
| Config | Median TTFT | Mean TTFT | Max TTFT |
|---|---|---|---|
| MLX default (f16 KV cache) | 422ms | 517ms | 1250ms |
| MLX + 8-bit KV cache | 320ms | 328ms | 539ms |
| Ollama 0.20.4, llama.cpp engine (flash attention + 8-bit KV) | 306ms | 326ms | 509ms |
Tuned MLX landed within about 4 percent of Ollama, well inside run-to-run noise. Neither slowed much as the conversation grew: time to first token rose about 80 to 110 ms from turn 2 to turn 25 while the prompt grew twentyfold, because both have working prefix caching. MLX's first request took 539 ms, and Ollama had three spikes near 500 ms. The one real gotcha: out-of-the-box mlx_lm.server is meaningfully slower (422ms vs 306ms here). The 8-bit KV cache closed almost the entire gap. If you benchmark raw mlx_lm.server against a tuned Ollama and conclude "MLX is slow," you are measuring the missing KV-cache flag, not the framework.
A second data point from July 2026, on Ollama 0.32.1 and the same machine: glm-4.7-flash (GGUF, llama.cpp) decoded at 77.8 tokens per second through Ollama against 59 through our own mlx_lm.server. Ollama's MLX engine was not active in that run either.
What this does and does not show: on these models, a tuned llama.cpp inside Ollama was as fast as or faster than raw mlx_lm. Third-party tests that pin both engines on one Mac find MLX ahead on dense 4-bit models and the gap shrinking or reversing on mixture-of-experts models (see Zach Rattner's 54-run comparison). We have not yet benchmarked Ollama's own MLX engine. Measure your model before switching for speed.
Since speed is close, decide on this
| What you care about | Better pick | Why |
|---|---|---|
| Simplest setup on a Mac | Ollama | One install: MLX where a model supports it, llama.cpp everywhere else. ollama pull and it runs. |
| Same setup on Macs and non-Mac machines | Ollama | One API and one model library on macOS, Linux and Windows. |
| Model coverage and updates | Ollama | Huge curated library plus Hugging Face GGUFs; new models land fast. |
| Tool-calling / agent reliability | Ollama | Most mature tool-use path across the most models. |
| A model that only exists as an MLX conversion on Hugging Face | mlx_lm | Runs any MLX-format checkpoint, not only what Ollama publishes. |
| Your own KV-cache quantization or draft-model flags | mlx_lm | Direct control over the server's flags. |
| Several large models resident on one Mac Studio | Either | Ollama's old 3-model limit is configurable with OLLAMA_MAX_LOADED_MODELS (use a positive number, not -1); memory is the real ceiling. |
The rule of thumb: default to Ollama, which now brings MLX along where it can. Reach for mlx_lm directly as a second backend when you need a model or a server flag Ollama does not offer. Speed alone is rarely the deciding factor.
You do not have to choose: the herd runs both
This is the part most comparisons miss, because they assume you pick one framework for everything. Ollama Herd runs both backends first-class. You can serve a coding model through mlx_lm.server on the node with the most memory, run general models and embeddings through Ollama elsewhere, and the router treats every backend uniformly behind one endpoint. Your client never knows or cares which framework answered.
So the real answer to "MLX or Ollama?" on a fleet is often both, chosen per model, per node, for the operational reason that fits, not for a speed difference that is not there.
pip install ollama-herd # or: brew install ollama-herd
herd # router; discovers Ollama and MLX nodes alike
Herd talks to Ollama over its HTTP API, so a model Ollama runs on its MLX engine is routed like any other; Herd does not need to know which engine Ollama picked. The mlx_lm.server backend is Apple Silicon only, while Linux and Windows machines join the same fleet through Ollama. See the Quickstart.
Frequently asked questions
Does Ollama use MLX on Apple Silicon?
Yes, for some models. Since v0.19 (March 2026) Ollama on Apple Silicon can run models on Apple's MLX framework, and the engine follows the model's format: an MLX (safetensors) version runs on MLX, a GGUF version runs on llama.cpp. The v0.40.0 release candidate (September 25, 2026) makes MLX the default on Apple Silicon wherever the MLX runtime supports the model's architecture. As of September 29, 2026 the latest stable release is v0.34.4, so on stable Ollama you still get MLX only by pulling an MLX version of a model.
How do I know if Ollama is using MLX?
On the v0.40 release candidate, ollama ps has a RUNNER column showing mlx, llamacpp or ggml, ollama show lists the runners a model offers, and ollama run --runner llamacpp MODEL forces an engine for one run. On earlier versions, check the tag you pulled: an MLX or safetensors tag runs on MLX, and anything GGUF runs on llama.cpp. There is no OLLAMA_MLX or OLLAMA_BACKEND environment variable.
Is MLX faster than Ollama on Apple Silicon?
It depends on the model, and on which engine Ollama is really running. In our April 2026 test on an M3 Ultra, tuned mlx_lm.server (8-bit KV cache) posted a median time to first token of 320ms against 306ms for Ollama 0.20.4 on its llama.cpp engine, inside measurement noise. In July 2026, on Ollama 0.32.1, glm-4.7-flash decoded at 77.8 tokens per second through Ollama's llama.cpp engine against 59 through mlx_lm.server on the same machine. Third-party tests find MLX ahead on dense 4-bit models and the gap shrinking or reversing on mixture-of-experts models, so measure your own model before switching for speed.
Should I use MLX or Ollama for local models on a Mac?
For most people, Ollama: it now runs MLX itself where a model supports it, and falls back to llama.cpp everywhere else, with the most mature tool-calling path and a simple pull and run workflow. Run mlx_lm yourself when you need a model that only exists as an MLX conversion on Hugging Face, or want direct control of server flags like KV-cache quantization and draft models.
Does Ollama Herd support both MLX and Ollama?
Yes. Ollama Herd routes to Ollama on any machine, and on Apple Silicon nodes it can also supervise mlx_lm.server and route mlx: models to it, all behind one endpoint. Because Herd talks to Ollama over its API, it does not need to know which engine Ollama picked for a model. Machines that are not Macs run Ollama with llama.cpp and join the same fleet.
Related Reading
- oMLX vs Ollama, an MLX server with batching and a persistent KV cache
- Models tested on Apple Silicon, run through both backends
- How much RAM you need, the sizing math
- Routing Engine, how the herd picks a backend and node