oMLX vs Ollama on a Mac
Which Mac server to run for Claude Code and Codex: oMLX's cached, batched MLX serving or Ollama's simplicity, with LM Studio and mlx_lm for context.
oMLX went from first release in February 2026 to over 22,000 GitHub stars by September by attacking one specific pain: a coding agent on a Mac re-reading its entire conversation every turn. It is an MLX inference server with a menu-bar app, continuous batching, and a KV cache that spills to your SSD and survives restarts. Ollama is the default everyone starts with. This guide compares them on the things that matter for agents, with sources, and says where each one falls short.
Checked September 29, 2026 against oMLX 0.6.4 (latest stable, August 29) and 0.7.0rc1 (September 24, which GitHub labels "Latest" and Homebrew installs), Ollama 0.34.4, mlx-lm 0.31.3, LM Studio 0.4.25, and vllm-mlx 0.5.0. All of them ship weekly; recheck before relying on a detail.
The short answer
- One Mac, long Claude Code or Codex sessions, MLX models: oMLX is built for exactly this. Its advantage is the second turn onward: it reuses the conversation's cached context instead of recomputing it, even after a restart, and it batches concurrent requests.
- You want the widest tool support, GGUF models, and the least fuss: Ollama. It works on Linux and Windows too, and far more apps have an "Ollama" setting. Its MLX runner, though, answers one request at a time.
- You want a GUI for finding and tuning models: LM Studio, which also batches and now saves prompt cache to disk on its MLX engine. Our Mac server comparison covers it in depth.
- You have more than one machine: none of these spreads requests across machines. That is a router's job, covered below.
What oMLX is
oMLX is an Apache-2.0 inference server for Apple Silicon, written largely by one developer (Jun Kim) with many outside contributors. It started from vllm-mlx 0.1.0 and has diverged a long way since. It runs MLX models only (no GGUF), including text models, vision models, OCR models, embeddings, and rerankers, and it has its own quantizer.
- Install: a DMG with in-app updates (it also installs an
omlxcommand), a Homebrew tap, or from source. It is not on PyPI. A plain source install skips the precompiled custom kernels for some model families, which then fall back to much slower paths; the DMG includes them. - Requirements: Apple Silicon and macOS 15 or later.
- Server: listens on
127.0.0.1:8000, with an admin panel at/admin. From 0.7.0rc1 it refuses to listen on a network address without an API key unless you override that by hand. - Model management: several models at once with least-recently-used eviction, pinning, per-model idle timeouts, and a memory ceiling (system RAM minus 8GB by default).
- Agents:
omlx launch claudeconfigures Claude Code, and the admin panel sets up Codex, OpenCode, OpenClaw, and others. It also adjusts the token counts it reports so Claude Code's auto-compact fires at the right point for smaller-context models.
Why coding agents are slow on local models
An agent like Claude Code resends the whole conversation every turn: system prompt, tool definitions, every file it read, every command output. By turn 20 that can be tens of thousands of tokens. Before the model writes a single new token it has to process all of them (prefill), and on a Mac, prefill is the slow part. If the server has to recompute that history every turn, each turn gets slower as the session grows.
Every serious server now tries to reuse the part of the prompt it has already seen. The differences are where that cache lives, how long it survives, and how well it copes when the prompt changes in the middle (which agents do constantly, for example when they trim old tool output).
oMLX's tiered KV cache, and who else does this now
oMLX keeps the cache in fixed-size blocks, shares blocks between requests with the same prefix, and writes blocks to the SSD so they outlive memory pressure and even a server restart. When a new request shares a prefix with anything cached, in RAM or on disk, it restores those blocks instead of recomputing them. The 0.7.0rc1 release adds partial-block reuse; the maintainer's own test on a 13.4K-token conversation cut the next turn's prefill from 1,174 tokens to 37.
The maintainer's headline claim is follow-up turns on long agent contexts dropping from 30 to 90 seconds to 1 to 3 seconds. Treat that as the upper end of what the design allows: it depends on the model, the Mac, and how much of the prompt actually repeats.
This is no longer unique to oMLX, which popularized it. vllm-mlx has an SSD cache tier, LM Studio's MLX engine saves prompt cache to disk (since June 2026), and Rapid-MLX advertises a RAM plus SSD cache. Ollama and mlx_lm.server reuse prefixes in RAM only, so a restart or eviction means recomputing.
The cost is disk: oMLX writes cache blocks as a normal part of operating, with an automatic size budget whose rules changed in 0.7.0rc1. If disk space or SSD writes matter to you, set a fixed limit.
Concurrency: batching versus one at a time
oMLX uses continuous batching and accepts up to 8 concurrent requests by default, so a second agent, a subagent, or a background title request does not wait behind the first. On Ollama, OLLAMA_NUM_PARALLEL defaults to 1 on the llama.cpp path, and the MLX runner processes requests one after another. LM Studio batches with 4 slots by default. This matters as soon as more than one thing talks to the same Mac; see our concurrency guide for Ollama's settings.
API support for Claude Code and Codex
| oMLX | Ollama | mlx_lm.server | LM Studio | |
|---|---|---|---|---|
| OpenAI chat completions | Yes | Yes | Yes | Yes |
Anthropic /v1/messages (Claude Code) | Yes, with token counting | Yes, no token counting | No | Yes |
OpenAI /v1/responses (Codex) | Yes, stateful (stores responses) | Yes, stateless | No | Yes |
Ollama /api/* | No | Yes | No | No |
| Embeddings and rerank | Both | Embeddings | No | Embeddings |
No Ollama API means tools that only know how to talk to Ollama need an OpenAI-compatible setting to use oMLX. Claude Code also changes its request format often, and oMLX has had to follow: several "422" errors after Claude Code updates were reported and fixed during 2026, and one was still open in late September (#3754).
Stability and memory: what the issue tracker says
oMLX moves fast, and its tracker shows it: about 1,000 of its roughly 2,100 issues are open. The recurring themes are specific, so here they are rather than a verdict:
- Memory pressure on 32 to 48GB Macs. The most-discussed issue is a kernel panic during Claude Code sessions on a 32GB M4 with a 35B model (#300); the fault is in macOS's GPU memory layer, so it is not necessarily oMLX's bug alone. Others report the memory guard refusing long contexts that llama.cpp handled (#2321) and the OS killing the process under pressure (#702).
- Cache edge cases. One report has a Claude Code session periodically re-prefilling everything despite caching (#2333); another has two identical cached prefills wedging the scheduler on an older release (#2330).
- macOS 27. Long-context inference reported ten times slower than on macOS 26.6 (#1835).
If you have 64GB or more and use models well inside that, most of these will not touch you. On a smaller Mac, size the model conservatively; our Mac memory guide has the arithmetic.
Benchmarks, and why they disagree
There is no neutral, current benchmark of oMLX against Ollama on an agent workload. The published numbers come from single testers on different Macs, models, and versions, and they point in different directions:
- A May 2026 test on an M3 Max 64GB found plain mlx-lm and vllm-mlx faster than oMLX 0.3.8 on decode speed, and oMLX well ahead of Ollama (zephel01). That version is old.
- A September 2026 test on an M5 Max measured one model on 32K-token prompts going from 15 to 78 tokens per second between oMLX 0.3.8 and 0.6.4 (jacar.es), which is why older comparisons understate it.
- A 48GB Mac mini test found oMLX using more memory than Ollama over a five-prompt session and crashing on a 33GB model, and recommended LM Studio as the default (Serve No Master).
The pattern that holds across them: oMLX's edge is warm follow-up turns and concurrency, not first-prompt speed. For a fresh single prompt, the engine and quantization matter more than the server; in our own test, a tuned Ollama and a tuned mlx_lm.server landed within noise of each other on time to first token (MLX vs Ollama). Measure your own model and your own sessions.
oMLX vs LM Studio, mlx_lm, and vllm-mlx
| oMLX | Ollama | LM Studio | mlx_lm.server | vllm-mlx | |
|---|---|---|---|---|---|
| Platforms | Apple Silicon, macOS 15+ | macOS, Linux, Windows | macOS 14+, Windows, Linux | Apple Silicon | Apple Silicon |
| Model format | MLX | GGUF; MLX runner | GGUF and MLX | MLX | MLX |
| Concurrency | Continuous batching, 8 by default | 1 by default; MLX runner serial | Continuous batching, 4 by default | Batching, one model per process | Continuous batching (opt-in) |
| Cache across requests | RAM + SSD, survives restart | RAM | RAM + disk (MLX engine) | RAM | RAM + SSD tier |
| Several models at once | Yes: LRU, pin, TTL | Yes | Yes | No | Yes |
| Interface | Menu-bar app, CLI, web admin | Service, CLI, app | Desktop app, CLI, headless daemon | CLI | CLI |
| License | Apache-2.0 | MIT | App proprietary (free); CLI and engine MIT | MIT | Apache-2.0 |
For a beginner who wants a GUI, LM Studio is the easier start. For MLX speed with agents and a polished Mac app, oMLX is the one built around that. For the thinnest possible server, mlx_lm.server.
More than one machine: where a router fits
oMLX and Ollama are both servers: each makes one machine answer requests. With a second machine, the question changes from "which server" to "which machine should take this request." oMLX's own multi-Mac mode splits a single model across Macs, is experimental, and is only available in source builds; it does not spread separate requests across machines. That distinction is the subject of our routing vs sharding guide.
Ollama Herd spreads requests. It puts one endpoint in front of Ollama on every machine you own (Macs, Linux boxes, and Windows PCs), finds them automatically, and sends each request to the machine that already has the model loaded and the memory to run it. It keeps a coding session on the machine that holds its cache, and trims long Claude Code histories before they reach the model.
What Herd does not do is replace a server's engine. It does not batch requests on a single Mac or keep a KV cache on your SSD; that is oMLX's territory. And it does not route to oMLX today: Herd's MLX support starts and manages mlx_lm.server itself, and Codex traffic through Herd is served by Ollama-backed models. You can run oMLX on one Mac and Herd plus Ollama across your other machines side by side. If you want a router in front of oMLX specifically, Olla lists oMLX as a supported backend.
If one Mac and one agent is your whole setup, you do not need a router. oMLX or plain Ollama is enough.
Frequently asked questions
Is oMLX faster than Ollama?
On follow-up turns of a long conversation, often yes, because it reuses cached context instead of recomputing it, and it handles concurrent requests where Ollama's MLX runner goes one at a time. On a single fresh prompt it is not reliably faster; published tests disagree, and oMLX has sped up a lot between versions. Test your own model on your own Mac.
Does oMLX work with Claude Code and Codex?
Yes. It serves the Anthropic Messages API, with token counting, for Claude Code and the OpenAI Responses API for Codex, including stored responses. omlx launch claude sets up Claude Code, and the admin panel configures Codex and other agents. Claude Code changes its request format often, and oMLX has had to follow with fixes.
Does oMLX have an Ollama-compatible API?
No. It exposes OpenAI- and Anthropic-style endpoints on port 8000, not Ollama's /api/chat or /api/tags. Tools that only know how to talk to Ollama need an OpenAI-compatible setting to use it.
Is the SSD cache unique to oMLX?
No, though oMLX made it popular. vllm-mlx has an SSD cache tier, LM Studio's MLX engine saves prompt cache to disk, and Rapid-MLX advertises a RAM plus SSD cache. oMLX's version is the most integrated into a Mac app.
Will oMLX fill my disk?
It writes KV cache blocks to disk as part of its design, with an automatic size budget and an option for a fixed limit. The sizing rules changed in 0.7.0rc1. If disk space or write volume matters to you, set a fixed limit.
Can I use oMLX with Ollama Herd?
Not as a supported setup today. Herd's MLX backend starts and manages mlx_lm.server itself, and Codex traffic through Herd is served by Ollama-backed models. You can run oMLX on one Mac and Herd plus Ollama across your other machines side by side, but Herd will not route to oMLX.
What do I need for more than one machine?
Decide whether you want to split one model across machines or spread requests across machines, each running its own models. oMLX has an experimental, source-build-only mode for the first. The second is what Ollama Herd does, across Macs, Linux, and Windows machines running Ollama.
Related reading
- Ollama vs LM Studio vs vLLM on a Mac, the other Mac servers in depth
- MLX vs Ollama, which engine you are actually running, and our time-to-first-token test
- Claude Code with local models, setup and context length
- Codex with local models, the Responses API path
- How much Mac memory you need, sizing models to avoid memory pressure
- Routing vs sharding, splitting one model versus spreading requests