Ollama vs LM Studio vs vLLM on a Mac

The Apple Silicon version of the comparison, checked September 2026: engines, concurrency, agent APIs, always-on setups, and what to do with more than one Mac.

Most "Ollama vs LM Studio vs vLLM" rankings benchmark on an NVIDIA GPU and dismiss vLLM on a Mac in one line. As of September 2026 that line is out of date: vLLM now serves models on the Apple Silicon GPU through an official-org plugin, LM Studio batches concurrent requests on MLX, and Ollama is moving its Mac path to MLX. This is the comparison for someone who owns a Mac, or several, and wants to know which server to run on it.

Checked September 29, 2026 against each project's releases and docs: Ollama 0.34.4 (latest stable; 0.35.0 and 0.40.0-rc0 are published as pre-releases), LM Studio 0.4.25, vLLM 0.30.0 with vllm-metal 0.30.0. All three ship every week or two, so recheck anything load-bearing.

The short answer

  • One person, chat plus the occasional coding agent: Ollama or LM Studio. Pick LM Studio if you want a GUI for finding, downloading, and tuning models. Pick Ollama if you live in the terminal and want the widest set of tools that already know how to talk to it.
  • Several agents or several people hitting one Mac at the same time: this is where they actually differ. LM Studio and vLLM (through vllm-metal) batch concurrent requests on MLX. Ollama's MLX runner serves one request at a time, and its llama.cpp path defaults to one request per model.
  • More than one Mac: decide first whether you need to split one model across machines or spread requests across machines. They are different tools; see below.
OllamaLM StudiovLLM (with vllm-metal)
LicenseMITProprietary app, free including for work; lms CLI and MLX engine are MITApache-2.0 (vLLM and vllm-metal)
Mac requirementmacOS 14+; Apple Silicon (Intel runs CPU-only)macOS 14+; Apple Silicon onlymacOS 15+; Apple Silicon; Python 3.12
Engines on Apple Siliconllama.cpp (GGUF), MLX runnerllama.cpp (GGUF), MLX, Splash (two models)MLX, via the vllm-metal plugin
InterfaceCLI, background service, chat appDesktop app, lms CLI, headless daemonCLI and server only
Concurrent requests to one model1 by default (llama.cpp); MLX runner serialContinuous batching, 4 slots by defaultContinuous batching, paged KV cache
Default port1143412348000
Claude Code (/v1/messages)Yes, since 0.14.0Yes, since 0.4.1Yes
Codex (/v1/responses)Yes (stateless)YesYes
Built-in multi-machineNoneLM Link: remote access to another machine's serverPipeline parallelism across Macs (vllm-metal); multi-node on GPU servers

What changed in 2026

If your picture of these three comes from 2025, it is stale. The Mac-relevant changes this year:

  • January 16: Ollama 0.14.0 adds an Anthropic-compatible /v1/messages endpoint, so Claude Code can talk to it directly.
  • January 28: LM Studio 0.4.0 introduces llmster, the app's core packaged as a headless daemon, and parallel requests through continuous batching (llama.cpp engine first).
  • January 30: LM Studio 0.4.1 adds its own Anthropic-compatible /v1/messages. The next release, 0.4.2, brings continuous batching to its MLX engine (text models).
  • February: LM Studio and Tailscale launch LM Link, for using models on another of your machines as if they were local.
  • March 30: Ollama 0.19 ships an MLX engine as a preview, initially for Macs with more than 32GB of unified memory.
  • September 19: LM Studio 0.4.25 adds Inco AI's Splash engine, a model-specific engine for two Qwen models on M3-or-newer Macs running macOS 26.4 or later.
  • September 22: the vLLM project announces vllm-metal's first official release (v0.28.0): vLLM's scheduler, paged KV cache, and OpenAI-compatible server on Apple Silicon, with MLX doing the math. vLLM 0.30.0 ships the same day, and vllm-metal 0.30.0 follows a day later, installable with Homebrew.
  • September 25: Ollama 0.40.0-rc0 (a pre-release) runs MLX-supported model architectures on MLX by default on Apple Silicon.

Ollama on a Mac

Ollama is the default for a reason: ollama pull, ollama run, and a server that is simply there on port 11434. Its model library and the number of tools with an "Ollama" option in their settings are unmatched, and it is MIT-licensed end to end.

What to know on a Mac in September 2026:

  • Two engines. GGUF models run on llama.cpp with Metal. The MLX runner, a preview since 0.19, becomes the default for supported architectures in the 0.40 pre-release. Expect the MLX path to change quickly until 0.40 is stable.
  • Concurrency is its weak spot. OLLAMA_NUM_PARALLEL defaults to 1, and it only affects the llama.cpp path: the MLX runner processes requests one after another regardless (ollama/ollama#17666, open since August 11, 2026). Raising the setting on llama.cpp works, but each parallel slot needs its own context memory.
  • Agent APIs are covered with gaps. Anthropic /v1/messages works for Claude Code, but Ollama documents no token-counting endpoint and no prompt caching. Its /v1/responses for Codex is stateless: no previous_response_id.

LM Studio on a Mac

LM Studio is the best desktop experience of the three: search Hugging Face, see which quantization fits your memory, load a model with specific settings, and chat, all without a terminal. The app is closed source but has been free for work use since July 8, 2025; its lms CLI and MLX engine are MIT-licensed.

  • Apple Silicon only, macOS 14 or later. It runs both GGUF (llama.cpp) and MLX models, and you choose per model.
  • Real concurrency. Continuous batching arrived in 0.4.0 and reached the MLX engine in 0.4.2. By default a model gets 4 parallel slots ("Max Concurrent Predictions" in the load settings).
  • Splash is a new engine for exactly two models (Qwen3.8-27B and Qwen3.6-35B-A3B) that needs an M3 or newer Mac, macOS 26.4+, and at least 36GB of memory. It is fast on those models (numbers below) and irrelevant for everything else.
  • API: OpenAI-compatible chat, responses, embeddings, and completions on port 1234, plus Anthropic /v1/messages.

vLLM on a Mac

vLLM is the serving engine most GPU clusters run, and the one most comparisons rule out for Macs. There are two separate things to know:

  • Core vLLM on macOS is experimental, CPU-only, and built from source. Not what you want.
  • vllm-metal is a community-maintained plugin under the official vllm-project GitHub organization. It runs vLLM's scheduler, paged KV cache, and API server with MLX and Metal doing the execution. It needs Apple Silicon, macOS 15 or later, and native arm64 Python 3.12, and installs with Homebrew (brew install vllm-project/vllm-metal/vllm-metal after tapping the repo) or an install script.

What you get is vLLM's serving machinery on a Mac: continuous batching, prefix caching, speculative decoding, GGUF checkpoints, LoRA adapters, and structured outputs, with vision, embedding, and speech-to-text models marked experimental. What you give up is convenience: there is no GUI and no model library, you pick Hugging Face checkpoints yourself, and the plugin's first official release is from September 22, 2026. It is the right tool if you are serving several clients at once and are comfortable running a Python server.

Which engine each one runs on Apple Silicon

On a Mac, all three are front ends over two open-source engines, so "which is faster" is often "which engine and which quantization."

llama.cpp (GGUF on Metal)MLXOther
OllamaYes, the long-standing defaultPreview since 0.19; default for supported models in 0.40 pre-release
LM StudioYesYes, through its MIT-licensed mlx-engineSplash (two Qwen models, M3+)
vLLMReads GGUF checkpoints, but executes on MLXYes, through vllm-metalCore macOS build: CPU only

For a single user, the engine matters less than people expect. In our own test on an M3 Ultra (a 25-turn coding session, same 4-bit model), a tuned Ollama and a tuned mlx_lm.server landed within measurement noise on time to first token. The details are in MLX vs Ollama.

Concurrency: one user, agents, or a team

This is the section that should decide it for most people, and it depends on how many requests arrive at once.

  • One person chatting. Requests arrive one at a time. Batching does nothing for you, and any of the three is fine.
  • Coding agents. Claude Code and Codex send overlapping requests (subagents, background summaries, parallel tool calls). On Ollama's MLX runner those queue behind each other. On LM Studio and vllm-metal they share the model in one batch, so the second request does not wait for the first to finish.
  • A small team on one Mac. This is the workload vLLM was designed for: a scheduler with admission control, a paged KV cache so memory grows with actual use, and prefix caching for shared system prompts. LM Studio's batching also helps here. Ollama will serve it, but mostly one request at a time per model.

Batching does not create memory bandwidth. Every concurrent request still shares one Mac's bandwidth, so the combined throughput goes up while each individual stream gets somewhat slower.

We have not run a three-way concurrency benchmark, and we have not found a neutral one. The published Mac numbers are vendor measurements:

EngineHardware and modelResultMeasured by
Splash (in LM Studio)48GB M5 Pro, Qwen3.8-27B74 tokens/s on one short-prompt request (54 at 32K context); 170 tokens/s combined with four concurrent requestsInco AI, the engine's developer, as reported on LM Studio's blog (September 18, 2026)
vllm-metalQwen3.6-35B-A3B, 4-bit (Mac model not stated)A batch of 8 requests (about 6,000 prompt tokens in total, 20 output tokens each) finished in 3.64s vs 3.87s for mlx_lmThe vLLM project blog (September 22, 2026)

Vendor numbers use the vendor's chosen model and settings. Treat them as evidence that batching works on a Mac, not as a ranking.

API compatibility for Claude Code and Codex

Claude Code speaks the Anthropic Messages API; Codex speaks the OpenAI Responses API. All three servers now expose both, with differences at the edges:

/v1/messages/v1/messages/count_tokens/v1/responsesOpenAI /v1/embeddings
OllamaYes (0.14.0+)Not supportedYes, stateless (0.13.3+)Yes
LM StudioYes (0.4.1+)Not documentedYesYes
vLLMYesYesYesYes

The vLLM row applies to vllm-metal too, since the plugin runs vLLM's own API server. Pointing Claude Code at any of them is the same two variables, with the port changed:

export ANTHROPIC_BASE_URL=http://localhost:11434   # Ollama; LM Studio is 1234, vLLM 8000
export ANTHROPIC_AUTH_TOKEN=local                 # any non-empty value
claude

Agent quality depends far more on the model than on the server: pick a model with reliable tool calling and give it a large context window. The Claude Code and Codex guides cover model choice and context settings in detail.

Headless and always-on

A Mac mini or Mac Studio in a closet serving models for the house is a common setup. How each one handles it:

  • Ollama is a background service by design. The Mac app starts it at login; with Homebrew, brew services start ollama keeps ollama serve running.
  • LM Studio offers two routes: the headless llmster daemon (lms daemon up), or the desktop app's setting to run the LLM server on login. It restores the last server state on launch, and just-in-time loading lets clients load models on demand.
  • vLLM is only a server (vllm serve <model>), one model per process. Keeping it running across reboots, for example with a launchd job, is up to you.

Whichever you pick, turn off system sleep on the serving Mac. See Deployment for the rest of the always-on checklist.

More than one machine

Two different problems get called "clustering," and they need different tools.

Problem 1: one model is too big for any single Mac. Then the model itself has to be split across machines.

  • vLLM splits models on GPU servers with tensor, pipeline, and expert parallelism, multi-node through Ray, and its docs recommend a fast interconnect such as InfiniBand. On Macs, vllm-metal lists pipeline parallelism across multiple Macs over MLX's ring backend among its shipped features.
  • Ollama and LM Studio have no built-in way to split one model across machines.
  • Other options for Macs are covered in Ollama Herd vs exo and Ollama Herd vs MLX distributed, and the trade-off itself in routing vs sharding.

Problem 2: several Macs, each able to hold the models you use, and more requests than one Mac should serve. Then you want whole requests spread across machines.

  • Ollama has nothing built in: each Mac is its own server with its own address.
  • LM Studio's LM Link lets one machine use models loaded on another of your machines over an encrypted Tailscale mesh. That is remote access to one server at a time, not load balancing. See Ollama Herd vs LM Link.
  • Ollama Herd is a router for this case. It sits in front of Ollama and mlx_lm.server nodes, presents them as one endpoint, and sends each whole request to the node best placed to serve it (model already loaded, memory, queue depth, health). It never splits a model, so each node needs enough memory for the models it serves. It does not route to LM Studio or vLLM servers. Ollama Herd vs vLLM covers where each fits.

Decision matrix

If you...Start with
Want a GUI to find, download, and try modelsLM Studio
Live in the terminal and want the most integrationsOllama
Need an open-source license for the whole stackOllama or vLLM
Have an Intel MacOllama (CPU only); the other two need Apple Silicon
Run one coding agent at a time against a local modelAny of the three; Ollama and LM Studio are the least setup
Run several agents in parallel against one MacLM Studio or vLLM (vllm-metal)
Serve a small team from one large MacvLLM (vllm-metal) if you are comfortable running a Python server; LM Studio headless otherwise
Want Qwen3.8-27B or Qwen3.6-35B-A3B as fast as possible on an M3+ MacLM Studio with Splash
Have several Macs, each able to hold your modelsOllama or MLX on each, behind a router such as Ollama Herd
Need one model larger than any single MacA tool that splits models; see the exo and MLX distributed comparisons

Related Reading