llama-swap vs Ollama vs llama.cpp Router Mode
Three ways to run more models than fit in memory on one machine, what each one controls, and what changes when you add a second machine.
"More models than memory" is the problem all three tools solve on one machine. Ollama loads and evicts models by itself. llama-swap starts and stops model servers according to rules you write. llama.cpp's own server gained a router mode in December 2025 that loads models on demand and evicts the least recently used. They differ in who decides, which engines they run, and how much control you get. None of them makes a cold load free.
Checked September 29, 2026 against each project's releases and docs: llama-swap v260 (September 26), llama.cpp build b11260, Ollama 0.34.4 (latest stable), and Ollama Herd 0.9.5. llama-swap ships about twice a week, so check its changelog for anything newer.
The short answer
- You want
ollama pulland nothing to configure: Ollama. It checks memory before loading, unloads idle models after five minutes, and speaks the Ollama, OpenAI, and Anthropic APIs. - You run llama.cpp and want the fewest moving parts: llama.cpp router mode. Point
llama-serverat a folder of GGUF files and it loads what is asked for, keeping up to four models. - You want exact control, or engines besides llama.cpp: llama-swap. Per-model launch flags, rules for which models may run together, idle timeouts, API keys, a web UI, and metrics, for llama.cpp, vLLM, stable-diffusion.cpp, whisper.cpp, ComfyUI, and more.
- The same few models keep evicting each other all day: no swapper fixes that. A second machine that keeps one of them loaded does. That is a placement problem, covered below.
| Ollama | llama.cpp + llama-swap | llama.cpp router mode | |
|---|---|---|---|
| Who decides what unloads | Ollama: memory fit and a 5-minute idle timer | You: group or matrix rules, optional idle TTL | Least recently used, once --models-max (default 4) is reached |
| Checks memory before loading | Yes | No, you declare what fits | No, the limit is a model count |
| Engines | Ollama's llama.cpp-based engine and its MLX runner | Anything it can launch: llama.cpp and forks, vLLM, stable-diffusion.cpp, whisper.cpp, ComfyUI, TabbyAPI | llama.cpp |
| Getting models | ollama pull | Download GGUF files yourself | Hugging Face cache, a folder, or a preset file |
| APIs | Ollama, OpenAI, Anthropic (subset) | OpenAI, plus Anthropic passed through to the server | OpenAI, Anthropic |
| Auth | None on the local server | API keys | API key |
| UI and metrics | Desktop app and CLI | Web UI with logs and captures, Prometheus | Web UI, /metrics |
| More than one machine | No | Peers you list in config | No |
| License | MIT | MIT | MIT |
What "model swapping" actually costs
Every tool here does the same physical thing when the model you asked for is not in memory: read gigabytes of weights from disk, allocate the KV cache, and only then start on your prompt. That cold load takes seconds for a small model and can take minutes for a large one. The tools differ in when they decide to pay it and what they throw out to make room, not in how fast the load itself is. Load speed comes from the engine, your disk, and whether the file is still in the OS page cache.
Two consequences worth keeping in mind. First, if your workload alternates between two models that cannot fit together, every switch pays a full load no matter which tool you pick. Second, the thing that removes the cost is keeping a model resident somewhere, which on one machine means having enough memory, and across machines means sending each request to where its model already lives. Our cold starts guide goes deeper on the latency side.
How Ollama loads and unloads models
Ollama manages memory for you. When a request arrives for a model that is not loaded, it estimates whether the model fits alongside what is already loaded; if not, it unloads idle models to make room, and if it still cannot fit, the request waits in a queue. Idle models unload after five minutes by default.
OLLAMA_KEEP_ALIVE(orkeep_aliveper request) sets the idle timer.-1keeps a model loaded until memory is needed.OLLAMA_MAX_LOADED_MODELScaps how many models stay loaded (default 3 per GPU, 3 on CPU).OLLAMA_NUM_PARALLEL(default 1) andOLLAMA_MAX_QUEUE(default 512) control concurrency and queueing. Our concurrency guide covers these in detail.
The trade is control. You do not choose launch flags per model the way you do with llama-server, and Ollama stores models as hashed blobs rather than GGUF files you can point other tools at. Those two points drive most of the moves from Ollama to llama.cpp.
How llama-swap works
llama-swap (about 5.8K GitHub stars, MIT, one active maintainer) is a single Go binary with one YAML file. Each model entry is a command that starts a server (usually llama-server with whatever flags you like) plus a health endpoint. When a request names a model, llama-swap stops whatever conflicts with it, starts the right server, waits until its health check returns 200, then forwards the request.
- You write the rules. By default one model runs at a time. Groups let you mark models that can run together or should never be unloaded, and a matrix mode (added April 2026) lets you express combinations like "the coder plus an embedder, or the big chat model alone," with eviction costs so it drops the cheapest set. llama-swap does not measure free memory: if your rules allow two models that do not fit, the second load fails.
- Idle unload is optional.
ttlper model orglobalTTL; both default to never. - Any server it can launch. llama.cpp and its forks are the best supported, but it also runs vLLM, stable-diffusion.cpp, whisper.cpp, audio servers, ComfyUI, and TabbyAPI, and routes image and audio endpoints to them.
- APIs: OpenAI chat, completions, responses, embeddings, audio, and images, plus Anthropic
/v1/messages, which it passes through to a server that speaks it (llama-server does). It has no Ollama API. - Operations: API keys, a web UI (playground, request captures, live logs, manual load and unload), Prometheus metrics, a per-model FIFO queue with priorities, and Docker images for CUDA, Vulkan, and more. Binaries for Linux, macOS, Windows, and FreeBSD.
- Selectors (July 2026) turn one model name into a choice made per request:
warmpicks the first listed target that is already loaded,spilloverfills targets in order up to a request count,pinalways picks the first.
The common complaints are the flip side of the control: you write and maintain the YAML, you pick quantizations by hand, and it is one more layer to debug when a server will not start. The project has since added a large knowledge base and an in-UI help agent, which answers the older "docs are thin" complaint.
llama.cpp router mode: the built-in alternative
Since December 2025, llama-server started without a model runs as a router. It finds models in the Hugging Face cache, a --models-dir folder, or an .ini preset file, starts each model in its own process when a request names it, and keeps up to --models-max models loaded (default 4, evicting the least recently used). --sleep-idle-seconds unloads after inactivity, and presets can mark models to load at startup.
For a llama.cpp-only setup this covers the basic need with no extra software, which is why "does router mode replace llama-swap?" has been the most common question about llama-swap this year. Users split on it: some switched and are happy, others stay on llama-swap for per-model flags and timeouts, the matrix, other engines, and the UI, logs, and metrics. Two limits to know: the cap is a count, not a memory budget, and router mode is newer, so expect its options to keep changing.
Not to be confused with llama.cpp's RPC server, which offloads computation to other machines. Its own README calls it a proof of concept that is fragile and insecure, not for open networks.
When llama-swap is the right call
- You run more than one engine: a llama.cpp chat model, a vLLM model, stable-diffusion.cpp for images, whisper.cpp for speech, behind one endpoint.
- You want per-model launch flags: GPU layer splits across two cards, context size, draft models, specific builds.
- You want to declare exactly which models coexist, rather than trusting an automatic policy.
- You need API keys, request captures, logs, and metrics on a shared box.
When Ollama is the right call
- You want
ollama pulland a server that is simply there, with nothing to write. - You want memory-aware eviction by default instead of rules you maintain.
- Your tools have an "Ollama" option in their settings. Many do, and llama-swap does not speak Ollama's API.
- You are on a Mac and want Ollama's MLX runner without managing separate servers (our Mac server comparison covers that choice).
When one machine is not enough
If two models you use all day cannot share memory, the fix is to stop swapping: keep one loaded on another machine and send each request to where its model already is. At that point the question changes from "which tool swaps best" to "which tool places requests," and the answer depends on what your machines run.
- Machines running llama.cpp: llama-swap's peers forward requests for listed models to another llama-swap, any compatible server, or a hosted API, and a
spilloverorwarmselector can prefer local capacity or an already-loaded target. Peers are a static list: you name each machine and the models it serves, and llama-swap cannot start or unload anything on the other side. - Machines running Ollama: Ollama Herd finds them on the local network with mDNS and scores every request across all of them: whether the model is already loaded (or was recently), whether it fits in the machine's free system RAM, queue depth, expected wait, and whether the context fits. It asks Ollama to keep models loaded and can pin models so they are reloaded if evicted.
| llama-swap peers | Ollama Herd | |
|---|---|---|
| Backends | Another llama-swap, any OpenAI-compatible server, hosted APIs | Ollama on each machine (plus mlx-lm servers Herd launches on Macs) |
| Finding machines | You list each peer and its models | mDNS discovery, no config |
| Choosing a machine | Config order: warm-first or spillover by request count | Scored per request: loaded model, RAM fit, queue depth, wait time, context fit, and more |
| Remote model lifecycle | None: it cannot load or unload on a peer | Keeps models loaded, reloads pinned models, pre-warms a second node when a queue builds |
| Auth | API keys; private peer links via Tailcat | None beyond an optional key on the Anthropic endpoint; built for trusted networks |
The two do not stack today. Herd nodes talk to Ollama's native API (model lists, loaded models, chat and generate), which neither llama-swap nor llama-server implements, so a llama.cpp box cannot join a Herd fleet. If you have both kinds of machines, run each as its own endpoint. Herd reads system RAM, not GPU memory, so on an NVIDIA machine choose models that fit the card.
Frequently asked questions
Does llama.cpp router mode replace llama-swap?
For many llama.cpp-only setups, yes. Since December 2025, llama-server can load models on demand from a folder or preset file and keep up to four loaded, evicting the least recently used. llama-swap still does more: it runs engines other than llama.cpp (vLLM, stable-diffusion.cpp, whisper.cpp, ComfyUI), lets you declare exactly which models may run together, sets per-model idle timeouts, and ships a web UI with logs and metrics.
Is llama-swap faster than Ollama?
llama-swap adds no inference of its own; speed comes from the server it launches. Users who moved report faster llama.cpp builds and more control over flags such as GPU layer splits, but results depend on model, build, and hardware. Measure your own model on your own machine before switching.
Can I put llama-swap in front of Ollama?
It is not what llama-swap is built for. It swaps by starting and stopping server processes, and its maintainer has said it does not drive other servers' own load and unload APIs. Ollama already loads and unloads models on its own, so the two would be managing the same memory twice.
Does llama-swap work across multiple machines?
Yes, within limits. Its peers feature forwards requests for listed models to another llama-swap or any compatible server, and a spillover selector can fill local capacity before using a peer. Peers are a static list: you name each machine and the models it serves, and llama-swap cannot start or unload anything on the other side.
Can Ollama Herd use llama-swap or llama.cpp as a node?
Not today. Herd nodes talk to Ollama's native API (model lists, loaded models, chat and generate), and neither llama-swap nor llama-server implements it. If your machines run Ollama, Herd routes across them; if they run llama.cpp, llama-swap peers or a gateway like LiteLLM are the closer fit.
Which one keeps my models loaded?
Ollama unloads idle models after five minutes unless you change OLLAMA_KEEP_ALIVE. llama-swap keeps a model until something else needs its slot, unless you set a ttl. Router mode keeps up to --models-max models and evicts the least recently used. Herd asks Ollama to keep models loaded and sends each request to a machine that already has the model in memory.
I have more models than memory. What should I do?
On one machine, pick the tool whose swapping rules match how you work, and accept a cold load on each swap. If the same few models alternate all day, a second machine that keeps one of them loaded removes the swap entirely; that is a placement problem rather than a swapping one.
Related reading
- Ollama Herd vs Ollama proxies, llama-swap alongside LoLLMs Hub, ollamaMQ, and others
- Ollama concurrent requests, the settings behind Ollama's loading and queueing
- Cold starts, what a cold load costs and how routing avoids it
- Ollama load balancing, the options for spreading requests across machines
- Ollama Herd vs LiteLLM, the gateway many people put in front of llama-swap