How Ollama Herd Works

Sits between your apps and the Ollama instances on your Macs, Linux servers, and Windows PCs. Point everything at one URL. The router handles the rest.

Two Commands, Zero Config

On your router machine:

pip install ollama-herd
herd

On each machine running Ollama (macOS, Linux, or Windows):

herd-node

Each node discovers the router via mDNS and starts sending heartbeats. No config files, no YAML, no Docker, no Kubernetes.

Need to skip mDNS? Use herd-node --router-url http://router-ip:11435

Ollama Herd dashboard Fleet Overview tab, two nodes visible: a 512 GB Mac Studio running Ollama and MLX backends with five models loaded plus image generation, speech-to-text, embedding, and vision services; and a 128 GB MacBook Pro M4 with Gemma and Qwen3-Coder models loaded.
Within ~60 seconds of herd-node starting, the router picks up each device and you see it here, memory, cores, loaded models, available services. One view across the whole fleet.

What Happens When a Request Arrives

Every model request, whether a chat completion, a vision prompt, or an Ollama embedding, on any of the four wire protocols, passes through a five-stage pipeline. Image generation, speech-to-text, and the native text-embedding service take a simpler path: the router picks the least busy node that offers the service, by free memory and CPU load.

Conceptual: the request path in v0.9.5
Claude Code
Codex
Open WebUI
Any OpenAI or Ollama client
one URL, four wire protocols
Anthropic Messages/v1/messages
OpenAI Responses/v1/responses
OpenAI Chat Completions/v1/chat/completions
Ollama API/api/chat, /api/generate
translated to one internal request
Herd router, port 11435 Nodes find it by mDNS, then send it a heartbeat every 5 seconds.
  • 1. Eliminate nodes that cannot serve this model right now
  • 2. Score the survivors on 8 signals; the highest score wins
  • 3. Queue on the winner's own node:model queue
  • 4 and 5, in the background every 5 seconds: pre-warm the runner-up and rebalance deep queues
each request goes to one node
Winning nodeOllama, or MLX on a Mac. Streams the whole response back.
Runner-up nodePre-warmed (model loaded, no request sent) once the winner has 3 or more requests queued or running
Other nodesIdle for this request; keep sending heartbeats
Every client talks to one router on port 11435, and each request runs whole on exactly one machine; the router never splits a request or a model across nodes. Heartbeats every 5 seconds keep the router's view of each machine current enough to score it.
1

Elimination

The router immediately removes nodes that can't serve the request: offline or not heartbeating, model not on disk, not enough memory to load the model (judged against system RAM), critical memory pressure, or hard-paused by opt-in capacity learning (in a meeting on macOS, sustained heavy CPU, low availability, or still in its first 7 days of observation). If nothing survives, the request enters a holding queue instead of failing: the router rescores it every 2 seconds for up to 30 seconds as node states change.

Stage 1 as the v0.9.5 scorer runs it
Every node the router knows about
a node is removed if any of these is true
Offlineno heartbeat for 30 seconds
Critical memory pressureas macOS or Linux reports it
Model not therenot loaded and not on disk
Won't fitnot loaded, and larger than free system RAM (never GPU VRAM)
Paused by capacity learningopt-in only: meeting (macOS), CPU over 85% for 2 minutes, availability under 0.2, or first 7 days
what is left
Survivors are scoredstage 2, eight signals
Nobody survivedholding queue: rescored every 2 seconds for up to 30 seconds. A model no node has is auto-pulled instead.
Elimination is a set of hard gates, not a score: a node that fails any one of them is never considered, however well it would have scored.
2

Scoring

Every surviving node gets scored across 8 weighted signals. A hot model on an idle machine with plenty of memory headroom, say a Mac Studio or a Linux server, scores 80+. A cold model on a busy laptop with rising CPU usage scores under 20. The highest score wins.

3

Queue and Execute

The winning node receives the request in its dedicated queue. Each node+model pair has its own queue with dynamic concurrency, the router knows how many parallel requests each device can handle without degrading performance.

4

Pre-Warm

If the primary node's queue is getting deep (3 or more requests queued or running, checked every 5 seconds), the router proactively loads the same model on the runner-up node. By the time the next request arrives, it's already hot.

5

Rebalance

A background process runs every 5 seconds, moving queued requests from overloaded nodes to nodes with spare capacity, but only where the model is already loaded, avoiding cold-load cascades.

Scoring Signals

Every surviving node gets scored across 8 weighted signals:

Signal What It Measures Weight
Model warmth Is the model already loaded (hot) or needs loading (cold)? +50 loaded, +30 unloaded in the last 30 minutes, +10 on disk
Memory fit How comfortably does the model fit in available system memory? Up to +20
Queue depth How many requests are already waiting on this node? Up to -30
Estimated wait time Using real latency history, how long until this request starts? Up to -25
Role affinity Does this machine match the model's weight class? Up to +25 (scaled by memory bandwidth)
Availability trend Is this device freeing up or getting busier? (Only nodes with capacity learning turned on) Up to +10
Context fit Can this node handle the requested context size? +15 to -15 (negative when the prompt would overflow)
Session affinity Does this node already hold the conversation's warm cache? Up to +20
Stacked bar chart of one gpt-oss:120b request scored by the v0.9.5 scorer. Mac Studio M3 Ultra, with the model loaded and 4 requests queued: model warmth +50, memory fit +20, role affinity +25, context fit +15, session affinity +6.7, queue depth -20, estimated wait -0.3, total 96. MacBook Pro M4 Max, idle but with the model only on disk: warmth +10, memory fit +15, role affinity +18.7, availability trend +7, total 51. Linux desktop with 64 GB: eliminated, the 65 GB model does not fit in 52 GB of free RAM.
A queue of 4 costs this Mac Studio 20 points, but a loaded model is worth 50, so it still beats an idle MacBook that would have to load 65 GB first, 96 to 51. Illustrative scenario; scores computed by the Ollama Herd v0.9.5 scorer (ScoringEngine.score_request, default settings).

The Fleet Gets Smarter Over Time

Ollama Herd isn't static. It learns:

  • Latency tables track per-node, per-model response times in SQLite. After a few days, the scoring engine knows exactly how fast each machine runs each model.
  • Capacity learner (opt-in, herd-node --learn-capacity) builds a 168-slot weekly behavioral model (one slot per hour). After a month, it knows your laptop is busy Tuesday mornings and the Windows desktop is idle on weekends.
  • Meeting detection (macOS, part of capacity learning) pauses nodes when cameras or microphones are active. No inference competes with your Zoom calls.
  • App fingerprinting (part of capacity learning, all platforms) classifies your workload (idle/light/moderate/heavy) without reading app names. Heavy workloads reduce a node's memory ceiling, shifting requests elsewhere.
  • Context optimizer tracks actual token usage per model and auto-adjusts context windows. Most models allocate far more context than they use, Herd right-sizes them to free memory for additional models.

All state persists across restarts. A fleet running for a month makes better routing decisions than one running for a day.

Multimodal Routing

The router handles five model types, each routed to the right node:

Model Type Protocol Example
LLM inference OpenAI + Ollama API Llama 3, Qwen 3, DeepSeek
Embeddings Ollama API nomic-embed-text
Image generation Custom API + OpenAI images FLUX via mflux, SD 3.5 via DiffusionKit (Apple Silicon); Ollama native image models
Speech-to-text Custom API Qwen3-ASR via MLX (Apple Silicon)
Vision OpenAI + Ollama API Gemma3, LLaVA, Llama3.2-Vision

LLM, embedding, and vision requests work on every platform, NVIDIA GPU or not. Speech-to-text and mflux/DiffusionKit image generation need Apple Silicon; Linux and Windows nodes simply don't advertise them, and the router sends those requests to a Mac that does. Ollama's native image models go through the normal Ollama path.

API Compatibility

Point any existing tool at the router, no code changes needed:

from openai import OpenAI

client = OpenAI(base_url="http://router-ip:11435/v1", api_key="not-needed")
response = client.chat.completions.create(
    model="llama3.3:70b",
    messages=[{"role": "user", "content": "Hello!"}],
    stream=True,
)

Or use the Ollama API directly:

curl http://router-ip:11435/api/chat -d '{
  "model": "llama3.3:70b",
  "messages": [{"role": "user", "content": "Hello!"}]
}'

The router speaks OpenAI Chat Completions, OpenAI Responses (Codex), Anthropic Messages (Claude Code), and native Ollama. Works with Open WebUI, LangChain, CrewAI, AutoGen, Aider, Continue.dev, LlamaIndex, LiteLLM, and any other OpenAI-compatible client.

What You Get

  • Zero wait time, requests run on the first available machine, not in a queue
  • Zero model swapping, every model stays loaded on its home machine
  • Zero babysitting, memory pressure handled automatically, plus meetings and heavy workloads once capacity learning is on
  • Zero client changes, point your existing tools at one URL
  • Compounding intelligence, the longer it runs, the better the routing decisions
Already running Ollama on one machine? See how Herd upgrades a single Ollama instance into a fleet, or compare with cloud APIs and Open WebUI. View all comparisons →

See the pipeline run on your own fleet

Two commands, zero config. Watch requests route themselves in the live dashboard.