Ollama Herd vs Single Ollama Instance

Single Ollama is perfect for one machine. The moment you add a second device, Ollama Herd turns your idle hardware into a unified AI fleet with zero-config discovery, intelligent routing, and model-aware load balancing.

What is Single Ollama?

Ollama is an open-source tool for running large language models locally on your machine. You install it, run ollama serve, and interact with models through a local API on port 11434. It handles model downloading, quantization, GPU acceleration, and memory management, all on a single device. Simple, fast, and free.

What is Ollama Herd?

Ollama Herd is an open-source smart AI router that turns the machines you already own (Apple Silicon Macs, Linux servers, and Windows PCs, with or without NVIDIA GPUs) into one endpoint. It routes LLMs, embeddings, image generation, speech-to-text, and vision with 8-signal scoring, mDNS auto-discovery, and an 8-tab real-time dashboard. Apple Silicon Macs get extras: an MLX backend, native image generation, and speech-to-text. Two commands to set up, zero config files. pip install ollama-herd or brew install ollama-herd.

Overview

Most people start with Ollama on one machine. It works great, ollama run llama3.3:70b and you're talking to a local model in seconds. No cloud, no API keys, no subscription.

The question isn't whether single Ollama works. It does. The question is what happens when:

  • You have a second machine sitting idle
  • You need to run multiple models concurrently
  • An agent pipeline makes 20 LLM calls in sequence and each one queues behind the last
  • You want image generation AND chat AND embeddings but one machine can't do all three well
  • One machine can't keep up with a long coding session plus everything else, and inference slows to a crawl
  • Someone else in your house/office wants to use AI at the same time

Ollama Herd doesn't replace Ollama. It connects multiple Ollama instances into one intelligent endpoint.

Conceptual: where each tool sits (a tool can span layers)
Clientschat UIs, coding agents, apps
Open WebUI Claude Code Codex Your apps and agents
Gateway / edgeauth, rate limits, cloud providers
Inference serverruns the model on one machine
One model, many machinessplits a model too big for one
Hardwarethe machines themselves
Macs Linux servers Windows PCs NVIDIA GPU servers
Herd does not replace Ollama: it sits one layer above it, and each request still runs on one machine's own Ollama (or mlx_lm.server on a Mac).

Feature Comparison

Feature Single Ollama Ollama Herd
Setupollama serveherd + herd-node (2 commands)
Devices1Unlimited (auto-discovered via mDNS)
Model routingManual (you pick the model)Automatic (best node selected per request)
Concurrent requestsQueued on one machineDistributed across fleet
Load balancingNone8-signal scoring (model loaded, memory, queue, latency, affinity, availability, context, session affinity)
FailoverNone, if it's down, it's downAuto-retry on different node before first chunk
Model fallbacksNoneClient-specified backup models tried automatically
Queue managementSingle queuePer node:model queues with rebalancing
Cold-load avoidanceNone, loads whatever you ask forPrefers nodes that already have the model loaded
Memory awarenessNone, loads until OOMScores by memory fit, dynamic ceiling
Meeting detectionNonePauses inference when camera/mic active (macOS)
Capacity learningNone168-slot weekly behavioral model per device
Image generationOllama native models onlymflux (FLUX.1) + DiffusionKit + Ollama native, capability-routed
Speech-to-textNot supportedQwen3-ASR routed to capable nodes
EmbeddingsSingle nodeRouted to nodes with embedding models
DashboardNone8-tab real-time UI with SSE
Health monitoringNone30+ automated health checks
Request tracingNoneEvery request traced to SQLite
Per-tag analyticsNoneTag requests, see usage by app/tool
Context optimizationManual num_ctxTracks actual usage, auto-adjusts to free KV cache memory
BenchmarkingNoneSmart benchmark across all model types
API compatibilityOllama API, plus OpenAI-compatible and Anthropic Messages endpointsThe same formats, served from one address for every machine
Thinking modelsManual configAuto-detected, token budget inflated 4x

When Single Ollama Is Enough

Be honest, not everyone needs Herd:

  • One machine, one user, light usage. If you run a few prompts a day on a machine with enough RAM for your model, single Ollama is perfect. No overhead, no complexity.
  • Single large model. If you only ever use one model and it fits comfortably in memory, there's nothing to route.
  • No concurrent demand. If requests never overlap (you wait for each response before asking the next question), queuing isn't a bottleneck.
  • No multimodal needs. If you only do text chat, no image gen, no transcription, no embeddings, the routing advantages are smaller.

The trigger to switch: The moment you have a second machine doing nothing, or the moment you run an agent framework that makes parallel LLM calls, single Ollama becomes the bottleneck.

Where Single Ollama Breaks Down

1. Idle hardware

You have a Mac Studio with 192GB and a MacBook Pro with 36GB. The Studio runs your big model. The MacBook does nothing. That's 36GB of unified memory, enough for a 32B model, contributing zero value.

With Herd: Both machines serve requests. Big models route to the Studio, small models to the MacBook. Every device contributes what it can.

2. Agent bottleneck

CrewAI, LangChain, OpenClaw, Aider, these frameworks make rapid sequential or parallel LLM calls. On single Ollama, each call queues behind the last. A 5-agent pipeline with 4 calls each = 20 requests, all serialized.

With Herd: Requests distribute across the fleet. The Mac Studio handles the reasoning model, the MacBook handles the summarizer, the Mini handles embeddings, simultaneously.

3. Your laptop is also your workstation

An agent pipeline is hammering Ollama on the MacBook you are also using for a video call. The call stutters, your editor lags, and token generation slows for everything at once.

With Herd: The rest of the fleet shares the load, so the laptop is not the only place requests can go. Turn on adaptive capacity (herd-node --learn-capacity) and a Mac in a video call, or one under sustained heavy CPU load, is paused until it frees up.

4. Model contention

You need llama3.3:70b for coding and qwen3:32b for chat. On a 64GB machine they don't fit together (roughly 43GB plus 20GB of weights, before KV cache), so loading one evicts the other constantly (model thrashing). Cold-loading a 70B model takes 15-30 seconds each time.

With Herd: The 70B model stays hot on the big machine. Embeddings run on the smaller machine. No eviction, no cold-loading delay.

5. No observability

Single Ollama has no dashboard, no health checks, no request tracing. When something is slow, you don't know why, is the model thrashing? Is memory pressure high? Is the queue backed up?

With Herd: 8-tab dashboard, 30+ health checks, SQLite traces you can query, per-tag analytics showing which tools consume the most resources.

The Upgrade Path

The beauty of Herd is that the upgrade from single Ollama is minimal:

# What you're doing now
ollama serve

# Add Herd (on the same machine or a different one)
pip install ollama-herd   # or: brew install ollama-herd
herd                      # starts the router

# On each machine (including this one)
herd-node                 # discovers the router via mDNS

Your existing Ollama installation, models, and configuration stay exactly the same. Herd sits in front of Ollama, not instead of it. Every tool that currently points at localhost:11434 just needs to point at router-ip:11435 instead.

What You Don't Get Without Herd

Running multiple Ollama instances without Herd means:

  • Manual model placement, you decide which model goes where, and update every client when it changes
  • Manual failover, if a machine goes down, clients break until you reconfigure them
  • No load awareness, you're guessing which machine is less busy
  • No workstation protection, nothing steps your MacBook back while you're on a video call
  • No queue management, requests pile up on whichever machine the client happens to target
  • No cross-device analytics, no unified view of what's happening across your fleet
  • No multimodal routing, you manually remember which machine has mflux, which has Qwen3-ASR
  • No capacity learning, the system never adapts to your usage patterns

You can solve some of these with nginx reverse proxy rules and manual scripting. But that's rebuilding what Herd already does, without the scoring engine, capacity learning, or health monitoring.

Cost

Single Ollama: free.
Ollama Herd: also free. Open source. MIT licensed.

The only cost is the 2 minutes it takes to install and start. If you have a second machine, there's no reason not to try it.

Bottom Line

Single Ollama is where everyone starts. It's great for what it is, local inference on one machine. Ollama Herd is where you go when you outgrow one machine, when agents need more throughput, when your spare hardware should be contributing instead of sitting idle.

The question isn't "should I switch from Ollama?", you're not switching. You're adding an orchestration layer that makes all your Ollama instances work together. Ollama is the engine. Herd is the fleet manager.

If you have one machine and no concurrent demand, stay with single Ollama. The moment you have two machines or an agent workload, Herd pays for itself in the first hour.

Getting Started

pip install ollama-herd    # or: brew install ollama-herd
herd                       # start the router
herd-node                  # on each device

FAQ

Do I need Ollama Herd if I only have one Mac?

Probably not. Single Ollama handles one machine beautifully. But the moment you add a second device or start running agent frameworks that make parallel LLM calls, you will hit the limits of a single queue on a single machine. Herd is a 2-minute install, so the barrier is low when the time comes.

Does Ollama Herd replace Ollama?

No. Herd sits in front of Ollama, not instead of it. Your existing Ollama installation, models, and configuration stay exactly the same. Herd is the orchestration layer that connects multiple Ollama instances into one smart endpoint.

Will my existing tools still work?

Yes. Herd exposes both the Ollama API and the OpenAI-compatible API. Any tool currently pointing at localhost:11434 just needs to point at your router's address on port 11435 instead.

How does Herd discover my other machines?

mDNS (multicast DNS) auto-discovery. Start herd-node on any device on your local network and the router finds it automatically. No IP addresses to configure, no config files to edit.

What happens if a node goes down?

Herd detects the failure and automatically retries the request on the next-best available node, before the first token is sent to the client. With single Ollama, if it is down, it is down.

See Also