How Ollama Herd Compares

An honest look at where Herd fits, and when you should use something else.

Quick Comparison

Feature Single Ollama DIY Scripts exo LiteLLM GPUStack Ollama Herd
Multi-device routing No Manual No (splits models) Yes (endpoints listed in config) Yes Yes
Zero-config setup Yes No Yes Config file Install + config 2 commands
mDNS auto-discovery No No Yes No Yes Yes
Memory pressure detection No No No No No Yes (macOS, Linux)
Meeting detection No No No No No Yes (macOS)
Capacity learning No No No No No 168-slot model (opt-in)
Per node:model queues No No No Rate limiting Yes Yes
Multi-signal scoring No No No Provider-level Engine selection 8 signals
Model fallbacks No No No Yes No Yes
Auto-retry on failure No No No Yes No Yes
Auto-pull missing models No No No No Yes Yes
Real-time dashboard No No Limited Admin panel Web UI SSE + 8 tabs
Request tagging/analytics No No No Yes No Yes
Runs on macOS, Linux, Windows Yes Depends macOS (Linux CPU-only) Yes Linux (GPU workers) Yes
OpenAI API compatible Yes Fragile Yes Yes Yes Yes
Ollama API compatible Yes Partial Yes Via config No Yes
Multimodal (images + STT) No No No No No Yes (STT on Apple Silicon)
Target user Single machine Tinkerers Model sharding Cloud gateway GPU clusters Personal fleet
Best for One machine, one user, simple setup Learning, prototyping with 2–3 machines Running one huge model across multiple GPUs Routing between cloud API providers Enterprise GPU cluster management A few Macs, Linux boxes, and PCs running mixed workloads

Which Tool Fits?

Go down the list and stop at the first yes. Most of these answers are not Herd, and that is the point: the right tool depends on the problem.

  1. Only one machine?
    Stay on single Ollama. A router adds nothing until there is a second machine or a second heavy user.
  2. Is one model too big for any single machine you own?
    You need sharding, not routing: exo, Apple's MLX distributed (Thunderbolt 5 Macs), LocalAI's worker mode, or OLOL. Herd sends whole requests to one machine, so it cannot serve that model; it can serve your other models side by side.
  3. Do you need cloud providers (OpenAI, Anthropic, Bedrock) behind the same endpoint?
    Use a gateway: LiteLLM or Bifrost, or Envoy AI Gateway on Kubernetes. Herd is optional behind it, as the gateway's local backend.
  4. Do your machines run engines other than Ollama (or MLX on a Mac)?
    Herd routes only to Ollama and mlx_lm.server. For vLLM, SGLang, llama.cpp, or LM Studio backends, look at Olla (11+ engines), llama-swap with peers, LocalAI, or GPUStack for Linux GPU servers.
  5. Can every machine that sends requests run the router itself, and do you want encrypted, authenticated node traffic?
    NVIDIA PAIR: PIN pairing, mutual TLS, GPU-utilization ranking, and signed installers, built for RTX PCs and DGX Spark.
  6. Otherwise: several machines running Ollama, used by people, called by anything on your network?
    Ollama Herd. One endpoint any device can call, placement on warm models, memory fit, queue depth, context fit, and session affinity, and extras on Apple Silicon. It has no authentication, so it assumes a trusted network.

Detailed Comparisons

vs. Olla

Olla is a high-performance multi-engine proxy and load balancer for self-hosted LLMs, built by TensorFoundry. It routes across the widest set of backends, vLLM, SGLang, LMDeploy, llama.cpp, Docker Model Runner, Ollama, and more, with production-proxy machinery like circuit breakers and KV-cache sticky sessions. It's our most direct competitor and a genuinely well-built project.

The difference is what each schedules around. Olla treats machines as interchangeable inference endpoints with a health status, listed in a config file. Herd treats them as workstations with conditions (warm models, memory pressure, heavy workloads, meetings on macOS, learned weekly availability), discovered automatically via mDNS. Olla is the better generic proxy; Herd is the better scheduler for machines people also use.

Choose Olla when: Your fleet mixes many inference engines and you want the broadest backend coverage.

Choose Herd when: Your fleet is real computers (Macs, Linux desktops, Windows PCs) that people also use for work, and routing around each machine's condition matters.

Read the full Olla comparison →

vs. NVIDIA PAIR

NVIDIA's Personal AI Router (PAIR) is an open-source (Apache-2.0) local inference router that NVIDIA announced as a beta in September 2026. It ships signed installers for Windows, Linux, and macOS, installs and manages Ollama or LM Studio for you, and pairs machines with a PIN so nodes talk over mutual TLS. Like Herd, it sends each whole request to one machine and never shards a model.

The difference is shape and scheduling. A PAIR endpoint only answers the machine it runs on, so every machine that sends requests must run PAIR and join the cluster, and it ranks nodes by pending jobs plus live GPU utilization. Herd is one router any device on your LAN can call, and it scores nodes on warm models, memory fit, queue depth, context fit, and session affinity, across LLMs, embeddings, image generation, and speech-to-text. Herd does not use GPU utilization and its endpoint has no authentication; PAIR is ahead on both.

Choose PAIR when: Your machines are NVIDIA RTX PCs or DGX Spark, you want a GUI installer and encrypted node-to-node traffic, and every client can run PAIR itself.

Choose Herd when: Devices or containers that cannot run a router need to call the fleet, you run Codex or Claude Code against local models, or you need multimodal routing or MLX on Apple Silicon.

Read the full NVIDIA PAIR comparison →

vs. Single Ollama

Running one Ollama instance is the starting point. It works great, until you have more than one machine or more than one concurrent user.

When single Ollama is enough:

  • You have one machine
  • You run one model at a time
  • You don't mind waiting in a queue

When you need Herd:

  • You have 2+ machines and want them to work together
  • Multiple tools hit Ollama simultaneously (agents, coding assistant, chat)
  • You're tired of model thrashing (loading/unloading models to free memory)
  • Your laptop is under memory pressure while you work and you want requests routed elsewhere

LM Link is LM Studio's private multi-device feature, connect your LM Studio install to other machines over a Tailscale mesh and access local models remotely. End-to-end encrypted, free up to 2 users / 10 devices in preview.

LM Link connects your LM Studio installs to each other. Ollama Herd routes across your whole team's mixed fleet, Mac, Linux, Windows, any Ollama- or MLX-compatible node, with intelligent scoring, three-layer context management for long Claude Code sessions, per-tier model mapping, and a real admin dashboard. LM Link is connectivity; Ollama Herd is orchestration.

Choose LM Link when: You're all-in on LM Studio as your model runner and just need remote access from your other machines.

Choose Herd when: Your fleet is heterogeneous (mixed OSes, multiple runtimes) or you need scoring/routing/compaction/admin features beyond "connect these devices."

Read the full LM Link comparison →

vs. exo

exo splits a single large model across multiple devices, by layers (pipeline parallelism) or within layers (tensor parallelism). If one machine can't fit a 405B model, the devices collectively run it.

exo and Herd solve different problems. exo answers "how do I run a model too big for one machine?" Herd answers "how do I route many requests to many models across many machines?" They're complementary: use exo for the one model too big for any single machine, and Herd for routing everything else across your fleet.

Status note, September 2026: exo's main branch has had one documentation commit since June 2026. It remains by far the most-starred project in Mac clustering and the tech is genuinely ahead, so this may well be temporary. If you are choosing today, check the repo yourself before depending on it.

Choose exo when: You need to run one model that's too large for any single device.

Choose Herd when: You have multiple devices that can each run their own models and you want intelligent routing across all of them.

vs. Apple's Distributed MLX

At WWDC26, Apple shipped first-party distributed inference for MLX (macOS 26.2+): RDMA over Thunderbolt 5, the open-source JACCL collective communication library, and MLX Distributed with tensor and pipeline parallelism. It runs frontier-scale models across Thunderbolt-connected Mac Studios, the same model-sharding problem exo solves, now from the platform vendor.

Same distinction as exo: Apple's stack dedicates a pure-Apple cluster to one huge model; Herd routes many models across a mixed fleet with 8-signal scoring. Use it for the one model that needs the whole cluster, and Herd for everything else.

Choose MLX Distributed when: You have Thunderbolt 5-connected Macs and one model too big for any of them.

Choose Herd when: Your fleet is mixed hardware, your workload is multi-model or multimodal, or you want request-level routing intelligence.

vs. LiteLLM

LiteLLM is an API gateway that provides a unified OpenAI-compatible interface to 100+ LLM providers: cloud APIs (OpenAI, Anthropic, Bedrock, Azure) and self-hosted servers such as Ollama and vLLM.

Different layer. LiteLLM can load-balance across several Ollama servers, but it treats them as endpoints you list in a config file. Herd routes between local devices it discovers on its own. LiteLLM has no concept of warm models, memory pressure, device health, or mDNS discovery. They work together naturally, Herd sits between LiteLLM and your local Ollama instances, giving LiteLLM a single "local" endpoint backed by an intelligent fleet.

Choose LiteLLM when: You need to route between cloud providers or want a unified API across OpenAI/Anthropic/etc.

Choose Herd when: You want your local devices to work together. Use both if you want local + cloud with intelligent routing at each layer.

vs. GPUStack

GPUStack is a GPU cluster manager for AI model deployment. It manages GPU resources across environments (on-prem, Kubernetes, cloud), auto-configures inference engines (vLLM, SGLang, TensorRT-LLM), and supports all GPU vendors.

GPUStack is more polished but more complex. It targets GPU cluster operators who want multi-engine support and enterprise features. Herd targets individuals and small teams who want zero-config fleet management with the Ollama they already use.

Choose GPUStack when: You're managing a Linux GPU cluster with mixed vendors and need multi-engine support.

Choose Herd when: You have a few personal devices running Ollama and want them to work together in 60 seconds.

Status note, August 2026: GPUStack dropped macOS and Windows support in v2.0.0. Their docs now state that macOS is not supported for worker nodes, so it is no longer an option for a Mac fleet at all.

vs. DIY Scripts

Many people write their own routing scripts, round-robin across Ollama instances, manually checking which node has capacity, or just SSH-ing into whichever machine seems free.

DIY works until it doesn't. You'll spend more time maintaining the scripts than using them. No warm-model awareness, no capacity learning, no auto-retry, no dashboard, no meeting detection. Every edge case becomes your problem.

Choose DIY when: You have very specific routing logic that no tool supports.

Choose Herd when: You want routing that handles the edge cases you haven't thought of yet.

Deep Dive Comparisons

Each comparison page covers feature tables, honest pros/cons, when to choose each tool, FAQs, and getting started guides.

The Competitive Landscape

Dozens of projects will route local AI requests for you. Once you look closely, they sort into seven groups by what they actually do, and only the first overlaps with Herd. In that one group, two projects come close, each built for a different kind of fleet. Every tool with a full comparison page is placed below.

Conceptual: every compared tool, grouped by what it does
Fleet routers and proxies

Pick one machine for each request. Herd's group.

Gateways and load balancers

Sit in front: auth, rate limits, cloud providers.

Platforms that own the stack

Run or manage the engines and schedule their own nodes.

One model across machines

Shard a model too big for any one device.

Single-machine servers

Run models on one machine. Backends, not rivals.

Chat interfaces

Frontends for people. A different layer.

Not a router at all

The two most common alternatives.

Twenty compared tools in seven groups; only the fleet routers do the same job as Herd, and several of the rest pair with it instead of replacing it.

Fleet routers & proxies, the group Herd lives in

This is where the real competition is. Most entrants are thin: round-robin dispatchers, model-aware proxies, per-user fair-share queues, each solving one slice of fleet coordination. Tools like ollamaMQ, ollamafarm, and LoLLMs Hub cover a slice well (see the proxy tools roundup). ollama_load_balancer races each chat request across several backends and keeps the fastest, and OLOL clusters Ollama over gRPC and can also shard a large model. Two projects compete with Herd on the whole problem: Olla, a genuinely strong, company-backed multi-engine proxy, and NVIDIA PAIR, NVIDIA's peer-to-peer router for RTX PCs, DGX Spark, and Macs, which ranks nodes by live GPU load but only serves machines that run PAIR themselves. Olla routes across more backend engines than Herd does, but it treats your machines as interchangeable servers in a config file. PAIR installs more easily and encrypts node-to-node traffic, but it schedules on job count and GPU load rather than on warm models, memory fit, or context. Herd is the only router that treats them as workstations with conditions (warm models, memory pressure, heavy workloads, meetings on macOS, learned weekly availability), discovered automatically with zero config. Olla is the better generic proxy, PAIR is the easier NVIDIA PC cluster, and Herd is the better fleet scheduler for a network of mixed machines that everything can call.

Gateways and load balancers, the layer in front

LiteLLM, Bifrost, and Envoy AI Gateway put many providers, mostly cloud APIs, behind one endpoint with keys, budgets, and rate limits. HAProxy is a general-purpose load balancer with a hardened edge. LiteLLM and HAProxy can balance several Ollama servers directly, but as endpoints listed in a config, without knowing which one has the model loaded. They pair naturally with Herd: the gateway handles auth and decides cloud or local, and Herd picks the local machine.

Platforms that own the stack, a different trade

LocalAI runs the models itself and, in its distributed mode (PostgreSQL and NATS), schedules them onto its own nodes with VRAM-aware placement and replica autoscaling. GPUStack manages Linux GPU servers and the engines on them (vLLM, SGLang, TensorRT-LLM). Both are stronger than Herd on GPU servers and multi-user access control. The trade is that every node runs their stack; Herd keeps the Ollama you already run and adds routing on top.

Distributed-inference clusters, a different problem

exo, Apple's WWDC26 distributed MLX stack, and Distributed Llama shard one model across several machines so you can run models too big for any single device; LocalAI's worker mode and OLOL can do it too. That's model parallelism, not request routing, and it's complementary: run a sharding tool for the model that needs it, and Herd for everything else. Herd doesn't route to those clusters today; they run as their own endpoint.

Single-machine servers, backends, not rivals

Ollama is what Herd routes to. vLLM is the serving engine for NVIDIA GPU servers, Docker Model Runner brings models into the Docker workflow, llama-swap swaps more models than fit in memory on one machine (and can forward listed models to peers), and vLLM-MLX and vMLX speed up one Apple Silicon machine. LM Studio's LM Link reaches one LM Studio server from other devices over Tailscale: remote access, not routing. Herd routes to Ollama and mlx_lm.server today, not to these other engines, so they run as their own endpoints.

Chat interfaces, a different layer

Open WebUI and friends are frontends. Point one at Herd and you get a beautiful chat UI backed by an intelligent fleet. Different layer, natural pairing.

Not a router at all

The two most common alternatives are not tools in this list. Cloud APIs trade hardware for per-token pricing and frontier models, and the usual answer is both. DIY scripts, usually an nginx round-robin, work on day one and drift as the fleet changes.

What Each Scheduler Looks At

Five tools choose a machine for each whole request. They differ in what they look at when choosing. This is taken from each tool's comparison page; where a page does not say, the cell says so rather than guessing.

Sourced: each tool's comparison page, September 2026
SignalOllama HerdNVIDIA PAIROllaLocalAI (distributed)llama-swap (peers)
Model is on the nodeYes (hard filter)Yes (hard filter)Partly (model discovery and unification across endpoints)Partly (places models on nodes itself)Partly (models you list per peer in its config)
Model already loadedYes (model warmth: loaded, recently loaded, on disk)No (shown in the UI, not used for routing)No (not part of placement)Yes (prefers a node that has it)Yes (prefers a peer that has it loaded)
Memory fit (system RAM)YesNoNoNot documentedNo (you configure which models fit)
VRAM / GPU memoryNo (memory fit uses system RAM)No (shown, not used for routing)NoYes (free VRAM per node)No
GPU utilizationNoYes (smoothed, busiest GPU)NoNot documentedNo
Queue depth / pending jobsYes (scaled by memory bandwidth)Yes (pending jobs)Not documentedPartly (idle and least-loaded nodes)Partly (spills over when local capacity is full)
Context fitYesNoNot documentedNot documentedNot documented
Session or prefix-cache affinityYes (session affinity)NoYes (KV-cache sticky sessions)Yes (prefix-cache-aware routing)Not documented
Workstation conditions (memory pressure, meetings)Yes (memory pressure; meeting detection is opt-in, macOS only)NoNo (endpoints are health up or down)No (not a documented focus)No (treats backends as always available)
Herd weighs the most signals of the five, including context fit and workstation conditions that none of the others document, but it ignores VRAM and GPU utilization, which LocalAI and PAIR use. ● uses it, ◐ partly, ○ does not, ? not documented on its comparison page. Sources: the Herd vs PAIR, Olla, LocalAI, and proxy tools pages.

What Makes Herd Unique

No other project combines all of these, and the combination is the point:

  1. 8-signal intelligent scoring with learned latency data
  2. Workstation-condition awareness, opt-in meeting detection (macOS) and app fingerprinting, because laptops aren't servers
  3. Adaptive capacity learning, a 168-slot model of your fleet's weekly rhythm
  4. mDNS zero-config discovery, no endpoint list to maintain, truly two commands
  5. Per node:model queue management with dynamic concurrency and auto-retry
  6. Multimodal routing, LLM, embeddings, image gen, speech-to-text, and vision as distinct service types
  7. OpenAI + Ollama + Anthropic Messages + OpenAI Responses, drop-in for any client, including Claude Code and Codex
  8. Every machine you own, macOS, Linux, and Windows, NVIDIA or not, with MLX, image generation, and speech-to-text extras on Apple Silicon
  9. Real-time 8-tab dashboard with fleet overview, trends, health, and analytics
  10. Completely free and open, MIT licensed, no paid tier, no token, no upsell, no accounts

Plenty of tools do one or two of these. Herd is the only one built around all of them at once, and the only one that starts from the premise that a group of computers people actually use for work are not interchangeable servers. That's the fleet Herd was built for, and it owns it.

And it's genuinely free. Herd is MIT-licensed top to bottom, with no paid tier, no per-seat pricing, no token, no metering, and no commercial product it's steering you toward. What you pip install is the whole thing, forever. Every feature on this site, including the enterprise ones, is open source. Optional deployment and support help is available if you want it, but the software itself has no gate.

The one built for the machines you actually use

You've seen the field. If your machines are workstations and not servers, start here.