Ollama Herd vs NVIDIA PAIR
NVIDIA's Personal AI Router joins your PCs into a peer cluster where every machine runs PAIR and nodes are ranked by job count and GPU load. Herd is one router any device on your LAN can call, scoring nodes on eight signals. Both send each request to one machine. Neither shards a model.
TL;DR
- Same basic idea. Both route each whole request to one machine that has the model. Neither pools VRAM or splits a model across machines.
- Different shape. PAIR is a peer cluster: every machine runs PAIR, and each PAIR endpoint answers only the machine it runs on. Herd is one router that any device on your network can call.
- Different scheduling. PAIR ranks nodes by pending jobs plus smoothed GPU utilization. Herd scores nodes on eight signals (warm model, memory fit, queue depth, wait time, model size fit, availability trend, context fit, session affinity) and does not look at GPU utilization.
- PAIR is ahead on GPU-aware ranking, security (PIN pairing and mutual TLS), signed installers with one-click engine setup, LM Studio support, and NVIDIA's backing.
- Herd is ahead on network-wide access, agent CLIs (Codex and Claude Code work today), image generation, speech-to-text, MLX on Apple Silicon, queues, fallbacks, health checks, and the web dashboard.
Version facts on this page are as of September 29, 2026: NVIDIA PAIR v0.1.1 (the latest published release, August 28, 2026) and Ollama Herd v0.9.5. Features on PAIR's development branch that have not shipped in a release are labeled "unreleased".
What is NVIDIA PAIR?
NVIDIA Personal AI Router (PAIR) is an open-source local inference router from NVIDIA, licensed Apache-2.0 and written in Go with an Electron desktop app. NVIDIA announced it as a beta on September 3, 2026. It connects Windows, Linux, and macOS machines on the same network into a cluster, manages Ollama and LM Studio on each one, and presents Ollama-compatible and OpenAI-compatible endpoints. Each request runs on one node chosen by a scheduler that combines queued work with GPU load. NVIDIA's product page lists GeForce RTX 20 Series and newer, RTX PRO, DGX Spark, and Macs with M4 or newer as supported hardware.
PAIR is well engineered and unusually candid about its limits: its architecture docs list exactly which signals the scheduler ignores. As of September 29, 2026 the repository had about 1,500 GitHub stars and a named team of eight NVIDIA engineers and managers.
What is Ollama Herd?
Ollama Herd is an open-source smart AI router that turns the machines you already own (Apple Silicon Macs, Linux servers, and Windows PCs, with or without NVIDIA GPUs) into one endpoint. It routes LLMs, embeddings, image generation, speech-to-text, and vision with an 8-signal scoring engine, mDNS auto-discovery, an 8-tab real-time dashboard, and OpenAI Chat Completions, OpenAI Responses, Ollama, and Anthropic Messages API compatibility. Backends are Ollama on every platform and mlx_lm.server on Apple Silicon, where Herd also adds native image generation and speech-to-text. pip install ollama-herd or brew install ollama-herd, then herd on one machine and herd-node on each device.
The Core Difference
Where the endpoint lives. PAIR has no central router. Every machine runs the same services, and the routing decision is made by the PAIR proxy on the machine where the request starts. That proxy serves plain HTTP to loopback only and refuses requests from other machines with a 403. Nodes talk to each other over mutual TLS after pairing. The result: to use the cluster from a machine, that machine must run PAIR and join the cluster. A container, a phone, a server running Open WebUI, or a laptop without PAIR cannot call it. PAIR's docs are explicit that a network-reachable endpoint is out of scope.
Herd is one router on port 11435 that listens on your LAN. Any device can point at it without installing anything. The trade is security: Herd's endpoint is plain, unauthenticated HTTP, so anything on your network can use it.
How the node is picked. Both first keep only the nodes that have the requested model. PAIR then sorts by pending jobs plus a GPU pressure score from the busiest GPU's smoothed utilization. PAIR's own docs note that it does not consider GPU model, available VRAM, latency, whether the model is already loaded, or how expensive a request is. Herd does not see GPU utilization at all, but it weighs whether the model is already warm, how comfortably it fits in memory, queue depth scaled by memory bandwidth, estimated wait, context headroom, and which node already holds a conversation's cache.
Where the Two Overlap
Being straight about what is not a difference:
- Whole-request routing, no sharding. Both send each request to one machine. If a model does not fit on any single machine, neither tool helps. That is a job for exo.
- Model-aware placement. Both only send a request to nodes that actually have the model, and both merge every node's model list into one.
- Ollama and OpenAI Chat Completions. Both proxy Ollama's native API and
/v1/chat/completions, with streaming. - mDNS discovery. Both find machines on the local network automatically.
- Cross-platform, free, open source. Both run on macOS, Linux, and Windows. PAIR is Apache-2.0, Herd is MIT.
Feature Comparison
As of September 29, 2026. PAIR column reflects the v0.1.1 release unless marked unreleased. Herd column reflects v0.9.5.
| Feature | NVIDIA PAIR | Ollama Herd |
|---|---|---|
| Topology | Peer cluster. Every machine runs PAIR; each endpoint serves its own machine only | One router on the LAN plus a light agent per node. Any device can call the router |
| Routing signals | Model present (hard filter), pending jobs, smoothed GPU utilization | Model present (hard filter), then 8 signals: warm model, memory fit, queue depth, wait time, model size fit, availability trend, context fit, session affinity |
| GPU awareness | Live GPU utilization (nvidia-smi on Linux, performance counters on Windows, ioreg on macOS). VRAM shown but not used for routing | No GPU utilization. VRAM not used for routing. Memory fit uses system memory |
| Model load state | Shown in the UI, not used for routing | Warm models scored highest |
| Discovery | mDNS, plus manual nodes by address (no desktop UI for manual nodes yet) | mDNS, or herd-node --router-url for any reachable address |
| Security | Six-digit PIN pairing, pinned certificates, mutual TLS between nodes, loopback-only local endpoints | Plain HTTP on the LAN, no authentication (optional key for the Anthropic endpoint only) |
| Backends | Ollama, LM Studio (installs, starts, updates them) | Ollama, mlx_lm.server |
| OpenAI Chat Completions | Yes, plus /v1/completions | Yes (/v1/completions not supported) |
| OpenAI Responses (Codex) | Not routed | Yes |
| Anthropic Messages (Claude Code) | Unreleased (merged to the development branch September 23, 2026), pass-through to the engine | Yes, with count_tokens, model auto-routing, and context management |
| Embeddings | Ollama /api/embed and OpenAI /v1/embeddings | Ollama /api/embed plus a native embedding server (OpenAI /v1/embeddings not supported) |
| Image generation | No | Yes (mflux and DiffusionKit on Apple Silicon, Ollama native models elsewhere) |
| Speech-to-text | No | Yes (Apple Silicon) |
| Queues and fallbacks | Retry on another node that has the model. Request queuing is on the roadmap | Per-node, per-model queues, auto-retry, client-specified fallback models |
| Health monitoring | Engine health probes, automatic restart of crashed services | 30 automated fleet health checks with recommendations |
| Dashboard | Desktop app: live GPU and memory per node, jobs, endpoints, settings. Terminal UI for headless machines | 8-tab web dashboard: overview, trends, model insights, tags, benchmarks, health, recommendations, settings |
| Install | Signed installers: Windows .exe, Linux .deb, macOS .dmg | pip or brew, command line only |
| Officially supported hardware | RTX 20 Series and newer, RTX PRO, DGX Spark, Mac M4 or newer | No hardware floor. Any machine that runs Ollama; MLX features on any Apple Silicon Mac |
| Usage telemetry | No usage reporting in the public source. Official installers are built with additional NVIDIA build configuration that is not in the public repository, so their behavior is not publicly verifiable | One anonymous daily summary, on by default, announced on first run, off with FLEET_NODE_TELEMETRY=false. See /telemetry |
| License | Apache-2.0 | MIT |
Where PAIR Wins
- GPU-aware load ranking. PAIR reads live GPU utilization on every node and folds it into its ranking, so a busy RTX card gets less work. Herd does not read GPU utilization at all, and it does not use VRAM when deciding where a model fits. On a fleet of NVIDIA PCs, that is a real gap in Herd.
- Security model. Pairing with a PIN, pinned certificates, and mutual TLS between nodes, with local endpoints that refuse network callers. Herd's router is plain HTTP with no authentication, so you are relying on your network being trusted.
- Installers and engine management. Signed installers for Windows, Linux, and macOS, one-click install of Ollama or LM Studio from the app, and update prompts. Herd is a command-line install through
piporbrew, and it expects you to install Ollama yourself. - LM Studio. PAIR routes LM Studio as a first-class engine. Herd does not support LM Studio.
- Drop-in for existing clients. PAIR takes over Ollama's usual port 11434 and moves Ollama behind it, so apps already pointed at the default Ollama address join the cluster without changes. With Herd you point clients at the router on port 11435.
- OpenAI embeddings and completions paths. PAIR routes
/v1/embeddingsand/v1/completions. Herd v0.9.5 supports neither (Herd embeddings go through Ollama's/api/embed). - NVIDIA's backing and hardware focus. A dedicated NVIDIA team, official support for RTX and DGX Spark, and launch coverage from NVIDIA's own blog. If your machines are Windows or Linux PCs with RTX cards, PAIR is built for exactly that.
Where Ollama Herd Wins
- One endpoint for the whole network. Any device can call Herd without installing anything: Open WebUI on a server, agents in containers, a phone, a teammate's laptop. PAIR requires every calling machine to run PAIR and join the cluster, and its maintainers have said a network-reachable endpoint is not planned.
- Richer placement. Herd prefers the node where the model is already warm, checks memory fit and context headroom, and keeps a conversation on the node holding its cache. PAIR's docs note that it can send a request to a node that must cold-load the model while a node with it loaded sits one place lower.
- Coding agents today. Codex works through Herd's OpenAI Responses API, and Claude Code works through its Anthropic Messages API, with model auto-routing, token counting, tool-schema fixup, and three-layer context management for long sessions. As of September 29, 2026, PAIR's Messages support is unreleased, and it does not route the Responses API.
- Multimodal routing. Image generation, speech-to-text, vision, and embeddings as distinct service types. PAIR routes chat and embeddings only.
- MLX on Apple Silicon. Herd routes to
mlx_lm.serveralongside Ollama, and supports any Apple Silicon Mac. NVIDIA lists M4 or newer for PAIR. - Fleet operations. Per-model queues, fallback models, 30 automated health checks, benchmarks, request tagging, and an 8-tab web dashboard with trends and recommendations.
- Reaching nodes across a VPN.
herd-node --router-urlaccepts any reachable address, including a Tailscale one. PAIR relies on mDNS, which does not cross Tailscale, and several of its open issues track VPN pairing problems.
When to Choose
| Scenario | Choose |
|---|---|
| Your machines are Windows or Linux PCs with NVIDIA RTX cards, or DGX Spark | NVIDIA PAIR |
| You want signed installers and one-click engine setup, no command line | NVIDIA PAIR |
| You use LM Studio | NVIDIA PAIR |
| Node-to-node encryption and authentication are requirements | NVIDIA PAIR |
| Every machine that sends requests can run PAIR itself | NVIDIA PAIR |
| Devices, containers, or servers that cannot run a router need to call the fleet | Ollama Herd |
| You run Codex or Claude Code against local models today | Ollama Herd |
| Multimodal workload (image generation, speech-to-text, vision) | Ollama Herd |
| Your fleet includes Macs older than M4, or you want MLX | Ollama Herd |
| Mixed hardware where warm models and memory fit matter more than GPU load | Ollama Herd |
Bottom Line
PAIR is the most serious new entrant in local request routing, and it is good at what it sets out to do: make a few NVIDIA PCs share work with a polished install and a sound security model. Its scheduler is deliberately simple today, and its endpoint is deliberately local, which means every machine that sends work must be part of the cluster.
Herd is built around a different idea: one router the whole network can call, with placement that knows which models are warm, what fits, and which node already holds a conversation, across chat, embeddings, images, and speech. Herd's gaps are real too: it does not use GPU utilization, it has no authentication, and it installs from the command line. If those matter most for your fleet, PAIR is the better fit.
Getting Started
If you already run Ollama or MLX on your machines, Herd discovers them automatically.
pip install ollama-herd # or: brew install ollama-herd
herd # start router
herd-node # on each device
FAQ
Does NVIDIA PAIR pool GPU memory or split a model across machines?
No. PAIR sends each request to one node and says so plainly: it does not pool GPU memory, combine GPUs, or shard a model across machines. Ollama Herd works the same way. Both are request routers, so every model you use must fit on at least one machine.
Can another computer on my network use a PAIR endpoint?
Not unless it runs PAIR too. A PAIR endpoint accepts plaintext requests only from the machine it runs on and refuses other machines on the network with a 403, by design. To use the cluster from a machine, you install PAIR there and pair it. Ollama Herd takes the opposite approach: the router listens on your LAN, so any device can call it without installing anything, but that endpoint has no authentication.
Does NVIDIA PAIR work with Claude Code or Codex?
As of September 29, 2026, the released PAIR (v0.1.1) routes Ollama and OpenAI Chat Completions requests but not the Anthropic Messages API that Claude Code uses. Messages support was merged to PAIR's development branch on September 23, 2026 as a pass-through to the engine, and is unreleased. PAIR does not route the OpenAI Responses API that Codex uses. Ollama Herd v0.9.5 serves both, with context management for long Claude Code sessions.
Does Ollama Herd use GPU utilization when routing?
No. As of v0.9.5, Herd scores nodes on warm models, memory fit, queue depth, wait time, model size fit, availability trend, context fit, and session affinity, but it does not read live GPU utilization or use VRAM in routing decisions. PAIR does rank nodes by smoothed GPU utilization, which is a real advantage on NVIDIA machines.
Can I run NVIDIA PAIR and Ollama Herd together?
We have not tested it. They solve overlapping problems, so most fleets should pick one. If you try both on one machine, note that PAIR moves Ollama to port 11435 and up by default, which is also Herd's default router port.
Are NVIDIA PAIR and Ollama Herd free?
Yes. NVIDIA PAIR is open source under Apache-2.0 and Ollama Herd is open source under MIT. Neither charges for use.
See Also
- Ollama Herd vs Olla, broad multi-engine proxy
- Ollama Herd vs LiteLLM, cloud API gateway
- Ollama Herd vs LM Link, LM Studio's remote access across devices
- Ollama Herd vs exo, splitting one model across machines