Ollama Herd vs NVIDIA PAIR

NVIDIA's Personal AI Router joins your PCs into a peer cluster where every machine runs PAIR and nodes are ranked by job count and GPU load. Herd is one router any device on your LAN can call, scoring nodes on eight signals. Both send each request to one machine. Neither shards a model.

TL;DR

  • Same basic idea. Both route each whole request to one machine that has the model. Neither pools VRAM or splits a model across machines.
  • Different shape. PAIR is a peer cluster: every machine runs PAIR, and each PAIR endpoint answers only the machine it runs on. Herd is one router that any device on your network can call.
  • Different scheduling. PAIR ranks nodes by pending jobs plus smoothed GPU utilization. Herd scores nodes on eight signals (warm model, memory fit, queue depth, wait time, model size fit, availability trend, context fit, session affinity) and does not look at GPU utilization.
  • PAIR is ahead on GPU-aware ranking, security (PIN pairing and mutual TLS), signed installers with one-click engine setup, LM Studio support, and NVIDIA's backing.
  • Herd is ahead on network-wide access, agent CLIs (Codex and Claude Code work today), image generation, speech-to-text, MLX on Apple Silicon, queues, fallbacks, health checks, and the web dashboard.

Version facts on this page are as of September 29, 2026: NVIDIA PAIR v0.1.1 (the latest published release, August 28, 2026) and Ollama Herd v0.9.5. Features on PAIR's development branch that have not shipped in a release are labeled "unreleased".

What is NVIDIA PAIR?

NVIDIA Personal AI Router (PAIR) is an open-source local inference router from NVIDIA, licensed Apache-2.0 and written in Go with an Electron desktop app. NVIDIA announced it as a beta on September 3, 2026. It connects Windows, Linux, and macOS machines on the same network into a cluster, manages Ollama and LM Studio on each one, and presents Ollama-compatible and OpenAI-compatible endpoints. Each request runs on one node chosen by a scheduler that combines queued work with GPU load. NVIDIA's product page lists GeForce RTX 20 Series and newer, RTX PRO, DGX Spark, and Macs with M4 or newer as supported hardware.

PAIR is well engineered and unusually candid about its limits: its architecture docs list exactly which signals the scheduler ignores. As of September 29, 2026 the repository had about 1,500 GitHub stars and a named team of eight NVIDIA engineers and managers.

What is Ollama Herd?

Ollama Herd is an open-source smart AI router that turns the machines you already own (Apple Silicon Macs, Linux servers, and Windows PCs, with or without NVIDIA GPUs) into one endpoint. It routes LLMs, embeddings, image generation, speech-to-text, and vision with an 8-signal scoring engine, mDNS auto-discovery, an 8-tab real-time dashboard, and OpenAI Chat Completions, OpenAI Responses, Ollama, and Anthropic Messages API compatibility. Backends are Ollama on every platform and mlx_lm.server on Apple Silicon, where Herd also adds native image generation and speech-to-text. pip install ollama-herd or brew install ollama-herd, then herd on one machine and herd-node on each device.

The Core Difference

Where the endpoint lives. PAIR has no central router. Every machine runs the same services, and the routing decision is made by the PAIR proxy on the machine where the request starts. That proxy serves plain HTTP to loopback only and refuses requests from other machines with a 403. Nodes talk to each other over mutual TLS after pairing. The result: to use the cluster from a machine, that machine must run PAIR and join the cluster. A container, a phone, a server running Open WebUI, or a laptop without PAIR cannot call it. PAIR's docs are explicit that a network-reachable endpoint is out of scope.

Herd is one router on port 11435 that listens on your LAN. Any device can point at it without installing anything. The trade is security: Herd's endpoint is plain, unauthenticated HTTP, so anything on your network can use it.

How the node is picked. Both first keep only the nodes that have the requested model. PAIR then sorts by pending jobs plus a GPU pressure score from the busiest GPU's smoothed utilization. PAIR's own docs note that it does not consider GPU model, available VRAM, latency, whether the model is already loaded, or how expensive a request is. Herd does not see GPU utilization at all, but it weighs whether the model is already warm, how comfortably it fits in memory, queue depth scaled by memory bandwidth, estimated wait, context headroom, and which node already holds a conversation's cache.

Conceptual: peer cluster vs central router
NVIDIA PAIR: peer cluster
Every machine that sends requests runs PAIR. Ranks nodes by pending jobs plus GPU utilization.
App on PC Acalls its own PAIR
Phone, container, or PC without PAIRrefused with a 403
loopback only, plain HTTP
PAIR on PC Amakes the routing decision
mutual TLS between paired nodes; each request goes to one node
PC APAIR + Ollama
PC BPAIR + LM Studio
Mac (M4 or newer)PAIR + Ollama
Ollama Herd: central router
One router any device can call. Does not use GPU utilization; no authentication.
Any device on the LANOpen WebUI server, container, phone, laptop
plain HTTP, no authentication
Herd router :11435warm model, memory fit, queue, context fit, session affinity
each request goes to one node
Macherd-node + Ollama or MLX
Linux PCherd-node + Ollama
Windows PCherd-node + Ollama
Both send each whole request to one machine: PAIR makes every calling machine join the cluster behind mutual TLS, while Herd gives the whole network one unauthenticated endpoint.

Where the Two Overlap

Being straight about what is not a difference:

  • Whole-request routing, no sharding. Both send each request to one machine. If a model does not fit on any single machine, neither tool helps. That is a job for exo.
  • Model-aware placement. Both only send a request to nodes that actually have the model, and both merge every node's model list into one.
  • Ollama and OpenAI Chat Completions. Both proxy Ollama's native API and /v1/chat/completions, with streaming.
  • mDNS discovery. Both find machines on the local network automatically.
  • Cross-platform, free, open source. Both run on macOS, Linux, and Windows. PAIR is Apache-2.0, Herd is MIT.

Feature Comparison

As of September 29, 2026. PAIR column reflects the v0.1.1 release unless marked unreleased. Herd column reflects v0.9.5.

Feature NVIDIA PAIR Ollama Herd
TopologyPeer cluster. Every machine runs PAIR; each endpoint serves its own machine onlyOne router on the LAN plus a light agent per node. Any device can call the router
Routing signalsModel present (hard filter), pending jobs, smoothed GPU utilizationModel present (hard filter), then 8 signals: warm model, memory fit, queue depth, wait time, model size fit, availability trend, context fit, session affinity
GPU awarenessLive GPU utilization (nvidia-smi on Linux, performance counters on Windows, ioreg on macOS). VRAM shown but not used for routingNo GPU utilization. VRAM not used for routing. Memory fit uses system memory
Model load stateShown in the UI, not used for routingWarm models scored highest
DiscoverymDNS, plus manual nodes by address (no desktop UI for manual nodes yet)mDNS, or herd-node --router-url for any reachable address
SecuritySix-digit PIN pairing, pinned certificates, mutual TLS between nodes, loopback-only local endpointsPlain HTTP on the LAN, no authentication (optional key for the Anthropic endpoint only)
BackendsOllama, LM Studio (installs, starts, updates them)Ollama, mlx_lm.server
OpenAI Chat CompletionsYes, plus /v1/completionsYes (/v1/completions not supported)
OpenAI Responses (Codex)Not routedYes
Anthropic Messages (Claude Code)Unreleased (merged to the development branch September 23, 2026), pass-through to the engineYes, with count_tokens, model auto-routing, and context management
EmbeddingsOllama /api/embed and OpenAI /v1/embeddingsOllama /api/embed plus a native embedding server (OpenAI /v1/embeddings not supported)
Image generationNoYes (mflux and DiffusionKit on Apple Silicon, Ollama native models elsewhere)
Speech-to-textNoYes (Apple Silicon)
Queues and fallbacksRetry on another node that has the model. Request queuing is on the roadmapPer-node, per-model queues, auto-retry, client-specified fallback models
Health monitoringEngine health probes, automatic restart of crashed services30 automated fleet health checks with recommendations
DashboardDesktop app: live GPU and memory per node, jobs, endpoints, settings. Terminal UI for headless machines8-tab web dashboard: overview, trends, model insights, tags, benchmarks, health, recommendations, settings
InstallSigned installers: Windows .exe, Linux .deb, macOS .dmgpip or brew, command line only
Officially supported hardwareRTX 20 Series and newer, RTX PRO, DGX Spark, Mac M4 or newerNo hardware floor. Any machine that runs Ollama; MLX features on any Apple Silicon Mac
Usage telemetryNo usage reporting in the public source. Official installers are built with additional NVIDIA build configuration that is not in the public repository, so their behavior is not publicly verifiableOne anonymous daily summary, on by default, announced on first run, off with FLEET_NODE_TELEMETRY=false. See /telemetry
LicenseApache-2.0MIT

Where PAIR Wins

  1. GPU-aware load ranking. PAIR reads live GPU utilization on every node and folds it into its ranking, so a busy RTX card gets less work. Herd does not read GPU utilization at all, and it does not use VRAM when deciding where a model fits. On a fleet of NVIDIA PCs, that is a real gap in Herd.
  2. Security model. Pairing with a PIN, pinned certificates, and mutual TLS between nodes, with local endpoints that refuse network callers. Herd's router is plain HTTP with no authentication, so you are relying on your network being trusted.
  3. Installers and engine management. Signed installers for Windows, Linux, and macOS, one-click install of Ollama or LM Studio from the app, and update prompts. Herd is a command-line install through pip or brew, and it expects you to install Ollama yourself.
  4. LM Studio. PAIR routes LM Studio as a first-class engine. Herd does not support LM Studio.
  5. Drop-in for existing clients. PAIR takes over Ollama's usual port 11434 and moves Ollama behind it, so apps already pointed at the default Ollama address join the cluster without changes. With Herd you point clients at the router on port 11435.
  6. OpenAI embeddings and completions paths. PAIR routes /v1/embeddings and /v1/completions. Herd v0.9.5 supports neither (Herd embeddings go through Ollama's /api/embed).
  7. NVIDIA's backing and hardware focus. A dedicated NVIDIA team, official support for RTX and DGX Spark, and launch coverage from NVIDIA's own blog. If your machines are Windows or Linux PCs with RTX cards, PAIR is built for exactly that.

Where Ollama Herd Wins

  1. One endpoint for the whole network. Any device can call Herd without installing anything: Open WebUI on a server, agents in containers, a phone, a teammate's laptop. PAIR requires every calling machine to run PAIR and join the cluster, and its maintainers have said a network-reachable endpoint is not planned.
  2. Richer placement. Herd prefers the node where the model is already warm, checks memory fit and context headroom, and keeps a conversation on the node holding its cache. PAIR's docs note that it can send a request to a node that must cold-load the model while a node with it loaded sits one place lower.
  3. Coding agents today. Codex works through Herd's OpenAI Responses API, and Claude Code works through its Anthropic Messages API, with model auto-routing, token counting, tool-schema fixup, and three-layer context management for long sessions. As of September 29, 2026, PAIR's Messages support is unreleased, and it does not route the Responses API.
  4. Multimodal routing. Image generation, speech-to-text, vision, and embeddings as distinct service types. PAIR routes chat and embeddings only.
  5. MLX on Apple Silicon. Herd routes to mlx_lm.server alongside Ollama, and supports any Apple Silicon Mac. NVIDIA lists M4 or newer for PAIR.
  6. Fleet operations. Per-model queues, fallback models, 30 automated health checks, benchmarks, request tagging, and an 8-tab web dashboard with trends and recommendations.
  7. Reaching nodes across a VPN. herd-node --router-url accepts any reachable address, including a Tailscale one. PAIR relies on mDNS, which does not cross Tailscale, and several of its open issues track VPN pairing problems.

When to Choose

Scenario Choose
Your machines are Windows or Linux PCs with NVIDIA RTX cards, or DGX SparkNVIDIA PAIR
You want signed installers and one-click engine setup, no command lineNVIDIA PAIR
You use LM StudioNVIDIA PAIR
Node-to-node encryption and authentication are requirementsNVIDIA PAIR
Every machine that sends requests can run PAIR itselfNVIDIA PAIR
Devices, containers, or servers that cannot run a router need to call the fleetOllama Herd
You run Codex or Claude Code against local models todayOllama Herd
Multimodal workload (image generation, speech-to-text, vision)Ollama Herd
Your fleet includes Macs older than M4, or you want MLXOllama Herd
Mixed hardware where warm models and memory fit matter more than GPU loadOllama Herd

Bottom Line

PAIR is the most serious new entrant in local request routing, and it is good at what it sets out to do: make a few NVIDIA PCs share work with a polished install and a sound security model. Its scheduler is deliberately simple today, and its endpoint is deliberately local, which means every machine that sends work must be part of the cluster.

Herd is built around a different idea: one router the whole network can call, with placement that knows which models are warm, what fits, and which node already holds a conversation, across chat, embeddings, images, and speech. Herd's gaps are real too: it does not use GPU utilization, it has no authentication, and it installs from the command line. If those matter most for your fleet, PAIR is the better fit.

Getting Started

If you already run Ollama or MLX on your machines, Herd discovers them automatically.

pip install ollama-herd    # or: brew install ollama-herd
herd                       # start router
herd-node                  # on each device

FAQ

Does NVIDIA PAIR pool GPU memory or split a model across machines?

No. PAIR sends each request to one node and says so plainly: it does not pool GPU memory, combine GPUs, or shard a model across machines. Ollama Herd works the same way. Both are request routers, so every model you use must fit on at least one machine.

Can another computer on my network use a PAIR endpoint?

Not unless it runs PAIR too. A PAIR endpoint accepts plaintext requests only from the machine it runs on and refuses other machines on the network with a 403, by design. To use the cluster from a machine, you install PAIR there and pair it. Ollama Herd takes the opposite approach: the router listens on your LAN, so any device can call it without installing anything, but that endpoint has no authentication.

Does NVIDIA PAIR work with Claude Code or Codex?

As of September 29, 2026, the released PAIR (v0.1.1) routes Ollama and OpenAI Chat Completions requests but not the Anthropic Messages API that Claude Code uses. Messages support was merged to PAIR's development branch on September 23, 2026 as a pass-through to the engine, and is unreleased. PAIR does not route the OpenAI Responses API that Codex uses. Ollama Herd v0.9.5 serves both, with context management for long Claude Code sessions.

Does Ollama Herd use GPU utilization when routing?

No. As of v0.9.5, Herd scores nodes on warm models, memory fit, queue depth, wait time, model size fit, availability trend, context fit, and session affinity, but it does not read live GPU utilization or use VRAM in routing decisions. PAIR does rank nodes by smoothed GPU utilization, which is a real advantage on NVIDIA machines.

Can I run NVIDIA PAIR and Ollama Herd together?

We have not tested it. They solve overlapping problems, so most fleets should pick one. If you try both on one machine, note that PAIR moves Ollama to port 11435 and up by default, which is also Herd's default router port.

Are NVIDIA PAIR and Ollama Herd free?

Yes. NVIDIA PAIR is open source under Apache-2.0 and Ollama Herd is open source under MIT. Neither charges for use.

See Also