Ollama Herd vs Ollama Proxy/Routing Tools

A whole field of tools has emerged to route local AI requests, round-robin dispatchers, model-aware proxies, per-user fair-share queues. Most solve one slice. Two compete on the whole problem: Olla, built for interchangeable servers, and NVIDIA PAIR, built for machines that all run PAIR. Herd is the only integrated solution with device-condition scoring, multimodal routing, mDNS auto-discovery, capacity learning, and a real dashboard.

What are Ollama Proxy and Routing Tools?

A growing ecosystem of open-source tools addresses pieces of the local-AI routing problem. Some are lean round-robin dispatchers; some add model-aware routing or per-user fair-share queues; some manage model loading on a single machine. Each solves one slice of fleet coordination. Very few attempt the whole problem, full routing intelligence, multimodal support, capacity learning, and observability, and of those that do, only one is a serious peer to Herd.

The One Serious Peer: Olla

Olla (built by TensorFoundry) is the standout of this field and the closest thing to a direct competitor Herd has. It's a high-performance Go proxy supporting more backend engines than Herd, vLLM, SGLang, LMDeploy, llama.cpp, Docker Model Runner, Ollama, and more, with production-proxy machinery like circuit breakers and KV-cache sticky sessions. If your fleet is a rack of dedicated, interchangeable inference servers across many engines, Olla is excellent.

The difference is philosophical. Olla treats machines as endpoints with a health status, listed in a config file. Herd treats them as workstations with conditions, is this Mac in a meeting (macOS, opt-in), is it under sustained heavy load, does it already have the model warm, discovered automatically via mDNS. Olla is the better generic proxy; Herd is the better scheduler for a fleet of real machines that people also work on. See the full Ollama Herd vs Olla comparison →

What is Ollama Herd?

Ollama Herd is an open-source smart AI router that turns the machines you already own (Apple Silicon Macs, Linux servers, and Windows PCs, with or without NVIDIA GPUs) into one endpoint. It routes LLMs, embeddings, image generation, speech-to-text, and vision with 8-signal scoring, mDNS auto-discovery, and an 8-tab real-time dashboard. Apple Silicon Macs get extras: an MLX backend, native image generation, and speech-to-text. Two commands to set up, zero config files. pip install ollama-herd or brew install ollama-herd.

The Players

GitHub star counts as of September 2026.

LoLLMs Hub (formerly ollama_proxy_server, ~650 stars), a Python gateway that has grown well past a proxy. It fronts Ollama, vLLM, llama.cpp, and OpenAI-compatible APIs; routes by priority, least-loaded backend, or keyword and semantic rules; and adds user management, analytics, HTTPS, and hierarchical hubs. Its strength is access control and policy routing; it does not read machine conditions such as memory pressure.

ollama_load_balancer (~26 stars), a lean Rust balancer. For chat it races the request across several health-scored backends and streams back the fastest, trading duplicate compute for low tail latency. See the full comparison.

llama-swap (~5.8K stars), the most popular tool here, and not an Ollama add-on. It is a Go proxy that starts and stops model servers on demand (llama.cpp and its forks, vLLM, stable-diffusion.cpp, ComfyUI, audio servers) behind OpenAI and Anthropic APIs, with API keys, a web UI, Prometheus metrics, and a swap matrix for running several models at once. It solves "more models than memory" on one machine, and since late 2025 it can also forward listed models to peers on other machines, preferring one that already has the model loaded or spilling over when local capacity is full. llama.cpp's own server now has a router mode that covers part of the single-machine ground. Our llama-swap vs Ollama guide compares all three.

SOLLOL (~6 stars), The most ambitious of the small tools. Context-aware scoring, priority queues, auto-discovery, and a dashboard. Closest to Herd's vision but far less mature, with minimal community validation.

OLOL (~47 stars), a Python gRPC clustering system for Ollama. It load balances with session affinity and can split a large model's layers across servers, but it does not score machine conditions. Last commit August 2025. See the full comparison.

Hive, Task queue for Ollama. Queues inference requests and dispatches them to available backends. More of a job scheduler than a router.

OllamaFlow, Label-based routing layer, now archived (its successor is a heavier multi-tenant control plane, Conductor). Included here for lineage.

The current-generation slice tools, a newer wave has more traction: ollamaMQ (per-user fair-share queues for shared labs, with load-aware backend selection), ollamafarm (multi-Ollama management with offline detection), and Conductor (OllamaFlow's successor, an alpha multi-tenant control plane whose optional policies can use host telemetry). Each is strong at its slice. None combine device-condition scoring, learned availability, multimodal routing, and mDNS discovery the way Herd does, and for the whole-problem comparison see Olla, NVIDIA PAIR, and LocalAI's distributed mode.

Conceptual: where each tool sits (a tool can span layers)
Clientschat UIs, coding agents, apps
Open WebUI Claude Code Codex Your apps and agents
Gateway / edgeauth, rate limits, cloud providers
Inference serverruns the model on one machine
One model, many machinessplits a model too big for one
Hardwarethe machines themselves
Macs Linux servers Windows PCs NVIDIA GPU servers
Most proxy tools sit in Herd's layer and each solves one slice of it; llama-swap is mainly a single-machine engine manager that can also forward to peers you list.

Feature Comparison

FeatureLoLLMs Hubollama_load_balancerSOLLOLllama-swapOllama Herd
Multi-instance routingYesYesYesYes (peers you list in config)Yes
Load balancingPriority, least-loaded, rulesHealth-scored raceScoredWarm-first or spillover, in config order8-signal scoring
Auto-discovery (mDNS)NoNoYesNoYes
API key authYesNoNoYesNo
Priority queuesNoNoYesYes (per-model priority)Yes
Real-time dashboardAdmin dashboard + analyticsNoBasicWeb UI + metrics8-tab dashboard
Capacity learningNoNoNoNoYes
Memory trackingNoNoPartialNo (you configure which models fit)System RAM only (fleet-wide)
Dynamic context optimizationNoNoNoNoYes
Smart benchmarkNoNoNoNoYes
Model swappingNoNoNoYesNo (routes instead)
Multimodal supportLLM, visionLLM onlyLLM onlyLLM, image, audio (any server it can launch)LLM, embed, img, STT
OpenAI API compatYesYesPartialYesYes
Ollama API compatYesPartialPartialNoYes
LanguagePythonRustPythonGoPython
Test coverageMinimalMinimalMinimalExtensive (1,000+ tests)1200+ tests
Health checksBasicHealth scoresBasicReadiness check per server30 checks
Active maintenanceSporadicLowLowActiveActive

What Each Does Well

  • LoLLMs Hub, The most complete access-control story here: users, API keys, HTTPS, analytics, and rule-based routing across several backend types. Good for teams sharing machines where who-can-use-what matters more than machine conditions.
  • ollama_load_balancer, Lowest tail latency on idle, identical machines, because it races each chat request and keeps the fastest stream. The cost is running every prompt more than once.
  • llama-swap, The best answer to "more models than memory" on one machine, for any server it can launch, with auth and a UI built in. Different problem than fleet routing.
  • SOLLOL, The most conceptually similar to Herd. Attempts scoring, queuing, discovery, and dashboard in one package. Shows that someone else sees the same problem space. Limited by tiny community and early maturity.
  • OLOL, The only tool here that can split one large model across machines, over gRPC. If you need a model bigger than any single box, it is the one to look at (or exo).
  • Hive, Task queue model is interesting for batch workloads. If you're processing 1,000 documents and want to queue them across backends, Hive's approach makes sense.
  • OllamaFlow, Label-based routing is useful for heterogeneous fleets where you want to manually direct certain models to certain hardware. Static but explicit.

Where They Fall Short

  • One signal at a time. Most route by round-robin or by which backend has the model. A few look further: ollamaMQ and llama-swap's warm selector prefer a backend with the model already loaded, and ollama_load_balancer keeps health scores. None besides SOLLOL's early attempt combine loaded models, memory fit, queue depth, latency history, and context fit in one score per request.
  • No sense of the machine's owner. They treat backends as always available. Herd drops a machine under memory pressure, and can opt in to learning when each machine is usually free (a 168-slot weekly model) and to pausing a Mac during meetings.
  • Little multimodal routing across machines. llama-swap can front image and audio servers, and LoLLMs Hub handles vision, but none route embeddings, image generation, and speech-to-text across a fleet. Herd routes all five model types.
  • No context protection. They pass context sizes through as-is, so a client asking for a different context can force a model reload. Herd strips those overrides by default when the loaded context already covers them.
  • Few fleet dashboards. LoLLMs Hub and llama-swap have real web UIs (analytics, logs, request captures) and SOLLOL has a basic one; the rest are CLI-only or status endpoints. Herd's 8-tab dashboard is built around the fleet: node health, per-request routing decisions, model distribution, and benchmark results.
  • Uneven testing. Most of the smaller tools have few tests; llama-swap is the exception, with over 1,000. Herd has 1200+ tests and 30+ health checks. For infrastructure that routes your AI workloads, test coverage matters.
  • No smart benchmark. None measure actual device inference capability. Herd's smart benchmark tests real throughput on each device.

The "DIY Alternative"

In theory, you could wire together LoLLMs Hub for auth and policy routing, ollama_load_balancer for dispatch, and llama-swap for model memory management on a llama.cpp machine. This gives you auth, distribution, and model swapping. But you still lack:

  • Intelligent scoring (model warmth, memory fit, queue depth, latency history)
  • Auto-discovery (manual config for every instance)
  • Multimodal routing (LLM only)
  • Capacity learning (static assumptions about each backend)
  • Unified dashboard (three separate tools, three separate monitoring stories)
  • Dynamic context optimization
  • Health checks and test coverage across the integration points

And you're maintaining three tools from three different authors with three different update cycles and no guarantee they play well together. Herd is the integrated answer. One install, zero config, all the intelligence.

When to Choose a Proxy Tool Instead

  • LoLLMs Hub, You need user accounts, API keys, or policy routing and Herd doesn't provide them. Auth is a genuine gap in Herd.
  • ollama_load_balancer, Your machines are identical and mostly idle, and you care about tail latency more than total capacity.
  • llama-swap, You have one machine and many models, especially llama.cpp or vLLM ones. Its peers feature also reaches other machines you list by hand. Herd is the fit when the machines run Ollama and you want them found and scored automatically; the two cannot be stacked today.
  • OllamaFlow, You want explicit manual control over which models go where, with human-defined labels rather than automatic scoring.

When to Choose Ollama Herd

  • You have multiple devices and want intelligent routing across all of them
  • You want zero-config setup, mDNS discovery, capacity learning, dynamic optimization
  • You need multimodal routing beyond just LLMs
  • You want a dashboard to see what your fleet is doing
  • You need production-grade reliability with real test coverage
  • You want routing that gets smarter over time as it learns your fleet

Bottom Line

The existence of 7+ Ollama routing tools validates that the problem is real. People with multiple Ollama instances need a way to coordinate them. Most of them solve one slice of it, and they are worth using for that slice.

Several of these tools beat Herd at their own slice: llama-swap at single-machine model swapping, LoLLMs Hub at access control, ollamaMQ at fairness between users. What none of them do is treat each machine as a workstation with conditions (warm models, memory fit, memory pressure, the owner's schedule) and find those machines with zero config. If that is your problem, Herd is built for it; if your problem is one of theirs, use theirs.

Getting Started

pip install ollama-herd    # or: brew install ollama-herd
herd                       # start the router
herd-node                  # on each device

Frequently Asked Questions

Can I use llama-swap alongside Ollama Herd?

Side by side, yes; stacked, no. llama-swap starts and stops model servers such as llama.cpp and vLLM on its machine, can forward listed models to peers, and exposes OpenAI and Anthropic APIs. Herd nodes serve Ollama and mlx-lm, so llama-swap cannot sit behind Herd as a node today. Run each as its own endpoint.

Does Ollama Herd support API key authentication?

Not yet. This is a genuine gap compared to LoLLMs Hub. If API key gating is your primary requirement, LoLLMs Hub (formerly ollama_proxy_server) handles that today. Herd is designed for trusted local networks.

Why not wire together multiple proxy tools?

You could combine LoLLMs Hub for auth, ollama_load_balancer for dispatch, and llama-swap for model management. But you still lack intelligent scoring, auto-discovery, multimodal routing, capacity learning, and a unified dashboard. You are also maintaining three tools from three authors with no integration guarantees.

How does Herd compare to SOLLOL?

SOLLOL is conceptually the closest to Herd, attempting scored routing, priority queues, and auto-discovery. However, it has minimal community validation (~6 stars), limited test coverage, and no multimodal support. Herd has 1200+ tests, 30+ health checks, and routes 5 model types.

Is ollama_load_balancer faster than Herd because it is written in Rust?

The proxy layer overhead is minimal in both. ollama_load_balancer adds very little latency of its own, but the routing decision itself (1–2ms) is negligible compared to actual inference time (seconds). Herd's scoring intelligence more than compensates by sending requests to the node that will complete fastest.

See Also