Ollama Herd vs exo
exo splits one model across your Macs. Herd routes your whole agent workload across every machine you own, reliably, in production. Different problems, and they work great together.
What is exo?
exo (~48K GitHub stars as of September 2026) is an open-source distributed inference framework built by EXO Labs. It splits a single large AI model across multiple devices using tensor and pipeline parallelism, allowing you to run models that would not fit on any one machine. exo targets Apple Silicon clusters connected via Thunderbolt or network, and exposes an OpenAI-compatible API endpoint. It is the leading project for model-sharding across consumer hardware.
What is Ollama Herd?
Ollama Herd is an open-source smart AI router that turns the machines you already own (Apple Silicon Macs, Linux servers, and Windows PCs, with or without NVIDIA GPUs) into one endpoint. It routes LLMs, embeddings, image generation, speech-to-text, and vision with 8-signal scoring, mDNS auto-discovery, and an 8-tab real-time dashboard. Apple Silicon Macs get extras: an MLX backend, native image generation, and speech-to-text. Two commands to set up, zero config files. pip install ollama-herd or brew install ollama-herd.
How exo Works
exo takes a model too large for one machine and splits it into shards distributed across multiple devices using two parallelism strategies:
- Pipeline parallelism: Splits model layers sequentially across devices. Device A runs layers 1-20, device B runs layers 21-40. Simple but sequential, each device waits for the previous one.
- Tensor parallelism: Splits individual layers across devices so they compute in parallel. Higher throughput but requires fast interconnect (Thunderbolt/RDMA).
exo auto-discovers peers on the network, builds a topology, and exposes an OpenAI-compatible API endpoint. Underneath, it uses MLX on Apple Silicon and supports heterogeneous device mixes (different Mac models, different memory sizes).
Performance numbers: ~1.8x speedup on 2 devices, ~3.2x speedup on 4 devices for single-request latency. Multi-request throughput scales better, a 3-device cluster handles ~2.2x the tokens per second of a single device.
Feature Comparison
| Feature | exo | Ollama Herd |
|---|---|---|
| Core approach | Model sharding (tensor/pipeline parallelism) | Request routing (8-signal scoring) |
| Primary use case | Run models too large for one device | Route requests to best available device |
| Model types | LLM-first (image gen behind a flag; no speech or embeddings endpoints) | LLMs, embeddings, image gen, STT, vision |
| Device discovery | Automatic peer discovery | mDNS auto-discovery |
| API compatibility | OpenAI Chat + Claude/Anthropic + Responses + Ollama | OpenAI + Ollama + Anthropic Messages (native, no format conversion) |
| Platforms | macOS GPU (Apple Silicon); Linux CPU-only, GPU in development | macOS + Linux + Windows |
| Production posture | 2026 reviews cite gaps in security, fault tolerance, ops tooling | 30+ health checks, auto-retry, holding queues, trace store |
| Backend | MLX, tinygrad | Ollama + MLX + fastembed |
| Queue management | None, single model focus | Per-node queue depth tracking |
| Health monitoring | Basic peer status | 30+ health checks, 8-signal scoring |
| Load balancing | N/A (all devices serve one model) | Adaptive capacity learning per model/node |
| Dashboard | Minimal web UI | 8-tab dashboard (fleet, models, routing, benchmarks) |
| Benchmarking | Manual | Smart benchmark with statistical analysis |
| Context optimization | None | Dynamic context window optimization |
| Meeting detection | None | Detects video calls on macOS, reduces load on busy Macs |
| Multi-model serving | One model at a time across the cluster | Many models across many nodes simultaneously |
| Setup | pip install exo + run on each device | pip install ollama-herd on one machine |
| Config required | None (auto-topology) | None (mDNS auto-discovery) |
| Interconnect | Benefits from Thunderbolt/RDMA | Standard network (WiFi or Ethernet) |
| Tests | Limited | 1200+ tests, 30+ health checks |
| License | Apache-2.0 | MIT |
Where exo Wins
- Running models that don't fit on one machine. If you need to run Llama 3.1 405B and your biggest Mac has 192GB RAM, exo is the only option. Herd can't help, it routes to nodes, it doesn't split models.
- Maximum single-request throughput for huge models. Tensor parallelism across Thunderbolt-connected Macs gives near-linear speedup for large model inference. A single request to a 70B model is faster on 2 exo nodes than on 1 Herd node.
- Simplicity of mental model for single-model use. If your entire use case is "run one big model as fast as possible," exo's model is simpler: shard it and go.
- Thunderbolt/RDMA optimization. exo has invested heavily in low-latency device-to-device communication, microsecond-scale latency with RDMA over Thunderbolt 5, aligned with Apple's WWDC26 distributed MLX stack.
Where Ollama Herd Wins
- Production reliability. Multiple 2026 reviews of exo converge on "not production-ready yet", gaps in security, fault tolerance, and operational tooling. Herd was built the other direction: 30+ automated health checks, transparent auto-retry before the first chunk, holding queues, adaptive capacity learning, live observability. For a team that needs a reliable local-AI endpoint today, this is the decisive difference.
- Platform reach. macOS + Linux (with GPU) + Windows. exo is macOS-GPU-focused; Linux runs CPU-only with GPU support still in development. If your fleet includes Linux GPU boxes or Windows workstations, Herd can use them.
- Multi-model, multi-user workloads. Real teams don't run one model. They run coding assistants, embeddings for RAG, image generation, and speech-to-text, often simultaneously. Herd routes each request to the best node for that specific model.
- Intelligent routing. 8-signal scoring (model warmth, memory fit, queue depth, estimated wait, role affinity, availability trend, context fit, session affinity) means requests go to the right machine, not just any machine.
- Multimodal support. Five model types (LLMs, embeddings, image gen, STT, vision) with type-aware routing. exo is LLM-first.
- Operational visibility. 8-tab dashboard showing fleet status, model distribution, routing decisions, and benchmark results. Know what's happening across your fleet at a glance.
- Adaptive capacity learning. Herd learns each node's actual performance per model over time and adjusts routing. No manual tuning needed.
- Meeting detection (macOS). Automatically reduces load on Macs running video calls. Small feature, huge quality-of-life improvement for real teams.
- Queue management. Tracks per-node queue depth and avoids piling requests on busy nodes. exo has no concept of request queuing, it's one model, one cluster.
- Ollama ecosystem. Works with the full Ollama model library and tooling. No special model format or conversion needed.
- Setup simplicity at fleet scale. Run
herdon one machine andherd-nodeon the others; mDNS finds them with no config files. exo also runs a process on every participating device, and its models only work while the whole cluster is up.
Why They're Complementary
exo and Herd solve different problems, and today they run side by side as two endpoints:
An exo cluster running a large sharded model exposes its own OpenAI-compatible endpoint. Herd does not route to it (Herd nodes serve Ollama and MLX models), so the two run side by side:
- Small/medium models get routed across individual nodes by Herd
- Huge models get served by the exo cluster at its own endpoint
- Two endpoints, and each application points at the one that fits its workload
When to Choose
| Scenario | Choose |
|---|---|
| Need to run a model too large for any single device | exo |
| Team of 2-10 people sharing a fleet of machines (Mac, Linux, Windows) | Ollama Herd |
| Multiple model types (LLM + embeddings + image gen) | Ollama Herd |
| Maximum throughput for one huge model | exo |
| Need operational visibility and health monitoring | Ollama Herd |
| Thunderbolt-connected Mac cluster for one workload | exo |
| WiFi/Ethernet fleet serving diverse workloads | Ollama Herd |
| Want both large model access and smart routing | Both, side by side |
Bottom Line
exo is a distributed compute layer, it makes small machines act like one big machine. Ollama Herd is a distributed routing layer, it makes many machines serve many users intelligently. They don't compete; they solve adjacent problems.
The typical exo user has 2-4 Macs hardwired together running one frontier model. The typical Herd user has 3-8 machines on a network (Macs, often alongside Linux or Windows boxes) running a dozen different models for a team. When you need both (large model access + fleet routing), run them side by side as two endpoints.
Getting Started
Ollama Herd works alongside exo, you can try it without changing your existing setup. If you already have Ollama running on your machines, Herd discovers them automatically and starts routing in under two minutes.
pip install ollama-herd # or: brew install ollama-herd
herd # start router
herd-node # on each device
Switching from exo to Ollama Herd
exo and Ollama Herd solve different problems, so you likely won't "switch", you may use both. But if you want fleet routing instead of model sharding:
- Install Ollama on each device, exo uses its own runtime, Herd uses Ollama. Run
ollama serveand pull your models on each machine. - Install Ollama Herd,
pip install ollama-herdon your router machine, thenherdto start. - Start node agents, run
herd-nodeon each device. They discover the router automatically via mDNS.
Your existing tools just need to point at http://router-ip:11435 instead of the exo endpoint. If you still want to run one massive model across devices, keep exo running for that model as its own endpoint and send only those requests to it; everything else goes through Herd.
FAQ
Is Ollama Herd a good alternative to exo?
They solve different problems. exo shards one large model across devices so you can run models that do not fit on a single machine. Ollama Herd routes requests across a fleet of devices, picking the best node for each request. If you need multi-model routing rather than single-model sharding, Herd is the right choice.
Can I use Ollama Herd with exo?
Side by side, yes. As a Herd backend, not today: Herd nodes serve Ollama and MLX (mlx-lm) models, and Herd does not route to an exo cluster's endpoint. Keep exo serving the one huge model at its own endpoint and send everything else through Herd.
How does Ollama Herd compare to exo for running multiple models?
Herd is built for multi-model workloads. It routes LLMs, embeddings, image generation, speech-to-text, and vision across your fleet simultaneously, picking the best device for each request type. exo focuses on running one model at a time across the cluster.
Does Ollama Herd require Thunderbolt connections?
No. Ollama Herd works over standard WiFi or Ethernet. It routes requests to devices rather than sharding model layers, so it does not need the high-bandwidth interconnect that exo benefits from.
Is Ollama Herd free?
Yes. Ollama Herd is open-source under the MIT license. No paid tiers, no API keys, no subscriptions.
See Also
- Ollama Herd vs Apple MLX Distributed, Apple's first-party take on the same sharding problem exo solves
- Ollama Herd vs GPUStack, enterprise GPU cluster manager with multi-backend support
- Ollama Herd vs vLLM, high-throughput LLM serving engine with PagedAttention
- Ollama Herd vs Single Ollama, why fleet routing beats running one instance