Start Here
- Quickstart, Install to first routed request in 60 seconds. Create a fleet from any mix of Mac, Linux, and Windows machines with two commands, send a request, and see it land on the right machine.
- Core Concepts, The mental model behind Ollama Herd. Nodes, heartbeats, scoring signals, queues, capacity modes, and how they fit together.
Learn the System
- Routing Engine, How the 5-stage scoring pipeline eliminates bad candidates, scores survivors across 8 signals, and picks a winner for every request.
- Adaptive Capacity, How your fleet learns when each device has spare compute (opt-in). Weekly behavioral models, meeting detection on macOS, app fingerprinting, and memory ceilings.
Put It to Work
- Claude Code CLI, How to run Claude Code with local models through an Ollama router that spans every machine you own. Native Anthropic Messages API, three-layer context management, per-tier model routing, tool-schema fixup. Fixes the "breaks at 30K tokens" failure mode on local Qwen3-Coder models.
- Claude Code Costs, How Claude Code limits and token usage work, which usage monitors help, and how routing routine work to local models cuts the real burn.
- Integrations, Connect Ollama Herd to Open WebUI, LangChain, CrewAI, OpenClaw, Aider, Continue.dev, LlamaIndex, and any OpenAI-compatible client.
- Deployment, Multi-node setup, monitoring, log analysis, health checks, graceful drain, and production tips.
- Load Balancing Ollama, How to distribute requests across multiple Ollama instances, from HAProxy configs to zero-config routing that avoids cold-load stalls.
- Codex CLI, How to run Codex CLI and the Codex desktop app with local models through Ollama, on any GPU: Macs, Linux, or Windows machines. One config block, no model map, verified end to end on a Mac fleet.
- OpenClaw, Run OpenClaw agents on your own machines. Install the Herd skill from ClawHub, point OpenClaw at the fleet, and give long agent sessions the context headroom they need.
- Mac Cluster, The two different things people mean by running LLMs across multiple Macs, routing vs sharding, which one you actually need, and where Herd is the wrong tool.
- Local Models, Which open LLMs, vision, and coding models actually run on Apple Silicon, tested first-hand on a real Mac fleet, including the ones that do not run locally yet.
- Mac Memory, How much unified memory each model needs, the sizing math, and which Mac runs a 30B, 70B, 120B, or 480B model.
- Load Ceiling, Why a 418 GB model fits on a 512 GB Mac's disk but won't load through MLX, the iogpu.wired_limit_mb footgun, and why your setup may differ.
- Routing vs Sharding, Splitting one model across machines and routing requests between them solve opposite problems. The memory math and how to tell which you need.
- MLX vs Ollama, We benchmarked both backends head to head on an M3 Ultra. They tie on speed, so the real choice is operational. Here is how to pick.
- Multimodal, Herd routes five model types from one endpoint: LLMs, embeddings, image generation, speech-to-text, and vision, with the image and speech extras on Apple Silicon. The pattern behind all of them.
- Image Generation, Route Flux and Stable Diffusion image generation through one endpoint: mflux and DiffusionKit on Apple Silicon Macs, plus Ollama-native image models.
- Troubleshooting, Why Claude Code breaks around 30K tokens, why models loop on tool calls, return empty, or hit Metal out-of-memory, and how to fix each.
- Local AI Cost, When running AI on hardware you already own actually beats cloud API pricing, the break-even math, and why fleet scale flips the answer.
- Desktop Chat Apps, Connect Page Assist, Jan, Chatbox, AnythingLLM, Cherry Studio, LibreChat, and other chat apps to your Herd router, with exact settings and known incompatibilities.
- Ollama vs LM Studio vs vLLM, Which local model server to run on a Mac in 2026: engines, concurrency, agent-CLI APIs, headless setups, and multiple machines.
- oMLX vs Ollama, oMLX's SSD KV cache and batching versus Ollama's simplicity, for Claude Code and Codex on a Mac.
- Open WebUI, Connect Open WebUI to several Ollama machines through one Herd endpoint with model-aware routing instead of random backend selection.
- Concurrency, Tune OLLAMA_NUM_PARALLEL, loaded models, and queue depth on one machine without OOM errors, and when to scale across a fleet.
- Cold Starts, Why random load balancing across Ollama servers throws away warm KV caches, and what cache-aware routing does instead.
- llama-swap vs Ollama, Three ways to run more models than fit in memory (Ollama, llama-swap, llama.cpp router mode), and what changes at two machines.
- RAG at Scale, Keep document-ingestion embeddings from blocking interactive chat by separating embedding and generation across a fleet.
- Reranking, Ollama has no rerank endpoint. How the cross-encoder step that fixes RAG ordering runs on the fastembed server your embedding nodes already have.
- Remote Access, Reach your Ollama fleet from anywhere over Tailscale without exposing a single inference node to the public internet.
- n8n, Run n8n automations, agents, and RAG against a fleet of local machines through one endpoint, with concurrency that spreads across nodes.
- API Reference, Every endpoint with request/response schemas, headers, error codes, and curl examples.
New here? Two commands and your fleet is live.