Guides

Everything you need to go from install to production fleet, on Macs, Linux servers, and Windows PCs. Start at the top and work down, or jump to what you need.

Start Here

  • Quickstart, Install to first routed request in 60 seconds. Create a fleet from any mix of Mac, Linux, and Windows machines with two commands, send a request, and see it land on the right machine.
  • Core Concepts, The mental model behind Ollama Herd. Nodes, heartbeats, scoring signals, queues, capacity modes, and how they fit together.

Learn the System

  • Routing Engine, How the 5-stage scoring pipeline eliminates bad candidates, scores survivors across 8 signals, and picks a winner for every request.
  • Adaptive Capacity, How your fleet learns when each device has spare compute (opt-in). Weekly behavioral models, meeting detection on macOS, app fingerprinting, and memory ceilings.

Put It to Work

  • Claude Code CLI, How to run Claude Code with local models through an Ollama router that spans every machine you own. Native Anthropic Messages API, three-layer context management, per-tier model routing, tool-schema fixup. Fixes the "breaks at 30K tokens" failure mode on local Qwen3-Coder models.
  • Claude Code Costs, How Claude Code limits and token usage work, which usage monitors help, and how routing routine work to local models cuts the real burn.
  • Integrations, Connect Ollama Herd to Open WebUI, LangChain, CrewAI, OpenClaw, Aider, Continue.dev, LlamaIndex, and any OpenAI-compatible client.
  • Deployment, Multi-node setup, monitoring, log analysis, health checks, graceful drain, and production tips.
  • Load Balancing Ollama, How to distribute requests across multiple Ollama instances, from HAProxy configs to zero-config routing that avoids cold-load stalls.
  • Codex CLI, How to run Codex CLI and the Codex desktop app with local models through Ollama, on any GPU: Macs, Linux, or Windows machines. One config block, no model map, verified end to end on a Mac fleet.
  • OpenClaw, Run OpenClaw agents on your own machines. Install the Herd skill from ClawHub, point OpenClaw at the fleet, and give long agent sessions the context headroom they need.
  • Mac Cluster, The two different things people mean by running LLMs across multiple Macs, routing vs sharding, which one you actually need, and where Herd is the wrong tool.
  • Local Models, Which open LLMs, vision, and coding models actually run on Apple Silicon, tested first-hand on a real Mac fleet, including the ones that do not run locally yet.
  • Mac Memory, How much unified memory each model needs, the sizing math, and which Mac runs a 30B, 70B, 120B, or 480B model.
  • Load Ceiling, Why a 418 GB model fits on a 512 GB Mac's disk but won't load through MLX, the iogpu.wired_limit_mb footgun, and why your setup may differ.
  • Routing vs Sharding, Splitting one model across machines and routing requests between them solve opposite problems. The memory math and how to tell which you need.
  • MLX vs Ollama, We benchmarked both backends head to head on an M3 Ultra. They tie on speed, so the real choice is operational. Here is how to pick.
  • Multimodal, Herd routes five model types from one endpoint: LLMs, embeddings, image generation, speech-to-text, and vision, with the image and speech extras on Apple Silicon. The pattern behind all of them.
  • Image Generation, Route Flux and Stable Diffusion image generation through one endpoint: mflux and DiffusionKit on Apple Silicon Macs, plus Ollama-native image models.
  • Troubleshooting, Why Claude Code breaks around 30K tokens, why models loop on tool calls, return empty, or hit Metal out-of-memory, and how to fix each.
  • Local AI Cost, When running AI on hardware you already own actually beats cloud API pricing, the break-even math, and why fleet scale flips the answer.
  • Desktop Chat Apps, Connect Page Assist, Jan, Chatbox, AnythingLLM, Cherry Studio, LibreChat, and other chat apps to your Herd router, with exact settings and known incompatibilities.
  • Ollama vs LM Studio vs vLLM, Which local model server to run on a Mac in 2026: engines, concurrency, agent-CLI APIs, headless setups, and multiple machines.
  • oMLX vs Ollama, oMLX's SSD KV cache and batching versus Ollama's simplicity, for Claude Code and Codex on a Mac.
  • Open WebUI, Connect Open WebUI to several Ollama machines through one Herd endpoint with model-aware routing instead of random backend selection.
  • Concurrency, Tune OLLAMA_NUM_PARALLEL, loaded models, and queue depth on one machine without OOM errors, and when to scale across a fleet.
  • Cold Starts, Why random load balancing across Ollama servers throws away warm KV caches, and what cache-aware routing does instead.
  • llama-swap vs Ollama, Three ways to run more models than fit in memory (Ollama, llama-swap, llama.cpp router mode), and what changes at two machines.
  • RAG at Scale, Keep document-ingestion embeddings from blocking interactive chat by separating embedding and generation across a fleet.
  • Reranking, Ollama has no rerank endpoint. How the cross-encoder step that fixes RAG ordering runs on the fastembed server your embedding nodes already have.
  • Remote Access, Reach your Ollama fleet from anywhere over Tailscale without exposing a single inference node to the public internet.
  • n8n, Run n8n automations, agents, and RAG against a fleet of local machines through one endpoint, with concurrency that spreads across nodes.
  • API Reference, Every endpoint with request/response schemas, headers, error codes, and curl examples.

New here? Two commands and your fleet is live.