Who Uses Ollama Herd

Real scenarios showing how developers, teams, and enthusiasts turn the machines they already own (Macs, Linux servers, Windows PCs, NVIDIA or not) into one smart AI endpoint.

Solo Developer with 2+ Machines

The pain: You have a desktop for heavy work (a Mac Studio, or a Linux or Windows tower with an NVIDIA card) and a laptop for portability. When you're running Aider or Continue.dev on the laptop, it heats up, fans spin, and inference slows down. Meanwhile the desktop sits idle. You keep SSH-ing between machines or manually switching base URLs.

With Herd: Point all your tools at http://router-ip:11435. The desktop handles the heavy models, the laptop handles quick tasks (7B–14B). When you're at your desk with both machines free, they share the load. Turn on capacity learning on a MacBook and a Zoom call pauses it automatically, sending requests to the desktop.

Example Setup

  • Mac Studio (192GB): Llama 3.3 70B + DeepSeek Coder 33B, always loaded
  • Or a Linux tower with an RTX 4090: Qwen3-Coder 30B, sized to fit the card
  • MacBook Pro (36GB) or Windows laptop: Qwen 2.5 7B + Nomic Embed, for lightweight tasks and RAG
  • Tools: Aider, Continue.dev, Open WebUI, all pointed at one URL

Agent-Heavy Workflows

The pain: You're running CrewAI crews, LangChain chains, or OpenClaw agents that fire dozens of concurrent LLM requests. A single Ollama instance queues them all sequentially. A 5-agent crew that should take 2 minutes takes 10 because every request waits in line.

With Herd: Concurrent requests fan out across your fleet. Agent #1 goes to the Mac Studio, agent #2 goes to the Linux GPU box, agent #3 goes to the Windows desktop. Throughput scales linearly with machines. Auto-retry means agent failures don't crash the crew, the router re-routes to the next best node.

Example Setup

  • 3 devices: Mac Studio + Linux box with an NVIDIA GPU + Windows desktop
  • Models: One large reasoning model (70B), one fast agent model (7B–14B), one embedding model
  • Framework: CrewAI / LangChain / OpenClaw, all using OpenAI SDK with base_url pointed at Herd

Small Team / Office

The pain: Your team has a handful of laptops (some Macs, some Windows) and one shared machine that can run the big models, a Mac Studio or a Linux GPU server. Everyone runs Ollama locally, but nobody's own machine is powerful enough for the big models. People share the big box by manually coordinating who's using it. No visibility into who's queued where.

With Herd: One router, all machines as nodes. Everyone points their tools at the same URL. The router handles contention, no manual coordination. The dashboard shows who's using what, queue depths, and per-tag analytics (via request tagging). The shared server handles the big models, personal laptops handle lightweight tasks.

Example Setup

  • Router: On the shared Mac Studio or Linux server
  • Nodes: The team's MacBooks and Windows laptops (each running herd-node)
  • Analytics: Per-app tagging, each developer's tools tagged for tracking
  • Dashboard: On a shared monitor or bookmarked URL

Home Lab Enthusiast

The pain: You've accumulated hardware, a Mac Mini, an older MacBook, maybe a Linux box with an NVIDIA GPU. You want a unified local AI setup but every tool assumes a single machine. Managing multiple Ollama instances manually is tedious.

With Herd: Every device joins the fleet automatically via mDNS. Mix and match platforms: macOS, Linux, Windows. The router knows which models and services each device has and routes accordingly. The NVIDIA box serves the models you pulled onto it, with Ollama running them on CUDA. Image generation routes to the Mac with mflux installed. Embeddings route to whichever node has the model loaded.

Example Setup

  • Mac Mini M2 (24GB): Small models + embeddings
  • Linux box with RTX 4090: Models that fit in the card's 24GB, run on CUDA by Ollama
  • Old MacBook (16GB): Lightweight agent tasks when it's not being used
  • Discovery: All found automatically, no config files

Multimodal AI Pipeline

The pain: You need LLM inference, embeddings for RAG, image generation, and speech-to-text. Each service runs on a different port, different machine, different API. Your application code is full of conditional routing logic.

With Herd: One endpoint handles all five model types. The router knows which nodes can serve which modality and routes accordingly. Your app talks to one URL for everything.

Example Setup

  • LLM: POST /v1/chat/completions or POST /api/chat, routed to best available node
  • Embeddings: POST /api/embed, routed to node with embedding model loaded
  • Image gen: POST /api/generate-image, routed to an Apple Silicon node with mflux or DiffusionKit (Ollama native image models also work)
  • Speech-to-text: POST /api/transcribe, routed to an Apple Silicon node with MLX and Qwen3-ASR
  • All through: http://router-ip:11435

Is This For You?

Herd is a great fit if:

  • You have 2 or more machines that can run Ollama (any mix of macOS, Linux, and Windows)
  • You run AI tools concurrently (agents, coding assistants, chat)
  • You want zero-config setup (no Docker, no Kubernetes, no YAML)
  • You care about privacy and want everything local
  • You're tired of model thrashing on a single machine

Herd is probably overkill if:

  • You have exactly one machine and no plans to add more
  • You run one model at a time with no concurrency needs
  • You're happy with single-machine Ollama performance

Getting started takes 60 seconds:

pip install ollama-herd
herd                    # on your router machine
herd-node               # on each device

Find your setup, then build it in 60 seconds

Whatever your fleet looks like, Herd discovers it over mDNS and starts routing automatically.