API Reference

Every endpoint with request/response schemas, headers, error codes, and examples.

The router runs on port 11435 by default. All endpoints accept JSON bodies and return JSON responses.

OpenAI-Compatible Endpoints

POST /v1/chat/completions

Chat completions with streaming and non-streaming support.

Request:

{
  "model": "llama3.3:70b",
  "messages": [
    {"role": "system", "content": "You are helpful."},
    {"role": "user", "content": "Hello!"}
  ],
  "stream": true,
  "temperature": 0.7,
  "max_tokens": 1024,
  "fallback_models": ["qwen2.5:32b", "qwen2.5:7b"],
  "metadata": {"tags": ["my-app"]},
  "user": "alice"
}
Field Type Default Description
model string required Model name
messages array [] Chat messages in OpenAI format
stream boolean false Enable SSE streaming
temperature float 0.7 Sampling temperature
max_tokens integer null Maximum tokens to generate
fallback_models array [] Backup models if primary unavailable
metadata.tags array [] Tags for per-tag analytics
user string null Stored as user:<value> tag

Response headers:

Header Description
X-Fleet-Node Node that handled the request
X-Fleet-Score Winning routing score
X-Fleet-Fallback Fallback model used (if applicable)
X-Fleet-Retries Retry count (if retries occurred)
X-Fleet-Affinity matched when routed back to the node already holding this conversation, new otherwise. Omitted when affinity did not apply
X-Fleet-Context-Overflow Context overflow warning
X-Thinking-Tokens Thinking tokens (chain-of-thought models, non-streaming)
X-Output-Tokens Output tokens (chain-of-thought models, non-streaming)
X-Budget-Used completion_tokens/num_predict budget check (non-streaming)
X-Done-Reason stop (natural) or length (budget exhausted)

Errors: 400 missing model, 404 model not found, 503 no node available.

Example:

curl http://localhost:11435/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model": "llama3.3:70b", "messages": [{"role": "user", "content": "Hello!"}], "stream": false}'

POST /v1/responses

OpenAI Responses API, the wire protocol OpenAI Codex speaks (Codex dropped Chat Completions in early 2026). Point Codex at the fleet and it routes across your machines. Unmapped model ids auto-route to the best loaded model, so no model map is required.

Request (standard Responses API shape):

{
  "model": "gpt-5-codex",
  "input": "Refactor this function.",
  "stream": true
}

Streams Server-Sent Events in Responses format. See the Codex guide for the full CLI and desktop-app setup.


POST /v1/images/generations

OpenAI-compatible image generation.

Request:

{
  "model": "z-image-turbo",
  "prompt": "a cat sitting on a laptop",
  "size": "1024x1024",
  "response_format": "b64_json"
}
Field Type Default Description
model string required Image model name
prompt string required Image description
size string 1024x1024 Dimensions (WIDTHxHEIGHT)
response_format string b64_json b64_json or url (raw PNG)
steps integer model default Inference steps
guidance float model default Guidance scale
seed integer random Reproducibility seed

Errors: 400 missing fields, 404 model unavailable, 502 generation failed, 503 disabled.


GET /v1/models

List all models across the fleet (LLM, image, embedding).

Response:

{
  "object": "list",
  "data": [
    {"id": "llama3.3:70b", "object": "model", "created": 1710000000, "owned_by": "ollama"},
    {"id": "z-image-turbo", "object": "model", "created": 1710000000, "owned_by": "mflux"}
  ]
}

Anthropic-Compatible Endpoints

POST /v1/messages

Anthropic Messages API, the wire protocol Claude Code speaks. Point Claude Code at the fleet with one env var and it runs on your local hardware. Unmapped claude-* ids auto-route to the best loaded model, and per-tier routing sends haiku to a fast model and sonnet/opus to a larger one. Images in the request auto-route to a vision-capable node.

Request (standard Messages API shape):

{
  "model": "claude-sonnet-4-5",
  "max_tokens": 4096,
  "system": "You are a coding assistant.",
  "messages": [{"role": "user", "content": "Explain this stack trace."}],
  "stream": true
}

Streams Anthropic SSE events. Three-layer context management (mechanical tool-result clearing, LLM summarization, pre-inference 413 cap) keeps long sessions from breaking on local models. See the Claude Code guide.


POST /v1/messages/count_tokens

Token-counting endpoint matching Anthropic's, used by Claude Code to size context before sending a request.


Ollama-Compatible Endpoints

POST /api/chat

Ollama-compatible chat. Streaming enabled by default (matches Ollama behavior).

Request:

{
  "model": "llama3.3:70b",
  "messages": [{"role": "user", "content": "Hello!"}],
  "stream": true,
  "options": {"temperature": 0.7, "num_predict": 1024},
  "fallback_models": ["qwen2.5:32b"],
  "metadata": {"tags": ["my-app"]}
}

Streaming response (NDJSON):

{"message":{"role":"assistant","content":"Hello"},"done":false}
{"message":{"role":"assistant","content":"!"},"done":false}
{"message":{"role":"assistant","content":""},"done":true,"total_duration":1234567890}

The metadata, fallback_models, and user fields are stripped before proxying to Ollama. The X-Herd-Tags header is also supported.


POST /api/generate

Ollama-compatible text generation. Uses prompt instead of messages.

{
  "model": "llama3.3:70b",
  "prompt": "Why is the sky blue?",
  "stream": true
}

POST /api/pull

Pull a model onto the fleet. Streams NDJSON progress matching Ollama's wire format.

Request:

{
  "name": "codestral",
  "stream": true,
  "node_id": "mac-studio"
}
Field Type Default Description
name string -- Model to pull (Ollama standard)
model string -- Model to pull (alias for name)
stream bool true Stream progress or block until done
node_id string auto Target node (auto-selects best if omitted)

Streaming response:

{"status":"pulling manifest"}
{"status":"pulling abc123...","digest":"sha256:abc123","total":5000000,"completed":2500000}
{"status":"success"}

Non-Ollama models (mflux, DiffusionKit, MLX) return a 400 with install instructions instead of pulling.

Errors: 400 missing name or non-Ollama model, 404 node not found, 409 already pulling, 503 no node with enough memory.

Examples:

# Stream pull progress
curl -N http://localhost:11435/api/pull -d '{"name": "codestral"}'

# Pull to specific node
curl -N http://localhost:11435/api/pull -d '{"name": "llama3.3:70b", "node_id": "mac-studio"}'

# Block until done
curl http://localhost:11435/api/pull -d '{"name": "phi4", "stream": false}'

GET /api/tags

List all models across the fleet with node information.

Response:

{
  "models": [
    {
      "name": "llama3.3:70b",
      "size": 42949672960,
      "details": {"fleet_nodes": ["mac-studio-ultra", "macbook-pro-m4"]}
    },
    {
      "name": "z-image-turbo",
      "size": 0,
      "details": {"fleet_nodes": ["mac-studio-ultra"], "type": "image"}
    }
  ]
}

GET /api/ps

List all currently loaded (hot) models.

Response:

{
  "models": [
    {"name": "llama3.3:70b", "size": 42949672960, "fleet_node": "mac-studio-ultra"},
    {"name": "qwen2.5:7b", "size": 4294967296, "fleet_node": "macbook-air-m2"}
  ]
}

POST /api/embed / POST /api/embeddings

Ollama-compatible embeddings. Both endpoints accept input or prompt.

{
  "model": "nomic-embed-text",
  "input": "The quick brown fox"
}

Errors: 400 missing model, 404 model not found, 503 no node available.

Changed in v0.10.0: inputs longer than 2,048 tokens are truncated, where they previously ran at up to 8,192. This is the fix for a 28 GB memory high-water mark in the node agent. If you were relying on the old limit, chunk before embedding. Sending "truncate": false now returns 400, which matches Ollama's behaviour.


POST /v1/embeddings

OpenAI-compatible embeddings, added in v0.9.6 so clients that only speak the OpenAI shape (AnythingLLM's Generic OpenAI embedder, Chatbox's Knowledge Base) can use the fleet. Routes through the same dispatcher as /api/embed, so it inherits fastembed routing, scoring and retries.

{
  "model": "nomic-embed-text",
  "input": ["The quick brown fox", "jumps over the lazy dog"]
}

input accepts a string or an array of strings. An empty string is rejected rather than skipped: the backend drops empties, which would shift every later index and hand you embeddings silently attached to the wrong inputs.

encoding_format: "base64" returns each vector as little-endian float32, the same encoding OpenAI uses.

Errors: 400 bad input or an empty string, 404 model not found, 502 the backend returned a different number of vectors than inputs, or ignored a dimensions you asked for, 503 no node available.


POST /v1/rerank

Reranks documents against a query, for the rerank step of a RAG pipeline (Open WebUI, Dify, LangChain, LlamaIndex). Ollama has no rerank endpoint. Herd serves this from the native fastembed server that nodes with the embedding extra already run, so it adds no dependency. Request and response follow the Jina/Cohere shape. Added v0.10.0.

curl http://localhost:11435/v1/rerank -H 'Content-Type: application/json' -d '{
  "query": "how do I stop herd now that launchd is installed?",
  "documents": ["launchctl bootout gui/$UID/...", "pkill is no longer a stop"],
  "top_n": 2,
  "return_documents": true
}'

Request: query and documents are required (documents takes strings or {"text": ...} objects). Optional: model, top_n, return_documents (default false).

Response: {model, results: [{index, relevance_score, document?}], usage}. Sorted best first, and index points back into the original documents. relevance_score is the cross-encoder logit through a sigmoid, so it lands in 0-1 like Cohere's and Jina's and a client's relevance threshold behaves the same. The serving node is in X-Fleet-Node, not the body.

Models download on first use and stay cached. ms-marco-minilm-l-6-v2 (80 MB) is the default; bge-reranker-base (1.04 GB) and jina-reranker-v2-base-multilingual (1.11 GB) are the quality picks, the latter for non-English.

Limits: at most 1,000 documents of 32,000 characters each, else 400.


POST /api/show

Ollama-compatible model metadata. Ollama's own desktop app calls this before every chat, so without it that app could not hold a conversation. Added v0.9.6.


GET /api/image-models

List all image models (mflux, DiffusionKit, Ollama native).

{
  "models": [
    {"name": "z-image-turbo", "type": "image", "backend": "mflux", "fleet_nodes": ["mac-studio"]},
    {"name": "x/z-image-turbo:latest", "type": "image", "backend": "ollama", "fleet_nodes": ["mac-studio"]}
  ]
}

POST /api/generate-image

Generate an image on the best available node. Routes to nodes running mflux or DiffusionKit (Apple Silicon only), or an Ollama-native image model (any platform where Ollama supports it). Requires image generation to be enabled (FLEET_IMAGE_GENERATION=true or the dashboard toggle).

Request:

{
  "model": "z-image-turbo",
  "prompt": "a neon-lit Tokyo alley at midnight",
  "width": 1024,
  "height": 1024,
  "steps": 4
}

Response: raw PNG bytes, with X-Fleet-Node, X-Fleet-Model, and X-Generation-Time headers. Also available as OpenAI-format POST /v1/images/generations. See the image generation guide.


POST /api/transcribe

Speech-to-text, Whisper-compatible. Routes an uploaded audio file to a node running an ASR model (Qwen3-ASR on MLX, so Apple Silicon nodes only). Requires transcription to be enabled (FLEET_TRANSCRIPTION=true or the settings API).

Request: multipart/form-data with an audio file field.

curl http://localhost:11435/api/transcribe \
  -F "audio=@meeting.wav"

Response:

{"text": "Full transcription of the audio..."}

Errors: 503 transcription disabled, 404 no node has an ASR model installed. See multimodal routing.


Fleet Management Endpoints

GET /fleet/status

Full fleet state, nodes, models, queues, hardware.

curl -s http://localhost:11435/fleet/status | python3 -m json.tool

Returns fleet summary (totals), nodes array (per-node details), and queues object (per node:model queue state).


GET /fleet/queue

Lightweight queue depths only. Designed for client-side backoff.

curl -s http://localhost:11435/fleet/queue | python3 -m json.tool

Dashboard API Endpoints

All return JSON. Used by the web dashboard and available for external monitoring.

Endpoint Method Description
/dashboard/api/health GET 30+ automated health checks
/dashboard/api/overview GET Fleet totals (requests, nodes, models)
/dashboard/api/traces GET Recent request traces (?limit=20)
/dashboard/api/usage GET Per-node, per-model daily usage stats
/dashboard/api/models GET Per-model aggregates (?days=7)
/dashboard/api/apps GET Per-tag analytics (?days=7)
/dashboard/api/apps/daily GET Per-tag daily breakdown
/dashboard/api/trends GET Hourly request/latency trends (?hours=24)
/dashboard/api/recommendations GET Model mix recommendations per node
/dashboard/api/settings GET Current configuration
/dashboard/api/settings POST Update runtime settings
/dashboard/api/model-management GET Per-node model details
/dashboard/api/pull POST Pull model to node (dashboard use)
/dashboard/api/delete POST Delete model from node
/dashboard/api/benchmarks GET List past benchmark runs
/dashboard/api/benchmarks/start POST Start benchmark (mode, duration, model_types)
/dashboard/api/benchmarks/progress GET Real-time benchmark progress
/dashboard/api/benchmarks/cancel POST Cancel running benchmark
/dashboard/api/benchmarks/{run_id} GET Fetch specific run results
/dashboard/api/context-usage GET Per-model context utilization stats
/dashboard/events GET Server-Sent Events stream for real-time updates

Context Usage API

GET /dashboard/api/context-usage

Per-model context window utilization statistics.

Query parameters:

Field Type Default Description
days integer 7 Days of history to analyze

Response (per model):

Field Description
model Model name
allocated_ctx Context window allocated by Ollama
recommended_ctx Optimal context size based on usage
utilization_pct Percentage of allocated context actually used (p99)
savings_pct Memory savings if context were right-sized
request_count Total requests in period
prompt_tokens Token distribution: avg, p50, p75, p95, p99, max
total_tokens Combined prompt+completion: p95, p99, max, max_24h

Request Tagging

Tags can come from three sources (merged and deduplicated):

Source Example
metadata.tags in body {"metadata": {"tags": ["my-app"]}}
X-Herd-Tags header X-Herd-Tags: my-app, production
user field in body {"user": "alice"} (stored as user:alice)

Common Response Headers

Every routed request includes:

Header Description
X-Fleet-Node Which node handled the request
X-Fleet-Score The winning routing score
X-Fleet-Affinity matched or new, whether session affinity kept the conversation on one node

These help with debugging, you can see exactly which node was chosen and why.

X-Fleet-Affinity reports our routing decision, not a backend cache hit. Whether the node actually reused its prefix cache depends on the backend:

  • MLX nodes (Apple Silicon) measure it. The response usage gains prompt_tokens_details.cached_tokens, the slice of the prompt that was skipped because the prefix was already cached. Divide by prompt_tokens for a real hit rate.
  • Ollama nodes cannot. It folds llama.cpp's cache_n back into prompt_n on purpose (ollama#16428), so no hit ratio exists to report. The field is omitted, not set to zero.

That distinction is deliberate and worth copying if you build something similar. cached_tokens: 0 means "measured, nothing was reused". An absent field means "cannot measure". Sending 0 for the second case quietly claims every Ollama request was a cache miss, which is a bug vLLM shipped and fixed (#44383) and SGLang still has.

Where you cannot measure, time to first token is the honest proxy: a matched follow-up much faster than its first turn is the cache doing its job.

Browser clients and CORS

A web page may only call another origin if that server sends CORS headers and answers the browser's preflight OPTIONS. Herd sent neither before v0.9.6, so browser-based clients (Hollama, TypingMind, Page Assist, the web builds of Chatbox and Msty) failed against it while working against Ollama.

It is off by default: with FLEET_CORS_ORIGINS unset no middleware is installed at all, so an existing router behaves byte-for-byte as it did.

FLEET_CORS_ORIGINS="http://localhost:*,chrome-extension://*"

Syntax follows Ollama's OLLAMA_ORIGINS: comma-separated origins, * usable inside an origin, a bare * allows any origin.

Two things worth knowing:

  • Herd adds no implicit origins. Ollama always allows a built-in set of localhost and app origins on top of yours. Herd does not, because the router listens on the whole LAN, so you name every origin you trust.
  • CORS does not lift mixed content. An HTTPS page still cannot call a plain-HTTP router. That is a browser rule about the scheme, not about permission.

The X-Fleet-* headers are in the exposed set, so a browser client can still read the routing metadata. CORS hides non-safelisted response headers by default, which would otherwise make them invisible to exactly the clients this exists for.


Next Steps