API Reference
Every endpoint with request/response schemas, headers, error codes, and examples.
The router runs on port 11435 by default. All endpoints accept JSON bodies and return JSON responses.
OpenAI-Compatible Endpoints
POST /v1/chat/completions
Chat completions with streaming and non-streaming support.
Request:
{
"model": "llama3.3:70b",
"messages": [
{"role": "system", "content": "You are helpful."},
{"role": "user", "content": "Hello!"}
],
"stream": true,
"temperature": 0.7,
"max_tokens": 1024,
"fallback_models": ["qwen2.5:32b", "qwen2.5:7b"],
"metadata": {"tags": ["my-app"]},
"user": "alice"
}
| Field | Type | Default | Description |
|---|---|---|---|
model |
string | required | Model name |
messages |
array | [] |
Chat messages in OpenAI format |
stream |
boolean | false |
Enable SSE streaming |
temperature |
float | 0.7 |
Sampling temperature |
max_tokens |
integer | null | Maximum tokens to generate |
fallback_models |
array | [] |
Backup models if primary unavailable |
metadata.tags |
array | [] |
Tags for per-tag analytics |
user |
string | null | Stored as user:<value> tag |
Response headers:
| Header | Description |
|---|---|
X-Fleet-Node |
Node that handled the request |
X-Fleet-Score |
Winning routing score |
X-Fleet-Fallback |
Fallback model used (if applicable) |
X-Fleet-Retries |
Retry count (if retries occurred) |
X-Fleet-Affinity |
matched when routed back to the node already holding this conversation, new otherwise. Omitted when affinity did not apply |
X-Fleet-Context-Overflow |
Context overflow warning |
X-Thinking-Tokens |
Thinking tokens (chain-of-thought models, non-streaming) |
X-Output-Tokens |
Output tokens (chain-of-thought models, non-streaming) |
X-Budget-Used |
completion_tokens/num_predict budget check (non-streaming) |
X-Done-Reason |
stop (natural) or length (budget exhausted) |
Errors: 400 missing model, 404 model not found, 503 no node available.
Example:
curl http://localhost:11435/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "llama3.3:70b", "messages": [{"role": "user", "content": "Hello!"}], "stream": false}'
POST /v1/responses
OpenAI Responses API, the wire protocol OpenAI Codex speaks (Codex dropped Chat Completions in early 2026). Point Codex at the fleet and it routes across your machines. Unmapped model ids auto-route to the best loaded model, so no model map is required.
Request (standard Responses API shape):
{
"model": "gpt-5-codex",
"input": "Refactor this function.",
"stream": true
}
Streams Server-Sent Events in Responses format. See the Codex guide for the full CLI and desktop-app setup.
POST /v1/images/generations
OpenAI-compatible image generation.
Request:
{
"model": "z-image-turbo",
"prompt": "a cat sitting on a laptop",
"size": "1024x1024",
"response_format": "b64_json"
}
| Field | Type | Default | Description |
|---|---|---|---|
model |
string | required | Image model name |
prompt |
string | required | Image description |
size |
string | 1024x1024 |
Dimensions (WIDTHxHEIGHT) |
response_format |
string | b64_json |
b64_json or url (raw PNG) |
steps |
integer | model default | Inference steps |
guidance |
float | model default | Guidance scale |
seed |
integer | random | Reproducibility seed |
Errors: 400 missing fields, 404 model unavailable, 502 generation failed, 503 disabled.
GET /v1/models
List all models across the fleet (LLM, image, embedding).
Response:
{
"object": "list",
"data": [
{"id": "llama3.3:70b", "object": "model", "created": 1710000000, "owned_by": "ollama"},
{"id": "z-image-turbo", "object": "model", "created": 1710000000, "owned_by": "mflux"}
]
}
Anthropic-Compatible Endpoints
POST /v1/messages
Anthropic Messages API, the wire protocol Claude Code speaks. Point Claude Code at the fleet with one env var and it runs on your local hardware. Unmapped claude-* ids auto-route to the best loaded model, and per-tier routing sends haiku to a fast model and sonnet/opus to a larger one. Images in the request auto-route to a vision-capable node.
Request (standard Messages API shape):
{
"model": "claude-sonnet-4-5",
"max_tokens": 4096,
"system": "You are a coding assistant.",
"messages": [{"role": "user", "content": "Explain this stack trace."}],
"stream": true
}
Streams Anthropic SSE events. Three-layer context management (mechanical tool-result clearing, LLM summarization, pre-inference 413 cap) keeps long sessions from breaking on local models. See the Claude Code guide.
POST /v1/messages/count_tokens
Token-counting endpoint matching Anthropic's, used by Claude Code to size context before sending a request.
Ollama-Compatible Endpoints
POST /api/chat
Ollama-compatible chat. Streaming enabled by default (matches Ollama behavior).
Request:
{
"model": "llama3.3:70b",
"messages": [{"role": "user", "content": "Hello!"}],
"stream": true,
"options": {"temperature": 0.7, "num_predict": 1024},
"fallback_models": ["qwen2.5:32b"],
"metadata": {"tags": ["my-app"]}
}
Streaming response (NDJSON):
{"message":{"role":"assistant","content":"Hello"},"done":false}
{"message":{"role":"assistant","content":"!"},"done":false}
{"message":{"role":"assistant","content":""},"done":true,"total_duration":1234567890}
The metadata, fallback_models, and user fields are stripped before proxying to Ollama. The X-Herd-Tags header is also supported.
POST /api/generate
Ollama-compatible text generation. Uses prompt instead of messages.
{
"model": "llama3.3:70b",
"prompt": "Why is the sky blue?",
"stream": true
}
POST /api/pull
Pull a model onto the fleet. Streams NDJSON progress matching Ollama's wire format.
Request:
{
"name": "codestral",
"stream": true,
"node_id": "mac-studio"
}
| Field | Type | Default | Description |
|---|---|---|---|
name |
string | -- | Model to pull (Ollama standard) |
model |
string | -- | Model to pull (alias for name) |
stream |
bool | true |
Stream progress or block until done |
node_id |
string | auto | Target node (auto-selects best if omitted) |
Streaming response:
{"status":"pulling manifest"}
{"status":"pulling abc123...","digest":"sha256:abc123","total":5000000,"completed":2500000}
{"status":"success"}
Non-Ollama models (mflux, DiffusionKit, MLX) return a 400 with install instructions instead of pulling.
Errors: 400 missing name or non-Ollama model, 404 node not found, 409 already pulling, 503 no node with enough memory.
Examples:
# Stream pull progress
curl -N http://localhost:11435/api/pull -d '{"name": "codestral"}'
# Pull to specific node
curl -N http://localhost:11435/api/pull -d '{"name": "llama3.3:70b", "node_id": "mac-studio"}'
# Block until done
curl http://localhost:11435/api/pull -d '{"name": "phi4", "stream": false}'
GET /api/tags
List all models across the fleet with node information.
Response:
{
"models": [
{
"name": "llama3.3:70b",
"size": 42949672960,
"details": {"fleet_nodes": ["mac-studio-ultra", "macbook-pro-m4"]}
},
{
"name": "z-image-turbo",
"size": 0,
"details": {"fleet_nodes": ["mac-studio-ultra"], "type": "image"}
}
]
}
GET /api/ps
List all currently loaded (hot) models.
Response:
{
"models": [
{"name": "llama3.3:70b", "size": 42949672960, "fleet_node": "mac-studio-ultra"},
{"name": "qwen2.5:7b", "size": 4294967296, "fleet_node": "macbook-air-m2"}
]
}
POST /api/embed / POST /api/embeddings
Ollama-compatible embeddings. Both endpoints accept input or prompt.
{
"model": "nomic-embed-text",
"input": "The quick brown fox"
}
Errors: 400 missing model, 404 model not found, 503 no node available.
Changed in v0.10.0: inputs longer than 2,048 tokens are truncated, where they previously ran at up to 8,192. This is the fix for a 28 GB memory high-water mark in the node agent. If you were relying on the old limit, chunk before embedding. Sending
"truncate": falsenow returns400, which matches Ollama's behaviour.
POST /v1/embeddings
OpenAI-compatible embeddings, added in v0.9.6 so clients that only speak the
OpenAI shape (AnythingLLM's Generic OpenAI embedder, Chatbox's Knowledge Base)
can use the fleet. Routes through the same dispatcher as /api/embed, so it
inherits fastembed routing, scoring and retries.
{
"model": "nomic-embed-text",
"input": ["The quick brown fox", "jumps over the lazy dog"]
}
input accepts a string or an array of strings. An empty string is rejected
rather than skipped: the backend drops empties, which would shift every later
index and hand you embeddings silently attached to the wrong inputs.
encoding_format: "base64" returns each vector as little-endian float32, the
same encoding OpenAI uses.
Errors: 400 bad input or an empty string, 404 model not found, 502 the
backend returned a different number of vectors than inputs, or ignored a
dimensions you asked for, 503 no node available.
POST /v1/rerank
Reranks documents against a query, for the rerank step of a RAG pipeline (Open
WebUI, Dify, LangChain, LlamaIndex). Ollama has no rerank endpoint. Herd
serves this from the native fastembed server that nodes with the embedding
extra already run, so it adds no dependency. Request and response follow the
Jina/Cohere shape. Added v0.10.0.
curl http://localhost:11435/v1/rerank -H 'Content-Type: application/json' -d '{
"query": "how do I stop herd now that launchd is installed?",
"documents": ["launchctl bootout gui/$UID/...", "pkill is no longer a stop"],
"top_n": 2,
"return_documents": true
}'
Request: query and documents are required (documents takes strings or
{"text": ...} objects). Optional: model, top_n, return_documents
(default false).
Response: {model, results: [{index, relevance_score, document?}], usage}.
Sorted best first, and index points back into the original documents.
relevance_score is the cross-encoder logit through a sigmoid, so it lands in
0-1 like Cohere's and Jina's and a client's relevance threshold behaves the
same. The serving node is in X-Fleet-Node, not the body.
Models download on first use and stay cached. ms-marco-minilm-l-6-v2 (80 MB)
is the default; bge-reranker-base (1.04 GB) and
jina-reranker-v2-base-multilingual (1.11 GB) are the quality picks, the
latter for non-English.
Limits: at most 1,000 documents of 32,000 characters each, else 400.
POST /api/show
Ollama-compatible model metadata. Ollama's own desktop app calls this before every chat, so without it that app could not hold a conversation. Added v0.9.6.
GET /api/image-models
List all image models (mflux, DiffusionKit, Ollama native).
{
"models": [
{"name": "z-image-turbo", "type": "image", "backend": "mflux", "fleet_nodes": ["mac-studio"]},
{"name": "x/z-image-turbo:latest", "type": "image", "backend": "ollama", "fleet_nodes": ["mac-studio"]}
]
}
POST /api/generate-image
Generate an image on the best available node. Routes to nodes running mflux or DiffusionKit (Apple Silicon only), or an Ollama-native image model (any platform where Ollama supports it). Requires image generation to be enabled (FLEET_IMAGE_GENERATION=true or the dashboard toggle).
Request:
{
"model": "z-image-turbo",
"prompt": "a neon-lit Tokyo alley at midnight",
"width": 1024,
"height": 1024,
"steps": 4
}
Response: raw PNG bytes, with X-Fleet-Node, X-Fleet-Model, and X-Generation-Time headers. Also available as OpenAI-format POST /v1/images/generations. See the image generation guide.
POST /api/transcribe
Speech-to-text, Whisper-compatible. Routes an uploaded audio file to a node running an ASR model (Qwen3-ASR on MLX, so Apple Silicon nodes only). Requires transcription to be enabled (FLEET_TRANSCRIPTION=true or the settings API).
Request: multipart/form-data with an audio file field.
curl http://localhost:11435/api/transcribe \
-F "audio=@meeting.wav"
Response:
{"text": "Full transcription of the audio..."}
Errors: 503 transcription disabled, 404 no node has an ASR model installed. See multimodal routing.
Fleet Management Endpoints
GET /fleet/status
Full fleet state, nodes, models, queues, hardware.
curl -s http://localhost:11435/fleet/status | python3 -m json.tool
Returns fleet summary (totals), nodes array (per-node details), and queues object (per node:model queue state).
GET /fleet/queue
Lightweight queue depths only. Designed for client-side backoff.
curl -s http://localhost:11435/fleet/queue | python3 -m json.tool
Dashboard API Endpoints
All return JSON. Used by the web dashboard and available for external monitoring.
| Endpoint | Method | Description |
|---|---|---|
/dashboard/api/health |
GET | 30+ automated health checks |
/dashboard/api/overview |
GET | Fleet totals (requests, nodes, models) |
/dashboard/api/traces |
GET | Recent request traces (?limit=20) |
/dashboard/api/usage |
GET | Per-node, per-model daily usage stats |
/dashboard/api/models |
GET | Per-model aggregates (?days=7) |
/dashboard/api/apps |
GET | Per-tag analytics (?days=7) |
/dashboard/api/apps/daily |
GET | Per-tag daily breakdown |
/dashboard/api/trends |
GET | Hourly request/latency trends (?hours=24) |
/dashboard/api/recommendations |
GET | Model mix recommendations per node |
/dashboard/api/settings |
GET | Current configuration |
/dashboard/api/settings |
POST | Update runtime settings |
/dashboard/api/model-management |
GET | Per-node model details |
/dashboard/api/pull |
POST | Pull model to node (dashboard use) |
/dashboard/api/delete |
POST | Delete model from node |
/dashboard/api/benchmarks |
GET | List past benchmark runs |
/dashboard/api/benchmarks/start |
POST | Start benchmark (mode, duration, model_types) |
/dashboard/api/benchmarks/progress |
GET | Real-time benchmark progress |
/dashboard/api/benchmarks/cancel |
POST | Cancel running benchmark |
/dashboard/api/benchmarks/{run_id} |
GET | Fetch specific run results |
/dashboard/api/context-usage |
GET | Per-model context utilization stats |
/dashboard/events |
GET | Server-Sent Events stream for real-time updates |
Context Usage API
GET /dashboard/api/context-usage
Per-model context window utilization statistics.
Query parameters:
| Field | Type | Default | Description |
|---|---|---|---|
days |
integer | 7 |
Days of history to analyze |
Response (per model):
| Field | Description |
|---|---|
model |
Model name |
allocated_ctx |
Context window allocated by Ollama |
recommended_ctx |
Optimal context size based on usage |
utilization_pct |
Percentage of allocated context actually used (p99) |
savings_pct |
Memory savings if context were right-sized |
request_count |
Total requests in period |
prompt_tokens |
Token distribution: avg, p50, p75, p95, p99, max |
total_tokens |
Combined prompt+completion: p95, p99, max, max_24h |
Request Tagging
Tags can come from three sources (merged and deduplicated):
| Source | Example |
|---|---|
metadata.tags in body |
{"metadata": {"tags": ["my-app"]}} |
X-Herd-Tags header |
X-Herd-Tags: my-app, production |
user field in body |
{"user": "alice"} (stored as user:alice) |
Common Response Headers
Every routed request includes:
| Header | Description |
|---|---|
X-Fleet-Node |
Which node handled the request |
X-Fleet-Score |
The winning routing score |
X-Fleet-Affinity |
matched or new, whether session affinity kept the conversation on one node |
These help with debugging, you can see exactly which node was chosen and why.
X-Fleet-Affinity reports our routing decision, not a backend cache hit.
Whether the node actually reused its prefix cache depends on the backend:
- MLX nodes (Apple Silicon) measure it. The response
usagegainsprompt_tokens_details.cached_tokens, the slice of the prompt that was skipped because the prefix was already cached. Divide byprompt_tokensfor a real hit rate. - Ollama nodes cannot. It folds llama.cpp's
cache_nback intoprompt_non purpose (ollama#16428), so no hit ratio exists to report. The field is omitted, not set to zero.
That distinction is deliberate and worth copying if you build something similar.
cached_tokens: 0 means "measured, nothing was reused". An absent field means
"cannot measure". Sending 0 for the second case quietly claims every Ollama
request was a cache miss, which is a bug vLLM shipped and fixed
(#44383) and SGLang still has.
Where you cannot measure, time to first token is the honest proxy: a matched
follow-up much faster than its first turn is the cache doing its job.
Browser clients and CORS
A web page may only call another origin if that server sends CORS headers and
answers the browser's preflight OPTIONS. Herd sent neither before v0.9.6, so
browser-based clients (Hollama, TypingMind, Page Assist, the web builds of
Chatbox and Msty) failed against it while working against Ollama.
It is off by default: with FLEET_CORS_ORIGINS unset no middleware is
installed at all, so an existing router behaves byte-for-byte as it did.
FLEET_CORS_ORIGINS="http://localhost:*,chrome-extension://*"
Syntax follows Ollama's OLLAMA_ORIGINS: comma-separated origins, * usable
inside an origin, a bare * allows any origin.
Two things worth knowing:
- Herd adds no implicit origins. Ollama always allows a built-in set of localhost and app origins on top of yours. Herd does not, because the router listens on the whole LAN, so you name every origin you trust.
- CORS does not lift mixed content. An HTTPS page still cannot call a plain-HTTP router. That is a browser rule about the scheme, not about permission.
The X-Fleet-* headers are in the exposed set, so a browser client can still
read the routing metadata. CORS hides non-safelisted response headers by
default, which would otherwise make them invisible to exactly the clients this
exists for.
Next Steps
- Integrations, How to connect your tools
- Deployment, Monitoring and log analysis
- Routing Engine, Understanding scoring decisions