Routing Engine

How the router makes smart decisions about where to send every request.

The Five-Stage Pipeline

Every incoming request passes through five stages. The first three decide where it runs; the last two run in the background, keeping the next requests fast.

Conceptual: one request through the v0.9.5 router
Incoming requestany of the four wire protocols, for one model
every node the router knows about
Stage 1: Eliminationremoves nodes that can't serve: offline, model not on disk, won't fit in free system RAM, critical memory pressure, paused
survivors (if none: holding queue, rescored every 2 s for up to 30 s)
Stage 2: Scoringeach survivor gets a total across 8 signals; the highest total wins
each request goes to one node
Stage 3: Queuethe request joins the winner's own node:model queue and runs there, whole
in the background, every 5 seconds
Stage 4: Pre-warmany queue at 3 or more: load that model on the runner-up
Stage 5: Rebalancemove waiting requests off deep queues to nodes where the model is already loaded
Elimination and scoring pick exactly one node for each request; pre-warming and rebalancing never split a request, they only prepare other nodes for the ones that follow.

The goal of every decision is to minimize total response time:

total_time = queue_wait_time + cold_load_time + inference_time

A request routed to a node with the model already loaded and an empty queue beats a request routed to an idle node that needs to load the model first.

Stage 1: Elimination

Before any scoring, the router removes nodes that can't serve the request. This is binary, pass or out.

ConditionOutcome If False
Node is online (heard from within the last 30 seconds)Eliminated
Model is on disk on this nodeEliminated
Node has enough free memory to load the modelEliminated
Node is not in hard-pause modeEliminated

Hard-pause triggers (regardless of learned baseline):

  • Memory pressure state is critical (reported on macOS and Linux)
  • With capacity learning turned on (opt-in): camera or microphone active (meeting in progress, macOS only), CPU above 85% for 2 minutes or more, availability score below 0.20, or the first 7 days of observation

If nothing survives: The request enters a holding queue. The router rescores it every 2 seconds, for up to 30 seconds, as node states change. The request waits rather than errors.

Stage 2: Scoring

Each surviving node gets scored across 8 weighted signals. Higher total wins.

Diverging bar chart of how far each routing signal can move a node's score in Ollama Herd v0.9.5: model warmth up to +50, memory fit up to +20, queue depth down to -30, estimated wait down to -25, role affinity up to +25, availability up to +10 (capacity learning nodes only), context fit from -15 to +15, session affinity up to +20.
Having the model already loaded (+50) outweighs any other single signal, while a deep queue (-30) and a long estimated wait (-25) can together cancel it. Sourced: the v0.9.5 scorer and its default settings.

Signal 1: Model Warmth (up to +50)

The most important signal. Cold-loading a 40GB model takes 15–30 seconds.

StatePoints
Model currently loaded in memory (hot)+50
Model on disk, loaded within last 30 min (likely OS-cached)+30
Model on disk, not recently used+10
Model not on disk (requires download)+0

The "recently loaded" tier exists because the operating system keeps recently read model files in its page cache (macOS is especially aggressive about this), so a model unloaded 20 minutes ago often reloads much faster.

Signal 2: Memory Fit (up to +20)

How comfortably the model fits given current utilization and the node's dynamic memory ceiling.

available = free system RAM, capped by the capacity ceiling if learning is on
fit_ratio = available / model_memory_gb

fit_ratio > 2.0     +20   comfortable headroom
fit_ratio 1.5-2.0   +15
fit_ratio 1.2-1.5   +8
fit_ratio 1.0-1.2   +3    tight -- risk of memory pressure
fit_ratio < 1.0     eliminated in Stage 1

A model that is already loaded gets the full +20. model_memory_gb is the model's measured resident size when the fleet has seen it loaded (weights plus KV cache), otherwise its weights. On a node with capacity learning turned on, the adaptive ceiling from the capacity learner also caps available: a 128GB laptop in learned-low mode has a 16GB ceiling.

Signal 3: Queue Depth Penalty (up to -30)

A hot model on a saturated node is less attractive than a warm model on an empty node.

depth = in_flight + pending
penalty = min(30, depth x 6)

A queue of 5 subtracts the maximum 30 points. This naturally spreads load when multiple nodes can serve the same model.

Signal 4: Estimated Wait Time Penalty (up to -25)

Queue depth alone is misleading. A queue of 3 on a fast 7B model completes in seconds. A queue of 3 on a slow 70B model takes minutes.

The router uses p75 historical latency per node per model:

est_wait = (in_flight + pending) x p75_latency_ms
penalty = min(25, est_wait_seconds / 10)

During the first 7 days before enough data is collected, the router uses a heuristic based on memory bandwidth and model size.

Signal 5: Role Affinity (up to +25)

Large models belong on fast machines. Small models should run on lighter hardware to preserve the fast machine's capacity. The router scores this from each node's memory bandwidth when it's known (the real bottleneck for prompt processing), and falls back to total memory when it isn't.

bandwidth bonus = 5 + bandwidth_GBps / 40, capped at 25
  100 GB/s  ~+7.5     400 GB/s  ~+15     800 GB/s  +25

Model over 20 GB:     full bandwidth bonus  (fast nodes win)
Model under 8 GB:     18 - 0.6 x bonus, at least +3  (slower nodes win)
In between:           0.6 x bonus

Bandwidth unknown (memory-tier fallback):
  Model over 20 GB:   128GB+ node +15, 32GB+ node +5, smaller +0
  Model under 8 GB:   32GB-or-less node +15, up to 128GB +8, larger +3

Machine brand never enters into it: a Linux server or Windows PC is placed by the same numbers as a Mac. Size thresholds are configurable (FLEET_SCORE_ROLE_LARGE_THRESHOLD_GB, FLEET_SCORE_ROLE_SMALL_THRESHOLD_GB).

Without affinity, every small-model request would drift to the most powerful machine, crowding out the large models it's uniquely suited for.

Signal 6: Availability Trend (up to +10)

Only nodes with capacity learning turned on (it is opt-in) get this signal; every other node scores 0 here. The bonus is the node's availability score, 0.0 to 1.0, times 10:

bonus = availability_score x 10

The availability score blends the learned baseline for this hour of the week, the current CPU and memory load, and whether CPU use has been rising or falling over the last 5 minutes (see Adaptive Capacity). Rising use lowers the score, which keeps a long inference request off a machine whose owner just sat down to start working.

Signal 7: Context Fit (+15 to -15)

Rewards nodes whose loaded context window comfortably handles the estimated request size. Penalizes nodes where the request might trigger a context resize (which causes a model reload).

Signal 8: Session Affinity (up to +20)

Pins a multi-turn conversation to the node that already holds its warm prefix cache. On turn N+1 that node re-uses the cached prefix, so it costs a few hundred prefill tokens instead of re-encoding the whole ~30K-token history, and it avoids the decode interference a fresh prefill inflicts on other streams sharing the node. The bonus is deliberately not decisive: a warm cache is worth a lot, but the bonus stays below Signal 1's +50 and shrinks as the pinned node's queue grows, so a saturated node loses to an idle peer. Added in v0.9.1.

session bonus = 20 / (1 + depth / 2)
depth = requests running or waiting for this model on that node

Signal 3 takes 6 points per queued request at the same time, so between two otherwise identical nodes the session's node leads by 20 with an empty queue, by 7.3 with one request, and trails by 2 once two requests are running or waiting there.

Line chart of the session node's score advantage over an identical idle peer by queue depth: +20 at depth 0, +7.3 at depth 1, -2 at depth 2, -10 at depth 3, -17.3 at depth 4. The affinity bonus falls from 20 to 5 while the queue penalty grows by 6 per request to a cap of 30.
Session affinity keeps a conversation on its node only while that node is lightly loaded: from a queue depth of 2, an otherwise identical idle peer outscores it. Modeled: signals 3 and 8 of the v0.9.5 scorer, equal nodes, wait-time penalty ignored.

Stage 3: Queue and Execute

The highest-scoring node wins. The request enters that node's dedicated queue for that model. Each node+model pair has its own queue with dynamic concurrency calculated from available memory.

Stage 4: Pre-Warm

Every 5 seconds, a background loop checks each queue: is it getting deep?

If a queue holds 3 or more requests, running plus waiting (the pre-warm threshold), the router loads the same model on the runner-up node by sending an empty generate request. By the time overflow requests arrive, the model is already hot.

A lock prevents duplicate pre-warm requests for the same model on the same node.

Stage 5: Rebalance

A background process runs every 5 seconds:

  1. Scans all queues for ones with 4 or more requests waiting (the rebalance threshold)
  2. For each overloaded queue, checks if another node has the same model hot with spare capacity
  3. Moves pending requests (not in-flight) to the better node
  4. Caps movement at 3 requests per cycle to prevent oscillation

The rebalancer only moves requests to nodes where the model is already loaded, it never triggers cold loads.

Fallback Chain

When a request lists fallback models, the router scores the primary and every fallback in the same pass, each through the full elimination and scoring pipeline:

  1. If any of them is already loaded on a surviving node, the first such model in list order wins. A loaded fallback is used rather than cold-loading the primary.
  2. Otherwise the first model in the list that some node can load wins, and that node loads it.
  3. If no node can serve any of them, the request waits in the holding queue, rescoring the whole list every 2 seconds for up to 30 seconds.
  4. If nothing frees up (and auto-pull cannot fetch the model), the router returns 503.

Configuration

All scoring weights are tunable via environment variables:

VariableDefaultEffect
FLEET_SCORE_MODEL_HOT50Increase to prefer hot models more aggressively
FLEET_SCORE_QUEUE_DEPTH_PENALTY_PER6Decrease to tolerate deeper queues
FLEET_SCORE_ROLE_AFFINITY_MAX15Shown on the dashboard but not read by the scorer in v0.9.5; role affinity is capped at +25 (bandwidth-aware, the default) or +15 (memory-tier fallback)
FLEET_PRE_WARM_THRESHOLD3Lower to pre-warm earlier
FLEET_REBALANCE_THRESHOLD4Lower to rebalance more aggressively

See the full configuration reference for all 44+ variables.

Next Steps