Routing Engine
How the router makes smart decisions about where to send every request.
The Five-Stage Pipeline
Every incoming request passes through five stages. The first three decide where it runs; the last two run in the background, keeping the next requests fast.
node:model queue and runs there, wholeThe goal of every decision is to minimize total response time:
total_time = queue_wait_time + cold_load_time + inference_time
A request routed to a node with the model already loaded and an empty queue beats a request routed to an idle node that needs to load the model first.
Stage 1: Elimination
Before any scoring, the router removes nodes that can't serve the request. This is binary, pass or out.
| Condition | Outcome If False |
|---|---|
| Node is online (heard from within the last 30 seconds) | Eliminated |
| Model is on disk on this node | Eliminated |
| Node has enough free memory to load the model | Eliminated |
| Node is not in hard-pause mode | Eliminated |
Hard-pause triggers (regardless of learned baseline):
- Memory pressure state is critical (reported on macOS and Linux)
- With capacity learning turned on (opt-in): camera or microphone active (meeting in progress, macOS only), CPU above 85% for 2 minutes or more, availability score below 0.20, or the first 7 days of observation
If nothing survives: The request enters a holding queue. The router rescores it every 2 seconds, for up to 30 seconds, as node states change. The request waits rather than errors.
Stage 2: Scoring
Each surviving node gets scored across 8 weighted signals. Higher total wins.
Signal 1: Model Warmth (up to +50)
The most important signal. Cold-loading a 40GB model takes 15–30 seconds.
| State | Points |
|---|---|
| Model currently loaded in memory (hot) | +50 |
| Model on disk, loaded within last 30 min (likely OS-cached) | +30 |
| Model on disk, not recently used | +10 |
| Model not on disk (requires download) | +0 |
The "recently loaded" tier exists because the operating system keeps recently read model files in its page cache (macOS is especially aggressive about this), so a model unloaded 20 minutes ago often reloads much faster.
Signal 2: Memory Fit (up to +20)
How comfortably the model fits given current utilization and the node's dynamic memory ceiling.
available = free system RAM, capped by the capacity ceiling if learning is on
fit_ratio = available / model_memory_gb
fit_ratio > 2.0 +20 comfortable headroom
fit_ratio 1.5-2.0 +15
fit_ratio 1.2-1.5 +8
fit_ratio 1.0-1.2 +3 tight -- risk of memory pressure
fit_ratio < 1.0 eliminated in Stage 1
A model that is already loaded gets the full +20. model_memory_gb is the model's measured resident size when the fleet has seen it loaded (weights plus KV cache), otherwise its weights. On a node with capacity learning turned on, the adaptive ceiling from the capacity learner also caps available: a 128GB laptop in learned-low mode has a 16GB ceiling.
Signal 3: Queue Depth Penalty (up to -30)
A hot model on a saturated node is less attractive than a warm model on an empty node.
depth = in_flight + pending
penalty = min(30, depth x 6)
A queue of 5 subtracts the maximum 30 points. This naturally spreads load when multiple nodes can serve the same model.
Signal 4: Estimated Wait Time Penalty (up to -25)
Queue depth alone is misleading. A queue of 3 on a fast 7B model completes in seconds. A queue of 3 on a slow 70B model takes minutes.
The router uses p75 historical latency per node per model:
est_wait = (in_flight + pending) x p75_latency_ms
penalty = min(25, est_wait_seconds / 10)
During the first 7 days before enough data is collected, the router uses a heuristic based on memory bandwidth and model size.
Signal 5: Role Affinity (up to +25)
Large models belong on fast machines. Small models should run on lighter hardware to preserve the fast machine's capacity. The router scores this from each node's memory bandwidth when it's known (the real bottleneck for prompt processing), and falls back to total memory when it isn't.
bandwidth bonus = 5 + bandwidth_GBps / 40, capped at 25
100 GB/s ~+7.5 400 GB/s ~+15 800 GB/s +25
Model over 20 GB: full bandwidth bonus (fast nodes win)
Model under 8 GB: 18 - 0.6 x bonus, at least +3 (slower nodes win)
In between: 0.6 x bonus
Bandwidth unknown (memory-tier fallback):
Model over 20 GB: 128GB+ node +15, 32GB+ node +5, smaller +0
Model under 8 GB: 32GB-or-less node +15, up to 128GB +8, larger +3
Machine brand never enters into it: a Linux server or Windows PC is placed by the same numbers as a Mac. Size thresholds are configurable (FLEET_SCORE_ROLE_LARGE_THRESHOLD_GB, FLEET_SCORE_ROLE_SMALL_THRESHOLD_GB).
Without affinity, every small-model request would drift to the most powerful machine, crowding out the large models it's uniquely suited for.
Signal 6: Availability Trend (up to +10)
Only nodes with capacity learning turned on (it is opt-in) get this signal; every other node scores 0 here. The bonus is the node's availability score, 0.0 to 1.0, times 10:
bonus = availability_score x 10
The availability score blends the learned baseline for this hour of the week, the current CPU and memory load, and whether CPU use has been rising or falling over the last 5 minutes (see Adaptive Capacity). Rising use lowers the score, which keeps a long inference request off a machine whose owner just sat down to start working.
Signal 7: Context Fit (+15 to -15)
Rewards nodes whose loaded context window comfortably handles the estimated request size. Penalizes nodes where the request might trigger a context resize (which causes a model reload).
Signal 8: Session Affinity (up to +20)
Pins a multi-turn conversation to the node that already holds its warm prefix cache. On turn N+1 that node re-uses the cached prefix, so it costs a few hundred prefill tokens instead of re-encoding the whole ~30K-token history, and it avoids the decode interference a fresh prefill inflicts on other streams sharing the node. The bonus is deliberately not decisive: a warm cache is worth a lot, but the bonus stays below Signal 1's +50 and shrinks as the pinned node's queue grows, so a saturated node loses to an idle peer. Added in v0.9.1.
session bonus = 20 / (1 + depth / 2)
depth = requests running or waiting for this model on that node
Signal 3 takes 6 points per queued request at the same time, so between two otherwise identical nodes the session's node leads by 20 with an empty queue, by 7.3 with one request, and trails by 2 once two requests are running or waiting there.
Stage 3: Queue and Execute
The highest-scoring node wins. The request enters that node's dedicated queue for that model. Each node+model pair has its own queue with dynamic concurrency calculated from available memory.
Stage 4: Pre-Warm
Every 5 seconds, a background loop checks each queue: is it getting deep?
If a queue holds 3 or more requests, running plus waiting (the pre-warm threshold), the router loads the same model on the runner-up node by sending an empty generate request. By the time overflow requests arrive, the model is already hot.
A lock prevents duplicate pre-warm requests for the same model on the same node.
Stage 5: Rebalance
A background process runs every 5 seconds:
- Scans all queues for ones with 4 or more requests waiting (the rebalance threshold)
- For each overloaded queue, checks if another node has the same model hot with spare capacity
- Moves pending requests (not in-flight) to the better node
- Caps movement at 3 requests per cycle to prevent oscillation
The rebalancer only moves requests to nodes where the model is already loaded, it never triggers cold loads.
Fallback Chain
When a request lists fallback models, the router scores the primary and every fallback in the same pass, each through the full elimination and scoring pipeline:
- If any of them is already loaded on a surviving node, the first such model in list order wins. A loaded fallback is used rather than cold-loading the primary.
- Otherwise the first model in the list that some node can load wins, and that node loads it.
- If no node can serve any of them, the request waits in the holding queue, rescoring the whole list every 2 seconds for up to 30 seconds.
- If nothing frees up (and auto-pull cannot fetch the model), the router returns 503.
Configuration
All scoring weights are tunable via environment variables:
| Variable | Default | Effect |
|---|---|---|
FLEET_SCORE_MODEL_HOT | 50 | Increase to prefer hot models more aggressively |
FLEET_SCORE_QUEUE_DEPTH_PENALTY_PER | 6 | Decrease to tolerate deeper queues |
FLEET_SCORE_ROLE_AFFINITY_MAX | 15 | Shown on the dashboard but not read by the scorer in v0.9.5; role affinity is capped at +25 (bandwidth-aware, the default) or +15 (memory-tier fallback) |
FLEET_PRE_WARM_THRESHOLD | 3 | Lower to pre-warm earlier |
FLEET_REBALANCE_THRESHOLD | 4 | Lower to rebalance more aggressively |
See the full configuration reference for all 44+ variables.
Next Steps
- Adaptive Capacity, How the capacity learner feeds into Signal 2 and Signal 6
- Deployment, Monitoring and tuning in production
- API Reference, Response headers that show routing decisions