Reranking on your own hardware

Vector search gets the right chunk into your top 20. A reranker gets it into your top 3. Ollama has no endpoint for that step, and this is how Herd fills it.

Ollama has no rerank endpoint. The request has been open since March 2024 (issue #3368, 113 comments and 381 reactions as of October 2026), there is no merged pull request, and the current API docs do not mention reranking at all. So if you are building RAG on Ollama, the rerank step is the one piece you cannot get from the same server as everything else.

Ollama Herd serves POST /v1/rerank as of v0.10.0. It runs on the native fastembed server that nodes installed with the embedding extra already run, so it adds no new dependency and no new process to supervise.

Why a reranker changes RAG results

Vector search is fast and approximate. It compares one embedding of your query against one embedding of each chunk, which means it never actually reads the two together. That is what makes it fast enough to scan a corpus, and it is also why the top 20 results are usually in roughly the right neighbourhood but not in the right order.

A cross-encoder reranker reads the query and the document in the same forward pass, so it can score the pair rather than compare two summaries of it. It is far too slow to run over a whole corpus and well suited to reordering the 20 to 100 candidates vector search already narrowed to. The usual pipeline:

  1. Retrieve 50 candidates by embedding similarity, which is cheap.
  2. Rerank those 50 with a cross-encoder, which is accurate.
  3. Pass the top 5 into the model's context.

Skipping step 2 is the most common reason a local RAG setup feels worse than it should: the right chunk was retrieved, it just was not in the first five.

Using it

Point your client at the router. The request and response follow the Jina and Cohere shape, which is what Open WebUI, Dify, LangChain and LlamaIndex already speak.

curl http://localhost:11435/v1/rerank -H 'Content-Type: application/json' -d '{
  "query": "how do I stop herd now that launchd is installed?",
  "documents": [
    "launchctl bootout gui/$UID/com.ollamaherd.node",
    "pkill is no longer a stop, launchd restarts it",
    "MLX setup requires re-running setup-mlx.sh"
  ],
  "top_n": 2,
  "return_documents": true
}'

query and documents are required. documents takes plain strings or {"text": ...} objects. model, top_n and return_documents (default false) are optional.

The response is {model, results: [{index, relevance_score, document?}], usage}, sorted best first. index points back into the array you sent, so you can reorder your own objects without asking for the documents back. relevance_score is the cross-encoder logit through a sigmoid, so it lands in 0 to 1 the way Cohere's and Jina's do and a client's relevance threshold behaves the same way.

Which node served it is in the X-Fleet-Node response header, not the body.

Which reranker to use

Weights download on first use and stay cached. Start with the default and only move up if you can measure the difference on your own corpus.

ModelSizeNotes
ms-marco-minilm-l-6-v2 (default)80 MBFastest. English
ms-marco-minilm-l-12-v2120 MBEnglish
jina-reranker-v1-tiny-en130 MBEnglish, long inputs
jina-reranker-v1-turbo-en150 MBEnglish, long inputs
bge-reranker-base1.04 GBQuality pick. English and Chinese
jina-reranker-v2-base-multilingual1.11 GBQuality pick. Multilingual

An 80 MB default is deliberate. A reranker runs over every candidate on every query, so its cost lands on the latency your user feels, not on a one-off index build. Reach for a gigabyte-class model when you have a measurement saying the small one is costing you answers.

What the fleet adds

Reranking is a good fit for routing because it is bursty and short. A query fires one rerank over 50 candidates and then nothing until the next query, so a node that is busy decoding a long completion is the wrong place to send it. Herd scores each request across its eight signals and sends the rerank to a node that can answer now, which is generally not the node currently streaming someone's 20K-token coding session.

It also means the reranker does not compete with inference for the same slots. The native text server is a separate process from Ollama, so a rerank does not consume an Ollama parallel slot the way an embedding through /api/embed would.

Limits and errors

  • At most 1,000 documents per request, each at most 32,000 characters. Beyond either, the request is a 400 rather than a silent truncation.
  • 503 means no node in the fleet runs a rerank-capable text server. Install the extra on at least one node: uv sync --extra embedding.
  • Rerankers and embedding models are separate registries, so you cannot accidentally send an embed request to a reranker or the reverse.

Try it on your own documents

If you already run Ollama Herd with the embedding extra on any node, reranking is already there. Send the curl above and you will get scores back. If you do not, the whole thing is two commands, free and MIT licensed:

# on the machine you want as the router
pip install ollama-herd
herd

# on every other machine
herd-node

Nodes find the router over mDNS on their own, so there is no config file to write. Add the embedding extra on at least one node (uv sync --extra embedding) and that node serves both embeddings and reranking.

The honest way to judge a reranker is on your own corpus. Take 20 queries you care about, retrieve 50 chunks for each, and compare the top 5 with and without the rerank step. That takes an afternoon and tells you more than any benchmark we could publish.

If you are already embedding, one thing changed in v0.10.0: inputs longer than 2,048 tokens are truncated, where they used to run at up to 8,192. That is the fix for a 28 GB memory high-water mark in the node agent. Chunk before embedding if you relied on the old limit, and note that "truncate": false now returns 400, matching Ollama.

Related Reading