Connect Desktop Chat Apps to Ollama Herd
The exact setting for each chat app that works with the router today, and an honest list of the ones that don't.
The Ollama Herd router answers on one address as two kinds of server at once: an Ollama server and an OpenAI-compatible server. That means most desktop chat apps connect to your whole fleet with a single setting, the same field you would use for a remote Ollama. This page gives the exact field for each app that works with Herd v0.9.5, and lists the apps that do not work yet, with the reason.
The short version: point the app at
http://<router-ip>:11435. If the app's provider is called Ollama, enter that URL as-is. If it is called OpenAI-compatible or custom, add/v1and type any placeholder API key.
Setup steps come from each app's source code at its latest release as of September 2026 (from the vendor's documentation for closed-source apps). Menus move between versions. After saving, use the app's Test or Check button where it has one, with one exception noted under Cherry Studio.
Before you start: find the router address
The router listens on port 11435 on every network interface of the machine where you ran herd. From a chat app on another machine, you need that machine's LAN address. On the router machine:
ipconfig getifaddr en0 # macOS: Wi-Fi on most Macs; try en1 for some Ethernet setups
hostname -I # Linux: the first address is usually the LAN one
ipconfig # Windows: read the IPv4 Address line
Then check the address from the machine where the chat app runs. If this command fails, the app will fail too, so fix it first (in Windows PowerShell, type curl.exe):
curl -s http://<router-ip>:11435/api/tags
You should get a JSON list of the models in your fleet. The dashboard at http://<router-ip>:11435/dashboard shows which nodes are online and, once you start chatting, which node served each request. If the chat app runs on the router machine itself, http://localhost:11435 works.
Two common mistakes:
- Port 11434 instead of 11435. Many apps prefill Ollama's default port. On a fleet machine, 11434 is that machine's own Ollama, so requests bypass routing entirely and you only see one machine's models.
- A firewall on the router machine. If
curlworks on the router but not from another machine, allow incoming connections on port11435: in the macOS firewall settings, withsudo ufw allow 11435/tcp(orfirewall-cmd) on Linux, or with a Windows Defender Firewall inbound rule on Windows. To reach the router from outside your home network, see secure remote access with Tailscale.
Which setting do I use?
Almost every chat app offers one of two provider types. Pick the row that matches the field in front of you:
| The app's provider field | URL to enter | API key |
|---|---|---|
| Ollama (native) | http://<router-ip>:11435 (no /v1) | None |
| OpenAI-compatible, custom, or "OpenAI API" | http://<router-ip>:11435/v1 | Any non-empty placeholder, such as sk-no-key |
Herd does not check API keys. Some apps refuse to save a provider with an empty key field, so the placeholder only exists to satisfy the app. When an app offers both types, either works; the Ollama one has less to get wrong (no key, no /v1).
Compatibility at a glance
| App | Runs on | Provider type | Base URL | Watch for |
|---|---|---|---|---|
| Page Assist | Chrome, Edge, Firefox extension | Ollama | :11435 | Odd dates in its model table (cosmetic) |
| LibreChat | Self-hosted (Docker) | Custom endpoint | :11435/v1/ | Docker networking |
| Witsy | Mac, Windows, Linux | Ollama | :11435 | Embedding models in the chat list |
| Jan | Mac, Windows, Linux | OpenAI-compatible | :11435/v1 | Must include /v1 |
| Chatbox (desktop) | Mac, Windows, Linux | Ollama | :11435 | Type http://; Knowledge Base unsupported |
| AnythingLLM | Desktop or Docker | Ollama | :11435 | Set the context window by hand |
| Cherry Studio | Mac, Windows, Linux | Ollama | :11435 | Its Check button fails; chat works |
| BoltAI | Mac, iOS | OpenAI-compatible | :11435/v1 | Placeholder API key |
| Msty Studio (desktop) | Mac, Windows, Linux | OpenAI-compatible | :11435/v1 | Web version cannot connect |
| Open WebUI | Self-hosted (Docker) | Ollama | :11435 | Covered in its own guide |
Setup, app by app
Page Assist (browser extension)
- Open Page Assist's settings (the gear icon in its side panel or web UI).
- Go to Ollama Settings.
- Set Ollama URL to
http://<router-ip>:11435, with no/v1, and save.
Fleet models then appear in the model picker. Page Assist is the one browser-based option that works, because an extension's host permissions let it call the router directly instead of being blocked the way a normal web page is (see not compatible yet). One cosmetic quirk: its Manage Models table shows odd digest and modified-date values, because Herd's model list does not include those fields. Chat is unaffected.
LibreChat (self-hosted)
Add Herd as a custom endpoint in librechat.yaml:
endpoints:
custom:
- name: "Ollama"
apiKey: "ollama" # must be non-empty; Herd ignores it
baseURL: "http://<router-ip>:11435/v1/"
models:
default: ["llama3.2:3b"] # shown until the fetch completes
fetch: true # pull the model list from the fleet
Restart LibreChat after editing the file. LibreChat's server makes these requests, not your browser, so browser cross-origin rules do not apply. If LibreChat runs in Docker on the same machine as the router, use http://host.docker.internal:11435/v1/ as the base URL; on Linux Docker that name needs a host-gateway mapping, shown in the Open WebUI guide's Docker section.
Witsy
- Open Settings → Models → Ollama.
- Set API Base URL to
http://<router-ip>:11435(no/v1). - Refresh the model list.
Side effect: embedding models from your fleet also appear in Witsy's chat model list, because Herd's model list covers every model the fleet holds. Pick a chat model; an embedding model cannot answer.
Jan
- Open Settings → Model Providers → Add Provider and choose the OpenAI-compatible type. Name it something like Herd.
- Set Base URL to
http://<router-ip>:11435/v1. The/v1is required here. - Enter any non-empty API key, such as
sk-no-key.
Jan's document chat computes its embeddings locally inside Jan, so it does not depend on the router for that.
Chatbox (desktop)
- Open Settings → Model Provider → Ollama.
- Set API Host to
http://<router-ip>:11435. Type thehttp://yourself: if you leave the scheme off, Chatbox assumeshttps://and the connection fails.
Chat works. Chatbox's Knowledge Base feature does not, because it needs an OpenAI-style /v1/embeddings endpoint, which Herd v0.9.5 does not serve. Chatbox's web version cannot reach Herd at all (see below); use the desktop app.
AnythingLLM
- In Settings, open the LLM preference page and choose Ollama.
- Enter Ollama Base URL by hand:
http://<router-ip>:11435, with no trailing slash. AnythingLLM's auto-detect only probes port 11434, so it will not find the router on its own. - Open the advanced settings and set Model context window to the context size your model actually runs with.
Step 3 matters. AnythingLLM normally reads a model's context size from Ollama's /api/show endpoint, which Herd does not serve, so without a value it silently falls back to 4,096 tokens and trims long chats and documents to fit. The same missing endpoint means AnythingLLM agents fall back to its prompt-based tool calling instead of the model's native tool calling.
For document chat, AnythingLLM's built-in embedder runs locally and needs nothing from the router. If you want the fleet to compute embeddings, choose the Ollama embedder with the same URL (it calls /api/embed, which Herd serves), not the Generic OpenAI embedder (it needs /v1/embeddings).
Cherry Studio
- Open Settings → Model Services → Ollama and enable it.
- Set the API address to
http://<router-ip>:11435. - Fetch or add models from the list, then send a test message.
The model list and chat work, but Cherry Studio's Check button reports a failure, because the check calls /api/show. Ignore the check result and confirm with a real message instead.
BoltAI (Mac and iOS)
- Open Settings → AI Providers → Add a Provider and choose the OpenAI-compatible type.
- Set the base URL to
http://<router-ip>:11435/v1. - Enter any non-empty API key.
On an iPhone or iPad, the device must be on the same network as the router, or connected to it over a private network such as Tailscale.
Msty Studio (desktop)
- Open Model Hub → Model Providers → Add Provider.
- Choose the OpenAI-compatible type and set the endpoint to
http://<router-ip>:11435/v1, with a placeholder key if asked.
This applies to the desktop app. Msty's web version runs in a browser tab and cannot reach Herd.
Open WebUI
Open WebUI works as an Ollama connection to http://<router-ip>:11435. It has enough moving parts (Docker networking, context overrides, direct connections that bypass routing) that it has its own page: Open WebUI with multiple Ollama servers.
Not compatible with v0.9.5
These clients do not work with the router in v0.9.5. In each case the cause is a specific endpoint or header the client depends on:
| Client | What happens | Why |
|---|---|---|
| Enchanted, Ollamac, Reins | The model list fails to load | They decode fields from Ollama's /api/tags (such as digest, modified date, and the standard model details) that Herd's model list does not return, so decoding fails. |
| Ollama's own desktop app | Chats fail | It calls /api/show before every chat, and Herd does not serve /api/show. |
| Hollama, TypingMind, and browser-based chat apps generally, including the web versions of Chatbox and Msty | Cannot connect | A web page may only call another origin if that server allows it with CORS headers. Herd sends no CORS headers and does not answer the browser's preflight request. Pages served over HTTPS are also blocked from calling a plain-HTTP address (mixed content). |
If you want a chat UI in the browser today, use one whose server makes the requests (Open WebUI or LibreChat), or the Page Assist extension. The same applies to any feature that needs OpenAI-style /v1/embeddings: use an app or setting that embeds locally or through Ollama's /api/embed.
Troubleshooting
- Empty model list. Run
curl -s http://<router-ip>:11435/api/tagsfrom the same machine as the app. If that is empty too, check the dashboard: no nodes are online, or none has models. - 404 errors on every request. The
/v1is wrong for the provider type: an Ollama field with/v1added, or an OpenAI-compatible field without it. Some apps add/v1themselves, so a URL ending in/v1/v1in an error message means you should remove yours. - Connection fails with an https address. The app added
https://for you. Typehttp://explicitly. - Long chats forget their beginning. The app is using a small default context window. Set it by hand if the app has the field (AnythingLLM needs this).
- Requests do not show on the dashboard. The app is talking to port 11434 (a single Ollama) instead of 11435.
For problems below the chat app (a node that will not join, a model that loads slowly), see troubleshooting local LLMs.
Related Reading
- Open WebUI with multiple Ollama servers, the full setup for the most popular self-hosted UI
- Quickstart, if you do not have a router running yet
- Secure remote access, to use these apps away from home
- API Reference, every endpoint the router serves