Connect Desktop Chat Apps to Ollama Herd

The exact setting for each chat app that works with the router today, and an honest list of the ones that don't.

The Ollama Herd router answers on one address as two kinds of server at once: an Ollama server and an OpenAI-compatible server. That means most desktop chat apps connect to your whole fleet with a single setting, the same field you would use for a remote Ollama. This page gives the exact field for each app that works with Herd v0.9.5, and lists the apps that do not work yet, with the reason.

The short version: point the app at http://<router-ip>:11435. If the app's provider is called Ollama, enter that URL as-is. If it is called OpenAI-compatible or custom, add /v1 and type any placeholder API key.

Setup steps come from each app's source code at its latest release as of September 2026 (from the vendor's documentation for closed-source apps). Menus move between versions. After saving, use the app's Test or Check button where it has one, with one exception noted under Cherry Studio.

Before you start: find the router address

The router listens on port 11435 on every network interface of the machine where you ran herd. From a chat app on another machine, you need that machine's LAN address. On the router machine:

ipconfig getifaddr en0     # macOS: Wi-Fi on most Macs; try en1 for some Ethernet setups
hostname -I                # Linux: the first address is usually the LAN one
ipconfig                   # Windows: read the IPv4 Address line

Then check the address from the machine where the chat app runs. If this command fails, the app will fail too, so fix it first (in Windows PowerShell, type curl.exe):

curl -s http://<router-ip>:11435/api/tags

You should get a JSON list of the models in your fleet. The dashboard at http://<router-ip>:11435/dashboard shows which nodes are online and, once you start chatting, which node served each request. If the chat app runs on the router machine itself, http://localhost:11435 works.

Two common mistakes:

  • Port 11434 instead of 11435. Many apps prefill Ollama's default port. On a fleet machine, 11434 is that machine's own Ollama, so requests bypass routing entirely and you only see one machine's models.
  • A firewall on the router machine. If curl works on the router but not from another machine, allow incoming connections on port 11435: in the macOS firewall settings, with sudo ufw allow 11435/tcp (or firewall-cmd) on Linux, or with a Windows Defender Firewall inbound rule on Windows. To reach the router from outside your home network, see secure remote access with Tailscale.

Which setting do I use?

Almost every chat app offers one of two provider types. Pick the row that matches the field in front of you:

The app's provider fieldURL to enterAPI key
Ollama (native)http://<router-ip>:11435 (no /v1)None
OpenAI-compatible, custom, or "OpenAI API"http://<router-ip>:11435/v1Any non-empty placeholder, such as sk-no-key

Herd does not check API keys. Some apps refuse to save a provider with an empty key field, so the placeholder only exists to satisfy the app. When an app offers both types, either works; the Ollama one has less to get wrong (no key, no /v1).

Compatibility at a glance

AppRuns onProvider typeBase URLWatch for
Page AssistChrome, Edge, Firefox extensionOllama:11435Odd dates in its model table (cosmetic)
LibreChatSelf-hosted (Docker)Custom endpoint:11435/v1/Docker networking
WitsyMac, Windows, LinuxOllama:11435Embedding models in the chat list
JanMac, Windows, LinuxOpenAI-compatible:11435/v1Must include /v1
Chatbox (desktop)Mac, Windows, LinuxOllama:11435Type http://; Knowledge Base unsupported
AnythingLLMDesktop or DockerOllama:11435Set the context window by hand
Cherry StudioMac, Windows, LinuxOllama:11435Its Check button fails; chat works
BoltAIMac, iOSOpenAI-compatible:11435/v1Placeholder API key
Msty Studio (desktop)Mac, Windows, LinuxOpenAI-compatible:11435/v1Web version cannot connect
Open WebUISelf-hosted (Docker)Ollama:11435Covered in its own guide

Setup, app by app

Page Assist (browser extension)

  1. Open Page Assist's settings (the gear icon in its side panel or web UI).
  2. Go to Ollama Settings.
  3. Set Ollama URL to http://<router-ip>:11435, with no /v1, and save.

Fleet models then appear in the model picker. Page Assist is the one browser-based option that works, because an extension's host permissions let it call the router directly instead of being blocked the way a normal web page is (see not compatible yet). One cosmetic quirk: its Manage Models table shows odd digest and modified-date values, because Herd's model list does not include those fields. Chat is unaffected.

LibreChat (self-hosted)

Add Herd as a custom endpoint in librechat.yaml:

endpoints:
  custom:
    - name: "Ollama"
      apiKey: "ollama"          # must be non-empty; Herd ignores it
      baseURL: "http://<router-ip>:11435/v1/"
      models:
        default: ["llama3.2:3b"]  # shown until the fetch completes
        fetch: true               # pull the model list from the fleet

Restart LibreChat after editing the file. LibreChat's server makes these requests, not your browser, so browser cross-origin rules do not apply. If LibreChat runs in Docker on the same machine as the router, use http://host.docker.internal:11435/v1/ as the base URL; on Linux Docker that name needs a host-gateway mapping, shown in the Open WebUI guide's Docker section.

Witsy

  1. Open Settings → Models → Ollama.
  2. Set API Base URL to http://<router-ip>:11435 (no /v1).
  3. Refresh the model list.

Side effect: embedding models from your fleet also appear in Witsy's chat model list, because Herd's model list covers every model the fleet holds. Pick a chat model; an embedding model cannot answer.

Jan

  1. Open Settings → Model Providers → Add Provider and choose the OpenAI-compatible type. Name it something like Herd.
  2. Set Base URL to http://<router-ip>:11435/v1. The /v1 is required here.
  3. Enter any non-empty API key, such as sk-no-key.

Jan's document chat computes its embeddings locally inside Jan, so it does not depend on the router for that.

Chatbox (desktop)

  1. Open Settings → Model Provider → Ollama.
  2. Set API Host to http://<router-ip>:11435. Type the http:// yourself: if you leave the scheme off, Chatbox assumes https:// and the connection fails.

Chat works. Chatbox's Knowledge Base feature does not, because it needs an OpenAI-style /v1/embeddings endpoint, which Herd v0.9.5 does not serve. Chatbox's web version cannot reach Herd at all (see below); use the desktop app.

AnythingLLM

  1. In Settings, open the LLM preference page and choose Ollama.
  2. Enter Ollama Base URL by hand: http://<router-ip>:11435, with no trailing slash. AnythingLLM's auto-detect only probes port 11434, so it will not find the router on its own.
  3. Open the advanced settings and set Model context window to the context size your model actually runs with.

Step 3 matters. AnythingLLM normally reads a model's context size from Ollama's /api/show endpoint, which Herd does not serve, so without a value it silently falls back to 4,096 tokens and trims long chats and documents to fit. The same missing endpoint means AnythingLLM agents fall back to its prompt-based tool calling instead of the model's native tool calling.

For document chat, AnythingLLM's built-in embedder runs locally and needs nothing from the router. If you want the fleet to compute embeddings, choose the Ollama embedder with the same URL (it calls /api/embed, which Herd serves), not the Generic OpenAI embedder (it needs /v1/embeddings).

Cherry Studio

  1. Open Settings → Model Services → Ollama and enable it.
  2. Set the API address to http://<router-ip>:11435.
  3. Fetch or add models from the list, then send a test message.

The model list and chat work, but Cherry Studio's Check button reports a failure, because the check calls /api/show. Ignore the check result and confirm with a real message instead.

BoltAI (Mac and iOS)

  1. Open Settings → AI Providers → Add a Provider and choose the OpenAI-compatible type.
  2. Set the base URL to http://<router-ip>:11435/v1.
  3. Enter any non-empty API key.

On an iPhone or iPad, the device must be on the same network as the router, or connected to it over a private network such as Tailscale.

Msty Studio (desktop)

  1. Open Model Hub → Model Providers → Add Provider.
  2. Choose the OpenAI-compatible type and set the endpoint to http://<router-ip>:11435/v1, with a placeholder key if asked.

This applies to the desktop app. Msty's web version runs in a browser tab and cannot reach Herd.

Open WebUI

Open WebUI works as an Ollama connection to http://<router-ip>:11435. It has enough moving parts (Docker networking, context overrides, direct connections that bypass routing) that it has its own page: Open WebUI with multiple Ollama servers.

Not compatible with v0.9.5

These clients do not work with the router in v0.9.5. In each case the cause is a specific endpoint or header the client depends on:

ClientWhat happensWhy
Enchanted, Ollamac, ReinsThe model list fails to loadThey decode fields from Ollama's /api/tags (such as digest, modified date, and the standard model details) that Herd's model list does not return, so decoding fails.
Ollama's own desktop appChats failIt calls /api/show before every chat, and Herd does not serve /api/show.
Hollama, TypingMind, and browser-based chat apps generally, including the web versions of Chatbox and MstyCannot connectA web page may only call another origin if that server allows it with CORS headers. Herd sends no CORS headers and does not answer the browser's preflight request. Pages served over HTTPS are also blocked from calling a plain-HTTP address (mixed content).

If you want a chat UI in the browser today, use one whose server makes the requests (Open WebUI or LibreChat), or the Page Assist extension. The same applies to any feature that needs OpenAI-style /v1/embeddings: use an app or setting that embeds locally or through Ollama's /api/embed.

Troubleshooting

  • Empty model list. Run curl -s http://<router-ip>:11435/api/tags from the same machine as the app. If that is empty too, check the dashboard: no nodes are online, or none has models.
  • 404 errors on every request. The /v1 is wrong for the provider type: an Ollama field with /v1 added, or an OpenAI-compatible field without it. Some apps add /v1 themselves, so a URL ending in /v1/v1 in an error message means you should remove yours.
  • Connection fails with an https address. The app added https:// for you. Type http:// explicitly.
  • Long chats forget their beginning. The app is using a small default context window. Set it by hand if the app has the field (AnythingLLM needs this).
  • Requests do not show on the dashboard. The app is talking to port 11434 (a single Ollama) instead of 11435.

For problems below the chat app (a node that will not join, a model that loads slowly), see troubleshooting local LLMs.

Related Reading