How to Run Codex CLI with Local Models
Point OpenAI Codex at local models on any machine that runs Ollama: a Mac, a Linux box with an NVIDIA GPU, a Windows PC, or all of them behind one endpoint. One provider block in config.toml, no model map to maintain, and the same routing that serves Claude Code. Verified end to end on the CLI and the desktop app.
TL;DR
Add this to ~/.codex/config.toml, then run codex:
model_provider = "herd"
[model_providers.herd]
name = "Ollama Herd"
base_url = "http://localhost:11435/v1"
wire_api = "responses"
No model map needed. Herd auto-routes whatever model id Codex sends to the best coding model you actually have loaded. Pull a model and go:
ollama pull qwen3-coder:30b
The same config block also works for the Codex desktop app, which reads the same file. See Quickstart if the herd is not running yet.
Before you start: Codex is served by Ollama-backed models only (not mlx: models), the endpoint is stateless (Codex's default mode, so no previous_response_id), and hosted tools such as web search are dropped. Details in Limitations.
How to run Codex CLI with a local model, step by step
- Install Ollama and pull a coding model on each machine you want to use. Mac, Linux, Windows and NVIDIA machines all work:
ollama pull qwen3-coder:30b. - Install and start Herd.
pip install ollama-herd, thenherdon one machine andherd-nodeon every machine running Ollama. Nodes find the router over mDNS. - Set the context length on each node to at least 64K (see context length below).
- Add the provider block from the TL;DR to
~/.codex/config.toml, withmodel_provider = "herd"above the first[table]header. - Run
codex, orcodex -m qwen3-coder:30bto force one model. - Confirm it routed locally. The
X-Fleet-Served-Modelresponse header names the model that answered:
curl -si localhost:11435/v1/responses \
-H 'content-type: application/json' \
-d '{"model":"gpt-5-codex","input":"say ok","stream":false}' | grep -i x-fleet-served-model
Ollama already runs Codex on one machine. When do you need Herd?
Ollama has its own Codex paths: ollama launch codex, Codex's built-in codex --oss mode, and an integration for the desktop app. If you have one machine and one model, use them. They are the simplest option.
Herd is for everything past that:
- Several machines behind one
base_url, with each request sent to the machine that already has the model loaded and the shortest queue. - Any model id resolved. Whatever id Codex sends, including ones you never chose (like the one the desktop app uses for chat titles), lands on the best coding model you have loaded.
- Images routed to a vision model even when your coding model cannot see.
- Fixes for the two tool calls local models reliably get wrong under Codex, covered in Troubleshooting.
Ollama's Codex paths per Ollama's Codex integration docs, checked 2026-09-29.
Context length: give Codex room
Codex's system prompt alone is about 27 KB, and a one-line fix cost about 109K tokens end to end in our runs, because every turn resends the whole conversation. Ollama sizes its default context by GPU memory (4K under 24 GB, 32K from 24 to 48 GB, 256K above that) and recommends at least 64000 tokens for coding agents.
Set OLLAMA_CONTEXT_LENGTH on every node to the largest context the fleet needs, for example 131072. Do not set it lower than a model needs: on our reference fleet, a too-small value collapsed prefix caching and took time to first token from about 1.0 s to about 6.3 s, with decode speed unchanged, so throughput graphs looked healthy while every request got slower.
# macOS (Ollama app): set it, then quit and reopen the app
launchctl setenv OLLAMA_CONTEXT_LENGTH 131072
# Linux (systemd)
sudo systemctl edit ollama # add: Environment="OLLAMA_CONTEXT_LENGTH=131072"
# Windows: set it as a system environment variable, then restart Ollama
Codex does not know a local model's window, because it keys model metadata off its own table of OpenAI names. OpenAI documents a model_context_window key in config.toml for telling it.
model_context_window is vendor-documented in Codex's advanced configuration docs; we have not tested it against Herd. Ollama figures per Ollama's context length docs, checked 2026-09-29.
Router on another machine, or a Linux and NVIDIA fleet
Nothing in the Codex path depends on the node being a Mac. Herd translates Codex's Responses API requests to Ollama's API, so a Linux box with an NVIDIA card, a Windows PC and a Mac all serve them the same way. When the router runs on a different machine from Codex, point base_url at it:
[model_providers.herd]
name = "Ollama Herd"
base_url = "http://ROUTER_IP:11435/v1"
wire_api = "responses"
Running a node in Docker? Set FLEET_NODE_OLLAMA_HOST to an address the router can reach. See running a node in a container, and secure remote access with Tailscale for reaching the router from outside your LAN.
Our end-to-end Codex runs were on Macs; the routing path is the same on every platform. Memory-fit scoring reads system RAM, not GPU VRAM. mlx: models are Apple Silicon only and not served on this endpoint anyway (see Limitations).
First, the wall you probably just hit
If you tried pointing Codex at Ollama, LM Studio, or any OpenAI-compatible server and it failed, this is why: Codex removed Chat Completions support in February 2026. wire_api = "responses" is the only valid value now, and it is the default.
That breaks most advice you will find. A server that implements only /v1/chat/completions cannot serve Codex at all, no matter how OpenAI-compatible it is otherwise. Guides still showing wire_api = "chat" predate the change.
Ollama Herd implements the Responses API directly at /v1/responses, so there is no translating gateway in the middle. That is what the wire_api = "responses" line above is telling Codex.
The config mistake that costs an hour
If you are appending to an existing ~/.codex/config.toml, model_provider = "herd" must go above the first [table] header. TOML assigns a bare key to whatever table precedes it, so pasting it at the bottom silently turns it into desktop.model_provider. Codex never sees it, no error is raised, and it quietly keeps using the default provider.
model_provider = "herd" # first line, before ANY [section]
[some.existing.section]
...
[model_providers.herd] # a table header, so this can live anywhere
Two more naming rules. The provider id cannot be openai, ollama or lmstudio, which Codex reserves for its built-in providers. And if a provider entry anywhere in the file still says wire_api = "chat", current Codex refuses to load the config, even when that provider is not the one you are using.
Keep your hosted setup as the default: a profile file
If you want Codex to keep using OpenAI by default and switch to your fleet on demand, Codex 0.134.0 and later reads profiles from separate files rather than [profiles.x] tables. Put the provider settings in ~/.codex/herd.config.toml instead of the main file:
# ~/.codex/herd.config.toml
model_provider = "herd"
[model_providers.herd]
name = "Ollama Herd"
base_url = "http://localhost:11435/v1"
wire_api = "responses"
Then run codex --profile herd when you want local models.
Profile files and the reserved provider ids are vendor-documented in Codex's advanced configuration docs; we verified Herd with the main config.toml, not with a profile file.
Using the Codex desktop app
No extra setup. The desktop app reads the same ~/.codex/config.toml as the CLI, so the one provider block above serves both surfaces. Configure it once, and the app routes to your fleet the next time you open it.
OpenAI now ships Codex inside the ChatGPT desktop app, for macOS, Windows and Linux (OpenAI's app docs). We verified the shared config.toml on a Mac during the v0.9.0 release in July 2026; we have not re-tested the Windows or Linux app.
Three things behave differently in the app, all verified against a real conversation:
- It sends more than one model id. The conversation uses the id from the in-app picker, and a second id fires alongside it for chat-title generation. Both auto-routed with no map. This is the clearest argument for auto-routing over a hand written model map: the map would have had to predict an id you never chose.
- The model picker may look empty. Codex decodes
/v1/modelsagainst its own undocumented schema, so the picker can fail to populate. This is cosmetic. Inference is unaffected, every turn still routes to your fleet, and the app carries its own model list regardless. Set the model inconfig.tomlrather than the picker. - There is no
-mflag. Where the CLI takescodex -m qwen3-coder:30b, the app relies on config plus auto-routing. Pin a model withFLEET_ANTHROPIC_MODEL_MAPif you want a specific one every time.
Everything else on this page, the provider block, the TOML gotcha, model resolution, and troubleshooting, applies identically to both surfaces.
Verified end to end, on both surfaces
Plenty of projects document Codex on local models. This one was run against a real client on a real fleet, and the findings are published rather than asserted.
- Codex CLI against the live fleet:
provider: herd, model idgpt-5.6-solauto-routed toqwen3-coder:30b,x-fleet-served-modelset on the response, traced asoriginal_format='responses'. - Codex desktop app, same config file, no extra setup: a real multi turn conversation served entirely by the fleet, three streaming turns of 6,209 to 8,579 prompt tokens each, all completed on one node via
qwen3-coder:30b. - Three different model ids auto-routed with no map:
gpt-5-codex,gpt-5.6-sol, andgpt-5.6-luna. That last one fired for chat title generation. A hand written map would have had to guess it existed, which is exactly the fragility auto-routing removes. - Real agentic coding, 4 tasks out of 4. Codex fixed two bugs and created a new module from scratch to satisfy a failing import, then ran pytest itself and reached green. Not model-specific:
qwen3-coder:30bandgpt-oss:120bboth drove the full loop. File editing and file creation both work.
Then a 26 hour soak across every code change in the release: 11,925 requests at 99.85% success, including 302 requests to /v1/responses with zero failures. The 18 failures all trace to a known slowness issue with one specific model, not to the Codex path.
Verification also found bugs unit tests had not, including tool calls being dropped in both directions and a model-listing schema that failed Codex's whole decode. Those are fixed. Spec-complete is not the same as client-verified.
Getting good results
Three things make the difference between a loop that finishes and one that wanders, all measured on local models:
- Name the tool in your prompt. This is the single biggest quality lever. "Diagnose the root cause and fix it" sent
qwen3-coder:30bexploring until it exhausted its budget. The same task phrased as "use the apply_patch command to fix X, then re-run pytest" succeeded. - If it stops after announcing an action, say "continue". About 3% of turns end with the model saying "Let me run pytest, then fix it" and stopping, which the protocol reads as done. It is stochastic, so a fresh chat will not help.
- Check the diff, not the model's summary. In one run that produced a correct fix, the model also claimed it could not run pytest, minutes after running it successfully. The code was right; the narration was not.
Budget for tokens too: a one-line fix cost about 109K tokens end to end, because agentic loops resend the whole conversation each turn and Codex's system prompt alone is around 27 KB.
How model selection works
Codex sends a model id such as gpt-5-codex. Herd resolves it in this order:
-
Did you pin it? Is there a
FLEET_ANTHROPIC_MODEL_MAPentry for this exact id?Yes: use it. The same map serves Claude Code and Codex. An entry pointing at anmlx:model gets a clear 503, because this endpoint serves Ollama models only. -
Does the id name a model your fleet actually has?Yes: use it as-is. This is why
codex -m qwen3-coder:30bworks. -
Is a suitable model already loaded somewhere on the fleet?Yes: use the best loaded one, no cold load. Coding models rank first,
mlx:models are left out, and a request with an image considers only vision models. -
Is a suitable model on disk?Yes: use the best on-disk one, accepting a cold load.
-
Does the map have a
"default"key?Yes: use it. Otherwise a clear 404 tells you to pull a model.
mlx: model; FLEET_ANTHROPIC_AUTO_ROUTE=false skips steps 3 and 4.To pin one id without affecting anything else:
export FLEET_ANTHROPIC_MODEL_MAP='{"gpt-5-codex": "qwen3-coder:30b"}'
To disable auto-routing and require an explicit map, set FLEET_ANTHROPIC_AUTO_ROUTE=false.
What the fleet gives you
- Multi-node routing. Every request is scored across all your machines, Macs, Linux servers and Windows PCs alike, on whether the model is already loaded, memory fit, queue depth, model affinity, and context fit, then sent to the best one.
- No model map to maintain. Auto-routing tracks what you have actually pulled instead of a list you keep in sync by hand.
- Routing around the machine you are using. Turn on adaptive capacity (
herd-node --learn-capacity) and a Mac in a video call, or a machine under sustained heavy CPU load, is paused, so your agent does not fight you for your own laptop. - Images route themselves. When Codex sends a picture, Herd routes that request to a vision-capable model on your fleet even when your coding model cannot see. Verified with
gemma3:27b. If you have no vision model,ollama pull gemma3:4bis about 6 GB. - Both of Codex's tool protocols work. Codex sends tools two different ways depending on the model slug, and the newer shape (openai/codex#31894) leaves the model with nothing callable even against OpenAI's own hosted models. Herd translates both, so you do not hit it.
- One endpoint for both agent CLIs. The same herd serves Codex here and Claude Code on
/v1/messages. - Observability. Every request lands in the trace store, and
X-Fleet-Served-Modeltells you exactly which model and node answered.
Choosing a model
Codex calls tools constantly, so tool-use quality matters more than chat quality. Coding-tuned models with a real tool-call parser hold up; general chat models drop calls and hallucinate arguments.
| Model | Agentic coding under Codex | Notes |
|---|---|---|
qwen3-coder:30b | Verified end to end | Best general-purpose pick. About 19 GB, 256K context |
gpt-oss:120b | Drives the loop, weak at converging | Reasoning model. Sustained 40+ tool-calling turns but explored without landing an edit; better for analysis than editing |
qwen3:32b | Untested | Strong reasoning, good tool use |
Newer coding families have since landed in Ollama's library: Qwen 3.6, Gemma 4, GLM-5, Kimi K2 Code, DeepSeek V4. See the current tool-capable list. Smaller or non-coding-tuned models tend to drop tool calls or invent arguments.
You do not need to tell Herd which one you picked. Auto-routing sends Codex to the best coding model you have loaded, so upgrading is ollama pull and nothing else.
Limitations, stated plainly
- MLX-backed models are not served on this path yet. Auto-routing skips
mlx:models for Codex so you always land on a working Ollama-backed model. An explicitmlx:mapping returns a clear503. Claude Code's/v1/messagesdoes serve MLX. - No stateful chaining. Herd does not persist responses, so
previous_response_idis rejected with a400. Codex's default stateless mode, resending the conversation each turn, is what is supported and what it actually does. - Hosted tools are dropped.
web_search,file_search, and MCP tool items have no local equivalent. Function tools pass through normally. - The model picker may not populate. Codex decodes
/v1/modelsagainst its own undocumented schema rather than OpenAI's. This is cosmetic, inference is unaffected, and every turn routes normally. Specify the model with-mor in config rather than the picker. - Codex will not recognise your model's name. It keys metadata off its own table of OpenAI slugs, so expect
Model metadata for <name> not found. Harmless: tool calling, editing, and multi-turn all work anyway.
Troubleshooting
404 not available on any node. You have not pulled a chat or coding model, or the one you pinned is not on the fleet. Run ollama pull qwen3-coder:30b, or check curl localhost:11435/v1/models.
Codex errors about the wire protocol. Confirm wire_api = "responses" in ~/.codex/config.toml. wire_api = "chat" was removed from Codex in February 2026.
400 previous_response_id not supported. Your client is using server-side conversation state. Stateless mode, which Codex uses by default, is what Herd supports.
Codex ignores your config entirely. Check that model_provider = "herd" sits above the first [table] header. See the config mistake above.
Tool calls come back as plain text. The model is too small or not coding-tuned. Switch to one of the recommended models above.
Codex insists it "cannot execute commands." Start a new chat. Once Codex has said it cannot run commands, it keeps concluding that even with working tools in front of it: in one measurement a 43-message conversation made zero tool calls while a fresh 6-message chat ran the command. Its explanation is not evidence about your setup, so check the herd log instead:
grep 'Responses\[' ~/.fleet-manager/logs/herd.jsonl | tail -3
# tools=3 custom=['exec'] -> tools ARE reaching the model
# tools=0 -> genuinely no tools; check your provider config
Commands fail in the sandbox without asking for approval. In Codex, the model has to request escalation itself. Local models sometimes write that request as JSON where Codex expects JavaScript, so it never reaches Codex. Herd repairs that shape automatically; grep 'emitted a JSON object instead of' ~/.fleet-manager/logs/herd.jsonl shows when the repair fired. Setting Full Access hides the symptom, but it is a workaround, not a fix.
The session reads files but never edits anything. Local models often call apply_patch as a tool, which Codex rejects, because apply_patch is a program on the sandbox path, not a tool. Herd rewrites that call into a shell command automatically and logs it at WARNING. If edits still do not land, name the tool in your prompt ("use the apply_patch command to fix X").
Frequently asked questions
Can Codex CLI use a local model?
Yes. Codex supports custom model providers via ~/.codex/config.toml. Point base_url at an endpoint that speaks the OpenAI Responses API and Codex will use it. Ollama Herd serves the Responses API at /v1/responses, so a four line provider block is all it takes. No model map is required because Herd auto-routes whatever model id Codex sends to the best coding model you have loaded.
Why does Codex fail against my Ollama or OpenAI-compatible endpoint?
Because Codex removed Chat Completions support in February 2026. wire_api = "responses" is now the only valid value, so an endpoint that only implements /v1/chat/completions cannot serve Codex at all. Most guides showing wire_api = "chat" are stale. You need a server that implements the Responses API, or a translating gateway in front of one.
Does the Codex desktop app work with local models too?
Yes. The desktop app reads the same ~/.codex/config.toml as the CLI, so a single provider block serves both surfaces with no extra setup. A real multi turn conversation in the app was served entirely by a local fleet during verification.
Do I need a model map to use Codex with Ollama Herd?
No. Herd auto-routes any model id Codex sends to the best coding model currently loaded on your fleet, falling back to the best one on disk. This matters because clients invent model ids: during verification the desktop app sent a second id for chat title generation that no hand written map would have predicted. Set FLEET_ANTHROPIC_MODEL_MAP only if you want to pin a specific id to a specific model.
What is the difference between running Codex on one machine and on a fleet?
On one machine Codex competes with everything else that machine is doing. Ollama Herd scores every request across all your machines, Macs, Linux servers and Windows PCs alike, on whether the model is already loaded, memory fit, queue depth, and context fit, then sends it to the best one. You also get one endpoint that serves Codex and Claude Code together.
Does the Codex desktop app need different setup than the CLI?
No. The desktop app reads the same ~/.codex/config.toml as the Codex CLI, so a single model_providers block configures both. There is no separate desktop configuration, and no extra steps once the CLI is working.
Why is the Codex model picker empty when using a custom provider?
Codex decodes the /v1/models response against its own undocumented schema rather than the standard OpenAI one, so a custom provider can leave the picker unpopulated. It is cosmetic: inference is unaffected and every turn still routes normally. Select the model in config.toml, or with the -m flag on the CLI, instead of the picker.
Does Codex work with local models on Linux or NVIDIA?
Yes. Herd translates Codex's Responses API requests to Ollama's API, so any machine that runs Ollama can serve them: a Linux box with an NVIDIA GPU, a Windows PC, or a Mac. Point base_url at the router's address. Our end-to-end Codex runs were on Macs, but the routing path is the same on every platform.
What context length does Codex need with a local model?
At least 64K. Codex's system prompt alone is about 27 KB and every turn resends the whole conversation, so a one-line fix cost about 109K tokens end to end in our runs. Ollama recommends at least 64000 tokens for coding agents. Set OLLAMA_CONTEXT_LENGTH on every node to the largest context the fleet needs, not lower.
Can I name my Codex provider ollama?
No. Codex reserves openai, ollama and lmstudio for its built-in providers, and a custom provider cannot reuse those IDs. That is why this guide names the provider herd.
Which variable pins a model for Codex?
FLEET_ANTHROPIC_MODEL_MAP. The name is historical: one resolver serves Claude Code and Codex, so the same map pins model ids from both, for example {"gpt-5-codex": "qwen3-coder:30b"}. Without it, Herd auto-routes every id to the best coding model you have loaded.
Related Reading
- Claude Code CLI, the same fleet for the other major agent CLI
- Quickstart, get the herd running first
- Routing Engine, the eight signals behind each decision
- OpenClaw, running local agents across your machines
- Deployment, running nodes on Linux, in containers, and against a remote Ollama