Skip to main content

LLM Router sidecar (OmniRoute)

Project Brain runs the open-source OmniRoute LLM gateway as a sidecar container. It is the LLM path for the deployment:

  • Coding agents (neusis-code, Claude Code, Cursor, Cline, Aider, …) point at the router's OpenAI/Anthropic-compatible endpoint with an API key minted from /admin/router.
  • Project Brain's own calls — assistant chat and the graph builder — go through the same gateway, using a key it mints for itself. There are no provider API keys in .env and none in Postgres; the admin picks the assistant model at /admin/ai, each project picks its graph builder model on its Settings tab, and the credential lives only in the router's encrypted store.
  • The existing external litellm-proxy deployment keeps working — nothing here replaces or reconfigures it.

Architecture​

neusis-code / Claude Code (developer laptop)
│ Bearer <key minted in /admin/router>
▼
https://APP_DOMAIN/router/v1 ──► Caddy ──► omniroute:20128 (sidecar)
(or https://ROUTER_DOMAIN/v1) ▲
│ provider API keys / OAuth,
PB web / worker ──────────────────────────────┤ stored encrypted in its volume
Bearer <internal key> /v1/chat/completions │
▼
OpenAI / Anthropic / Google / …

PB web ──► omniroute:20128 /api/keys (mint/revoke/reveal keys,
Authorization: Bearer <encrypted token>) ledger in router_api_keys)

Project Brain's own key​

Onboarding mints one extra API key named Project Brain (internal) and stores it encrypted on router_gateway_credentials.internal_key_enc. It is deliberately not in the router_api_keys ledger: that table drives the admin key list and member assignment, and this key belongs to the platform rather than a person. Routers onboarded before this existed get one on the next visit to /admin/router or /admin/ai.

resolveAiSelection() and resolveGraphBuilderSelection() (packages/db/src/ai-selection.ts) pair that key with the assistant's model or the project's graph builder model and hand both to the caller. There is no environment fallback: if the router is not onboarded or no model is selected, the assistant answers 503 and a graph run is recorded as llm_not_configured, each naming the page that fixes it. That is deliberate — a silent fallback to a directly-held provider key is exactly what this consolidation removes.

One gateway quirk is worth knowing: a chat-completions request whose body omits stream is answered with SSE, not JSON. apps/web/lib/llm.ts fills in stream: false so both streamText and generateText work. Without it, non-streaming calls fail with "Invalid JSON response".

The raw agent key is generated by OmniRoute and shown once in the admin UI. Project Brain stores only the sidecar's key id and masked prefix (router_api_keys table) for attribution, auditing, and revocation.

Setup​

  1. Pick the public route. Default needs no DNS change: the main Caddy site serves the router's /v1 API at https://<APP_DOMAIN>/router/v1. Alternatively, set ROUTER_DOMAIN and add an A record (e.g. router.pbtest.neusis.ai) pointing at the same host for a subdomain.

  2. Env — in .env (path route shown; for the subdomain also set ROUTER_DOMAIN):

    ROUTER_PUBLIC_URL=https://pbtest.neusis.ai/router
    # ROUTER_INTERNAL_URL defaults to http://omniroute:20128
  3. Start — docker compose up -d omniroute (it is part of the default stack; recreate caddy so it picks up ROUTER_DOMAIN).

  4. Initialize in Project Brain. Sign in as a router admin (a platform super-admin or project-creator), open /admin/router, and select Initialize router. Project Brain uses the sidecar's supported cold-boot API to generate its private management password, exchange it for an admin-scoped control token, and encrypt both with the platform credential-encryption layer. Neither secret is returned to the browser.

    If the sidecar volume was initialized previously, the page asks for its existing management password once and adopts it without resetting it.

The sidecar dashboard is not part of normal setup. Port 20128 remains bound to loopback for diagnostics, and Caddy continues to return 404 for the dashboard and management API on every public hostname. ROUTER_ADMIN_TOKEN remains an optional deployment override for backwards compatibility and emergency recovery; new installations do not need to set it.

Connecting providers​

Upstream connectors are managed from /admin/router (super-admin or project-creator). /router is the member surface; router admins manage and connect agents from /admin/router instead (visiting /router as a router admin 404s). The UI deliberately never mentions OmniRoute — users just see "Router". Named connectors are offered (Google AI Studio, OpenAI API, Claude Code, ChatGPT (Codex), OpenRouter, Groq, Mistral, DeepSeek, xAI Grok, Antigravity) plus an "Other" entry where the admin types any gateway provider id directly (the gateway rejects invalid ids). Key connectors with special flows:

  • Google AI Studio — API key (AIza… from aistudio.google.com/apikey); maps to the gateway's gemini provider.
  • OpenAI API — API key (sk-…); maps to openai.
  • Claude Code — auth token: claude setup-token on a signed-in machine prints an sk-ant-oat01-… subscription token (a regular sk-ant-api… key also works). Maps to the gateway's anthropic provider, whose validator accepts the oat token inline. (claude is not a valid /api/providers id on 3.8.49 — do not "fix" the mapping without re-testing.)
  • ChatGPT (Codex) — auth token: run codex login on your own machine, then paste the contents of ~/.codex/auth.json; PB flattens the token fields and forwards to POST /api/oauth/codex/import, which validates the refresh token upstream and stores access + refresh (survives access-token expiry — cf. the litellm-proxy ChatGPT-OAuth recovery runbook this replaces).
  • Antigravity — auth token: the Google consent flow needs a loopback redirect a server can't provide, so the user runs npx omniroute@latest login antigravity on their own machine and pastes the printed omniroute-cred-v1.… blob; PB forwards it to POST /api/oauth/antigravity/paste-credentials. The @latest pin matters: a stale npx cache can serve an old helper whose blob format the gateway rejects (rm -rf ~/.npm/_npx clears it).

Default-model rule: a connector's seeded defaults contain only models that work with that connector's own credential out of the box — e.g. Antigravity seeds Gemini tiers only, because its claude-* entries need a separate Anthropic key added on the Antigravity side. Beware that the gateway's catalog is not an entitlement check: codex/gpt-5.6-sol is listed but OpenAI rejects it for ChatGPT accounts, and one such 400 triggers a ~2-minute credential cooldown that blocks every codex model. When a connector's models suddenly all fail with model_cooldown, suspect one bad model id being retried, remove it, and wait out the cooldown.

Several keys for one provider, and rotation​

Connecting a second key for a provider you already have adds an account, it does not replace the first. This is what gives the router something to rotate: one selected model can be served by any of that provider's accounts, and the gateway both spreads traffic across them and puts an account that returns a 429 into cooldown. Project Brain implements no rotation of its own — it never sits in the request path.

This only works because each apikey connection is named after its own credential mask (first8****last4). The gateway upserts apikey connections by (provider, name), so a shared name would make every new key silently overwrite the previous one. That is also why the admin-facing Name field is stored Project Brain side, in router_provider_labels, rather than sent as the connection name — a name typed twice would otherwise collapse two keys into one connection and quietly undo the rotation. The name is optional and display-only; the account list shows it above the masked key, and collapses to a dropdown past two accounts.

Credentials go straight to the sidecar's encrypted store — Project Brain never persists them — and connect/disconnect are audit-logged (platform.router_provider.connect / .disconnect) without the credential material. The Agent API keys section (including the agent endpoint URL) only appears once at least one connector is connected. Claude Code and Antigravity route subscription accounts — provider-ToS gray zone; evaluate on test deployments.

Workload models​

The Assistant chat model is chosen at /admin/ai (super admin). The Graph builder model is chosen per project, on the project's Settings tab, by any member of that project. Both pickers offer the agent-model selection below, the same curated list, one source of truth. Making a model available to Project Brain means adding it on /admin/router first. Both APIs re-check the choice server-side against router_agent_models, so a stale form cannot save a model the router will not serve.

Codex credentials: one login per router​

The ChatGPT (Codex) connector imports the refresh token from ~/.codex/auth.json. OpenAI rotates that token: every refresh returns a new one and the old one stops working. Two copies of the same auth.json therefore race, and whichever refreshes second gets "refresh token consumed" or "authentication token has been invalidated", and stays dead until someone runs codex login again and pastes the new file. This is what happens when a developer copies the auth.json that is already connected on the shared router into a local one.

The account is not the problem, the file is. One ChatGPT account can be logged in on many devices at once, and every codex login produces its own token. So a developer who wants Codex in a local deployment does this:

  1. Run codex login on the laptop (own account or the shared team account).
  2. Paste the contents of ~/.codex/auth.json into the local router's ChatGPT (Codex) connector.
  3. Move the file out of the way (mv ~/.codex/auth.json ~/.codex/auth.json.imported) so Codex CLI on the same laptop does not keep refreshing the token the router now owns. To keep using Codex CLI locally, run codex login again: that is a separate token and the two never conflict.

Never copy an auth.json from one machine or router to another. The other option is to not connect Codex locally at all and use the shared router as an upstream, see the next section.

Neusis Router upstream (temporary)​

The Neusis Router (temporary) connector on /admin/router makes the shared pbtest router an upstream of a local router. Paste an agent API key minted on https://pbtest.neusis.ai/admin/router; the local router creates an OpenAI-compatible provider node pointed at https://pbtest.neusis.ai/router/v1 (override with NEUSIS_ROUTER_UPSTREAM_URL), syncs the shared router's model list into the local catalog (a node's models only appear after that sync), and seeds the Codex and Gemini tiers it finds there as <node id>/<model>. Pick those on a project's Settings tab like any other model; the Manage picker offers everything the shared router serves. Every request then travels local router -> shared router -> provider, and the shared router's quota and rate limits apply.

Set the project's Parallel requests to 1 when the graph builder uses one of these models. The shared router admits one structurally heavy request at a time (its OMNIROUTE_CHAT_MAX_HEAVY_IN_FLIGHT is at the default); a second in-flight document chunk gets a 503, the local router then cools the connection down for about 90 seconds, and the remaining chunks fail with 429. Raising that value on the shared deployment lifts the limit.

This exists only so local deployment tests can run without their own provider credentials. It is tracked for removal as Neusis-AI-Org/hnt#57; do not build on it.

For more headroom on the shared router, connect more ChatGPT accounts, each with its own codex login. The gateway spreads traffic across them.

The token does not expire on a fixed schedule. Codex refreshes it when it has been more than about 8 days since the last refresh, and the access token lives about 10 days. A credential that sits in a router that sees traffic refreshes itself; one that sits idle for longer than that goes stale and needs a reconnect (Dashboard -> Provider -> Reconnect, or delete and re-add the connector). The gateway keeps a dead account marked active, so the admin page cannot show this; the router's logs and dashboard can.

Models for graph builds​

The graph builder sends document chunks of about token_budget tokens (60k by default) and expects up to max_output_tokens (32k) of JSON back. A 400k-token document corpus is roughly 7 to 10 such requests per full build. Any model with at least 128k of context works; what matters more is that it returns long, well-formed JSON and that the provider does not rate-limit a handful of large requests in flight.

Free options. A build of that size fits inside these free tiers if you keep parallel requests at 1 or 2 and the token budget at or below 60k. Limits as of September 2026; all are per account.

Provider (connector)ModelContextFree limits
Mistral (Mistral)mistral-large-latest, mistral-medium-latest128kabout 1B tokens/month, roughly 500k tokens/min, 1 to 2 requests/s. The most generous.
Google AI Studio (Google AI Studio)gemini-2.5-flash1M10 requests/min, 250k tokens/min, a few hundred requests/day
OpenRouter (OpenRouter)nvidia/nemotron-3-ultra-550b-a55b:free, qwen/qwen3-coder:free, google/gemma-4-31b-it:free262k to 1M20 requests/min, 50 requests/day (1,000/day once $10 of credits was ever bought)

Mistral's Experiment plan needs phone verification and no card. On Gemini the 250k tokens per minute cap is what binds, so keep parallel requests at 1 with a 60k budget or at 2 with 40k. OpenRouter's 50 requests per day covers one build of a 400k corpus, but a retry storm from truncated chunks can exhaust it; set max_output_tokens high (65536 where the model allows) so chunks do not split.

Two things that do not help on OpenRouter: extra accounts or keys, because the docs say capacity is governed per account and "making additional accounts or API keys will not affect your rate limits"; and stealth models such as stealth/ox-alpha, which are short-lived previews that disappear from the catalog without notice (it was gone on 2026-09-02). Project Brain hides openrouter/stealth/* from the model picker and refuses to add them.

The project Settings tab shows the published limits of the chosen model (context window, response limit, requests per minute and day, price) from apps/web/lib/model-limits.ts, with the date each number was checked. Extend that file when you add a model people will pick for builds.

Paid open-weight models on OpenRouter have no request cap and cost well under a dollar for a full build of that size. Prices are per million tokens, input then output, as of September 2026.

Model (OpenRouter id)ContextPrice in / outNote
deepseek/deepseek-v3.2164k$0.21 / $0.31best value, reliable JSON
qwen/qwen3-235b-a22b-2507262k$0.09 / $0.35cheapest
minimax/minimax-m2.7205k$0.24 / $0.96very large output ceiling
z-ai/glm-4.7205k$0.40 / $1.75strong on code, 131k output
moonshotai/kimi-k2.5262k$0.45 / $2.25strongest of the set

A good starting point is DeepSeek V3.2 or Qwen3 235B through the OpenRouter connector with the default token budget, max_output_tokens at 32768 (raise to 65536 on MiniMax or GLM if chunks truncate) and 3 to 5 parallel requests. The Codex subscription also works (about 272k usable input) but shares its quota with every coding agent on the router.

Selections made before this consolidation were cleared by migration 0013: their provider used the old google/anthropic/openai vocabulary and their model ids lacked the router's <provider>/ prefix, so neither could be translated.

Agent models​

Connecting a connector seeds a default set of 4-5 agent-visible models (one high-reasoning tier, medium and low/fast tiers, plus another strong all-rounder — never auto/* smart-routing ids). The admin adjusts the selection in the Agent models card on /admin/router: remove a chip or add any model the router serves for that connector. The generated agent configs list exactly this selection (router_agent_models table); disconnecting a connector's last connection clears its selection, and reconnecting re-seeds the defaults. Adds/removes are audit-logged (platform.router_model.add / .remove).

Compression​

Agent traffic through /router/v1 is compressed before it reaches upstream providers. OmniRoute ships the engines but leaves them off on a fresh install (enabled:false, defaultMode:"off"), so Project Brain turns them on during onboarding with ROUTER_COMPRESSION_SETTINGS (apps/web/lib/router-bootstrap.ts): the stacked rtk -> caveman pipeline. RTK strips command-output noise from tool results (keeping failures, errors, changed files and the tail); Caveman condenses the remaining prose. The write is best-effort — if it fails, onboarding still succeeds and the router simply runs uncompressed.

Choosing engines​

Router admins (super-admin or project-creator) pick the pipeline in the Compression card on /admin/router.

The card is generated from the gateway's own catalog (GET /api/compression/engines), which describes each engine and its editable fields — types, labels, defaults, bounds. Project Brain hardcodes no engine list and no knob list, so an image upgrade that adds engines or fields appears in the UI with no code change. apps/web/lib/router-compression.ts holds the only PB-side policy: which engines are offered, and what a fresh router starts with.

Engines always run in the gateway's own stackPriority order, not the order they were ticked — structural filtering first (session-dedup, ccr, rtk, headroom), prose compression last (relevance, caveman, aggressive, ultra). Per-engine values are written to each pipeline step's config, which upstream's buildStepOptions merges over the global settings.<engine> bag, so a step setting always wins.

Selecting no engines is a valid choice and is how compression is turned off from the UI: it writes enabled:false / defaultMode:"off" and never sends an empty stackedPipeline (the gateway normalizes an empty array back to its built-in default, which would silently re-enable what was just cleared).

Also editable: the auto-trigger threshold (autoTriggerTokens — compress only above N tokens; 0 means always), cache lifetime (cacheMinutes), and preserve system prompt. Changes are audit-logged as platform.router_compression.update.

Engines that are withheld​

Some catalog entries are deliberately not offered. The rule is derived from the catalog itself, so new experimental engines are withheld by default:

  • stable: false — upstream's experimental marker. Today: omniglyph, which renders context to PNG images for one specific model.
  • supportsPreview: false — cannot be dry-run before being enabled. Today these are exactly the engines with external requirements: llmlingua lazily downloads a ~57 MB ONNX model into DATA_DIR, and llm is inert until an operator wires a separate chat-completion backend.
  • An explicit short list: ionizer and read-lifecycle, both documented upstream as lossy and default-off.

Applying it to an existing router​

The setting lives in the sidecar's own SQLite (omniroute_data volume), which PB never writes directly, and onboarding does not re-run. For a router bootstrapped before this shipped, the /admin/router card writes it on save, or:

curl https://<app-domain>/api/admin/router/compression # settings + catalog
curl -X POST https://<app-domain>/api/admin/router/compression \
-H 'Content-Type: application/json' \
-d '{"engineIds":["rtk","caveman"]}'

Turning it down​

Compression is lossy by design. Escape hatches, in increasing order of scope:

  • Per request: send x-omniroute-compression: off (or engine:rtk) — this header outranks every other setting, but cannot turn compression on when the master switch is off.
  • Per model/endpoint: the exclusions array on the settings endpoint.
  • Gentler filtering: lower rtkConfig.intensity, or raise maxLinesPerResult / maxCharsPerResult in the card's RTK options.
  • Only compress big contexts: set the auto-trigger threshold (e.g. 32000). Worth considering if prompt-cache hit rates drop — compression rewrites history every turn, which busts provider prefix caches.
  • Deselect every engine, which is the kill switch.

A known interaction​

RTK classifies a tool result by command type and filters the whole message. A tool result that mixes formats — say a JSON array followed by test output — can be matched by the test filter, which then discards the JSON along with the passing-test noise. Measured: a payload with 25 JSON rows plus jest output compressed 1372 -> 127 tokens, and the model correctly named the failing test but reported "no JSON provided". Headroom alone on the same payload preserved both facts. RTK runs before Headroom (stackPriority 10 vs 15), so Headroom cannot rescue what RTK already dropped. If agents need mixed tool output kept intact, lower RTK's intensity or deselect it.

Verify without spending tokens via POST /api/compression/preview and POST /api/context/rtk/test; GET /api/analytics/compression reports observed savings. Live requests echo the applied plan in the X-OmniRoute-Compression: <mode>; source=<source> response header.

Minting keys for users​

Router admins (super-admin or project-creator) open /admin/router, enter a name for the key (who/what it is for), and hand the generated key to the user. Revoking a key cuts the agent off immediately. Creation and revocation are recorded in the audit log (platform.router_key.create / platform.router_key.revoke).

Connecting agents​

The Connect a coding agent button on /admin/router (and on /router for members) opens a dialog that generates a ready-to-paste config per agent: opencode, Antigravity, Codex CLI, neusis-code, and Claude Code up top, plus Cursor, Cline, Roo Code, Kilo Code, Continue, Aider, Zed, Crush, Goose, and Qwen Code under "More agents". Any other agent that accepts an OpenAI-compatible base URL + API key works the same way. Agent-specific notes:

  • neusis-code / opencode — add a custom provider with baseURL: https://<ROUTER_DOMAIN>/v1 and the minted key (or use the @omniroute/opencode-provider plugin). neusis-code's release binary is rebranded at build time, so its config lives at ~/.config/neusiscode/neusiscode.json (or neusiscode.json in the project root) — not the opencode paths.

  • Claude Code — ANTHROPIC_BASE_URL=https://<ROUTER_DOMAIN> and ANTHROPIC_AUTH_TOKEN=<key> (OmniRoute serves the Anthropic wire format at /v1/messages).

  • Google Antigravity (IDE and CLI) — both clients can override their Cloud Code backend URL, so a small local proxy (a standalone build of the antigravity-add-model proxy) injects router models into Antigravity's own model picker and translates Cloud Code ⇄ OpenAI chat completions; Google-model traffic passes through untouched. Per developer machine:

    1. Build the proxy once: clone the repo, npm install && npx tsc, copy dist/ to ~/.antigravity-model-proxy/, add a run.js that calls startProxy(), and stub electron / electron-log in its node_modules with plain-Node shims (the upstream project ships as an Electron-app patcher; only its proxy is used here). Two local patches to dist/proxy.js are required and must be reapplied after any rebuild: delete the transfer-encoding header in the buffered-response path (Electron's strict HTTP parser otherwise rejects every response, which silently breaks IDE sign-in), and aggregate SSE bodies in the non-streaming handler (the router replies with SSE even for stream: false).
    2. Save the models: the panel-generated config goes to ~/.gemini/antigravity/custom_models.json (hot-reloaded by the proxy).
    3. Start the proxy before launching Antigravity: cd ~/.antigravity-model-proxy && node run.js — it listens on 127.0.0.1:50999.
    4. IDE: set "jetski.cloudCodeUrl": "http://127.0.0.1:50999" in the Antigravity IDE settings.json (survives IDE auto-updates).
    5. CLI: launch with CLOUD_CODE_URL=http://127.0.0.1:50999 agy — alias it as agyx and open the CLI that way; plain agy keeps talking to Google directly and will not show the models.

    Models appear as custom-<name> slugs (agyx models to verify). The Cloud Code override is an undocumented hook in both clients and may change in future releases. Antigravity can also be an upstream of the router through subscription OAuth, which carries provider-ToS risk — evaluate that connector on test deployments only.

Keeping agent configs current​

The base URL and the API key never change, so a developer pastes those once. The model list is the part that goes stale: an admin adding a connector on /admin/router does nothing for a config generated last month. How much that matters depends on the agent:

The model list is…AgentsA stale config means
load-bearing (the file enumerates models)opencode, neusis-code, Zed, Crush, Continue, Antigravity, Cursornew models are invisible; removed ones still offered, then 404 at the gateway
documentation only (the id is passed at run time)Claude Code, Codex, Aider, Goose, Qwen Codeonly the default model is stale; any current id still works

Two mechanisms keep it current, both reading router_agent_models live, so there is nothing to redeploy and nothing to re-paste.

Curated /v1/models​

Project Brain answers GET <router base>/v1/models itself. Caddy routes that one exact path to web:3000; every other /v1 request still goes straight to the sidecar, so PB stays out of the chat request path.

This exists because the gateway's own listing is not a usable model picker. Measured on a test deployment with four connectors attached: 1,220 models in 38 s, against 21 curated ones served instantly from Postgres. It also lists models the admin deliberately did not publish.

Agents that discover models — Cline, Roo Code, Kilo Code, and anything else using an "OpenAI Compatible" provider — therefore track /admin/router with no user action at all.

Two consequences worth knowing:

  • The gateway's raw /v1/models is no longer reachable publicly. PB's own calls are unaffected: listRouterModels() goes to omniroute:20128 through ROUTER_INTERNAL_URL and never passes through Caddy.
  • If web is down, model discovery fails while chat keeps working. That is the intended trade.

In local development there is no Caddy, so ROUTER_PUBLIC_URL points at the sidecar directly and this path is only reachable at the app's own origin (/api/router/v1/models).

One-command refresh​

For everything that needs a file, GET <router base>/config/<agent> renders it live, authenticated by the developer's own router key — no session, so it works from a laptop shell. The Keep it up to date block on /router (and /admin/router) shows both lines, with the key filled into the first:

# once, in your shell profile
export NEUSIS_ROUTER_KEY='sk-…'

# whenever an admin publishes new models
curl -fsS -K - "https://<app-domain>/router/config/claude-code" \
--create-dirs -o ~/.neusis/router-claude-code.sh <<EOF
header = "Authorization: Bearer $NEUSIS_ROUTER_KEY"
EOF

Two exposures are avoided here, and both are easy to reintroduce. The key comes from the environment, so the line a developer re-runs does not accumulate in shell history. And the header goes to curl on stdin rather than as an argument: -H "Authorization: Bearer $VAR" reads as safe, but the shell expands it before exec, leaving the key in curl's argv for the whole request where any local user running ps can read it. The heredoc delimiter is unquoted on purpose; <<'EOF' would stop the expansion and send curl the literal variable name.

The bearer token is the key that ends up in the config, so nothing is revealed from the sidecar and no new secret is minted.

Where the file lands depends on who owns it:

  • Router-owned paths are overwritten in place. ~/.neusis/router-<agent>.sh for the env-var agents (Claude Code, Goose, Qwen Code, Aider) — source it from your shell profile — and ~/.gemini/antigravity/custom_models.json, which the Antigravity proxy hot-reloads, so a refresh lands without restarting anything.
  • An agent's own config file is never clobbered. ~/.config/opencode/opencode.json, ~/.codex/config.toml, ~/.continue/config.yaml, ~/.config/zed/settings.json and ~/.config/crush/crush.json hold the developer's other providers, so the command writes a .neusis sibling to be merged by hand. buildAgentConfigFile marks this with exclusive: false; the UI reads that flag rather than hardcoding a list.

Cursor, Cline, Roo Code and Kilo Code have no file — they are configured in a settings UI — so the endpoint answers 409 and points at /v1/models instead.

Codex is the one agent whose file cannot carry the key: it reads <NAME>_API_KEY from the environment, so that export stays a shell-profile line the refresh does not manage.

Keeping the model list current​

Providers change their catalogs constantly. OpenCode Zen rotates the models it marks as free, OpenRouter adds and drops :free variants, and stealth previews disappear without notice. Today the published list (router_agent_models) is written when a connector is connected and by hand afterwards; nothing re-checks it against the gateway catalog. A model the provider removed stays published until it 404s, and a new free model is invisible until an admin finds it. The generated agent configs for opencode, neusis-code, Zed, Crush, Continue and Antigravity enumerate every model id, so any change means regenerating the file on every developer machine.

This is a solved problem elsewhere. Four patterns cover it:

  • Stable aliases the gateway resolves. LiteLLM model group aliases with fallback chains, OpenRouter's openrouter/auto and openrouter/free, Vercel AI Gateway aliases, and OmniRoute's combos and auto/best-* routers. The client names the alias once; the operator changes what it resolves to.
  • Discovery at run time. The OpenAI-compatible GET /v1/models is the source of truth. Cline, Roo Code, Kilo Code and Cursor query it, and Project Brain already serves a curated per-key version (see "Curated /v1/models" above), so those tools are current with no action. opencode does not query it natively, but the opencode-models-discovery plugin does for providers marked dynamic: true, with a /reload-models command.
  • Wildcard pass-through. LiteLLM provider/* routes any id the provider serves without listing it. OmniRoute behaves the same: anything in a connector's catalog is callable, and the published list is the only filter.
  • Scheduled sync plus shared metadata. Gateways re-pull provider lists on a timer (OmniRoute every 24 hours), and catalogs such as models.dev and OpenRouter's models API carry pricing, context and free flags, so nobody maintains a table by hand.

Not built yet. Recorded here so the implementation follows one design.

  1. Aliases first. Project Brain owns five combos on the gateway and publishes them like models: neusis/coding, neusis/fast, neusis/free, neusis/graph-builder and neusis/assistant. Each is an ordered fallback list an admin edits in Router. Generated configs list the aliases first and the concrete ids after, so a file generated today keeps working when the backing model changes. Project Settings and /admin/ai can pick an alias. neusis/free is recomputed from whatever is free right now.
  2. Hourly catalog reconciliation as a router-catalog-tick job in the worker scheduler, next to sync-tick. It pulls the gateway catalog with pricing, context and free flags, then: flags published rows that vanished (unavailable_since on router_agent_models; hidden from /v1/models, configs and pickers, shown as "no longer served" in Router, deleted after 30 days); publishes new free models on connectors whose policy is "free", which is the default for OpenCode Zen and OpenRouter (a per-connector setting in Router); rewrites alias membership; records a catalog snapshot that replaces the hand-maintained apps/web/lib/model-limits.ts for pricing and context; and writes one audit entry and one admin banner per change set.
  3. Configs. Keep the curated /v1/models and the refresh command from "One-command refresh" above. Document the opencode discovery plugin as the zero-maintenance path for opencode users. Treat auto/* and openrouter/openrouter/free as a distinct class in isHiddenModel (apps/web/lib/router-gateway.ts) so a smart router cannot be published by accident next to the aliases.

An implementation touches: router_agent_models (new unavailable_since), a router_connector_settings table for the per-connector publish policy, the scheduler job, combo management in router-gateway.ts, alias-first output in router-agent-config.ts, and the retirement of model-limits.ts.

Notes and caveats​

  • The compose service pins diegosouzapw/omniroute:3.8.49. The -web image variant (~300 MB larger, bundles Chromium) is required only for web-cookie subscription providers (chatgpt-web, claude-web, gemini-web); switch the image tag if you need those.
  • Port 20128 is bound to loopback only; all public traffic goes through Caddy.
  • OmniRoute keeps its SQLite DB, settings, and AES-256-GCM-encrypted provider credentials in the omniroute_data volume — include it in any backup plan if you configure providers you care about.
  • Swapping the gateway implementation later (e.g. LiteLLM's /key/generate) only requires reimplementing apps/web/lib/router-gateway.ts and the compose service — the UI, API routes, and DB ledger are gateway-agnostic. One extra obligation now comes with that: Caddy serves /v1/models from Project Brain, so a replacement gateway must keep tolerating a client that never calls its own catalog endpoint.
  • router_api_keys.key_hash holds a sha256 of each minted key so the key-authenticated endpoints recognise a caller in one indexed lookup instead of a sidecar round trip. Keys minted before that column existed are backfilled on first use from the gateway's reveal API (ALLOW_API_KEY_REVEAL), once per key; revoked keys are left alone.
  • The graphify sidecar reaches its model through the router, so omniroute is in that container's NO_PROXY and joined to the graphify_internal network. It has to be exempt from the egress proxy: squid denies private-IP destinations, which is what a sibling container is. The consequence is that GRAPHIFY_EGRESS_ALLOWED_HOSTS no longer bounds graph-builder LLM traffic — the connector list in /admin/router does, one hop later. The allowlist still governs the sidecar's other outbound requests, and the proxy refuses to start on an empty list.