LLM Router sidecar (OmniRoute)
Project Brain runs the open-source OmniRoute LLM gateway as a sidecar container. It is the LLM path for the deployment:
- Coding agents (neusis-code, Claude Code, Cursor, Cline, Aider, …) point at the router's OpenAI/Anthropic-compatible endpoint with an API key minted from /admin/router.
- Project Brain's own calls — assistant chat and the graph builder — go through
the same gateway, using a key it mints for itself. There are no provider API
keys in
.envand none in Postgres; the admin picks the assistant model at /admin/ai, each project picks its graph builder model on its Settings tab, and the credential lives only in the router's encrypted store. - The existing external litellm-proxy deployment keeps working — nothing here replaces or reconfigures it.
Architecture
neusis-code / Claude Code (developer laptop)
│ Bearer <key minted in /admin/router>
▼
https://APP_DOMAIN/router/v1 ──► Caddy ──► omniroute:20128 (sidecar)
(or https://ROUTER_DOMAIN/v1) ▲
│ provider API keys / OAuth,
PB web / worker ──────────────────────────────┤ stored encrypted in its volume
Bearer <internal key> /v1/chat/completions │
▼
OpenAI / Anthropic / Google / …
PB web ──► omniroute:20128 /api/keys (mint/revoke/reveal keys,
Authorization: Bearer <encrypted token>) ledger in router_api_keys)
Project Brain's own key
Onboarding mints one extra API key named Project Brain (internal) and stores
it encrypted on router_gateway_credentials.internal_key_enc. It is
deliberately not in the router_api_keys ledger: that table drives the
admin key list and member assignment, and this key belongs to the platform
rather than a person. Routers onboarded before this existed get one on the next
visit to /admin/router or /admin/ai.
resolveAiSelection() and resolveGraphBuilderSelection()
(packages/db/src/ai-selection.ts) pair that key with the assistant's model or
the project's graph builder model and hand both to the caller. There is no environment
fallback: if the router is not onboarded or no model is selected, the assistant
answers 503 and a graph run is recorded as llm_not_configured, each naming the
page that fixes it. That is deliberate — a silent fallback to a directly-held
provider key is exactly what this consolidation removes.
One gateway quirk is worth knowing: a chat-completions request whose body omits
stream is answered with SSE, not JSON. apps/web/lib/llm.ts fills in
stream: false so both streamText and generateText work. Without it,
non-streaming calls fail with "Invalid JSON response".
The raw agent key is generated by OmniRoute and shown once in the admin UI.
Project Brain stores only the sidecar's key id and masked prefix
(router_api_keys table) for attribution, auditing, and revocation.
Setup
-
Pick the public route. Default needs no DNS change: the main Caddy site serves the router's /v1 API at
https://<APP_DOMAIN>/router/v1. Alternatively, setROUTER_DOMAINand add an A record (e.g.router.pbtest.neusis.ai) pointing at the same host for a subdomain. -
Env — in
.env(path route shown; for the subdomain also setROUTER_DOMAIN):ROUTER_PUBLIC_URL=https://pbtest.neusis.ai/router# ROUTER_INTERNAL_URL defaults to http://omniroute:20128 -
Start —
docker compose up -d omniroute(it is part of the default stack; recreatecaddyso it picks upROUTER_DOMAIN). -
Initialize in Project Brain. Sign in as a router admin (a platform super-admin or project-creator), open /admin/router, and select Initialize router. Project Brain uses the sidecar's supported cold-boot API to generate its private management password, exchange it for an admin-scoped control token, and encrypt both with the platform credential-encryption layer. Neither secret is returned to the browser.
If the sidecar volume was initialized previously, the page asks for its existing management password once and adopts it without resetting it.
The sidecar dashboard is not part of normal setup. Port 20128 remains bound to
loopback for diagnostics, and Caddy continues to return 404 for the dashboard
and management API on every public hostname. ROUTER_ADMIN_TOKEN remains an
optional deployment override for backwards compatibility and emergency
recovery; new installations do not need to set it.
Connecting providers
Upstream connectors are managed from /admin/router (super-admin or project-creator). /router is the member surface; router admins manage and connect agents from /admin/router instead (visiting /router as a router admin 404s). The UI deliberately never mentions OmniRoute — users just see "Router". Named connectors are offered (Google AI Studio, OpenAI API, Claude Code, ChatGPT (Codex), OpenRouter, Groq, Mistral, DeepSeek, xAI Grok, Antigravity) plus an "Other" entry where the admin types any gateway provider id directly (the gateway rejects invalid ids). Key connectors with special flows:
- Google AI Studio — API key (
AIza…from aistudio.google.com/apikey); maps to the gateway'sgeminiprovider. - OpenAI API — API key (
sk-…); maps toopenai. - Claude Code — auth token:
claude setup-tokenon a signed-in machine prints ansk-ant-oat01-…subscription token (a regularsk-ant-api…key also works). Maps to the gateway'santhropicprovider, whose validator accepts the oat token inline. (claudeis not a valid/api/providersid on 3.8.49 — do not "fix" the mapping without re-testing.) - ChatGPT (Codex) — auth token: run
codex loginon your own machine, then paste the contents of~/.codex/auth.json; PB flattens the token fields and forwards toPOST /api/oauth/codex/import, which validates the refresh token upstream and stores access + refresh (survives access-token expiry — cf. the litellm-proxy ChatGPT-OAuth recovery runbook this replaces). - Antigravity — auth token: the Google consent flow needs a loopback
redirect a server can't provide, so the user runs
npx omniroute@latest login antigravityon their own machine and pastes the printedomniroute-cred-v1.…blob; PB forwards it toPOST /api/oauth/antigravity/paste-credentials. The@latestpin matters: a stale npx cache can serve an old helper whose blob format the gateway rejects (rm -rf ~/.npm/_npxclears it).
Default-model rule: a connector's seeded defaults contain only models that
work with that connector's own credential out of the box — e.g. Antigravity
seeds Gemini tiers only, because its claude-* entries need a separate
Anthropic key added on the Antigravity side. Beware that the gateway's
catalog is not an entitlement check: codex/gpt-5.6-sol is listed but
OpenAI rejects it for ChatGPT accounts, and one such 400 triggers a ~2-minute
credential cooldown that blocks every codex model. When a connector's models
suddenly all fail with model_cooldown, suspect one bad model id being
retried, remove it, and wait out the cooldown.
Several keys for one provider, and rotation
Connecting a second key for a provider you already have adds an account, it does not replace the first. This is what gives the router something to rotate: one selected model can be served by any of that provider's accounts, and the gateway both spreads traffic across them and puts an account that returns a 429 into cooldown. Project Brain implements no rotation of its own — it never sits in the request path.
This only works because each apikey connection is named after its own credential
mask (first8****last4). The gateway upserts apikey connections by
(provider, name), so a shared name would make every new key silently overwrite
the previous one. That is also why the admin-facing Name field is stored
Project Brain side, in router_provider_labels, rather than sent as the
connection name — a name typed twice would otherwise collapse two keys into one
connection and quietly undo the rotation. The name is optional and display-only;
the account list shows it above the masked key, and collapses to a dropdown past
two accounts.
Credentials go straight to the sidecar's encrypted store — Project Brain
never persists them — and connect/disconnect are audit-logged
(platform.router_provider.connect / .disconnect) without the credential
material. The Agent API keys section (including the agent endpoint URL)
only appears once at least one connector is connected. Claude Code and
Antigravity route subscription accounts — provider-ToS gray zone; evaluate on
test deployments.
Workload models
The Assistant chat model is chosen at /admin/ai (super admin). The
Graph builder model is chosen per project, on the project's Settings tab, by
any member of that project. Both pickers offer the agent-model selection below,
the same curated list, one source of truth. Making a model available to Project
Brain means adding it on /admin/router first. Both APIs re-check the choice
server-side against router_agent_models, so a stale form cannot save a model
the router will not serve.
Codex credentials: one login per router
The ChatGPT (Codex) connector imports the refresh token from ~/.codex/auth.json.
OpenAI rotates that token: every refresh returns a new one and the old one stops
working. Two copies of the same auth.json therefore race, and whichever
refreshes second gets "refresh token consumed" or "authentication token has
been invalidated", and stays dead until someone runs codex login again and
pastes the new file. This is what happens when a developer copies the
auth.json that is already connected on the shared router into a local one.
The account is not the problem, the file is. One ChatGPT account can be logged
in on many devices at once, and every codex login produces its own token.
So a developer who wants Codex in a local deployment does this:
- Run
codex loginon the laptop (own account or the shared team account). - Paste the contents of
~/.codex/auth.jsoninto the local router's ChatGPT (Codex) connector. - Move the file out of the way (
mv ~/.codex/auth.json ~/.codex/auth.json.imported) so Codex CLI on the same laptop does not keep refreshing the token the router now owns. To keep using Codex CLI locally, runcodex loginagain: that is a separate token and the two never conflict.
Never copy an auth.json from one machine or router to another. The other
option is to not connect Codex locally at all and use the shared router as an
upstream, see the next section.
Neusis Router upstream (temporary)
The Neusis Router (temporary) connector on /admin/router makes the shared
pbtest router an upstream of a local router. Paste an agent API key minted on
https://pbtest.neusis.ai/admin/router; the local router creates an
OpenAI-compatible provider node pointed at https://pbtest.neusis.ai/router/v1
(override with NEUSIS_ROUTER_UPSTREAM_URL), syncs the shared router's model
list into the local catalog (a node's models only appear after that sync), and
seeds the Codex and Gemini tiers it finds there as <node id>/<model>. Pick
those on a project's Settings tab like any other model; the Manage picker
offers everything the shared router serves. Every request then travels local router -> shared
router -> provider, and the shared router's quota and rate limits apply.
Set the project's Parallel requests to 1 when the graph builder uses one
of these models. The shared router admits one structurally heavy request at a
time (its OMNIROUTE_CHAT_MAX_HEAVY_IN_FLIGHT is at the default); a second
in-flight document chunk gets a 503, the local router then cools the
connection down for about 90 seconds, and the remaining chunks fail with 429.
Raising that value on the shared deployment lifts the limit.
This exists only so local deployment tests can run without their own provider credentials. It is tracked for removal as Neusis-AI-Org/hnt#57; do not build on it.
For more headroom on the shared router, connect more ChatGPT accounts, each with
its own codex login. The gateway spreads traffic across them.
The token does not expire on a fixed schedule. Codex refreshes it when it has been more than about 8 days since the last refresh, and the access token lives about 10 days. A credential that sits in a router that sees traffic refreshes itself; one that sits idle for longer than that goes stale and needs a reconnect (Dashboard -> Provider -> Reconnect, or delete and re-add the connector). The gateway keeps a dead account marked active, so the admin page cannot show this; the router's logs and dashboard can.
Models for graph builds
The graph builder sends document chunks of about token_budget tokens (60k by
default) and expects up to max_output_tokens (32k) of JSON back. A 400k-token
document corpus is roughly 7 to 10 such requests per full build. Any model with
at least 128k of context works; what matters more is that it returns long,
well-formed JSON and that the provider does not rate-limit a handful of large
requests in flight.
Free options. A build of that size fits inside these free tiers if you keep parallel requests at 1 or 2 and the token budget at or below 60k. Limits as of September 2026; all are per account.
| Provider (connector) | Model | Context | Free limits |
|---|---|---|---|
| Mistral (Mistral) | mistral-large-latest, mistral-medium-latest | 128k | about 1B tokens/month, roughly 500k tokens/min, 1 to 2 requests/s. The most generous. |
| Google AI Studio (Google AI Studio) | gemini-2.5-flash | 1M | 10 requests/min, 250k tokens/min, a few hundred requests/day |
| OpenRouter (OpenRouter) | nvidia/nemotron-3-ultra-550b-a55b:free, qwen/qwen3-coder:free, google/gemma-4-31b-it:free | 262k to 1M | 20 requests/min, 50 requests/day (1,000/day once $10 of credits was ever bought) |
Mistral's Experiment plan needs phone verification and no card. On Gemini the
250k tokens per minute cap is what binds, so keep parallel requests at 1 with a
60k budget or at 2 with 40k. OpenRouter's 50 requests per day covers one build
of a 400k corpus, but a retry storm from truncated chunks can exhaust it; set
max_output_tokens high (65536 where the model allows) so chunks do not split.
Two things that do not help on OpenRouter: extra accounts or keys, because the
docs say capacity is governed per account and "making additional accounts or
API keys will not affect your rate limits"; and stealth models such as
stealth/ox-alpha, which are short-lived previews that disappear from the
catalog without notice (it was gone on 2026-09-02). Project Brain hides
openrouter/stealth/* from the model picker and refuses to add them.
The project Settings tab shows the published limits of the chosen model
(context window, response limit, requests per minute and day, price) from
apps/web/lib/model-limits.ts, with the date each number was checked. Extend
that file when you add a model people will pick for builds.
Paid open-weight models on OpenRouter have no request cap and cost well under a dollar for a full build of that size. Prices are per million tokens, input then output, as of September 2026.
| Model (OpenRouter id) | Context | Price in / out | Note |
|---|---|---|---|
deepseek/deepseek-v3.2 | 164k | $0.21 / $0.31 | best value, reliable JSON |
qwen/qwen3-235b-a22b-2507 | 262k | $0.09 / $0.35 | cheapest |
minimax/minimax-m2.7 | 205k | $0.24 / $0.96 | very large output ceiling |
z-ai/glm-4.7 | 205k | $0.40 / $1.75 | strong on code, 131k output |
moonshotai/kimi-k2.5 | 262k | $0.45 / $2.25 | strongest of the set |
A good starting point is DeepSeek V3.2 or Qwen3 235B through the OpenRouter
connector with the default token budget, max_output_tokens at 32768 (raise to
65536 on MiniMax or GLM if chunks truncate) and 3 to 5 parallel requests. The
Codex subscription also works (about 272k usable input) but shares its quota
with every coding agent on the router.
Selections made before this consolidation were cleared by migration 0013:
their provider used the old google/anthropic/openai vocabulary and their model
ids lacked the router's <provider>/ prefix, so neither could be translated.
Agent models
Connecting a connector seeds a default set of 4-5 agent-visible models (one
high-reasoning tier, medium and low/fast tiers, plus another strong
all-rounder — never auto/* smart-routing ids). The admin adjusts the
selection in the Agent models card on /admin/router: remove a chip or add
any model the router serves for that connector. The generated agent configs
list exactly this selection (router_agent_models table); disconnecting a
connector's last connection clears its selection, and reconnecting re-seeds
the defaults. Adds/removes are audit-logged (platform.router_model.add /
.remove).
Compression
Agent traffic through /router/v1 is compressed before it reaches upstream
providers. OmniRoute ships the engines but leaves them off on a fresh install
(enabled:false, defaultMode:"off"), so Project Brain turns them on during
onboarding with ROUTER_COMPRESSION_SETTINGS
(apps/web/lib/router-bootstrap.ts): the stacked rtk -> caveman pipeline.
RTK strips command-output noise from tool results (keeping failures, errors,
changed files and the tail); Caveman condenses the remaining prose. The write
is best-effort — if it fails, onboarding still succeeds and the router simply
runs uncompressed.
Choosing engines
Router admins (super-admin or project-creator) pick the pipeline in the Compression card on /admin/router.
The card is generated from the gateway's own catalog
(GET /api/compression/engines), which describes each engine and its
editable fields — types, labels, defaults, bounds. Project Brain hardcodes no
engine list and no knob list, so an image upgrade that adds engines or fields
appears in the UI with no code change. apps/web/lib/router-compression.ts
holds the only PB-side policy: which engines are offered, and what a fresh
router starts with.
Engines always run in the gateway's own stackPriority order, not the order
they were ticked — structural filtering first (session-dedup, ccr, rtk,
headroom), prose compression last (relevance, caveman, aggressive, ultra).
Per-engine values are written to each pipeline step's config, which
upstream's buildStepOptions merges over the global settings.<engine> bag,
so a step setting always wins.
Selecting no engines is a valid choice and is how compression is turned
off from the UI: it writes enabled:false / defaultMode:"off" and never
sends an empty stackedPipeline (the gateway normalizes an empty array back
to its built-in default, which would silently re-enable what was just
cleared).
Also editable: the auto-trigger threshold (autoTriggerTokens — compress
only above N tokens; 0 means always), cache lifetime (cacheMinutes), and
preserve system prompt. Changes are audit-logged as
platform.router_compression.update.
Engines that are withheld
Some catalog entries are deliberately not offered. The rule is derived from the catalog itself, so new experimental engines are withheld by default:
stable: false— upstream's experimental marker. Today:omniglyph, which renders context to PNG images for one specific model.supportsPreview: false— cannot be dry-run before being enabled. Today these are exactly the engines with external requirements:llmlingualazily downloads a ~57 MB ONNX model intoDATA_DIR, andllmis inert until an operator wires a separate chat-completion backend.- An explicit short list:
ionizerandread-lifecycle, both documented upstream as lossy and default-off.
Applying it to an existing router
The setting lives in the sidecar's own SQLite (omniroute_data volume), which
PB never writes directly, and onboarding does not re-run. For a router
bootstrapped before this shipped, the /admin/router card writes it on save, or:
curl https://<app-domain>/api/admin/router/compression # settings + catalog
curl -X POST https://<app-domain>/api/admin/router/compression \
-H 'Content-Type: application/json' \
-d '{"engineIds":["rtk","caveman"]}'
Turning it down
Compression is lossy by design. Escape hatches, in increasing order of scope:
- Per request: send
x-omniroute-compression: off(orengine:rtk) — this header outranks every other setting, but cannot turn compression on when the master switch is off. - Per model/endpoint: the
exclusionsarray on the settings endpoint. - Gentler filtering: lower
rtkConfig.intensity, or raisemaxLinesPerResult/maxCharsPerResultin the card's RTK options. - Only compress big contexts: set the auto-trigger threshold (e.g. 32000). Worth considering if prompt-cache hit rates drop — compression rewrites history every turn, which busts provider prefix caches.
- Deselect every engine, which is the kill switch.
A known interaction
RTK classifies a tool result by command type and filters the whole message.
A tool result that mixes formats — say a JSON array followed by test output —
can be matched by the test filter, which then discards the JSON along with the
passing-test noise. Measured: a payload with 25 JSON rows plus jest output
compressed 1372 -> 127 tokens, and the model correctly named the failing test
but reported "no JSON provided". Headroom alone on the same payload preserved
both facts. RTK runs before Headroom (stackPriority 10 vs 15), so Headroom
cannot rescue what RTK already dropped. If agents need mixed tool output kept
intact, lower RTK's intensity or deselect it.
Verify without spending tokens via POST /api/compression/preview and
POST /api/context/rtk/test; GET /api/analytics/compression reports
observed savings. Live requests echo the applied plan in the
X-OmniRoute-Compression: <mode>; source=<source> response header.
Minting keys for users
Router admins (super-admin or project-creator) open /admin/router, enter
a name for the key (who/what it is
for), and hand the generated key to the user. Revoking a key cuts the agent
off immediately. Creation and revocation are recorded in the audit log
(platform.router_key.create / platform.router_key.revoke).
Connecting agents
The Connect a coding agent button on /admin/router (and on /router for members) opens a dialog that generates a ready-to-paste config per agent: opencode, Antigravity, Codex CLI, neusis-code, and Claude Code up top, plus Cursor, Cline, Roo Code, Kilo Code, Continue, Aider, Zed, Crush, Goose, and Qwen Code under "More agents". Any other agent that accepts an OpenAI-compatible base URL + API key works the same way. Agent-specific notes:
-
neusis-code / opencode — add a custom provider with
baseURL: https://<ROUTER_DOMAIN>/v1and the minted key (or use the@omniroute/opencode-providerplugin). neusis-code's release binary is rebranded at build time, so its config lives at~/.config/neusiscode/neusiscode.json(orneusiscode.jsonin the project root) — not the opencode paths. -
Claude Code —
ANTHROPIC_BASE_URL=https://<ROUTER_DOMAIN>andANTHROPIC_AUTH_TOKEN=<key>(OmniRoute serves the Anthropic wire format at/v1/messages). -
Google Antigravity (IDE and CLI) — both clients can override their Cloud Code backend URL, so a small local proxy (a standalone build of the antigravity-add-model proxy) injects router models into Antigravity's own model picker and translates Cloud Code ⇄ OpenAI chat completions; Google-model traffic passes through untouched. Per developer machine:
- Build the proxy once: clone the repo,
npm install && npx tsc, copydist/to~/.antigravity-model-proxy/, add arun.jsthat callsstartProxy(), and stubelectron/electron-login itsnode_moduleswith plain-Node shims (the upstream project ships as an Electron-app patcher; only its proxy is used here). Two local patches todist/proxy.jsare required and must be reapplied after any rebuild: delete thetransfer-encodingheader in the buffered-response path (Electron's strict HTTP parser otherwise rejects every response, which silently breaks IDE sign-in), and aggregate SSE bodies in the non-streaming handler (the router replies with SSE even forstream: false). - Save the models: the panel-generated config goes to
~/.gemini/antigravity/custom_models.json(hot-reloaded by the proxy). - Start the proxy before launching Antigravity:
cd ~/.antigravity-model-proxy && node run.js— it listens on127.0.0.1:50999. - IDE: set
"jetski.cloudCodeUrl": "http://127.0.0.1:50999"in the Antigravity IDEsettings.json(survives IDE auto-updates). - CLI: launch with
CLOUD_CODE_URL=http://127.0.0.1:50999 agy— alias it asagyxand open the CLI that way; plainagykeeps talking to Google directly and will not show the models.
Models appear as
custom-<name>slugs (agyx modelsto verify). The Cloud Code override is an undocumented hook in both clients and may change in future releases. Antigravity can also be an upstream of the router through subscription OAuth, which carries provider-ToS risk — evaluate that connector on test deployments only. - Build the proxy once: clone the repo,
Keeping agent configs current
The base URL and the API key never change, so a developer pastes those once. The model list is the part that goes stale: an admin adding a connector on /admin/router does nothing for a config generated last month. How much that matters depends on the agent:
| The model list is… | Agents | A stale config means |
|---|---|---|
| load-bearing (the file enumerates models) | opencode, neusis-code, Zed, Crush, Continue, Antigravity, Cursor | new models are invisible; removed ones still offered, then 404 at the gateway |
| documentation only (the id is passed at run time) | Claude Code, Codex, Aider, Goose, Qwen Code | only the default model is stale; any current id still works |
Two mechanisms keep it current, both reading router_agent_models live, so
there is nothing to redeploy and nothing to re-paste.
Curated /v1/models
Project Brain answers GET <router base>/v1/models itself. Caddy routes that
one exact path to web:3000; every other /v1 request still goes straight to
the sidecar, so PB stays out of the chat request path.
This exists because the gateway's own listing is not a usable model picker. Measured on a test deployment with four connectors attached: 1,220 models in 38 s, against 21 curated ones served instantly from Postgres. It also lists models the admin deliberately did not publish.
Agents that discover models — Cline, Roo Code, Kilo Code, and anything else using an "OpenAI Compatible" provider — therefore track /admin/router with no user action at all.
Two consequences worth knowing:
- The gateway's raw
/v1/modelsis no longer reachable publicly. PB's own calls are unaffected:listRouterModels()goes toomniroute:20128throughROUTER_INTERNAL_URLand never passes through Caddy. - If
webis down, model discovery fails while chat keeps working. That is the intended trade.
In local development there is no Caddy, so ROUTER_PUBLIC_URL points at the
sidecar directly and this path is only reachable at the app's own origin
(/api/router/v1/models).
One-command refresh
For everything that needs a file, GET <router base>/config/<agent> renders it
live, authenticated by the developer's own router key — no session, so it works
from a laptop shell. The Keep it up to date block on /router (and
/admin/router) shows both lines, with the key filled into the first:
# once, in your shell profile
export NEUSIS_ROUTER_KEY='sk-…'
# whenever an admin publishes new models
curl -fsS -K - "https://<app-domain>/router/config/claude-code" \
--create-dirs -o ~/.neusis/router-claude-code.sh <<EOF
header = "Authorization: Bearer $NEUSIS_ROUTER_KEY"
EOF
Two exposures are avoided here, and both are easy to reintroduce. The key comes
from the environment, so the line a developer re-runs does not accumulate in
shell history. And the header goes to curl on stdin rather than as an argument:
-H "Authorization: Bearer $VAR" reads as safe, but the shell expands it before
exec, leaving the key in curl's argv for the whole request where any local
user running ps can read it. The heredoc delimiter is unquoted on purpose;
<<'EOF' would stop the expansion and send curl the literal variable name.
The bearer token is the key that ends up in the config, so nothing is revealed from the sidecar and no new secret is minted.
Where the file lands depends on who owns it:
- Router-owned paths are overwritten in place.
~/.neusis/router-<agent>.shfor the env-var agents (Claude Code, Goose, Qwen Code, Aider) — source it from your shell profile — and~/.gemini/antigravity/custom_models.json, which the Antigravity proxy hot-reloads, so a refresh lands without restarting anything. - An agent's own config file is never clobbered.
~/.config/opencode/opencode.json,~/.codex/config.toml,~/.continue/config.yaml,~/.config/zed/settings.jsonand~/.config/crush/crush.jsonhold the developer's other providers, so the command writes a.neusissibling to be merged by hand.buildAgentConfigFilemarks this withexclusive: false; the UI reads that flag rather than hardcoding a list.
Cursor, Cline, Roo Code and Kilo Code have no file — they are configured in a
settings UI — so the endpoint answers 409 and points at /v1/models instead.
Codex is the one agent whose file cannot carry the key: it reads
<NAME>_API_KEY from the environment, so that export stays a shell-profile
line the refresh does not manage.
Keeping the model list current
Providers change their catalogs constantly. OpenCode Zen rotates the models it
marks as free, OpenRouter adds and drops :free variants, and stealth previews
disappear without notice. Today the published list (router_agent_models) is
written when a connector is connected and by hand afterwards; nothing re-checks
it against the gateway catalog. A model the provider removed stays published
until it 404s, and a new free model is invisible until an admin finds it. The
generated agent configs for opencode, neusis-code, Zed, Crush, Continue and
Antigravity enumerate every model id, so any change means regenerating the
file on every developer machine.
This is a solved problem elsewhere. Four patterns cover it:
- Stable aliases the gateway resolves. LiteLLM model group aliases with
fallback chains, OpenRouter's
openrouter/autoandopenrouter/free, Vercel AI Gateway aliases, and OmniRoute's combos andauto/best-*routers. The client names the alias once; the operator changes what it resolves to. - Discovery at run time. The OpenAI-compatible
GET /v1/modelsis the source of truth. Cline, Roo Code, Kilo Code and Cursor query it, and Project Brain already serves a curated per-key version (see "Curated /v1/models" above), so those tools are current with no action. opencode does not query it natively, but theopencode-models-discoveryplugin does for providers markeddynamic: true, with a/reload-modelscommand. - Wildcard pass-through. LiteLLM
provider/*routes any id the provider serves without listing it. OmniRoute behaves the same: anything in a connector's catalog is callable, and the published list is the only filter. - Scheduled sync plus shared metadata. Gateways re-pull provider lists on a timer (OmniRoute every 24 hours), and catalogs such as models.dev and OpenRouter's models API carry pricing, context and free flags, so nobody maintains a table by hand.
Recommended for Project Brain
Not built yet. Recorded here so the implementation follows one design.
- Aliases first. Project Brain owns five combos on the gateway and
publishes them like models:
neusis/coding,neusis/fast,neusis/free,neusis/graph-builderandneusis/assistant. Each is an ordered fallback list an admin edits in Router. Generated configs list the aliases first and the concrete ids after, so a file generated today keeps working when the backing model changes. Project Settings and /admin/ai can pick an alias.neusis/freeis recomputed from whatever is free right now. - Hourly catalog reconciliation as a
router-catalog-tickjob in the worker scheduler, next tosync-tick. It pulls the gateway catalog with pricing, context and free flags, then: flags published rows that vanished (unavailable_sinceonrouter_agent_models; hidden from/v1/models, configs and pickers, shown as "no longer served" in Router, deleted after 30 days); publishes new free models on connectors whose policy is "free", which is the default for OpenCode Zen and OpenRouter (a per-connector setting in Router); rewrites alias membership; records a catalog snapshot that replaces the hand-maintainedapps/web/lib/model-limits.tsfor pricing and context; and writes one audit entry and one admin banner per change set. - Configs. Keep the curated
/v1/modelsand the refresh command from "One-command refresh" above. Document the opencode discovery plugin as the zero-maintenance path for opencode users. Treatauto/*andopenrouter/openrouter/freeas a distinct class inisHiddenModel(apps/web/lib/router-gateway.ts) so a smart router cannot be published by accident next to the aliases.
An implementation touches: router_agent_models (new unavailable_since), a
router_connector_settings table for the per-connector publish policy, the
scheduler job, combo management in router-gateway.ts, alias-first output in
router-agent-config.ts, and the retirement of model-limits.ts.
Notes and caveats
- The compose service pins
diegosouzapw/omniroute:3.8.49. The-webimage variant (~300 MB larger, bundles Chromium) is required only for web-cookie subscription providers (chatgpt-web,claude-web,gemini-web); switch the image tag if you need those. - Port 20128 is bound to loopback only; all public traffic goes through Caddy.
- OmniRoute keeps its SQLite DB, settings, and AES-256-GCM-encrypted provider
credentials in the
omniroute_datavolume — include it in any backup plan if you configure providers you care about. - Swapping the gateway implementation later (e.g. LiteLLM's
/key/generate) only requires reimplementingapps/web/lib/router-gateway.tsand the compose service — the UI, API routes, and DB ledger are gateway-agnostic. One extra obligation now comes with that: Caddy serves/v1/modelsfrom Project Brain, so a replacement gateway must keep tolerating a client that never calls its own catalog endpoint. router_api_keys.key_hashholds a sha256 of each minted key so the key-authenticated endpoints recognise a caller in one indexed lookup instead of a sidecar round trip. Keys minted before that column existed are backfilled on first use from the gateway's reveal API (ALLOW_API_KEY_REVEAL), once per key; revoked keys are left alone.- The graphify sidecar reaches its model through the router, so
omnirouteis in that container'sNO_PROXYand joined to thegraphify_internalnetwork. It has to be exempt from the egress proxy: squid denies private-IP destinations, which is what a sibling container is. The consequence is thatGRAPHIFY_EGRESS_ALLOWED_HOSTSno longer bounds graph-builder LLM traffic — the connector list in /admin/router does, one hop later. The allowlist still governs the sidecar's other outbound requests, and the proxy refuses to start on an empty list.