Graph pipeline (Graphify)
How ingested content becomes the knowledge graph, and the settings that govern it. The security boundary around the sidecar is in docs/security.md; build triggers, artifacts, and egress setup are in docs/INSTALL.md. This document covers the platform settings, the environment knobs, and the failure modes that recur.
Project settings
Every project has its own graph settings on its Settings tab (Project ->
Settings). Any member of the project can view and change them. They are stored
as one row per project in project_graphify_settings and take effect on the
next sync or graph build, nothing needs restarting. A project that has never
saved the form runs on the defaults below.
The page has four boxes:
- Graph builder model: the connector and model the document pass calls. The list is the models an admin selected in Router. "No model" is allowed and is fine for code-only builds.
- Graph scope: code only or code and documents, plus the mirrored-code switch.
- Model calls: how the document pass calls the model. Hidden when the scope is code only.
- Sync approvals: what a sync does with deletions.
| Setting | Column | Default |
|---|---|---|
| Graph builder connector | model_provider | null, meaning no model |
| Graph builder model | model_id | null, meaning no model |
| Graph scope | code_only | false, meaning code and documents |
| Build from mirrored code only | code_connector_only | false |
| Tokens per request | token_budget | 60000 |
| Parallel requests | max_concurrency | 3 |
| Max tokens per response | max_output_tokens | 32768 |
| Request timeout | api_timeout_s | 300 |
| Deep extraction | deep_extraction | false |
| Deletions | auto_accept_proposals | false, meaning hold for review |
| Max deletions per sync run | max_deletions_per_run | 25 |
Nothing on this table falls back to an environment variable. An operator
editing .env and a member editing the form would otherwise disagree about
what a build is doing.
Changes go through POST /api/projects/<slug>/settings. Any project member may
call it; a non-member gets a 404. Changes are audit-logged as
project.settings.update. There is no GET; the form is server-rendered.
The assistant's model is still platform-wide, at Admin -> Assistant model.
Upgrading from a release with platform-wide Graphify settings copies the old
values and the old graph builder model into every existing project (migration
0019), so nothing changes until someone edits a project.
Changing scope and rebuilding
Three settings change what the graph contains: the graph scope, the
mirrored-code switch, and deep extraction. Every build records the settings it
ran with on its graph_runs row. When one of those three differs from the last
successful build, the project header shows Settings changed since last
build with a Rebuild button, and the Settings page shows a Rebuild now
button after saving. Only authors (admins, editors, contributors) can rebuild;
viewers see the message without the button.
That rebuild is a fresh build: it wipes every graphify output (graph.json, cache, report, wiki) before extracting, and removes from the repo any file the old build committed that the new one did not write. A normal or force-full rebuild is not enough here because graphify merges into the existing graph, so switching from code and documents to code only would keep every document node. A fresh build costs the same as a first build.
The worker applies this on every trigger, not just the button: if a sync's automatic rebuild finds the settings changed, it promotes itself to a fresh build. So the prompt never clears with the old nodes still in the graph.
The numeric knobs (tokens per request, parallel requests, max tokens per response, request timeout) and the model choice never show the prompt. They change cost and speed, not what the graph contains, and simply apply on the next build.
Graph scope (code_only)
Graphify routes files purely by extension. CODE_EXTENSIONS (.py, .c,
.json, …) go through tree-sitter at zero tokens; DOC_EXTENSIONS (.md,
.mdx, .qmd, .skill, .txt, .rst, .html, .yaml, .yml) go to the
graph-builder model as "semantic" files.
- Code and documents (the default) sends doc files to the graph-builder model, which costs tokens on every rebuild. Source files still use tree-sitter.
- Code only runs
graphify extract --code-only. Zero model calls, so no model needs to be configured at all.
Because the default needs a model, a project that has not picked one on its
Settings tab will see graph builds fail with llm_not_configured rather than
quietly producing a code-only graph. Switch the scope to code only to build
without a model.
The clustering pass names one community per model call, 286 calls on a
mid-sized code-and-documents graph, so its budget
(GRAPHIFY_CLUSTER_ONLY_TIMEOUT_S, default 900 seconds) is minutes, not
seconds. If that pass fails or runs out of time, the sidecar runs it again
with --no-label so GRAPH_REPORT.md and graph.html always exist; without
them the UI reads the project as unbuilt, and a fresh rebuild would have
deleted the previous copies.
Code-only mode also passes --no-label to the clustering pass. That flag is not
cosmetic: graphify cluster-only regenerates the report and community names,
and its community naming auto-detects an LLM backend and calls graphify's
hardcoded gpt-4.1-mini default. The platform's graph_builder selection only
flows into extract, never into cluster-only, so without --no-label an
"AST-only" build still makes model calls. Deterministic hub-based names take
their place.
Model call tuning
Five settings apply only when the scope is code and documents; an AST-only
build makes no model calls for them to govern, so the panel hides them. All of
them are passed straight through to graphify extract.
- Tokens per request (
--token-budget, default 60000, upstream's own). Files packed into one request. Higher means fewer, larger calls. Past roughly 2.5 times the output ceiling, dense chunks start overrunning it and getting split and retried, which costs more than it saves. The form derives its warning from that ratio rather than hardcoding a number, so raising the ceiling raises the point at which the budget is flagged; against the default 32768 ceiling it lands on the documented 80000. - Parallel requests (
--max-concurrency, default 3, where upstream uses 4). Capped at 10, and the ceiling is the router's rather than any provider's. OmniRoute admitsOMNIROUTE_CHAT_MAX_HEAVY_IN_FLIGHT(10) "heavy" bodies at once, meaning over 256KB or roughly 32k estimated tokens, and answers the rest withchat_admission_busy. At the default token budget every semantic chunk is heavy, so the router binds before the provider does and a 503 there reads exactly like a provider rate limit. Raise the compose value and this cap together, or neither. - Max tokens per response (
GRAPHIFY_MAX_OUTPUT_TOKENS, default 32768, above graphify's own 8k/16k backend defaults). The ceiling graphify sends the provider asmax_completion_tokens. A response that reaches it comes back withfinish_reason: "length", truncated mid-JSON and unparseable, so graphify discards it and bisects the chunk. The generated tokens are billed and wasted, which makes this the direct lever on that cost. The real ceiling is the selected model's own output maximum: ask for more than it allows and the provider answers 400 rather than a longer response. Unlike the other three this is an environment variable to graphify, not a flag, so the sidecar assigns it into the subprocess environment. - Request timeout (
--api-timeout, default 300, below upstream's 600 on purpose). Graphify bisects on context-window errors, so a tight per-call timeout no longer aborts a run and fails a wedged request twice as fast. Raise it when a slow gateway sits in front of the provider. - Deep extraction (
--mode deep, off by default). Looks for relationships the text implies but never states. A richer graph for noticeably more tokens.
Values are bounds-checked twice, in the project settings API and again in the sidecar, because those are separate trust boundaries and a bad value wedges every build.
Switching back to code only does not remove documents already in the graph
on an ordinary rebuild: --code-only skips the doc files rather than deleting
their nodes (verified on a forced rebuild: a 4,257-node graph stayed at 4,257).
That is why the scope change triggers the fresh rebuild described above, which
starts from an empty graph.
Build from mirrored code only (code_connector_only)
Restricts a build's input to the repositories the GitHub Code connector mirrors
into raw/code/{owner}/{repo}/. Everything else that has been ingested, Jira,
Confluence, synthesized documents, is left out of the build. Nothing is
deleted; clearing the setting restores the full graph on the next rebuild.
This is orthogonal to Graph scope. code_only filters by file extension
within whatever the build can see, while code_connector_only decides what it
can see at all. Setting both gives an AST-only graph over mirrored source and nothing
else.
It is implemented by widening the .graphifyignore the worker writes before
every build (buildGraphifyIgnore in
apps/worker/src/workers/graphify/prepareGraphifyWorkspace.ts), so there is no
exclude-glob field to maintain and no per-connector configuration. The excludes
are enumerated from the working tree rather than written as a gitignore
negation (/* plus !raw/code/): the sidecar's ignore parser is upstream code
we do not control, and every pattern the base list already relies on is a plain
path, so enumerating stays inside proven behaviour.
"If available" is literal. When nothing has been mirrored yet, the build falls back to the full file set rather than producing an empty graph. Turning the setting on before connecting the GitHub Code connector therefore changes nothing.
Deletions and the per-run cap
These were one checkbox, which welded together two unrelated questions: who approves a deletion, and how many a single run may stage. They are now separate settings, so a project can have unreviewed deletions and still bound the blast radius of one run.
Deletions decides the approval gate. Creates and updates are always
bulk-approved, because they are idempotent under checksum dedup and review adds
nothing. On hold for review (the default) a run that stages deletions waits at
staged until someone approves or rejects them, while its creates and updates
commit immediately. On accept automatically deletions commit in the same
writer job as everything else. That suits mirrored content whose source of
truth is upstream, where a prune is as mechanical as the create before it.
Max deletions per sync run caps how many delete proposals one run stages.
The No limit checkbox on the Settings tab stores it as 0. It applies in
both modes. Purging 433 stale
documents at 25 a run is about 18 cycles, so someone doing a large one-off
prune raises it or clears it, independently of whether review is on.
The cap is a rate limit, not a safety check. Protection against a catastrophic false positive, an upstream API returning empty and the code concluding "everything is gone", comes from guards neither setting can switch off:
- the empty seen-set guard, treating an empty result as an upstream hiccup rather than a legitimately empty scope;
- the
truncatedcheck in the caller, so a truncated sync never reconciles; - the ambiguity guard, skipping projects with several integrations of one type, where the path prefix cannot say which owns a file;
- explicitly listed items, so anything in
include_itemsorexclude_itemsis never deleted.
Note the cap is applied as min(n, cap) and drains. An earlier version staged
nothing when the count exceeded the cap, so the count never fell and a project
over the limit could never prune anything. That was observed as 66 orphaned Jira
tickets re-detected and re-skipped every 15 minutes for weeks.
One asymmetry worth knowing: the settings change is audit-logged, but
auto-accepted proposals are recorded only as reviewedBy: null plus worker log
lines.
How a build runs
worker: read the project's settings, compare with the last successful run
├─ promote to a fresh build if a graph-shaping setting changed
worker: clone/fetch the KB repo, scrub the authenticated remote
├─ skip if the source SHA is unchanged → 'source_unchanged'
│ (never when the build is fresh)
├─ skip if nothing but scaffolding is present → 'no_graph_inputs'
├─ write .graphifyignore (widened when code_connector_only is on)
├─ chown -R 10001:10001 (the read-only sidecar must write graphify-out/)
├─ record the settings on the graph_runs row
└─ POST /run to graphify-sidecar:8080
[fresh: rm -rf graphify-out/] [full: rm -rf graphify-out/cache/]
graphify extract [--code-only] [--backend/--model]
graphify cluster-only [--no-label] (falls back to --no-label if it fails)
graphify export wiki
worker: collect graphify-out/, read stats, commit through the writer
└─ fresh: also delete outputs HEAD still tracks that this build did not write
The sidecar receives a workspace path and model settings, never GitHub credentials. The worker scrubs the authenticated git remote and verifies the scrub before handing over; a failed verification fails the run.
.graphifyignore always removes Project Brain's own scaffolding (.git/,
graphify-out/, .brain/, synthesized/, the seed README.md and
agent-guide.md). The same list decides whether a checkout has any graph input
at all.
When no model is configured
resolveGraphBuilderSelection(settings) reads the model from the project's
settings and has no environment fallback by design: a deployment whose router
is not onboarded fails loudly rather than quietly calling a provider directly
with a different key.
- Code and documents: the run is marked
failedwitherrorMessage: llm_not_configured, visible in the UI, not a crashed job. - Code only: the failure is logged as a warning and the build continues.
There is no model to miss, extraction is tree-sitter and clustering runs with
--no-label.
The provider is always reported as openai because the router speaks the OpenAI
wire format regardless of which upstream serves the request; the real provider
is the prefix of the model id. That is why the gateway's provider id never
reaches the sidecar, and why a missing openrouter entry in the sidecar's
PROVIDER_TO_BACKEND map is not a bug.
Environment variables
Graph scope, token budget, concurrency, request timeout and extraction depth
are project settings, not environment variables; the sections above cover them.
The GRAPHIFY_TOKEN_BUDGET, GRAPHIFY_MAX_CONCURRENCY and
GRAPHIFY_API_TIMEOUT_S variables survive only as the sidecar's fallback for a
worker image older than those fields, and GRAPHIFY_CODE_ONLY likewise.
.env.example documents GRAPHIFY_SIDECAR_TOKEN, GRAPHIFY_TIMEOUT_S,
GRAPHIFY_MAX_CONCURRENCY, GRAPHIFY_TOKEN_BUDGET,
GRAPHIFY_EGRESS_ALLOWED_HOSTS, GRAPHIFY_EGRESS_PROXY, and
GRAPHIFY_AUTO_REBUILD_DELAY_MS. The rest live in compose or have code
defaults:
| Variable | Default | Notes |
|---|---|---|
GRAPHIFY_OLLAMA_PARALLEL | 1 | Graphify forces its ollama backend serial as a single-GPU guard. Our "ollama" endpoint is the router, not a GPU, so parallel chunks are safe, without this, --max-concurrency is silently ignored. |
GRAPHIFY_MAX_OUTPUT_TOKENS | 32768 | Raises graphify's per-chunk output cap from its 8k/16k default; dense pages otherwise trip a split-and-retry cascade. |
GRAPHIFY_CLUSTER_ONLY_TIMEOUT_S / GRAPHIFY_WIKI_TIMEOUT_S | , | Post-step budgets; see INSTALL.md. |
GRAPHIFY_WORKSPACE_DIR / GRAPHIFY_WORKSPACE_ROOT | /workspace | Shared volume between worker and sidecar. |
GRAPHIFY_SIDECAR_UID / GID | 10001 | Must match the sidecar image's user. |
GRAPHIFY_TIMEOUT_S has two derived ceilings, the worker's job timeout and the
sidecar HTTP timeout are both computed from it. Raising one without the other
just moves which timeout fires first.
OMNIROUTE_CHAT_MAX_HEAVY_IN_FLIGHT (on the router service) should be sized to
GRAPHIFY_MAX_CONCURRENCY. Its default of 1 silently serialized graphify's
parallel semantic chunks.
Operational notes
Upgrade ordering. The sidecar validates request bodies against a strict allowlist, so an unknown field is a 400. Rebuild the sidecar image before shipping a worker that sends a new field.
Aborting a build. The sidecar keeps running its graphify extract
subprocess even when the worker-side job is killed. Restart the sidecar
container too, or the zombie finishes and may keep making model calls.
Three distinct OOM signatures, all of which look like "stuck at Reporting":
- Redis killed. Worker logs flood with
ENOTFOUND redis/ECONNREFUSED, anddocker inspect -f '{{.RestartCount}}' project-brain-redis-1climbs. Writer jobs carry whole file bodies through Redis, so large mirror syncs are the trigger. - Worker Node heap. Look for
FATAL ERROR: Reached heap limit.tsx watchrespawns the process, so the container still looks healthy; only the log tail shows it. - Sidecar Python. The run reports
errorMessage = sidecar_status_500, with its log ending cleanly right after "AST extraction: 100%" and no error text. Confirm withoom_kill > 0in the sidecar cgroup'smemory.events. Memory scales with corpus size.
BullMQ traps. A graphify-auto-<projectId> job left in any state,
failed included, silently blocks every future auto-rebuild: jobId dedupe
ignores adds whose id still exists. And because the graphify worker runs at
concurrency 1, a single orphaned entry in the active list wedges the queue
indefinitely with no error and no progress. Prefer draining the queue before
restarting the worker.
Large artifacts. graph.json can exceed GitHub's ~40MB createBlob ceiling
(base64 adds ~33%); files over MAX_FILE_BYTES (30MB) are skipped from the
commit with a warning rather than failing the writer job. This is why the
build-committed marker is graphify-out/graph.html, which always fits, and why
build stats are read from the workspace copy. Above 32MB the stats come from
GRAPH_REPORT.md counts instead of parsing graph.json, and the per-community
fingerprint that drives the UI's community delta is absent.
Egress. GRAPHIFY_EGRESS_ALLOWED_HOSTS no longer bounds graph-builder model
traffic, the sidecar reaches its model through the router, which is exempt from
the proxy, and the router makes the upstream call one hop later. The connector
list on /admin/router is what bounds that. The allowlist still governs every
other outbound request the sidecar makes, and the proxy refuses to start on an
empty list. After a hard VM restart squid can crash-loop on a stale pidfile;
docker compose up -d --force-recreate graphify-egress-proxy clears it.
Open items
- Rebuilds skip
graph.jsonwhen it exceeds the commit cap, so a committed copy goes stale one build behind. Either add a git-push fallback for over-cap artifacts, or drop the committedgraph.jsonto avoid the confusion. packages/staging'sauto_rulesengine remains unwired with zero call sites; the per-project auto-accept boolean supersedes it for this use.- Migration
0019stamps each project's last successful run with the settings the deployment was using at upgrade time (backfilled: truein the snapshot), so the "settings changed" prompt works right after the upgrade. Older runs keepsettings = null.