Skip to main content

Graph pipeline (Graphify)

How ingested content becomes the knowledge graph, and the settings that govern it. The security boundary around the sidecar is in docs/security.md; build triggers, artifacts, and egress setup are in docs/INSTALL.md. This document covers the platform settings, the environment knobs, and the failure modes that recur.

Project settings​

Every project has its own graph settings on its Settings tab (Project -> Settings). Any member of the project can view and change them. They are stored as one row per project in project_graphify_settings and take effect on the next sync or graph build, nothing needs restarting. A project that has never saved the form runs on the defaults below.

The page has four boxes:

  • Graph builder model: the connector and model the document pass calls. The list is the models an admin selected in Router. "No model" is allowed and is fine for code-only builds.
  • Graph scope: code only or code and documents, plus the mirrored-code switch.
  • Model calls: how the document pass calls the model. Hidden when the scope is code only.
  • Sync approvals: what a sync does with deletions.
SettingColumnDefault
Graph builder connectormodel_providernull, meaning no model
Graph builder modelmodel_idnull, meaning no model
Graph scopecode_onlyfalse, meaning code and documents
Build from mirrored code onlycode_connector_onlyfalse
Tokens per requesttoken_budget60000
Parallel requestsmax_concurrency3
Max tokens per responsemax_output_tokens32768
Request timeoutapi_timeout_s300
Deep extractiondeep_extractionfalse
Deletionsauto_accept_proposalsfalse, meaning hold for review
Max deletions per sync runmax_deletions_per_run25

Nothing on this table falls back to an environment variable. An operator editing .env and a member editing the form would otherwise disagree about what a build is doing.

Changes go through POST /api/projects/<slug>/settings. Any project member may call it; a non-member gets a 404. Changes are audit-logged as project.settings.update. There is no GET; the form is server-rendered.

The assistant's model is still platform-wide, at Admin -> Assistant model. Upgrading from a release with platform-wide Graphify settings copies the old values and the old graph builder model into every existing project (migration 0019), so nothing changes until someone edits a project.

Changing scope and rebuilding​

Three settings change what the graph contains: the graph scope, the mirrored-code switch, and deep extraction. Every build records the settings it ran with on its graph_runs row. When one of those three differs from the last successful build, the project header shows Settings changed since last build with a Rebuild button, and the Settings page shows a Rebuild now button after saving. Only authors (admins, editors, contributors) can rebuild; viewers see the message without the button.

That rebuild is a fresh build: it wipes every graphify output (graph.json, cache, report, wiki) before extracting, and removes from the repo any file the old build committed that the new one did not write. A normal or force-full rebuild is not enough here because graphify merges into the existing graph, so switching from code and documents to code only would keep every document node. A fresh build costs the same as a first build.

The worker applies this on every trigger, not just the button: if a sync's automatic rebuild finds the settings changed, it promotes itself to a fresh build. So the prompt never clears with the old nodes still in the graph.

The numeric knobs (tokens per request, parallel requests, max tokens per response, request timeout) and the model choice never show the prompt. They change cost and speed, not what the graph contains, and simply apply on the next build.

Graph scope (code_only)​

Graphify routes files purely by extension. CODE_EXTENSIONS (.py, .c, .json, …) go through tree-sitter at zero tokens; DOC_EXTENSIONS (.md, .mdx, .qmd, .skill, .txt, .rst, .html, .yaml, .yml) go to the graph-builder model as "semantic" files.

  • Code and documents (the default) sends doc files to the graph-builder model, which costs tokens on every rebuild. Source files still use tree-sitter.
  • Code only runs graphify extract --code-only. Zero model calls, so no model needs to be configured at all.

Because the default needs a model, a project that has not picked one on its Settings tab will see graph builds fail with llm_not_configured rather than quietly producing a code-only graph. Switch the scope to code only to build without a model.

The clustering pass names one community per model call, 286 calls on a mid-sized code-and-documents graph, so its budget (GRAPHIFY_CLUSTER_ONLY_TIMEOUT_S, default 900 seconds) is minutes, not seconds. If that pass fails or runs out of time, the sidecar runs it again with --no-label so GRAPH_REPORT.md and graph.html always exist; without them the UI reads the project as unbuilt, and a fresh rebuild would have deleted the previous copies.

Code-only mode also passes --no-label to the clustering pass. That flag is not cosmetic: graphify cluster-only regenerates the report and community names, and its community naming auto-detects an LLM backend and calls graphify's hardcoded gpt-4.1-mini default. The platform's graph_builder selection only flows into extract, never into cluster-only, so without --no-label an "AST-only" build still makes model calls. Deterministic hub-based names take their place.

Model call tuning​

Five settings apply only when the scope is code and documents; an AST-only build makes no model calls for them to govern, so the panel hides them. All of them are passed straight through to graphify extract.

  • Tokens per request (--token-budget, default 60000, upstream's own). Files packed into one request. Higher means fewer, larger calls. Past roughly 2.5 times the output ceiling, dense chunks start overrunning it and getting split and retried, which costs more than it saves. The form derives its warning from that ratio rather than hardcoding a number, so raising the ceiling raises the point at which the budget is flagged; against the default 32768 ceiling it lands on the documented 80000.
  • Parallel requests (--max-concurrency, default 3, where upstream uses 4). Capped at 10, and the ceiling is the router's rather than any provider's. OmniRoute admits OMNIROUTE_CHAT_MAX_HEAVY_IN_FLIGHT (10) "heavy" bodies at once, meaning over 256KB or roughly 32k estimated tokens, and answers the rest with chat_admission_busy. At the default token budget every semantic chunk is heavy, so the router binds before the provider does and a 503 there reads exactly like a provider rate limit. Raise the compose value and this cap together, or neither.
  • Max tokens per response (GRAPHIFY_MAX_OUTPUT_TOKENS, default 32768, above graphify's own 8k/16k backend defaults). The ceiling graphify sends the provider as max_completion_tokens. A response that reaches it comes back with finish_reason: "length", truncated mid-JSON and unparseable, so graphify discards it and bisects the chunk. The generated tokens are billed and wasted, which makes this the direct lever on that cost. The real ceiling is the selected model's own output maximum: ask for more than it allows and the provider answers 400 rather than a longer response. Unlike the other three this is an environment variable to graphify, not a flag, so the sidecar assigns it into the subprocess environment.
  • Request timeout (--api-timeout, default 300, below upstream's 600 on purpose). Graphify bisects on context-window errors, so a tight per-call timeout no longer aborts a run and fails a wedged request twice as fast. Raise it when a slow gateway sits in front of the provider.
  • Deep extraction (--mode deep, off by default). Looks for relationships the text implies but never states. A richer graph for noticeably more tokens.

Values are bounds-checked twice, in the project settings API and again in the sidecar, because those are separate trust boundaries and a bad value wedges every build.

Switching back to code only does not remove documents already in the graph on an ordinary rebuild: --code-only skips the doc files rather than deleting their nodes (verified on a forced rebuild: a 4,257-node graph stayed at 4,257). That is why the scope change triggers the fresh rebuild described above, which starts from an empty graph.

Build from mirrored code only (code_connector_only)​

Restricts a build's input to the repositories the GitHub Code connector mirrors into raw/code/{owner}/{repo}/. Everything else that has been ingested, Jira, Confluence, synthesized documents, is left out of the build. Nothing is deleted; clearing the setting restores the full graph on the next rebuild.

This is orthogonal to Graph scope. code_only filters by file extension within whatever the build can see, while code_connector_only decides what it can see at all. Setting both gives an AST-only graph over mirrored source and nothing else.

It is implemented by widening the .graphifyignore the worker writes before every build (buildGraphifyIgnore in apps/worker/src/workers/graphify/prepareGraphifyWorkspace.ts), so there is no exclude-glob field to maintain and no per-connector configuration. The excludes are enumerated from the working tree rather than written as a gitignore negation (/* plus !raw/code/): the sidecar's ignore parser is upstream code we do not control, and every pattern the base list already relies on is a plain path, so enumerating stays inside proven behaviour.

"If available" is literal. When nothing has been mirrored yet, the build falls back to the full file set rather than producing an empty graph. Turning the setting on before connecting the GitHub Code connector therefore changes nothing.

Deletions and the per-run cap​

These were one checkbox, which welded together two unrelated questions: who approves a deletion, and how many a single run may stage. They are now separate settings, so a project can have unreviewed deletions and still bound the blast radius of one run.

Deletions decides the approval gate. Creates and updates are always bulk-approved, because they are idempotent under checksum dedup and review adds nothing. On hold for review (the default) a run that stages deletions waits at staged until someone approves or rejects them, while its creates and updates commit immediately. On accept automatically deletions commit in the same writer job as everything else. That suits mirrored content whose source of truth is upstream, where a prune is as mechanical as the create before it.

Max deletions per sync run caps how many delete proposals one run stages. The No limit checkbox on the Settings tab stores it as 0. It applies in both modes. Purging 433 stale documents at 25 a run is about 18 cycles, so someone doing a large one-off prune raises it or clears it, independently of whether review is on.

The cap is a rate limit, not a safety check. Protection against a catastrophic false positive, an upstream API returning empty and the code concluding "everything is gone", comes from guards neither setting can switch off:

  • the empty seen-set guard, treating an empty result as an upstream hiccup rather than a legitimately empty scope;
  • the truncated check in the caller, so a truncated sync never reconciles;
  • the ambiguity guard, skipping projects with several integrations of one type, where the path prefix cannot say which owns a file;
  • explicitly listed items, so anything in include_items or exclude_items is never deleted.

Note the cap is applied as min(n, cap) and drains. An earlier version staged nothing when the count exceeded the cap, so the count never fell and a project over the limit could never prune anything. That was observed as 66 orphaned Jira tickets re-detected and re-skipped every 15 minutes for weeks.

One asymmetry worth knowing: the settings change is audit-logged, but auto-accepted proposals are recorded only as reviewedBy: null plus worker log lines.

How a build runs​

worker: read the project's settings, compare with the last successful run
├─ promote to a fresh build if a graph-shaping setting changed
worker: clone/fetch the KB repo, scrub the authenticated remote
├─ skip if the source SHA is unchanged → 'source_unchanged'
│ (never when the build is fresh)
├─ skip if nothing but scaffolding is present → 'no_graph_inputs'
├─ write .graphifyignore (widened when code_connector_only is on)
├─ chown -R 10001:10001 (the read-only sidecar must write graphify-out/)
├─ record the settings on the graph_runs row
└─ POST /run to graphify-sidecar:8080
[fresh: rm -rf graphify-out/] [full: rm -rf graphify-out/cache/]
graphify extract [--code-only] [--backend/--model]
graphify cluster-only [--no-label] (falls back to --no-label if it fails)
graphify export wiki
worker: collect graphify-out/, read stats, commit through the writer
└─ fresh: also delete outputs HEAD still tracks that this build did not write

The sidecar receives a workspace path and model settings, never GitHub credentials. The worker scrubs the authenticated git remote and verifies the scrub before handing over; a failed verification fails the run.

.graphifyignore always removes Project Brain's own scaffolding (.git/, graphify-out/, .brain/, synthesized/, the seed README.md and agent-guide.md). The same list decides whether a checkout has any graph input at all.

When no model is configured​

resolveGraphBuilderSelection(settings) reads the model from the project's settings and has no environment fallback by design: a deployment whose router is not onboarded fails loudly rather than quietly calling a provider directly with a different key.

  • Code and documents: the run is marked failed with errorMessage: llm_not_configured, visible in the UI, not a crashed job.
  • Code only: the failure is logged as a warning and the build continues. There is no model to miss, extraction is tree-sitter and clustering runs with --no-label.

The provider is always reported as openai because the router speaks the OpenAI wire format regardless of which upstream serves the request; the real provider is the prefix of the model id. That is why the gateway's provider id never reaches the sidecar, and why a missing openrouter entry in the sidecar's PROVIDER_TO_BACKEND map is not a bug.

Environment variables​

Graph scope, token budget, concurrency, request timeout and extraction depth are project settings, not environment variables; the sections above cover them. The GRAPHIFY_TOKEN_BUDGET, GRAPHIFY_MAX_CONCURRENCY and GRAPHIFY_API_TIMEOUT_S variables survive only as the sidecar's fallback for a worker image older than those fields, and GRAPHIFY_CODE_ONLY likewise.

.env.example documents GRAPHIFY_SIDECAR_TOKEN, GRAPHIFY_TIMEOUT_S, GRAPHIFY_MAX_CONCURRENCY, GRAPHIFY_TOKEN_BUDGET, GRAPHIFY_EGRESS_ALLOWED_HOSTS, GRAPHIFY_EGRESS_PROXY, and GRAPHIFY_AUTO_REBUILD_DELAY_MS. The rest live in compose or have code defaults:

VariableDefaultNotes
GRAPHIFY_OLLAMA_PARALLEL1Graphify forces its ollama backend serial as a single-GPU guard. Our "ollama" endpoint is the router, not a GPU, so parallel chunks are safe, without this, --max-concurrency is silently ignored.
GRAPHIFY_MAX_OUTPUT_TOKENS32768Raises graphify's per-chunk output cap from its 8k/16k default; dense pages otherwise trip a split-and-retry cascade.
GRAPHIFY_CLUSTER_ONLY_TIMEOUT_S / GRAPHIFY_WIKI_TIMEOUT_S,Post-step budgets; see INSTALL.md.
GRAPHIFY_WORKSPACE_DIR / GRAPHIFY_WORKSPACE_ROOT/workspaceShared volume between worker and sidecar.
GRAPHIFY_SIDECAR_UID / GID10001Must match the sidecar image's user.

GRAPHIFY_TIMEOUT_S has two derived ceilings, the worker's job timeout and the sidecar HTTP timeout are both computed from it. Raising one without the other just moves which timeout fires first.

OMNIROUTE_CHAT_MAX_HEAVY_IN_FLIGHT (on the router service) should be sized to GRAPHIFY_MAX_CONCURRENCY. Its default of 1 silently serialized graphify's parallel semantic chunks.

Operational notes​

Upgrade ordering. The sidecar validates request bodies against a strict allowlist, so an unknown field is a 400. Rebuild the sidecar image before shipping a worker that sends a new field.

Aborting a build. The sidecar keeps running its graphify extract subprocess even when the worker-side job is killed. Restart the sidecar container too, or the zombie finishes and may keep making model calls.

Three distinct OOM signatures, all of which look like "stuck at Reporting":

  • Redis killed. Worker logs flood with ENOTFOUND redis / ECONNREFUSED, and docker inspect -f '{{.RestartCount}}' project-brain-redis-1 climbs. Writer jobs carry whole file bodies through Redis, so large mirror syncs are the trigger.
  • Worker Node heap. Look for FATAL ERROR: Reached heap limit. tsx watch respawns the process, so the container still looks healthy; only the log tail shows it.
  • Sidecar Python. The run reports errorMessage = sidecar_status_500, with its log ending cleanly right after "AST extraction: 100%" and no error text. Confirm with oom_kill > 0 in the sidecar cgroup's memory.events. Memory scales with corpus size.

BullMQ traps. A graphify-auto-<projectId> job left in any state, failed included, silently blocks every future auto-rebuild: jobId dedupe ignores adds whose id still exists. And because the graphify worker runs at concurrency 1, a single orphaned entry in the active list wedges the queue indefinitely with no error and no progress. Prefer draining the queue before restarting the worker.

Large artifacts. graph.json can exceed GitHub's ~40MB createBlob ceiling (base64 adds ~33%); files over MAX_FILE_BYTES (30MB) are skipped from the commit with a warning rather than failing the writer job. This is why the build-committed marker is graphify-out/graph.html, which always fits, and why build stats are read from the workspace copy. Above 32MB the stats come from GRAPH_REPORT.md counts instead of parsing graph.json, and the per-community fingerprint that drives the UI's community delta is absent.

Egress. GRAPHIFY_EGRESS_ALLOWED_HOSTS no longer bounds graph-builder model traffic, the sidecar reaches its model through the router, which is exempt from the proxy, and the router makes the upstream call one hop later. The connector list on /admin/router is what bounds that. The allowlist still governs every other outbound request the sidecar makes, and the proxy refuses to start on an empty list. After a hard VM restart squid can crash-loop on a stale pidfile; docker compose up -d --force-recreate graphify-egress-proxy clears it.

Open items​

  • Rebuilds skip graph.json when it exceeds the commit cap, so a committed copy goes stale one build behind. Either add a git-push fallback for over-cap artifacts, or drop the committed graph.json to avoid the confusion.
  • packages/staging's auto_rules engine remains unwired with zero call sites; the per-project auto-accept boolean supersedes it for this use.
  • Migration 0019 stamps each project's last successful run with the settings the deployment was using at upgrade time (backfilled: true in the snapshot), so the "settings changed" prompt works right after the upgrade. Older runs keep settings = null.