Graph pipeline findings — 2026-08-26 debugging session
Working notes from getting a zero-LLM, code-AST-only graph build working for
projectbrain-astropy-swe-bench-lite (astropy mirror, 1,345 code files), and from
the infrastructure failures found along the way. Each section lists what happened,
what was changed (⚙️ = change already applied on this branch / environment), and
what is still open (⬜ = decision or follow-up needed).
Verified end state: full rebuild in ~2 min, 27,551 nodes / 57,784 edges, 100%
from code files, zero LLM calls (confirmed at OmniRoute, the only egress),
output committed including the 41MB graph.json.
1. Zero-LLM (AST-only) graph builds
How graphify decides what hits the LLM: routing is purely by file extension
(graphify/detect.py in the sidecar image). CODE_EXTENSIONS (.py, .c, .json, …)
go through tree-sitter at zero tokens; DOC_EXTENSIONS
(.md .mdx .qmd .skill .txt .rst .html .yaml .yml) go to the LLM as "semantic"
files.
- ⚙️
GRAPHIFY_CODE_ONLY=1(already in.env, passed through compose) makes the sidecar rungraphify extract --code-only: semantic pass skipped entirely, no API key needed. - ✅ Productised (2026-08-27). This is now Admin → LLM Settings → Graphify →
Graph scope, stored in
platform_graphify_settings.code_onlyand sent on each sidecar request; the env var is only the fallback when the setting is unset. Code-only builds no longer require a configured graph-builder model. - ⚙️ Source scope for the github_code integration got exclude globs
(
docs/**,licenses/**,CHANGES/**, plus the nine doc extensions) so doc files never land inraw/code/at all. 433 doc files were deleted from the repo (sync delete proposals).
The LLM leak that remained (the gpt-4.1-mini mystery): even in code-only
mode the sidecar's post-extract step ran graphify cluster-only <workdir> with
no --model/--backend/--no-label. Community naming then auto-detects a
backend and uses graphify's hardcoded openai default gpt-4.1-mini
(graphify/llm.py:155) — the platform graph_builder model selection only
flows into extract, never into cluster-only.
- ⚙️ Fix in
docker/graphify-sidecar/server.py: append--no-labelto the cluster-only call whenCODE_ONLY. Communities keep deterministic hub-based names. Requires sidecar image rebuild to take effect (docker compose build graphify-sidecar && docker compose up -d graphify-sidecar) — already done in this environment. — corrected 2026-08-27: not an issue.PROVIDER_TO_BACKENDinserver.pyhas noopenrouterentryresolveAiSelection(packages/db/src/ai-selection.ts:62-67) hardcodesprovider: 'openai'pointed at the router, so the gateway provider id inplatform_ai_settingsnever reaches the sidecar; the openai branch (ollama backend override +OLLAMA_MODELpin) already handles every selection. Thegpt-4.1-minicalls came solely from the unpinnedcluster-onlyinvocation above.
2. Sync delete flow (purging already-mirrored files)
- Orphan deletions are rate-limited by
MAX_DELETE_PROPOSALS_PER_RUN = 25(packages/integrations/src/sync/orphans.ts) and every batch needs review in the UI. Purging 433 docs at 25/run = ~18 review cycles. - ⚙️ The cap was temporarily bumped to 500 for the purge and has been reverted to 25 — nothing to do, listed for the record.
- ✅ Productised (2026-08-27). Admin → LLM Settings → Graphify → Accept
sync proposals automatically approves deletions with everything else,
commits them in the same writer job, and lifts the per-run cap, so a purge
of any size drains in one run. Stored in
platform_graphify_settings.auto_accept_proposals. - The
auto_rulesengine (packages/staging) remains unwired with zero call sites; the platform-wide boolean supersedes it for this use.
3. GitHub commit limits (the missing graph.json)
- The repo-writer ships oversized files via
createBlob; its comment claimed a 100MB ceiling but GitHub rejects ~40MB+ payloads ("Sorry, your input was too large to process") — the 41MBgraph.jsonis 55MB as base64. - ⚙️
collectGraphifyOutput.tsMAX_FILE_BYTESlowered 50MB → 30MB, so oversized artifacts are skipped from the API commit with a warning instead of failing the whole writer job three times. - ⚙️ The UI's "graph built" marker (
GRAPH_ARTIFACT_PATHinpackages/shared/src/graphify-paths.ts) wasgraphify-out/graph.json— with the file skipped, the pipeline card showed "N files ready to build" forever. Repointed tographify-out/graph.html(988KB, always commits, and is what the UI actually renders). The worker now reads build stats from the workspace copy of graph.json instead of the commit list (apps/worker/.../graphify.ts). - ⚙️ One-off:
graph.jsonwas pushed to the repo via real git push from inside the worker container (shallow clone + installation token; git's per-file limit is 100MB). Pattern lives in the session scratchpad scripts. - ⬜ Durable follow-up: future rebuilds skip
graph.jsonagain, so the committed copy goes stale one build behind whenever the graph changes. If a current in-repo graph.json matters, add a git-push fallback to the graphify worker for artifacts over the API cap (clone → cp → commit → push, exactly like the one-off). If it doesn't matter, remove the committed graph.json to avoid confusion.
4. Redis OOM crash-loop (the cause of every "stuck at Reporting" build)
Symptom chain: build finishes in the sidecar → redis is OOM-killed at that
moment → worker's BullMQ bookkeeping hangs (ENOTFOUND redis /
ECONNREFUSED) → UI shows "Reporting, NN:NN elapsed" forever → a worker
restart makes BullMQ retry the stalled job, which opens a new graph_run row,
often immediately failed ("job stalled more than allowable limit") and left
orphaned at running.
- Root cause is memory, in two layers:
- Container layer: BullMQ writer jobs carry entire repo file bodies
(a 1,778-file mirror commit ≈ 50MB payload). Redis dataset grew to
~160-320MB; AOF history + allocator fragmentation put process RSS at
2-4× the dataset; the save-fork spike then clipped whatever
mem_limitwas set (crash-looped at 256m with 278 restarts, again at 768m with 121). - Host layer: the Docker VM had only 4GB total against ~8.5GB of summed container limits; the VM-wide OOM killer picked redis at every build+sync peak even when redis was under its own cgroup limit.
- Container layer: BullMQ writer jobs carry entire repo file bodies
(a 1,778-file mirror commit ≈ 50MB payload). Redis dataset grew to
~160-320MB; AOF history + allocator fragmentation put process RSS at
2-4× the dataset; the save-fork spike then clipped whatever
- ⚙️
docker-compose.ymlredismem_limitraised 256m → 768m → 1280m, AOF compacted (BGREWRITEAOF), allocator purged. - ⚙️ Docker Desktop VM raised from 4GB (user applied ~8GB / 6 CPUs).
- ⬜ This Mac is an 8GB M1 — an 8GB VM overcommits the machine; dial the Docker Desktop memory setting back to 5GB. CPU count is irrelevant to this failure mode.
- ⬜ Consider trimming what rides through redis: writer-queue retention already
ages out payloads, but
BGREWRITEAOFafter large mirror syncs (orauto-aof-rewrite-percentagetuning) would keep boot-time RSS down.
Recognizing a recurrence: docker inspect -f '{{.RestartCount}}' project-brain-redis-1
climbing, getaddrinfo ENOTFOUND redis / ECONNREFUSED floods in worker logs,
"container is restarting" on docker exec, and pipeline cards frozen at
Reporting with a climbing clock.
4b. Worker Node heap OOM (a second "stuck at Reporting" cause)
Distinct from the redis kills: with redis perfectly healthy, a build froze at
"Reporting" because the worker's Node process itself crashed with
FATAL ERROR: Reached heap limit Allocation failed - JavaScript heap out of memory (~830MB, V8's default cap of ~half the container's 1.5G limit) while
assembling the graphify commit — the ~60MB graphify-out file set as JS strings
plus the 41MB graph.json stats parse held simultaneously. tsx watch respawns
the process, so docker ps still shows the container healthy/Up — only the
log tail reveals the crash. The dead process orphans its BullMQ job and
graph_run row exactly like a redis kill does.
- ⚙️
docker-compose.ymlworker env:NODE_OPTIONS: --max-old-space-size=1152(verified live: 1176MB heap limit). - ⚙️
JSON.parseof a largegraph.jsonwas the other half. Reading it for build stats costs roughly 10× the file size in heap; django's 67MB graph killed the worker outright even at the raised limit (child process gone, tsx supervisor still up, so the container looked healthy while the queue sat wedged).readGraphStats(collectGraphifyOutput.ts) now parses graph.json only under 32MB and otherwise reads the node/edge/community counts out ofGRAPH_REPORT.md, which is a few hundred KB. The per-community fingerprint (UI community-delta) is simply absent for graphs above that line. - ⬜ Structural reduction if it recurs: stop committing
cache/(the bulk of the ~60MB commit payload) or stream the collect instead of holding all file bodies at once.
4c. Sidecar OOM on large corpora (sidecar_status_500)
Third memory variant (hit 2026-08-27 on the django mirror, ~48k nodes /
~107k edges): the graphify Python process inside the sidecar is OOM-killed
during the in-process graph merge/clustering phase. Signature:
error_message = sidecar_status_500 on the run, the stored log ends cleanly
right after "AST extraction: 100%" with no error text, and the sidecar
cgroup shows it plainly — memory.peak exactly at the limit and
oom_kill > 0 in /sys/fs/cgroup/memory.events. Astropy (~27k nodes) fit in
1g; django didn't.
- ⚙️
docker-compose.ymlgraphify-sidecarmem_limitraised 1g → 1536m. - ⬜ Memory scales with corpus size; unpinned repos (django tracks branch head) keep growing. If it recurs on a bigger corpus, raise further or prune scope.
5. BullMQ bookkeeping traps (found while cleaning up)
- A
graphify-auto-<projectId>job left in any state (includingfailed) silently blocks every future auto-rebuild add — BullMQ jobId dedupe ignores adds whose id still exists.writer.tsdocuments this; it bit us when a manually-killed job landed in the failed set. Clean withZREM bull:graphify:failed <id>+DEL bull:graphify:<id>. - Stalled-job retries re-run with the job's stored data; if the job hash was
deleted, the retry fires with empty data and fails on
projects.id = ''— a confusing but harmless signature. - Orphaned
graph_runsrows atrunning(their job died with the worker/redis) need manual resolution: markfailedwith an error message, or delete the row if the run did no work; the UI otherwise shows a phantom in-progress build with an ever-climbing timer. - Restarting the worker mid-build orphans the job in BullMQ's
activelist, and because the graphify worker runs at concurrency 1 that one dead entry blocks every later build indefinitely — no error, no progress, just a queue that never moves. Hit again on 2026-08-27 while deploying. Recognise it byLLEN bull:graphify:active= 1 with an idle sidecar and no worker log lines; clear withLREM bull:graphify:active 0 <jobId>+DEL bull:graphify:<jobId>, then restart the worker so the waiting jobs are picked up. Prefer draining the queue before restarting the worker. (/api/projects/:slug/graph/rebuildreapsrunningrows older than 35 minutes, so the UI self-heals; the queue entry does not.)
6. Misc environment notes
- GitHub REST vs git push for large files: REST = no clone, but base64 (+33%) uncompressed on the wire and a ~40MB blob ceiling; git push = local clone + pack CPU, but compressed upload and a 100MB file limit. The writer's REST/tree design is right for many small files; git push is the tool for the rare oversized artifact.
- The Next.js dev server 404s every route (while
/still 307s) after being SIGKILLed mid-write — corrupted.next. Fix:rm -rf apps/web/.nextand restart. (pnpm exec next devfromapps/web, notpnpm dev, on this machine.) - A hard VM restart leaves squid's stale pidfile in the egress proxy →
"Squid is already running" crash-loop. Fix:
docker compose up -d --force-recreate graphify-egress-proxy. pb-eval-mysql(a different project's container, no restart policy) does not come back after a VM restart — start it manually if needed.- The graphify sidecar keeps running its
graphify extractsubprocess even if the worker-side job is killed — restart the sidecar container too when aborting a build, or the zombie finishes and may make LLM calls.
7. Uncommitted changes on this branch from the session
| File | Change | Keep? |
|---|---|---|
docker/graphify-sidecar/server.py | --no-label on cluster-only when CODE_ONLY | ✅ bug fix |
docker-compose.yml | redis mem_limit: 1280m; worker NODE_OPTIONS: --max-old-space-size=1152; sidecar mem_limit: 1536m | ✅ bug fixes |
apps/worker/src/workers/graphify/collectGraphifyOutput.ts | 30MB per-file commit cap | ✅ bug fix |
apps/worker/src/workers/graphify.ts | stats read from workspace graph.json | ✅ bug fix |
packages/shared/src/graphify-paths.ts | artifact marker → graph.html | ✅ bug fix |
packages/integrations/src/sync/orphans.ts | cap bump | already reverted |
DB-side (not in git): astropy project's integration scope has the doc exclude globs; graph_runs history contains manually-failed/cleaned rows from the incident timeline.
8. Follow-up feature work (2026-08-27)
The two recurring sources of manual intervention above were turned into settings rather than left as tribal knowledge:
- Admin → LLM Settings → Graphify (
platform_graphify_settings, one row): Accept sync proposals automatically and Graph scope (code only / code and documents / server default). Route:POST /api/admin/graphify-settings. codeOnlynow travels on every sidecar request rather than living in the container's environment, so switching modes needs no redeploy — but note the sidecar'sALLOWED_KEYSis a strict allowlist, so rebuild the sidecar image before shipping a worker that sends the new field.- The GitHub Code scope form gained repo and branch pickers backed by
GET /api/integrations/:id/reposandGET /api/integrations/:id/repos/:owner/:repo/branches, abranchkey on the scope schema, andcommit → branch → default branchresolution inresolveRef. Repo rows are encoded asowner/repo#branch@commit(parseRepoLine/serializeRepoLineinpackages/integrations/src/ui).