Skip to main content

Graph pipeline findings — 2026-08-26 debugging session

Working notes from getting a zero-LLM, code-AST-only graph build working for projectbrain-astropy-swe-bench-lite (astropy mirror, 1,345 code files), and from the infrastructure failures found along the way. Each section lists what happened, what was changed (⚙️ = change already applied on this branch / environment), and what is still open (⬜ = decision or follow-up needed).

Verified end state: full rebuild in ~2 min, 27,551 nodes / 57,784 edges, 100% from code files, zero LLM calls (confirmed at OmniRoute, the only egress), output committed including the 41MB graph.json.


1. Zero-LLM (AST-only) graph builds​

How graphify decides what hits the LLM: routing is purely by file extension (graphify/detect.py in the sidecar image). CODE_EXTENSIONS (.py, .c, .json, …) go through tree-sitter at zero tokens; DOC_EXTENSIONS (.md .mdx .qmd .skill .txt .rst .html .yaml .yml) go to the LLM as "semantic" files.

  • ⚙️ GRAPHIFY_CODE_ONLY=1 (already in .env, passed through compose) makes the sidecar run graphify extract --code-only: semantic pass skipped entirely, no API key needed.
  • ✅ Productised (2026-08-27). This is now Admin → LLM Settings → Graphify → Graph scope, stored in platform_graphify_settings.code_only and sent on each sidecar request; the env var is only the fallback when the setting is unset. Code-only builds no longer require a configured graph-builder model.
  • ⚙️ Source scope for the github_code integration got exclude globs (docs/**, licenses/**, CHANGES/**, plus the nine doc extensions) so doc files never land in raw/code/ at all. 433 doc files were deleted from the repo (sync delete proposals).

The LLM leak that remained (the gpt-4.1-mini mystery): even in code-only mode the sidecar's post-extract step ran graphify cluster-only <workdir> with no --model/--backend/--no-label. Community naming then auto-detects a backend and uses graphify's hardcoded openai default gpt-4.1-mini (graphify/llm.py:155) — the platform graph_builder model selection only flows into extract, never into cluster-only.

  • ⚙️ Fix in docker/graphify-sidecar/server.py: append --no-label to the cluster-only call when CODE_ONLY. Communities keep deterministic hub-based names. Requires sidecar image rebuild to take effect (docker compose build graphify-sidecar && docker compose up -d graphify-sidecar) — already done in this environment.
  • PROVIDER_TO_BACKEND in server.py has no openrouter entry — corrected 2026-08-27: not an issue. resolveAiSelection (packages/db/src/ai-selection.ts:62-67) hardcodes provider: 'openai' pointed at the router, so the gateway provider id in platform_ai_settings never reaches the sidecar; the openai branch (ollama backend override + OLLAMA_MODEL pin) already handles every selection. The gpt-4.1-mini calls came solely from the unpinned cluster-only invocation above.

2. Sync delete flow (purging already-mirrored files)​

  • Orphan deletions are rate-limited by MAX_DELETE_PROPOSALS_PER_RUN = 25 (packages/integrations/src/sync/orphans.ts) and every batch needs review in the UI. Purging 433 docs at 25/run = ~18 review cycles.
  • ⚙️ The cap was temporarily bumped to 500 for the purge and has been reverted to 25 — nothing to do, listed for the record.
  • ✅ Productised (2026-08-27). Admin → LLM Settings → Graphify → Accept sync proposals automatically approves deletions with everything else, commits them in the same writer job, and lifts the per-run cap, so a purge of any size drains in one run. Stored in platform_graphify_settings.auto_accept_proposals.
  • The auto_rules engine (packages/staging) remains unwired with zero call sites; the platform-wide boolean supersedes it for this use.

3. GitHub commit limits (the missing graph.json)​

  • The repo-writer ships oversized files via createBlob; its comment claimed a 100MB ceiling but GitHub rejects ~40MB+ payloads ("Sorry, your input was too large to process") — the 41MB graph.json is 55MB as base64.
  • ⚙️ collectGraphifyOutput.ts MAX_FILE_BYTES lowered 50MB → 30MB, so oversized artifacts are skipped from the API commit with a warning instead of failing the whole writer job three times.
  • ⚙️ The UI's "graph built" marker (GRAPH_ARTIFACT_PATH in packages/shared/src/graphify-paths.ts) was graphify-out/graph.json — with the file skipped, the pipeline card showed "N files ready to build" forever. Repointed to graphify-out/graph.html (988KB, always commits, and is what the UI actually renders). The worker now reads build stats from the workspace copy of graph.json instead of the commit list (apps/worker/.../graphify.ts).
  • ⚙️ One-off: graph.json was pushed to the repo via real git push from inside the worker container (shallow clone + installation token; git's per-file limit is 100MB). Pattern lives in the session scratchpad scripts.
  • ⬜ Durable follow-up: future rebuilds skip graph.json again, so the committed copy goes stale one build behind whenever the graph changes. If a current in-repo graph.json matters, add a git-push fallback to the graphify worker for artifacts over the API cap (clone → cp → commit → push, exactly like the one-off). If it doesn't matter, remove the committed graph.json to avoid confusion.

4. Redis OOM crash-loop (the cause of every "stuck at Reporting" build)​

Symptom chain: build finishes in the sidecar → redis is OOM-killed at that moment → worker's BullMQ bookkeeping hangs (ENOTFOUND redis / ECONNREFUSED) → UI shows "Reporting, NN:NN elapsed" forever → a worker restart makes BullMQ retry the stalled job, which opens a new graph_run row, often immediately failed ("job stalled more than allowable limit") and left orphaned at running.

  • Root cause is memory, in two layers:
    1. Container layer: BullMQ writer jobs carry entire repo file bodies (a 1,778-file mirror commit ≈ 50MB payload). Redis dataset grew to ~160-320MB; AOF history + allocator fragmentation put process RSS at 2-4× the dataset; the save-fork spike then clipped whatever mem_limit was set (crash-looped at 256m with 278 restarts, again at 768m with 121).
    2. Host layer: the Docker VM had only 4GB total against ~8.5GB of summed container limits; the VM-wide OOM killer picked redis at every build+sync peak even when redis was under its own cgroup limit.
  • ⚙️ docker-compose.yml redis mem_limit raised 256m → 768m → 1280m, AOF compacted (BGREWRITEAOF), allocator purged.
  • ⚙️ Docker Desktop VM raised from 4GB (user applied ~8GB / 6 CPUs).
  • ⬜ This Mac is an 8GB M1 — an 8GB VM overcommits the machine; dial the Docker Desktop memory setting back to 5GB. CPU count is irrelevant to this failure mode.
  • ⬜ Consider trimming what rides through redis: writer-queue retention already ages out payloads, but BGREWRITEAOF after large mirror syncs (or auto-aof-rewrite-percentage tuning) would keep boot-time RSS down.

Recognizing a recurrence: docker inspect -f '{{.RestartCount}}' project-brain-redis-1 climbing, getaddrinfo ENOTFOUND redis / ECONNREFUSED floods in worker logs, "container is restarting" on docker exec, and pipeline cards frozen at Reporting with a climbing clock.

4b. Worker Node heap OOM (a second "stuck at Reporting" cause)​

Distinct from the redis kills: with redis perfectly healthy, a build froze at "Reporting" because the worker's Node process itself crashed with FATAL ERROR: Reached heap limit Allocation failed - JavaScript heap out of memory (~830MB, V8's default cap of ~half the container's 1.5G limit) while assembling the graphify commit — the ~60MB graphify-out file set as JS strings plus the 41MB graph.json stats parse held simultaneously. tsx watch respawns the process, so docker ps still shows the container healthy/Up — only the log tail reveals the crash. The dead process orphans its BullMQ job and graph_run row exactly like a redis kill does.

  • ⚙️ docker-compose.yml worker env: NODE_OPTIONS: --max-old-space-size=1152 (verified live: 1176MB heap limit).
  • ⚙️ JSON.parse of a large graph.json was the other half. Reading it for build stats costs roughly 10× the file size in heap; django's 67MB graph killed the worker outright even at the raised limit (child process gone, tsx supervisor still up, so the container looked healthy while the queue sat wedged). readGraphStats (collectGraphifyOutput.ts) now parses graph.json only under 32MB and otherwise reads the node/edge/community counts out of GRAPH_REPORT.md, which is a few hundred KB. The per-community fingerprint (UI community-delta) is simply absent for graphs above that line.
  • ⬜ Structural reduction if it recurs: stop committing cache/ (the bulk of the ~60MB commit payload) or stream the collect instead of holding all file bodies at once.

4c. Sidecar OOM on large corpora (sidecar_status_500)​

Third memory variant (hit 2026-08-27 on the django mirror, ~48k nodes / ~107k edges): the graphify Python process inside the sidecar is OOM-killed during the in-process graph merge/clustering phase. Signature: error_message = sidecar_status_500 on the run, the stored log ends cleanly right after "AST extraction: 100%" with no error text, and the sidecar cgroup shows it plainly — memory.peak exactly at the limit and oom_kill > 0 in /sys/fs/cgroup/memory.events. Astropy (~27k nodes) fit in 1g; django didn't.

  • ⚙️ docker-compose.yml graphify-sidecar mem_limit raised 1g → 1536m.
  • ⬜ Memory scales with corpus size; unpinned repos (django tracks branch head) keep growing. If it recurs on a bigger corpus, raise further or prune scope.

5. BullMQ bookkeeping traps (found while cleaning up)​

  • A graphify-auto-<projectId> job left in any state (including failed) silently blocks every future auto-rebuild add — BullMQ jobId dedupe ignores adds whose id still exists. writer.ts documents this; it bit us when a manually-killed job landed in the failed set. Clean with ZREM bull:graphify:failed <id> + DEL bull:graphify:<id>.
  • Stalled-job retries re-run with the job's stored data; if the job hash was deleted, the retry fires with empty data and fails on projects.id = '' — a confusing but harmless signature.
  • Orphaned graph_runs rows at running (their job died with the worker/redis) need manual resolution: mark failed with an error message, or delete the row if the run did no work; the UI otherwise shows a phantom in-progress build with an ever-climbing timer.
  • Restarting the worker mid-build orphans the job in BullMQ's active list, and because the graphify worker runs at concurrency 1 that one dead entry blocks every later build indefinitely — no error, no progress, just a queue that never moves. Hit again on 2026-08-27 while deploying. Recognise it by LLEN bull:graphify:active = 1 with an idle sidecar and no worker log lines; clear with LREM bull:graphify:active 0 <jobId> + DEL bull:graphify:<jobId>, then restart the worker so the waiting jobs are picked up. Prefer draining the queue before restarting the worker. (/api/projects/:slug/graph/rebuild reaps running rows older than 35 minutes, so the UI self-heals; the queue entry does not.)

6. Misc environment notes​

  • GitHub REST vs git push for large files: REST = no clone, but base64 (+33%) uncompressed on the wire and a ~40MB blob ceiling; git push = local clone + pack CPU, but compressed upload and a 100MB file limit. The writer's REST/tree design is right for many small files; git push is the tool for the rare oversized artifact.
  • The Next.js dev server 404s every route (while / still 307s) after being SIGKILLed mid-write — corrupted .next. Fix: rm -rf apps/web/.next and restart. (pnpm exec next dev from apps/web, not pnpm dev, on this machine.)
  • A hard VM restart leaves squid's stale pidfile in the egress proxy → "Squid is already running" crash-loop. Fix: docker compose up -d --force-recreate graphify-egress-proxy.
  • pb-eval-mysql (a different project's container, no restart policy) does not come back after a VM restart — start it manually if needed.
  • The graphify sidecar keeps running its graphify extract subprocess even if the worker-side job is killed — restart the sidecar container too when aborting a build, or the zombie finishes and may make LLM calls.

7. Uncommitted changes on this branch from the session​

FileChangeKeep?
docker/graphify-sidecar/server.py--no-label on cluster-only when CODE_ONLY✅ bug fix
docker-compose.ymlredis mem_limit: 1280m; worker NODE_OPTIONS: --max-old-space-size=1152; sidecar mem_limit: 1536m✅ bug fixes
apps/worker/src/workers/graphify/collectGraphifyOutput.ts30MB per-file commit cap✅ bug fix
apps/worker/src/workers/graphify.tsstats read from workspace graph.json✅ bug fix
packages/shared/src/graphify-paths.tsartifact marker → graph.html✅ bug fix
packages/integrations/src/sync/orphans.tscap bumpalready reverted

DB-side (not in git): astropy project's integration scope has the doc exclude globs; graph_runs history contains manually-failed/cleaned rows from the incident timeline.

8. Follow-up feature work (2026-08-27)​

The two recurring sources of manual intervention above were turned into settings rather than left as tribal knowledge:

  • Admin → LLM Settings → Graphify (platform_graphify_settings, one row): Accept sync proposals automatically and Graph scope (code only / code and documents / server default). Route: POST /api/admin/graphify-settings.
  • codeOnly now travels on every sidecar request rather than living in the container's environment, so switching modes needs no redeploy — but note the sidecar's ALLOWED_KEYS is a strict allowlist, so rebuild the sidecar image before shipping a worker that sends the new field.
  • The GitHub Code scope form gained repo and branch pickers backed by GET /api/integrations/:id/repos and GET /api/integrations/:id/repos/:owner/:repo/branches, a branch key on the scope schema, and commit → branch → default branch resolution in resolveRef. Repo rows are encoded as owner/repo#branch@commit (parseRepoLine / serializeRepoLine in packages/integrations/src/ui).