You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Is your feature request related to a problem? Please describe.
Analyzing vscode (18,391 files, 9,351 modules, 174,767 callables) costs 18.1 GB at -a 1 and 24.3 GB at -a 2. The correct per-program dataflow fix in #111 pushes past what a single process can hold and is killed by a JSC heap OOM. The analyzer is close to its ceiling on a repository that is large but not exceptional, and nothing in the pipeline bounds growth.
There are three independent ceilings, hit in this order as a repository grows. #111 is only the first.
#
Ceiling
What holds it
Measured on vscode
1
Program materialization
ts-morph ASTs + tsc checker state, per program
18.1 GB at -a 1; OOM ~29 GB once all 92 programs materialize (#111)
2
Resident output tree
symbol table + 1,149,984 edge objects + per-callable graphs
24.3 GB at -a 2
3
Emit
src/utils/serialize.ts:26 — JSON.stringify(application) builds the entire 1.03 GB output as a single string
survives today; hard wall at roughly 2-4x
Ceiling 3 has already caused one failure: finalizeAnalysis used a JSON.stringify deep-copy roundtrip to strip internal fields and OOM'd at vscode -a 4. Moving that copy to structuredClone fixed the intermediate, but the final write is still one giant string.
The proposed path, in order. Each step is independently shippable and independently useful.
Step 1 — Streaming emit (ceiling 3). Stream the envelope per module rather than building one JSON.stringify string. Independent of everything else, needed regardless, and removes roughly half the emit-time peak plus the string-length wall. Cheapest real win.
Step 2 — Bound the resident program set (ceiling 1). What makes L3/L4 populate 0.7% of callables on multi-program repos — dataflow uses one root tsconfig #111 landable. Three candidate shapes: serialize the call-graph and extraction phases so each Project can be disposed after use; an LRU project pool holding at most N programs; or true subprocess workers. Pick one in that work item's design.
Step 3 — Incremental reuse.<cache_dir>/graphs_summaries.json is written on every run and read by nothing — one write site, no read. The content-hash cache and summary persistence are scaffolded but unwired. For repeated runs on a large repository this dominates every other saving, because most files do not change between runs.
Step 4 — Shard by program, union by id. The genuine scale-out, and the schema was built for it: can:// ids are path- and signature-derived and span-ordered, never discovery-ordered, so the same file yields the same id in any run. Analyze program by program, emit a shard each, union the results. Peak memory becomes the largest single program rather than the whole repository.
Rationale for this order: steps 1 and 2 buy headroom; only steps 3 and 4 change the growth curve. Step 1 is first because it is cheap, independent, and blocks no other design. Step 4 is last because it is the one that deserves a deliberate design rather than being reached for early.
Describe alternatives you've considered
Not stated in the original issue.
Additional context
Scope boundary
This issue is the sequencing decision for the four options below, not any one of their implementations. Each option becomes its own work item, filed when it is picked up.
Not in scope: the per-program correctness fix itself (#111) and multi-program resolution (#56). This issue is about making that correctness affordable.
Caveats and known risks
Step 2 costs the concurrency win. Extraction is deliberately started before the call-graph solve so the two overlap (src/core.ts); serializing them to allow disposal gives that up. An LRU pool keeps the overlap but pays recompute on eviction.
Step 4 changes cross-program edges. Calls that leave a shard degrade to phantoms/externals. That is the existing behaviour for genuinely external callees, but sharding will apply it to some first-party edges that resolve today. The blast radius needs measuring before committing.
Step 4 needs a JSON merge step. Neo4j gets it nearly free (MERGE dedupes by id, which is what the neutral id tier already buys); analysis.json consumers do not.
--target-files does not currently help. It filters which files are built, but every program is still constructed, so it does not cut the dominant cost.
This issue closes when the path is ratified and the first work item is filed — not when all four land. Each step carries its own definition of done, and each must be verified by re-running the vscode measurements above rather than by fixture tests alone, since every ceiling here is invisible at fixture scale.
Is your feature request related to a problem? Please describe.
Analyzing vscode (18,391 files, 9,351 modules, 174,767 callables) costs 18.1 GB at
-a 1and 24.3 GB at-a 2. The correct per-program dataflow fix in #111 pushes past what a single process can hold and is killed by a JSC heap OOM. The analyzer is close to its ceiling on a repository that is large but not exceptional, and nothing in the pipeline bounds growth.There are three independent ceilings, hit in this order as a repository grows. #111 is only the first.
-a 1; OOM ~29 GB once all 92 programs materialize (#111)-a 2src/utils/serialize.ts:26—JSON.stringify(application)builds the entire 1.03 GB output as a single stringCeiling 3 has already caused one failure:
finalizeAnalysisused aJSON.stringifydeep-copy roundtrip to strip internal fields and OOM'd at vscode-a 4. Moving that copy tostructuredClonefixed the intermediate, but the final write is still one giant string.Reference measurements, analyzer v1.1.0,
--no-build:-a 1analysis.json-a 2-a 4(root-only indexing, current)-a 4(per-program, #111 branch)Describe the solution you'd like
The proposed path, in order. Each step is independently shippable and independently useful.
JSON.stringifystring. Independent of everything else, needed regardless, and removes roughly half the emit-time peak plus the string-length wall. Cheapest real win.<cache_dir>/graphs_summaries.jsonis written on every run and read by nothing — one write site, no read. The content-hash cache and summary persistence are scaffolded but unwired. For repeated runs on a large repository this dominates every other saving, because most files do not change between runs.can://ids are path- and signature-derived and span-ordered, never discovery-ordered, so the same file yields the same id in any run. Analyze program by program, emit a shard each, union the results. Peak memory becomes the largest single program rather than the whole repository.Rationale for this order: steps 1 and 2 buy headroom; only steps 3 and 4 change the growth curve. Step 1 is first because it is cheap, independent, and blocks no other design. Step 4 is last because it is the one that deserves a deliberate design rather than being reached for early.
Describe alternatives you've considered
Not stated in the original issue.
Additional context
Scope boundary
This issue is the sequencing decision for the four options below, not any one of their implementations. Each option becomes its own work item, filed when it is picked up.
Not in scope: the per-program correctness fix itself (#111) and multi-program resolution (#56). This issue is about making that correctness affordable.
Caveats and known risks
src/core.ts); serializing them to allow disposal gives that up. An LRU pool keeps the overlap but pays recompute on eviction.-j Ndoes not bound memory. BunWorkers are threads in one process, so each worker gets its own JSC heap but total process RSS is unaffected.-j 4died at 29.2 GB on the L3/L4 populate 0.7% of callables on multi-program repos — dataflow uses one root tsconfig #111 branch. Any subprocess-worker design must use real child processes.MERGEdedupes by id, which is what the neutral id tier already buys);analysis.jsonconsumers do not.--target-filesdoes not currently help. It filters which files are built, but every program is still constructed, so it does not cut the dominant cost.-a 1/-a 2, which work today. The forcing function is-a 3/-a 4correctness (L3/L4 populate 0.7% of callables on multi-program repos — dataflow uses one root tsconfig #111), not-a 1/-a 2.Definition of done
This issue closes when the path is ratified and the first work item is filed — not when all four land. Each step carries its own definition of done, and each must be verified by re-running the vscode measurements above rather than by fixture tests alone, since every ceiling here is invisible at fixture scale.