From 3a9223285bcf1d1ce48c5cbe9a81e373752611cc Mon Sep 17 00:00:00 2001 From: lambertpw Date: Mon, 28 Sep 2026 22:11:03 -0400 Subject: [PATCH 01/17] docs: design record for v1 -> v2 schema migration (L1 + L2) Roadmap, cross-repo spec, schema decisions, and drafted parent/child issue bodies for migrating codeanalyzer-go and its python-sdk Go models to the v2 canonical schema. Reach: L1 + L2 to parity; L3/L4 deferred to a later train. Group A vocabulary (can:// ids, span.bytes UTF-8) decided once here per the parity clause, since L3/L4 reuse it verbatim. Full transcript in docs/design/specs/v2-l1-emission.md and CLAUDE.md. --- CLAUDE.md | 47 ++++++++++++ docs/design/issues/child-1-l1-emission.md | 40 ++++++++++ docs/design/issues/parent-v2-migration.md | 41 +++++++++++ docs/design/roadmap.md | 77 +++++++++++++++++++ docs/design/specs/v2-l1-emission.md | 90 +++++++++++++++++++++++ 5 files changed, 295 insertions(+) create mode 100644 CLAUDE.md create mode 100644 docs/design/issues/child-1-l1-emission.md create mode 100644 docs/design/issues/parent-v2-migration.md create mode 100644 docs/design/roadmap.md create mode 100644 docs/design/specs/v2-l1-emission.md diff --git a/CLAUDE.md b/CLAUDE.md new file mode 100644 index 0000000..1694459 --- /dev/null +++ b/CLAUDE.md @@ -0,0 +1,47 @@ +# codeanalyzer-go — agent notes + +The Go static-analysis backend for CLDK (`cango`). Emits the canonical CLDK +`analysis.json`. Parser/resolver/call-graph builder compute the facts; the emission +layer serializes them into the schema shape. + +## Schema decisions + +Decisions made with the user during the v1 → v2 migration design loop +(`designing-cldk-changes`), spine order. Each is a leaf-level addition to the shared +v2 vocabulary (parity clause: add at the leaves, never rename shared names). The +committed spec is the full transcript: `docs/design/specs/v2-l1-emission.md`. + +### Identity & positioning (Group A — reused verbatim by L3/L4) +- **`application.id`** = `can://go/`, where `` is a slug derived from the + **go.mod module path** (deterministic, no flag). L3/L4 build ids on this unchanged. +- **Callable signature** = existing `signatureOf()` output as the last `can://` path + segment; receiver folded in for methods. One canonicalizer, unchanged from v1. +- **`span.bytes`** = **UTF-8 byte offsets** into `module.source` (Go strings are + natively UTF-8; avoids the UTF-16/rune slicing mismatch TS hit in cants#179). + +### `type` node +- **`kind`** ∈ `struct | interface | alias | defined` — Go's four type-declaration + shapes. Collapses v1's `is_interface` boolean. (`alias` = `type X = Y`; + `defined` = `type X Y` over a non-struct, e.g. `type Celsius float64`.) +- **Methods** (declared outside the type in Go source) are resolved to their receiver + type and placed in that `type.callables{}`; receiver recorded as a field. The tree + reflects containment, not source location. +- **Embedding vs satisfaction:** `base_types[]` = embedded type ids (struct/interface + embedding — the explicit spine); `interfaces[]` = interfaces the type is *computed* + to satisfy via method-set matching (Go's implicit/structural implementation). + +### `callable` node +- **`error_channel[]`** populated from `error`-typed return values (Go's error idiom); + those returns remain in `return_type` too. Maps Go onto the shared field. +- **Closures / function literals** → nested `callables{}` on the enclosing callable + (replaces v1 `InnerCallables`); each closure gets its own `can://` id. + +### `call` body node +- Typed **`is_goroutine`** and **`is_deferred`** boolean fields (`defer` is net-new vs + v1). `is_constructor_call` **dropped** — Go has no constructors. +- `callee` is the sanctioned `null → id` refinement slot (null at L1, backfilled at L2). + +### Deferred (not this train) +- L3 (CFG/CDG/DDG) and L4 (SDG param edges) — reuse Group A ids/spans unchanged. +- Go-specific `cfg` edge kinds (e.g. `defer_resume`, `select`) are recorded when L3 + is designed, not now. diff --git a/docs/design/issues/child-1-l1-emission.md b/docs/design/issues/child-1-l1-emission.md new file mode 100644 index 0000000..7cb806d --- /dev/null +++ b/docs/design/issues/child-1-l1-emission.md @@ -0,0 +1,40 @@ + + +**Is your feature request related to a problem? Please describe.** + +The Go analyzer serializes the v1 shape. L1 is the structural keystone of the v2 +migration: it introduces the additive containment tree and every Group A identity +decision the later levels build on. Do L1 first and get the symbol-table gate green +before touching the L2 call graph. The parser/resolver stay; only the emission layer +changes (golden rule: replace what serializes, keep what computes). + +**Describe the solution you'd like.** + +- [ ] `application.id` = `can://go/`, `` derived from the go.mod module path +- [ ] Durable `can://` ids on module/type/callable; `signatureOf()` is the last segment +- [ ] `span: {start:[l,c], end:[l,c], bytes:[from,to]}` with UTF-8 byte offsets; drop flat `start_line`/`end_line` +- [ ] `source` emitted once per module; drop per-callable `code` +- [ ] `type.kind` ∈ `struct|interface|alias|defined`; methods grouped under receiver type's `callables{}` +- [ ] `base_types[]` = embedded ids, `interfaces[]` = computed structural satisfaction +- [ ] `body{}` holds `call` nodes with `callee:null`, typed `is_goroutine`/`is_deferred`; closures nest in `callables{}` + +**Describe alternatives you've considered.** + +Does NOT emit the `call_graph` (that is the L2 child), nor any L3/L4 `body` statements +or edges. `body{}` at L1 holds only `call` nodes. + +**Additional context.** + +- Emitter approach: walk the existing in-memory model, add a v2 emitter behind a flag; keep v1 emitter as the compat shim during transition. +- One genuinely new datum: thread parser byte offsets into `span.bytes`. +- Structural interface satisfaction needs method-set matching — cost lives here, not in the resolver. +- Validate each node against the SDK v2 Go models as they land; superset check vs v1 output. +- Schema decisions recorded in repo `CLAUDE.md`; full transcript in the spec. diff --git a/docs/design/issues/parent-v2-migration.md b/docs/design/issues/parent-v2-migration.md new file mode 100644 index 0000000..3496ca0 --- /dev/null +++ b/docs/design/issues/parent-v2-migration.md @@ -0,0 +1,41 @@ + + +**Is your feature request related to a problem? Please describe.** + +`codeanalyzer-go` emits the v1 schema (`internal/schema/schema.go`): two-key +`{symbol_table, call_graph}`, flat `start_line`/`end_line`, per-callable `code`, +`is_*` booleans, `signature` as id. Every sibling analyzer has moved to the v2 +keystone (`can://` ids, additive tree, `span` with byte offsets, `source` per module, +`{src,dst,prov}` edges). Go is the last on v1, so `CLDK(language="go")` cannot share +the one-model SDK surface. This is a coordinated schema major across two repos: +`codeanalyzer-go` (emission rewrite) and `python-sdk` (Go model remap), released in +lockstep. Reach for this train: L1 + L2 to parity; L3/L4 deferred. + +**Describe the solution you'd like.** + +- [ ] `codeanalyzer-go`: L1 emission — additive tree, `source` per module, `can://` ids, `body{}` with `call` nodes +- [ ] `codeanalyzer-go`: L2 emission — `call_graph: [{src,dst,prov,weight}]` at application scope +- [ ] `codeanalyzer-go`: Neo4j projection relabelled to v2 node/edge families +- [ ] `python-sdk`: Go v2 model remap, every public accessor name/return type unchanged +- [ ] Release v2.0.0 and pin analyzer in SDK **only once both are cut** +- [ ] Superset gate green: v2 output contains every v1 fact modulo sanctioned drops + +**Describe alternatives you've considered.** + +Does NOT build L3 (CFG/CDG/DDG) or L4 (SDG). Those are a later train reusing this +train's `can://` ids and edge shape unchanged. Does NOT add framework detection. + +**Additional context.** + +- Design transcript: `docs/design/specs/v2-l1-emission.md` (linked, not pasted). +- Group A vocabulary (`can://` ids, `span.bytes` UTF-8) is reused verbatim by L3/L4 — coined once here per the parity clause. +- Sanctioned drops: `code` (slice from `source`), `is_*` booleans (folded into `kind`), `is_constructor_call` (Go has no constructors). +- Ordering risk: SDK's v1 Go models will not parse v2 output; do not pin the new major prematurely. +- go.mod-derived `` slug must be deterministic across runs, else ids churn. diff --git a/docs/design/roadmap.md b/docs/design/roadmap.md new file mode 100644 index 0000000..28e3ada --- /dev/null +++ b/docs/design/roadmap.md @@ -0,0 +1,77 @@ +# CLDK Go analyzer — roadmap + +> Local staging copy. Canonical home is `codellm-devkit/.github → docs/design/roadmap.md`; +> port there when that repo is available. One roadmap, amended in place — never a second file. + +## Initiative: v1 → v2 schema migration (codeanalyzer-go + python-sdk) + +A schema **major**: move the Go analyzer off the v1 keystone +(`{symbol_table, call_graph}`, flat `start_line`/`end_line`, per-callable `code`, `is_*` +booleans, `signature` as id) onto the v2 canonical schema (additive tree, `can://` ids, +`span` with byte offsets, `source` per module, typed edge lists). The SDK moves in +lockstep as a coordinated major release. **Reach for this train: L1 + L2 to parity.** +L3/L4 deferred (see Not now). + +Golden rule carried from `schema-migration.md`: keep everything that *computes* facts +(parser, resolver, call-graph builder); replace only what *serializes* them. + +## Starting now (exactly one) + +- **Group A + L1 emission** — the identity, span, and `body`-with-call-nodes decisions. + This is the keystone rung; everything below reuses its vocabulary. → `designing-cldk-changes`. + +Everything else below has no issue filed yet, by design. + +## Collision groups (the reason this is planning, not design) + +Contract decisions that share v2 vocabulary — **each group is one design session**, even +when its members ship in different release trains. Coining a shared term twice coins it +wrong permanently (parity clause). + +- **A — Identity & positioning.** `application.id` (`can:///`), the `can://` + durable-id grammar (≥ callable), the `@line:col` ordinal ids (< callable), and `span` + with byte offsets. **L3/L4 in a later train reuse this id/span shape verbatim** — so it + must be settled now, in one session, not re-decided per level. This is the group the + gate exists to protect. +- **B — Node shape.** `type.kind` collapsing the `is_interface`/`is_enum`/… booleans; + structured `decorators[]` (`{name,args,span}`) replacing flat annotation strings; + generalized `error_channel[]`; drop per-callable `code`, add `source` once per module. +- **C — Edge shape.** Rename edge keys `source/target/type` → `src/dst/prov`; `call_graph` + as a list at application scope. **Every future edge family (cfg/cdg/ddg/param_*) inherits + this record shape** — coin `{src, dst, prov, weight}` correctly now. +- **D — Neo4j projection.** Node/edge family names must match A/B/C exactly; it is a + relabel of the same tree, always full-depth. + +## Dependency order (DAG) + +``` +A (identity + span) ─┬─► L1 emission (tree, source, ids, body-with-call-nodes) +B (node shape) ─┤ +C (edge shape) ─┴─► L2 emission (call_graph rename @ application scope) + │ +A/B/C ─► D (Neo4j relabel) ─► SDK v2 model remap ─► pin SDK→analyzer (once BOTH cut) + │ + (later train) L3 ─► L4 ──────┘ reuse A & C unchanged +``` + +## Release trains + +- **Train 1 — v2.0.0 (coordinated major, lockstep).** Analyzer L1+L2 emission rewrite + + Neo4j relabel; SDK v2 model remap keeping every public accessor name/return type + identical (two-layer model). **Ordering constraint (first-class on the parent):** pin the + analyzer version in the SDK *only once both are released* — until then the SDK's old + models won't parse v2 output. Superset gate: v2 output must contain every fact v1 did, + modulo the sanctioned drops (`code`, `is_*` booleans folded into `kind`). +- **Train 2+ (deferred).** L3 (CFG/CDG/DDG, AST-only, per-callable parallel — net-new for + Go) then L4 (SDG `param_in`/`param_out`/`summary`, needs a points-to oracle). Both reuse + group A & C vocabulary; no new schema-major coordination beyond what Train 1 establishes. + +## Not now (considered, deliberately excluded) + +- **L3 / L4 dataflow** — deferred by decision this session; net-new heavy construction, not + required for v2 parity. Revisit as Train 2. +- **Points-to oracle** — L4 prerequisite; parked with L4. +- **`--materialize-expressions` / `--materialize-basic-blocks`** — optional body node kinds; + out of scope for the parity migration. +- **New framework detection / CodeQL tier expansion** — orthogonal enrichment axis + (provenance-merged evidence), not a schema level; does not belong on the migration train. diff --git a/docs/design/specs/v2-l1-emission.md b/docs/design/specs/v2-l1-emission.md new file mode 100644 index 0000000..209bf07 --- /dev/null +++ b/docs/design/specs/v2-l1-emission.md @@ -0,0 +1,90 @@ +# Spec: v2 schema migration — Group A + L1 emission (codeanalyzer-go) + +> **Cross-repo spec.** Canonical home is `codellm-devkit/.github → docs/design/specs/`, +> beside the coordinating parent issue. Staged locally until that repo is available. +> This is the design transcript — the committed provenance of the design loop, linked +> (not pasted) from the parent issue. + +## Contract-Impact Triage + +**Does this change schema v2 output?** Yes — it is the schema-major migration that +introduces the v2 shape for Go: `application.id`, the `can://` id grammar, `@line:col` +ordinal ids (defined here, populated by L3/L4), `span` with UTF-8 byte offsets, the +additive tree, `source` per module, and `body{}` with `call` nodes. + +| Change type | Analyzers | SDKs | Docs | +| --- | --- | --- | --- | +| Schema v2 migration | **codeanalyzer-go** (emission rewrite) | **python-sdk** (Go model remap, lockstep) | CLAUDE.md, this spec | + +Arrived from `planning-cldk-work` with collision **Group A** known (identity + span); +its vocabulary is designed once here and reused verbatim by the deferred L3/L4 train. + +## Design-loop decisions (the transcript) + +Every decision below was made with the user, node by node in spine order. Full copy in +`codeanalyzer-go/CLAUDE.md` § Schema decisions. Anchored on the keystone +(`canonical-schema.md`) and the completed **TypeScript** v2 migration (closest structural +analog: module functions + interfaces + structural types). + +### Root envelope (settled spine, kept as-is) +`{ schema_version:"2.0.0", language:"go", max_level, analyzer:{name,version}, +application:{ id, kind:"application", symbol_table, call_graph, param_in, param_out } }`. + +### Identity & positioning (Group A) +| Decision | Choice | +| --- | --- | +| `application.id` | `can://go/`, `` = slug from **go.mod module path** (no flag) | +| Callable signature | existing `signatureOf()` as last id segment; receiver folded in for methods | +| `span.bytes` | **UTF-8 byte offsets** into `module.source` | + +### `module` +`kind:"module"`, `package`, `source` (whole file once), `imports[]`, `types{}`, +`functions{}`, `content_hash`. Drops per-callable `code` (sliced from `source`). + +### `type` +| Decision | Choice | +| --- | --- | +| `kind` | `struct \| interface \| alias \| defined` (collapses v1 `is_interface`) | +| Method placement | resolved to receiver type's `callables{}`; receiver as a field | +| `base_types[]` | embedded type ids (struct/interface embedding) | +| `interfaces[]` | computed structural satisfaction (method-set matching) | + +### `callable` +| Decision | Choice | +| --- | --- | +| `error_channel[]` | from `error`-typed returns (kept in `return_type` too) | +| Closures | nested `callables{}` on enclosing callable (replaces `InnerCallables`) | + +### `call` (body node, L1) +| Decision | Choice | +| --- | --- | +| Markers | typed `is_goroutine`, `is_deferred`; **drop** `is_constructor_call` | +| `callee` | sanctioned `null → id` refinement slot (null at L1, backfilled L2) | + +### Edges (Group C, settled here for L2) +`call_graph: [{src, dst, prov, weight}]` at application scope (rename from v1 +`{source, target, type, provenance}`). + +## Release plan + +- **Train 1 — v2.0.0 (coordinated major, lockstep).** codeanalyzer-go emission rewrite + (L1 tree/source/ids/body-with-calls, then L2 call_graph) + Neo4j relabel; python-sdk + Go model remap to v2 keeping every public accessor name/return type identical. +- **Ordering constraint (first-class):** pin the analyzer version in the SDK **only once + both are released**. Until then the SDK's old Go models won't parse v2 output. +- **Superset gate:** v2 output must contain every fact v1 did, modulo sanctioned drops + (`code`, `is_*` booleans folded into `kind`, `is_constructor_call`). +- **Deferred (Train 2+):** L3 (CFG/CDG/DDG), L4 (SDG). Reuse Group A ids/spans + Group C + edge shape unchanged; no new schema-major coordination. + +## Decomposition (tracking shape — to be confirmed with user) + +Recommended: **parent (in `codellm-devkit/.github`) + one sub-issue per PR**, because the +work spans two repos on their own clocks and needs a coordination record: +1. codeanalyzer-go: L1 emission (tree, source, ids, body-with-call-nodes) +2. codeanalyzer-go: L2 emission (call_graph rename @ app scope) +3. codeanalyzer-go: Neo4j relabel +4. python-sdk: Go v2 model remap (accessors stable) +5. release + pin (once both cut) + +Children filed just-in-time as picked up; this spec records the full plan. From b1ceabaf0e921d012083733764ac0a6cace7ce6e Mon Sep 17 00:00:00 2001 From: lambertpw Date: Mon, 28 Sep 2026 22:41:23 -0400 Subject: [PATCH 02/17] feat(cli): conform to CLDK CLI contract Add the flags the CLDK CLI contract requires so the SDK facade can shell out to the Go backend with the same code path as the other language backends: - --emit : output projection selector. neo4j and schema return a clear non-zero "not yet implemented" error (never a silent fallback to JSON) until the Neo4j child lands. schema needs no --input, matching the contract. - --app-name: application anchor for can://go/ ids and the Neo4j :Application node; defaults to the input dir base name. - --neo4j-uri/user/password/database: accepted and validated now with explicit-flag > env-var > default precedence; consumed by the Neo4j child. - -j/--jobs: worker parallelism, defaulting to CPU cores. No v2 emission logic yet; default output stays the v1 JSON shape. --- cmd/codeanalyzer/main.go | 101 +++++++++++++++++++++++++++------- cmd/codeanalyzer/main_test.go | 34 ++++++++++++ internal/options/options.go | 27 +++++++++ 3 files changed, 141 insertions(+), 21 deletions(-) diff --git a/cmd/codeanalyzer/main.go b/cmd/codeanalyzer/main.go index 66b3187..49bda45 100644 --- a/cmd/codeanalyzer/main.go +++ b/cmd/codeanalyzer/main.go @@ -8,6 +8,8 @@ import ( "encoding/json" "fmt" "os" + "path/filepath" + "runtime" "github.com/spf13/cobra" @@ -30,17 +32,24 @@ func main() { func rootCmd() *cobra.Command { var ( - inputPath string - outputDir string - format string - level int - targetFiles []string - skipTests bool - eager bool - cacheDir string - useCodeQL bool - verbosity int - showVersion bool + inputPath string + outputDir string + format string + emit string + appName string + level int + targetFiles []string + skipTests bool + eager bool + cacheDir string + jobs int + useCodeQL bool + verbosity int + showVersion bool + neo4jURI string + neo4jUser string + neo4jPassword string + neo4jDatabase string ) cmd := &cobra.Command{ @@ -57,6 +66,21 @@ via CLDK(language="go").analysis(project_path=...).`, cmd.Println("cango " + version) return nil } + + // --emit selects the output projection. neo4j/schema are validated + // here and rejected with a non-zero exit until the Neo4j child lands + // (never a silent fallback to JSON). schema needs no --input. + switch options.EmitTarget(emit) { + case options.EmitJSON: + // valid — falls through to the JSON path below. + case options.EmitNeo4j: + return fmt.Errorf("--emit neo4j is not yet implemented; use --emit json") + case options.EmitSchema: + return fmt.Errorf("--emit schema is not yet implemented; use --emit json") + default: + return fmt.Errorf("unsupported --emit target %q; supported: json, neo4j, schema", emit) + } + if inputPath == "" { return fmt.Errorf("--input / -i is required") } @@ -74,18 +98,32 @@ via CLDK(language="go").analysis(project_path=...).`, home, _ := os.UserHomeDir() cacheDir = home + "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/.cldk/go-cache" } + if jobs < 1 { + jobs = runtime.NumCPU() + } + // application anchor: explicit flag > input dir base name. + if appName == "" { + appName = filepath.Base(filepath.Clean(inputPath)) + } opts := options.AnalysisOptions{ - InputPath: inputPath, - OutputDir: outputDir, - Format: format, - Level: options.AnalysisLevel(level), - TargetFiles: targetFiles, - SkipTests: skipTests, - Eager: eager, - CacheDir: cacheDir, - UseCodeQL: useCodeQL, - Verbose: verbosity > 0, + InputPath: inputPath, + OutputDir: outputDir, + Format: format, + Emit: options.EmitTarget(emit), + AppName: appName, + Level: options.AnalysisLevel(level), + TargetFiles: targetFiles, + SkipTests: skipTests, + Eager: eager, + CacheDir: cacheDir, + Jobs: jobs, + UseCodeQL: useCodeQL, + Verbose: verbosity > 0, + Neo4jURI: firstNonEmpty(neo4jURI, os.Getenv("NEO4J_URI")), + Neo4jUser: firstNonEmpty(neo4jUser, os.Getenv("NEO4J_USERNAME"), "neo4j"), + Neo4jPassword: firstNonEmpty(neo4jPassword, os.Getenv("NEO4J_PASSWORD"), "neo4j"), + Neo4jDatabase: firstNonEmpty(neo4jDatabase, os.Getenv("NEO4J_DATABASE")), } analyzer := core.New(opts) @@ -112,6 +150,9 @@ via CLDK(language="go").analysis(project_path=...).`, f.StringVarP(&inputPath, "input", "i", "", "Project root to analyze (required)") f.StringVarP(&outputDir, "output", "o", "", "Output directory for analysis.json (default: stdout)") f.StringVarP(&format, "format", "f", "json", "Output format: json|msgpack") + f.StringVar(&emit, "emit", "json", "Output projection: json|neo4j|schema") + f.StringVar(&appName, "app-name", "", + "Application anchor name for can:// ids and Neo4j :Application (default: input dir name)") f.IntVarP(&level, "analysis-level", "a", 1, "Analysis level: 1=symbol table only, 2=+resolver call graph") f.StringSliceVarP(&targetFiles, "target-files", "t", nil, @@ -119,9 +160,27 @@ via CLDK(language="go").analysis(project_path=...).`, f.BoolVar(&skipTests, "skip-tests", true, "Skip *_test.go files") f.BoolVar(&eager, "eager", false, "Force clean rebuild (ignore cache)") f.StringVarP(&cacheDir, "cache-dir", "c", "", "Cache directory (default: ~/.cldk/go-cache)") + f.IntVarP(&jobs, "jobs", "j", 0, "Worker parallelism (default: CPU cores)") f.BoolVar(&useCodeQL, "codeql", false, "Enable CodeQL framework-based call graph (level 2, stub)") f.CountVarP(&verbosity, "verbose", "v", "Verbosity (repeat for more detail)") f.BoolVar(&showVersion, "version", false, "Print version and exit") + // Neo4j projection targets (validated now, consumed by the Neo4j child). + f.StringVar(&neo4jURI, "neo4j-uri", "", "Live Bolt push target (env NEO4J_URI); omit to write graph.cypher") + f.StringVar(&neo4jUser, "neo4j-user", "", "Neo4j username (env NEO4J_USERNAME, default neo4j)") + f.StringVar(&neo4jPassword, "neo4j-password", "", "Neo4j password (env NEO4J_PASSWORD, default neo4j)") + f.StringVar(&neo4jDatabase, "neo4j-database", "", "Neo4j database (env NEO4J_DATABASE, optional)") + return cmd } + +// firstNonEmpty returns the first non-empty string in vals, or "" if all are +// empty. It encodes the contract's precedence: explicit flag > env var > default. +func firstNonEmpty(vals ...string) string { + for _, v := range vals { + if v != "" { + return v + } + } + return "" +} diff --git a/cmd/codeanalyzer/main_test.go b/cmd/codeanalyzer/main_test.go index acc08bd..ea5ff3b 100644 --- a/cmd/codeanalyzer/main_test.go +++ b/cmd/codeanalyzer/main_test.go @@ -59,6 +59,40 @@ func TestRootCmd_UnknownFormatReturnsError(t *testing.T) { } } +func TestRootCmd_UnknownEmitReturnsError(t *testing.T) { + td := cliTestdataDir() + _, _, err := runCmd("--input", filepath.Join(td, "greeter"), "--emit", "bogus") + if err == nil { + t.Fatal("expected error for unknown --emit value, got nil") + } +} + +func TestRootCmd_EmitNeo4jNotImplemented(t *testing.T) { + td := cliTestdataDir() + _, _, err := runCmd("--input", filepath.Join(td, "greeter"), "--emit", "neo4j") + if err == nil { + t.Fatal("expected non-zero exit for --emit neo4j, got nil") + } + if !strings.Contains(err.Error(), "not yet implemented") { + t.Errorf("error should say 'not yet implemented'; got %q", err.Error()) + } +} + +func TestRootCmd_EmitSchemaNotImplementedWithoutInput(t *testing.T) { + // --emit schema needs no --input; it must still fail (not yet implemented) + // rather than complain about a missing --input. + _, _, err := runCmd("--emit", "schema") + if err == nil { + t.Fatal("expected non-zero exit for --emit schema, got nil") + } + if strings.Contains(err.Error(), "required") { + t.Errorf("--emit schema should not require --input; got %q", err.Error()) + } + if !strings.Contains(err.Error(), "not yet implemented") { + t.Errorf("error should say 'not yet implemented'; got %q", err.Error()) + } +} + // ── --version ──────────────────────────────────────────────────────────────── func TestRootCmd_VersionFlag(t *testing.T) { diff --git a/internal/options/options.go b/internal/options/options.go index 97e6782..051fe6c 100644 --- a/internal/options/options.go +++ b/internal/options/options.go @@ -11,6 +11,18 @@ const ( LevelCallGraph AnalysisLevel = 2 ) +// EmitTarget selects the analyzer's output projection (cli-contract.md § Neo4j). +type EmitTarget string + +const ( + // EmitJSON writes the analysis.json projection (the default). + EmitJSON EmitTarget = "json" + // EmitNeo4j writes/pushes the Neo4j graph projection at full implemented depth. + EmitNeo4j EmitTarget = "neo4j" + // EmitSchema writes the static schema.neo4j.json contract (needs no --input). + EmitSchema EmitTarget = "schema" +) + // AnalysisOptions is the configuration surface passed from the CLI into Analyzer. type AnalysisOptions struct { // InputPath is the project root to analyze. @@ -19,6 +31,11 @@ type AnalysisOptions struct { OutputDir string // Format is the serialization format: "json" or "msgpack". Format string + // Emit selects the output projection: json (default), neo4j, or schema. + Emit EmitTarget + // AppName anchors application identity (can://go/ and the Neo4j + // :Application node). Defaults to the input directory's base name. + AppName string // AnalysisLevel controls symbol-table-only (1) vs + call graph (2). Level AnalysisLevel // TargetFiles restricts analysis to specific files (incremental mode). @@ -29,8 +46,18 @@ type AnalysisOptions struct { Eager bool // CacheDir is where per-file caches and intermediate data are stored. CacheDir string + // Jobs is the worker parallelism (default: CPU cores). Output must be + // byte-identical across Jobs values. + Jobs int // UseCodeQL enables the framework-based (Tier-2) CodeQL call graph. UseCodeQL bool // Verbose enables verbose logging. Verbose bool + + // Neo4j projection targets (consumed by the Neo4j emitter; validated now). + // Precedence for each: explicit flag > env var > default. + Neo4jURI string // live Bolt push target; empty → write graph.cypher. Env NEO4J_URI + Neo4jUser string // env NEO4J_USERNAME, default "neo4j" + Neo4jPassword string // env NEO4J_PASSWORD, default "neo4j" + Neo4jDatabase string // env NEO4J_DATABASE, optional } From 34da145c05c2d8314f4c03a9a0a8038097a6501c Mon Sep 17 00:00:00 2001 From: lambertpw Date: Mon, 28 Sep 2026 22:42:40 -0400 Subject: [PATCH 03/17] feat(schema): add v2 canonical schema types Declare the v2 canonical-schema Go types in internal/schema/v2, a faithful transcript of the keystone (canonical-schema.md) and the Go-specific leaf decisions in CLAUDE.md. Models the L1 additive containment tree (application -> module -> type -> callable -> body) with typed edge overlays, plus the L2 refinement/edge slots so the shape is stable: - Span{start,end,bytes} with UTF-8 byte offsets into module.source. - Module carries source once; no per-node code field. - Type.kind (struct|interface|alias|defined) replaces v1 is_interface; base_types (embedding) vs interfaces (computed satisfaction). - Callable.error_channel, nested callables{} for closures, metrics map. - BodyNode 'call' with the sanctioned callee null->id refinement slot (pointer, the one allowed null) and is_goroutine/is_deferred. - Edge{src,dst,prov,weight} for call_graph. No emitter or wiring yet; these types are not referenced by any code path. L3/L4 fields (cfg/cdg/ddg/param_*) are deliberately not declared. --- internal/schema/v2/schema.go | 196 +++++++++++++++++++++++++++++++++++ 1 file changed, 196 insertions(+) create mode 100644 internal/schema/v2/schema.go diff --git a/internal/schema/v2/schema.go b/internal/schema/v2/schema.go new file mode 100644 index 0000000..66042b7 --- /dev/null +++ b/internal/schema/v2/schema.go @@ -0,0 +1,196 @@ +// Package v2 defines the canonical CLDK schema (v2) types that codeanalyzer-go +// emits. It is a faithful transcript of the keystone +// (designing-cldk-changes/references/canonical-schema.md) and the Go-specific +// leaf additions recorded in the repo-root CLAUDE.md § Schema decisions. +// +// The model is ONE additive containment tree +// (application → module → type → callable → body) with typed edge overlays +// laid over it (a Code Property Graph). This package declares the L1 tree plus +// the L2 refinement/edge slots; L3/L4 fields (cfg/cdg/ddg/param_*) are not +// declared here — they arrive in their own train and reuse these ids/spans +// unchanged. +// +// Conventions from the keystone, held here: +// - snake_case JSON keys everywhere, so one set of SDK models parses every +// analyzer. +// - a fact is present or absent; there is no null, EXCEPT the one sanctioned +// `callee: null` refinement slot on a call node (null at L1, backfilled at +// L2). Absence is the "no fact" encoding — `omitempty` on optional fields. +// - open-vocabulary fields (prov, tags) are plain strings. +package v2 + +// SchemaVersion is the canonical schema major this package targets. +const SchemaVersion = "2.0.0" + +// Language is the analyzer's language tag in the manifest and `can:///…`. +const Language = "go" + +// ─── Positioning ─────────────────────────────────────────────────────────── + +// Span is the one universal node attribute: where in source the node lives. +// line:col addresses and displays; bytes slices module.source in O(1). Byte +// offsets are UTF-8 offsets into module.source (Go source is natively UTF-8). +type Span struct { + // Start is [line, col] of the first byte (1-based line, 1-based col). + Start [2]int `json:"start"` + // End is [line, col] one past the last byte. + End [2]int `json:"end"` + // Bytes is [from, to) UTF-8 byte offsets into module.source. + Bytes [2]int `json:"bytes"` +} + +// ─── Root envelope ─────────────────────────────────────────────────────────── + +// Analysis is the top-level payload written to analysis.json / stdout. Its +// siblings of `application` carry the manifest; consumers read `max_level` +// rather than sniffing for keys. +type Analysis struct { + SchemaVersion string `json:"schema_version"` + Language string `json:"language"` + MaxLevel int `json:"max_level"` + Analyzer AnalyzerTag `json:"analyzer"` + Application Application `json:"application"` +} + +// AnalyzerTag identifies the producing analyzer and version. +type AnalyzerTag struct { + Name string `json:"name"` + Version string `json:"version"` +} + +// Application is the root node of the containment tree. +type Application struct { + // ID is `can:///` — the app segment disambiguates apps in one + // language. is the --app-name value (default: input dir base name). + ID string `json:"id"` + Kind string `json:"kind"` // always "application" + // SymbolTable is the L1 named map keyed by relative file path. + SymbolTable map[string]Module `json:"symbol_table"` + // CallGraph holds the L2 cross-function edges (callable → callable). Empty + // at L1. + CallGraph []Edge `json:"call_graph"` +} + +// ─── module (per-file compilation unit) ─────────────────────────────────────── + +// Module is one Go source file: the L1 per-file container. +type Module struct { + ID string `json:"id"` + Kind string `json:"kind"` // always "module" + Span Span `json:"span"` + // Package is the Go package the file belongs to. + Package string `json:"package"` + // Source is the WHOLE file's text, stored once. Every node's text slices + // from this via its span.bytes; there is no per-node `code`. + Source string `json:"source"` + // Imports are the file's import declarations. + Imports []Import `json:"imports"` + // Types is the named map of type declarations in the file. + Types map[string]Type `json:"types"` + // Functions is the named map of module-level callables (keyed by signature). + Functions map[string]Callable `json:"functions"` + // ContentHash supports incremental caching; it is not identity. + ContentHash string `json:"content_hash,omitempty"` +} + +// Import is a single import declaration. +type Import struct { + Name string `json:"name"` // imported package name + Path string `json:"path"` // import path, e.g. "hash/fnv" + Alias string `json:"alias,omitempty"` // present only when aliased + Span Span `json:"span"` +} + +// ─── type ───────────────────────────────────────────────────────────────────── + +// Type is a named Go type. Kind is the specific type-declaration shape, not a +// pile of is_* booleans. +type Type struct { + ID string `json:"id"` + Kind string `json:"kind"` // struct | interface | alias | defined + Span Span `json:"span"` + // BaseTypes are embedded type ids (struct/interface embedding — the explicit + // spine). + BaseTypes []string `json:"base_types,omitempty"` + // Interfaces are ids of interfaces this type is computed to satisfy via + // method-set matching (Go's structural implementation). + Interfaces []string `json:"interfaces,omitempty"` + // Callables are the type's methods, resolved to their receiver type + // regardless of source location. Keyed by signature. + Callables map[string]Callable `json:"callables"` + // Fields are the struct fields, keyed by name. + Fields map[string]Field `json:"fields,omitempty"` +} + +// Field is a struct field. +type Field struct { + ID string `json:"id"` + Kind string `json:"kind"` // always "field" + Type string `json:"type"` + Span Span `json:"span"` +} + +// ─── callable ───────────────────────────────────────────────────────────────── + +// Callable is a function, method, or closure. span.bytes → get_method_body = +// module.source[bytes]. +type Callable struct { + ID string `json:"id"` + Kind string `json:"kind"` // function | method | lambda + Span Span `json:"span"` + // Signature is the human-readable last path segment of ID (one signatureOf()). + Signature string `json:"signature"` + // Parameters are the ordered formal parameters. + Parameters []Parameter `json:"parameters"` + // ReturnType is the joined return type, e.g. "(int, error)". + ReturnType string `json:"return_type,omitempty"` + // ErrorChannel is the generalized error surface: populated from error-typed + // returns (those returns remain in ReturnType too). + ErrorChannel []string `json:"error_channel,omitempty"` + // Metrics is an extensible metrics map (e.g. {"cyclomatic": 3}). + Metrics map[string]int `json:"metrics,omitempty"` + // Body holds call nodes at L1 (keyed by local id), completing at L3. + Body map[string]BodyNode `json:"body"` + // Callables are nested closures / function literals, each with its own + // can:// id (replaces v1 InnerCallables). + Callables map[string]Callable `json:"callables,omitempty"` +} + +// Parameter is a single formal parameter. +type Parameter struct { + Name string `json:"name"` + Type string `json:"type"` + Span Span `json:"span"` + IsVariadic bool `json:"is_variadic,omitempty"` +} + +// ─── body nodes ──────────────────────────────────────────────────────────────── + +// BodyNode is a node inside a callable's body, keyed by its local id +// (line:col for real nodes). At L1 the only kind emitted is "call"; L3 adds the +// remaining statement kinds. +type BodyNode struct { + Kind string `json:"kind"` // "call" at L1 + Span Span `json:"span"` + // Callee is the sanctioned null→id refinement slot on a call node: null at + // L1, backfilled to a callable id at L2. A pointer so it serializes as JSON + // null (the one place null is allowed) rather than being omitted. + Callee *string `json:"callee,omitempty"` + // IsGoroutine is true when the call is preceded by `go`. + IsGoroutine bool `json:"is_goroutine,omitempty"` + // IsDeferred is true when the call is preceded by `defer` (net-new vs v1). + IsDeferred bool `json:"is_deferred,omitempty"` +} + +// ─── edges ───────────────────────────────────────────────────────────────────── + +// Edge is a typed edge overlay record. The list it lives in IS its type (there +// is no `type` field). src/dst are node ids; no dangling endpoints. +type Edge struct { + Src string `json:"src"` + Dst string `json:"dst"` + // Prov is open-vocabulary provenance, e.g. ["go/types"], ["go/types","codeql"]. + Prov []string `json:"prov,omitempty"` + // Weight accumulates when merging edges from multiple backends. + Weight int `json:"weight,omitempty"` +} From 3790749dc164b1d25d92123166d0ebe827d81faa Mon Sep 17 00:00:00 2001 From: lambertpw Date: Mon, 28 Sep 2026 22:43:42 -0400 Subject: [PATCH 04/17] feat(schema): can:// id builder and span offsets Add the Group A identity/positioning helpers in internal/schema/v2, pure functions so they are testable in isolation and reused verbatim by the deferred L3/L4 train: - can:// durable-id builders (AppID/ModuleID/TypeID/CallableID) that assemble the containment path can://////, with signatureOf() output as the last segment. - @line:col ordinal-id builders (OrdinalID/TagID/LocalID) for body nodes below the callable line. - NewSpan + Span.Slice: build a span from line/col and UTF-8 byte offsets, and slice module.source in O(1) (the replacement for the v1 per-node code field), degrading to "" on a malformed span. Unit tests reproduce the keystone worked-example id chain and cover path normalization, module-level function parents, ordinal/tag ids, and span slicing edge cases. Not wired into any emitter yet. --- internal/schema/v2/ids.go | 128 +++++++++++++++++++++++++++++++++ internal/schema/v2/ids_test.go | 93 ++++++++++++++++++++++++ 2 files changed, 221 insertions(+) create mode 100644 internal/schema/v2/ids.go create mode 100644 internal/schema/v2/ids_test.go diff --git a/internal/schema/v2/ids.go b/internal/schema/v2/ids.go new file mode 100644 index 0000000..d8d192d --- /dev/null +++ b/internal/schema/v2/ids.go @@ -0,0 +1,128 @@ +package v2 + +import "strings" + +// Group A vocabulary: the `can://` durable-id grammar (>= callable) and the +// `@line:col` ordinal-id grammar (< callable), plus span construction. These +// helpers are pure so they can be tested in isolation and reused verbatim by +// the deferred L3/L4 train (which must not re-decide id or span shape). +// +// Durable id grammar (a containment path with an app segment so multiple apps +// in one language don't collide): +// +// can:////// +// can://go/myapp/src/util.go/Hasher/Hash(string)uint64 +// +// Ordinal id grammar (statements and synthetic vertices, addressed WITHIN their +// callable): +// +// @: e.g. …/Hash(string)uint64@15:2 +// @ e.g. …/Hash(string)uint64@entry +// +// The delimiters `/`, `@`, `:` are fixed by the keystone; do not substitute. +const ( + scheme = "can://" + // idSep joins containment segments in a durable id. + idSep = "/" + // ordinalSep separates a callable id from an ordinal (line:col or @tag). + ordinalSep = "@" +) + +// AppID builds the application id: can:///. lang is fixed to the +// package Language constant so callers cannot drift it. +func AppID(app string) string { + return scheme + Language + idSep + app +} + +// ModuleID builds a module id from the application id and the module's file +// path relative to the input root (forward-slashed, no leading slash, no ".."). +func ModuleID(appID, relPath string) string { + return appID + idSep + normalizePath(relPath) +} + +// TypeID builds a type id under its module. sig is the type's signatureOf() +// output — the last (and only) segment the type contributes. +func TypeID(moduleID, sig string) string { + return moduleID + idSep + sig +} + +// CallableID builds a callable id under its parent. For a method the parent is +// its receiver Type id; for a module-level function the parent is the Module +// id. sig is the callable's signatureOf() output — the last path segment. +func CallableID(parentID, sig string) string { + return parentID + idSep + sig +} + +// OrdinalID addresses a real body node (statement / call) within its callable +// by source position: @:. Both line and col are +// required — a bare line is not unique within a callable. +func OrdinalID(callableID string, line, col int) string { + return callableID + ordinalSep + itoa(line) + ":" + itoa(col) +} + +// TagID addresses a synthetic vertex within its callable by tag: +// @ (e.g. "entry", "exit", "formal_in:0"). +func TagID(callableID, tag string) string { + return callableID + ordinalSep + tag +} + +// LocalID is the key a real body node is stored under inside body{}: the bare +// ":" ordinal (the id relative to the enclosing callable). +func LocalID(line, col int) string { + return itoa(line) + ":" + itoa(col) +} + +// NewSpan builds a Span from 1-based line/col positions and the [from, to) UTF-8 +// byte offsets into module.source. Byte offsets are what make source slicing +// O(1); line:col is what addresses and displays. +func NewSpan(startLine, startCol, endLine, endCol, fromByte, toByte int) Span { + return Span{ + Start: [2]int{startLine, startCol}, + End: [2]int{endLine, endCol}, + Bytes: [2]int{fromByte, toByte}, + } +} + +// Slice returns the node's source text: module source sliced by the span's byte +// offsets. It is the O(1) replacement for the v1 per-node `code` field. Out-of- +// range or inverted offsets yield "" rather than panicking, so a malformed span +// degrades gracefully. +func (s Span) Slice(source string) string { + from, to := s.Bytes[0], s.Bytes[1] + if from < 0 || to > len(source) || from > to { + return "" + } + return source[from:to] +} + +// normalizePath makes a relative file path safe for a can:// segment: forward +// slashes, no leading slash. Callers are responsible for passing a path already +// relative to the input root (no ".."). +func normalizePath(p string) string { + p = strings.ReplaceAll(p, "\\", "/") + return strings.TrimPrefix(p, "/") +} + +// itoa is a tiny non-allocating-path integer formatter for id assembly, kept +// local so this file has no fmt dependency for its hot path. +func itoa(n int) string { + if n == 0 { + return "0" + } + neg := n < 0 + if neg { + n = -n + } + var buf [20]byte + i := len(buf) + for n > 0 { + i-- + buf[i] = byte('0' + n%10) + n /= 10 + } + if neg { + i-- + buf[i] = '-' + } + return string(buf[i:]) +} diff --git a/internal/schema/v2/ids_test.go b/internal/schema/v2/ids_test.go new file mode 100644 index 0000000..f7ec330 --- /dev/null +++ b/internal/schema/v2/ids_test.go @@ -0,0 +1,93 @@ +package v2 + +import "testing" + +func TestAppID(t *testing.T) { + if got, want := AppID("myapp"), "can://go/myapp"; got != want { + t.Errorf("AppID = %q, want %q", got, want) + } +} + +func TestDurableIDChain(t *testing.T) { + // Reproduces the keystone worked example, segment by segment. + app := AppID("myapp") + mod := ModuleID(app, "src/util.go") + typ := TypeID(mod, "Hasher") + call := CallableID(typ, "Hash(string)uint64") + + cases := []struct{ got, want string }{ + {mod, "can://go/myapp/src/util.go"}, + {typ, "can://go/myapp/src/util.go/Hasher"}, + {call, "can://go/myapp/src/util.go/Hasher/Hash(string)uint64"}, + } + for _, c := range cases { + if c.got != c.want { + t.Errorf("id = %q, want %q", c.got, c.want) + } + } +} + +func TestModuleIDNormalizesPath(t *testing.T) { + app := AppID("myapp") + // Backslashes become forward slashes; a leading slash is trimmed. + if got, want := ModuleID(app, "\\pkg\\a.go"), "can://go/myapp/pkg/a.go"; got != want { + t.Errorf("ModuleID = %q, want %q", got, want) + } + if got, want := ModuleID(app, "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/pkg/a.go"), "can://go/myapp/pkg/a.go"; got != want { + t.Errorf("ModuleID = %q, want %q", got, want) + } +} + +func TestModuleLevelFunctionParent(t *testing.T) { + // A module-level function hangs off the module id, not a type. + mod := ModuleID(AppID("myapp"), "src/util.go") + fn := CallableID(mod, "New64()") + if want := "can://go/myapp/src/util.go/New64()"; fn != want { + t.Errorf("CallableID = %q, want %q", fn, want) + } +} + +func TestOrdinalAndTagIDs(t *testing.T) { + call := "can://go/myapp/src/util.go/Hasher/Hash(string)uint64" + if got, want := OrdinalID(call, 15, 2), call+"@15:2"; got != want { + t.Errorf("OrdinalID = %q, want %q", got, want) + } + if got, want := TagID(call, "entry"), call+"@entry"; got != want { + t.Errorf("TagID = %q, want %q", got, want) + } + if got, want := LocalID(16, 2), "16:2"; got != want { + t.Errorf("LocalID = %q, want %q", got, want) + } +} + +func TestSpanSlice(t *testing.T) { + src := "package util\nfunc f() {}\n" + // bytes [13,25) covers "func f() {}\n" + s := NewSpan(2, 1, 2, 12, 13, 25) + if got, want := s.Slice(src), "func f() {}\n"; got != want { + t.Errorf("Slice = %q, want %q", got, want) + } +} + +func TestSpanSliceOutOfRangeIsEmpty(t *testing.T) { + src := "abc" + cases := []Span{ + {Bytes: [2]int{0, 99}}, // to past end + {Bytes: [2]int{-1, 2}}, // negative from + {Bytes: [2]int{2, 1}}, // inverted + } + for i, s := range cases { + if got := s.Slice(src); got != "" { + t.Errorf("case %d: Slice = %q, want empty", i, got) + } + } +} + +func TestItoa(t *testing.T) { + cases := map[int]string{0: "0", 7: "7", 42: "42", 1000: "1000", -3: "-3"} + for in, want := range cases { + if got := itoa(in); got != want { + t.Errorf("itoa(%d) = %q, want %q", in, got, want) + } + } +} From d3affba6715791e1aabc6b7b729bf63120d33bd1 Mon Sep 17 00:00:00 2001 From: lambertpw Date: Mon, 28 Sep 2026 22:54:29 -0400 Subject: [PATCH 05/17] feat(emit): source line-to-byte-offset indexer for v2 spans Add the v2emit package with lineIndex: reads a file's source once and maps 1-based (line, col) positions to UTF-8 byte offsets, the one new datum the v2 span needs. go/token columns are byte columns, so no rune/byte reconciliation is required. This keeps the v1 model diff-free (per the migration spec): the emitter derives span.bytes from the existing line/col positions instead of threading offsets through internal/schema. lineEndOffset/lineLen give a line-only v1 position (callables/types carry no columns) a precise end. Positions outside the file clamp to [0, len(source)] rather than panicking. Tests cover line-start tabulation, offset lookup, clamping, the last-line-without-newline case, and empty source. Emitter tree walk lands next; nothing wired yet. --- internal/syntactic_analysis/v2emit/offsets.go | 94 +++++++++++++++++++ .../syntactic_analysis/v2emit/offsets_test.go | 69 ++++++++++++++ 2 files changed, 163 insertions(+) create mode 100644 internal/syntactic_analysis/v2emit/offsets.go create mode 100644 internal/syntactic_analysis/v2emit/offsets_test.go diff --git a/internal/syntactic_analysis/v2emit/offsets.go b/internal/syntactic_analysis/v2emit/offsets.go new file mode 100644 index 0000000..e9d3160 --- /dev/null +++ b/internal/syntactic_analysis/v2emit/offsets.go @@ -0,0 +1,94 @@ +// Package v2emit transforms the v1 in-memory model (internal/schema) into the +// v2 canonical schema (internal/schema/v2). It is an EMISSION layer only: it +// computes no new facts beyond source byte offsets, it just re-serializes what +// the parser/resolver already produced into the additive-tree shape. +// +// The one genuinely new datum is span byte offsets. The v1 model records +// line/col but not UTF-8 byte offsets (and only line-level positions for +// callables/types). Rather than change the v1 model, the emitter reads each +// file's source once and derives byte offsets from the existing line/col +// positions via a lineIndex. That same source string is emitted once per module +// as module.source, off which every node's text slices. +package v2emit + +// lineIndex maps 1-based (line, col) positions in a source file to UTF-8 byte +// offsets. Built once per file. Columns are byte columns within the line, which +// is what go/token.Position reports (Position.Column counts bytes, not runes), +// so no rune/byte reconciliation is needed here. +type lineIndex struct { + source string + // lineStart[i] is the byte offset where 1-based line (i+1) begins. + lineStart []int +} + +// newLineIndex builds a lineIndex over the given source text. +func newLineIndex(source string) *lineIndex { + // Line 1 starts at offset 0; each '\n' begins the next line at the byte + // immediately after it. + starts := make([]int, 0, 1+countByte(source, '\n')) + starts = append(starts, 0) + for i := 0; i < len(source); i++ { + if source[i] == '\n' { + starts = append(starts, i+1) + } + } + return &lineIndex{source: source, lineStart: starts} +} + +// offset converts a 1-based (line, col) position to a byte offset into source. +// col is a 1-based byte column (go/token semantics). Positions outside the +// file's range are clamped to [0, len(source)] so a bad position yields a +// harmless in-range offset rather than a panic. +func (li *lineIndex) offset(line, col int) int { + if len(li.lineStart) == 0 { + return 0 + } + if line < 1 { + line = 1 + } + if line > len(li.lineStart) { + line = len(li.lineStart) + } + base := li.lineStart[line-1] + if col < 1 { + col = 1 + } + off := base + (col - 1) + if off > len(li.source) { + off = len(li.source) + } + return off +} + +// lineEndOffset returns the byte offset of the end of a 1-based line: the +// position just before its terminating '\n' (or end of file for the last line). +// Used to give a line-only position (no column in the v1 model) a precise end. +func (li *lineIndex) lineEndOffset(line int) int { + if line < 1 || len(li.lineStart) == 0 { + return 0 + } + if line >= len(li.lineStart) { + return len(li.source) + } + // Next line begins one byte after this line's '\n'; back up over the '\n'. + return li.lineStart[line] - 1 +} + +// lineLen returns the number of bytes on a 1-based line, excluding the newline. +// It gives a line-only end position a byte column (used to synthesize end.col). +func (li *lineIndex) lineLen(line int) int { + if line < 1 || line > len(li.lineStart) { + return 0 + } + return li.lineEndOffset(line) - li.lineStart[line-1] +} + +func countByte(s string, b byte) int { + n := 0 + for i := 0; i < len(s); i++ { + if s[i] == b { + n++ + } + } + return n +} diff --git a/internal/syntactic_analysis/v2emit/offsets_test.go b/internal/syntactic_analysis/v2emit/offsets_test.go new file mode 100644 index 0000000..69b61b9 --- /dev/null +++ b/internal/syntactic_analysis/v2emit/offsets_test.go @@ -0,0 +1,69 @@ +package v2emit + +import "testing" + +const sample = "package util\n" + // line 1: bytes 0..12 (\n at 12) + "\n" + // line 2: byte 13 (\n at 13) + "func f() {}\n" + // line 3: bytes 14..25 (\n at 25) + "var x int" // line 4: bytes 26..34, no trailing newline + +func TestLineStarts(t *testing.T) { + li := newLineIndex(sample) + want := []int{0, 13, 14, 26} + if len(li.lineStart) != len(want) { + t.Fatalf("lineStart len = %d, want %d (%v)", len(li.lineStart), len(want), li.lineStart) + } + for i, w := range want { + if li.lineStart[i] != w { + t.Errorf("lineStart[%d] = %d, want %d", i, li.lineStart[i], w) + } + } +} + +func TestOffset(t *testing.T) { + li := newLineIndex(sample) + // line 3 ("func f() {}") col 1 → byte 14; col 6 ('f') → 19 + if got := li.offset(3, 1); got != 14 { + t.Errorf("offset(3,1) = %d, want 14", got) + } + if got := li.offset(3, 6); got != 19 { + t.Errorf("offset(3,6) = %d, want 19", got) + } + // The slice from a callable start to its end reproduces the source text. + start := li.offset(3, 1) + end := li.lineEndOffset(3) + if got := sample[start:end]; got != "func f() {}" { + t.Errorf("slice = %q, want %q", got, "func f() {}") + } +} + +func TestOffsetClamping(t *testing.T) { + li := newLineIndex(sample) + if got := li.offset(0, 0); got != 0 { + t.Errorf("offset(0,0) = %d, want 0 (clamped)", got) + } + if got := li.offset(999, 999); got != len(sample) { + t.Errorf("offset(999,999) = %d, want %d (clamped to EOF)", got, len(sample)) + } +} + +func TestLineEndAndLen(t *testing.T) { + li := newLineIndex(sample) + // last line has no trailing newline → ends at EOF + if got := li.lineEndOffset(4); got != len(sample) { + t.Errorf("lineEndOffset(4) = %d, want %d", got, len(sample)) + } + if got := li.lineLen(3); got != len("func f() {}") { + t.Errorf("lineLen(3) = %d, want %d", got, len("func f() {}")) + } + if got := li.lineLen(4); got != len("var x int") { + t.Errorf("lineLen(4) = %d, want %d", got, len("var x int")) + } +} + +func TestEmptySource(t *testing.T) { + li := newLineIndex("") + if got := li.offset(1, 1); got != 0 { + t.Errorf("offset on empty source = %d, want 0", got) + } +} From 79433931f479d789ae82ef854238374904ae7ace Mon Sep 17 00:00:00 2001 From: lambertpw Date: Mon, 28 Sep 2026 22:56:46 -0400 Subject: [PATCH 06/17] feat(emit): v2 L1 emitter over the v1 model Emit() transforms the v1 GoApplication into the v2 Analysis payload: a pure re-serialization walking application -> module -> type -> callable -> body, computing no new facts beyond span byte offsets. - Module carries source once (read from projectDir/relPath); every node span slices off it via the lineIndex-derived byte offsets. - can:// ids assembled from the id builder, signatureOf() as last segment; methods resolve under their receiver type's callables{}. - type.kind collapses is_interface (struct|interface); base_types kept; error-typed returns lift into error_channel; cyclomatic into metrics; closures nest in callables{}. - body holds one 'call' node per call site keyed by line:col, with the callee null->id slot carried straight from the v1 pointer. - call_graph renamed to {src,dst,prov,weight}. Also drop omitempty from BodyNode.Callee so the sanctioned refinement slot always serializes (null at L1, id at L2) as the keystone requires. Tests build the greeter fixture through the real symbol-table builder and assert the envelope, module source/ids, span slicing, call-node bodies, callee-null serialization, dropped is_constructor_call, and type/method kinds. Not wired to the CLI yet. --- internal/schema/v2/schema.go | 7 +- internal/syntactic_analysis/v2emit/emit.go | 284 ++++++++++++++++++ .../syntactic_analysis/v2emit/emit_test.go | 185 ++++++++++++ 3 files changed, 473 insertions(+), 3 deletions(-) create mode 100644 internal/syntactic_analysis/v2emit/emit.go create mode 100644 internal/syntactic_analysis/v2emit/emit_test.go diff --git a/internal/schema/v2/schema.go b/internal/schema/v2/schema.go index 66042b7..c3acb27 100644 --- a/internal/schema/v2/schema.go +++ b/internal/schema/v2/schema.go @@ -173,9 +173,10 @@ type BodyNode struct { Kind string `json:"kind"` // "call" at L1 Span Span `json:"span"` // Callee is the sanctioned null→id refinement slot on a call node: null at - // L1, backfilled to a callable id at L2. A pointer so it serializes as JSON - // null (the one place null is allowed) rather than being omitted. - Callee *string `json:"callee,omitempty"` + // L1, backfilled to a callable id at L2. A pointer with NO omitempty so it + // always serializes — as JSON null at L1 (the one place null is allowed), + // as an id once backfilled. The keystone requires the field to be present. + Callee *string `json:"callee"` // IsGoroutine is true when the call is preceded by `go`. IsGoroutine bool `json:"is_goroutine,omitempty"` // IsDeferred is true when the call is preceded by `defer` (net-new vs v1). diff --git a/internal/syntactic_analysis/v2emit/emit.go b/internal/syntactic_analysis/v2emit/emit.go new file mode 100644 index 0000000..d2501d2 --- /dev/null +++ b/internal/syntactic_analysis/v2emit/emit.go @@ -0,0 +1,284 @@ +package v2emit + +import ( + "os" + "path/filepath" + "sort" + + "github.com/codellm-devkit/codeanalyzer-go/internal/schema" + v2 "github.com/codellm-devkit/codeanalyzer-go/internal/schema/v2" +) + +// Emit transforms the v1 GoApplication into the v2 Analysis payload. appName is +// the application anchor (--app-name); projectDir is the absolute input root, +// used to read each module's source for span slicing. maxLevel records how +// deeply the tree was populated (1 = symbol table, 2 = + call_graph). +// +// This is a pure re-serialization: every fact comes from the v1 model, except +// span byte offsets, which are derived from source via lineIndex. +func Emit(app *schema.GoApplication, appName, projectDir string, maxLevel int) *v2.Analysis { + appID := v2.AppID(appName) + + out := &v2.Analysis{ + SchemaVersion: v2.SchemaVersion, + Language: v2.Language, + MaxLevel: maxLevel, + Analyzer: v2.AnalyzerTag{Name: "codeanalyzer-go"}, + Application: v2.Application{ + ID: appID, + Kind: "application", + SymbolTable: make(map[string]v2.Module, len(app.SymbolTable)), + CallGraph: emitCallGraph(app.CallGraph), + }, + } + + // Deterministic module order does not affect the JSON map, but iterating + // sorted keeps any incidental ordering (and logs) stable across runs. + for _, relPath := range sortedFileKeys(app.SymbolTable) { + file := app.SymbolTable[relPath] + out.Application.SymbolTable[relPath] = emitModule(appID, projectDir, relPath, file) + } + return out +} + +// emitModule builds a v2 module from a v1 GoFile, reading the file's source so +// spans can slice off it. A file that cannot be read still emits its structure +// (with empty source and zeroed byte spans) rather than dropping the module. +func emitModule(appID, projectDir, relPath string, file schema.GoFile) v2.Module { + source := readSource(projectDir, relPath) + li := newLineIndex(source) + modID := v2.ModuleID(appID, relPath) + + mod := v2.Module{ + ID: modID, + Kind: "module", + Span: wholeFileSpan(li, source), + Package: file.PackageName, + Source: source, + Imports: emitImports(li, file.Imports), + Types: make(map[string]v2.Type, len(file.Types)), + Functions: make(map[string]v2.Callable, len(file.Functions)), + } + if file.ContentHash != nil { + mod.ContentHash = *file.ContentHash + } + + for name, t := range file.Types { + mod.Types[name] = emitType(modID, li, t) + } + for sig, fn := range file.Functions { + mod.Functions[sig] = emitCallable(modID, li, fn) + } + return mod +} + +// emitType maps a v1 GoType to a v2 type node, collapsing is_interface into the +// kind and splitting embedding (base_types) from computed satisfaction. Methods +// resolve into the type's callables{}. +func emitType(modID string, li *lineIndex, t schema.GoType) v2.Type { + typeID := v2.TypeID(modID, t.Signature) + out := v2.Type{ + ID: typeID, + Kind: typeKind(t), + Span: lineSpan(li, t.StartLine, t.EndLine), + BaseTypes: nonEmpty(t.BaseTypes), + Callables: make(map[string]v2.Callable, len(t.Methods)), + } + if len(t.Fields) > 0 { + out.Fields = make(map[string]v2.Field, len(t.Fields)) + for _, f := range t.Fields { + out.Fields[f.Name] = v2.Field{ + ID: typeID + "/" + f.Name, + Kind: "field", + Type: f.Type, + Span: lineSpan(li, f.StartLine, f.EndLine), + } + } + } + for sig, m := range t.Methods { + out.Callables[sig] = emitCallable(typeID, li, m) + } + return out +} + +// emitCallable maps a v1 GoCallable to a v2 callable node, including its body +// (call nodes) and nested closures. +func emitCallable(parentID string, li *lineIndex, c schema.GoCallable) v2.Callable { + callID := v2.CallableID(parentID, c.Signature) + out := v2.Callable{ + ID: callID, + Kind: callableKind(c), + Span: lineSpan(li, c.StartLine, c.EndLine), + Signature: c.Signature, + Parameters: emitParams(li, c.Parameters), + ReturnType: c.ReturnType, + Body: emitBody(callID, li, c.CallSites), + } + if ch := errorChannel(c.ReturnTypes); len(ch) > 0 { + out.ErrorChannel = ch + } + if c.CyclomaticComplexity > 0 { + out.Metrics = map[string]int{"cyclomatic": c.CyclomaticComplexity} + } + if len(c.InnerCallables) > 0 { + out.Callables = make(map[string]v2.Callable, len(c.InnerCallables)) + for sig, inner := range c.InnerCallables { + out.Callables[sig] = emitCallable(callID, li, inner) + } + } + return out +} + +// emitBody builds the L1 body map: one `call` node per call site, keyed by its +// line:col local id, with the sanctioned callee null->id refinement slot. +func emitBody(callableID string, li *lineIndex, sites []schema.GoCallsite) map[string]v2.BodyNode { + body := make(map[string]v2.BodyNode, len(sites)) + for _, cs := range sites { + local := v2.LocalID(cs.StartLine, cs.StartColumn) + body[local] = v2.BodyNode{ + Kind: "call", + Span: colSpan(li, cs.StartLine, cs.StartColumn, cs.EndLine, cs.EndColumn), + Callee: cs.CalleeSignature, // *string: null at L1, backfilled at L2 + IsGoroutine: cs.IsGoroutine, + } + } + return body +} + +func emitParams(li *lineIndex, params []schema.GoParameter) []v2.Parameter { + out := make([]v2.Parameter, 0, len(params)) + for _, p := range params { + out = append(out, v2.Parameter{ + Name: p.Name, + Type: p.Type, + Span: lineSpan(li, p.StartLine, p.EndLine), + IsVariadic: p.IsVariadic, + }) + } + return out +} + +func emitImports(li *lineIndex, imports []schema.GoImport) []v2.Import { + out := make([]v2.Import, 0, len(imports)) + for _, imp := range imports { + out = append(out, v2.Import{ + Name: packageBase(imp.Module), + Path: imp.Module, + Alias: imp.Alias, + Span: lineSpan(li, imp.StartLine, imp.EndLine), + }) + } + return out +} + +// emitCallGraph maps v1 identity-only edges to the v2 {src,dst,prov,weight} +// shape (rename of source/target/provenance). +func emitCallGraph(edges []schema.GoCallEdge) []v2.Edge { + out := make([]v2.Edge, 0, len(edges)) + for _, e := range edges { + out = append(out, v2.Edge{ + Src: e.Source, + Dst: e.Target, + Prov: nonEmpty(e.Provenance), + Weight: e.Weight, + }) + } + return out +} + +// ─── kind mapping ──────────────────────────────────────────────────────────── + +// typeKind collapses the v1 is_interface boolean into the v2 kind vocabulary. +// v1 only distinguishes struct vs interface; alias/defined await richer v1 +// facts, so a non-interface type maps to struct for now. +func typeKind(t schema.GoType) string { + if t.IsInterface { + return "interface" + } + return "struct" +} + +// callableKind classifies a callable: a non-empty receiver type marks a method. +func callableKind(c schema.GoCallable) string { + if c.ReceiverType != "" { + return "method" + } + return "function" +} + +// errorChannel extracts error-typed returns into the generalized error surface. +func errorChannel(returnTypes []string) []string { + var ch []string + for _, rt := range returnTypes { + if rt == "error" { + ch = append(ch, rt) + } + } + return ch +} + +// ─── span construction ───────────────────────────────────────────────────────── + +// wholeFileSpan is the module node's span: the entire file. +func wholeFileSpan(li *lineIndex, source string) v2.Span { + lastLine := len(li.lineStart) + if lastLine == 0 { + lastLine = 1 + } + return v2.NewSpan(1, 1, lastLine, li.lineLen(lastLine)+1, 0, len(source)) +} + +// lineSpan builds a span from a line-only v1 position: start at column 1, end at +// the last line's end. Byte offsets bound the whole line range. +func lineSpan(li *lineIndex, startLine, endLine int) v2.Span { + if endLine < startLine { + endLine = startLine + } + endCol := li.lineLen(endLine) + 1 + return v2.NewSpan( + startLine, 1, endLine, endCol, + li.offset(startLine, 1), li.lineEndOffset(endLine), + ) +} + +// colSpan builds a span from a v1 position that carries columns (call sites). +func colSpan(li *lineIndex, startLine, startCol, endLine, endCol int) v2.Span { + return v2.NewSpan( + startLine, startCol, endLine, endCol, + li.offset(startLine, startCol), li.offset(endLine, endCol), + ) +} + +// ─── helpers ───────────────────────────────────────────────────────────────── + +func readSource(projectDir, relPath string) string { + data, err := os.ReadFile(filepath.Join(projectDir, relPath)) + if err != nil { + return "" + } + return string(data) +} + +func sortedFileKeys(m map[string]schema.GoFile) []string { + keys := make([]string, 0, len(m)) + for k := range m { + keys = append(keys, k) + } + sort.Strings(keys) + return keys +} + +// packageBase returns the last path segment of an import path as the package +// name, e.g. "hash/fnv" -> "fnv". +func packageBase(importPath string) string { + return filepath.Base(importPath) +} + +// nonEmpty returns s unchanged, or nil if empty, so omitempty drops the field +// rather than emitting an empty array (absence = no fact). +func nonEmpty(s []string) []string { + if len(s) == 0 { + return nil + } + return s +} diff --git a/internal/syntactic_analysis/v2emit/emit_test.go b/internal/syntactic_analysis/v2emit/emit_test.go new file mode 100644 index 0000000..deb8a66 --- /dev/null +++ b/internal/syntactic_analysis/v2emit/emit_test.go @@ -0,0 +1,185 @@ +package v2emit + +import ( + "encoding/json" + "path/filepath" + "runtime" + "strings" + "testing" + + "github.com/codellm-devkit/codeanalyzer-go/internal/schema" + v2 "github.com/codellm-devkit/codeanalyzer-go/internal/schema/v2" + "github.com/codellm-devkit/codeanalyzer-go/internal/syntactic_analysis" +) + +// greeterDir returns the absolute path to the greeter testdata fixture. +func greeterDir(t *testing.T) string { + t.Helper() + _, thisFile, _, _ := runtime.Caller(0) + abs, err := filepath.Abs(filepath.Join(filepath.Dir(thisFile), "..", "..", "..", "testdata", "greeter")) + if err != nil { + t.Fatalf("resolving fixture dir: %v", err) + } + return abs +} + +// buildGreeter runs the real symbol-table builder over the greeter fixture. +func buildGreeter(t *testing.T) (*schema.GoApplication, string) { + t.Helper() + dir := greeterDir(t) + b := syntactic_analysis.NewSymbolTableBuilder(dir) + st, err := b.Build(nil, true) + if err != nil { + t.Fatalf("symbol table build failed: %v", err) + } + return &schema.GoApplication{SymbolTable: st, CallGraph: []schema.GoCallEdge{}}, dir +} + +func TestEmit_Envelope(t *testing.T) { + app, dir := buildGreeter(t) + out := Emit(app, "greeter", dir, 1) + + if out.SchemaVersion != "2.0.0" { + t.Errorf("schema_version = %q, want 2.0.0", out.SchemaVersion) + } + if out.Language != "go" { + t.Errorf("language = %q, want go", out.Language) + } + if out.MaxLevel != 1 { + t.Errorf("max_level = %d, want 1", out.MaxLevel) + } + if out.Application.ID != "can://go/greeter" { + t.Errorf("application.id = %q, want can://go/greeter", out.Application.ID) + } + if out.Application.Kind != "application" { + t.Errorf("application.kind = %q, want application", out.Application.Kind) + } + if len(out.Application.SymbolTable) == 0 { + t.Fatal("symbol_table is empty") + } +} + +func TestEmit_ModuleSourceAndIDs(t *testing.T) { + app, dir := buildGreeter(t) + out := Emit(app, "greeter", dir, 1) + + mod, ok := out.Application.SymbolTable["main.go"] + if !ok { + t.Fatalf("main.go not in symbol_table; keys=%v", keys(out.Application.SymbolTable)) + } + if mod.Kind != "module" { + t.Errorf("module.kind = %q, want module", mod.Kind) + } + if mod.ID != "can://go/greeter/main.go" { + t.Errorf("module.id = %q, want can://go/greeter/main.go", mod.ID) + } + if !strings.HasPrefix(mod.Source, "package main") { + t.Errorf("module.source should hold the file text; got prefix %q", head(mod.Source, 20)) + } + // The module span slices the whole file. + if got := mod.Span.Slice(mod.Source); got != mod.Source { + t.Errorf("module span does not cover whole source (%d vs %d bytes)", len(got), len(mod.Source)) + } +} + +func TestEmit_CallableBodyAndSpanSlice(t *testing.T) { + app, dir := buildGreeter(t) + out := Emit(app, "greeter", dir, 1) + + mod := out.Application.SymbolTable["main.go"] + var main v2.Callable + found := false + for _, fn := range mod.Functions { + if fn.Signature != "" && strings.HasSuffix(fn.Signature, ".main") { + main, found = fn, true + } + } + if !found { + t.Fatalf("main function not found; functions=%v", keys(mod.Functions)) + } + if main.Kind != "function" { + t.Errorf("main.kind = %q, want function", main.Kind) + } + // Its span should slice to source beginning with "func main". + if body := main.Span.Slice(mod.Source); !strings.HasPrefix(body, "func main") { + t.Errorf("main span slice should start with 'func main'; got %q", head(body, 20)) + } + // main() has call sites → body call nodes, each with callee null at L1. + if len(main.Body) == 0 { + t.Fatal("main should have call nodes in its body") + } + for local, node := range main.Body { + if node.Kind != "call" { + t.Errorf("body[%s].kind = %q, want call", local, node.Kind) + } + if node.Callee != nil { + t.Errorf("body[%s].callee should be null at L1; got %q", local, *node.Callee) + } + if !strings.Contains(local, ":") { + t.Errorf("body key %q should be line:col", local) + } + } +} + +func TestEmit_CalleeSerializesAsNull(t *testing.T) { + app, dir := buildGreeter(t) + out := Emit(app, "greeter", dir, 1) + data, err := json.Marshal(out) + if err != nil { + t.Fatalf("marshal failed: %v", err) + } + if !strings.Contains(string(data), `"callee":null`) { + t.Error("expected at least one call node to serialize callee as null at L1") + } + // The v1-only keys must be gone. + for _, banned := range []string{`"symbol_table"`, `"call_graph"`} { + if !strings.Contains(string(data), banned) { + t.Errorf("expected v2 payload to contain %s", banned) + } + } + if strings.Contains(string(data), `"is_constructor_call"`) { + t.Error("v2 payload should not contain dropped is_constructor_call") + } +} + +func TestEmit_TypeKindAndMethods(t *testing.T) { + app, dir := buildGreeter(t) + out := Emit(app, "greeter", dir, 1) + + // Find the Greeter type wherever it lives; assert struct kind + methods. + var found bool + for _, mod := range out.Application.SymbolTable { + for name, ty := range mod.Types { + if ty.Kind != "struct" && ty.Kind != "interface" { + t.Errorf("type %s kind = %q, want struct|interface", name, ty.Kind) + } + if strings.Contains(ty.ID, "://") == false { + t.Errorf("type %s id should be a can:// id; got %q", name, ty.ID) + } + for sig, m := range ty.Callables { + if m.Kind != "method" { + t.Errorf("callable %s under type %s kind = %q, want method", sig, name, m.Kind) + } + found = true + } + } + } + if !found { + t.Skip("no type methods in fixture; envelope/module coverage still asserted elsewhere") + } +} + +func head(s string, n int) string { + if len(s) < n { + return s + } + return s[:n] +} + +func keys[V any](m map[string]V) []string { + out := make([]string, 0, len(m)) + for k := range m { + out = append(out, k) + } + return out +} From c94d44173c51717da26ed62292851c5d397406eb Mon Sep 17 00:00:00 2001 From: lambertpw Date: Mon, 28 Sep 2026 23:01:08 -0400 Subject: [PATCH 07/17] feat(cli): --analysis-schema 2 selects v2 output Wire the v2 emitter behind an opt-in flag. Default stays the v1 shape so the current SDK keeps parsing output during the migration; --analysis- schema 2 emits the canonical v2 tree. - Add options.SchemaVersion (1 default, 2 v2) and AnalyzerVersion. - core.RenderJSON centralizes the v1/v2 marshal decision so the stdout and file paths agree; WriteOutput now takes opts and routes through it. An unknown schema returns a non-zero error, never a silent fallback. - --analysis-schema flag; --app-name/version threaded into the v2 manifest (v2 surfaces the analyzer version, which v1 never carried). Update WriteOutput callers (now opts-based) in core tests. CLI tests assert the default is v1 (top-level symbol_table, no schema_version), that --analysis-schema 2 emits schema_version 2.0.0 with application.id can://go/greeter, and that an unknown schema errors. --- cmd/codeanalyzer/main.go | 83 ++++++++++--------- cmd/codeanalyzer/main_test.go | 63 ++++++++++++++ internal/core/analyzer.go | 49 +++++++---- internal/core/analyzer_test.go | 8 +- internal/core/testsetup_test.go | 3 +- internal/options/options.go | 6 ++ internal/syntactic_analysis/v2emit/emit.go | 7 +- .../syntactic_analysis/v2emit/emit_test.go | 10 +-- 8 files changed, 162 insertions(+), 67 deletions(-) diff --git a/cmd/codeanalyzer/main.go b/cmd/codeanalyzer/main.go index 49bda45..5b0b40f 100644 --- a/cmd/codeanalyzer/main.go +++ b/cmd/codeanalyzer/main.go @@ -5,7 +5,6 @@ package main import ( - "encoding/json" "fmt" "os" "path/filepath" @@ -32,24 +31,25 @@ func main() { func rootCmd() *cobra.Command { var ( - inputPath string - outputDir string - format string - emit string - appName string - level int - targetFiles []string - skipTests bool - eager bool - cacheDir string - jobs int - useCodeQL bool - verbosity int - showVersion bool - neo4jURI string - neo4jUser string - neo4jPassword string - neo4jDatabase string + inputPath string + outputDir string + format string + emit string + appName string + level int + analysisSchema int + targetFiles []string + skipTests bool + eager bool + cacheDir string + jobs int + useCodeQL bool + verbosity int + showVersion bool + neo4jURI string + neo4jUser string + neo4jPassword string + neo4jDatabase string ) cmd := &cobra.Command{ @@ -107,23 +107,25 @@ via CLDK(language="go").analysis(project_path=...).`, } opts := options.AnalysisOptions{ - InputPath: inputPath, - OutputDir: outputDir, - Format: format, - Emit: options.EmitTarget(emit), - AppName: appName, - Level: options.AnalysisLevel(level), - TargetFiles: targetFiles, - SkipTests: skipTests, - Eager: eager, - CacheDir: cacheDir, - Jobs: jobs, - UseCodeQL: useCodeQL, - Verbose: verbosity > 0, - Neo4jURI: firstNonEmpty(neo4jURI, os.Getenv("NEO4J_URI")), - Neo4jUser: firstNonEmpty(neo4jUser, os.Getenv("NEO4J_USERNAME"), "neo4j"), - Neo4jPassword: firstNonEmpty(neo4jPassword, os.Getenv("NEO4J_PASSWORD"), "neo4j"), - Neo4jDatabase: firstNonEmpty(neo4jDatabase, os.Getenv("NEO4J_DATABASE")), + InputPath: inputPath, + OutputDir: outputDir, + Format: format, + Emit: options.EmitTarget(emit), + AppName: appName, + Level: options.AnalysisLevel(level), + SchemaVersion: analysisSchema, + AnalyzerVersion: version, + TargetFiles: targetFiles, + SkipTests: skipTests, + Eager: eager, + CacheDir: cacheDir, + Jobs: jobs, + UseCodeQL: useCodeQL, + Verbose: verbosity > 0, + Neo4jURI: firstNonEmpty(neo4jURI, os.Getenv("NEO4J_URI")), + Neo4jUser: firstNonEmpty(neo4jUser, os.Getenv("NEO4J_USERNAME"), "neo4j"), + Neo4jPassword: firstNonEmpty(neo4jPassword, os.Getenv("NEO4J_PASSWORD"), "neo4j"), + Neo4jDatabase: firstNonEmpty(neo4jDatabase, os.Getenv("NEO4J_DATABASE")), } analyzer := core.New(opts) @@ -133,16 +135,17 @@ via CLDK(language="go").analysis(project_path=...).`, } // When no --output dir is given, write JSON to cobra's output - // writer so tests can capture it via cmd.SetOut. + // writer so tests can capture it via cmd.SetOut. Both paths render + // through core so the v1/v2 schema selection stays in one place. if outputDir == "" { - data, err := json.Marshal(app) + data, err := core.RenderJSON(app, opts) if err != nil { return err } _, err = cmd.OutOrStdout().Write(data) return err } - return core.WriteOutput(app, outputDir, format) + return core.WriteOutput(app, opts) }, } @@ -155,6 +158,8 @@ via CLDK(language="go").analysis(project_path=...).`, "Application anchor name for can:// ids and Neo4j :Application (default: input dir name)") f.IntVarP(&level, "analysis-level", "a", 1, "Analysis level: 1=symbol table only, 2=+resolver call graph") + f.IntVar(&analysisSchema, "analysis-schema", 1, + "Output schema major: 1=legacy v1 shape (default), 2=canonical v2 tree") f.StringSliceVarP(&targetFiles, "target-files", "t", nil, "Restrict analysis to specific files (incremental mode)") f.BoolVar(&skipTests, "skip-tests", true, "Skip *_test.go files") diff --git a/cmd/codeanalyzer/main_test.go b/cmd/codeanalyzer/main_test.go index ea5ff3b..3843261 100644 --- a/cmd/codeanalyzer/main_test.go +++ b/cmd/codeanalyzer/main_test.go @@ -196,6 +196,69 @@ func TestRootCmd_Level2ProducesCallGraph(t *testing.T) { } } +// ── --analysis-schema ────────────────────────────────────────────────────────── + +func TestRootCmd_DefaultSchemaIsV1(t *testing.T) { + td := cliTestdataDir() + out, _, err := runCmd("--input", filepath.Join(td, "greeter"), "--cache-dir", t.TempDir()) + if err != nil { + t.Fatalf("command failed: %v", err) + } + var v map[string]interface{} + if jsonErr := json.Unmarshal([]byte(out), &v); jsonErr != nil { + t.Fatalf("stdout is not valid JSON: %v", jsonErr) + } + if _, ok := v["symbol_table"]; !ok { + t.Error("default output should be the v1 shape (top-level symbol_table)") + } + if _, ok := v["schema_version"]; ok { + t.Error("default output should not carry schema_version (that is v2)") + } +} + +func TestRootCmd_AnalysisSchema2EmitsV2(t *testing.T) { + td := cliTestdataDir() + out, _, err := runCmd( + "--input", filepath.Join(td, "greeter"), + "--analysis-schema", "2", + "--cache-dir", t.TempDir(), + ) + if err != nil { + t.Fatalf("command failed: %v", err) + } + var v struct { + SchemaVersion string `json:"schema_version"` + Language string `json:"language"` + Application struct { + ID string `json:"id"` + } `json:"application"` + } + if jsonErr := json.Unmarshal([]byte(out), &v); jsonErr != nil { + t.Fatalf("stdout is not valid JSON: %v", jsonErr) + } + if v.SchemaVersion != "2.0.0" { + t.Errorf("schema_version = %q, want 2.0.0", v.SchemaVersion) + } + if v.Language != "go" { + t.Errorf("language = %q, want go", v.Language) + } + if v.Application.ID != "can://go/greeter" { + t.Errorf("application.id = %q, want can://go/greeter", v.Application.ID) + } +} + +func TestRootCmd_UnknownSchemaReturnsError(t *testing.T) { + td := cliTestdataDir() + _, _, err := runCmd( + "--input", filepath.Join(td, "greeter"), + "--analysis-schema", "9", + "--cache-dir", t.TempDir(), + ) + if err == nil { + t.Fatal("expected error for unknown --analysis-schema value, got nil") + } +} + // ── --skip-tests ───────────────────────────────────────────────────────────── func TestRootCmd_SkipTestsFalseIncludesTestFiles(t *testing.T) { diff --git a/internal/core/analyzer.go b/internal/core/analyzer.go index 2b9f410..d889a2f 100644 --- a/internal/core/analyzer.go +++ b/internal/core/analyzer.go @@ -5,11 +5,11 @@ // discipline of codeanalyzer-python/codeanalyzer/core.py. // // Phase order: -// 1. Project materialization (go mod download) -// 2. Symbol table construction (syntactic_analysis) -// 3. Resolver-based call graph (semantic_analysis) — if level >= 2 -// 4. Pass pipeline (analysis/registry) -// 5. Optional CodeQL enrichment (semantic_analysis/codeql) — if --codeql +// 1. Project materialization (go mod download) +// 2. Symbol table construction (syntactic_analysis) +// 3. Resolver-based call graph (semantic_analysis) — if level >= 2 +// 4. Pass pipeline (analysis/registry) +// 5. Optional CodeQL enrichment (semantic_analysis/codeql) — if --codeql package core import ( @@ -25,6 +25,7 @@ import ( "github.com/codellm-devkit/codeanalyzer-go/internal/semantic_analysis" "github.com/codellm-devkit/codeanalyzer-go/internal/semantic_analysis/codeql" "github.com/codellm-devkit/codeanalyzer-go/internal/syntactic_analysis" + "github.com/codellm-devkit/codeanalyzer-go/internal/syntactic_analysis/v2emit" "github.com/codellm-devkit/codeanalyzer-go/internal/utils" ) @@ -170,10 +171,14 @@ func (a *Analyzer) saveCache(app *schema.GoApplication) error { return os.WriteFile(cachePath, data, 0o644) } -// WriteOutput writes the GoApplication to outputDir/analysis.json (or stdout -// when outputDir is empty). Only "json" is supported; "msgpack" and other -// values return an explicit error rather than silently falling back to JSON. -func WriteOutput(app *schema.GoApplication, outputDir, format string) error { +// RenderJSON serializes the analysis to JSON in the schema selected by +// opts.SchemaVersion: 1 (default) emits the legacy v1 GoApplication shape; 2 +// emits the canonical v2 tree via the v2emit emitter. This is the single place +// the v1/v2 decision is made, so both the stdout and file output paths agree. +// Only "json" is supported; other formats return an explicit error rather than +// silently falling back. +func RenderJSON(app *schema.GoApplication, opts options.AnalysisOptions) ([]byte, error) { + format := opts.Format if format == "" { format = "json" } @@ -181,21 +186,35 @@ func WriteOutput(app *schema.GoApplication, outputDir, format string) error { case "json": // only supported format case "msgpack": - return fmt.Errorf("msgpack output is not yet implemented; use --format json") + return nil, fmt.Errorf("msgpack output is not yet implemented; use --format json") default: - return fmt.Errorf("unsupported output format %q; supported: json", format) + return nil, fmt.Errorf("unsupported output format %q; supported: json", format) } - data, err := json.Marshal(app) + switch opts.SchemaVersion { + case 0, 1: + return json.Marshal(app) + case 2: + payload := v2emit.Emit(app, opts.AppName, opts.InputPath, int(opts.Level), opts.AnalyzerVersion) + return json.Marshal(payload) + default: + return nil, fmt.Errorf("unsupported --analysis-schema %d; supported: 1, 2", opts.SchemaVersion) + } +} + +// WriteOutput serializes the analysis (in the selected schema) and writes it to +// outputDir/analysis.json, or to stdout when outputDir is empty. +func WriteOutput(app *schema.GoApplication, opts options.AnalysisOptions) error { + data, err := RenderJSON(app, opts) if err != nil { return err } - if outputDir == "" { + if opts.OutputDir == "" { _, err = os.Stdout.Write(data) return err } - if err := utils.EnsureDir(outputDir); err != nil { + if err := utils.EnsureDir(opts.OutputDir); err != nil { return err } - return os.WriteFile(filepath.Join(outputDir, "analysis.json"), data, 0o644) + return os.WriteFile(filepath.Join(opts.OutputDir, "analysis.json"), data, 0o644) } diff --git a/internal/core/analyzer_test.go b/internal/core/analyzer_test.go index 1e95011..45dacf2 100644 --- a/internal/core/analyzer_test.go +++ b/internal/core/analyzer_test.go @@ -137,7 +137,7 @@ func TestCallGraph_CallSitesBackfilled(t *testing.T) { func TestWriteOutput_ValidJSON(t *testing.T) { outDir := t.TempDir() - if err := core.WriteOutput(sharedGreeterL2, outDir, "json"); err != nil { + if err := core.WriteOutput(sharedGreeterL2, options.AnalysisOptions{OutputDir: outDir, Format: "json"}); err != nil { t.Fatalf("WriteOutput: %v", err) } data, err := os.ReadFile(filepath.Join(outDir, "analysis.json")) @@ -155,7 +155,7 @@ func TestWriteOutput_ValidJSON(t *testing.T) { func TestWriteOutput_EmptyFormatDefaultsToJSON(t *testing.T) { outDir := t.TempDir() - if err := core.WriteOutput(sharedGreeterL1, outDir, ""); err != nil { + if err := core.WriteOutput(sharedGreeterL1, options.AnalysisOptions{OutputDir: outDir, Format: ""}); err != nil { t.Fatalf("WriteOutput with empty format: %v", err) } if _, err := os.Stat(filepath.Join(outDir, "analysis.json")); err != nil { @@ -165,14 +165,14 @@ func TestWriteOutput_EmptyFormatDefaultsToJSON(t *testing.T) { func TestWriteOutput_MsgpackNotImplemented(t *testing.T) { outDir := t.TempDir() - if err := core.WriteOutput(sharedGreeterL1, outDir, "msgpack"); err == nil { + if err := core.WriteOutput(sharedGreeterL1, options.AnalysisOptions{OutputDir: outDir, Format: "msgpack"}); err == nil { t.Fatal("expected error for --format msgpack, got nil") } } func TestWriteOutput_UnknownFormatErrors(t *testing.T) { outDir := t.TempDir() - if err := core.WriteOutput(sharedGreeterL1, outDir, "csv"); err == nil { + if err := core.WriteOutput(sharedGreeterL1, options.AnalysisOptions{OutputDir: outDir, Format: "csv"}); err == nil { t.Fatal("expected error for unknown format, got nil") } } diff --git a/internal/core/testsetup_test.go b/internal/core/testsetup_test.go index 85f3b72..2b14d32 100644 --- a/internal/core/testsetup_test.go +++ b/internal/core/testsetup_test.go @@ -66,6 +66,7 @@ func runTestMain(m *testing.M) int { opts := options.AnalysisOptions{ InputPath: f.path, OutputDir: outDir, + Format: "json", Level: f.level, SkipTests: true, CacheDir: cacheDir, @@ -75,7 +76,7 @@ func runTestMain(m *testing.M) int { fmt.Fprintf(os.Stderr, "testsetup %s: Analyze: %v\n", f.name, err) return 1 } - if err := core.WriteOutput(app, outDir, "json"); err != nil { + if err := core.WriteOutput(app, opts); err != nil { fmt.Fprintf(os.Stderr, "testsetup %s: WriteOutput: %v\n", f.name, err) return 1 } diff --git a/internal/options/options.go b/internal/options/options.go index 051fe6c..e19d542 100644 --- a/internal/options/options.go +++ b/internal/options/options.go @@ -38,6 +38,12 @@ type AnalysisOptions struct { AppName string // AnalysisLevel controls symbol-table-only (1) vs + call graph (2). Level AnalysisLevel + // SchemaVersion selects the output schema major: 1 = the legacy v1 shape + // (default, compat shim during the migration), 2 = the canonical v2 tree. + SchemaVersion int + // AnalyzerVersion is the analyzer's own version, recorded in the v2 + // manifest's analyzer{} tag. Set by the CLI from its build-stamped value. + AnalyzerVersion string // TargetFiles restricts analysis to specific files (incremental mode). TargetFiles []string // SkipTests skips test files (files ending in _test.go). diff --git a/internal/syntactic_analysis/v2emit/emit.go b/internal/syntactic_analysis/v2emit/emit.go index d2501d2..f4ba1f0 100644 --- a/internal/syntactic_analysis/v2emit/emit.go +++ b/internal/syntactic_analysis/v2emit/emit.go @@ -12,18 +12,19 @@ import ( // Emit transforms the v1 GoApplication into the v2 Analysis payload. appName is // the application anchor (--app-name); projectDir is the absolute input root, // used to read each module's source for span slicing. maxLevel records how -// deeply the tree was populated (1 = symbol table, 2 = + call_graph). +// deeply the tree was populated (1 = symbol table, 2 = + call_graph); +// analyzerVersion is stamped into the manifest's analyzer{} tag. // // This is a pure re-serialization: every fact comes from the v1 model, except // span byte offsets, which are derived from source via lineIndex. -func Emit(app *schema.GoApplication, appName, projectDir string, maxLevel int) *v2.Analysis { +func Emit(app *schema.GoApplication, appName, projectDir string, maxLevel int, analyzerVersion string) *v2.Analysis { appID := v2.AppID(appName) out := &v2.Analysis{ SchemaVersion: v2.SchemaVersion, Language: v2.Language, MaxLevel: maxLevel, - Analyzer: v2.AnalyzerTag{Name: "codeanalyzer-go"}, + Analyzer: v2.AnalyzerTag{Name: "codeanalyzer-go", Version: analyzerVersion}, Application: v2.Application{ ID: appID, Kind: "application", diff --git a/internal/syntactic_analysis/v2emit/emit_test.go b/internal/syntactic_analysis/v2emit/emit_test.go index deb8a66..6dd39ee 100644 --- a/internal/syntactic_analysis/v2emit/emit_test.go +++ b/internal/syntactic_analysis/v2emit/emit_test.go @@ -37,7 +37,7 @@ func buildGreeter(t *testing.T) (*schema.GoApplication, string) { func TestEmit_Envelope(t *testing.T) { app, dir := buildGreeter(t) - out := Emit(app, "greeter", dir, 1) + out := Emit(app, "greeter", dir, 1, "test") if out.SchemaVersion != "2.0.0" { t.Errorf("schema_version = %q, want 2.0.0", out.SchemaVersion) @@ -61,7 +61,7 @@ func TestEmit_Envelope(t *testing.T) { func TestEmit_ModuleSourceAndIDs(t *testing.T) { app, dir := buildGreeter(t) - out := Emit(app, "greeter", dir, 1) + out := Emit(app, "greeter", dir, 1, "test") mod, ok := out.Application.SymbolTable["main.go"] if !ok { @@ -84,7 +84,7 @@ func TestEmit_ModuleSourceAndIDs(t *testing.T) { func TestEmit_CallableBodyAndSpanSlice(t *testing.T) { app, dir := buildGreeter(t) - out := Emit(app, "greeter", dir, 1) + out := Emit(app, "greeter", dir, 1, "test") mod := out.Application.SymbolTable["main.go"] var main v2.Callable @@ -123,7 +123,7 @@ func TestEmit_CallableBodyAndSpanSlice(t *testing.T) { func TestEmit_CalleeSerializesAsNull(t *testing.T) { app, dir := buildGreeter(t) - out := Emit(app, "greeter", dir, 1) + out := Emit(app, "greeter", dir, 1, "test") data, err := json.Marshal(out) if err != nil { t.Fatalf("marshal failed: %v", err) @@ -144,7 +144,7 @@ func TestEmit_CalleeSerializesAsNull(t *testing.T) { func TestEmit_TypeKindAndMethods(t *testing.T) { app, dir := buildGreeter(t) - out := Emit(app, "greeter", dir, 1) + out := Emit(app, "greeter", dir, 1, "test") // Find the Greeter type wherever it lives; assert struct kind + methods. var found bool From e8a958122f96f1890a79824c29a2a9f3ec145240 Mon Sep 17 00:00:00 2001 From: lambertpw Date: Mon, 28 Sep 2026 23:05:05 -0400 Subject: [PATCH 08/17] test: L1 v2 fixture gate over the multipackage fixture Lock the v2 L1 output shape against a known multi-file fixture so the level cannot silently regress. Asserts concrete values for every item on the L1 coverage checklist (testing-and-validation.md): - multi-file compilation unit: four modules keyed by relative path, each with source-once and a whole-file span that slices back - exported + unexported symbols (Worker.Run vs Worker.execute) - interface kind collapse (Processor -> kind=interface) - error_channel from an error-typed return (Processor.Process) - variadic parameter (Combine(results ...Result)) - language-specific call flag (is_goroutine on `go w.execute(...)`) - sanctioned callee:null refinement slot present at L1 - can:// durable ids at module/type/callable depth Plus the superset gate: walk the same v1 model the emitter walked and assert every file/type/method/function reappears in v2, and that the sanctioned drops (is_constructor_call, per-node code) are absent from the serialized payload. --- .../syntactic_analysis/v2emit/gate_test.go | 311 ++++++++++++++++++ 1 file changed, 311 insertions(+) create mode 100644 internal/syntactic_analysis/v2emit/gate_test.go diff --git a/internal/syntactic_analysis/v2emit/gate_test.go b/internal/syntactic_analysis/v2emit/gate_test.go new file mode 100644 index 0000000..db1317a --- /dev/null +++ b/internal/syntactic_analysis/v2emit/gate_test.go @@ -0,0 +1,311 @@ +package v2emit + +// L1 conformance gate for the v2 emitter. +// +// This is the fixture gate the codeanalyzer-backend ladder requires before L1 +// advances: it locks the v2 L1 OUTPUT SHAPE against a known multi-file fixture, +// asserting concrete values (not just "a key exists") for every item on the L1 +// coverage checklist in +// designing-cldk-changes/references/testing-and-validation.md: +// +// - a multi-file compilation unit (four modules keyed by relative path) +// - exported AND unexported symbols (Worker.Run vs Worker.execute) +// - an interface type collapsed into kind=interface (Processor) +// - error_channel populated from an error-typed return (Processor.Process) +// - a variadic parameter (Combine(results ...Result)) +// - a language-specific call flag (is_goroutine on `go w.execute(...)`) +// - the sanctioned callee:null refinement slot present at L1 +// - can:// durable ids at module/type/callable depth +// - the SUPERSET gate: every v1 fact survives into v2 modulo sanctioned drops +// (is_constructor_call is dropped; per-node `code` folds into module.source). +// +// If any of these regress, this gate fails and the level does not advance. + +import ( + "encoding/json" + "path/filepath" + "runtime" + "strings" + "testing" + + "github.com/codellm-devkit/codeanalyzer-go/internal/schema" + v2 "github.com/codellm-devkit/codeanalyzer-go/internal/schema/v2" + "github.com/codellm-devkit/codeanalyzer-go/internal/syntactic_analysis" +) + +// multipackageDir returns the absolute path to the multipackage testdata fixture. +func multipackageDir(t *testing.T) string { + t.Helper() + _, thisFile, _, _ := runtime.Caller(0) + abs, err := filepath.Abs(filepath.Join(filepath.Dir(thisFile), "..", "..", "..", "testdata", "multipackage")) + if err != nil { + t.Fatalf("resolving fixture dir: %v", err) + } + return abs +} + +// buildMultipackageV2 runs the real symbol-table builder over the multipackage +// fixture and emits the v2 L1 payload. skipTests=true keeps *_test.go out so the +// module set is deterministic. +func buildMultipackageV2(t *testing.T) *v2.Analysis { + t.Helper() + dir := multipackageDir(t) + b := syntactic_analysis.NewSymbolTableBuilder(dir) + st, err := b.Build(nil, true) + if err != nil { + t.Fatalf("symbol table build failed: %v", err) + } + app := &schema.GoApplication{SymbolTable: st, CallGraph: []schema.GoCallEdge{}} + return Emit(app, "multipackage", dir, 1, "gate-test") +} + +// mustModule fetches a module by relative path or fails. +func mustModule(t *testing.T, out *v2.Analysis, rel string) v2.Module { + t.Helper() + mod, ok := out.Application.SymbolTable[rel] + if !ok { + t.Fatalf("module %q missing; modules=%v", rel, keys(out.Application.SymbolTable)) + } + return mod +} + +// TestGateL1_ModuleSet locks the exact multi-file compilation unit: the four +// non-test source files, each a module keyed by its relative path. +func TestGateL1_ModuleSet(t *testing.T) { + out := buildMultipackageV2(t) + + want := []string{ + "main.go", + "server/middleware.go", + "server/server.go", + "worker/worker.go", + } + for _, rel := range want { + mod := mustModule(t, out, rel) + if mod.Kind != "module" { + t.Errorf("%s: kind = %q, want module", rel, mod.Kind) + } + wantID := "can://go/multipackage/" + rel + if mod.ID != wantID { + t.Errorf("%s: id = %q, want %q", rel, mod.ID, wantID) + } + if mod.Source == "" { + t.Errorf("%s: source is empty; every module carries its whole file text", rel) + } + // The module span must slice back to the whole source (source-once invariant). + if got := mod.Span.Slice(mod.Source); got != mod.Source { + t.Errorf("%s: module span does not cover whole source (%d vs %d bytes)", rel, len(got), len(mod.Source)) + } + } +} + +// TestGateL1_InterfaceKind locks the Processor interface: is_interface collapses +// into kind=interface, and its method resolves as a callable under the type. +func TestGateL1_InterfaceKind(t *testing.T) { + out := buildMultipackageV2(t) + worker := mustModule(t, out, "worker/worker.go") + + proc, ok := worker.Types["Processor"] + if !ok { + t.Fatalf("Processor type missing; types=%v", keys(worker.Types)) + } + if proc.Kind != "interface" { + t.Errorf("Processor.kind = %q, want interface", proc.Kind) + } + const procMethod = "example.com/multipackage/worker.Processor.Process" + m, ok := proc.Callables[procMethod] + if !ok { + t.Fatalf("Processor.Process missing; callables=%v", keys(proc.Callables)) + } + if m.Kind != "method" { + t.Errorf("Process.kind = %q, want method", m.Kind) + } + // error_channel is populated from the (Result, error) return. + if !contains(m.ErrorChannel, "error") { + t.Errorf("Process.error_channel = %v, want to contain \"error\"", m.ErrorChannel) + } +} + +// TestGateL1_ExportedAndUnexported locks that both an exported (Run) and an +// unexported (execute) method survive, each a method callable under Worker. +func TestGateL1_ExportedAndUnexported(t *testing.T) { + out := buildMultipackageV2(t) + worker := mustModule(t, out, "worker/worker.go") + + wtype, ok := worker.Types["Worker"] + if !ok { + t.Fatalf("Worker type missing; types=%v", keys(worker.Types)) + } + const ( + runID = "example.com/multipackage/worker.Worker.Run" + execID = "example.com/multipackage/worker.Worker.execute" + ) + for _, sig := range []string{runID, execID} { + m, ok := wtype.Callables[sig] + if !ok { + t.Fatalf("method %q missing; callables=%v", sig, keys(wtype.Callables)) + } + if m.Kind != "method" { + t.Errorf("%s.kind = %q, want method", sig, m.Kind) + } + } +} + +// TestGateL1_VariadicParameter locks the variadic-parameter fact: +// Combine(results ...Result) → is_variadic=true on the sole parameter. +func TestGateL1_VariadicParameter(t *testing.T) { + out := buildMultipackageV2(t) + worker := mustModule(t, out, "worker/worker.go") + + const combineID = "example.com/multipackage/worker.Combine" + comb, ok := worker.Functions[combineID] + if !ok { + t.Fatalf("Combine function missing; functions=%v", keys(worker.Functions)) + } + if len(comb.Parameters) != 1 { + t.Fatalf("Combine should have 1 parameter; got %d", len(comb.Parameters)) + } + if !comb.Parameters[0].IsVariadic { + t.Errorf("Combine parameter %q should be variadic", comb.Parameters[0].Name) + } +} + +// TestGateL1_GoroutineCallNode locks the language-specific call flag: Worker.Run +// contains `go w.execute(...)`, which must surface as a body call node with +// is_goroutine=true AND callee null at L1. +func TestGateL1_GoroutineCallNode(t *testing.T) { + out := buildMultipackageV2(t) + worker := mustModule(t, out, "worker/worker.go") + + run, ok := worker.Types["Worker"].Callables["example.com/multipackage/worker.Worker.Run"] + if !ok { + t.Fatal("Worker.Run missing") + } + if len(run.Body) == 0 { + t.Fatal("Worker.Run should have body call nodes") + } + sawGoroutine := false + for local, node := range run.Body { + if node.Kind != "call" { + t.Errorf("body[%s].kind = %q, want call", local, node.Kind) + } + if node.Callee != nil { + t.Errorf("body[%s].callee should be null at L1; got %q", local, *node.Callee) + } + if node.IsGoroutine { + sawGoroutine = true + } + } + if !sawGoroutine { + t.Error("Worker.Run should have a call node with is_goroutine=true (go w.execute(...))") + } +} + +// TestGateL1_CrossPackageStructure locks that a cross-package caller (main.go +// calling into server/ and worker/) is present as its own module with body call +// nodes — the compilation unit spans packages, not a single file. +func TestGateL1_CrossPackageStructure(t *testing.T) { + out := buildMultipackageV2(t) + main := mustModule(t, out, "main.go") + + if main.Package != "main" { + t.Errorf("main.go package = %q, want main", main.Package) + } + // main() drives the cross-package calls; its body must carry call nodes. + var mainFn v2.Callable + found := false + for _, fn := range main.Functions { + if strings.HasSuffix(fn.Signature, ".main") { + mainFn, found = fn, true + } + } + if !found { + t.Fatalf("main function missing; functions=%v", keys(main.Functions)) + } + if len(mainFn.Body) == 0 { + t.Error("main() should have body call nodes for its cross-package calls") + } +} + +// TestGateL1_Envelope locks the manifest envelope at L1. +func TestGateL1_Envelope(t *testing.T) { + out := buildMultipackageV2(t) + + if out.SchemaVersion != v2.SchemaVersion { + t.Errorf("schema_version = %q, want %q", out.SchemaVersion, v2.SchemaVersion) + } + if out.Language != "go" { + t.Errorf("language = %q, want go", out.Language) + } + if out.MaxLevel != 1 { + t.Errorf("max_level = %d, want 1", out.MaxLevel) + } + if out.Application.ID != "can://go/multipackage" { + t.Errorf("application.id = %q, want can://go/multipackage", out.Application.ID) + } + // L1: no call_graph edges yet. + if len(out.Application.CallGraph) != 0 { + t.Errorf("L1 call_graph should be empty; got %d edges", len(out.Application.CallGraph)) + } +} + +// TestGateL1_SupersetOfV1 is the superset gate: every v1 fact must survive into +// v2 modulo the sanctioned drops. It walks the SAME v1 model the emitter walked +// and asserts each callable/type/function reappears in v2, and that the +// sanctioned drops are gone from the serialized payload. +func TestGateL1_SupersetOfV1(t *testing.T) { + dir := multipackageDir(t) + b := syntactic_analysis.NewSymbolTableBuilder(dir) + st, err := b.Build(nil, true) + if err != nil { + t.Fatalf("symbol table build failed: %v", err) + } + app := &schema.GoApplication{SymbolTable: st, CallGraph: []schema.GoCallEdge{}} + out := Emit(app, "multipackage", dir, 1, "gate-test") + + // Every v1 file, type, method and function reappears in v2. + for rel, v1file := range st { + mod, ok := out.Application.SymbolTable[rel] + if !ok { + t.Errorf("v1 file %q dropped from v2 symbol_table", rel) + continue + } + for name := range v1file.Types { + if _, ok := mod.Types[name]; !ok { + t.Errorf("%s: v1 type %q dropped from v2", rel, name) + } + for msig := range v1file.Types[name].Methods { + if _, ok := mod.Types[name].Callables[msig]; !ok { + t.Errorf("%s: v1 method %q dropped from v2 type %q", rel, msig, name) + } + } + } + for sig := range v1file.Functions { + if _, ok := mod.Functions[sig]; !ok { + t.Errorf("%s: v1 function %q dropped from v2", rel, sig) + } + } + } + + // Sanctioned drops must be absent from the serialized payload. + data, err := json.Marshal(out) + if err != nil { + t.Fatalf("marshal failed: %v", err) + } + if strings.Contains(string(data), `"is_constructor_call"`) { + t.Error("v2 payload should not carry the dropped is_constructor_call") + } + // The sanctioned callee:null slot must be present at L1. + if !strings.Contains(string(data), `"callee":null`) { + t.Error("expected the sanctioned callee:null refinement slot at L1") + } +} + +func contains(ss []string, want string) bool { + for _, s := range ss { + if s == want { + return true + } + } + return false +} From b133324f6de3493a6464096fa720731c2e9e5376 Mon Sep 17 00:00:00 2001 From: lambertpw Date: Mon, 28 Sep 2026 23:11:00 -0400 Subject: [PATCH 09/17] =?UTF-8?q?feat(emit):=20v2=20L2=20refinement=20?= =?UTF-8?q?=E2=80=94=20can://=20ids=20on=20callees=20and=20edges?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit L2 is a pure refinement over the L1 tree: it fills the sanctioned callee null->id slot and translates call_graph edge endpoints onto real callable ids. Both need one thing v1 does not have: a mapping from v1 signature identity to v2 can:// node id. Add buildSigIndex, a one-pass signature->id index built with the same id builders emitCallable uses, so an indexed id is byte-identical to the id the callable node carries (no dangling endpoints by construction). Then: - emitBody: backfill body.callee to the callee's can:// id when the resolver-set signature names an in-tree callable; keep it null for external/stdlib callees (the honest-unresolved fallback). - emitCallGraph: map src/dst from v1 signatures to can:// ids; drop any edge whose endpoint does not resolve rather than emit a dangling one (defensive — the v1 resolver only emits in-project targets). At L1 (no resolver run) callee stays null, so L1 output is unchanged. --- internal/syntactic_analysis/v2emit/emit.go | 72 ++++++++++++++----- .../syntactic_analysis/v2emit/sigindex.go | 59 +++++++++++++++ 2 files changed, 113 insertions(+), 18 deletions(-) create mode 100644 internal/syntactic_analysis/v2emit/sigindex.go diff --git a/internal/syntactic_analysis/v2emit/emit.go b/internal/syntactic_analysis/v2emit/emit.go index f4ba1f0..1fb7d7e 100644 --- a/internal/syntactic_analysis/v2emit/emit.go +++ b/internal/syntactic_analysis/v2emit/emit.go @@ -20,6 +20,11 @@ import ( func Emit(app *schema.GoApplication, appName, projectDir string, maxLevel int, analyzerVersion string) *v2.Analysis { appID := v2.AppID(appName) + // The signature→can:// id index translates v1 identity (signature strings) + // into v2 node ids. It is what the L2 refinement uses to backfill body + // callees and to map call_graph edge endpoints onto real callable ids. + idx := buildSigIndex(appID, app.SymbolTable) + out := &v2.Analysis{ SchemaVersion: v2.SchemaVersion, Language: v2.Language, @@ -29,7 +34,7 @@ func Emit(app *schema.GoApplication, appName, projectDir string, maxLevel int, a ID: appID, Kind: "application", SymbolTable: make(map[string]v2.Module, len(app.SymbolTable)), - CallGraph: emitCallGraph(app.CallGraph), + CallGraph: emitCallGraph(app.CallGraph, idx), }, } @@ -37,7 +42,7 @@ func Emit(app *schema.GoApplication, appName, projectDir string, maxLevel int, a // sorted keeps any incidental ordering (and logs) stable across runs. for _, relPath := range sortedFileKeys(app.SymbolTable) { file := app.SymbolTable[relPath] - out.Application.SymbolTable[relPath] = emitModule(appID, projectDir, relPath, file) + out.Application.SymbolTable[relPath] = emitModule(appID, projectDir, relPath, file, idx) } return out } @@ -45,7 +50,7 @@ func Emit(app *schema.GoApplication, appName, projectDir string, maxLevel int, a // emitModule builds a v2 module from a v1 GoFile, reading the file's source so // spans can slice off it. A file that cannot be read still emits its structure // (with empty source and zeroed byte spans) rather than dropping the module. -func emitModule(appID, projectDir, relPath string, file schema.GoFile) v2.Module { +func emitModule(appID, projectDir, relPath string, file schema.GoFile, idx sigIndex) v2.Module { source := readSource(projectDir, relPath) li := newLineIndex(source) modID := v2.ModuleID(appID, relPath) @@ -65,10 +70,10 @@ func emitModule(appID, projectDir, relPath string, file schema.GoFile) v2.Module } for name, t := range file.Types { - mod.Types[name] = emitType(modID, li, t) + mod.Types[name] = emitType(modID, li, t, idx) } for sig, fn := range file.Functions { - mod.Functions[sig] = emitCallable(modID, li, fn) + mod.Functions[sig] = emitCallable(modID, li, fn, idx) } return mod } @@ -76,7 +81,7 @@ func emitModule(appID, projectDir, relPath string, file schema.GoFile) v2.Module // emitType maps a v1 GoType to a v2 type node, collapsing is_interface into the // kind and splitting embedding (base_types) from computed satisfaction. Methods // resolve into the type's callables{}. -func emitType(modID string, li *lineIndex, t schema.GoType) v2.Type { +func emitType(modID string, li *lineIndex, t schema.GoType, idx sigIndex) v2.Type { typeID := v2.TypeID(modID, t.Signature) out := v2.Type{ ID: typeID, @@ -97,14 +102,14 @@ func emitType(modID string, li *lineIndex, t schema.GoType) v2.Type { } } for sig, m := range t.Methods { - out.Callables[sig] = emitCallable(typeID, li, m) + out.Callables[sig] = emitCallable(typeID, li, m, idx) } return out } // emitCallable maps a v1 GoCallable to a v2 callable node, including its body // (call nodes) and nested closures. -func emitCallable(parentID string, li *lineIndex, c schema.GoCallable) v2.Callable { +func emitCallable(parentID string, li *lineIndex, c schema.GoCallable, idx sigIndex) v2.Callable { callID := v2.CallableID(parentID, c.Signature) out := v2.Callable{ ID: callID, @@ -113,7 +118,7 @@ func emitCallable(parentID string, li *lineIndex, c schema.GoCallable) v2.Callab Signature: c.Signature, Parameters: emitParams(li, c.Parameters), ReturnType: c.ReturnType, - Body: emitBody(callID, li, c.CallSites), + Body: emitBody(callID, li, c.CallSites, idx), } if ch := errorChannel(c.ReturnTypes); len(ch) > 0 { out.ErrorChannel = ch @@ -124,28 +129,50 @@ func emitCallable(parentID string, li *lineIndex, c schema.GoCallable) v2.Callab if len(c.InnerCallables) > 0 { out.Callables = make(map[string]v2.Callable, len(c.InnerCallables)) for sig, inner := range c.InnerCallables { - out.Callables[sig] = emitCallable(callID, li, inner) + out.Callables[sig] = emitCallable(callID, li, inner, idx) } } return out } -// emitBody builds the L1 body map: one `call` node per call site, keyed by its -// line:col local id, with the sanctioned callee null->id refinement slot. -func emitBody(callableID string, li *lineIndex, sites []schema.GoCallsite) map[string]v2.BodyNode { +// emitBody builds the body map: one `call` node per call site, keyed by its +// line:col local id, carrying the sanctioned callee null->id refinement slot. +// +// callee resolution: +// - L1: the v1 site has no backfilled signature, so callee stays null. +// - L2: the resolver backfilled the site's callee signature. When that +// signature names an in-tree callable, callee becomes its can:// id; when +// it names an external/stdlib callee (no in-tree node), callee stays null — +// the honest-unresolved fallback the L2 gate expects. +func emitBody(callableID string, li *lineIndex, sites []schema.GoCallsite, idx sigIndex) map[string]v2.BodyNode { body := make(map[string]v2.BodyNode, len(sites)) for _, cs := range sites { local := v2.LocalID(cs.StartLine, cs.StartColumn) body[local] = v2.BodyNode{ Kind: "call", Span: colSpan(li, cs.StartLine, cs.StartColumn, cs.EndLine, cs.EndColumn), - Callee: cs.CalleeSignature, // *string: null at L1, backfilled at L2 + Callee: calleeID(cs.CalleeSignature, idx), IsGoroutine: cs.IsGoroutine, } } return body } +// calleeID translates a v1 backfilled callee signature into the v2 can:// id +// refinement slot: nil (JSON null) when unresolved or external, a pointer to the +// in-tree callable id when resolved. Never returns the raw v1 signature — a +// call node's callee is always a node id or null. +func calleeID(sig *string, idx sigIndex) *string { + if sig == nil { + return nil + } + id, ok := idx.idFor(*sig) + if !ok { + return nil + } + return &id +} + func emitParams(li *lineIndex, params []schema.GoParameter) []v2.Parameter { out := make([]v2.Parameter, 0, len(params)) for _, p := range params { @@ -173,13 +200,22 @@ func emitImports(li *lineIndex, imports []schema.GoImport) []v2.Import { } // emitCallGraph maps v1 identity-only edges to the v2 {src,dst,prov,weight} -// shape (rename of source/target/provenance). -func emitCallGraph(edges []schema.GoCallEdge) []v2.Edge { +// shape: the source/target/provenance rename, PLUS the endpoint translation from +// v1 signatures to can:// callable ids. An edge whose src or dst does not resolve +// to an in-tree node is dropped rather than emitted with a dangling endpoint — +// the L2 gate requires every endpoint to be a real callable id. (The v1 resolver +// only emits edges to in-project targets, so this drop is defensive.) +func emitCallGraph(edges []schema.GoCallEdge, idx sigIndex) []v2.Edge { out := make([]v2.Edge, 0, len(edges)) for _, e := range edges { + src, srcOK := idx.idFor(e.Source) + dst, dstOK := idx.idFor(e.Target) + if !srcOK || !dstOK { + continue + } out = append(out, v2.Edge{ - Src: e.Source, - Dst: e.Target, + Src: src, + Dst: dst, Prov: nonEmpty(e.Provenance), Weight: e.Weight, }) diff --git a/internal/syntactic_analysis/v2emit/sigindex.go b/internal/syntactic_analysis/v2emit/sigindex.go new file mode 100644 index 0000000..f0e8cd8 --- /dev/null +++ b/internal/syntactic_analysis/v2emit/sigindex.go @@ -0,0 +1,59 @@ +package v2emit + +import ( + "github.com/codellm-devkit/codeanalyzer-go/internal/schema" + v2 "github.com/codellm-devkit/codeanalyzer-go/internal/schema/v2" +) + +// sigIndex maps a v1 callable signature to its v2 can:// callable id. +// +// v1 identifies every callable by its signatureOf() string and uses that string +// as call-graph edge endpoints and as a call site's backfilled callee. v2 +// addresses callables by durable can:// id. The L2 refinement therefore needs a +// signature→id translation: this index is built in one pass over the same tree +// the emitter walks, using the SAME id builders emitCallable uses, so an id here +// is byte-identical to the id the callable node carries. That is what lets the +// L2 gate assert "no dangling endpoints" — every edge endpoint resolves to a +// real node id. +// +// v1 signatures are globally unique (a method signature embeds its receiver +// type), so a flat signature→id map has no collisions. +type sigIndex map[string]string + +// buildSigIndex walks the symbol table and records, for every callable +// (functions, methods, and nested closures), signature → can:// id. It mirrors +// emitModule/emitType/emitCallable's parent-id computation exactly. +func buildSigIndex(appID string, symbolTable map[string]schema.GoFile) sigIndex { + idx := make(sigIndex) + for relPath, file := range symbolTable { + modID := v2.ModuleID(appID, relPath) + for _, fn := range file.Functions { + idx.addCallable(modID, fn) + } + for _, t := range file.Types { + typeID := v2.TypeID(modID, t.Signature) + for _, m := range t.Methods { + idx.addCallable(typeID, m) + } + } + } + return idx +} + +// addCallable records one callable and recurses into its closures, matching the +// id nesting emitCallable produces. +func (idx sigIndex) addCallable(parentID string, c schema.GoCallable) { + callID := v2.CallableID(parentID, c.Signature) + idx[c.Signature] = callID + for _, inner := range c.InnerCallables { + idx.addCallable(callID, inner) + } +} + +// idFor returns the can:// id for a v1 signature, or ("", false) when the +// signature names a callable outside the project (stdlib/external) — those +// resolve to no in-tree node and must not become an edge endpoint. +func (idx sigIndex) idFor(sig string) (string, bool) { + id, ok := idx[sig] + return id, ok +} From 62990c11ce615654e62d367cd842f4312cb92914 Mon Sep 17 00:00:00 2001 From: lambertpw Date: Mon, 28 Sep 2026 23:13:10 -0400 Subject: [PATCH 10/17] test: L2 v2 fixture gate over the multipackage fixture MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Lock the v2 L2 output shape by building the fixture through the real resolver (the -a 2 wiring) and asserting the L2 checklist (testing-and-validation.md): - no dangling endpoints: every edge src/dst resolves to a real callable id in the tree - every edge carries a non-empty prov naming the resolver - a NAMED expected edge (Worker.Run -> Worker.execute) by exact can:// id pair, so the gate proves correctness not just non-emptiness - a NAMED cross-package edge (main -> server.New) - callee backfilled to an in-tree can:// id on the goroutine call site - the L1 ⊆ L2 superset: emitting the same model at level 1 vs 2 leaves every module/type/callable id, span, kind and body node identical, changing only callee null -> id (the one sanctioned mutation). --- .../syntactic_analysis/v2emit/gate_l2_test.go | 243 ++++++++++++++++++ 1 file changed, 243 insertions(+) create mode 100644 internal/syntactic_analysis/v2emit/gate_l2_test.go diff --git a/internal/syntactic_analysis/v2emit/gate_l2_test.go b/internal/syntactic_analysis/v2emit/gate_l2_test.go new file mode 100644 index 0000000..8a50883 --- /dev/null +++ b/internal/syntactic_analysis/v2emit/gate_l2_test.go @@ -0,0 +1,243 @@ +package v2emit + +// L2 conformance gate for the v2 emitter. +// +// L2 is a pure refinement over L1: it backfills the callee null->id slot and +// emits the call_graph edge list, both endpoints callable ids. This gate builds +// the multipackage fixture through the REAL resolver (the same wiring the +// analyzer uses at -a 2), emits v2, and asserts every item on the L2 checklist +// in designing-cldk-changes/references/testing-and-validation.md: +// +// - no dangling endpoints — every edge src/dst is a real callable id +// - every edge carries a non-empty prov naming the resolver +// - callee is a backfilled can:// id on resolved sites, null on unresolved +// - a NAMED expected edge is present (exact src,dst pair, not "graph non-empty") +// - at least one cross-package edge is present +// - the L1 ⊆ L2 superset holds: the tree is unchanged except the callee fill. + +import ( + "testing" + + "github.com/codellm-devkit/codeanalyzer-go/internal/schema" + v2 "github.com/codellm-devkit/codeanalyzer-go/internal/schema/v2" + "github.com/codellm-devkit/codeanalyzer-go/internal/semantic_analysis" + "github.com/codellm-devkit/codeanalyzer-go/internal/syntactic_analysis" +) + +// buildMultipackageL2 builds the fixture AND runs the resolver, so the emitted +// v2 payload carries L2 facts (backfilled callees, call_graph edges). +func buildMultipackageL2(t *testing.T) *v2.Analysis { + t.Helper() + dir := multipackageDir(t) + b := syntactic_analysis.NewSymbolTableBuilder(dir) + st, err := b.Build(nil, true) + if err != nil { + t.Fatalf("symbol table build failed: %v", err) + } + cg := semantic_analysis.NewCallGraphBuilder(dir, b.Fset(), b.Pkgs()) + edges := cg.Build(st) // mutates st in place: backfills callee_signature + app := &schema.GoApplication{SymbolTable: st, CallGraph: edges} + return Emit(app, "multipackage", dir, 2, "gate-test") +} + +// allCallableIDs collects every callable/type id in the emitted tree. +func allCallableIDs(out *v2.Analysis) map[string]bool { + ids := map[string]bool{} + var walk func(c v2.Callable) + walk = func(c v2.Callable) { + ids[c.ID] = true + for _, inner := range c.Callables { + walk(inner) + } + } + for _, mod := range out.Application.SymbolTable { + for _, fn := range mod.Functions { + walk(fn) + } + for _, ty := range mod.Types { + ids[ty.ID] = true + for _, m := range ty.Callables { + walk(m) + } + } + } + return ids +} + +// TestGateL2_MaxLevelAndEdgesPresent locks that L2 emits a non-empty call graph. +func TestGateL2_MaxLevelAndEdgesPresent(t *testing.T) { + out := buildMultipackageL2(t) + if out.MaxLevel != 2 { + t.Errorf("max_level = %d, want 2", out.MaxLevel) + } + if len(out.Application.CallGraph) == 0 { + t.Fatal("L2 should emit call_graph edges; got none") + } +} + +// TestGateL2_NoDanglingEndpoints locks that every edge endpoint resolves to a +// real callable id, and every edge names its resolver in prov. +func TestGateL2_NoDanglingEndpoints(t *testing.T) { + out := buildMultipackageL2(t) + ids := allCallableIDs(out) + + for i, e := range out.Application.CallGraph { + if !ids[e.Src] { + t.Errorf("edge[%d] src is a dangling endpoint: %q", i, e.Src) + } + if !ids[e.Dst] { + t.Errorf("edge[%d] dst is a dangling endpoint: %q", i, e.Dst) + } + if len(e.Prov) == 0 { + t.Errorf("edge[%d] (%s -> %s) has empty prov", i, e.Src, e.Dst) + } + } +} + +// TestGateL2_NamedEdgePresent asserts an EXACT expected edge, so the gate proves +// correctness, not merely a non-empty graph: Worker.Run calls Worker.execute +// (the `go w.execute(...)` goroutine), both as can:// ids. +func TestGateL2_NamedEdgePresent(t *testing.T) { + out := buildMultipackageL2(t) + + const ( + src = "can://go/multipackage/worker/worker.go/example.com/multipackage/worker.Worker/example.com/multipackage/worker.Worker.Run" + dst = "can://go/multipackage/worker/worker.go/example.com/multipackage/worker.Worker/example.com/multipackage/worker.Worker.execute" + ) + found := false + for _, e := range out.Application.CallGraph { + if e.Src == src && e.Dst == dst { + found = true + break + } + } + if !found { + t.Errorf("expected edge %s -> %s not found in call_graph", src, dst) + } +} + +// TestGateL2_CrossPackageEdge locks a specific cross-package edge: main() (in +// package main, main.go) calls server.New (in package server, server/server.go). +// Asserting the exact pair proves a resolved cross-module edge, not just that +// the graph is non-empty. +func TestGateL2_CrossPackageEdge(t *testing.T) { + out := buildMultipackageL2(t) + + const ( + src = "can://go/multipackage/main.go/example.com/multipackage.main" + dst = "can://go/multipackage/server/server.go/example.com/multipackage/server.New" + ) + found := false + for _, e := range out.Application.CallGraph { + if e.Src == src && e.Dst == dst { + found = true + break + } + } + if !found { + t.Errorf("expected cross-package edge %s -> %s not found", src, dst) + } +} + +// TestGateL2_CalleeBackfilled locks the refinement: the goroutine call site in +// Worker.Run has its callee backfilled to Worker.execute's can:// id (not null, +// not the raw v1 signature). +func TestGateL2_CalleeBackfilled(t *testing.T) { + out := buildMultipackageL2(t) + ids := allCallableIDs(out) + + run, ok := out.Application.SymbolTable["worker/worker.go"]. + Types["Worker"].Callables["example.com/multipackage/worker.Worker.Run"] + if !ok { + t.Fatal("Worker.Run missing") + } + if len(run.Body) == 0 { + t.Fatal("Worker.Run should have body call nodes") + } + sawBackfilled := false + for local, node := range run.Body { + if node.Callee == nil { + continue + } + if !ids[*node.Callee] { + t.Errorf("body[%s].callee = %q is not an in-tree callable id", local, *node.Callee) + } + sawBackfilled = true + } + if !sawBackfilled { + t.Error("Worker.Run should have at least one backfilled callee at L2") + } +} + +// TestGateL2_SupersetOfL1 locks L1 ⊆ L2: emitting the SAME model at level 1 vs +// level 2 changes nothing but the callee slot. Every module/type/callable id and +// span is identical; only body callees may go null -> id. +func TestGateL2_SupersetOfL1(t *testing.T) { + dir := multipackageDir(t) + b := syntactic_analysis.NewSymbolTableBuilder(dir) + st, err := b.Build(nil, true) + if err != nil { + t.Fatalf("symbol table build failed: %v", err) + } + // L1 snapshot BEFORE the resolver mutates call sites. + l1 := Emit(&schema.GoApplication{SymbolTable: st, CallGraph: nil}, "multipackage", dir, 1, "gate-test") + + // Now run the resolver (mutates st) and emit L2. + cg := semantic_analysis.NewCallGraphBuilder(dir, b.Fset(), b.Pkgs()) + edges := cg.Build(st) + l2 := Emit(&schema.GoApplication{SymbolTable: st, CallGraph: edges}, "multipackage", dir, 2, "gate-test") + + if len(l1.Application.SymbolTable) != len(l2.Application.SymbolTable) { + t.Fatalf("module count changed L1->L2: %d vs %d", + len(l1.Application.SymbolTable), len(l2.Application.SymbolTable)) + } + for rel, m1 := range l1.Application.SymbolTable { + m2 := l2.Application.SymbolTable[rel] + if m1.ID != m2.ID || m1.Span != m2.Span || m1.Source != m2.Source { + t.Errorf("%s: module identity/span/source changed L1->L2", rel) + } + assertCallablesEqualModuloCallee(t, rel, m1.Functions, m2.Functions) + for name, t1 := range m1.Types { + t2 := m2.Types[name] + if t1.ID != t2.ID || t1.Span != t2.Span || t1.Kind != t2.Kind { + t.Errorf("%s type %s: identity/span/kind changed L1->L2", rel, name) + } + assertCallablesEqualModuloCallee(t, rel+" "+name, t1.Callables, t2.Callables) + } + } +} + +// assertCallablesEqualModuloCallee checks two callable maps are identical except +// that a body node's callee may have gone from null (L1) to a non-null id (L2). +func assertCallablesEqualModuloCallee(t *testing.T, where string, a, b map[string]v2.Callable) { + t.Helper() + if len(a) != len(b) { + t.Errorf("%s: callable count changed L1->L2: %d vs %d", where, len(a), len(b)) + return + } + for sig, c1 := range a { + c2, ok := b[sig] + if !ok { + t.Errorf("%s: callable %q dropped L1->L2", where, sig) + continue + } + if c1.ID != c2.ID || c1.Span != c2.Span || c1.Kind != c2.Kind { + t.Errorf("%s callable %s: identity/span/kind changed L1->L2", where, sig) + } + if len(c1.Body) != len(c2.Body) { + t.Errorf("%s callable %s: body node count changed L1->L2", where, sig) + continue + } + for local, n1 := range c1.Body { + n2 := c2.Body[local] + if n1.Kind != n2.Kind || n1.Span != n2.Span || n1.IsGoroutine != n2.IsGoroutine { + t.Errorf("%s callable %s body[%s]: non-callee field changed L1->L2", where, sig, local) + } + // The one sanctioned mutation: callee null -> id only. + if n1.Callee != nil { + t.Errorf("%s callable %s body[%s]: L1 callee should be null", where, sig, local) + } + } + assertCallablesEqualModuloCallee(t, where+" "+sig, c1.Callables, c2.Callables) + } +} From 67e051eb2d5cc8f297d9ec93f65ffebbfc41c5b5 Mon Sep 17 00:00:00 2001 From: lambertpw Date: Tue, 29 Sep 2026 00:02:08 -0400 Subject: [PATCH 11/17] design: source_file on callable for cross-file methods The real-app eval (dae/openbao/cockroach) surfaced a datamodel collision: v2 nests a method under its receiver type, but Go allows a method to be declared in a different file than its type, so span.bytes computed against the nesting module's source collapse to empty (~3292 cases; cobra clean). Design-loop decision (Go-only for now; cross-language parity deferred): add an optional callable.source_file field, present only when the declaring file differs from the nesting module. span.bytes then index that file's module.source; text = symbol_table[source_file].source[span.bytes] when present, else the nesting module. Containment stays receiver-based. Records the decision in the committed spec and CLAUDE.md; python-sdk source_file support is a tracked follow-up on the SDK rung. --- CLAUDE.md | 17 +++++++- docs/design/specs/v2-l1-emission.md | 66 ++++++++++++++++++++++++++--- 2 files changed, 75 insertions(+), 8 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 1694459..7b3f774 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -12,8 +12,9 @@ v2 vocabulary (parity clause: add at the leaves, never rename shared names). The committed spec is the full transcript: `docs/design/specs/v2-l1-emission.md`. ### Identity & positioning (Group A — reused verbatim by L3/L4) -- **`application.id`** = `can://go/`, where `` is a slug derived from the - **go.mod module path** (deterministic, no flag). L3/L4 build ids on this unchanged. +- **`application.id`** = `can://go/`, where `` is the **`--app-name`** value, + defaulting to the input directory's base name (per the CLI contract; the SDK's Neo4j + backend must use the same anchor). L3/L4 build ids on this unchanged. - **Callable signature** = existing `signatureOf()` output as the last `can://` path segment; receiver folded in for methods. One canonicalizer, unchanged from v1. - **`span.bytes`** = **UTF-8 byte offsets** into `module.source` (Go strings are @@ -35,6 +36,18 @@ committed spec is the full transcript: `docs/design/specs/v2-l1-emission.md`. those returns remain in `return_type` too. Maps Go onto the shared field. - **Closures / function literals** → nested `callables{}` on the enclosing callable (replaces v1 `InnerCallables`); each closure gets its own `can://` id. +- **`source_file`** (optional; added in the 2026-09-29 design loop): Go allows a method + to be declared in a *different file* than its receiver type. The method node stays + nested under its receiver type (containment-by-receiver, above), but its `span.bytes` + then index the **declaring file's** `module.source`, not the nesting module's. When + the declaring file ≠ the nesting module, the callable carries `source_file` = that + file's path (a `symbol_table` key). **Text recovery:** a node's text = + `symbol_table[source_file].source[span.bytes]` if `source_file` present, else the + nesting module's `source`. Emitted only when it differs (absent = same file). This is + the leaf-level fix for the cross-file-method empty-span collision the real-app eval + surfaced. Cross-language parity (C# partial classes, Rust `impl`, Ruby reopened + classes, TS declaration merging) is **deferred** — promote `source_file` into the + shared keystone vocabulary when a second language needs it. ### `call` body node - Typed **`is_goroutine`** and **`is_deferred`** boolean fields (`defer` is net-new vs diff --git a/docs/design/specs/v2-l1-emission.md b/docs/design/specs/v2-l1-emission.md index 209bf07..5ff7201 100644 --- a/docs/design/specs/v2-l1-emission.md +++ b/docs/design/specs/v2-l1-emission.md @@ -33,7 +33,7 @@ application:{ id, kind:"application", symbol_table, call_graph, param_in, param_ ### Identity & positioning (Group A) | Decision | Choice | | --- | --- | -| `application.id` | `can://go/`, `` = slug from **go.mod module path** (no flag) | +| `application.id` | `can://go/`, `` = **`--app-name`** (default: input dir base name), per CLI contract | | Callable signature | existing `signatureOf()` as last id segment; receiver folded in for methods | | `span.bytes` | **UTF-8 byte offsets** into `module.source` | @@ -54,6 +54,39 @@ application:{ id, kind:"application", symbol_table, call_graph, param_in, param_ | --- | --- | | `error_channel[]` | from `error`-typed returns (kept in `return_type` too) | | Closures | nested `callables{}` on enclosing callable (replaces `InnerCallables`) | +| `source_file` | **optional**; emitted **only when** the callable's declaring file differs from the module it is nested under (a method whose receiver type lives in another file). Value = declaring file path relative to input root, same shape as a `symbol_table` key | + +#### Cross-file methods — the `source_file` amendment (design loop, 2026-09-29) + +The eval on real Go apps (dae/openbao/cockroach, ~3292 empty spans; cobra clean) +exposed a datamodel collision: v2 nests a method under its **receiver type** +(containment-by-receiver, § `type` above), but Go lets a method be declared in a +*different file* than its type. The keystone and the **Java** reference both fix +`span.bytes` as offsets into the **nesting module's** `source` (Java's SDK slices a +node against the `JCompilationUnit` it is hydrated under). Java can't produce the +conflict — its methods are lexically inside their type's file — so there was no +precedent; this is a genuine Go divergence brought to the user. + +Decision (Go-only for now; parity deferred — see below): +- **Containment unchanged.** The method node stays under its receiver type's + `callables{}`. The tree still reflects containment-by-receiver. +- **New leaf-level optional field `callable.source_file`**, present only when the + declaring file ≠ the nesting module. Parity-safe (an additive leaf field, never a + rename of shared vocabulary). +- **`span.bytes` for a cross-file method index the `source_file` module's `source`**, + not the nesting module's. +- **Text-recovery rule (the SDK contract):** a node's text = + `symbol_table[source_file].source[span.bytes]` when `source_file` is present, else + `.source[span.bytes]`. The declaring file is always another module + already in `symbol_table` (Go analyzes every `.go` file), so its `source` is present + to slice — the rule is fully data-driven, no extra I/O. + +**Deferred — cross-language parity.** C# partial classes, Rust `impl` blocks, Ruby +reopened classes, and TS declaration merging all hit the same +member-declared-apart-from-its-type shape. `source_file` is scoped to Go here; when a +second language needs it, promote the same field + text-recovery rule into the shared +keystone vocabulary in a design pass (the parity clause forbids a divergent second +spelling). Recorded so the promotion is a known follow-up, not a rediscovery. ### `call` (body node, L1) | Decision | Choice | @@ -65,6 +98,21 @@ application:{ id, kind:"application", symbol_table, call_graph, param_in, param_ `call_graph: [{src, dst, prov, weight}]` at application scope (rename from v1 `{source, target, type, provenance}`). +### CLI conformance (cli-contract.md) +The current `cango` CLI is non-conformant. Bring it to the contract as part of this +train (the SDK facade depends on a uniform flag surface across backends): +- **`--emit `** — output target; `json` default. `neo4j`/`schema` + return "not yet implemented" (non-zero) until the Neo4j child lands, never silent + fallback. +- **`--app-name `** — `application.id`/`:Application` anchor; default = input dir + base name. Precedence: explicit flag > env > default. +- **`--neo4j-uri` / `--neo4j-user` / `--neo4j-password` / `--neo4j-database`** — accepted + and validated now; consumed by the Neo4j child. +- **`-j, --jobs `** — worker parallelism (default CPU cores); output byte-identical + across `-j` values. +- Flag validation: unknown/unimplemented value → non-zero exit + clear message. stdout = + data channel only; all diagnostics to stderr. + ## Release plan - **Train 1 — v2.0.0 (coordinated major, lockstep).** codeanalyzer-go emission rewrite @@ -76,15 +124,21 @@ application:{ id, kind:"application", symbol_table, call_graph, param_in, param_ (`code`, `is_*` booleans folded into `kind`, `is_constructor_call`). - **Deferred (Train 2+):** L3 (CFG/CDG/DDG), L4 (SDG). Reuse Group A ids/spans + Group C edge shape unchanged; no new schema-major coordination. +- **`source_file` SDK follow-up (Train 1, SDK rung):** python-sdk's Go model must add + the optional `source_file` field and its `slice()`/`get_method_body` must resolve + text against `symbol_table[source_file].source` when present (else the nesting + module). Deferred to the SDK rung — the analyzer ships the field first; the SDK + catches up on its own clock (analyzer release must precede the SDK rung regardless). ## Decomposition (tracking shape — to be confirmed with user) Recommended: **parent (in `codellm-devkit/.github`) + one sub-issue per PR**, because the work spans two repos on their own clocks and needs a coordination record: -1. codeanalyzer-go: L1 emission (tree, source, ids, body-with-call-nodes) -2. codeanalyzer-go: L2 emission (call_graph rename @ app scope) -3. codeanalyzer-go: Neo4j relabel -4. python-sdk: Go v2 model remap (accessors stable) -5. release + pin (once both cut) +1. codeanalyzer-go: CLI contract conformance (`--emit`, `--app-name`, `--neo4j-*`, `-j`) +2. codeanalyzer-go: L1 emission (tree, source, ids, body-with-call-nodes) +3. codeanalyzer-go: L2 emission (call_graph rename @ app scope) +4. codeanalyzer-go: Neo4j projection (implements the deferred `--emit neo4j`/`schema`) +5. python-sdk: Go v2 model remap (accessors stable) +6. release + pin (once both cut) Children filed just-in-time as picked up; this spec records the full plan. From ef204e0563f1db90499ceccaae6d7102ce50ea37 Mon Sep 17 00:00:00 2001 From: lambertpw Date: Tue, 29 Sep 2026 00:10:56 -0400 Subject: [PATCH 12/17] fix(emit): source_file for methods declared apart from their type A Go method may be declared in a different file than its receiver type. The v2 tree nests it under the receiver type's module, but its line/col positions are relative to its OWN file, so computing span.bytes against the nesting module's source collapsed them to an empty [n,n] slice (~3292 cases across dae/openbao/cockroach; cobra clean). Compute a cross-file method's span against its declaring file (c.Path) via a per-file lineIndex cache, and record that file in the new optional callable.source_file field. Text recovery: symbol_table[source_file].source [span.bytes] when set, else the nesting module. Closures inherit their enclosing callable's file. Containment stays receiver-based. Adds TestGateL1_CrossFileMethodSourceFile over the multipackage fixture (Server declared in server.go, its Describe method in middleware.go): asserts source_file names the declaring file, the span is non-empty, and it slices to the method text; plus the negative case (in-file method carries no source_file). Design: docs/design/specs/v2-l1-emission.md, CLAUDE.md. --- internal/schema/v2/schema.go | 8 ++ internal/syntactic_analysis/v2emit/emit.go | 86 ++++++++++++++++--- .../syntactic_analysis/v2emit/gate_test.go | 59 +++++++++++++ 3 files changed, 141 insertions(+), 12 deletions(-) diff --git a/internal/schema/v2/schema.go b/internal/schema/v2/schema.go index c3acb27..f4a5dcf 100644 --- a/internal/schema/v2/schema.go +++ b/internal/schema/v2/schema.go @@ -140,6 +140,14 @@ type Callable struct { Span Span `json:"span"` // Signature is the human-readable last path segment of ID (one signatureOf()). Signature string `json:"signature"` + // SourceFile names the callable's declaring file (a symbol_table key) when it + // differs from the module this callable is nested under — the Go case where a + // method is declared in a different file than its receiver type. When set, + // Span.Bytes index symbol_table[SourceFile].source, NOT the nesting module's + // source; text = symbol_table[SourceFile].source[Span.Bytes]. Absent (the + // common case) means the declaring file is the nesting module. See CLAUDE.md + // § Schema decisions → callable.source_file and docs/design/specs/v2-l1-emission.md. + SourceFile string `json:"source_file,omitempty"` // Parameters are the ordered formal parameters. Parameters []Parameter `json:"parameters"` // ReturnType is the joined return type, e.g. "(int, error)". diff --git a/internal/syntactic_analysis/v2emit/emit.go b/internal/syntactic_analysis/v2emit/emit.go index 1fb7d7e..4e08785 100644 --- a/internal/syntactic_analysis/v2emit/emit.go +++ b/internal/syntactic_analysis/v2emit/emit.go @@ -25,6 +25,13 @@ func Emit(app *schema.GoApplication, appName, projectDir string, maxLevel int, a // callees and to map call_graph edge endpoints onto real callable ids. idx := buildSigIndex(appID, app.SymbolTable) + // One lineIndex per source file, read once and shared. A method whose + // receiver type lives in a different file is nested under that type's module + // but must have its span computed against ITS OWN file (CLAUDE.md § Schema + // decisions → callable.source_file); the cache lets emitCallable reach any + // file's index without re-reading it per callable. + srcs := newSourceCache(projectDir) + out := &v2.Analysis{ SchemaVersion: v2.SchemaVersion, Language: v2.Language, @@ -42,7 +49,7 @@ func Emit(app *schema.GoApplication, appName, projectDir string, maxLevel int, a // sorted keeps any incidental ordering (and logs) stable across runs. for _, relPath := range sortedFileKeys(app.SymbolTable) { file := app.SymbolTable[relPath] - out.Application.SymbolTable[relPath] = emitModule(appID, projectDir, relPath, file, idx) + out.Application.SymbolTable[relPath] = emitModule(appID, relPath, file, idx, srcs) } return out } @@ -50,10 +57,14 @@ func Emit(app *schema.GoApplication, appName, projectDir string, maxLevel int, a // emitModule builds a v2 module from a v1 GoFile, reading the file's source so // spans can slice off it. A file that cannot be read still emits its structure // (with empty source and zeroed byte spans) rather than dropping the module. -func emitModule(appID, projectDir, relPath string, file schema.GoFile, idx sigIndex) v2.Module { - source := readSource(projectDir, relPath) - li := newLineIndex(source) - modID := v2.ModuleID(appID, relPath) +// +// modRel is this module's own file path (a symbol_table key). It is threaded to +// callables so a method declared in another file can be detected and given a +// span against its own source plus a source_file pointer. +func emitModule(appID, modRel string, file schema.GoFile, idx sigIndex, srcs *sourceCache) v2.Module { + li := srcs.index(modRel) + source := li.source + modID := v2.ModuleID(appID, modRel) mod := v2.Module{ ID: modID, @@ -70,18 +81,20 @@ func emitModule(appID, projectDir, relPath string, file schema.GoFile, idx sigIn } for name, t := range file.Types { - mod.Types[name] = emitType(modID, li, t, idx) + mod.Types[name] = emitType(modID, modRel, t, idx, srcs) } for sig, fn := range file.Functions { - mod.Functions[sig] = emitCallable(modID, li, fn, idx) + mod.Functions[sig] = emitCallable(modID, modRel, li, fn, idx, srcs) } return mod } // emitType maps a v1 GoType to a v2 type node, collapsing is_interface into the // kind and splitting embedding (base_types) from computed satisfaction. Methods -// resolve into the type's callables{}. -func emitType(modID string, li *lineIndex, t schema.GoType, idx sigIndex) v2.Type { +// resolve into the type's callables{}. modRel is the module the type lives in; +// its methods may be declared in other files (handled in emitCallable). +func emitType(modID, modRel string, t schema.GoType, idx sigIndex, srcs *sourceCache) v2.Type { + li := srcs.index(modRel) typeID := v2.TypeID(modID, t.Signature) out := v2.Type{ ID: typeID, @@ -102,20 +115,37 @@ func emitType(modID string, li *lineIndex, t schema.GoType, idx sigIndex) v2.Typ } } for sig, m := range t.Methods { - out.Callables[sig] = emitCallable(typeID, li, m, idx) + out.Callables[sig] = emitCallable(typeID, modRel, li, m, idx, srcs) } return out } // emitCallable maps a v1 GoCallable to a v2 callable node, including its body // (call nodes) and nested closures. -func emitCallable(parentID string, li *lineIndex, c schema.GoCallable, idx sigIndex) v2.Callable { +// +// modRel and modLI are the nesting module's path and lineIndex. When the +// callable's declaring file (c.Path) differs — a method whose receiver type is +// in another file — its span is computed against its OWN file's lineIndex and +// source_file records that file, per CLAUDE.md § Schema decisions. Nested +// closures share their enclosing callable's file (Go has no cross-file +// closures), so they carry the resolved file down unchanged. +func emitCallable(parentID, modRel string, modLI *lineIndex, c schema.GoCallable, idx sigIndex, srcs *sourceCache) v2.Callable { callID := v2.CallableID(parentID, c.Signature) + + // Resolve the file this callable's positions are relative to. + li := modLI + sourceFile := "" + if c.Path != "" && c.Path != modRel { + li = srcs.index(c.Path) + sourceFile = c.Path + } + out := v2.Callable{ ID: callID, Kind: callableKind(c), Span: lineSpan(li, c.StartLine, c.EndLine), Signature: c.Signature, + SourceFile: sourceFile, Parameters: emitParams(li, c.Parameters), ReturnType: c.ReturnType, Body: emitBody(callID, li, c.CallSites, idx), @@ -128,8 +158,14 @@ func emitCallable(parentID string, li *lineIndex, c schema.GoCallable, idx sigIn } if len(c.InnerCallables) > 0 { out.Callables = make(map[string]v2.Callable, len(c.InnerCallables)) + // Closures live in the same file as their enclosing callable, so the + // resolved file (modRel-or-c.Path) becomes their nesting file. + enclRel := modRel + if sourceFile != "" { + enclRel = sourceFile + } for sig, inner := range c.InnerCallables { - out.Callables[sig] = emitCallable(callID, li, inner, idx) + out.Callables[sig] = emitCallable(callID, enclRel, li, inner, idx, srcs) } } return out @@ -296,6 +332,32 @@ func readSource(projectDir, relPath string) string { return string(data) } +// sourceCache reads each source file once and memoizes its lineIndex, keyed by +// path relative to projectDir. The emitter needs a file's index in two places — +// when emitting the module itself, and when a method declared in that file is +// nested under a type in a DIFFERENT module — so the cache avoids re-reading a +// file per cross-file method. A file that cannot be read caches an empty index +// (so spans degrade to zero rather than re-attempting the read). +type sourceCache struct { + projectDir string + byPath map[string]*lineIndex +} + +func newSourceCache(projectDir string) *sourceCache { + return &sourceCache{projectDir: projectDir, byPath: make(map[string]*lineIndex)} +} + +// index returns the lineIndex for relPath, reading and indexing the file on +// first use. The returned index carries the file's source (li.source). +func (c *sourceCache) index(relPath string) *lineIndex { + if li, ok := c.byPath[relPath]; ok { + return li + } + li := newLineIndex(readSource(c.projectDir, relPath)) + c.byPath[relPath] = li + return li +} + func sortedFileKeys(m map[string]schema.GoFile) []string { keys := make([]string, 0, len(m)) for k := range m { diff --git a/internal/syntactic_analysis/v2emit/gate_test.go b/internal/syntactic_analysis/v2emit/gate_test.go index db1317a..b021a9d 100644 --- a/internal/syntactic_analysis/v2emit/gate_test.go +++ b/internal/syntactic_analysis/v2emit/gate_test.go @@ -227,6 +227,65 @@ func TestGateL1_CrossPackageStructure(t *testing.T) { } } +// TestGateL1_CrossFileMethodSourceFile locks the source_file amendment +// (CLAUDE.md § Schema decisions → callable.source_file): a method declared in a +// different file than its receiver type stays nested under the type, but its +// span.bytes index its OWN file's source and it carries source_file naming that +// file. In the fixture, Server is declared in server/server.go while its method +// (s Server) Describe() is declared in server/middleware.go. +// +// The text-recovery rule is the contract: a node's text = +// symbol_table[source_file].source[span.bytes] when source_file is present. This +// test slices the declaring module's source by the span and asserts it yields +// the method text — the check that regressed to "" before the fix (empty span +// against the wrong file's source). +func TestGateL1_CrossFileMethodSourceFile(t *testing.T) { + out := buildMultipackageV2(t) + + // Server's type node lives in server/server.go. + server := mustModule(t, out, "server/server.go") + stype, ok := server.Types["Server"] + if !ok { + t.Fatalf("Server type missing; types=%v", keys(server.Types)) + } + const describeID = "example.com/multipackage/server.Server.Describe" + m, ok := stype.Callables[describeID] + if !ok { + t.Fatalf("Server.Describe missing; callables=%v", keys(stype.Callables)) + } + + // It is declared in a different file, so source_file must name that file. + const declFile = "server/middleware.go" + if m.SourceFile != declFile { + t.Errorf("Describe.source_file = %q, want %q (declared apart from its type)", m.SourceFile, declFile) + } + + // The span must be non-empty (the pre-fix bug collapsed it to [n,n]). + if m.Span.Bytes[0] >= m.Span.Bytes[1] { + t.Errorf("Describe span is empty %v — cross-file method span not computed against its own file", m.Span.Bytes) + } + + // Text-recovery rule: slice the DECLARING module's source by the span. + decl := mustModule(t, out, declFile) + text := m.Span.Slice(decl.Source) + if !strings.Contains(text, "func (s Server) Describe()") { + t.Errorf("Describe span over %s does not slice to the method text; got %q", declFile, text) + } + + // Negative: an in-file method (Server.Addr, declared in server.go with its + // type) must NOT carry source_file — absence encodes "same file". + addr, ok := stype.Callables["example.com/multipackage/server.Server.Addr"] + if !ok { + t.Fatalf("Server.Addr missing; callables=%v", keys(stype.Callables)) + } + if addr.SourceFile != "" { + t.Errorf("Addr.source_file = %q, want empty (declared in its type's own file)", addr.SourceFile) + } + if got := addr.Span.Slice(server.Source); !strings.Contains(got, "func (s *Server) Addr()") { + t.Errorf("Addr span over its own module does not slice to the method text; got %q", got) + } +} + // TestGateL1_Envelope locks the manifest envelope at L1. func TestGateL1_Envelope(t *testing.T) { out := buildMultipackageV2(t) From 06d1cbb85b8581d1661dfccfdeac9ec0fa2b5ee2 Mon Sep 17 00:00:00 2001 From: lambertpw Date: Tue, 29 Sep 2026 00:55:26 -0400 Subject: [PATCH 13/17] fix(symtab): skip decls whose file is outside the input root cgo synthesizes wrapper functions (_Cfunc_*, _Cgo_*, _cgo_runtime_*) into a generated file under $GOCACHE; go/packages resolves their token.Pos to that cache path. Emitting them produced callables whose source_file pointed at the build cache and whose span could not be sliced against any project module (ISSUE-2 in the real-app eval: 85 such callables on cockroach, 106 validation failures at L1 and L2; cobra/dae/openbao have no cgo and were unaffected). These are toolchain artifacts, not project source. buildCallable now returns nil when the declaration's file is not within projectDir, at the point its (cache) file is first known; both callers already skip nil. Adds utils.IsWithin and a unit test. Verified: cockroach L1 106->0 failures, 85 dropped signatures all cgo glue, real pkg/geo/geos project code preserved. --- internal/syntactic_analysis/symbol_table.go | 12 ++++++++++ internal/utils/fs.go | 18 ++++++++++++++ internal/utils/fs_test.go | 26 +++++++++++++++++++++ 3 files changed, 56 insertions(+) diff --git a/internal/syntactic_analysis/symbol_table.go b/internal/syntactic_analysis/symbol_table.go index c7f48ce..93d32ab 100644 --- a/internal/syntactic_analysis/symbol_table.go +++ b/internal/syntactic_analysis/symbol_table.go @@ -492,6 +492,18 @@ func (b *SymbolTableBuilder) buildCallable( decl *ast.FuncDecl, ) *schema.GoCallable { name := decl.Name.Name + + // Skip declarations whose file is outside the input root. cgo synthesizes + // wrapper functions (_Cfunc_*, _Cgo_*, _cgo_runtime_*) into a generated file + // under $GOCACHE; go/packages resolves their token.Pos to that cache path. + // They are toolchain artifacts, not project source — emitting them yields + // callables whose span cannot be sliced against any project module. Drop + // them here, at the point their (cache) file is first known. + if declFile := b.fset.File(decl.Pos()); declFile == nil || + !utils.IsWithin(b.projectDir, declFile.Name()) { + return nil + } + isExported := unicode.IsUpper(rune(name[0])) pos := b.fset.Position(decl.Pos()) end := b.fset.Position(decl.End()) diff --git a/internal/utils/fs.go b/internal/utils/fs.go index e46da2e..541127c 100644 --- a/internal/utils/fs.go +++ b/internal/utils/fs.go @@ -35,6 +35,24 @@ func RelativePath(root, path string) string { return rel } +// IsWithin reports whether path lies inside root (or is root itself). It is the +// test for "this declaration's file is part of the analyzed project": a path +// whose relative form escapes root (starts with ".." or is absolute) is outside. +// cgo synthesizes wrapper functions (_Cfunc_*, _Cgo_*) into files under +// $GOCACHE, so their declaring file resolves outside the input root — IsWithin +// is how the symbol-table builder tells those toolchain artifacts from project +// source. +func IsWithin(root, path string) bool { + rel, err := filepath.Rel(root, path) + if err != nil { + return false + } + if rel == ".." || strings.HasPrefix(rel, ".."+string(filepath.Separator)) { + return false + } + return !filepath.IsAbs(rel) +} + // FileHash returns the SHA-256 hex digest of the file at path. func FileHash(path string) (string, error) { f, err := os.Open(path) diff --git a/internal/utils/fs_test.go b/internal/utils/fs_test.go index 91df3d4..4e12166 100644 --- a/internal/utils/fs_test.go +++ b/internal/utils/fs_test.go @@ -53,6 +53,32 @@ func TestIsVendored(t *testing.T) { } } +// ── IsWithin ───────────────────────────────────────────────────────────────── + +func TestIsWithin(t *testing.T) { + root := filepath.Join(string(filepath.Separator), "home", "user", "proj") + tests := []struct { + path string + want bool + }{ + // Inside the project root. + {filepath.Join(root, "main.go"), true}, + {filepath.Join(root, "pkg", "geo", "geos", "geos.go"), true}, + {root, true}, // root itself resolves to "." + // Outside: the go-build cache cgo synthesizes into — the ISSUE-2 shape. + {filepath.Join(string(filepath.Separator), "home", "user", "Library", "Caches", "go-build", "bb", "hash-d"), false}, + // A sibling directory sharing a prefix but not nested under root. + {filepath.Join(string(filepath.Separator), "home", "user", "proj-other", "x.go"), false}, + // A parent directory. + {filepath.Join(string(filepath.Separator), "home", "user"), false}, + } + for _, tc := range tests { + if got := utils.IsWithin(root, tc.path); got != tc.want { + t.Errorf("IsWithin(%q, %q) = %v, want %v", root, tc.path, got, tc.want) + } + } +} + // ── FileHash ────────────────────────────────────────────────────────────────── func TestFileHash_Deterministic(t *testing.T) { From adb07ff5a7078fec88a2436099027720690272b3 Mon Sep 17 00:00:00 2001 From: lambertpw Date: Wed, 30 Sep 2026 13:21:44 -0400 Subject: [PATCH 14/17] docs: v2 analysis.json structure reference (keys only) A quick key-skeleton map of the canonical v2 output: the node tree, the recursive callable node, span shape, presence rules (always-present vs omitempty vs emitted-when-true), and can:// id shapes. Authored from the JSON tags in internal/schema/v2/schema.go; links to v2-l1-emission.md for the decisions. --- docs/design/specs/v2-schema-structure.md | 108 +++++++++++++++++++++++ 1 file changed, 108 insertions(+) create mode 100644 docs/design/specs/v2-schema-structure.md diff --git a/docs/design/specs/v2-schema-structure.md b/docs/design/specs/v2-schema-structure.md new file mode 100644 index 0000000..f833b43 --- /dev/null +++ b/docs/design/specs/v2-schema-structure.md @@ -0,0 +1,108 @@ +# v2 `analysis.json` — structure reference (keys only) + +The key skeleton of the canonical CLDK v2 output (`--analysis-schema 2`). This is a +quick map of the node tree and every field; the *decisions* behind the shape live in +[`v2-l1-emission.md`](./v2-l1-emission.md) and `CLAUDE.md` § Schema decisions, and the +authoritative Go definitions are `internal/schema/v2/schema.go`. Field names here are +the JSON tags from that file. + +## Tree + +``` +├─ schema_version # "2.0.0" +├─ language # "go" +├─ max_level # 1 (symbol table) | 2 (+ call graph) +├─ analyzer +│ ├─ name # "codeanalyzer-go" +│ └─ version +└─ application + ├─ id # "can://go/" + ├─ kind # "application" + ├─ symbol_table # { "": module } + │ └─ + │ ├─ id + │ ├─ kind # "module" + │ ├─ span { start, end, bytes } + │ ├─ package + │ ├─ source # whole file text, once per module + │ ├─ content_hash # optional + │ ├─ imports[] { name, path, alias?, span } + │ ├─ types # { "": type } + │ │ └─ + │ │ ├─ id + │ │ ├─ kind # struct | interface | alias | defined + │ │ ├─ span { start, end, bytes } + │ │ ├─ base_types[] # optional: embedded type ids (explicit) + │ │ ├─ interfaces[] # optional: interfaces computed-satisfied + │ │ ├─ fields # optional: { "": field } + │ │ │ └─ { id, kind, type, span } + │ │ └─ callables # { "": callable } + │ │ └─ # ← callable node, defined below + │ └─ functions # { "": callable } + │ └─ # ← callable node, defined below + └─ call_graph[] # L2+; empty [] at L1 + └─ + ├─ src # can:// callable id + ├─ dst # can:// callable id + ├─ prov[] # optional: e.g. ["go/types"] + └─ weight # optional +``` + +## `` node (recursive) + +The same shape whether it sits under a type's `callables` (methods), a module's +`functions` (package-level functions), or another callable's `callables` (closures). + +``` + +├─ id +├─ kind # function | method | lambda +├─ signature +├─ source_file # optional: set only for a method declared in a +│ # different file than its receiver type; its +│ # span.bytes then index symbol_table[source_file].source +├─ span { start, end, bytes } +├─ parameters[] { name, type, span, is_variadic? } +├─ return_type # optional +├─ error_channel[] # optional: error-typed returns +├─ metrics # optional: { cyclomatic: N } +├─ callables # optional: nested closures (same shape) +└─ body # { "": call node } + └─ + ├─ kind # "call" + ├─ span { start, end, bytes } + ├─ callee # null | can:// id (key always present) + ├─ is_goroutine # optional: present only when true + └─ is_deferred # optional: present only when true +``` + +## `span` + +``` +span +├─ start [line, col] +├─ end [line, col] +└─ bytes [from, to] # UTF-8 byte offsets into the owning module.source +``` + +## Presence rules + +- **Always present:** every `id`, `kind`, `span`; the envelope fields; `symbol_table`, + `types`, `callables`, `functions`, `body` maps (may be empty `{}`); `call_graph` (empty + `[]` at L1); a body node's `callee` key (value `null` when unresolved/external). +- **`omitempty` — emitted only when non-empty:** `content_hash`, `imports[].alias`, + `base_types`, `interfaces`, `fields`, `source_file`, `return_type`, `error_channel`, + `metrics`, a callable's nested `callables`, `parameters[].is_variadic`, edge `prov` / + `weight`. +- **Emitted only when `true`:** `is_goroutine`, `is_deferred`. + +## `id` shape + +``` +application can://go/ +module can://go// +type can://go/// +callable can://go//// + (functions omit the segment) +body call keyed by ":" within the callable's body map +``` From af3510adbce0956c6dffb58ff9716476deded498 Mon Sep 17 00:00:00 2001 From: lambertpw Date: Wed, 30 Sep 2026 13:26:15 -0400 Subject: [PATCH 15/17] docs(readme): document v2 schema and analysis.json generation - Add a step-by-step 'Generating analysis.json for a Go app' section (--analysis-schema 2 -a 2 -o , --app-name, verify, cgo note). - Document the canonical v2 output shape alongside the legacy v1 shape, with a full JSON example verified against real emitter output. - Fix the stale CLI options block (add --analysis-schema, --app-name). --- README.md | 161 +++++++++++++++++++++++++++++++++++++++++++++++++++++- 1 file changed, 159 insertions(+), 2 deletions(-) diff --git a/README.md b/README.md index 6cf4285..da1235a 100644 --- a/README.md +++ b/README.md @@ -131,6 +131,8 @@ Aliases: Flags: -a, --analysis-level int Analysis level: 1=symbol table only, 2=+resolver call graph (default 1) + --analysis-schema int Output schema major: 1=legacy v1 shape (default), 2=canonical v2 tree (default 1) + --app-name string Application anchor name for can:// ids and Neo4j :Application (default: input dir name) -c, --cache-dir string Cache directory (default: ~/.cldk/go-cache) --codeql Enable CodeQL framework-based call graph (level 2, stub) --eager Force clean rebuild (ignore cache) @@ -178,6 +180,58 @@ cango -i ./my-go-project --eager cango -i ./my-go-project -a 2 -vv ``` +## Generating `analysis.json` for a Go app + +The end-to-end steps to produce an `analysis.json` file for any Go project. + +**1. Make sure you have `cango`.** Either install a released binary (see +[Installation](#installation)) or build from source: + +```bash +git clone https://github.com/codellm-devkit/codeanalyzer-go +cd codeanalyzer-go +go build -o cango ./cmd/codeanalyzer +``` + +**2. Point it at the project root** (the directory containing the app's `go.mod`). +`cango` resolves imports with the Go type checker, so the target must be a normal Go +module. Write the output to a directory with `-o`; `cango` names the file +`analysis.json` inside it. + +```bash +# Canonical v2 schema, Level 2 (symbol table + call graph): +cango -i /path/to/go/app --analysis-schema 2 -a 2 -o /path/to/output/ +# → writes /path/to/output/analysis.json +``` + +> **Pick the schema explicitly.** `--analysis-schema` defaults to `1` (the legacy v1 +> shape). Pass `--analysis-schema 2` for the current **canonical v2** tree (the shape +> the Python SDK's v2 loader and the sections below describe). Omit `-o` to stream the +> JSON to stdout instead of writing a file. + +**3. (Optional) name the application anchor.** `--app-name` sets the `` in every +`can://go//…` id and defaults to the input directory's base name. Set it when the +directory name isn't the identity you want: + +```bash +cango -i /path/to/go/app --analysis-schema 2 -a 2 --app-name myapp -o ./out/ +``` + +**4. Verify.** A successful run exits `0` and produces a JSON document whose envelope +carries `"schema_version": "2.0.0"`, `"language": "go"`, `"max_level": 2`, and an +`application` tree under `application.symbol_table`: + +```bash +cango -i /path/to/go/app --analysis-schema 2 -a 2 -o ./out/ && echo "exit=$?" +python3 -c "import json; d=json.load(open('out/analysis.json')); print(d['schema_version'], d['max_level'], len(d['application']['symbol_table']), 'modules')" +``` + +**Notes for larger apps.** Analysis parallelism defaults to your CPU count (tune with +`-j`); a warm cache (`--cache-dir`, default `~/.cldk/go-cache`) makes re-runs fast, and +`--eager` forces a clean rebuild. Projects that use **cgo** (`import "C"`) are supported: +the toolchain-synthesized wrapper functions are correctly excluded from the output (they +are build artifacts, not project source), so only your own declarations are emitted. + ## Analysis levels | Level | Flag | What runs | Status | @@ -192,7 +246,110 @@ cango -i ./my-go-project -a 2 -vv ## Output schema -The root object is `GoApplication`: +`cango` emits one of two schema shapes, selected by `--analysis-schema`: + +- **`--analysis-schema 2`** — the **canonical CLDK v2** shape (detailed below); the shape the + v2 Python SDK loader consumes and the [step-by-step section + above](#generating-analysisjson-for-a-go-app) produces. +- **`--analysis-schema 1`** (default) — the **legacy v1** `GoApplication` shape, kept for + backward compatibility. + +### Canonical v2 shape (`--analysis-schema 2`) + +The document is a manifest envelope wrapping one `application` containment tree. Every +node carries a `can://go//…` `id`, a `kind`, and a `span`: + +```json +{ + "schema_version": "2.0.0", + "language": "go", + "max_level": 2, + "analyzer": { "name": "codeanalyzer-go", "version": "0.1.0" }, + "application": { + "id": "can://go/myapp", + "kind": "application", + "symbol_table": { + "pkg/greeter/greeter.go": { + "id": "can://go/myapp/pkg/greeter/greeter.go", + "kind": "module", + "package": "greeter", + "span": { "start": [1, 1], "end": [40, 2], "bytes": [0, 812] }, + "source": "package greeter\n\n...", + "content_hash": "…", + "imports": [ { "name": "fmt", "path": "fmt", "span": {…} } ], + "types": { + "Greeter": { + "id": "can://go/myapp/pkg/greeter/greeter.go/example.com/pkg/greeter.Greeter", + "kind": "struct", + "span": { "start": [5, 1], "end": [7, 2], "bytes": [17, 63] }, + "fields": { + "Prefix": { "id": "…/Prefix", "kind": "field", "type": "string", "span": {…} } + }, + "callables": { + "example.com/pkg/greeter.Greeter.Greet": { + "id": "can://go/myapp/pkg/greeter/greeter.go/example.com/pkg/greeter.Greeter/example.com/pkg/greeter.Greeter.Greet", + "kind": "method", + "signature": "example.com/pkg/greeter.Greeter.Greet", + "span": { "start": [9, 1], "end": [11, 2], "bytes": [65, 140] }, + "parameters": [ { "name": "name", "type": "string", "span": {…} } ], + "return_type": "string", + "error_channel": ["error"], + "metrics": { "cyclomatic": 1 }, + "body": { + "10:2": { + "kind": "call", + "span": { "start": [10, 2], "end": [10, 30], "bytes": [90, 118] }, + "callee": "can://go/myapp/pkg/greeter/greeter.go/…/example.com/pkg/greeter.format" + } + } + } + } + } + }, + "functions": { + "example.com/main.main": { "id": "…", "kind": "function", "signature": "…", "body": {…} } + } + } + }, + "call_graph": [ + { + "src": "can://go/myapp/main.go/…/example.com/main.main", + "dst": "can://go/myapp/pkg/greeter/greeter.go/…/example.com/pkg/greeter.Greeter.Greet", + "prov": ["go/types"], + "weight": 1 + } + ] + } +} +``` + +Key v2 schema properties: +- **Envelope** — `schema_version` (`"2.0.0"`), `language` (`"go"`), `max_level` (1 or 2), + and `analyzer{name,version}` wrap a single `application` node. +- **One containment tree** — `application → module → type → callable → body`, keyed as + `symbol_table` (modules, by **project-relative file path**), `types`, `callables`, `body`. +- **`id`** — every node carries a durable `can://go////` id; + `` is the `--app-name` anchor. Body call nodes use `` keyed by `line:col`. +- **`kind`** — `module` · `type` kinds `struct | interface | alias | defined` (collapses v1's + `is_interface`) · `callable` kinds `function | method | lambda` · body `call`. +- **`base_types: []`** (type, optional) — embedded type ids (struct/interface embedding — the + explicit spine), as opposed to interfaces the type is *computed* to satisfy. +- **`span`** — `{start:[line,col], end:[line,col], bytes:[from,to]}`; `bytes` are **UTF-8 byte + offsets** into the owning `module.source` (source is stored once per module, every node slices it). +- **`source_file`** (callable, optional) — set when a method is declared in a *different file* + than its receiver type; the method stays nested under the type, and its `span.bytes` index + `symbol_table[source_file].source` instead of the nesting module's. +- **`error_channel: []`** — populated from `error`-typed returns (the returns also stay in + `return_type`). **Closures** nest as `callables{}` on the enclosing callable. +- **`body` call nodes** — `callee` is the sanctioned `null → id` slot: `null` at L1 (and for + external/stdlib callees), a `can://` node id once resolved at L2. `is_goroutine` / `is_deferred` + are boolean flags emitted **only when true** (a plain call omits them), for `go f()` / `defer f()`. +- **`call_graph` edges** — `{src, dst, prov, weight}`; `src`/`dst` are `can://` node ids that + exist in the tree (not raw signatures), `prov` is resolver provenance, e.g. `["go/types"]`. + +### Legacy v1 shape (`--analysis-schema 1`, default) + +The v1 root object is `GoApplication`: ```json { @@ -226,7 +383,7 @@ The root object is `GoApplication`: } ``` -Key schema properties: +Key v1 schema properties: - `symbol_table` — keyed by **file path relative to the project root** (never absolute) - `classes` — JSON key for types (spine compatibility with Java/Python schemas); value is `GoType` - `module_name` — JSON key for the Go package name (spine compatibility) From 0289c1c9c96f983caf227fe9f8a70cbfcea6ebff Mon Sep 17 00:00:00 2001 From: lambertpw Date: Thu, 1 Oct 2026 15:46:31 -0400 Subject: [PATCH 16/17] fix(emit): sort call_graph edges for deterministic output The v2 call_graph edge list was emitted in Go-map iteration order, so two runs of the same input (even at the same -j) produced byte-differing analysis.json, failing the release determinism gate. Sort the emitted edges by (src, dst, prov, weight) in emitCallGraph, matching the sorted-key discipline already used for module output. Add a determinism gate test that emits a reversed edge slice and asserts identical serialized call_graph. --- .../v2emit/determinism_test.go | 76 +++++++++++++++++++ internal/syntactic_analysis/v2emit/emit.go | 23 ++++++ 2 files changed, 99 insertions(+) create mode 100644 internal/syntactic_analysis/v2emit/determinism_test.go diff --git a/internal/syntactic_analysis/v2emit/determinism_test.go b/internal/syntactic_analysis/v2emit/determinism_test.go new file mode 100644 index 0000000..a7c758a --- /dev/null +++ b/internal/syntactic_analysis/v2emit/determinism_test.go @@ -0,0 +1,76 @@ +package v2emit + +// Determinism gate for the v2 emitter. +// +// release-gates.md requires v2 output to be byte-identical across runs and +// across -j values. The symbol-table tree is already order-stable (modules +// iterated via sortedFileKeys), but the call_graph edge list is assembled by +// ranging a Go map in CallGraphBuilder.Build — random iteration order — and +// emitCallGraph preserved that order. Two runs of the same input therefore +// produced call_graph arrays in different orders. +// +// This gate emits the SAME resolved model twice and asserts the serialized +// call_graph is identical both times, so a regression to unsorted edge output +// fails here rather than only in an external CLI diff. + +import ( + "encoding/json" + "testing" + + "github.com/codellm-devkit/codeanalyzer-go/internal/schema" + v2 "github.com/codellm-devkit/codeanalyzer-go/internal/schema/v2" + "github.com/codellm-devkit/codeanalyzer-go/internal/semantic_analysis" + "github.com/codellm-devkit/codeanalyzer-go/internal/syntactic_analysis" +) + +// TestDeterminism_CallGraphStableAcrossEmits locks that emitting the same model +// twice yields a byte-identical call_graph. emitCallGraph must impose a total +// order on the edges rather than inherit map-iteration order. +func TestDeterminism_CallGraphStableAcrossEmits(t *testing.T) { + dir := multipackageDir(t) + b := syntactic_analysis.NewSymbolTableBuilder(dir) + st, err := b.Build(nil, true) + if err != nil { + t.Fatalf("symbol table build failed: %v", err) + } + cg := semantic_analysis.NewCallGraphBuilder(dir, b.Fset(), b.Pkgs()) + edges := cg.Build(st) + app := &schema.GoApplication{SymbolTable: st, CallGraph: edges} + + // Emit twice from the SAME model. If the edge order is derived from the + // edges slice without a stable sort, repeated marshals still match (same + // slice), so to actually exercise order-independence we marshal a shuffled + // copy of the edge slice too. + first := mustMarshalCallGraph(t, Emit(app, "multipackage", dir, 2, "det-test")) + + // Reverse the input edge order and re-emit: a correct emitter sorts, so the + // serialized call_graph must be unchanged. + rev := make([]schema.GoCallEdge, len(edges)) + for i, e := range edges { + rev[len(edges)-1-i] = e + } + appRev := &schema.GoApplication{SymbolTable: st, CallGraph: rev} + second := mustMarshalCallGraph(t, Emit(appRev, "multipackage", dir, 2, "det-test")) + + if first != second { + t.Errorf("call_graph not order-independent: emitCallGraph must sort edges to a total order\nfirst=%s\nsecond=%s", first, second) + } +} + +func mustMarshalCallGraph(t *testing.T, out *v2.Analysis) string { + t.Helper() + type cgOnly struct { + Application struct { + CallGraph json.RawMessage `json:"call_graph"` + } `json:"application"` + } + data, err := json.Marshal(out) + if err != nil { + t.Fatalf("marshal failed: %v", err) + } + var c cgOnly + if err := json.Unmarshal(data, &c); err != nil { + t.Fatalf("unmarshal failed: %v", err) + } + return string(c.Application.CallGraph) +} diff --git a/internal/syntactic_analysis/v2emit/emit.go b/internal/syntactic_analysis/v2emit/emit.go index 4e08785..5b61f9b 100644 --- a/internal/syntactic_analysis/v2emit/emit.go +++ b/internal/syntactic_analysis/v2emit/emit.go @@ -4,6 +4,7 @@ import ( "os" "path/filepath" "sort" + "strings" "github.com/codellm-devkit/codeanalyzer-go/internal/schema" v2 "github.com/codellm-devkit/codeanalyzer-go/internal/schema/v2" @@ -256,6 +257,28 @@ func emitCallGraph(edges []schema.GoCallEdge, idx sigIndex) []v2.Edge { Weight: e.Weight, }) } + + // Impose a total order on the edges so the output is byte-identical across + // runs and across -j values (release-gates.md determinism gate). The upstream + // edges slice is assembled by ranging a Go map (random iteration order), so + // without this sort two runs of the same input emit the call_graph in + // different orders. Sort on the full tuple — (src, dst) is the identity but + // prov/weight tie-break so merged edges from multiple backends also order + // stably. + sort.Slice(out, func(i, j int) bool { + a, b := out[i], out[j] + if a.Src != b.Src { + return a.Src < b.Src + } + if a.Dst != b.Dst { + return a.Dst < b.Dst + } + ap, bp := strings.Join(a.Prov, ","), strings.Join(b.Prov, ",") + if ap != bp { + return ap < bp + } + return a.Weight < b.Weight + }) return out } From 58a628eba03fe20fee1bc2db65597b65d085451c Mon Sep 17 00:00:00 2001 From: lambertpw Date: Sat, 3 Oct 2026 11:57:47 -0400 Subject: [PATCH 17/17] docs: tracking records for the v2 migration epic and L1/L2 children Epic codellm-devkit/.github#94 and its two codeanalyzer-go children (#8 L1 tree emission, #9 L2 call_graph refinement), plus the PR body. Corrects child-1's app-id anchor to --app-name (matches the code and CLAUDE.md) and aligns issue headings to the live feature_request form. Co-Authored-By: Claude Opus 4.8 --- docs/design/issues/child-1-l1-emission.md | 8 +-- docs/design/issues/child-2-l2-refinement.md | 39 ++++++++++++ docs/design/issues/epic-v2-migration.md | 67 +++++++++++++++++++++ docs/design/issues/parent-v2-migration.md | 6 +- docs/design/issues/pr-body-v2-l1-l2.md | 52 ++++++++++++++++ 5 files changed, 165 insertions(+), 7 deletions(-) create mode 100644 docs/design/issues/child-2-l2-refinement.md create mode 100644 docs/design/issues/epic-v2-migration.md create mode 100644 docs/design/issues/pr-body-v2-l1-l2.md diff --git a/docs/design/issues/child-1-l1-emission.md b/docs/design/issues/child-1-l1-emission.md index 7cb806d..0a33e82 100644 --- a/docs/design/issues/child-1-l1-emission.md +++ b/docs/design/issues/child-1-l1-emission.md @@ -16,9 +16,9 @@ decision the later levels build on. Do L1 first and get the symbol-table gate gr before touching the L2 call graph. The parser/resolver stay; only the emission layer changes (golden rule: replace what serializes, keep what computes). -**Describe the solution you'd like.** +**Describe the solution you'd like** -- [ ] `application.id` = `can://go/`, `` derived from the go.mod module path +- [ ] `application.id` = `can://go/`, `` = `--app-name` (default: input dir base name) - [ ] Durable `can://` ids on module/type/callable; `signatureOf()` is the last segment - [ ] `span: {start:[l,c], end:[l,c], bytes:[from,to]}` with UTF-8 byte offsets; drop flat `start_line`/`end_line` - [ ] `source` emitted once per module; drop per-callable `code` @@ -26,12 +26,12 @@ changes (golden rule: replace what serializes, keep what computes). - [ ] `base_types[]` = embedded ids, `interfaces[]` = computed structural satisfaction - [ ] `body{}` holds `call` nodes with `callee:null`, typed `is_goroutine`/`is_deferred`; closures nest in `callables{}` -**Describe alternatives you've considered.** +**Describe alternatives you've considered** Does NOT emit the `call_graph` (that is the L2 child), nor any L3/L4 `body` statements or edges. `body{}` at L1 holds only `call` nodes. -**Additional context.** +**Additional context** - Emitter approach: walk the existing in-memory model, add a v2 emitter behind a flag; keep v1 emitter as the compat shim during transition. - One genuinely new datum: thread parser byte offsets into `span.bytes`. diff --git a/docs/design/issues/child-2-l2-refinement.md b/docs/design/issues/child-2-l2-refinement.md new file mode 100644 index 0000000..51c0ee5 --- /dev/null +++ b/docs/design/issues/child-2-l2-refinement.md @@ -0,0 +1,39 @@ + + +**Is your feature request related to a problem? Please describe.** + +L1 emits the v2 containment tree with `call` body nodes whose `callee` is `null` — the +sanctioned refinement slot. L2 fills it: resolve each call site and each call-graph edge +endpoint from a v1 `signatureOf()` string to the durable `can://` id of the callable it +names. L2 is pure refinement over the L1 tree — it computes no new facts, it only +translates identity. The parser/resolver stay; only the emission layer gains the +signature→id index (golden rule: replace what serializes, keep what computes). + +**Describe the solution you'd like** + +- [ ] `buildSigIndex` maps every v1 callable signature (functions, methods, nested closures) → its `can://` id, using the SAME id builders `emitCallable` uses (byte-identical ids) +- [ ] `callee` on each `call` body node backfilled from `null` to the resolved `can://` id +- [ ] `call_graph` emitted at application scope as `[{src, dst, prov, weight}]` +- [ ] External/stdlib callees (outside the project) resolve to no endpoint — never a dangling edge +- [ ] L2 gate green: every edge endpoint and every non-null `callee` resolves to a real in-tree node id +- [ ] Determinism: `call_graph` edge order stable across runs (sorted before emit) + +**Describe alternatives you've considered** + +Does NOT compute any new call-graph facts — it reuses the v1 call graph and only +translates endpoints to `can://` ids. Does NOT emit L3/L4 `body` statements or edges. + +**Additional context** + +- v1 signatures are globally unique (a method signature embeds its receiver), so the flat signature→id map has no collisions. +- The index mirrors `emitModule`/`emitType`/`emitCallable` parent-id computation exactly; drift there would silently dangle endpoints. +- Determinism risk: edge order was nondeterministic before sorting — the determinism gate fails without a stable sort. +- Depends on the L1 child: L2 refines the tree L1 produces; file/land L1 first. +- Full transcript in the spec; schema decisions in repo `CLAUDE.md`. diff --git a/docs/design/issues/epic-v2-migration.md b/docs/design/issues/epic-v2-migration.md new file mode 100644 index 0000000..1072e35 --- /dev/null +++ b/docs/design/issues/epic-v2-migration.md @@ -0,0 +1,67 @@ + + +## Spec + +https://github.com/codellm-devkit/codeanalyzer-go/blob/main/docs/design/specs/v2-l1-emission.md + +## Summary + +Migrate `codeanalyzer-go` from the legacy v1 output to the **canonical schema v2** (one +additive tree: `application → module → type → callable → body` with typed edge overlays) +and bring it to **analysis level 2** — L1 structural containment + L2 identity-only +`call_graph`, in both projections (`analysis.json` + Neo4j). This is a **schema-major** +change (new envelope, `can://` ids, `span` with UTF-8 byte offsets, `source` per module, +`body` call nodes, `{src,dst,prov}` edges). The `python-sdk` Go wiring lands in lockstep +behind a frozen public API. **Reach for this train: L1 + L2 to parity; L3/L4 deferred.** + +Golden rule: keep everything that *computes* facts (parser, resolver, call-graph +builder); replace only what *serializes* them. + +## Affected repos (from Contract-Impact Triage) + +- `codeanalyzer-go` — analyzer emission (L1 + L2) + Neo4j v2 projection — backend rung +- `python-sdk` — net-new Go model wiring to v2 (`cldk/analysis/` has no `go` package today) — frontend rung +- docs (user-facing schema/levels) — later, `finishing-cldk-work` + +## Design decisions + +- **D1** Pure canonical v2 (drop per-callable `code`/flat `start_line`/`end_line`; recover text by slicing `module.source`) +- **D2** `application.id` = `can://go/`, `` = `--app-name` (default: input dir base name) +- **D3** `span.bytes` = UTF-8 byte offsets into `module.source` (Go strings are UTF-8; avoids the UTF-16/rune mismatch) +- **D4** Single type `kind` ∈ `struct|interface|alias|defined` (collapses v1 `is_interface`); methods nest under receiver type's `callables{}` +- **D5** `base_types[]` = embedded ids (explicit spine) vs `interfaces[]` = computed structural satisfaction (method-set matching) +- **D6** `error_channel[]` from `error`-typed returns; closures → nested `callables{}`; `is_constructor_call` dropped (Go has none) +- **D7** `call` body nodes carry typed `is_goroutine`/`is_deferred`; `callee` is the null→id refinement slot (null at L1, backfilled at L2) +- **D8** `source_file` on a callable when a method is declared apart from its receiver type (cross-file-method span fix) +- **Scope guard:** analyzer is a **pure graph provider** — no slicing/taint in the analyzer (those are SDK queries) + +## Release plan + +- Analyzer = **major** release (breaking output). L1 and L2 are independently shippable behind `--analysis-schema 2`; v1 stays the default until the SDK migrates. +- SDK = **major** release; pins the analyzer **only after** the analyzer v2 release is cut (old SDK Go wiring can't parse v2 until then — and there is no Go wiring yet, so this is net-new). +- Neo4j projection and the JSON schema move in lockstep. + +## Deferred (not this train) + +- **L3** (CFG/CDG/DDG) and **L4** (SDG `param_in`/`param_out`/`summary`) — a later train reusing this train's `can://` ids and `{src,dst,prov}` edge shape verbatim (Group A/C vocabulary, coined once here per the parity clause). +- Go-specific `cfg` edge kinds (`defer_resume`, `select`) — recorded when L3 is designed. +- Points-to oracle (L4 prerequisite); `--materialize-expressions`/`--materialize-basic-blocks`; framework detection. + +## Definition of done (epic-level) + +- Every sub-issue closed and its PR's gate green. +- Analyzer output validates against the SDK v2 Go models; the `L1 ⊆ L2` superset gate holds; parity clause holds (no renamed/repurposed shared vocabulary). +- Superset gate: v2 output carries every v1 fact modulo the sanctioned drops (`code`, `is_*` booleans folded into `kind`, `is_constructor_call`). +- `--analysis-schema 2` output is deterministic across runs (ids and edge order stable). +- SDK public API unchanged (major bump + documented semantic shifts); analyzer↔SDK versions pinned in lockstep. +- Docs / CHANGELOG updated. +- The two switch-over steps: (1) flip the default output from v1 to v2; (2) remove the v1 emitter once `python-sdk` has migrated and nothing reads v1 — both filed as their own units when due. diff --git a/docs/design/issues/parent-v2-migration.md b/docs/design/issues/parent-v2-migration.md index 3496ca0..c0bd54d 100644 --- a/docs/design/issues/parent-v2-migration.md +++ b/docs/design/issues/parent-v2-migration.md @@ -18,7 +18,7 @@ the one-model SDK surface. This is a coordinated schema major across two repos: `codeanalyzer-go` (emission rewrite) and `python-sdk` (Go model remap), released in lockstep. Reach for this train: L1 + L2 to parity; L3/L4 deferred. -**Describe the solution you'd like.** +**Describe the solution you'd like** - [ ] `codeanalyzer-go`: L1 emission — additive tree, `source` per module, `can://` ids, `body{}` with `call` nodes - [ ] `codeanalyzer-go`: L2 emission — `call_graph: [{src,dst,prov,weight}]` at application scope @@ -27,12 +27,12 @@ lockstep. Reach for this train: L1 + L2 to parity; L3/L4 deferred. - [ ] Release v2.0.0 and pin analyzer in SDK **only once both are cut** - [ ] Superset gate green: v2 output contains every v1 fact modulo sanctioned drops -**Describe alternatives you've considered.** +**Describe alternatives you've considered** Does NOT build L3 (CFG/CDG/DDG) or L4 (SDG). Those are a later train reusing this train's `can://` ids and edge shape unchanged. Does NOT add framework detection. -**Additional context.** +**Additional context** - Design transcript: `docs/design/specs/v2-l1-emission.md` (linked, not pasted). - Group A vocabulary (`can://` ids, `span.bytes` UTF-8) is reused verbatim by L3/L4 — coined once here per the parity clause. diff --git a/docs/design/issues/pr-body-v2-l1-l2.md b/docs/design/issues/pr-body-v2-l1-l2.md new file mode 100644 index 0000000..9258919 --- /dev/null +++ b/docs/design/issues/pr-body-v2-l1-l2.md @@ -0,0 +1,52 @@ +Migrate `codeanalyzer-go` to emit the canonical CLDK **v2** `analysis.json` for analysis levels **L1** (additive containment tree) and **L2** (call-graph refinement). v1 output is unchanged and remains the default; v2 is opt-in behind `--analysis-schema 2`. + +## Motivation and Context + +Closes #8 +Closes #9 + +Part of epic codellm-devkit/.github#94. Go was the last analyzer on v1; this moves it onto the shared v2 keystone (`can://` ids, additive tree, `span` with UTF-8 byte offsets, `source` per module, `{src,dst,prov}` edges) so `CLDK(language="go")` can share the one-model SDK surface. Design transcript: `docs/design/specs/v2-l1-emission.md`. + +## How Has This Been Tested? + +Gates run on this commit (analyzer matrix; `testdata/multipackage` fixture): + +``` +# Fixture suite (forced, uncached) +$ go test -count=1 ./... +ok cmd/codeanalyzer ok internal/core ok internal/schema/v2 ok internal/syntactic_analysis/v2emit (all green, exit 0) + +# Determinism: -j 1 vs -j 8, v2 L2 +$ cango -i testdata/multipackage --analysis-schema 2 -a 2 -j 1 -o j1/ ; ... -j 8 -o jN/ +$ diff j1/analysis.json jN/analysis.json # byte-identical -> PASS + +# Monotonicity: v2 L1 ⊆ v2 L2 +L1 node ids: 33; L2 node ids: 33; L1 ids missing from L2: 0 +L1 call_graph: 0 edges; L2 call_graph: 9 edges # L2 adds, never rewrites -> PASS +``` + +Schema-conformance gate (validate against the SDK CPG model) is **deferred**: `python-sdk` has no Go model yet (`cldk/analysis/` has no `go` package) — that is the separate SDK-wiring child of epic #94. Structural self-check passed: valid JSON, `schema_version 2.0.0`, `application.id can://go/multipackage`, edge keys exactly `src/dst/prov/weight`. + +## Breaking Changes + +None for existing users: v1 is still the default output. v2 is additive and opt-in behind `--analysis-schema 2`. The v2 contract itself is a new major, consumed by nothing yet (no SDK Go pin); the default flip and v1-emitter removal are later units, filed when due. + +## Types of changes +- [ ] Bug fix (non-breaking change which fixes an issue) +- [x] New feature (non-breaking change which adds functionality) +- [ ] Breaking change (fix or feature that would cause existing functionality to change) +- [x] Documentation update + +## Checklist +- [x] I have read the [Codellm-Devkit Documentation](https://codellm-devkit.info) +- [x] My code follows the repository's style guidelines +- [x] New and existing tests pass locally +- [x] I have added appropriate error handling +- [x] I have added or updated documentation as needed + +## Additional context + +- L1 computes no new facts beyond span offsets; L2 is pure refinement (signature→`can://` id via `buildSigIndex`). Parser/resolver unchanged (golden rule: replace what serializes, keep what computes). +- `span.bytes` are UTF-8 byte offsets into `module.source` (Go strings are UTF-8). +- `source_file` carried on a callable only when a method is declared apart from its receiver type. +- L3/L4 deferred to a later train; Group A ids/spans already designed for verbatim reuse.