diff --git a/CLAUDE.md b/CLAUDE.md new file mode 100644 index 0000000..7b3f774 --- /dev/null +++ b/CLAUDE.md @@ -0,0 +1,60 @@ +# codeanalyzer-go — agent notes + +The Go static-analysis backend for CLDK (`cango`). Emits the canonical CLDK +`analysis.json`. Parser/resolver/call-graph builder compute the facts; the emission +layer serializes them into the schema shape. + +## Schema decisions + +Decisions made with the user during the v1 → v2 migration design loop +(`designing-cldk-changes`), spine order. Each is a leaf-level addition to the shared +v2 vocabulary (parity clause: add at the leaves, never rename shared names). The +committed spec is the full transcript: `docs/design/specs/v2-l1-emission.md`. + +### Identity & positioning (Group A — reused verbatim by L3/L4) +- **`application.id`** = `can://go/`, where `` is the **`--app-name`** value, + defaulting to the input directory's base name (per the CLI contract; the SDK's Neo4j + backend must use the same anchor). L3/L4 build ids on this unchanged. +- **Callable signature** = existing `signatureOf()` output as the last `can://` path + segment; receiver folded in for methods. One canonicalizer, unchanged from v1. +- **`span.bytes`** = **UTF-8 byte offsets** into `module.source` (Go strings are + natively UTF-8; avoids the UTF-16/rune slicing mismatch TS hit in cants#179). + +### `type` node +- **`kind`** ∈ `struct | interface | alias | defined` — Go's four type-declaration + shapes. Collapses v1's `is_interface` boolean. (`alias` = `type X = Y`; + `defined` = `type X Y` over a non-struct, e.g. `type Celsius float64`.) +- **Methods** (declared outside the type in Go source) are resolved to their receiver + type and placed in that `type.callables{}`; receiver recorded as a field. The tree + reflects containment, not source location. +- **Embedding vs satisfaction:** `base_types[]` = embedded type ids (struct/interface + embedding — the explicit spine); `interfaces[]` = interfaces the type is *computed* + to satisfy via method-set matching (Go's implicit/structural implementation). + +### `callable` node +- **`error_channel[]`** populated from `error`-typed return values (Go's error idiom); + those returns remain in `return_type` too. Maps Go onto the shared field. +- **Closures / function literals** → nested `callables{}` on the enclosing callable + (replaces v1 `InnerCallables`); each closure gets its own `can://` id. +- **`source_file`** (optional; added in the 2026-09-29 design loop): Go allows a method + to be declared in a *different file* than its receiver type. The method node stays + nested under its receiver type (containment-by-receiver, above), but its `span.bytes` + then index the **declaring file's** `module.source`, not the nesting module's. When + the declaring file ≠ the nesting module, the callable carries `source_file` = that + file's path (a `symbol_table` key). **Text recovery:** a node's text = + `symbol_table[source_file].source[span.bytes]` if `source_file` present, else the + nesting module's `source`. Emitted only when it differs (absent = same file). This is + the leaf-level fix for the cross-file-method empty-span collision the real-app eval + surfaced. Cross-language parity (C# partial classes, Rust `impl`, Ruby reopened + classes, TS declaration merging) is **deferred** — promote `source_file` into the + shared keystone vocabulary when a second language needs it. + +### `call` body node +- Typed **`is_goroutine`** and **`is_deferred`** boolean fields (`defer` is net-new vs + v1). `is_constructor_call` **dropped** — Go has no constructors. +- `callee` is the sanctioned `null → id` refinement slot (null at L1, backfilled at L2). + +### Deferred (not this train) +- L3 (CFG/CDG/DDG) and L4 (SDG param edges) — reuse Group A ids/spans unchanged. +- Go-specific `cfg` edge kinds (e.g. `defer_resume`, `select`) are recorded when L3 + is designed, not now. diff --git a/README.md b/README.md index 6cf4285..da1235a 100644 --- a/README.md +++ b/README.md @@ -131,6 +131,8 @@ Aliases: Flags: -a, --analysis-level int Analysis level: 1=symbol table only, 2=+resolver call graph (default 1) + --analysis-schema int Output schema major: 1=legacy v1 shape (default), 2=canonical v2 tree (default 1) + --app-name string Application anchor name for can:// ids and Neo4j :Application (default: input dir name) -c, --cache-dir string Cache directory (default: ~/.cldk/go-cache) --codeql Enable CodeQL framework-based call graph (level 2, stub) --eager Force clean rebuild (ignore cache) @@ -178,6 +180,58 @@ cango -i ./my-go-project --eager cango -i ./my-go-project -a 2 -vv ``` +## Generating `analysis.json` for a Go app + +The end-to-end steps to produce an `analysis.json` file for any Go project. + +**1. Make sure you have `cango`.** Either install a released binary (see +[Installation](#installation)) or build from source: + +```bash +git clone https://github.com/codellm-devkit/codeanalyzer-go +cd codeanalyzer-go +go build -o cango ./cmd/codeanalyzer +``` + +**2. Point it at the project root** (the directory containing the app's `go.mod`). +`cango` resolves imports with the Go type checker, so the target must be a normal Go +module. Write the output to a directory with `-o`; `cango` names the file +`analysis.json` inside it. + +```bash +# Canonical v2 schema, Level 2 (symbol table + call graph): +cango -i /path/to/go/app --analysis-schema 2 -a 2 -o /path/to/output/ +# → writes /path/to/output/analysis.json +``` + +> **Pick the schema explicitly.** `--analysis-schema` defaults to `1` (the legacy v1 +> shape). Pass `--analysis-schema 2` for the current **canonical v2** tree (the shape +> the Python SDK's v2 loader and the sections below describe). Omit `-o` to stream the +> JSON to stdout instead of writing a file. + +**3. (Optional) name the application anchor.** `--app-name` sets the `` in every +`can://go//…` id and defaults to the input directory's base name. Set it when the +directory name isn't the identity you want: + +```bash +cango -i /path/to/go/app --analysis-schema 2 -a 2 --app-name myapp -o ./out/ +``` + +**4. Verify.** A successful run exits `0` and produces a JSON document whose envelope +carries `"schema_version": "2.0.0"`, `"language": "go"`, `"max_level": 2`, and an +`application` tree under `application.symbol_table`: + +```bash +cango -i /path/to/go/app --analysis-schema 2 -a 2 -o ./out/ && echo "exit=$?" +python3 -c "import json; d=json.load(open('out/analysis.json')); print(d['schema_version'], d['max_level'], len(d['application']['symbol_table']), 'modules')" +``` + +**Notes for larger apps.** Analysis parallelism defaults to your CPU count (tune with +`-j`); a warm cache (`--cache-dir`, default `~/.cldk/go-cache`) makes re-runs fast, and +`--eager` forces a clean rebuild. Projects that use **cgo** (`import "C"`) are supported: +the toolchain-synthesized wrapper functions are correctly excluded from the output (they +are build artifacts, not project source), so only your own declarations are emitted. + ## Analysis levels | Level | Flag | What runs | Status | @@ -192,7 +246,110 @@ cango -i ./my-go-project -a 2 -vv ## Output schema -The root object is `GoApplication`: +`cango` emits one of two schema shapes, selected by `--analysis-schema`: + +- **`--analysis-schema 2`** — the **canonical CLDK v2** shape (detailed below); the shape the + v2 Python SDK loader consumes and the [step-by-step section + above](#generating-analysisjson-for-a-go-app) produces. +- **`--analysis-schema 1`** (default) — the **legacy v1** `GoApplication` shape, kept for + backward compatibility. + +### Canonical v2 shape (`--analysis-schema 2`) + +The document is a manifest envelope wrapping one `application` containment tree. Every +node carries a `can://go//…` `id`, a `kind`, and a `span`: + +```json +{ + "schema_version": "2.0.0", + "language": "go", + "max_level": 2, + "analyzer": { "name": "codeanalyzer-go", "version": "0.1.0" }, + "application": { + "id": "can://go/myapp", + "kind": "application", + "symbol_table": { + "pkg/greeter/greeter.go": { + "id": "can://go/myapp/pkg/greeter/greeter.go", + "kind": "module", + "package": "greeter", + "span": { "start": [1, 1], "end": [40, 2], "bytes": [0, 812] }, + "source": "package greeter\n\n...", + "content_hash": "…", + "imports": [ { "name": "fmt", "path": "fmt", "span": {…} } ], + "types": { + "Greeter": { + "id": "can://go/myapp/pkg/greeter/greeter.go/example.com/pkg/greeter.Greeter", + "kind": "struct", + "span": { "start": [5, 1], "end": [7, 2], "bytes": [17, 63] }, + "fields": { + "Prefix": { "id": "…/Prefix", "kind": "field", "type": "string", "span": {…} } + }, + "callables": { + "example.com/pkg/greeter.Greeter.Greet": { + "id": "can://go/myapp/pkg/greeter/greeter.go/example.com/pkg/greeter.Greeter/example.com/pkg/greeter.Greeter.Greet", + "kind": "method", + "signature": "example.com/pkg/greeter.Greeter.Greet", + "span": { "start": [9, 1], "end": [11, 2], "bytes": [65, 140] }, + "parameters": [ { "name": "name", "type": "string", "span": {…} } ], + "return_type": "string", + "error_channel": ["error"], + "metrics": { "cyclomatic": 1 }, + "body": { + "10:2": { + "kind": "call", + "span": { "start": [10, 2], "end": [10, 30], "bytes": [90, 118] }, + "callee": "can://go/myapp/pkg/greeter/greeter.go/…/example.com/pkg/greeter.format" + } + } + } + } + } + }, + "functions": { + "example.com/main.main": { "id": "…", "kind": "function", "signature": "…", "body": {…} } + } + } + }, + "call_graph": [ + { + "src": "can://go/myapp/main.go/…/example.com/main.main", + "dst": "can://go/myapp/pkg/greeter/greeter.go/…/example.com/pkg/greeter.Greeter.Greet", + "prov": ["go/types"], + "weight": 1 + } + ] + } +} +``` + +Key v2 schema properties: +- **Envelope** — `schema_version` (`"2.0.0"`), `language` (`"go"`), `max_level` (1 or 2), + and `analyzer{name,version}` wrap a single `application` node. +- **One containment tree** — `application → module → type → callable → body`, keyed as + `symbol_table` (modules, by **project-relative file path**), `types`, `callables`, `body`. +- **`id`** — every node carries a durable `can://go////` id; + `` is the `--app-name` anchor. Body call nodes use `` keyed by `line:col`. +- **`kind`** — `module` · `type` kinds `struct | interface | alias | defined` (collapses v1's + `is_interface`) · `callable` kinds `function | method | lambda` · body `call`. +- **`base_types: []`** (type, optional) — embedded type ids (struct/interface embedding — the + explicit spine), as opposed to interfaces the type is *computed* to satisfy. +- **`span`** — `{start:[line,col], end:[line,col], bytes:[from,to]}`; `bytes` are **UTF-8 byte + offsets** into the owning `module.source` (source is stored once per module, every node slices it). +- **`source_file`** (callable, optional) — set when a method is declared in a *different file* + than its receiver type; the method stays nested under the type, and its `span.bytes` index + `symbol_table[source_file].source` instead of the nesting module's. +- **`error_channel: []`** — populated from `error`-typed returns (the returns also stay in + `return_type`). **Closures** nest as `callables{}` on the enclosing callable. +- **`body` call nodes** — `callee` is the sanctioned `null → id` slot: `null` at L1 (and for + external/stdlib callees), a `can://` node id once resolved at L2. `is_goroutine` / `is_deferred` + are boolean flags emitted **only when true** (a plain call omits them), for `go f()` / `defer f()`. +- **`call_graph` edges** — `{src, dst, prov, weight}`; `src`/`dst` are `can://` node ids that + exist in the tree (not raw signatures), `prov` is resolver provenance, e.g. `["go/types"]`. + +### Legacy v1 shape (`--analysis-schema 1`, default) + +The v1 root object is `GoApplication`: ```json { @@ -226,7 +383,7 @@ The root object is `GoApplication`: } ``` -Key schema properties: +Key v1 schema properties: - `symbol_table` — keyed by **file path relative to the project root** (never absolute) - `classes` — JSON key for types (spine compatibility with Java/Python schemas); value is `GoType` - `module_name` — JSON key for the Go package name (spine compatibility) diff --git a/cmd/codeanalyzer/main.go b/cmd/codeanalyzer/main.go index 66b3187..5b0b40f 100644 --- a/cmd/codeanalyzer/main.go +++ b/cmd/codeanalyzer/main.go @@ -5,9 +5,10 @@ package main import ( - "encoding/json" "fmt" "os" + "path/filepath" + "runtime" "github.com/spf13/cobra" @@ -30,17 +31,25 @@ func main() { func rootCmd() *cobra.Command { var ( - inputPath string - outputDir string - format string - level int - targetFiles []string - skipTests bool - eager bool - cacheDir string - useCodeQL bool - verbosity int - showVersion bool + inputPath string + outputDir string + format string + emit string + appName string + level int + analysisSchema int + targetFiles []string + skipTests bool + eager bool + cacheDir string + jobs int + useCodeQL bool + verbosity int + showVersion bool + neo4jURI string + neo4jUser string + neo4jPassword string + neo4jDatabase string ) cmd := &cobra.Command{ @@ -57,6 +66,21 @@ via CLDK(language="go").analysis(project_path=...).`, cmd.Println("cango " + version) return nil } + + // --emit selects the output projection. neo4j/schema are validated + // here and rejected with a non-zero exit until the Neo4j child lands + // (never a silent fallback to JSON). schema needs no --input. + switch options.EmitTarget(emit) { + case options.EmitJSON: + // valid — falls through to the JSON path below. + case options.EmitNeo4j: + return fmt.Errorf("--emit neo4j is not yet implemented; use --emit json") + case options.EmitSchema: + return fmt.Errorf("--emit schema is not yet implemented; use --emit json") + default: + return fmt.Errorf("unsupported --emit target %q; supported: json, neo4j, schema", emit) + } + if inputPath == "" { return fmt.Errorf("--input / -i is required") } @@ -74,18 +98,34 @@ via CLDK(language="go").analysis(project_path=...).`, home, _ := os.UserHomeDir() cacheDir = home + "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/.cldk/go-cache" } + if jobs < 1 { + jobs = runtime.NumCPU() + } + // application anchor: explicit flag > input dir base name. + if appName == "" { + appName = filepath.Base(filepath.Clean(inputPath)) + } opts := options.AnalysisOptions{ - InputPath: inputPath, - OutputDir: outputDir, - Format: format, - Level: options.AnalysisLevel(level), - TargetFiles: targetFiles, - SkipTests: skipTests, - Eager: eager, - CacheDir: cacheDir, - UseCodeQL: useCodeQL, - Verbose: verbosity > 0, + InputPath: inputPath, + OutputDir: outputDir, + Format: format, + Emit: options.EmitTarget(emit), + AppName: appName, + Level: options.AnalysisLevel(level), + SchemaVersion: analysisSchema, + AnalyzerVersion: version, + TargetFiles: targetFiles, + SkipTests: skipTests, + Eager: eager, + CacheDir: cacheDir, + Jobs: jobs, + UseCodeQL: useCodeQL, + Verbose: verbosity > 0, + Neo4jURI: firstNonEmpty(neo4jURI, os.Getenv("NEO4J_URI")), + Neo4jUser: firstNonEmpty(neo4jUser, os.Getenv("NEO4J_USERNAME"), "neo4j"), + Neo4jPassword: firstNonEmpty(neo4jPassword, os.Getenv("NEO4J_PASSWORD"), "neo4j"), + Neo4jDatabase: firstNonEmpty(neo4jDatabase, os.Getenv("NEO4J_DATABASE")), } analyzer := core.New(opts) @@ -95,16 +135,17 @@ via CLDK(language="go").analysis(project_path=...).`, } // When no --output dir is given, write JSON to cobra's output - // writer so tests can capture it via cmd.SetOut. + // writer so tests can capture it via cmd.SetOut. Both paths render + // through core so the v1/v2 schema selection stays in one place. if outputDir == "" { - data, err := json.Marshal(app) + data, err := core.RenderJSON(app, opts) if err != nil { return err } _, err = cmd.OutOrStdout().Write(data) return err } - return core.WriteOutput(app, outputDir, format) + return core.WriteOutput(app, opts) }, } @@ -112,16 +153,39 @@ via CLDK(language="go").analysis(project_path=...).`, f.StringVarP(&inputPath, "input", "i", "", "Project root to analyze (required)") f.StringVarP(&outputDir, "output", "o", "", "Output directory for analysis.json (default: stdout)") f.StringVarP(&format, "format", "f", "json", "Output format: json|msgpack") + f.StringVar(&emit, "emit", "json", "Output projection: json|neo4j|schema") + f.StringVar(&appName, "app-name", "", + "Application anchor name for can:// ids and Neo4j :Application (default: input dir name)") f.IntVarP(&level, "analysis-level", "a", 1, "Analysis level: 1=symbol table only, 2=+resolver call graph") + f.IntVar(&analysisSchema, "analysis-schema", 1, + "Output schema major: 1=legacy v1 shape (default), 2=canonical v2 tree") f.StringSliceVarP(&targetFiles, "target-files", "t", nil, "Restrict analysis to specific files (incremental mode)") f.BoolVar(&skipTests, "skip-tests", true, "Skip *_test.go files") f.BoolVar(&eager, "eager", false, "Force clean rebuild (ignore cache)") f.StringVarP(&cacheDir, "cache-dir", "c", "", "Cache directory (default: ~/.cldk/go-cache)") + f.IntVarP(&jobs, "jobs", "j", 0, "Worker parallelism (default: CPU cores)") f.BoolVar(&useCodeQL, "codeql", false, "Enable CodeQL framework-based call graph (level 2, stub)") f.CountVarP(&verbosity, "verbose", "v", "Verbosity (repeat for more detail)") f.BoolVar(&showVersion, "version", false, "Print version and exit") + // Neo4j projection targets (validated now, consumed by the Neo4j child). + f.StringVar(&neo4jURI, "neo4j-uri", "", "Live Bolt push target (env NEO4J_URI); omit to write graph.cypher") + f.StringVar(&neo4jUser, "neo4j-user", "", "Neo4j username (env NEO4J_USERNAME, default neo4j)") + f.StringVar(&neo4jPassword, "neo4j-password", "", "Neo4j password (env NEO4J_PASSWORD, default neo4j)") + f.StringVar(&neo4jDatabase, "neo4j-database", "", "Neo4j database (env NEO4J_DATABASE, optional)") + return cmd } + +// firstNonEmpty returns the first non-empty string in vals, or "" if all are +// empty. It encodes the contract's precedence: explicit flag > env var > default. +func firstNonEmpty(vals ...string) string { + for _, v := range vals { + if v != "" { + return v + } + } + return "" +} diff --git a/cmd/codeanalyzer/main_test.go b/cmd/codeanalyzer/main_test.go index acc08bd..3843261 100644 --- a/cmd/codeanalyzer/main_test.go +++ b/cmd/codeanalyzer/main_test.go @@ -59,6 +59,40 @@ func TestRootCmd_UnknownFormatReturnsError(t *testing.T) { } } +func TestRootCmd_UnknownEmitReturnsError(t *testing.T) { + td := cliTestdataDir() + _, _, err := runCmd("--input", filepath.Join(td, "greeter"), "--emit", "bogus") + if err == nil { + t.Fatal("expected error for unknown --emit value, got nil") + } +} + +func TestRootCmd_EmitNeo4jNotImplemented(t *testing.T) { + td := cliTestdataDir() + _, _, err := runCmd("--input", filepath.Join(td, "greeter"), "--emit", "neo4j") + if err == nil { + t.Fatal("expected non-zero exit for --emit neo4j, got nil") + } + if !strings.Contains(err.Error(), "not yet implemented") { + t.Errorf("error should say 'not yet implemented'; got %q", err.Error()) + } +} + +func TestRootCmd_EmitSchemaNotImplementedWithoutInput(t *testing.T) { + // --emit schema needs no --input; it must still fail (not yet implemented) + // rather than complain about a missing --input. + _, _, err := runCmd("--emit", "schema") + if err == nil { + t.Fatal("expected non-zero exit for --emit schema, got nil") + } + if strings.Contains(err.Error(), "required") { + t.Errorf("--emit schema should not require --input; got %q", err.Error()) + } + if !strings.Contains(err.Error(), "not yet implemented") { + t.Errorf("error should say 'not yet implemented'; got %q", err.Error()) + } +} + // ── --version ──────────────────────────────────────────────────────────────── func TestRootCmd_VersionFlag(t *testing.T) { @@ -162,6 +196,69 @@ func TestRootCmd_Level2ProducesCallGraph(t *testing.T) { } } +// ── --analysis-schema ────────────────────────────────────────────────────────── + +func TestRootCmd_DefaultSchemaIsV1(t *testing.T) { + td := cliTestdataDir() + out, _, err := runCmd("--input", filepath.Join(td, "greeter"), "--cache-dir", t.TempDir()) + if err != nil { + t.Fatalf("command failed: %v", err) + } + var v map[string]interface{} + if jsonErr := json.Unmarshal([]byte(out), &v); jsonErr != nil { + t.Fatalf("stdout is not valid JSON: %v", jsonErr) + } + if _, ok := v["symbol_table"]; !ok { + t.Error("default output should be the v1 shape (top-level symbol_table)") + } + if _, ok := v["schema_version"]; ok { + t.Error("default output should not carry schema_version (that is v2)") + } +} + +func TestRootCmd_AnalysisSchema2EmitsV2(t *testing.T) { + td := cliTestdataDir() + out, _, err := runCmd( + "--input", filepath.Join(td, "greeter"), + "--analysis-schema", "2", + "--cache-dir", t.TempDir(), + ) + if err != nil { + t.Fatalf("command failed: %v", err) + } + var v struct { + SchemaVersion string `json:"schema_version"` + Language string `json:"language"` + Application struct { + ID string `json:"id"` + } `json:"application"` + } + if jsonErr := json.Unmarshal([]byte(out), &v); jsonErr != nil { + t.Fatalf("stdout is not valid JSON: %v", jsonErr) + } + if v.SchemaVersion != "2.0.0" { + t.Errorf("schema_version = %q, want 2.0.0", v.SchemaVersion) + } + if v.Language != "go" { + t.Errorf("language = %q, want go", v.Language) + } + if v.Application.ID != "can://go/greeter" { + t.Errorf("application.id = %q, want can://go/greeter", v.Application.ID) + } +} + +func TestRootCmd_UnknownSchemaReturnsError(t *testing.T) { + td := cliTestdataDir() + _, _, err := runCmd( + "--input", filepath.Join(td, "greeter"), + "--analysis-schema", "9", + "--cache-dir", t.TempDir(), + ) + if err == nil { + t.Fatal("expected error for unknown --analysis-schema value, got nil") + } +} + // ── --skip-tests ───────────────────────────────────────────────────────────── func TestRootCmd_SkipTestsFalseIncludesTestFiles(t *testing.T) { diff --git a/docs/design/issues/child-1-l1-emission.md b/docs/design/issues/child-1-l1-emission.md new file mode 100644 index 0000000..0a33e82 --- /dev/null +++ b/docs/design/issues/child-1-l1-emission.md @@ -0,0 +1,40 @@ + + +**Is your feature request related to a problem? Please describe.** + +The Go analyzer serializes the v1 shape. L1 is the structural keystone of the v2 +migration: it introduces the additive containment tree and every Group A identity +decision the later levels build on. Do L1 first and get the symbol-table gate green +before touching the L2 call graph. The parser/resolver stay; only the emission layer +changes (golden rule: replace what serializes, keep what computes). + +**Describe the solution you'd like** + +- [ ] `application.id` = `can://go/`, `` = `--app-name` (default: input dir base name) +- [ ] Durable `can://` ids on module/type/callable; `signatureOf()` is the last segment +- [ ] `span: {start:[l,c], end:[l,c], bytes:[from,to]}` with UTF-8 byte offsets; drop flat `start_line`/`end_line` +- [ ] `source` emitted once per module; drop per-callable `code` +- [ ] `type.kind` ∈ `struct|interface|alias|defined`; methods grouped under receiver type's `callables{}` +- [ ] `base_types[]` = embedded ids, `interfaces[]` = computed structural satisfaction +- [ ] `body{}` holds `call` nodes with `callee:null`, typed `is_goroutine`/`is_deferred`; closures nest in `callables{}` + +**Describe alternatives you've considered** + +Does NOT emit the `call_graph` (that is the L2 child), nor any L3/L4 `body` statements +or edges. `body{}` at L1 holds only `call` nodes. + +**Additional context** + +- Emitter approach: walk the existing in-memory model, add a v2 emitter behind a flag; keep v1 emitter as the compat shim during transition. +- One genuinely new datum: thread parser byte offsets into `span.bytes`. +- Structural interface satisfaction needs method-set matching — cost lives here, not in the resolver. +- Validate each node against the SDK v2 Go models as they land; superset check vs v1 output. +- Schema decisions recorded in repo `CLAUDE.md`; full transcript in the spec. diff --git a/docs/design/issues/child-2-l2-refinement.md b/docs/design/issues/child-2-l2-refinement.md new file mode 100644 index 0000000..51c0ee5 --- /dev/null +++ b/docs/design/issues/child-2-l2-refinement.md @@ -0,0 +1,39 @@ + + +**Is your feature request related to a problem? Please describe.** + +L1 emits the v2 containment tree with `call` body nodes whose `callee` is `null` — the +sanctioned refinement slot. L2 fills it: resolve each call site and each call-graph edge +endpoint from a v1 `signatureOf()` string to the durable `can://` id of the callable it +names. L2 is pure refinement over the L1 tree — it computes no new facts, it only +translates identity. The parser/resolver stay; only the emission layer gains the +signature→id index (golden rule: replace what serializes, keep what computes). + +**Describe the solution you'd like** + +- [ ] `buildSigIndex` maps every v1 callable signature (functions, methods, nested closures) → its `can://` id, using the SAME id builders `emitCallable` uses (byte-identical ids) +- [ ] `callee` on each `call` body node backfilled from `null` to the resolved `can://` id +- [ ] `call_graph` emitted at application scope as `[{src, dst, prov, weight}]` +- [ ] External/stdlib callees (outside the project) resolve to no endpoint — never a dangling edge +- [ ] L2 gate green: every edge endpoint and every non-null `callee` resolves to a real in-tree node id +- [ ] Determinism: `call_graph` edge order stable across runs (sorted before emit) + +**Describe alternatives you've considered** + +Does NOT compute any new call-graph facts — it reuses the v1 call graph and only +translates endpoints to `can://` ids. Does NOT emit L3/L4 `body` statements or edges. + +**Additional context** + +- v1 signatures are globally unique (a method signature embeds its receiver), so the flat signature→id map has no collisions. +- The index mirrors `emitModule`/`emitType`/`emitCallable` parent-id computation exactly; drift there would silently dangle endpoints. +- Determinism risk: edge order was nondeterministic before sorting — the determinism gate fails without a stable sort. +- Depends on the L1 child: L2 refines the tree L1 produces; file/land L1 first. +- Full transcript in the spec; schema decisions in repo `CLAUDE.md`. diff --git a/docs/design/issues/epic-v2-migration.md b/docs/design/issues/epic-v2-migration.md new file mode 100644 index 0000000..1072e35 --- /dev/null +++ b/docs/design/issues/epic-v2-migration.md @@ -0,0 +1,67 @@ + + +## Spec + +https://github.com/codellm-devkit/codeanalyzer-go/blob/main/docs/design/specs/v2-l1-emission.md + +## Summary + +Migrate `codeanalyzer-go` from the legacy v1 output to the **canonical schema v2** (one +additive tree: `application → module → type → callable → body` with typed edge overlays) +and bring it to **analysis level 2** — L1 structural containment + L2 identity-only +`call_graph`, in both projections (`analysis.json` + Neo4j). This is a **schema-major** +change (new envelope, `can://` ids, `span` with UTF-8 byte offsets, `source` per module, +`body` call nodes, `{src,dst,prov}` edges). The `python-sdk` Go wiring lands in lockstep +behind a frozen public API. **Reach for this train: L1 + L2 to parity; L3/L4 deferred.** + +Golden rule: keep everything that *computes* facts (parser, resolver, call-graph +builder); replace only what *serializes* them. + +## Affected repos (from Contract-Impact Triage) + +- `codeanalyzer-go` — analyzer emission (L1 + L2) + Neo4j v2 projection — backend rung +- `python-sdk` — net-new Go model wiring to v2 (`cldk/analysis/` has no `go` package today) — frontend rung +- docs (user-facing schema/levels) — later, `finishing-cldk-work` + +## Design decisions + +- **D1** Pure canonical v2 (drop per-callable `code`/flat `start_line`/`end_line`; recover text by slicing `module.source`) +- **D2** `application.id` = `can://go/`, `` = `--app-name` (default: input dir base name) +- **D3** `span.bytes` = UTF-8 byte offsets into `module.source` (Go strings are UTF-8; avoids the UTF-16/rune mismatch) +- **D4** Single type `kind` ∈ `struct|interface|alias|defined` (collapses v1 `is_interface`); methods nest under receiver type's `callables{}` +- **D5** `base_types[]` = embedded ids (explicit spine) vs `interfaces[]` = computed structural satisfaction (method-set matching) +- **D6** `error_channel[]` from `error`-typed returns; closures → nested `callables{}`; `is_constructor_call` dropped (Go has none) +- **D7** `call` body nodes carry typed `is_goroutine`/`is_deferred`; `callee` is the null→id refinement slot (null at L1, backfilled at L2) +- **D8** `source_file` on a callable when a method is declared apart from its receiver type (cross-file-method span fix) +- **Scope guard:** analyzer is a **pure graph provider** — no slicing/taint in the analyzer (those are SDK queries) + +## Release plan + +- Analyzer = **major** release (breaking output). L1 and L2 are independently shippable behind `--analysis-schema 2`; v1 stays the default until the SDK migrates. +- SDK = **major** release; pins the analyzer **only after** the analyzer v2 release is cut (old SDK Go wiring can't parse v2 until then — and there is no Go wiring yet, so this is net-new). +- Neo4j projection and the JSON schema move in lockstep. + +## Deferred (not this train) + +- **L3** (CFG/CDG/DDG) and **L4** (SDG `param_in`/`param_out`/`summary`) — a later train reusing this train's `can://` ids and `{src,dst,prov}` edge shape verbatim (Group A/C vocabulary, coined once here per the parity clause). +- Go-specific `cfg` edge kinds (`defer_resume`, `select`) — recorded when L3 is designed. +- Points-to oracle (L4 prerequisite); `--materialize-expressions`/`--materialize-basic-blocks`; framework detection. + +## Definition of done (epic-level) + +- Every sub-issue closed and its PR's gate green. +- Analyzer output validates against the SDK v2 Go models; the `L1 ⊆ L2` superset gate holds; parity clause holds (no renamed/repurposed shared vocabulary). +- Superset gate: v2 output carries every v1 fact modulo the sanctioned drops (`code`, `is_*` booleans folded into `kind`, `is_constructor_call`). +- `--analysis-schema 2` output is deterministic across runs (ids and edge order stable). +- SDK public API unchanged (major bump + documented semantic shifts); analyzer↔SDK versions pinned in lockstep. +- Docs / CHANGELOG updated. +- The two switch-over steps: (1) flip the default output from v1 to v2; (2) remove the v1 emitter once `python-sdk` has migrated and nothing reads v1 — both filed as their own units when due. diff --git a/docs/design/issues/parent-v2-migration.md b/docs/design/issues/parent-v2-migration.md new file mode 100644 index 0000000..c0bd54d --- /dev/null +++ b/docs/design/issues/parent-v2-migration.md @@ -0,0 +1,41 @@ + + +**Is your feature request related to a problem? Please describe.** + +`codeanalyzer-go` emits the v1 schema (`internal/schema/schema.go`): two-key +`{symbol_table, call_graph}`, flat `start_line`/`end_line`, per-callable `code`, +`is_*` booleans, `signature` as id. Every sibling analyzer has moved to the v2 +keystone (`can://` ids, additive tree, `span` with byte offsets, `source` per module, +`{src,dst,prov}` edges). Go is the last on v1, so `CLDK(language="go")` cannot share +the one-model SDK surface. This is a coordinated schema major across two repos: +`codeanalyzer-go` (emission rewrite) and `python-sdk` (Go model remap), released in +lockstep. Reach for this train: L1 + L2 to parity; L3/L4 deferred. + +**Describe the solution you'd like** + +- [ ] `codeanalyzer-go`: L1 emission — additive tree, `source` per module, `can://` ids, `body{}` with `call` nodes +- [ ] `codeanalyzer-go`: L2 emission — `call_graph: [{src,dst,prov,weight}]` at application scope +- [ ] `codeanalyzer-go`: Neo4j projection relabelled to v2 node/edge families +- [ ] `python-sdk`: Go v2 model remap, every public accessor name/return type unchanged +- [ ] Release v2.0.0 and pin analyzer in SDK **only once both are cut** +- [ ] Superset gate green: v2 output contains every v1 fact modulo sanctioned drops + +**Describe alternatives you've considered** + +Does NOT build L3 (CFG/CDG/DDG) or L4 (SDG). Those are a later train reusing this +train's `can://` ids and edge shape unchanged. Does NOT add framework detection. + +**Additional context** + +- Design transcript: `docs/design/specs/v2-l1-emission.md` (linked, not pasted). +- Group A vocabulary (`can://` ids, `span.bytes` UTF-8) is reused verbatim by L3/L4 — coined once here per the parity clause. +- Sanctioned drops: `code` (slice from `source`), `is_*` booleans (folded into `kind`), `is_constructor_call` (Go has no constructors). +- Ordering risk: SDK's v1 Go models will not parse v2 output; do not pin the new major prematurely. +- go.mod-derived `` slug must be deterministic across runs, else ids churn. diff --git a/docs/design/issues/pr-body-v2-l1-l2.md b/docs/design/issues/pr-body-v2-l1-l2.md new file mode 100644 index 0000000..9258919 --- /dev/null +++ b/docs/design/issues/pr-body-v2-l1-l2.md @@ -0,0 +1,52 @@ +Migrate `codeanalyzer-go` to emit the canonical CLDK **v2** `analysis.json` for analysis levels **L1** (additive containment tree) and **L2** (call-graph refinement). v1 output is unchanged and remains the default; v2 is opt-in behind `--analysis-schema 2`. + +## Motivation and Context + +Closes #8 +Closes #9 + +Part of epic codellm-devkit/.github#94. Go was the last analyzer on v1; this moves it onto the shared v2 keystone (`can://` ids, additive tree, `span` with UTF-8 byte offsets, `source` per module, `{src,dst,prov}` edges) so `CLDK(language="go")` can share the one-model SDK surface. Design transcript: `docs/design/specs/v2-l1-emission.md`. + +## How Has This Been Tested? + +Gates run on this commit (analyzer matrix; `testdata/multipackage` fixture): + +``` +# Fixture suite (forced, uncached) +$ go test -count=1 ./... +ok cmd/codeanalyzer ok internal/core ok internal/schema/v2 ok internal/syntactic_analysis/v2emit (all green, exit 0) + +# Determinism: -j 1 vs -j 8, v2 L2 +$ cango -i testdata/multipackage --analysis-schema 2 -a 2 -j 1 -o j1/ ; ... -j 8 -o jN/ +$ diff j1/analysis.json jN/analysis.json # byte-identical -> PASS + +# Monotonicity: v2 L1 ⊆ v2 L2 +L1 node ids: 33; L2 node ids: 33; L1 ids missing from L2: 0 +L1 call_graph: 0 edges; L2 call_graph: 9 edges # L2 adds, never rewrites -> PASS +``` + +Schema-conformance gate (validate against the SDK CPG model) is **deferred**: `python-sdk` has no Go model yet (`cldk/analysis/` has no `go` package) — that is the separate SDK-wiring child of epic #94. Structural self-check passed: valid JSON, `schema_version 2.0.0`, `application.id can://go/multipackage`, edge keys exactly `src/dst/prov/weight`. + +## Breaking Changes + +None for existing users: v1 is still the default output. v2 is additive and opt-in behind `--analysis-schema 2`. The v2 contract itself is a new major, consumed by nothing yet (no SDK Go pin); the default flip and v1-emitter removal are later units, filed when due. + +## Types of changes +- [ ] Bug fix (non-breaking change which fixes an issue) +- [x] New feature (non-breaking change which adds functionality) +- [ ] Breaking change (fix or feature that would cause existing functionality to change) +- [x] Documentation update + +## Checklist +- [x] I have read the [Codellm-Devkit Documentation](https://codellm-devkit.info) +- [x] My code follows the repository's style guidelines +- [x] New and existing tests pass locally +- [x] I have added appropriate error handling +- [x] I have added or updated documentation as needed + +## Additional context + +- L1 computes no new facts beyond span offsets; L2 is pure refinement (signature→`can://` id via `buildSigIndex`). Parser/resolver unchanged (golden rule: replace what serializes, keep what computes). +- `span.bytes` are UTF-8 byte offsets into `module.source` (Go strings are UTF-8). +- `source_file` carried on a callable only when a method is declared apart from its receiver type. +- L3/L4 deferred to a later train; Group A ids/spans already designed for verbatim reuse. diff --git a/docs/design/roadmap.md b/docs/design/roadmap.md new file mode 100644 index 0000000..28e3ada --- /dev/null +++ b/docs/design/roadmap.md @@ -0,0 +1,77 @@ +# CLDK Go analyzer — roadmap + +> Local staging copy. Canonical home is `codellm-devkit/.github → docs/design/roadmap.md`; +> port there when that repo is available. One roadmap, amended in place — never a second file. + +## Initiative: v1 → v2 schema migration (codeanalyzer-go + python-sdk) + +A schema **major**: move the Go analyzer off the v1 keystone +(`{symbol_table, call_graph}`, flat `start_line`/`end_line`, per-callable `code`, `is_*` +booleans, `signature` as id) onto the v2 canonical schema (additive tree, `can://` ids, +`span` with byte offsets, `source` per module, typed edge lists). The SDK moves in +lockstep as a coordinated major release. **Reach for this train: L1 + L2 to parity.** +L3/L4 deferred (see Not now). + +Golden rule carried from `schema-migration.md`: keep everything that *computes* facts +(parser, resolver, call-graph builder); replace only what *serializes* them. + +## Starting now (exactly one) + +- **Group A + L1 emission** — the identity, span, and `body`-with-call-nodes decisions. + This is the keystone rung; everything below reuses its vocabulary. → `designing-cldk-changes`. + +Everything else below has no issue filed yet, by design. + +## Collision groups (the reason this is planning, not design) + +Contract decisions that share v2 vocabulary — **each group is one design session**, even +when its members ship in different release trains. Coining a shared term twice coins it +wrong permanently (parity clause). + +- **A — Identity & positioning.** `application.id` (`can:///`), the `can://` + durable-id grammar (≥ callable), the `@line:col` ordinal ids (< callable), and `span` + with byte offsets. **L3/L4 in a later train reuse this id/span shape verbatim** — so it + must be settled now, in one session, not re-decided per level. This is the group the + gate exists to protect. +- **B — Node shape.** `type.kind` collapsing the `is_interface`/`is_enum`/… booleans; + structured `decorators[]` (`{name,args,span}`) replacing flat annotation strings; + generalized `error_channel[]`; drop per-callable `code`, add `source` once per module. +- **C — Edge shape.** Rename edge keys `source/target/type` → `src/dst/prov`; `call_graph` + as a list at application scope. **Every future edge family (cfg/cdg/ddg/param_*) inherits + this record shape** — coin `{src, dst, prov, weight}` correctly now. +- **D — Neo4j projection.** Node/edge family names must match A/B/C exactly; it is a + relabel of the same tree, always full-depth. + +## Dependency order (DAG) + +``` +A (identity + span) ─┬─► L1 emission (tree, source, ids, body-with-call-nodes) +B (node shape) ─┤ +C (edge shape) ─┴─► L2 emission (call_graph rename @ application scope) + │ +A/B/C ─► D (Neo4j relabel) ─► SDK v2 model remap ─► pin SDK→analyzer (once BOTH cut) + │ + (later train) L3 ─► L4 ──────┘ reuse A & C unchanged +``` + +## Release trains + +- **Train 1 — v2.0.0 (coordinated major, lockstep).** Analyzer L1+L2 emission rewrite + + Neo4j relabel; SDK v2 model remap keeping every public accessor name/return type + identical (two-layer model). **Ordering constraint (first-class on the parent):** pin the + analyzer version in the SDK *only once both are released* — until then the SDK's old + models won't parse v2 output. Superset gate: v2 output must contain every fact v1 did, + modulo the sanctioned drops (`code`, `is_*` booleans folded into `kind`). +- **Train 2+ (deferred).** L3 (CFG/CDG/DDG, AST-only, per-callable parallel — net-new for + Go) then L4 (SDG `param_in`/`param_out`/`summary`, needs a points-to oracle). Both reuse + group A & C vocabulary; no new schema-major coordination beyond what Train 1 establishes. + +## Not now (considered, deliberately excluded) + +- **L3 / L4 dataflow** — deferred by decision this session; net-new heavy construction, not + required for v2 parity. Revisit as Train 2. +- **Points-to oracle** — L4 prerequisite; parked with L4. +- **`--materialize-expressions` / `--materialize-basic-blocks`** — optional body node kinds; + out of scope for the parity migration. +- **New framework detection / CodeQL tier expansion** — orthogonal enrichment axis + (provenance-merged evidence), not a schema level; does not belong on the migration train. diff --git a/docs/design/specs/v2-l1-emission.md b/docs/design/specs/v2-l1-emission.md new file mode 100644 index 0000000..5ff7201 --- /dev/null +++ b/docs/design/specs/v2-l1-emission.md @@ -0,0 +1,144 @@ +# Spec: v2 schema migration — Group A + L1 emission (codeanalyzer-go) + +> **Cross-repo spec.** Canonical home is `codellm-devkit/.github → docs/design/specs/`, +> beside the coordinating parent issue. Staged locally until that repo is available. +> This is the design transcript — the committed provenance of the design loop, linked +> (not pasted) from the parent issue. + +## Contract-Impact Triage + +**Does this change schema v2 output?** Yes — it is the schema-major migration that +introduces the v2 shape for Go: `application.id`, the `can://` id grammar, `@line:col` +ordinal ids (defined here, populated by L3/L4), `span` with UTF-8 byte offsets, the +additive tree, `source` per module, and `body{}` with `call` nodes. + +| Change type | Analyzers | SDKs | Docs | +| --- | --- | --- | --- | +| Schema v2 migration | **codeanalyzer-go** (emission rewrite) | **python-sdk** (Go model remap, lockstep) | CLAUDE.md, this spec | + +Arrived from `planning-cldk-work` with collision **Group A** known (identity + span); +its vocabulary is designed once here and reused verbatim by the deferred L3/L4 train. + +## Design-loop decisions (the transcript) + +Every decision below was made with the user, node by node in spine order. Full copy in +`codeanalyzer-go/CLAUDE.md` § Schema decisions. Anchored on the keystone +(`canonical-schema.md`) and the completed **TypeScript** v2 migration (closest structural +analog: module functions + interfaces + structural types). + +### Root envelope (settled spine, kept as-is) +`{ schema_version:"2.0.0", language:"go", max_level, analyzer:{name,version}, +application:{ id, kind:"application", symbol_table, call_graph, param_in, param_out } }`. + +### Identity & positioning (Group A) +| Decision | Choice | +| --- | --- | +| `application.id` | `can://go/`, `` = **`--app-name`** (default: input dir base name), per CLI contract | +| Callable signature | existing `signatureOf()` as last id segment; receiver folded in for methods | +| `span.bytes` | **UTF-8 byte offsets** into `module.source` | + +### `module` +`kind:"module"`, `package`, `source` (whole file once), `imports[]`, `types{}`, +`functions{}`, `content_hash`. Drops per-callable `code` (sliced from `source`). + +### `type` +| Decision | Choice | +| --- | --- | +| `kind` | `struct \| interface \| alias \| defined` (collapses v1 `is_interface`) | +| Method placement | resolved to receiver type's `callables{}`; receiver as a field | +| `base_types[]` | embedded type ids (struct/interface embedding) | +| `interfaces[]` | computed structural satisfaction (method-set matching) | + +### `callable` +| Decision | Choice | +| --- | --- | +| `error_channel[]` | from `error`-typed returns (kept in `return_type` too) | +| Closures | nested `callables{}` on enclosing callable (replaces `InnerCallables`) | +| `source_file` | **optional**; emitted **only when** the callable's declaring file differs from the module it is nested under (a method whose receiver type lives in another file). Value = declaring file path relative to input root, same shape as a `symbol_table` key | + +#### Cross-file methods — the `source_file` amendment (design loop, 2026-09-29) + +The eval on real Go apps (dae/openbao/cockroach, ~3292 empty spans; cobra clean) +exposed a datamodel collision: v2 nests a method under its **receiver type** +(containment-by-receiver, § `type` above), but Go lets a method be declared in a +*different file* than its type. The keystone and the **Java** reference both fix +`span.bytes` as offsets into the **nesting module's** `source` (Java's SDK slices a +node against the `JCompilationUnit` it is hydrated under). Java can't produce the +conflict — its methods are lexically inside their type's file — so there was no +precedent; this is a genuine Go divergence brought to the user. + +Decision (Go-only for now; parity deferred — see below): +- **Containment unchanged.** The method node stays under its receiver type's + `callables{}`. The tree still reflects containment-by-receiver. +- **New leaf-level optional field `callable.source_file`**, present only when the + declaring file ≠ the nesting module. Parity-safe (an additive leaf field, never a + rename of shared vocabulary). +- **`span.bytes` for a cross-file method index the `source_file` module's `source`**, + not the nesting module's. +- **Text-recovery rule (the SDK contract):** a node's text = + `symbol_table[source_file].source[span.bytes]` when `source_file` is present, else + `.source[span.bytes]`. The declaring file is always another module + already in `symbol_table` (Go analyzes every `.go` file), so its `source` is present + to slice — the rule is fully data-driven, no extra I/O. + +**Deferred — cross-language parity.** C# partial classes, Rust `impl` blocks, Ruby +reopened classes, and TS declaration merging all hit the same +member-declared-apart-from-its-type shape. `source_file` is scoped to Go here; when a +second language needs it, promote the same field + text-recovery rule into the shared +keystone vocabulary in a design pass (the parity clause forbids a divergent second +spelling). Recorded so the promotion is a known follow-up, not a rediscovery. + +### `call` (body node, L1) +| Decision | Choice | +| --- | --- | +| Markers | typed `is_goroutine`, `is_deferred`; **drop** `is_constructor_call` | +| `callee` | sanctioned `null → id` refinement slot (null at L1, backfilled L2) | + +### Edges (Group C, settled here for L2) +`call_graph: [{src, dst, prov, weight}]` at application scope (rename from v1 +`{source, target, type, provenance}`). + +### CLI conformance (cli-contract.md) +The current `cango` CLI is non-conformant. Bring it to the contract as part of this +train (the SDK facade depends on a uniform flag surface across backends): +- **`--emit `** — output target; `json` default. `neo4j`/`schema` + return "not yet implemented" (non-zero) until the Neo4j child lands, never silent + fallback. +- **`--app-name `** — `application.id`/`:Application` anchor; default = input dir + base name. Precedence: explicit flag > env > default. +- **`--neo4j-uri` / `--neo4j-user` / `--neo4j-password` / `--neo4j-database`** — accepted + and validated now; consumed by the Neo4j child. +- **`-j, --jobs `** — worker parallelism (default CPU cores); output byte-identical + across `-j` values. +- Flag validation: unknown/unimplemented value → non-zero exit + clear message. stdout = + data channel only; all diagnostics to stderr. + +## Release plan + +- **Train 1 — v2.0.0 (coordinated major, lockstep).** codeanalyzer-go emission rewrite + (L1 tree/source/ids/body-with-calls, then L2 call_graph) + Neo4j relabel; python-sdk + Go model remap to v2 keeping every public accessor name/return type identical. +- **Ordering constraint (first-class):** pin the analyzer version in the SDK **only once + both are released**. Until then the SDK's old Go models won't parse v2 output. +- **Superset gate:** v2 output must contain every fact v1 did, modulo sanctioned drops + (`code`, `is_*` booleans folded into `kind`, `is_constructor_call`). +- **Deferred (Train 2+):** L3 (CFG/CDG/DDG), L4 (SDG). Reuse Group A ids/spans + Group C + edge shape unchanged; no new schema-major coordination. +- **`source_file` SDK follow-up (Train 1, SDK rung):** python-sdk's Go model must add + the optional `source_file` field and its `slice()`/`get_method_body` must resolve + text against `symbol_table[source_file].source` when present (else the nesting + module). Deferred to the SDK rung — the analyzer ships the field first; the SDK + catches up on its own clock (analyzer release must precede the SDK rung regardless). + +## Decomposition (tracking shape — to be confirmed with user) + +Recommended: **parent (in `codellm-devkit/.github`) + one sub-issue per PR**, because the +work spans two repos on their own clocks and needs a coordination record: +1. codeanalyzer-go: CLI contract conformance (`--emit`, `--app-name`, `--neo4j-*`, `-j`) +2. codeanalyzer-go: L1 emission (tree, source, ids, body-with-call-nodes) +3. codeanalyzer-go: L2 emission (call_graph rename @ app scope) +4. codeanalyzer-go: Neo4j projection (implements the deferred `--emit neo4j`/`schema`) +5. python-sdk: Go v2 model remap (accessors stable) +6. release + pin (once both cut) + +Children filed just-in-time as picked up; this spec records the full plan. diff --git a/docs/design/specs/v2-schema-structure.md b/docs/design/specs/v2-schema-structure.md new file mode 100644 index 0000000..f833b43 --- /dev/null +++ b/docs/design/specs/v2-schema-structure.md @@ -0,0 +1,108 @@ +# v2 `analysis.json` — structure reference (keys only) + +The key skeleton of the canonical CLDK v2 output (`--analysis-schema 2`). This is a +quick map of the node tree and every field; the *decisions* behind the shape live in +[`v2-l1-emission.md`](./v2-l1-emission.md) and `CLAUDE.md` § Schema decisions, and the +authoritative Go definitions are `internal/schema/v2/schema.go`. Field names here are +the JSON tags from that file. + +## Tree + +``` +├─ schema_version # "2.0.0" +├─ language # "go" +├─ max_level # 1 (symbol table) | 2 (+ call graph) +├─ analyzer +│ ├─ name # "codeanalyzer-go" +│ └─ version +└─ application + ├─ id # "can://go/" + ├─ kind # "application" + ├─ symbol_table # { "": module } + │ └─ + │ ├─ id + │ ├─ kind # "module" + │ ├─ span { start, end, bytes } + │ ├─ package + │ ├─ source # whole file text, once per module + │ ├─ content_hash # optional + │ ├─ imports[] { name, path, alias?, span } + │ ├─ types # { "": type } + │ │ └─ + │ │ ├─ id + │ │ ├─ kind # struct | interface | alias | defined + │ │ ├─ span { start, end, bytes } + │ │ ├─ base_types[] # optional: embedded type ids (explicit) + │ │ ├─ interfaces[] # optional: interfaces computed-satisfied + │ │ ├─ fields # optional: { "": field } + │ │ │ └─ { id, kind, type, span } + │ │ └─ callables # { "": callable } + │ │ └─ # ← callable node, defined below + │ └─ functions # { "": callable } + │ └─ # ← callable node, defined below + └─ call_graph[] # L2+; empty [] at L1 + └─ + ├─ src # can:// callable id + ├─ dst # can:// callable id + ├─ prov[] # optional: e.g. ["go/types"] + └─ weight # optional +``` + +## `` node (recursive) + +The same shape whether it sits under a type's `callables` (methods), a module's +`functions` (package-level functions), or another callable's `callables` (closures). + +``` + +├─ id +├─ kind # function | method | lambda +├─ signature +├─ source_file # optional: set only for a method declared in a +│ # different file than its receiver type; its +│ # span.bytes then index symbol_table[source_file].source +├─ span { start, end, bytes } +├─ parameters[] { name, type, span, is_variadic? } +├─ return_type # optional +├─ error_channel[] # optional: error-typed returns +├─ metrics # optional: { cyclomatic: N } +├─ callables # optional: nested closures (same shape) +└─ body # { "": call node } + └─ + ├─ kind # "call" + ├─ span { start, end, bytes } + ├─ callee # null | can:// id (key always present) + ├─ is_goroutine # optional: present only when true + └─ is_deferred # optional: present only when true +``` + +## `span` + +``` +span +├─ start [line, col] +├─ end [line, col] +└─ bytes [from, to] # UTF-8 byte offsets into the owning module.source +``` + +## Presence rules + +- **Always present:** every `id`, `kind`, `span`; the envelope fields; `symbol_table`, + `types`, `callables`, `functions`, `body` maps (may be empty `{}`); `call_graph` (empty + `[]` at L1); a body node's `callee` key (value `null` when unresolved/external). +- **`omitempty` — emitted only when non-empty:** `content_hash`, `imports[].alias`, + `base_types`, `interfaces`, `fields`, `source_file`, `return_type`, `error_channel`, + `metrics`, a callable's nested `callables`, `parameters[].is_variadic`, edge `prov` / + `weight`. +- **Emitted only when `true`:** `is_goroutine`, `is_deferred`. + +## `id` shape + +``` +application can://go/ +module can://go// +type can://go/// +callable can://go//// + (functions omit the segment) +body call keyed by ":" within the callable's body map +``` diff --git a/internal/core/analyzer.go b/internal/core/analyzer.go index 2b9f410..d889a2f 100644 --- a/internal/core/analyzer.go +++ b/internal/core/analyzer.go @@ -5,11 +5,11 @@ // discipline of codeanalyzer-python/codeanalyzer/core.py. // // Phase order: -// 1. Project materialization (go mod download) -// 2. Symbol table construction (syntactic_analysis) -// 3. Resolver-based call graph (semantic_analysis) — if level >= 2 -// 4. Pass pipeline (analysis/registry) -// 5. Optional CodeQL enrichment (semantic_analysis/codeql) — if --codeql +// 1. Project materialization (go mod download) +// 2. Symbol table construction (syntactic_analysis) +// 3. Resolver-based call graph (semantic_analysis) — if level >= 2 +// 4. Pass pipeline (analysis/registry) +// 5. Optional CodeQL enrichment (semantic_analysis/codeql) — if --codeql package core import ( @@ -25,6 +25,7 @@ import ( "github.com/codellm-devkit/codeanalyzer-go/internal/semantic_analysis" "github.com/codellm-devkit/codeanalyzer-go/internal/semantic_analysis/codeql" "github.com/codellm-devkit/codeanalyzer-go/internal/syntactic_analysis" + "github.com/codellm-devkit/codeanalyzer-go/internal/syntactic_analysis/v2emit" "github.com/codellm-devkit/codeanalyzer-go/internal/utils" ) @@ -170,10 +171,14 @@ func (a *Analyzer) saveCache(app *schema.GoApplication) error { return os.WriteFile(cachePath, data, 0o644) } -// WriteOutput writes the GoApplication to outputDir/analysis.json (or stdout -// when outputDir is empty). Only "json" is supported; "msgpack" and other -// values return an explicit error rather than silently falling back to JSON. -func WriteOutput(app *schema.GoApplication, outputDir, format string) error { +// RenderJSON serializes the analysis to JSON in the schema selected by +// opts.SchemaVersion: 1 (default) emits the legacy v1 GoApplication shape; 2 +// emits the canonical v2 tree via the v2emit emitter. This is the single place +// the v1/v2 decision is made, so both the stdout and file output paths agree. +// Only "json" is supported; other formats return an explicit error rather than +// silently falling back. +func RenderJSON(app *schema.GoApplication, opts options.AnalysisOptions) ([]byte, error) { + format := opts.Format if format == "" { format = "json" } @@ -181,21 +186,35 @@ func WriteOutput(app *schema.GoApplication, outputDir, format string) error { case "json": // only supported format case "msgpack": - return fmt.Errorf("msgpack output is not yet implemented; use --format json") + return nil, fmt.Errorf("msgpack output is not yet implemented; use --format json") default: - return fmt.Errorf("unsupported output format %q; supported: json", format) + return nil, fmt.Errorf("unsupported output format %q; supported: json", format) } - data, err := json.Marshal(app) + switch opts.SchemaVersion { + case 0, 1: + return json.Marshal(app) + case 2: + payload := v2emit.Emit(app, opts.AppName, opts.InputPath, int(opts.Level), opts.AnalyzerVersion) + return json.Marshal(payload) + default: + return nil, fmt.Errorf("unsupported --analysis-schema %d; supported: 1, 2", opts.SchemaVersion) + } +} + +// WriteOutput serializes the analysis (in the selected schema) and writes it to +// outputDir/analysis.json, or to stdout when outputDir is empty. +func WriteOutput(app *schema.GoApplication, opts options.AnalysisOptions) error { + data, err := RenderJSON(app, opts) if err != nil { return err } - if outputDir == "" { + if opts.OutputDir == "" { _, err = os.Stdout.Write(data) return err } - if err := utils.EnsureDir(outputDir); err != nil { + if err := utils.EnsureDir(opts.OutputDir); err != nil { return err } - return os.WriteFile(filepath.Join(outputDir, "analysis.json"), data, 0o644) + return os.WriteFile(filepath.Join(opts.OutputDir, "analysis.json"), data, 0o644) } diff --git a/internal/core/analyzer_test.go b/internal/core/analyzer_test.go index 1e95011..45dacf2 100644 --- a/internal/core/analyzer_test.go +++ b/internal/core/analyzer_test.go @@ -137,7 +137,7 @@ func TestCallGraph_CallSitesBackfilled(t *testing.T) { func TestWriteOutput_ValidJSON(t *testing.T) { outDir := t.TempDir() - if err := core.WriteOutput(sharedGreeterL2, outDir, "json"); err != nil { + if err := core.WriteOutput(sharedGreeterL2, options.AnalysisOptions{OutputDir: outDir, Format: "json"}); err != nil { t.Fatalf("WriteOutput: %v", err) } data, err := os.ReadFile(filepath.Join(outDir, "analysis.json")) @@ -155,7 +155,7 @@ func TestWriteOutput_ValidJSON(t *testing.T) { func TestWriteOutput_EmptyFormatDefaultsToJSON(t *testing.T) { outDir := t.TempDir() - if err := core.WriteOutput(sharedGreeterL1, outDir, ""); err != nil { + if err := core.WriteOutput(sharedGreeterL1, options.AnalysisOptions{OutputDir: outDir, Format: ""}); err != nil { t.Fatalf("WriteOutput with empty format: %v", err) } if _, err := os.Stat(filepath.Join(outDir, "analysis.json")); err != nil { @@ -165,14 +165,14 @@ func TestWriteOutput_EmptyFormatDefaultsToJSON(t *testing.T) { func TestWriteOutput_MsgpackNotImplemented(t *testing.T) { outDir := t.TempDir() - if err := core.WriteOutput(sharedGreeterL1, outDir, "msgpack"); err == nil { + if err := core.WriteOutput(sharedGreeterL1, options.AnalysisOptions{OutputDir: outDir, Format: "msgpack"}); err == nil { t.Fatal("expected error for --format msgpack, got nil") } } func TestWriteOutput_UnknownFormatErrors(t *testing.T) { outDir := t.TempDir() - if err := core.WriteOutput(sharedGreeterL1, outDir, "csv"); err == nil { + if err := core.WriteOutput(sharedGreeterL1, options.AnalysisOptions{OutputDir: outDir, Format: "csv"}); err == nil { t.Fatal("expected error for unknown format, got nil") } } diff --git a/internal/core/testsetup_test.go b/internal/core/testsetup_test.go index 85f3b72..2b14d32 100644 --- a/internal/core/testsetup_test.go +++ b/internal/core/testsetup_test.go @@ -66,6 +66,7 @@ func runTestMain(m *testing.M) int { opts := options.AnalysisOptions{ InputPath: f.path, OutputDir: outDir, + Format: "json", Level: f.level, SkipTests: true, CacheDir: cacheDir, @@ -75,7 +76,7 @@ func runTestMain(m *testing.M) int { fmt.Fprintf(os.Stderr, "testsetup %s: Analyze: %v\n", f.name, err) return 1 } - if err := core.WriteOutput(app, outDir, "json"); err != nil { + if err := core.WriteOutput(app, opts); err != nil { fmt.Fprintf(os.Stderr, "testsetup %s: WriteOutput: %v\n", f.name, err) return 1 } diff --git a/internal/options/options.go b/internal/options/options.go index 97e6782..e19d542 100644 --- a/internal/options/options.go +++ b/internal/options/options.go @@ -11,6 +11,18 @@ const ( LevelCallGraph AnalysisLevel = 2 ) +// EmitTarget selects the analyzer's output projection (cli-contract.md § Neo4j). +type EmitTarget string + +const ( + // EmitJSON writes the analysis.json projection (the default). + EmitJSON EmitTarget = "json" + // EmitNeo4j writes/pushes the Neo4j graph projection at full implemented depth. + EmitNeo4j EmitTarget = "neo4j" + // EmitSchema writes the static schema.neo4j.json contract (needs no --input). + EmitSchema EmitTarget = "schema" +) + // AnalysisOptions is the configuration surface passed from the CLI into Analyzer. type AnalysisOptions struct { // InputPath is the project root to analyze. @@ -19,8 +31,19 @@ type AnalysisOptions struct { OutputDir string // Format is the serialization format: "json" or "msgpack". Format string + // Emit selects the output projection: json (default), neo4j, or schema. + Emit EmitTarget + // AppName anchors application identity (can://go/ and the Neo4j + // :Application node). Defaults to the input directory's base name. + AppName string // AnalysisLevel controls symbol-table-only (1) vs + call graph (2). Level AnalysisLevel + // SchemaVersion selects the output schema major: 1 = the legacy v1 shape + // (default, compat shim during the migration), 2 = the canonical v2 tree. + SchemaVersion int + // AnalyzerVersion is the analyzer's own version, recorded in the v2 + // manifest's analyzer{} tag. Set by the CLI from its build-stamped value. + AnalyzerVersion string // TargetFiles restricts analysis to specific files (incremental mode). TargetFiles []string // SkipTests skips test files (files ending in _test.go). @@ -29,8 +52,18 @@ type AnalysisOptions struct { Eager bool // CacheDir is where per-file caches and intermediate data are stored. CacheDir string + // Jobs is the worker parallelism (default: CPU cores). Output must be + // byte-identical across Jobs values. + Jobs int // UseCodeQL enables the framework-based (Tier-2) CodeQL call graph. UseCodeQL bool // Verbose enables verbose logging. Verbose bool + + // Neo4j projection targets (consumed by the Neo4j emitter; validated now). + // Precedence for each: explicit flag > env var > default. + Neo4jURI string // live Bolt push target; empty → write graph.cypher. Env NEO4J_URI + Neo4jUser string // env NEO4J_USERNAME, default "neo4j" + Neo4jPassword string // env NEO4J_PASSWORD, default "neo4j" + Neo4jDatabase string // env NEO4J_DATABASE, optional } diff --git a/internal/schema/v2/ids.go b/internal/schema/v2/ids.go new file mode 100644 index 0000000..d8d192d --- /dev/null +++ b/internal/schema/v2/ids.go @@ -0,0 +1,128 @@ +package v2 + +import "strings" + +// Group A vocabulary: the `can://` durable-id grammar (>= callable) and the +// `@line:col` ordinal-id grammar (< callable), plus span construction. These +// helpers are pure so they can be tested in isolation and reused verbatim by +// the deferred L3/L4 train (which must not re-decide id or span shape). +// +// Durable id grammar (a containment path with an app segment so multiple apps +// in one language don't collide): +// +// can:////// +// can://go/myapp/src/util.go/Hasher/Hash(string)uint64 +// +// Ordinal id grammar (statements and synthetic vertices, addressed WITHIN their +// callable): +// +// @: e.g. …/Hash(string)uint64@15:2 +// @ e.g. …/Hash(string)uint64@entry +// +// The delimiters `/`, `@`, `:` are fixed by the keystone; do not substitute. +const ( + scheme = "can://" + // idSep joins containment segments in a durable id. + idSep = "/" + // ordinalSep separates a callable id from an ordinal (line:col or @tag). + ordinalSep = "@" +) + +// AppID builds the application id: can:///. lang is fixed to the +// package Language constant so callers cannot drift it. +func AppID(app string) string { + return scheme + Language + idSep + app +} + +// ModuleID builds a module id from the application id and the module's file +// path relative to the input root (forward-slashed, no leading slash, no ".."). +func ModuleID(appID, relPath string) string { + return appID + idSep + normalizePath(relPath) +} + +// TypeID builds a type id under its module. sig is the type's signatureOf() +// output — the last (and only) segment the type contributes. +func TypeID(moduleID, sig string) string { + return moduleID + idSep + sig +} + +// CallableID builds a callable id under its parent. For a method the parent is +// its receiver Type id; for a module-level function the parent is the Module +// id. sig is the callable's signatureOf() output — the last path segment. +func CallableID(parentID, sig string) string { + return parentID + idSep + sig +} + +// OrdinalID addresses a real body node (statement / call) within its callable +// by source position: @:. Both line and col are +// required — a bare line is not unique within a callable. +func OrdinalID(callableID string, line, col int) string { + return callableID + ordinalSep + itoa(line) + ":" + itoa(col) +} + +// TagID addresses a synthetic vertex within its callable by tag: +// @ (e.g. "entry", "exit", "formal_in:0"). +func TagID(callableID, tag string) string { + return callableID + ordinalSep + tag +} + +// LocalID is the key a real body node is stored under inside body{}: the bare +// ":" ordinal (the id relative to the enclosing callable). +func LocalID(line, col int) string { + return itoa(line) + ":" + itoa(col) +} + +// NewSpan builds a Span from 1-based line/col positions and the [from, to) UTF-8 +// byte offsets into module.source. Byte offsets are what make source slicing +// O(1); line:col is what addresses and displays. +func NewSpan(startLine, startCol, endLine, endCol, fromByte, toByte int) Span { + return Span{ + Start: [2]int{startLine, startCol}, + End: [2]int{endLine, endCol}, + Bytes: [2]int{fromByte, toByte}, + } +} + +// Slice returns the node's source text: module source sliced by the span's byte +// offsets. It is the O(1) replacement for the v1 per-node `code` field. Out-of- +// range or inverted offsets yield "" rather than panicking, so a malformed span +// degrades gracefully. +func (s Span) Slice(source string) string { + from, to := s.Bytes[0], s.Bytes[1] + if from < 0 || to > len(source) || from > to { + return "" + } + return source[from:to] +} + +// normalizePath makes a relative file path safe for a can:// segment: forward +// slashes, no leading slash. Callers are responsible for passing a path already +// relative to the input root (no ".."). +func normalizePath(p string) string { + p = strings.ReplaceAll(p, "\\", "/") + return strings.TrimPrefix(p, "/") +} + +// itoa is a tiny non-allocating-path integer formatter for id assembly, kept +// local so this file has no fmt dependency for its hot path. +func itoa(n int) string { + if n == 0 { + return "0" + } + neg := n < 0 + if neg { + n = -n + } + var buf [20]byte + i := len(buf) + for n > 0 { + i-- + buf[i] = byte('0' + n%10) + n /= 10 + } + if neg { + i-- + buf[i] = '-' + } + return string(buf[i:]) +} diff --git a/internal/schema/v2/ids_test.go b/internal/schema/v2/ids_test.go new file mode 100644 index 0000000..f7ec330 --- /dev/null +++ b/internal/schema/v2/ids_test.go @@ -0,0 +1,93 @@ +package v2 + +import "testing" + +func TestAppID(t *testing.T) { + if got, want := AppID("myapp"), "can://go/myapp"; got != want { + t.Errorf("AppID = %q, want %q", got, want) + } +} + +func TestDurableIDChain(t *testing.T) { + // Reproduces the keystone worked example, segment by segment. + app := AppID("myapp") + mod := ModuleID(app, "src/util.go") + typ := TypeID(mod, "Hasher") + call := CallableID(typ, "Hash(string)uint64") + + cases := []struct{ got, want string }{ + {mod, "can://go/myapp/src/util.go"}, + {typ, "can://go/myapp/src/util.go/Hasher"}, + {call, "can://go/myapp/src/util.go/Hasher/Hash(string)uint64"}, + } + for _, c := range cases { + if c.got != c.want { + t.Errorf("id = %q, want %q", c.got, c.want) + } + } +} + +func TestModuleIDNormalizesPath(t *testing.T) { + app := AppID("myapp") + // Backslashes become forward slashes; a leading slash is trimmed. + if got, want := ModuleID(app, "\\pkg\\a.go"), "can://go/myapp/pkg/a.go"; got != want { + t.Errorf("ModuleID = %q, want %q", got, want) + } + if got, want := ModuleID(app, "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/pkg/a.go"), "can://go/myapp/pkg/a.go"; got != want { + t.Errorf("ModuleID = %q, want %q", got, want) + } +} + +func TestModuleLevelFunctionParent(t *testing.T) { + // A module-level function hangs off the module id, not a type. + mod := ModuleID(AppID("myapp"), "src/util.go") + fn := CallableID(mod, "New64()") + if want := "can://go/myapp/src/util.go/New64()"; fn != want { + t.Errorf("CallableID = %q, want %q", fn, want) + } +} + +func TestOrdinalAndTagIDs(t *testing.T) { + call := "can://go/myapp/src/util.go/Hasher/Hash(string)uint64" + if got, want := OrdinalID(call, 15, 2), call+"@15:2"; got != want { + t.Errorf("OrdinalID = %q, want %q", got, want) + } + if got, want := TagID(call, "entry"), call+"@entry"; got != want { + t.Errorf("TagID = %q, want %q", got, want) + } + if got, want := LocalID(16, 2), "16:2"; got != want { + t.Errorf("LocalID = %q, want %q", got, want) + } +} + +func TestSpanSlice(t *testing.T) { + src := "package util\nfunc f() {}\n" + // bytes [13,25) covers "func f() {}\n" + s := NewSpan(2, 1, 2, 12, 13, 25) + if got, want := s.Slice(src), "func f() {}\n"; got != want { + t.Errorf("Slice = %q, want %q", got, want) + } +} + +func TestSpanSliceOutOfRangeIsEmpty(t *testing.T) { + src := "abc" + cases := []Span{ + {Bytes: [2]int{0, 99}}, // to past end + {Bytes: [2]int{-1, 2}}, // negative from + {Bytes: [2]int{2, 1}}, // inverted + } + for i, s := range cases { + if got := s.Slice(src); got != "" { + t.Errorf("case %d: Slice = %q, want empty", i, got) + } + } +} + +func TestItoa(t *testing.T) { + cases := map[int]string{0: "0", 7: "7", 42: "42", 1000: "1000", -3: "-3"} + for in, want := range cases { + if got := itoa(in); got != want { + t.Errorf("itoa(%d) = %q, want %q", in, got, want) + } + } +} diff --git a/internal/schema/v2/schema.go b/internal/schema/v2/schema.go new file mode 100644 index 0000000..f4a5dcf --- /dev/null +++ b/internal/schema/v2/schema.go @@ -0,0 +1,205 @@ +// Package v2 defines the canonical CLDK schema (v2) types that codeanalyzer-go +// emits. It is a faithful transcript of the keystone +// (designing-cldk-changes/references/canonical-schema.md) and the Go-specific +// leaf additions recorded in the repo-root CLAUDE.md § Schema decisions. +// +// The model is ONE additive containment tree +// (application → module → type → callable → body) with typed edge overlays +// laid over it (a Code Property Graph). This package declares the L1 tree plus +// the L2 refinement/edge slots; L3/L4 fields (cfg/cdg/ddg/param_*) are not +// declared here — they arrive in their own train and reuse these ids/spans +// unchanged. +// +// Conventions from the keystone, held here: +// - snake_case JSON keys everywhere, so one set of SDK models parses every +// analyzer. +// - a fact is present or absent; there is no null, EXCEPT the one sanctioned +// `callee: null` refinement slot on a call node (null at L1, backfilled at +// L2). Absence is the "no fact" encoding — `omitempty` on optional fields. +// - open-vocabulary fields (prov, tags) are plain strings. +package v2 + +// SchemaVersion is the canonical schema major this package targets. +const SchemaVersion = "2.0.0" + +// Language is the analyzer's language tag in the manifest and `can:///…`. +const Language = "go" + +// ─── Positioning ─────────────────────────────────────────────────────────── + +// Span is the one universal node attribute: where in source the node lives. +// line:col addresses and displays; bytes slices module.source in O(1). Byte +// offsets are UTF-8 offsets into module.source (Go source is natively UTF-8). +type Span struct { + // Start is [line, col] of the first byte (1-based line, 1-based col). + Start [2]int `json:"start"` + // End is [line, col] one past the last byte. + End [2]int `json:"end"` + // Bytes is [from, to) UTF-8 byte offsets into module.source. + Bytes [2]int `json:"bytes"` +} + +// ─── Root envelope ─────────────────────────────────────────────────────────── + +// Analysis is the top-level payload written to analysis.json / stdout. Its +// siblings of `application` carry the manifest; consumers read `max_level` +// rather than sniffing for keys. +type Analysis struct { + SchemaVersion string `json:"schema_version"` + Language string `json:"language"` + MaxLevel int `json:"max_level"` + Analyzer AnalyzerTag `json:"analyzer"` + Application Application `json:"application"` +} + +// AnalyzerTag identifies the producing analyzer and version. +type AnalyzerTag struct { + Name string `json:"name"` + Version string `json:"version"` +} + +// Application is the root node of the containment tree. +type Application struct { + // ID is `can:///` — the app segment disambiguates apps in one + // language. is the --app-name value (default: input dir base name). + ID string `json:"id"` + Kind string `json:"kind"` // always "application" + // SymbolTable is the L1 named map keyed by relative file path. + SymbolTable map[string]Module `json:"symbol_table"` + // CallGraph holds the L2 cross-function edges (callable → callable). Empty + // at L1. + CallGraph []Edge `json:"call_graph"` +} + +// ─── module (per-file compilation unit) ─────────────────────────────────────── + +// Module is one Go source file: the L1 per-file container. +type Module struct { + ID string `json:"id"` + Kind string `json:"kind"` // always "module" + Span Span `json:"span"` + // Package is the Go package the file belongs to. + Package string `json:"package"` + // Source is the WHOLE file's text, stored once. Every node's text slices + // from this via its span.bytes; there is no per-node `code`. + Source string `json:"source"` + // Imports are the file's import declarations. + Imports []Import `json:"imports"` + // Types is the named map of type declarations in the file. + Types map[string]Type `json:"types"` + // Functions is the named map of module-level callables (keyed by signature). + Functions map[string]Callable `json:"functions"` + // ContentHash supports incremental caching; it is not identity. + ContentHash string `json:"content_hash,omitempty"` +} + +// Import is a single import declaration. +type Import struct { + Name string `json:"name"` // imported package name + Path string `json:"path"` // import path, e.g. "hash/fnv" + Alias string `json:"alias,omitempty"` // present only when aliased + Span Span `json:"span"` +} + +// ─── type ───────────────────────────────────────────────────────────────────── + +// Type is a named Go type. Kind is the specific type-declaration shape, not a +// pile of is_* booleans. +type Type struct { + ID string `json:"id"` + Kind string `json:"kind"` // struct | interface | alias | defined + Span Span `json:"span"` + // BaseTypes are embedded type ids (struct/interface embedding — the explicit + // spine). + BaseTypes []string `json:"base_types,omitempty"` + // Interfaces are ids of interfaces this type is computed to satisfy via + // method-set matching (Go's structural implementation). + Interfaces []string `json:"interfaces,omitempty"` + // Callables are the type's methods, resolved to their receiver type + // regardless of source location. Keyed by signature. + Callables map[string]Callable `json:"callables"` + // Fields are the struct fields, keyed by name. + Fields map[string]Field `json:"fields,omitempty"` +} + +// Field is a struct field. +type Field struct { + ID string `json:"id"` + Kind string `json:"kind"` // always "field" + Type string `json:"type"` + Span Span `json:"span"` +} + +// ─── callable ───────────────────────────────────────────────────────────────── + +// Callable is a function, method, or closure. span.bytes → get_method_body = +// module.source[bytes]. +type Callable struct { + ID string `json:"id"` + Kind string `json:"kind"` // function | method | lambda + Span Span `json:"span"` + // Signature is the human-readable last path segment of ID (one signatureOf()). + Signature string `json:"signature"` + // SourceFile names the callable's declaring file (a symbol_table key) when it + // differs from the module this callable is nested under — the Go case where a + // method is declared in a different file than its receiver type. When set, + // Span.Bytes index symbol_table[SourceFile].source, NOT the nesting module's + // source; text = symbol_table[SourceFile].source[Span.Bytes]. Absent (the + // common case) means the declaring file is the nesting module. See CLAUDE.md + // § Schema decisions → callable.source_file and docs/design/specs/v2-l1-emission.md. + SourceFile string `json:"source_file,omitempty"` + // Parameters are the ordered formal parameters. + Parameters []Parameter `json:"parameters"` + // ReturnType is the joined return type, e.g. "(int, error)". + ReturnType string `json:"return_type,omitempty"` + // ErrorChannel is the generalized error surface: populated from error-typed + // returns (those returns remain in ReturnType too). + ErrorChannel []string `json:"error_channel,omitempty"` + // Metrics is an extensible metrics map (e.g. {"cyclomatic": 3}). + Metrics map[string]int `json:"metrics,omitempty"` + // Body holds call nodes at L1 (keyed by local id), completing at L3. + Body map[string]BodyNode `json:"body"` + // Callables are nested closures / function literals, each with its own + // can:// id (replaces v1 InnerCallables). + Callables map[string]Callable `json:"callables,omitempty"` +} + +// Parameter is a single formal parameter. +type Parameter struct { + Name string `json:"name"` + Type string `json:"type"` + Span Span `json:"span"` + IsVariadic bool `json:"is_variadic,omitempty"` +} + +// ─── body nodes ──────────────────────────────────────────────────────────────── + +// BodyNode is a node inside a callable's body, keyed by its local id +// (line:col for real nodes). At L1 the only kind emitted is "call"; L3 adds the +// remaining statement kinds. +type BodyNode struct { + Kind string `json:"kind"` // "call" at L1 + Span Span `json:"span"` + // Callee is the sanctioned null→id refinement slot on a call node: null at + // L1, backfilled to a callable id at L2. A pointer with NO omitempty so it + // always serializes — as JSON null at L1 (the one place null is allowed), + // as an id once backfilled. The keystone requires the field to be present. + Callee *string `json:"callee"` + // IsGoroutine is true when the call is preceded by `go`. + IsGoroutine bool `json:"is_goroutine,omitempty"` + // IsDeferred is true when the call is preceded by `defer` (net-new vs v1). + IsDeferred bool `json:"is_deferred,omitempty"` +} + +// ─── edges ───────────────────────────────────────────────────────────────────── + +// Edge is a typed edge overlay record. The list it lives in IS its type (there +// is no `type` field). src/dst are node ids; no dangling endpoints. +type Edge struct { + Src string `json:"src"` + Dst string `json:"dst"` + // Prov is open-vocabulary provenance, e.g. ["go/types"], ["go/types","codeql"]. + Prov []string `json:"prov,omitempty"` + // Weight accumulates when merging edges from multiple backends. + Weight int `json:"weight,omitempty"` +} diff --git a/internal/syntactic_analysis/symbol_table.go b/internal/syntactic_analysis/symbol_table.go index c7f48ce..93d32ab 100644 --- a/internal/syntactic_analysis/symbol_table.go +++ b/internal/syntactic_analysis/symbol_table.go @@ -492,6 +492,18 @@ func (b *SymbolTableBuilder) buildCallable( decl *ast.FuncDecl, ) *schema.GoCallable { name := decl.Name.Name + + // Skip declarations whose file is outside the input root. cgo synthesizes + // wrapper functions (_Cfunc_*, _Cgo_*, _cgo_runtime_*) into a generated file + // under $GOCACHE; go/packages resolves their token.Pos to that cache path. + // They are toolchain artifacts, not project source — emitting them yields + // callables whose span cannot be sliced against any project module. Drop + // them here, at the point their (cache) file is first known. + if declFile := b.fset.File(decl.Pos()); declFile == nil || + !utils.IsWithin(b.projectDir, declFile.Name()) { + return nil + } + isExported := unicode.IsUpper(rune(name[0])) pos := b.fset.Position(decl.Pos()) end := b.fset.Position(decl.End()) diff --git a/internal/syntactic_analysis/v2emit/determinism_test.go b/internal/syntactic_analysis/v2emit/determinism_test.go new file mode 100644 index 0000000..a7c758a --- /dev/null +++ b/internal/syntactic_analysis/v2emit/determinism_test.go @@ -0,0 +1,76 @@ +package v2emit + +// Determinism gate for the v2 emitter. +// +// release-gates.md requires v2 output to be byte-identical across runs and +// across -j values. The symbol-table tree is already order-stable (modules +// iterated via sortedFileKeys), but the call_graph edge list is assembled by +// ranging a Go map in CallGraphBuilder.Build — random iteration order — and +// emitCallGraph preserved that order. Two runs of the same input therefore +// produced call_graph arrays in different orders. +// +// This gate emits the SAME resolved model twice and asserts the serialized +// call_graph is identical both times, so a regression to unsorted edge output +// fails here rather than only in an external CLI diff. + +import ( + "encoding/json" + "testing" + + "github.com/codellm-devkit/codeanalyzer-go/internal/schema" + v2 "github.com/codellm-devkit/codeanalyzer-go/internal/schema/v2" + "github.com/codellm-devkit/codeanalyzer-go/internal/semantic_analysis" + "github.com/codellm-devkit/codeanalyzer-go/internal/syntactic_analysis" +) + +// TestDeterminism_CallGraphStableAcrossEmits locks that emitting the same model +// twice yields a byte-identical call_graph. emitCallGraph must impose a total +// order on the edges rather than inherit map-iteration order. +func TestDeterminism_CallGraphStableAcrossEmits(t *testing.T) { + dir := multipackageDir(t) + b := syntactic_analysis.NewSymbolTableBuilder(dir) + st, err := b.Build(nil, true) + if err != nil { + t.Fatalf("symbol table build failed: %v", err) + } + cg := semantic_analysis.NewCallGraphBuilder(dir, b.Fset(), b.Pkgs()) + edges := cg.Build(st) + app := &schema.GoApplication{SymbolTable: st, CallGraph: edges} + + // Emit twice from the SAME model. If the edge order is derived from the + // edges slice without a stable sort, repeated marshals still match (same + // slice), so to actually exercise order-independence we marshal a shuffled + // copy of the edge slice too. + first := mustMarshalCallGraph(t, Emit(app, "multipackage", dir, 2, "det-test")) + + // Reverse the input edge order and re-emit: a correct emitter sorts, so the + // serialized call_graph must be unchanged. + rev := make([]schema.GoCallEdge, len(edges)) + for i, e := range edges { + rev[len(edges)-1-i] = e + } + appRev := &schema.GoApplication{SymbolTable: st, CallGraph: rev} + second := mustMarshalCallGraph(t, Emit(appRev, "multipackage", dir, 2, "det-test")) + + if first != second { + t.Errorf("call_graph not order-independent: emitCallGraph must sort edges to a total order\nfirst=%s\nsecond=%s", first, second) + } +} + +func mustMarshalCallGraph(t *testing.T, out *v2.Analysis) string { + t.Helper() + type cgOnly struct { + Application struct { + CallGraph json.RawMessage `json:"call_graph"` + } `json:"application"` + } + data, err := json.Marshal(out) + if err != nil { + t.Fatalf("marshal failed: %v", err) + } + var c cgOnly + if err := json.Unmarshal(data, &c); err != nil { + t.Fatalf("unmarshal failed: %v", err) + } + return string(c.Application.CallGraph) +} diff --git a/internal/syntactic_analysis/v2emit/emit.go b/internal/syntactic_analysis/v2emit/emit.go new file mode 100644 index 0000000..5b61f9b --- /dev/null +++ b/internal/syntactic_analysis/v2emit/emit.go @@ -0,0 +1,406 @@ +package v2emit + +import ( + "os" + "path/filepath" + "sort" + "strings" + + "github.com/codellm-devkit/codeanalyzer-go/internal/schema" + v2 "github.com/codellm-devkit/codeanalyzer-go/internal/schema/v2" +) + +// Emit transforms the v1 GoApplication into the v2 Analysis payload. appName is +// the application anchor (--app-name); projectDir is the absolute input root, +// used to read each module's source for span slicing. maxLevel records how +// deeply the tree was populated (1 = symbol table, 2 = + call_graph); +// analyzerVersion is stamped into the manifest's analyzer{} tag. +// +// This is a pure re-serialization: every fact comes from the v1 model, except +// span byte offsets, which are derived from source via lineIndex. +func Emit(app *schema.GoApplication, appName, projectDir string, maxLevel int, analyzerVersion string) *v2.Analysis { + appID := v2.AppID(appName) + + // The signature→can:// id index translates v1 identity (signature strings) + // into v2 node ids. It is what the L2 refinement uses to backfill body + // callees and to map call_graph edge endpoints onto real callable ids. + idx := buildSigIndex(appID, app.SymbolTable) + + // One lineIndex per source file, read once and shared. A method whose + // receiver type lives in a different file is nested under that type's module + // but must have its span computed against ITS OWN file (CLAUDE.md § Schema + // decisions → callable.source_file); the cache lets emitCallable reach any + // file's index without re-reading it per callable. + srcs := newSourceCache(projectDir) + + out := &v2.Analysis{ + SchemaVersion: v2.SchemaVersion, + Language: v2.Language, + MaxLevel: maxLevel, + Analyzer: v2.AnalyzerTag{Name: "codeanalyzer-go", Version: analyzerVersion}, + Application: v2.Application{ + ID: appID, + Kind: "application", + SymbolTable: make(map[string]v2.Module, len(app.SymbolTable)), + CallGraph: emitCallGraph(app.CallGraph, idx), + }, + } + + // Deterministic module order does not affect the JSON map, but iterating + // sorted keeps any incidental ordering (and logs) stable across runs. + for _, relPath := range sortedFileKeys(app.SymbolTable) { + file := app.SymbolTable[relPath] + out.Application.SymbolTable[relPath] = emitModule(appID, relPath, file, idx, srcs) + } + return out +} + +// emitModule builds a v2 module from a v1 GoFile, reading the file's source so +// spans can slice off it. A file that cannot be read still emits its structure +// (with empty source and zeroed byte spans) rather than dropping the module. +// +// modRel is this module's own file path (a symbol_table key). It is threaded to +// callables so a method declared in another file can be detected and given a +// span against its own source plus a source_file pointer. +func emitModule(appID, modRel string, file schema.GoFile, idx sigIndex, srcs *sourceCache) v2.Module { + li := srcs.index(modRel) + source := li.source + modID := v2.ModuleID(appID, modRel) + + mod := v2.Module{ + ID: modID, + Kind: "module", + Span: wholeFileSpan(li, source), + Package: file.PackageName, + Source: source, + Imports: emitImports(li, file.Imports), + Types: make(map[string]v2.Type, len(file.Types)), + Functions: make(map[string]v2.Callable, len(file.Functions)), + } + if file.ContentHash != nil { + mod.ContentHash = *file.ContentHash + } + + for name, t := range file.Types { + mod.Types[name] = emitType(modID, modRel, t, idx, srcs) + } + for sig, fn := range file.Functions { + mod.Functions[sig] = emitCallable(modID, modRel, li, fn, idx, srcs) + } + return mod +} + +// emitType maps a v1 GoType to a v2 type node, collapsing is_interface into the +// kind and splitting embedding (base_types) from computed satisfaction. Methods +// resolve into the type's callables{}. modRel is the module the type lives in; +// its methods may be declared in other files (handled in emitCallable). +func emitType(modID, modRel string, t schema.GoType, idx sigIndex, srcs *sourceCache) v2.Type { + li := srcs.index(modRel) + typeID := v2.TypeID(modID, t.Signature) + out := v2.Type{ + ID: typeID, + Kind: typeKind(t), + Span: lineSpan(li, t.StartLine, t.EndLine), + BaseTypes: nonEmpty(t.BaseTypes), + Callables: make(map[string]v2.Callable, len(t.Methods)), + } + if len(t.Fields) > 0 { + out.Fields = make(map[string]v2.Field, len(t.Fields)) + for _, f := range t.Fields { + out.Fields[f.Name] = v2.Field{ + ID: typeID + "/" + f.Name, + Kind: "field", + Type: f.Type, + Span: lineSpan(li, f.StartLine, f.EndLine), + } + } + } + for sig, m := range t.Methods { + out.Callables[sig] = emitCallable(typeID, modRel, li, m, idx, srcs) + } + return out +} + +// emitCallable maps a v1 GoCallable to a v2 callable node, including its body +// (call nodes) and nested closures. +// +// modRel and modLI are the nesting module's path and lineIndex. When the +// callable's declaring file (c.Path) differs — a method whose receiver type is +// in another file — its span is computed against its OWN file's lineIndex and +// source_file records that file, per CLAUDE.md § Schema decisions. Nested +// closures share their enclosing callable's file (Go has no cross-file +// closures), so they carry the resolved file down unchanged. +func emitCallable(parentID, modRel string, modLI *lineIndex, c schema.GoCallable, idx sigIndex, srcs *sourceCache) v2.Callable { + callID := v2.CallableID(parentID, c.Signature) + + // Resolve the file this callable's positions are relative to. + li := modLI + sourceFile := "" + if c.Path != "" && c.Path != modRel { + li = srcs.index(c.Path) + sourceFile = c.Path + } + + out := v2.Callable{ + ID: callID, + Kind: callableKind(c), + Span: lineSpan(li, c.StartLine, c.EndLine), + Signature: c.Signature, + SourceFile: sourceFile, + Parameters: emitParams(li, c.Parameters), + ReturnType: c.ReturnType, + Body: emitBody(callID, li, c.CallSites, idx), + } + if ch := errorChannel(c.ReturnTypes); len(ch) > 0 { + out.ErrorChannel = ch + } + if c.CyclomaticComplexity > 0 { + out.Metrics = map[string]int{"cyclomatic": c.CyclomaticComplexity} + } + if len(c.InnerCallables) > 0 { + out.Callables = make(map[string]v2.Callable, len(c.InnerCallables)) + // Closures live in the same file as their enclosing callable, so the + // resolved file (modRel-or-c.Path) becomes their nesting file. + enclRel := modRel + if sourceFile != "" { + enclRel = sourceFile + } + for sig, inner := range c.InnerCallables { + out.Callables[sig] = emitCallable(callID, enclRel, li, inner, idx, srcs) + } + } + return out +} + +// emitBody builds the body map: one `call` node per call site, keyed by its +// line:col local id, carrying the sanctioned callee null->id refinement slot. +// +// callee resolution: +// - L1: the v1 site has no backfilled signature, so callee stays null. +// - L2: the resolver backfilled the site's callee signature. When that +// signature names an in-tree callable, callee becomes its can:// id; when +// it names an external/stdlib callee (no in-tree node), callee stays null — +// the honest-unresolved fallback the L2 gate expects. +func emitBody(callableID string, li *lineIndex, sites []schema.GoCallsite, idx sigIndex) map[string]v2.BodyNode { + body := make(map[string]v2.BodyNode, len(sites)) + for _, cs := range sites { + local := v2.LocalID(cs.StartLine, cs.StartColumn) + body[local] = v2.BodyNode{ + Kind: "call", + Span: colSpan(li, cs.StartLine, cs.StartColumn, cs.EndLine, cs.EndColumn), + Callee: calleeID(cs.CalleeSignature, idx), + IsGoroutine: cs.IsGoroutine, + } + } + return body +} + +// calleeID translates a v1 backfilled callee signature into the v2 can:// id +// refinement slot: nil (JSON null) when unresolved or external, a pointer to the +// in-tree callable id when resolved. Never returns the raw v1 signature — a +// call node's callee is always a node id or null. +func calleeID(sig *string, idx sigIndex) *string { + if sig == nil { + return nil + } + id, ok := idx.idFor(*sig) + if !ok { + return nil + } + return &id +} + +func emitParams(li *lineIndex, params []schema.GoParameter) []v2.Parameter { + out := make([]v2.Parameter, 0, len(params)) + for _, p := range params { + out = append(out, v2.Parameter{ + Name: p.Name, + Type: p.Type, + Span: lineSpan(li, p.StartLine, p.EndLine), + IsVariadic: p.IsVariadic, + }) + } + return out +} + +func emitImports(li *lineIndex, imports []schema.GoImport) []v2.Import { + out := make([]v2.Import, 0, len(imports)) + for _, imp := range imports { + out = append(out, v2.Import{ + Name: packageBase(imp.Module), + Path: imp.Module, + Alias: imp.Alias, + Span: lineSpan(li, imp.StartLine, imp.EndLine), + }) + } + return out +} + +// emitCallGraph maps v1 identity-only edges to the v2 {src,dst,prov,weight} +// shape: the source/target/provenance rename, PLUS the endpoint translation from +// v1 signatures to can:// callable ids. An edge whose src or dst does not resolve +// to an in-tree node is dropped rather than emitted with a dangling endpoint — +// the L2 gate requires every endpoint to be a real callable id. (The v1 resolver +// only emits edges to in-project targets, so this drop is defensive.) +func emitCallGraph(edges []schema.GoCallEdge, idx sigIndex) []v2.Edge { + out := make([]v2.Edge, 0, len(edges)) + for _, e := range edges { + src, srcOK := idx.idFor(e.Source) + dst, dstOK := idx.idFor(e.Target) + if !srcOK || !dstOK { + continue + } + out = append(out, v2.Edge{ + Src: src, + Dst: dst, + Prov: nonEmpty(e.Provenance), + Weight: e.Weight, + }) + } + + // Impose a total order on the edges so the output is byte-identical across + // runs and across -j values (release-gates.md determinism gate). The upstream + // edges slice is assembled by ranging a Go map (random iteration order), so + // without this sort two runs of the same input emit the call_graph in + // different orders. Sort on the full tuple — (src, dst) is the identity but + // prov/weight tie-break so merged edges from multiple backends also order + // stably. + sort.Slice(out, func(i, j int) bool { + a, b := out[i], out[j] + if a.Src != b.Src { + return a.Src < b.Src + } + if a.Dst != b.Dst { + return a.Dst < b.Dst + } + ap, bp := strings.Join(a.Prov, ","), strings.Join(b.Prov, ",") + if ap != bp { + return ap < bp + } + return a.Weight < b.Weight + }) + return out +} + +// ─── kind mapping ──────────────────────────────────────────────────────────── + +// typeKind collapses the v1 is_interface boolean into the v2 kind vocabulary. +// v1 only distinguishes struct vs interface; alias/defined await richer v1 +// facts, so a non-interface type maps to struct for now. +func typeKind(t schema.GoType) string { + if t.IsInterface { + return "interface" + } + return "struct" +} + +// callableKind classifies a callable: a non-empty receiver type marks a method. +func callableKind(c schema.GoCallable) string { + if c.ReceiverType != "" { + return "method" + } + return "function" +} + +// errorChannel extracts error-typed returns into the generalized error surface. +func errorChannel(returnTypes []string) []string { + var ch []string + for _, rt := range returnTypes { + if rt == "error" { + ch = append(ch, rt) + } + } + return ch +} + +// ─── span construction ───────────────────────────────────────────────────────── + +// wholeFileSpan is the module node's span: the entire file. +func wholeFileSpan(li *lineIndex, source string) v2.Span { + lastLine := len(li.lineStart) + if lastLine == 0 { + lastLine = 1 + } + return v2.NewSpan(1, 1, lastLine, li.lineLen(lastLine)+1, 0, len(source)) +} + +// lineSpan builds a span from a line-only v1 position: start at column 1, end at +// the last line's end. Byte offsets bound the whole line range. +func lineSpan(li *lineIndex, startLine, endLine int) v2.Span { + if endLine < startLine { + endLine = startLine + } + endCol := li.lineLen(endLine) + 1 + return v2.NewSpan( + startLine, 1, endLine, endCol, + li.offset(startLine, 1), li.lineEndOffset(endLine), + ) +} + +// colSpan builds a span from a v1 position that carries columns (call sites). +func colSpan(li *lineIndex, startLine, startCol, endLine, endCol int) v2.Span { + return v2.NewSpan( + startLine, startCol, endLine, endCol, + li.offset(startLine, startCol), li.offset(endLine, endCol), + ) +} + +// ─── helpers ───────────────────────────────────────────────────────────────── + +func readSource(projectDir, relPath string) string { + data, err := os.ReadFile(filepath.Join(projectDir, relPath)) + if err != nil { + return "" + } + return string(data) +} + +// sourceCache reads each source file once and memoizes its lineIndex, keyed by +// path relative to projectDir. The emitter needs a file's index in two places — +// when emitting the module itself, and when a method declared in that file is +// nested under a type in a DIFFERENT module — so the cache avoids re-reading a +// file per cross-file method. A file that cannot be read caches an empty index +// (so spans degrade to zero rather than re-attempting the read). +type sourceCache struct { + projectDir string + byPath map[string]*lineIndex +} + +func newSourceCache(projectDir string) *sourceCache { + return &sourceCache{projectDir: projectDir, byPath: make(map[string]*lineIndex)} +} + +// index returns the lineIndex for relPath, reading and indexing the file on +// first use. The returned index carries the file's source (li.source). +func (c *sourceCache) index(relPath string) *lineIndex { + if li, ok := c.byPath[relPath]; ok { + return li + } + li := newLineIndex(readSource(c.projectDir, relPath)) + c.byPath[relPath] = li + return li +} + +func sortedFileKeys(m map[string]schema.GoFile) []string { + keys := make([]string, 0, len(m)) + for k := range m { + keys = append(keys, k) + } + sort.Strings(keys) + return keys +} + +// packageBase returns the last path segment of an import path as the package +// name, e.g. "hash/fnv" -> "fnv". +func packageBase(importPath string) string { + return filepath.Base(importPath) +} + +// nonEmpty returns s unchanged, or nil if empty, so omitempty drops the field +// rather than emitting an empty array (absence = no fact). +func nonEmpty(s []string) []string { + if len(s) == 0 { + return nil + } + return s +} diff --git a/internal/syntactic_analysis/v2emit/emit_test.go b/internal/syntactic_analysis/v2emit/emit_test.go new file mode 100644 index 0000000..6dd39ee --- /dev/null +++ b/internal/syntactic_analysis/v2emit/emit_test.go @@ -0,0 +1,185 @@ +package v2emit + +import ( + "encoding/json" + "path/filepath" + "runtime" + "strings" + "testing" + + "github.com/codellm-devkit/codeanalyzer-go/internal/schema" + v2 "github.com/codellm-devkit/codeanalyzer-go/internal/schema/v2" + "github.com/codellm-devkit/codeanalyzer-go/internal/syntactic_analysis" +) + +// greeterDir returns the absolute path to the greeter testdata fixture. +func greeterDir(t *testing.T) string { + t.Helper() + _, thisFile, _, _ := runtime.Caller(0) + abs, err := filepath.Abs(filepath.Join(filepath.Dir(thisFile), "..", "..", "..", "testdata", "greeter")) + if err != nil { + t.Fatalf("resolving fixture dir: %v", err) + } + return abs +} + +// buildGreeter runs the real symbol-table builder over the greeter fixture. +func buildGreeter(t *testing.T) (*schema.GoApplication, string) { + t.Helper() + dir := greeterDir(t) + b := syntactic_analysis.NewSymbolTableBuilder(dir) + st, err := b.Build(nil, true) + if err != nil { + t.Fatalf("symbol table build failed: %v", err) + } + return &schema.GoApplication{SymbolTable: st, CallGraph: []schema.GoCallEdge{}}, dir +} + +func TestEmit_Envelope(t *testing.T) { + app, dir := buildGreeter(t) + out := Emit(app, "greeter", dir, 1, "test") + + if out.SchemaVersion != "2.0.0" { + t.Errorf("schema_version = %q, want 2.0.0", out.SchemaVersion) + } + if out.Language != "go" { + t.Errorf("language = %q, want go", out.Language) + } + if out.MaxLevel != 1 { + t.Errorf("max_level = %d, want 1", out.MaxLevel) + } + if out.Application.ID != "can://go/greeter" { + t.Errorf("application.id = %q, want can://go/greeter", out.Application.ID) + } + if out.Application.Kind != "application" { + t.Errorf("application.kind = %q, want application", out.Application.Kind) + } + if len(out.Application.SymbolTable) == 0 { + t.Fatal("symbol_table is empty") + } +} + +func TestEmit_ModuleSourceAndIDs(t *testing.T) { + app, dir := buildGreeter(t) + out := Emit(app, "greeter", dir, 1, "test") + + mod, ok := out.Application.SymbolTable["main.go"] + if !ok { + t.Fatalf("main.go not in symbol_table; keys=%v", keys(out.Application.SymbolTable)) + } + if mod.Kind != "module" { + t.Errorf("module.kind = %q, want module", mod.Kind) + } + if mod.ID != "can://go/greeter/main.go" { + t.Errorf("module.id = %q, want can://go/greeter/main.go", mod.ID) + } + if !strings.HasPrefix(mod.Source, "package main") { + t.Errorf("module.source should hold the file text; got prefix %q", head(mod.Source, 20)) + } + // The module span slices the whole file. + if got := mod.Span.Slice(mod.Source); got != mod.Source { + t.Errorf("module span does not cover whole source (%d vs %d bytes)", len(got), len(mod.Source)) + } +} + +func TestEmit_CallableBodyAndSpanSlice(t *testing.T) { + app, dir := buildGreeter(t) + out := Emit(app, "greeter", dir, 1, "test") + + mod := out.Application.SymbolTable["main.go"] + var main v2.Callable + found := false + for _, fn := range mod.Functions { + if fn.Signature != "" && strings.HasSuffix(fn.Signature, ".main") { + main, found = fn, true + } + } + if !found { + t.Fatalf("main function not found; functions=%v", keys(mod.Functions)) + } + if main.Kind != "function" { + t.Errorf("main.kind = %q, want function", main.Kind) + } + // Its span should slice to source beginning with "func main". + if body := main.Span.Slice(mod.Source); !strings.HasPrefix(body, "func main") { + t.Errorf("main span slice should start with 'func main'; got %q", head(body, 20)) + } + // main() has call sites → body call nodes, each with callee null at L1. + if len(main.Body) == 0 { + t.Fatal("main should have call nodes in its body") + } + for local, node := range main.Body { + if node.Kind != "call" { + t.Errorf("body[%s].kind = %q, want call", local, node.Kind) + } + if node.Callee != nil { + t.Errorf("body[%s].callee should be null at L1; got %q", local, *node.Callee) + } + if !strings.Contains(local, ":") { + t.Errorf("body key %q should be line:col", local) + } + } +} + +func TestEmit_CalleeSerializesAsNull(t *testing.T) { + app, dir := buildGreeter(t) + out := Emit(app, "greeter", dir, 1, "test") + data, err := json.Marshal(out) + if err != nil { + t.Fatalf("marshal failed: %v", err) + } + if !strings.Contains(string(data), `"callee":null`) { + t.Error("expected at least one call node to serialize callee as null at L1") + } + // The v1-only keys must be gone. + for _, banned := range []string{`"symbol_table"`, `"call_graph"`} { + if !strings.Contains(string(data), banned) { + t.Errorf("expected v2 payload to contain %s", banned) + } + } + if strings.Contains(string(data), `"is_constructor_call"`) { + t.Error("v2 payload should not contain dropped is_constructor_call") + } +} + +func TestEmit_TypeKindAndMethods(t *testing.T) { + app, dir := buildGreeter(t) + out := Emit(app, "greeter", dir, 1, "test") + + // Find the Greeter type wherever it lives; assert struct kind + methods. + var found bool + for _, mod := range out.Application.SymbolTable { + for name, ty := range mod.Types { + if ty.Kind != "struct" && ty.Kind != "interface" { + t.Errorf("type %s kind = %q, want struct|interface", name, ty.Kind) + } + if strings.Contains(ty.ID, "://") == false { + t.Errorf("type %s id should be a can:// id; got %q", name, ty.ID) + } + for sig, m := range ty.Callables { + if m.Kind != "method" { + t.Errorf("callable %s under type %s kind = %q, want method", sig, name, m.Kind) + } + found = true + } + } + } + if !found { + t.Skip("no type methods in fixture; envelope/module coverage still asserted elsewhere") + } +} + +func head(s string, n int) string { + if len(s) < n { + return s + } + return s[:n] +} + +func keys[V any](m map[string]V) []string { + out := make([]string, 0, len(m)) + for k := range m { + out = append(out, k) + } + return out +} diff --git a/internal/syntactic_analysis/v2emit/gate_l2_test.go b/internal/syntactic_analysis/v2emit/gate_l2_test.go new file mode 100644 index 0000000..8a50883 --- /dev/null +++ b/internal/syntactic_analysis/v2emit/gate_l2_test.go @@ -0,0 +1,243 @@ +package v2emit + +// L2 conformance gate for the v2 emitter. +// +// L2 is a pure refinement over L1: it backfills the callee null->id slot and +// emits the call_graph edge list, both endpoints callable ids. This gate builds +// the multipackage fixture through the REAL resolver (the same wiring the +// analyzer uses at -a 2), emits v2, and asserts every item on the L2 checklist +// in designing-cldk-changes/references/testing-and-validation.md: +// +// - no dangling endpoints — every edge src/dst is a real callable id +// - every edge carries a non-empty prov naming the resolver +// - callee is a backfilled can:// id on resolved sites, null on unresolved +// - a NAMED expected edge is present (exact src,dst pair, not "graph non-empty") +// - at least one cross-package edge is present +// - the L1 ⊆ L2 superset holds: the tree is unchanged except the callee fill. + +import ( + "testing" + + "github.com/codellm-devkit/codeanalyzer-go/internal/schema" + v2 "github.com/codellm-devkit/codeanalyzer-go/internal/schema/v2" + "github.com/codellm-devkit/codeanalyzer-go/internal/semantic_analysis" + "github.com/codellm-devkit/codeanalyzer-go/internal/syntactic_analysis" +) + +// buildMultipackageL2 builds the fixture AND runs the resolver, so the emitted +// v2 payload carries L2 facts (backfilled callees, call_graph edges). +func buildMultipackageL2(t *testing.T) *v2.Analysis { + t.Helper() + dir := multipackageDir(t) + b := syntactic_analysis.NewSymbolTableBuilder(dir) + st, err := b.Build(nil, true) + if err != nil { + t.Fatalf("symbol table build failed: %v", err) + } + cg := semantic_analysis.NewCallGraphBuilder(dir, b.Fset(), b.Pkgs()) + edges := cg.Build(st) // mutates st in place: backfills callee_signature + app := &schema.GoApplication{SymbolTable: st, CallGraph: edges} + return Emit(app, "multipackage", dir, 2, "gate-test") +} + +// allCallableIDs collects every callable/type id in the emitted tree. +func allCallableIDs(out *v2.Analysis) map[string]bool { + ids := map[string]bool{} + var walk func(c v2.Callable) + walk = func(c v2.Callable) { + ids[c.ID] = true + for _, inner := range c.Callables { + walk(inner) + } + } + for _, mod := range out.Application.SymbolTable { + for _, fn := range mod.Functions { + walk(fn) + } + for _, ty := range mod.Types { + ids[ty.ID] = true + for _, m := range ty.Callables { + walk(m) + } + } + } + return ids +} + +// TestGateL2_MaxLevelAndEdgesPresent locks that L2 emits a non-empty call graph. +func TestGateL2_MaxLevelAndEdgesPresent(t *testing.T) { + out := buildMultipackageL2(t) + if out.MaxLevel != 2 { + t.Errorf("max_level = %d, want 2", out.MaxLevel) + } + if len(out.Application.CallGraph) == 0 { + t.Fatal("L2 should emit call_graph edges; got none") + } +} + +// TestGateL2_NoDanglingEndpoints locks that every edge endpoint resolves to a +// real callable id, and every edge names its resolver in prov. +func TestGateL2_NoDanglingEndpoints(t *testing.T) { + out := buildMultipackageL2(t) + ids := allCallableIDs(out) + + for i, e := range out.Application.CallGraph { + if !ids[e.Src] { + t.Errorf("edge[%d] src is a dangling endpoint: %q", i, e.Src) + } + if !ids[e.Dst] { + t.Errorf("edge[%d] dst is a dangling endpoint: %q", i, e.Dst) + } + if len(e.Prov) == 0 { + t.Errorf("edge[%d] (%s -> %s) has empty prov", i, e.Src, e.Dst) + } + } +} + +// TestGateL2_NamedEdgePresent asserts an EXACT expected edge, so the gate proves +// correctness, not merely a non-empty graph: Worker.Run calls Worker.execute +// (the `go w.execute(...)` goroutine), both as can:// ids. +func TestGateL2_NamedEdgePresent(t *testing.T) { + out := buildMultipackageL2(t) + + const ( + src = "can://go/multipackage/worker/worker.go/example.com/multipackage/worker.Worker/example.com/multipackage/worker.Worker.Run" + dst = "can://go/multipackage/worker/worker.go/example.com/multipackage/worker.Worker/example.com/multipackage/worker.Worker.execute" + ) + found := false + for _, e := range out.Application.CallGraph { + if e.Src == src && e.Dst == dst { + found = true + break + } + } + if !found { + t.Errorf("expected edge %s -> %s not found in call_graph", src, dst) + } +} + +// TestGateL2_CrossPackageEdge locks a specific cross-package edge: main() (in +// package main, main.go) calls server.New (in package server, server/server.go). +// Asserting the exact pair proves a resolved cross-module edge, not just that +// the graph is non-empty. +func TestGateL2_CrossPackageEdge(t *testing.T) { + out := buildMultipackageL2(t) + + const ( + src = "can://go/multipackage/main.go/example.com/multipackage.main" + dst = "can://go/multipackage/server/server.go/example.com/multipackage/server.New" + ) + found := false + for _, e := range out.Application.CallGraph { + if e.Src == src && e.Dst == dst { + found = true + break + } + } + if !found { + t.Errorf("expected cross-package edge %s -> %s not found", src, dst) + } +} + +// TestGateL2_CalleeBackfilled locks the refinement: the goroutine call site in +// Worker.Run has its callee backfilled to Worker.execute's can:// id (not null, +// not the raw v1 signature). +func TestGateL2_CalleeBackfilled(t *testing.T) { + out := buildMultipackageL2(t) + ids := allCallableIDs(out) + + run, ok := out.Application.SymbolTable["worker/worker.go"]. + Types["Worker"].Callables["example.com/multipackage/worker.Worker.Run"] + if !ok { + t.Fatal("Worker.Run missing") + } + if len(run.Body) == 0 { + t.Fatal("Worker.Run should have body call nodes") + } + sawBackfilled := false + for local, node := range run.Body { + if node.Callee == nil { + continue + } + if !ids[*node.Callee] { + t.Errorf("body[%s].callee = %q is not an in-tree callable id", local, *node.Callee) + } + sawBackfilled = true + } + if !sawBackfilled { + t.Error("Worker.Run should have at least one backfilled callee at L2") + } +} + +// TestGateL2_SupersetOfL1 locks L1 ⊆ L2: emitting the SAME model at level 1 vs +// level 2 changes nothing but the callee slot. Every module/type/callable id and +// span is identical; only body callees may go null -> id. +func TestGateL2_SupersetOfL1(t *testing.T) { + dir := multipackageDir(t) + b := syntactic_analysis.NewSymbolTableBuilder(dir) + st, err := b.Build(nil, true) + if err != nil { + t.Fatalf("symbol table build failed: %v", err) + } + // L1 snapshot BEFORE the resolver mutates call sites. + l1 := Emit(&schema.GoApplication{SymbolTable: st, CallGraph: nil}, "multipackage", dir, 1, "gate-test") + + // Now run the resolver (mutates st) and emit L2. + cg := semantic_analysis.NewCallGraphBuilder(dir, b.Fset(), b.Pkgs()) + edges := cg.Build(st) + l2 := Emit(&schema.GoApplication{SymbolTable: st, CallGraph: edges}, "multipackage", dir, 2, "gate-test") + + if len(l1.Application.SymbolTable) != len(l2.Application.SymbolTable) { + t.Fatalf("module count changed L1->L2: %d vs %d", + len(l1.Application.SymbolTable), len(l2.Application.SymbolTable)) + } + for rel, m1 := range l1.Application.SymbolTable { + m2 := l2.Application.SymbolTable[rel] + if m1.ID != m2.ID || m1.Span != m2.Span || m1.Source != m2.Source { + t.Errorf("%s: module identity/span/source changed L1->L2", rel) + } + assertCallablesEqualModuloCallee(t, rel, m1.Functions, m2.Functions) + for name, t1 := range m1.Types { + t2 := m2.Types[name] + if t1.ID != t2.ID || t1.Span != t2.Span || t1.Kind != t2.Kind { + t.Errorf("%s type %s: identity/span/kind changed L1->L2", rel, name) + } + assertCallablesEqualModuloCallee(t, rel+" "+name, t1.Callables, t2.Callables) + } + } +} + +// assertCallablesEqualModuloCallee checks two callable maps are identical except +// that a body node's callee may have gone from null (L1) to a non-null id (L2). +func assertCallablesEqualModuloCallee(t *testing.T, where string, a, b map[string]v2.Callable) { + t.Helper() + if len(a) != len(b) { + t.Errorf("%s: callable count changed L1->L2: %d vs %d", where, len(a), len(b)) + return + } + for sig, c1 := range a { + c2, ok := b[sig] + if !ok { + t.Errorf("%s: callable %q dropped L1->L2", where, sig) + continue + } + if c1.ID != c2.ID || c1.Span != c2.Span || c1.Kind != c2.Kind { + t.Errorf("%s callable %s: identity/span/kind changed L1->L2", where, sig) + } + if len(c1.Body) != len(c2.Body) { + t.Errorf("%s callable %s: body node count changed L1->L2", where, sig) + continue + } + for local, n1 := range c1.Body { + n2 := c2.Body[local] + if n1.Kind != n2.Kind || n1.Span != n2.Span || n1.IsGoroutine != n2.IsGoroutine { + t.Errorf("%s callable %s body[%s]: non-callee field changed L1->L2", where, sig, local) + } + // The one sanctioned mutation: callee null -> id only. + if n1.Callee != nil { + t.Errorf("%s callable %s body[%s]: L1 callee should be null", where, sig, local) + } + } + assertCallablesEqualModuloCallee(t, where+" "+sig, c1.Callables, c2.Callables) + } +} diff --git a/internal/syntactic_analysis/v2emit/gate_test.go b/internal/syntactic_analysis/v2emit/gate_test.go new file mode 100644 index 0000000..b021a9d --- /dev/null +++ b/internal/syntactic_analysis/v2emit/gate_test.go @@ -0,0 +1,370 @@ +package v2emit + +// L1 conformance gate for the v2 emitter. +// +// This is the fixture gate the codeanalyzer-backend ladder requires before L1 +// advances: it locks the v2 L1 OUTPUT SHAPE against a known multi-file fixture, +// asserting concrete values (not just "a key exists") for every item on the L1 +// coverage checklist in +// designing-cldk-changes/references/testing-and-validation.md: +// +// - a multi-file compilation unit (four modules keyed by relative path) +// - exported AND unexported symbols (Worker.Run vs Worker.execute) +// - an interface type collapsed into kind=interface (Processor) +// - error_channel populated from an error-typed return (Processor.Process) +// - a variadic parameter (Combine(results ...Result)) +// - a language-specific call flag (is_goroutine on `go w.execute(...)`) +// - the sanctioned callee:null refinement slot present at L1 +// - can:// durable ids at module/type/callable depth +// - the SUPERSET gate: every v1 fact survives into v2 modulo sanctioned drops +// (is_constructor_call is dropped; per-node `code` folds into module.source). +// +// If any of these regress, this gate fails and the level does not advance. + +import ( + "encoding/json" + "path/filepath" + "runtime" + "strings" + "testing" + + "github.com/codellm-devkit/codeanalyzer-go/internal/schema" + v2 "github.com/codellm-devkit/codeanalyzer-go/internal/schema/v2" + "github.com/codellm-devkit/codeanalyzer-go/internal/syntactic_analysis" +) + +// multipackageDir returns the absolute path to the multipackage testdata fixture. +func multipackageDir(t *testing.T) string { + t.Helper() + _, thisFile, _, _ := runtime.Caller(0) + abs, err := filepath.Abs(filepath.Join(filepath.Dir(thisFile), "..", "..", "..", "testdata", "multipackage")) + if err != nil { + t.Fatalf("resolving fixture dir: %v", err) + } + return abs +} + +// buildMultipackageV2 runs the real symbol-table builder over the multipackage +// fixture and emits the v2 L1 payload. skipTests=true keeps *_test.go out so the +// module set is deterministic. +func buildMultipackageV2(t *testing.T) *v2.Analysis { + t.Helper() + dir := multipackageDir(t) + b := syntactic_analysis.NewSymbolTableBuilder(dir) + st, err := b.Build(nil, true) + if err != nil { + t.Fatalf("symbol table build failed: %v", err) + } + app := &schema.GoApplication{SymbolTable: st, CallGraph: []schema.GoCallEdge{}} + return Emit(app, "multipackage", dir, 1, "gate-test") +} + +// mustModule fetches a module by relative path or fails. +func mustModule(t *testing.T, out *v2.Analysis, rel string) v2.Module { + t.Helper() + mod, ok := out.Application.SymbolTable[rel] + if !ok { + t.Fatalf("module %q missing; modules=%v", rel, keys(out.Application.SymbolTable)) + } + return mod +} + +// TestGateL1_ModuleSet locks the exact multi-file compilation unit: the four +// non-test source files, each a module keyed by its relative path. +func TestGateL1_ModuleSet(t *testing.T) { + out := buildMultipackageV2(t) + + want := []string{ + "main.go", + "server/middleware.go", + "server/server.go", + "worker/worker.go", + } + for _, rel := range want { + mod := mustModule(t, out, rel) + if mod.Kind != "module" { + t.Errorf("%s: kind = %q, want module", rel, mod.Kind) + } + wantID := "can://go/multipackage/" + rel + if mod.ID != wantID { + t.Errorf("%s: id = %q, want %q", rel, mod.ID, wantID) + } + if mod.Source == "" { + t.Errorf("%s: source is empty; every module carries its whole file text", rel) + } + // The module span must slice back to the whole source (source-once invariant). + if got := mod.Span.Slice(mod.Source); got != mod.Source { + t.Errorf("%s: module span does not cover whole source (%d vs %d bytes)", rel, len(got), len(mod.Source)) + } + } +} + +// TestGateL1_InterfaceKind locks the Processor interface: is_interface collapses +// into kind=interface, and its method resolves as a callable under the type. +func TestGateL1_InterfaceKind(t *testing.T) { + out := buildMultipackageV2(t) + worker := mustModule(t, out, "worker/worker.go") + + proc, ok := worker.Types["Processor"] + if !ok { + t.Fatalf("Processor type missing; types=%v", keys(worker.Types)) + } + if proc.Kind != "interface" { + t.Errorf("Processor.kind = %q, want interface", proc.Kind) + } + const procMethod = "example.com/multipackage/worker.Processor.Process" + m, ok := proc.Callables[procMethod] + if !ok { + t.Fatalf("Processor.Process missing; callables=%v", keys(proc.Callables)) + } + if m.Kind != "method" { + t.Errorf("Process.kind = %q, want method", m.Kind) + } + // error_channel is populated from the (Result, error) return. + if !contains(m.ErrorChannel, "error") { + t.Errorf("Process.error_channel = %v, want to contain \"error\"", m.ErrorChannel) + } +} + +// TestGateL1_ExportedAndUnexported locks that both an exported (Run) and an +// unexported (execute) method survive, each a method callable under Worker. +func TestGateL1_ExportedAndUnexported(t *testing.T) { + out := buildMultipackageV2(t) + worker := mustModule(t, out, "worker/worker.go") + + wtype, ok := worker.Types["Worker"] + if !ok { + t.Fatalf("Worker type missing; types=%v", keys(worker.Types)) + } + const ( + runID = "example.com/multipackage/worker.Worker.Run" + execID = "example.com/multipackage/worker.Worker.execute" + ) + for _, sig := range []string{runID, execID} { + m, ok := wtype.Callables[sig] + if !ok { + t.Fatalf("method %q missing; callables=%v", sig, keys(wtype.Callables)) + } + if m.Kind != "method" { + t.Errorf("%s.kind = %q, want method", sig, m.Kind) + } + } +} + +// TestGateL1_VariadicParameter locks the variadic-parameter fact: +// Combine(results ...Result) → is_variadic=true on the sole parameter. +func TestGateL1_VariadicParameter(t *testing.T) { + out := buildMultipackageV2(t) + worker := mustModule(t, out, "worker/worker.go") + + const combineID = "example.com/multipackage/worker.Combine" + comb, ok := worker.Functions[combineID] + if !ok { + t.Fatalf("Combine function missing; functions=%v", keys(worker.Functions)) + } + if len(comb.Parameters) != 1 { + t.Fatalf("Combine should have 1 parameter; got %d", len(comb.Parameters)) + } + if !comb.Parameters[0].IsVariadic { + t.Errorf("Combine parameter %q should be variadic", comb.Parameters[0].Name) + } +} + +// TestGateL1_GoroutineCallNode locks the language-specific call flag: Worker.Run +// contains `go w.execute(...)`, which must surface as a body call node with +// is_goroutine=true AND callee null at L1. +func TestGateL1_GoroutineCallNode(t *testing.T) { + out := buildMultipackageV2(t) + worker := mustModule(t, out, "worker/worker.go") + + run, ok := worker.Types["Worker"].Callables["example.com/multipackage/worker.Worker.Run"] + if !ok { + t.Fatal("Worker.Run missing") + } + if len(run.Body) == 0 { + t.Fatal("Worker.Run should have body call nodes") + } + sawGoroutine := false + for local, node := range run.Body { + if node.Kind != "call" { + t.Errorf("body[%s].kind = %q, want call", local, node.Kind) + } + if node.Callee != nil { + t.Errorf("body[%s].callee should be null at L1; got %q", local, *node.Callee) + } + if node.IsGoroutine { + sawGoroutine = true + } + } + if !sawGoroutine { + t.Error("Worker.Run should have a call node with is_goroutine=true (go w.execute(...))") + } +} + +// TestGateL1_CrossPackageStructure locks that a cross-package caller (main.go +// calling into server/ and worker/) is present as its own module with body call +// nodes — the compilation unit spans packages, not a single file. +func TestGateL1_CrossPackageStructure(t *testing.T) { + out := buildMultipackageV2(t) + main := mustModule(t, out, "main.go") + + if main.Package != "main" { + t.Errorf("main.go package = %q, want main", main.Package) + } + // main() drives the cross-package calls; its body must carry call nodes. + var mainFn v2.Callable + found := false + for _, fn := range main.Functions { + if strings.HasSuffix(fn.Signature, ".main") { + mainFn, found = fn, true + } + } + if !found { + t.Fatalf("main function missing; functions=%v", keys(main.Functions)) + } + if len(mainFn.Body) == 0 { + t.Error("main() should have body call nodes for its cross-package calls") + } +} + +// TestGateL1_CrossFileMethodSourceFile locks the source_file amendment +// (CLAUDE.md § Schema decisions → callable.source_file): a method declared in a +// different file than its receiver type stays nested under the type, but its +// span.bytes index its OWN file's source and it carries source_file naming that +// file. In the fixture, Server is declared in server/server.go while its method +// (s Server) Describe() is declared in server/middleware.go. +// +// The text-recovery rule is the contract: a node's text = +// symbol_table[source_file].source[span.bytes] when source_file is present. This +// test slices the declaring module's source by the span and asserts it yields +// the method text — the check that regressed to "" before the fix (empty span +// against the wrong file's source). +func TestGateL1_CrossFileMethodSourceFile(t *testing.T) { + out := buildMultipackageV2(t) + + // Server's type node lives in server/server.go. + server := mustModule(t, out, "server/server.go") + stype, ok := server.Types["Server"] + if !ok { + t.Fatalf("Server type missing; types=%v", keys(server.Types)) + } + const describeID = "example.com/multipackage/server.Server.Describe" + m, ok := stype.Callables[describeID] + if !ok { + t.Fatalf("Server.Describe missing; callables=%v", keys(stype.Callables)) + } + + // It is declared in a different file, so source_file must name that file. + const declFile = "server/middleware.go" + if m.SourceFile != declFile { + t.Errorf("Describe.source_file = %q, want %q (declared apart from its type)", m.SourceFile, declFile) + } + + // The span must be non-empty (the pre-fix bug collapsed it to [n,n]). + if m.Span.Bytes[0] >= m.Span.Bytes[1] { + t.Errorf("Describe span is empty %v — cross-file method span not computed against its own file", m.Span.Bytes) + } + + // Text-recovery rule: slice the DECLARING module's source by the span. + decl := mustModule(t, out, declFile) + text := m.Span.Slice(decl.Source) + if !strings.Contains(text, "func (s Server) Describe()") { + t.Errorf("Describe span over %s does not slice to the method text; got %q", declFile, text) + } + + // Negative: an in-file method (Server.Addr, declared in server.go with its + // type) must NOT carry source_file — absence encodes "same file". + addr, ok := stype.Callables["example.com/multipackage/server.Server.Addr"] + if !ok { + t.Fatalf("Server.Addr missing; callables=%v", keys(stype.Callables)) + } + if addr.SourceFile != "" { + t.Errorf("Addr.source_file = %q, want empty (declared in its type's own file)", addr.SourceFile) + } + if got := addr.Span.Slice(server.Source); !strings.Contains(got, "func (s *Server) Addr()") { + t.Errorf("Addr span over its own module does not slice to the method text; got %q", got) + } +} + +// TestGateL1_Envelope locks the manifest envelope at L1. +func TestGateL1_Envelope(t *testing.T) { + out := buildMultipackageV2(t) + + if out.SchemaVersion != v2.SchemaVersion { + t.Errorf("schema_version = %q, want %q", out.SchemaVersion, v2.SchemaVersion) + } + if out.Language != "go" { + t.Errorf("language = %q, want go", out.Language) + } + if out.MaxLevel != 1 { + t.Errorf("max_level = %d, want 1", out.MaxLevel) + } + if out.Application.ID != "can://go/multipackage" { + t.Errorf("application.id = %q, want can://go/multipackage", out.Application.ID) + } + // L1: no call_graph edges yet. + if len(out.Application.CallGraph) != 0 { + t.Errorf("L1 call_graph should be empty; got %d edges", len(out.Application.CallGraph)) + } +} + +// TestGateL1_SupersetOfV1 is the superset gate: every v1 fact must survive into +// v2 modulo the sanctioned drops. It walks the SAME v1 model the emitter walked +// and asserts each callable/type/function reappears in v2, and that the +// sanctioned drops are gone from the serialized payload. +func TestGateL1_SupersetOfV1(t *testing.T) { + dir := multipackageDir(t) + b := syntactic_analysis.NewSymbolTableBuilder(dir) + st, err := b.Build(nil, true) + if err != nil { + t.Fatalf("symbol table build failed: %v", err) + } + app := &schema.GoApplication{SymbolTable: st, CallGraph: []schema.GoCallEdge{}} + out := Emit(app, "multipackage", dir, 1, "gate-test") + + // Every v1 file, type, method and function reappears in v2. + for rel, v1file := range st { + mod, ok := out.Application.SymbolTable[rel] + if !ok { + t.Errorf("v1 file %q dropped from v2 symbol_table", rel) + continue + } + for name := range v1file.Types { + if _, ok := mod.Types[name]; !ok { + t.Errorf("%s: v1 type %q dropped from v2", rel, name) + } + for msig := range v1file.Types[name].Methods { + if _, ok := mod.Types[name].Callables[msig]; !ok { + t.Errorf("%s: v1 method %q dropped from v2 type %q", rel, msig, name) + } + } + } + for sig := range v1file.Functions { + if _, ok := mod.Functions[sig]; !ok { + t.Errorf("%s: v1 function %q dropped from v2", rel, sig) + } + } + } + + // Sanctioned drops must be absent from the serialized payload. + data, err := json.Marshal(out) + if err != nil { + t.Fatalf("marshal failed: %v", err) + } + if strings.Contains(string(data), `"is_constructor_call"`) { + t.Error("v2 payload should not carry the dropped is_constructor_call") + } + // The sanctioned callee:null slot must be present at L1. + if !strings.Contains(string(data), `"callee":null`) { + t.Error("expected the sanctioned callee:null refinement slot at L1") + } +} + +func contains(ss []string, want string) bool { + for _, s := range ss { + if s == want { + return true + } + } + return false +} diff --git a/internal/syntactic_analysis/v2emit/offsets.go b/internal/syntactic_analysis/v2emit/offsets.go new file mode 100644 index 0000000..e9d3160 --- /dev/null +++ b/internal/syntactic_analysis/v2emit/offsets.go @@ -0,0 +1,94 @@ +// Package v2emit transforms the v1 in-memory model (internal/schema) into the +// v2 canonical schema (internal/schema/v2). It is an EMISSION layer only: it +// computes no new facts beyond source byte offsets, it just re-serializes what +// the parser/resolver already produced into the additive-tree shape. +// +// The one genuinely new datum is span byte offsets. The v1 model records +// line/col but not UTF-8 byte offsets (and only line-level positions for +// callables/types). Rather than change the v1 model, the emitter reads each +// file's source once and derives byte offsets from the existing line/col +// positions via a lineIndex. That same source string is emitted once per module +// as module.source, off which every node's text slices. +package v2emit + +// lineIndex maps 1-based (line, col) positions in a source file to UTF-8 byte +// offsets. Built once per file. Columns are byte columns within the line, which +// is what go/token.Position reports (Position.Column counts bytes, not runes), +// so no rune/byte reconciliation is needed here. +type lineIndex struct { + source string + // lineStart[i] is the byte offset where 1-based line (i+1) begins. + lineStart []int +} + +// newLineIndex builds a lineIndex over the given source text. +func newLineIndex(source string) *lineIndex { + // Line 1 starts at offset 0; each '\n' begins the next line at the byte + // immediately after it. + starts := make([]int, 0, 1+countByte(source, '\n')) + starts = append(starts, 0) + for i := 0; i < len(source); i++ { + if source[i] == '\n' { + starts = append(starts, i+1) + } + } + return &lineIndex{source: source, lineStart: starts} +} + +// offset converts a 1-based (line, col) position to a byte offset into source. +// col is a 1-based byte column (go/token semantics). Positions outside the +// file's range are clamped to [0, len(source)] so a bad position yields a +// harmless in-range offset rather than a panic. +func (li *lineIndex) offset(line, col int) int { + if len(li.lineStart) == 0 { + return 0 + } + if line < 1 { + line = 1 + } + if line > len(li.lineStart) { + line = len(li.lineStart) + } + base := li.lineStart[line-1] + if col < 1 { + col = 1 + } + off := base + (col - 1) + if off > len(li.source) { + off = len(li.source) + } + return off +} + +// lineEndOffset returns the byte offset of the end of a 1-based line: the +// position just before its terminating '\n' (or end of file for the last line). +// Used to give a line-only position (no column in the v1 model) a precise end. +func (li *lineIndex) lineEndOffset(line int) int { + if line < 1 || len(li.lineStart) == 0 { + return 0 + } + if line >= len(li.lineStart) { + return len(li.source) + } + // Next line begins one byte after this line's '\n'; back up over the '\n'. + return li.lineStart[line] - 1 +} + +// lineLen returns the number of bytes on a 1-based line, excluding the newline. +// It gives a line-only end position a byte column (used to synthesize end.col). +func (li *lineIndex) lineLen(line int) int { + if line < 1 || line > len(li.lineStart) { + return 0 + } + return li.lineEndOffset(line) - li.lineStart[line-1] +} + +func countByte(s string, b byte) int { + n := 0 + for i := 0; i < len(s); i++ { + if s[i] == b { + n++ + } + } + return n +} diff --git a/internal/syntactic_analysis/v2emit/offsets_test.go b/internal/syntactic_analysis/v2emit/offsets_test.go new file mode 100644 index 0000000..69b61b9 --- /dev/null +++ b/internal/syntactic_analysis/v2emit/offsets_test.go @@ -0,0 +1,69 @@ +package v2emit + +import "testing" + +const sample = "package util\n" + // line 1: bytes 0..12 (\n at 12) + "\n" + // line 2: byte 13 (\n at 13) + "func f() {}\n" + // line 3: bytes 14..25 (\n at 25) + "var x int" // line 4: bytes 26..34, no trailing newline + +func TestLineStarts(t *testing.T) { + li := newLineIndex(sample) + want := []int{0, 13, 14, 26} + if len(li.lineStart) != len(want) { + t.Fatalf("lineStart len = %d, want %d (%v)", len(li.lineStart), len(want), li.lineStart) + } + for i, w := range want { + if li.lineStart[i] != w { + t.Errorf("lineStart[%d] = %d, want %d", i, li.lineStart[i], w) + } + } +} + +func TestOffset(t *testing.T) { + li := newLineIndex(sample) + // line 3 ("func f() {}") col 1 → byte 14; col 6 ('f') → 19 + if got := li.offset(3, 1); got != 14 { + t.Errorf("offset(3,1) = %d, want 14", got) + } + if got := li.offset(3, 6); got != 19 { + t.Errorf("offset(3,6) = %d, want 19", got) + } + // The slice from a callable start to its end reproduces the source text. + start := li.offset(3, 1) + end := li.lineEndOffset(3) + if got := sample[start:end]; got != "func f() {}" { + t.Errorf("slice = %q, want %q", got, "func f() {}") + } +} + +func TestOffsetClamping(t *testing.T) { + li := newLineIndex(sample) + if got := li.offset(0, 0); got != 0 { + t.Errorf("offset(0,0) = %d, want 0 (clamped)", got) + } + if got := li.offset(999, 999); got != len(sample) { + t.Errorf("offset(999,999) = %d, want %d (clamped to EOF)", got, len(sample)) + } +} + +func TestLineEndAndLen(t *testing.T) { + li := newLineIndex(sample) + // last line has no trailing newline → ends at EOF + if got := li.lineEndOffset(4); got != len(sample) { + t.Errorf("lineEndOffset(4) = %d, want %d", got, len(sample)) + } + if got := li.lineLen(3); got != len("func f() {}") { + t.Errorf("lineLen(3) = %d, want %d", got, len("func f() {}")) + } + if got := li.lineLen(4); got != len("var x int") { + t.Errorf("lineLen(4) = %d, want %d", got, len("var x int")) + } +} + +func TestEmptySource(t *testing.T) { + li := newLineIndex("") + if got := li.offset(1, 1); got != 0 { + t.Errorf("offset on empty source = %d, want 0", got) + } +} diff --git a/internal/syntactic_analysis/v2emit/sigindex.go b/internal/syntactic_analysis/v2emit/sigindex.go new file mode 100644 index 0000000..f0e8cd8 --- /dev/null +++ b/internal/syntactic_analysis/v2emit/sigindex.go @@ -0,0 +1,59 @@ +package v2emit + +import ( + "github.com/codellm-devkit/codeanalyzer-go/internal/schema" + v2 "github.com/codellm-devkit/codeanalyzer-go/internal/schema/v2" +) + +// sigIndex maps a v1 callable signature to its v2 can:// callable id. +// +// v1 identifies every callable by its signatureOf() string and uses that string +// as call-graph edge endpoints and as a call site's backfilled callee. v2 +// addresses callables by durable can:// id. The L2 refinement therefore needs a +// signature→id translation: this index is built in one pass over the same tree +// the emitter walks, using the SAME id builders emitCallable uses, so an id here +// is byte-identical to the id the callable node carries. That is what lets the +// L2 gate assert "no dangling endpoints" — every edge endpoint resolves to a +// real node id. +// +// v1 signatures are globally unique (a method signature embeds its receiver +// type), so a flat signature→id map has no collisions. +type sigIndex map[string]string + +// buildSigIndex walks the symbol table and records, for every callable +// (functions, methods, and nested closures), signature → can:// id. It mirrors +// emitModule/emitType/emitCallable's parent-id computation exactly. +func buildSigIndex(appID string, symbolTable map[string]schema.GoFile) sigIndex { + idx := make(sigIndex) + for relPath, file := range symbolTable { + modID := v2.ModuleID(appID, relPath) + for _, fn := range file.Functions { + idx.addCallable(modID, fn) + } + for _, t := range file.Types { + typeID := v2.TypeID(modID, t.Signature) + for _, m := range t.Methods { + idx.addCallable(typeID, m) + } + } + } + return idx +} + +// addCallable records one callable and recurses into its closures, matching the +// id nesting emitCallable produces. +func (idx sigIndex) addCallable(parentID string, c schema.GoCallable) { + callID := v2.CallableID(parentID, c.Signature) + idx[c.Signature] = callID + for _, inner := range c.InnerCallables { + idx.addCallable(callID, inner) + } +} + +// idFor returns the can:// id for a v1 signature, or ("", false) when the +// signature names a callable outside the project (stdlib/external) — those +// resolve to no in-tree node and must not become an edge endpoint. +func (idx sigIndex) idFor(sig string) (string, bool) { + id, ok := idx[sig] + return id, ok +} diff --git a/internal/utils/fs.go b/internal/utils/fs.go index e46da2e..541127c 100644 --- a/internal/utils/fs.go +++ b/internal/utils/fs.go @@ -35,6 +35,24 @@ func RelativePath(root, path string) string { return rel } +// IsWithin reports whether path lies inside root (or is root itself). It is the +// test for "this declaration's file is part of the analyzed project": a path +// whose relative form escapes root (starts with ".." or is absolute) is outside. +// cgo synthesizes wrapper functions (_Cfunc_*, _Cgo_*) into files under +// $GOCACHE, so their declaring file resolves outside the input root — IsWithin +// is how the symbol-table builder tells those toolchain artifacts from project +// source. +func IsWithin(root, path string) bool { + rel, err := filepath.Rel(root, path) + if err != nil { + return false + } + if rel == ".." || strings.HasPrefix(rel, ".."+string(filepath.Separator)) { + return false + } + return !filepath.IsAbs(rel) +} + // FileHash returns the SHA-256 hex digest of the file at path. func FileHash(path string) (string, error) { f, err := os.Open(path) diff --git a/internal/utils/fs_test.go b/internal/utils/fs_test.go index 91df3d4..4e12166 100644 --- a/internal/utils/fs_test.go +++ b/internal/utils/fs_test.go @@ -53,6 +53,32 @@ func TestIsVendored(t *testing.T) { } } +// ── IsWithin ───────────────────────────────────────────────────────────────── + +func TestIsWithin(t *testing.T) { + root := filepath.Join(string(filepath.Separator), "home", "user", "proj") + tests := []struct { + path string + want bool + }{ + // Inside the project root. + {filepath.Join(root, "main.go"), true}, + {filepath.Join(root, "pkg", "geo", "geos", "geos.go"), true}, + {root, true}, // root itself resolves to "." + // Outside: the go-build cache cgo synthesizes into — the ISSUE-2 shape. + {filepath.Join(string(filepath.Separator), "home", "user", "Library", "Caches", "go-build", "bb", "hash-d"), false}, + // A sibling directory sharing a prefix but not nested under root. + {filepath.Join(string(filepath.Separator), "home", "user", "proj-other", "x.go"), false}, + // A parent directory. + {filepath.Join(string(filepath.Separator), "home", "user"), false}, + } + for _, tc := range tests { + if got := utils.IsWithin(root, tc.path); got != tc.want { + t.Errorf("IsWithin(%q, %q) = %v, want %v", root, tc.path, got, tc.want) + } + } +} + // ── FileHash ────────────────────────────────────────────────────────────────── func TestFileHash_Deterministic(t *testing.T) {