Skip to content

V1 to v2 schema migration - #7

Closed
lamwassi wants to merge 16 commits into
codellm-devkit:mainfrom
lamwassi:v1-to-v2-schema-migration
Closed

lamwassi wants to merge 16 commits into
codellm-devkit:mainfrom
lamwassi:v1-to-v2-schema-migration

Conversation

@lamwassi

@lamwassi lamwassi commented Oct 2, 2026

Copy link
Copy Markdown

v2 canonical schema migration (L1 + L2) for the Go backend

What

Migrates the cango Go backend to emit the canonical CLDK v2 analysis.json,
covering analysis levels L1 (additive containment tree) and L2 (edge/callee
refinement). v1 output is unchanged and remains the default; v2 is opt-in behind
--analysis-schema 2, so the current SDK keeps parsing output throughout the migration.
L3/L4 are deferred to a later train (Group A ids/spans are already designed for reuse).

Design

Decided in the designing-cldk-changes loop and committed as the full transcript in
docs/design/specs/v2-l1-emission.md (+ roadmap, cross-repo spec, drafted parent/child
issue bodies). Per the parity clause, the shared Group A vocabulary (can:// ids,
span.bytes as UTF-8 byte offsets) is decided once here because L3/L4 reuse it verbatim.

Changes

Schema & identity (internal/schema/v2)

  • v2 canonical-schema Go types: application → module → type → callable → body tree with
    typed edge overlays and the L2 refinement slots.
  • type.kind (struct|interface|alias|defined) collapses v1's is_interface;
    base_types (embedding) vs interfaces (computed satisfaction); error_channel from
    error-typed returns; nested callables{} for closures; is_constructor_call dropped.
  • Pure can:// id builders and Span with Slice (O(1) text recovery off
    module.source, replacing v1's per-node code field).

Emitter (internal/syntactic_analysis/v2emit)

  • lineIndex: reads source once, maps 1-based (line, col) → UTF-8 byte offsets (go/token
    columns are byte columns, so no rune/byte reconciliation). Keeps the v1 model diff-free.
  • L1: Emit() re-serializes the v1 GoApplication into the v2 tree, computing no new
    facts beyond span offsets. Module source carried once; methods nest under their receiver
    type.
  • L2: pure refinement over the L1 tree via buildSigIndex (v1 signature → v2 can://
    id, byte-identical to the node's own id, so no dangling endpoints). Backfills
    body.callee null→id for in-tree callees (null kept for external/stdlib), and translates
    call_graph edge endpoints onto real ids.

CLI (cmd/codeanalyzer, internal/options, internal/core)

  • Conforms to the CLDK CLI contract: --emit <json|neo4j|schema>, --app-name,
    --neo4j-*, -j/--jobs. Unimplemented projections return a clear non-zero error (never
    a silent fallback).
  • --analysis-schema 2 selects v2; default stays v1. core.RenderJSON centralizes the
    v1/v2 marshal decision so stdout and file paths agree; unknown schema errors out.

Cross-file method fix (source_file)

  • Go allows a method to be declared in a different file than its receiver type. The v2 tree
    nests it under the receiver's module, but its positions index its own file — collapsing
    span.bytes to an empty slice (~3292 cases across dae/openbao/cockroach). Added optional
    callable.source_file: when the declaring file differs from the nesting module,
    span.bytes index that file and text recovers via symbol_table[source_file].source.
    Cross-language parity deferred (tracked follow-up on the SDK rung).

cgo artifact fix

  • cgo synthesizes wrapper functions into a generated file under $GOCACHE; emitting them
    produced callables whose span couldn't slice against any project module (ISSUE-2: 85
    callables / 106 failures on cockroach). buildCallable now returns nil for declarations
    outside the input root. Adds utils.IsWithin.

Docs

  • README: "Generating analysis.json for a Go app" walkthrough, v2 output shape alongside
    legacy v1, corrected CLI options. Plus a keys-only v2 structure reference.

Tests

  • L1 fixture gate: concrete values for every L1 checklist item (multi-file units,
    exported/unexported, interface-kind collapse, error_channel, variadics, is_goroutine,
    callee:null slot, can:// ids) + a superset gate (every v1 file/type/method reappears;
    sanctioned drops absent).
  • L2 fixture gate: no dangling endpoints, every edge carries prov, named expected edges
    by exact id pair (incl. cross-package), goroutine-site callee backfill, and the L1 ⊆ L2
    superset (only callee null→id changes).
  • Cross-file source_file gate (positive + in-file negative); lineIndex, id builder, and
    IsWithin unit tests.
  • Verified on the real-app eval: cockroach L1 106→0 failures.

Compatibility

Default output stays v1 (top-level symbol_table, no schema_version). v2 is strictly
opt-in. No changes to the v1 model.

Roadmap, cross-repo spec, schema decisions, and drafted parent/child issue
bodies for migrating codeanalyzer-go and its python-sdk Go models to the v2
canonical schema. Reach: L1 + L2 to parity; L3/L4 deferred to a later train.

Group A vocabulary (can:// ids, span.bytes UTF-8) decided once here per the
parity clause, since L3/L4 reuse it verbatim. Full transcript in
docs/design/specs/v2-l1-emission.md and CLAUDE.md.
Add the flags the CLDK CLI contract requires so the SDK facade can shell
out to the Go backend with the same code path as the other language
backends:

- --emit <json|neo4j|schema>: output projection selector. neo4j and
  schema return a clear non-zero "not yet implemented" error (never a
  silent fallback to JSON) until the Neo4j child lands. schema needs no
  --input, matching the contract.
- --app-name: application anchor for can://go/<app> ids and the Neo4j
  :Application node; defaults to the input dir base name.
- --neo4j-uri/user/password/database: accepted and validated now with
  explicit-flag > env-var > default precedence; consumed by the Neo4j
  child.
- -j/--jobs: worker parallelism, defaulting to CPU cores.

No v2 emission logic yet; default output stays the v1 JSON shape.
Declare the v2 canonical-schema Go types in internal/schema/v2, a
faithful transcript of the keystone (canonical-schema.md) and the
Go-specific leaf decisions in CLAUDE.md.

Models the L1 additive containment tree
(application -> module -> type -> callable -> body) with typed edge
overlays, plus the L2 refinement/edge slots so the shape is stable:

- Span{start,end,bytes} with UTF-8 byte offsets into module.source.
- Module carries source once; no per-node code field.
- Type.kind (struct|interface|alias|defined) replaces v1 is_interface;
  base_types (embedding) vs interfaces (computed satisfaction).
- Callable.error_channel, nested callables{} for closures, metrics map.
- BodyNode 'call' with the sanctioned callee null->id refinement slot
  (pointer, the one allowed null) and is_goroutine/is_deferred.
- Edge{src,dst,prov,weight} for call_graph.

No emitter or wiring yet; these types are not referenced by any code
path. L3/L4 fields (cfg/cdg/ddg/param_*) are deliberately not declared.
Add the Group A identity/positioning helpers in internal/schema/v2,
pure functions so they are testable in isolation and reused verbatim by
the deferred L3/L4 train:

- can:// durable-id builders (AppID/ModuleID/TypeID/CallableID) that
  assemble the containment path can://<lang>/<app>/<file>/<type>/<sig>,
  with signatureOf() output as the last segment.
- @line:col ordinal-id builders (OrdinalID/TagID/LocalID) for body nodes
  below the callable line.
- NewSpan + Span.Slice: build a span from line/col and UTF-8 byte
  offsets, and slice module.source in O(1) (the replacement for the v1
  per-node code field), degrading to "" on a malformed span.

Unit tests reproduce the keystone worked-example id chain and cover path
normalization, module-level function parents, ordinal/tag ids, and span
slicing edge cases.

Not wired into any emitter yet.
Add the v2emit package with lineIndex: reads a file's source once and
maps 1-based (line, col) positions to UTF-8 byte offsets, the one new
datum the v2 span needs. go/token columns are byte columns, so no
rune/byte reconciliation is required.

This keeps the v1 model diff-free (per the migration spec): the emitter
derives span.bytes from the existing line/col positions instead of
threading offsets through internal/schema. lineEndOffset/lineLen give a
line-only v1 position (callables/types carry no columns) a precise end.

Positions outside the file clamp to [0, len(source)] rather than
panicking. Tests cover line-start tabulation, offset lookup, clamping,
the last-line-without-newline case, and empty source.

Emitter tree walk lands next; nothing wired yet.
Emit() transforms the v1 GoApplication into the v2 Analysis payload: a
pure re-serialization walking application -> module -> type -> callable
-> body, computing no new facts beyond span byte offsets.

- Module carries source once (read from projectDir/relPath); every node
  span slices off it via the lineIndex-derived byte offsets.
- can:// ids assembled from the id builder, signatureOf() as last
  segment; methods resolve under their receiver type's callables{}.
- type.kind collapses is_interface (struct|interface); base_types kept;
  error-typed returns lift into error_channel; cyclomatic into metrics;
  closures nest in callables{}.
- body holds one 'call' node per call site keyed by line:col, with the
  callee null->id slot carried straight from the v1 pointer.
- call_graph renamed to {src,dst,prov,weight}.

Also drop omitempty from BodyNode.Callee so the sanctioned refinement
slot always serializes (null at L1, id at L2) as the keystone requires.

Tests build the greeter fixture through the real symbol-table builder
and assert the envelope, module source/ids, span slicing, call-node
bodies, callee-null serialization, dropped is_constructor_call, and
type/method kinds. Not wired to the CLI yet.
Wire the v2 emitter behind an opt-in flag. Default stays the v1 shape so
the current SDK keeps parsing output during the migration; --analysis-
schema 2 emits the canonical v2 tree.

- Add options.SchemaVersion (1 default, 2 v2) and AnalyzerVersion.
- core.RenderJSON centralizes the v1/v2 marshal decision so the stdout
  and file paths agree; WriteOutput now takes opts and routes through it.
  An unknown schema returns a non-zero error, never a silent fallback.
- --analysis-schema flag; --app-name/version threaded into the v2
  manifest (v2 surfaces the analyzer version, which v1 never carried).

Update WriteOutput callers (now opts-based) in core tests. CLI tests
assert the default is v1 (top-level symbol_table, no schema_version),
that --analysis-schema 2 emits schema_version 2.0.0 with
application.id can://go/greeter, and that an unknown schema errors.
Lock the v2 L1 output shape against a known multi-file fixture so the
level cannot silently regress. Asserts concrete values for every item on
the L1 coverage checklist (testing-and-validation.md):

  - multi-file compilation unit: four modules keyed by relative path,
    each with source-once and a whole-file span that slices back
  - exported + unexported symbols (Worker.Run vs Worker.execute)
  - interface kind collapse (Processor -> kind=interface)
  - error_channel from an error-typed return (Processor.Process)
  - variadic parameter (Combine(results ...Result))
  - language-specific call flag (is_goroutine on `go w.execute(...)`)
  - sanctioned callee:null refinement slot present at L1
  - can:// durable ids at module/type/callable depth

Plus the superset gate: walk the same v1 model the emitter walked and
assert every file/type/method/function reappears in v2, and that the
sanctioned drops (is_constructor_call, per-node code) are absent from the
serialized payload.
L2 is a pure refinement over the L1 tree: it fills the sanctioned
callee null->id slot and translates call_graph edge endpoints onto
real callable ids. Both need one thing v1 does not have: a mapping from
v1 signature identity to v2 can:// node id.

Add buildSigIndex, a one-pass signature->id index built with the same id
builders emitCallable uses, so an indexed id is byte-identical to the id
the callable node carries (no dangling endpoints by construction). Then:

  - emitBody: backfill body.callee to the callee's can:// id when the
    resolver-set signature names an in-tree callable; keep it null for
    external/stdlib callees (the honest-unresolved fallback).
  - emitCallGraph: map src/dst from v1 signatures to can:// ids; drop any
    edge whose endpoint does not resolve rather than emit a dangling one
    (defensive — the v1 resolver only emits in-project targets).

At L1 (no resolver run) callee stays null, so L1 output is unchanged.
Lock the v2 L2 output shape by building the fixture through the real
resolver (the -a 2 wiring) and asserting the L2 checklist
(testing-and-validation.md):

  - no dangling endpoints: every edge src/dst resolves to a real
    callable id in the tree
  - every edge carries a non-empty prov naming the resolver
  - a NAMED expected edge (Worker.Run -> Worker.execute) by exact
    can:// id pair, so the gate proves correctness not just non-emptiness
  - a NAMED cross-package edge (main -> server.New)
  - callee backfilled to an in-tree can:// id on the goroutine call site
  - the L1 ⊆ L2 superset: emitting the same model at level 1 vs 2 leaves
    every module/type/callable id, span, kind and body node identical,
    changing only callee null -> id (the one sanctioned mutation).
The real-app eval (dae/openbao/cockroach) surfaced a datamodel collision:
v2 nests a method under its receiver type, but Go allows a method to be
declared in a different file than its type, so span.bytes computed against
the nesting module's source collapse to empty (~3292 cases; cobra clean).

Design-loop decision (Go-only for now; cross-language parity deferred):
add an optional callable.source_file field, present only when the declaring
file differs from the nesting module. span.bytes then index that file's
module.source; text = symbol_table[source_file].source[span.bytes] when
present, else the nesting module. Containment stays receiver-based.

Records the decision in the committed spec and CLAUDE.md; python-sdk
source_file support is a tracked follow-up on the SDK rung.
A Go method may be declared in a different file than its receiver type.
The v2 tree nests it under the receiver type's module, but its line/col
positions are relative to its OWN file, so computing span.bytes against the
nesting module's source collapsed them to an empty [n,n] slice (~3292 cases
across dae/openbao/cockroach; cobra clean).

Compute a cross-file method's span against its declaring file (c.Path) via a
per-file lineIndex cache, and record that file in the new optional
callable.source_file field. Text recovery: symbol_table[source_file].source
[span.bytes] when set, else the nesting module. Closures inherit their
enclosing callable's file. Containment stays receiver-based.

Adds TestGateL1_CrossFileMethodSourceFile over the multipackage fixture
(Server declared in server.go, its Describe method in middleware.go):
asserts source_file names the declaring file, the span is non-empty, and it
slices to the method text; plus the negative case (in-file method carries no
source_file). Design: docs/design/specs/v2-l1-emission.md, CLAUDE.md.
cgo synthesizes wrapper functions (_Cfunc_*, _Cgo_*, _cgo_runtime_*) into a
generated file under $GOCACHE; go/packages resolves their token.Pos to that
cache path. Emitting them produced callables whose source_file pointed at the
build cache and whose span could not be sliced against any project module
(ISSUE-2 in the real-app eval: 85 such callables on cockroach, 106 validation
failures at L1 and L2; cobra/dae/openbao have no cgo and were unaffected).

These are toolchain artifacts, not project source. buildCallable now returns
nil when the declaration's file is not within projectDir, at the point its
(cache) file is first known; both callers already skip nil. Adds utils.IsWithin
and a unit test.

Verified: cockroach L1 106->0 failures, 85 dropped signatures all cgo glue,
real pkg/geo/geos project code preserved.
A quick key-skeleton map of the canonical v2 output: the node tree, the recursive
callable node, span shape, presence rules (always-present vs omitempty vs
emitted-when-true), and can:// id shapes. Authored from the JSON tags in
internal/schema/v2/schema.go; links to v2-l1-emission.md for the decisions.
- Add a step-by-step 'Generating analysis.json for a Go app' section
  (--analysis-schema 2 -a 2 -o <dir>, --app-name, verify, cgo note).
- Document the canonical v2 output shape alongside the legacy v1 shape,
  with a full JSON example verified against real emitter output.
- Fix the stale CLI options block (add --analysis-schema, --app-name).
The v2 call_graph edge list was emitted in Go-map iteration order, so two
runs of the same input (even at the same -j) produced byte-differing
analysis.json, failing the release determinism gate. Sort the emitted edges
by (src, dst, prov, weight) in emitCallGraph, matching the sorted-key
discipline already used for module output. Add a determinism gate test that
emits a reversed edge slice and asserts identical serialized call_graph.
@lamwassi
lamwassi marked this pull request as ready for review October 2, 2026 15:06
@lamwassi
lamwassi marked this pull request as draft October 2, 2026 15:07
@lamwassi lamwassi closed this Oct 3, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant