All notable changes to this project will be documented in this file.
The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
locateplaces a position on a decorator line on the decorated callable, on both Python backends (python-sdk#408).PyCallable.start_lineis thedefline (ast.FunctionDef.lineno) and Python's AST puts decorators above it, so@http.route(...)sat inside no callable span and came back asmodule_scope. Both backends now read where the decorator is applied:_find_innermostadmits the callable whosedecorators[].spanstarts at or before the line and above itsdef;_LOCATE_QUERYgains oneEXISTSdisjunct overPY_DECORATED_BY.start_line, which codeanalyzer-python 1.5.2 emits. Ranking is unchanged (the def-based width), so a decorator on a nested callable resolves to the nested one. Three things deliberately stay as they were: a position between two callables with no decorator over it is stillmodule_scope(this is not a nearest-callable fallback); a class decorator is stillmodule_scope(a result withtypeset andcallableunset is a new shape and needs its own decision); and an analysis or graph from codeanalyzer-python 1.5.1 or earlier, which records no decorator span, answers exactly as before. Measured on odoo-slim-19 re-emitted by 1.5.3, 40 positions, median of 5, same graph both ways: 71.0 ms placing 20/40 → 70.2 ms placing 40/40; the plan still seekspysymbol_id.
codeanalyzer-pythonpin1.5.2→1.5.3(dependenciesand[tool.backend-versions],python-sdk#410). Two analyzer fixes, neither moving the schema (schema_versionand the graphSCHEMA_VERSIONstay2.0.0;PyNeo4jBackend._ANALYZER_FLOORstays1.5.0):- the analyzer no longer ingests its own output (codeanalyzer-python#207).
discover_artifactshad no idea where the run writes, so an output or cache directory inside the input made run N inventory run N-1'sanalysis.jsonand embed it whole. This SDK's defaultcache_diris<project>/.codeanalyzer, inside the input, so every local run was exposed. - one body node per call site when two calls share a start position (codeanalyzer-python#215).
getattr(o, n)(x)starts the outer application and the innergetattrat the same column, and the position-keyedbodykept the inner one. Keys now carry a/2,/3disambiguator, outermost first;body_key_columnalready splits on/and needed no change. A key that used to collide now names the outer call, and a call whose callee is itself a call carriescallee_signature: nullwithmethod_name: "<unknown>".
- the analyzer no longer ingests its own output (codeanalyzer-python#207).
The Python tier offline (tests/analysis/python, mocked Neo4j and the in-process analyzer):
482 passed, 168 skipped, on 1.5.2 and on 1.5.3 alike — every skip is a live-Neo4j tier. No fixture
expectation moved.
The Java facade learns the view layer codeanalyzer-java 3.3.0–3.3.2 added, and the 2.0 line takes up
everything rc.6 shipped from main (the 3.2.0 text model and the sibling-analyzer pins) — this is
the first release cut from release/2.0 that contains rc.6.
Java view dispatches (python-sdk#404, spec codellm-devkit/.github
docs/design/specs/2026-09-11-java-view-templates-and-dispatch.md). JSP, Facelets and Thymeleaf
templates are artifacts with roles: ["view-template"], and the analyzer now says which code reaches
one. Three accessors on JavaAnalysis, both backends:
get_view_dispatches(view=None) -> List[JViewDispatch]— every resolved dispatch: the body node that hands the request over (aforward/include/sendRedirectcall, aModelAndViewconstruction orsetViewName, or a Spring controller'sreturn) and the artifact it reaches, withvianaming the mechanism andprovthe tier.literalanddataflowmean exactly one target;table(3.3.1) is a may-dispatch over a static string table, one edge per entry, from the same site.viewfilters by a segment-aligned suffix of the artifact's path.get_view_dispatchers(path) -> List[JCallableOverview]— the same edges resolved to the callables that own the dispatching sites.get_unresolved_view_dispatches() -> List[JViewDispatchUnresolved]— what closed on nothing (a variable target, a servlet URL, a view name matching two templates), with its reason and the tiers attempted. JSON-only: the graph has no node for a target that resolved to nothing, so the Neo4j backend refuses this one rather than answering[].
Two models, JViewDispatch and JViewDispatchUnresolved, on JApplication.view_dispatches /
.view_dispatches_unresolved; the Neo4j backend reads J_DISPATCHES_TO back into the former.
One rule stated because it is the only one of its kind: this layer is gated on the analyzer
generation. The 3.1.0 config trio is probed by an always-present sibling key; 3.3.x added none —
both lists are written only when non-empty and the edge type is declared only once an edge exists —
so a pre-3.3.0 analysis is refused from analyzer.version (wire) / analyzer_version (graph) instead,
because [] off a 3.2.0 analysis would say "reaches no view" where the truth is that nothing looked.
Measured on daytrader8 with 3.3.3: 37 edges at level 1–2 (3 literal, 34 table), 54 at level 4,
19 of 23 JSPs reached; the TradeConfig.webUI table's 17 on-disk pages all reached.
codeanalyzer-javapin3.2.0→3.3.3(thejavaextra and[tool.backend-versions]). 3.3.0 adds the view-template role,J_DISPATCHES_TO, andargument_expronreturnbody nodes; 3.3.1 the table tier; 3.3.2 narrows the Jakarta entrypoint finder to lifecycle methods, so helpers taking servlet-typed parameters are no longer entrypoints and the interprocedural config-use and view-dispatch tiers bind their parameters (daytrader8 entrypoint marks 135 → 196: 21 helpers dropped, 82init/destroy/doFilter/onMessagegained); 3.3.3 fixes theNullPointerException3.3.0–3.3.2 threw on any unbuilt project with a dispatch call, which this SDK'stest_java_degradation.pywas the first to hit (codeanalyzer-java#265).release/2.0absorbedmainat rc.6.
The Java suite offline on both backends over a1 with the layer injected (edges, suffix filter, dispatcher resolution, the JSON-only refusal on the graph, the pre-3.3.0 refusal, the empty answer on a 3.3.x payload that reaches nothing, and the answered-once check across backends); the public surface pins the three signatures; e2e against the 3.3.3 wheel on daytrader8 at level 1 and 2 asserts the 37 edges, the 19 views, the three unresolved URLs and the 38 view-template artifacts.
The Java Neo4j backend stops being the lossy one about source text. Until now a graph-backed Java
analysis could answer where a node is far better than what it says: text existed only as
:JCallable.code, so a callable came back as its whole declaration where the local backend returned
the body block, a body node had no text at all, and a module-scope locate returned "". The
analyzer closed that gap in 3.2.0; this release takes it up, and moves the sibling pins to the
releases that closed the same gap for Python and TypeScript.
The Java Neo4j attach floor moves to codeanalyzer-java 3.2.0. A graph emitted by 3.1.x — served silently until now — is refused at attach, naming what was found and the floor.
The floor moves rather than degrading gracefully because the degradation would be silent and would
look like data. A 3.1.x graph has no :JModule.source and no byte offsets, so every text answer
would come back empty or as the old whole-declaration string, from accessors this release documents
as returning the same text as the local backend. An agent cannot tell "this node has no text" from
"this graph predates the property" — so the check belongs at attach, once, where it can say so.
Migration: re-emit with codeanalyzer-java 3.2.0 --emit neo4j. Wipe the database first when the
graph predates 3.1.1, which moved the can:// grammar.
Text is the same on both Java backends. JCallable.code, get_source, LocateResult.source and
describe answer the same slice of the same file at every granularity — a callable's body block, a
type's or field's or local's declaration, a single statement or call site, and a module's whole text at
module scope. Each is a byte slice of :JModule.source taken at the node's own offsets, so the live
suites assert equality rather than endswith or "the same length".
Three positions still have no text, on both backends and for the same reason each time — there is
nothing in the file to point at: an implicit compiler-generated <init>(), an @external callable, and
a resolve_value ref (a formal_in vertex is a dataflow position, not a region of the file).
module_source_unavailable survives as a data diagnostic: it now fires on a module whose text the
analyzer could not read or decode, and on nothing else.
Two removals fall out of it. :JCallable.code is no longer read — the slice replaces it — and
body_end_byte was never needed: across both committed fixtures, all 1,182 callables carrying both
spans end their body exactly where their declaration ends, because a Java declaration's last character
is its body's closing brace. That is asserted in the fixture builder rather than written down.
Analyzer pins move to codeanalyzer-python==1.5.2 and codeanalyzer-typescript==1.6.0
(codeanalyzer-java is already at 3.2.0). Both are Neo4j-projection-only conformance fixes: python
1.5.2 adds :PyModule.source, all six span properties wherever a span is declared,
:PyVariable.value_json, :PyBodyNode.callee_signature, PY_IMPORTS.positions_json and a span on
PY_DECORATED_BY; typescript 1.6.0 adds :TSModule.source and the four new span properties across ten
labels. Neither changes analysis.json — the TypeScript fixtures regenerated with the 1.6.0 wheel are
byte-identical to the 1.5.3 generation at all four levels apart from the analyzer.version stamp.
Neither sibling floor moves. They stay 1.5.0 for Python and 1.5.2 for TypeScript: a floor move is what would let the SDK read that new text, and that is uptake work the size of this release's Java half, tracked in #396 and #391. The pin is what the SDK installs; the floor is what it requires of a graph. Every sentence in the docs that read "the graph does not carry this" now says the SDK does not read it, and names the issue that will.
One thing does improve for free on a re-emitted Python graph: a call site's columns. The
reconstruction already reads start_column with a -1 default, so 1.5.2 supplies real values and an
older graph still answers -1 — data-driven, no version branch.
The mocked, backend-contract and E2E tiers are green, and the Java E2E ran against the real 3.2.0 jar. The live Java tiers — the ones that assert byte-for-byte equality across the two backends on daytrader8 — assert the new behaviour but have not been run against a 3.2.0-emitted graph: the verification corpus holds a 3.1.1 graph, which this release's own floor now refuses. Re-emitting it is the outstanding verification debt, recorded rather than glossed.
Two changes, and they are the two halves of one question: what should happen when an analyzer ships a field the SDK has not learned about yet. Until now the answer was "refuse the entire payload." That answer caught a real bug within seconds — and it also meant a backward-compatible analyzer release could not be taken up at all until the SDK was edited first. This release takes up the fix and changes the answer.
The schema-v2 mirrors relax from extra="forbid" to extra="ignore". A Java or TypeScript
analysis.json carrying a field these models do not declare now parses, with that field dropped,
where it previously raised.
What it gives up, stated plainly because it is the whole cost: forbid was the drift detector. Under
ignore an addition is absorbed silently, and the first symptom is a wrong answer from something
reading a field the SDK never learned. The intended replacement is a comparison against each
analyzer's published schema, reporting what is unmodelled instead of refusing to parse.
Scoped narrower than the whole codebase, on a distinction that matters:
- relaxed —
_Baseincldk/models/javaandcldk/models/typescript, the analyzer mirrors validated from the wire, plusJCompilationUnit, which overridesmodel_configwholesale for its alias settings and so never inherited the change (a subclass config replaces rather than merges); - kept
forbid—JCallableOverviewandJClassOverviewinprojections.py. Nothing callsmodel_validateon them; this SDK builds them from graph rows. Strictness there catches our own typo'd kwarg, not the analyzer's additions.
ignore rather than allow, deliberately: allow keeps unknown fields in model_extra and so
widens model_dump_json() with whatever the analyzer emitted, and several tests assert properties
of dumps. A field the SDK does not model must not be able to change what a dump contains.
What this does not cost: extra does not govern required fields, so "a 1.x analysis.json is
refused, not parsed" survives intact — a v1 payload still fails, because it lacks what v2 requires
rather than carrying extras.
Analyzer pins move to codeanalyzer-python==1.5.1, codeanalyzer-java==3.1.2 and
codeanalyzer-typescript==1.5.3. All three shipped the same fix in lockstep: param_in and
param_out edges now name the bound formal in var. The schemas had declared that property since
the L4 layer landed while the projections wrote nothing, so a consumer predicate on it evaluated to
null on every edge crossing a call boundary — and under Cypher's three-valued logic an all() over
that null excludes the whole path. An interprocedural flow therefore read as a proved absence of
flow, indistinguishable from a real negative.
Measured on a graph emitted by rc.4's pinned codeanalyzer-python 1.5.0: var was null on 4 of 4
PY_PARAM_IN edges and 6 of 6 PY_PARAM_OUT, while all 44 PY_DDG carried it. So this was the
shipped state of rc.4, not a legacy-graph edge case. After the bump: 4 of 4 and 6 of 6 carry it.
No analyzer floor moves. Each delta is the var fix plus release plumbing, with no change to the
can:// grammar. Raising a floor would refuse graphs that still work, since the fix only adds a
property — and a capability difference belongs in a data-measured probe rather than a version literal.
One change, and it moves the identity of every node: the application segment of a can:// id is
now outermost, so a single can://<app>/ prefix scopes a whole application across every language
it is written in. That is what a polyglot repository needs in order to be one thing rather than
several, and it is what the previous grammar could not express.
It is breaking for anyone holding an emitted graph or a cached analysis. All three analyzers moved in step, the SDK pins them exactly, and each graph floor refuses the release immediately below it by name — because those releases attach cleanly and then answer nothing.
The can:// id grammar puts the application outermost, and the SDK now reads and writes only
that form, on all three backends:
before can://<lang>/<app>/<file>/<type>/<callable-signature>
after can://<app>/<lang>/<file>/<type>/<callable-signature>
Two shared namespaces also changed shape, in every analyzer:
- externals are language-neutral —
can://<app>/@external/<module>/<name>, no language segment, so sibling analyzers over the same repository name a library symbol identically; - artifacts sit inside the application prefix —
can://<app>/artifact/<path>, so a destructive push reclaims them instead of letting them accumulate.
Analyzer pins and graph floors moved together, and the floors are exact:
| backend | pin | Neo4j graph floor | refuses, by name |
|---|---|---|---|
| Python | codeanalyzer-python==1.5.0 |
1.5.0 | 1.4.1 |
| TypeScript | codeanalyzer-typescript==1.5.2 |
1.5.2 | 1.5.0 and 1.5.1 |
| Java | codeanalyzer-java==3.1.1 |
3.1.1 | 3.1.0 |
The floors are exactly the flip release, not "that release or newer with the older one tolerated". The releases immediately below each one are in the wild carrying the old grammar; each carries every relationship type the vocabulary probe looks for, so it attaches cleanly and then answers every prefix-scoped statement with zero rows — a silent empty that reads as "this codebase has nothing". Refusing on the version stamp is the only thing between a caller and that.
Java's floor moves furthest — 3.0.1 to 3.1.1 — because the intervening releases were already degraded for other reasons: a 3.0.x graph carries no config-read edges, no entrypoint report, and before 3.0.3 a port lattice joined to nothing. It was readable but answered three separate refusals; now it is one clear “re-emit” at attach.
TypeScript's floor is 1.5.2 rather than 1.5.1, the release that actually flipped the grammar,
because 1.5.1 shipped with its internal version constant left at 1.5.0 and therefore stamps itself
1.5.0 on the graph. There is no version test that admits a 1.5.1 graph and refuses a genuine
old-grammar 1.5.0 one, so the floor sits at the first release whose stamp tells the truth.
Migration. Re-emit every attached graph with the pinned analyzer, and wipe the database first: old and new ids do not collide, so a re-push onto an existing graph leaves a disconnected old copy that the new application-prefix delete cannot reach.
- TypeScript's Neo4j scoping is one prefix, not two. With the language outermost the analyzer's
two namespaces were two disjoint top-level prefixes, and every scoped statement spelled them as an
OR; with the application outermost they are two children ofcan://<app>/. The dual-namespace branching is removed rather than left dead — it is what produced a scoped lookup returning 40 rows where 113 was right, and what left the now language-neutral@externalghosts outside both prefixes. The two language prefixes survive in exactly one place,TSNeo4jBackend._module_key, which has to strip the language segment to reach the file key beneath it; nothing is scoped on them. - The application root is matched by its
can://<app>id on all three backends (:PyApplication,:Application,:JApplication). All three analyzers moved the root's merge key toid;namesurvives as a display property and carries no uniqueness constraint any more. - The Python backend separates the application scope (
can://<app>/, which every statement carries — it has to admit the language-neutral@externalghosts) from the code prefix (can://<app>/python/), which only_module_keyuses to recover a repo-relative module key. The Java backend makes the same split, for the same reason.
TypeScript and Java now answer the same query surface Python has: addressing, per-callable control
and data flow, slices, call-graph and value-flow predicates, entrypoints, and the repository-artifact
layer — identically whether you read an analysis.json or an attached Neo4j graph, including on the
paths where a backend has to refuse.
Four legs of work: TypeScript on schema v2 and its query surface, Java on schema v2 (driven by
the analyzer wheel, with degraded runs reported rather than silent) and its query surface. Design
records: docs/design/specs/2026-09-06-leg-2.5-typescript.md and
docs/design/specs/2026-09-06-leg-3-java.md.
Both analyzers moved with it. codeanalyzer-java 3.1.0 connects the level-4 port lattice, so Java's interprocedural value questions answer for the first time, and adds config-read provenance and the entrypoint report. codeanalyzer-typescript 1.5.0 corrects source spans to UTF-8 byte offsets and gives declaration-merged names one id per facet, so a class and an interface of the same name are both reachable instead of one shadowing the other.
Where a backend cannot answer, it says so and says why. Every such refusal is measured from the data in front of it, never from an analyzer version string, so a re-emitted graph starts answering on its own. What remains unanswerable is listed under Known limitations rather than returned as an empty result.
- Java:
get_call_graph()nodes are"<type fqn>.<signature>"strings, not(signature, klass)tuples. Take the owning class fromcg.nodes[key]["method_detail"].klassrather than splitting the key. - Java: single-file
source_codemode is gone. Analyze the project directory instead. - Java: a local or anonymous class's qualified name carries its declaring callable (
p.Outer.m(int).$anon$0). Without it, sibling callables collide. Take class keys fromget_all_classes()rather than composing them. - A graph emitted below the analyzer floor is refused at attach with
GraphSchemaMismatchinstead of answering every query with zero rows: codeanalyzer-java 3.0.1 and codeanalyzer-typescript 1.3.0. Re-emit; there is no in-place upgrade.TSNeo4jBackendalso speaks a graph vocabulary that shares nothing with 0.4.3's. - Values that changed shape:
JGraphEdgesandTSCallEdgeare now{src, dst, prov, weight}; Java'scalling_linesis a sorted list of absolute file lines; Java'sget_config_keys()is keyed by the artifact-relative key; Java'sget_test_methods()reads the analyzer's annotations rather than re-parsing source;TSCallablelostpath/call_sites/accessed_symbols/local_variables/code_start_lineandTSModulelostfile_path/module_name, since v2 keys modules by path and stores source once;TSCallableOverview.from_callabletakes a required keyword-onlypath. - Removed:
TypeScriptAnalysis.get_entry_point_methodsandget_service_entry_point_methods, which only ever raised — the working entrypoint accessors below replace them. - TypeScript
TSSpan.bytesare UTF-8 byte offsets, on every node at every level (codeanalyzer-typescript 1.5.0, cants#179). They were UTF-16 code units, which Python sliced as code points and the Neo4j projection sliced as bytes — three units that agreed only on ASCII.TSCallable.codeand every accessor built on it (get_source,get_method_bodies,describe,locate(...).source) now decode the module's UTF-8 bytes, the way Java'sJCompilationUnit.slice()already did: encoded once per module, a plain index when the file is ASCII. On a non-ASCII file, text read through 1.3.0/1.4.0 was wrong and is now right — re-emit rather than compare offsets across versions. Note thatschema_versionstays2.0.0through this break: the contract version is not a signal that a breaking change landed.
- The agent-facing query surface on
TypeScriptAnalysis— 29 accessors, each withPythonAnalysis's signature and semantics. Addressing (locate,locate_many,resolve_callable,resolve_value,get_source,describe,has_resolution_edges); per-callable graphs and dataflow (get_cfg/get_cdg/get_ddg,slice_backward/slice_forward/backward_cone,reaches,callers_of/callees_of,paths_between/call_paths_between,flows_to_call/flows_to_argument); entrypoints (get_entrypoints,get_entrypoint_classes,get_entrypoint_coverage,get_config_readers); and the five repository-artifact getters. Both backends answer identically, including on the miss paths. - The agent-facing query surface on
JavaAnalysis— 38 accessors, each withPythonAnalysis's signature and semantics. Addressing (locate,locate_many,resolve_callable,resolve_value,get_source,describe,has_resolution_edges); per-callable graphs and dataflow (get_cfg/get_cdg/get_ddg,slice_backward/slice_forward/backward_cone,reaches,callers_of/callees_of,paths_between/call_paths_between,flows_to_call/flows_to_argument); entrypoints and the bulk projections (get_entrypoints,get_entrypoint_classes,get_entrypoint_coverage,get_callables_overview,get_method_bodies,get_decorated_callables,get_callsites_for,get_external_symbols); the six repository-artifact getters; and the type-kind leaf accessorsget_interfaces/get_enums/get_enum_members/get_records. Both backends answer identically, including on the miss paths. cldk.models.java.JCallableOverviewandJClassOverview, the projections those bulk accessors return. They carry the addressable"<type fqn>.<signature>"key, never acan://id.- Six Java accessors that only ever raised
NotImplementedErrornow answer (#366), at the signature they have always been published with:get_imports()(the project's distinct import targets, sorted — the Neo4j projection aggregates a module's imports per target, so file order is not recoverable and the set is what both backends can give),get_variables()(each callable's local variables, keyed by the"<type fqn>.<signature>"call-graph key and ordered by(line, name); fields and parameters keep their own accessors, and an unexpected keyword now raisesTypeErrorinstead of being ignored),get_class_hierarchy()(anx.DiGraph, subclass → supertype, each edge carryingtype="EXTENDS"or"IMPLEMENTS"— read off each declaration's ownbase_types/interfaces, which is why out-of-project supertypes are in it),get_methods_with_annotations()(grouped by the spelling the caller passed, matched by the J-5 marker rule, each entry{class, signature, method_name, body}),get_call_targets()(the declared names some call site actually writes — simple-name matching, no overload resolution) andget_calling_lines()(sorted, distinct absolute file lines, off the call graph's owncalling_lines). Both backends answer identically; none of the six issues any new Cypher.get_service_entry_point_classes/get_service_entry_point_methodsandremove_all_commentsstill raise. - Java reaches analysis levels 3 and 4 — control flow, control and data dependence, and the interprocedural graph. The level now reaches the analyzer, which it never did before.
- A
javainstall extra.pip install "cldk[java]"brings the analyzer and its bundled JVM;pip install "cldk[all]"reproduces the previous behaviour. Barecldkno longer carries either. The Python and TypeScript analyzers are still installed unconditionally; #340 moves them into extras of their own. - JavaScript modules are in scope for TypeScript analysis, under their own id prefix.
- A degraded Java level-3 or level-4 run is reported rather than silent. Those levels need compiled classes; without them the analyzer emits a declared-only graph and still reports the level. The SDK now surfaces the analyzer's own warnings and records them beside the payload, so a cached run still knows. The graph is returned either way.
cldk.models.typescript.TSClassOverview, the class-level projectionget_entrypoint_classesreturns.
- Java's four interprocedural value accessors now answer.
slice_forward,paths_between,flows_to_callandflows_to_argumentrefused on Java because codeanalyzer-java emitted the level-4 port lattice disconnected from the statement dependence graph — aformal_inhad out-degree zero, so every forward answer was empty whatever the program did. codeanalyzer-java 3.0.3 joins the two layers (codeanalyzer-java#227), and the SDK's guard reads the data rather than a version, so all four answer: a value can be followed out of a parameter, across call boundaries and into a callee's parameter, with each hop labelleddata/argument/return/control. Measured on the whole of daytrader8: 5,083ddgedges cross between the port lattice and the statement graph in all four directions, where there were none. To get this, re-analyse (or re-emit your Neo4j graph) with the pinned analyzer; an older graph stays attachable and keeps refusing, as does any analysis run under--l3-engine wala. The same release dropsddgedges whose endpoint was never emitted as a body node (codeanalyzer-java#228), soget_ddg()overanalysis.jsonand over Neo4j now report the same edges — 10,430 on daytrader8, set for set, against a 5,434/5,347 split before. - Java's code-to-config layer and entrypoint report now answer. codeanalyzer-java 3.1.0 adds both, so
get_config_uses(),get_config_readers(key)andget_unresolved_config_reads()return real edges instead of[], andget_entrypoint_coverage()returns the pass's own record instead ofentrypoint_report_unavailable. The two config tiers are surfaced, not flattened:prov == ["literal"]is a string literal at the call site (codeanalyzer-java#233),"dataflow"is a key reached over the L3 DDG or the L4 call graph (#237) — a derived answer, weaker evidence, and it widens with the analysis level. On aPyConfigReadthe list is every tier attempted, so["literal", "dataflow"]means the dataflow tier ran and still could not name the key. Measured on daytrader8: 13 resolved uses (all["literal"]) and 16 unresolved reads —["literal"]at level 1,["literal", "dataflow"]at level 4; the graph collapses those 16 into 8 edges, sinceJ_READS_CONFIG_UNRESOLVEDis discriminated by(key, reason)and carries no site. To get this, re-analyse (or re-emit your Neo4j graph) with the pinned analyzer; an older analysis keeps refusing, and the probe is measured from the data rather than from a version string — see the known limitation below for what it measures and why it cannot be the config layer's own absence. - Pins:
codeanalyzer-java2.4.1 → 3.1.0,codeanalyzer-typescript0.4.3 → 1.5.0. - The Java analyzer ships as a wheel, not a jar in this repo. The 35 MB checked-in jar, the Temurin download
in
_jdk.py, and the release workflow's jar injection are gone; noJAVA_HOMEis read or set, and no JDK is downloaded. The published wheel drops from about 35 MB to 320 KB. - Both languages' backends inherit the generic
AnalysisBackend, so each answers the shared artifact, dependency and configuration accessors. - Where a backend cannot answer, it says so instead of returning an empty value. On Neo4j: Java's file-keyed
comment accessors, and TypeScript's
get_extended_classes/get_implemented_interfaceswhen the relationship type is absent — and itsget_imports,get_exports,get_method_parametersandget_unresolved_config_readson a graph emitted before codeanalyzer-typescript 1.4.0, which carries none of the data they read (below). The remaining documented gaps, and where the two backends legitimately differ, are listed indocs/agent-api-reference.md. - TypeScript's
get_imports,get_exports,get_method_parametersandget_unresolved_config_readsanswer over Neo4j, on a graph emitted by codeanalyzer-typescript 1.4.0 or newer: the projection gainedTS_IMPORTS/TS_RE_EXPORTSedges,:TSModule.exports_json,:TSCallable.parameters_jsonandTS_READS_CONFIG_UNRESOLVED(codeanalyzer-typescript#182). Parameters also reachTSCallable.parametersand exports reachTSModule.exportson a rebuilt node. A 1.3.0 graph is still attachable and those four still refuse there rather than answer an empty that would read as a fact; the decision is measured from the application's own data — is any carrier present — never from the analyzer's version string, so a re-emitted graph starts answering with no SDK change.TSImport/TSExportgained the analyzer's newresolved_module. Migration: re-emit withcodeanalyzer-typescript>=1.5.0 --emit neo4j. - Three more codeanalyzer-typescript ceilings are gone in 1.5.0, verified against the re-emitted reference
graph rather than taken from the release notes.
:TSCallable.codeis no longer one line short over Neo4j (cants#179) — the fourxfail(strict=True)marks that pinned it now XPASS and are removed. Declaration merging mints one id per facet (cants#177), so aclass X+interface Xpair is two nodes and both are served, where before one facet was lost to whichever declaration the emitter saw last; on superset-frontend that is 7 signatures, and no node carries two declaration labels any more. The analyzer also refuses a missing or non-directory input instead of exiting 0 (cants#181) and no longer clones the whole envelope per module (cants#180); neither had an SDK workaround to remove. - Every Cypher statement is scoped per bound variable, not per statement — both endpoints of a call edge, a
quantified path's far end, a slice's reached body nodes, and every interior node of a variable-length or
shortest-path walk (
all(n IN nodes(p) …)on the slice, path and reachability queries). A statement that matched one endpoint by a signature two applications both declare, or that walked through an intermediate it never predicated, could previously return the other application's node. LocateResult.bodyis a language-neutralBodyRef, socldk/analysis/commons/no longer imports a language package for it.- An
AmbiguousNamenow names only the ways out that are actually open — on all three languages. TheNarrow it with …sentence offersin_class=/in_module=only when the caller has not already passed that keyword and the listed matches disagree on it, so two overloads of one class, or two__init__s of one class, are no longer told to narrow by a keyword that provably cannot split them. The last clause is always offered and is language-specific:more of the dotted pathon Python and TypeScript, and on Javathe full signature, exactly as one of the listed matches spells it(more of the dotted path cannot separate two overloads).AmbiguousName.candidatesis unchanged. Python'sresolve_callable,callers_of,callees_ofandresolve_withintherefore produce a different message for the same input than in rc.2; assert on.candidates, not on the sentence. - Internal: the language-neutral query helpers moved from
cldk/analysis/python/tocldk/analysis/commons/; Python re-imports every name unchanged.
-
get_cfg/get_cdg/get_ddgon the Python Neo4j backend no longer drop self-loop edges. The per-callable query bound the containment relationship twice, so Cypher's relationship-uniqueness rule silently discarded every edge whose endpoints are the same body node — a statement that reads a variable it also redefines. 64,702 such edges exist on the odoo reference graph; one callable returned 28,146 of its 28,394.totalwas counted from the same match, so the short page reported itself complete. (#349) -
JavaAnalysis.get_method_parameters()is annotatedList[JCallableParameter], which is what it has always returned. -
get_call_graph()on Java no longer re-parses each method body once per edge — 145.7s to 41.0s on a 4,100-file project. -
Two documentation errors:
EntrypointCoverage.unresolvedis adict[str, int], not alist[str]; and theget_config_keys()example used a key form that never worked. -
get_source()on the in-process TypeScript backend named the application by itscan://id where the Neo4j backend named it plainly; both now name the application.get_external_symbols()andget_source()over Neo4j scoped externals by thecan://typescript/prefix alone, dropping any external a JavaScript module owns; both now carry the two-prefix scope.
- Java CRUD accessors raise: schema v2 does not carry CRUD yet (codeanalyzer-java#187).
- The Java graph's
JCallable.codeis the declaration slice where the JSON's is the body block, and the graph cannot recover the body block (codeanalyzer-java#176). - Import bindings read over the TypeScript Neo4j backend are narrower than in process.
TS_IMPORTSfolds every binding between a module pair into one edge of sorted sets, so an entry's alias,import_kindand span are not recoverable, and the emitter drops a relative specifier that resolved to no emitted module. Exports, parameters and unresolved config reads are lossless, except that the config-read edge carries nositeand collapses sites sharing a(callee, key, reason)triple. - TypeScript's DDG has one provenance tier; Java's has two; Python's has three.
get_config_keys()is still keyed by acan://id on Python and TypeScript, where Java now uses the artifact-relative key (#346).- Java's
slice_forward,paths_between,flows_to_callandflows_to_argumentstill raise on an analysis emitted before codeanalyzer-java 3.0.3, or by any version under--l3-engine wala: no dependence edge leaves aformal_inthere, so every forward answer would be empty whatever the program does. A Neo4j graph emitted by 3.0.1 or 3.0.2 is still attachable (the floor is 3.0.1) and still refuses — re-emit it. - Java's
flows_to_argumentis not per-argument precise. codeanalyzer-java feeds every actual of a call site from the one statement containing it, not from the reaching definition of that argument, so on a reached call site every argument answersTruetogether. Paths are complete; per-argument precision is not what the analyzer promises. - Java's entrypoint report and code-to-config layer need codeanalyzer-java 3.1.0. On an older analysis —
a cached
analysis.jsonis reused whatever wrote it, and a 3.0.x Neo4j graph is still attachable, the floor being 3.0.1 —get_entrypoint_coverage()returnsentrypoint_report_unavailablerather than fabricated coverage, andget_config_uses()/get_config_readers()/get_unresolved_config_reads()raise rather than answer[]. The probe is one fact measured from the data: 3.1.0 writes an entrypoint report on every run and 3.0.x writes none of the three overlays. It deliberately is not the config layer's own absence, which is ambiguous — the analyzer writesconfig_uses/config_reads_unresolvedonly when non-empty, and the graph declaresJ_USES_CONFIG/J_READS_CONFIG_UNRESOLVEDonly once an edge exists, so a clean 3.1.0 analysis of a project that reads no configuration carries neither and must still answer.get_entrypoints()andget_entrypoint_classes()carry the real marks at every generation, as doesget_config_keys(). - Java's
get_external_symbols()answers over Neo4j and raises locally: the analyzer homes out-of-project call targets only under--external-calls, which--emit neo4jforces and a local run does not. - The Java graph carries a
switchbody-node kind, which is outsideSliceNode.KINDS(that vocabulary is codeanalyzer-python's, and Python has no switch statement). It is reported as the analyzer spells it. - Java's
get_ddg()over Neo4j carries 276points-toedges (on daytrader8) that the same analyzer'sanalysis.jsondoes not, and that is a difference of what each source was asked. codeanalyzer-java 3.1.0 makes the level-4points-tolayer depend on--external-calls, which--emit neo4jforces on and which the SDK's local run does not pass; 3.0.3 produced the same 10,430 edges either way. Measured on daytrader8, same tree, four runs: 3.0.3-a 4→ 10,430 (1,134points-to); 3.1.0-a 4→ 10,154 (858); 3.1.0-a 4 --external-calls→ 10,430, set for set identical to the graph; 3.1.0--emit neo4j→ 10,430. The graph is a strict superset and the payload has nothing the graph lacks, soslice_forwardand the other forward walks can reach further over Neo4j. Reported upstream: the flag is documented as controlling only whether out-of-project call targets are homed asexternal_symbols. - The Java Neo4j projection still carries no comment nodes. codeanalyzer-java 3.1.0 closes
codeanalyzer-java#231's config-read and entrypoint-report halves and not its comment half:
:JCommentis a declared label with zero nodes on a graph emitted by it, and only a declaration'sdocstringreaches the graph. Soremove_all_commentskeeps raising, and the file-keyed comment accessors keep refusing over Neo4j while the javadoc-only ones narrow, exactly as before.
Python legs 1, 1.5 and 1.6 of the CLDK 2.0 agent-facing query facade (see
docs/design/specs/2026-09-03-agent-facing-query-facade.md,
docs/design/specs/2026-09-05-leg-1.5-bounded-queries-and-dataflow.md and
docs/design/specs/2026-09-06-leg-1.6-id-prefix-scoping.md). Targets 2.0.0-rc.2.
-
Pinned
codeanalyzer-python1.4.0 → 1.4.1 (pyproject.tomldependenciesand[tool.backend-versions]). 1.4.1 removes the_modulenode property from every graph it emits (upstream #183), so the Neo4j backend no longer scopes on it: every statement is scoped to the application by itscan://id prefix (n.id STARTS WITH 'can://python/<app>/'), the rule the analyzer's own destructive statements use, and a callable's repo-relativepathis derived from its id and verified against the application's module keys rather than projected from a property. No accessor changes name, signature, return type or value; the_modulescoping is simply gone. Ghosts (:PyExternal) fall inside the prefix, so every call-graph walk pins its traversal source to:PyCallableby label -- an@externalnode is still reached and never traversed through. -
The schema probe reads the graph's analyzer generation. Attaching to a graph whose
:PyApplication.analyzer_versionis below 1.4.0 (or missing) raisesGraphSchemaMismatchnaming the version found and the floor -- before this it was served with silent empties. A 1.4.0 graph and a 1.4.1 graph are served identically and silently: both carry the unique:PySymbol(id)range index, solocate/locate_manyandresolve_callableanchor on(c:PyCallable:PySymbol)and the prefix seeks on either generation (40 positions: 381 → 46 ms on 1.4.1, 427 → 53 ms on 1.4.0;resolve_callable19.2 → 15.3 ms on 1.4.1, a wash on 1.4.0). 1.4.1's:PyCanNode(id)index was measured and rejected -- it spans all 955,961 application nodes, so seeking it maderesolve_callable10× slower -- and no statement names it. Statements whose prefix is the whole application stay on the:PyCallablelabel scan, which is faster there. -
BREAKING: Neo4j graph vocabulary migration. The Python Neo4j backend (
cldk.analysis.python.neo4j.PyNeo4jBackend) now queriesPyBodyNode/PY_HAS_BODY_NODEinstead of the pre-1.4.0PyCallSite/PY_HAS_CALLSITE/PySymbolvocabulary, matching whatcodeanalyzer-python>=1.4.0 actually emits. Signature-level backwards compatibility does not extend to graph generation. A graph built bycodeanalyzer-python0.3.x that used to answer every query with zero rows (silently, indistinguishable from "this codebase has no callables") now raisesGraphSchemaMismatchat attach time instead, naming the expected/found/missing relationship types and the analyzer generation implied. Migration: re-ingest the project withcodeanalyzer-python>=1.4.0(codeanalyzer-python --emit neo4j) before attaching this SDK version to it; there is no in-place graph upgrade. -
Pinned
codeanalyzer-python0.3.1 → 1.4.0 (pyproject.toml), the source of the vocabulary migration above. The in-process (local) backend picks this up automatically; only the Neo4j backend needs a re-ingested graph. -
Neo4j enumeration is no longer an N+1 fan-out.
get_symbol_table(),get_classes()and everything built on them (get_modules(),get_application_view(),get_call_graph_json()) used to rebuild each declaration with one query per child collection per parent node — 73,669 round trips to describe a 1,626-module application, at ~5 ms each. Each collection is now fetched once for the scope being read and served from a by-parent index, so a bulk accessor pays at most eleven round trips however many modules, classes and callables it walks. Measured on a 970k-node graph (Odoo, 2,364 files):get_classes()348s → 10.4s,get_symbol_table()382s → 11.9s,get_call_graph_json()422s → 26.1s. No signature, return type or result changes. -
BREAKING (value, not signature):
PyCallableOverview.pathis now repo-relative. It used to be the absolute path on whichever machine ran the analysis (/Users/someone/checkout/addons/account/models/onboarding.py); it is now the repo-relative module key (addons/account/models/onboarding.py) — the same string aslocate().module.path,PyClassOverview.pathand the keys ofget_symbol_table(), which it previously could not be joined against. Both backends changed together: the Neo4j projections behindget_callables_overview(),get_decorated_callables(),get_entrypoints()andget_config_readers()now project the module key rather than:PyCallable.path(from the_moduleproperty on a 1.4.0 graph in leg 1; derived from the node'scan://id since leg 1.6), and the local backend projects the symbol table key rather thanPyCallable.path. Migration: a caller that stored or persisted these paths will see different strings for the same callable, and one that stripped a project-root prefix off them must stop. -
BREAKING (value, not signature):
LocateResult.node_idis now the analyzer's own body-node id. It used to be composed by the SDK as<signature>@<body key>(addons.onboarding.models.onboarding_onboarding_step.OnboardingOnboardingStep._compute_current_progress@51:34), a string that named nothing in the graph; it is now the node's own id (can://python/<app>/…/OnboardingOnboardingStep/_compute_current_progress(self)@51:34) — read off:PyBodyNode.idover Neo4j, and composed fromPyCallable.idby the emitter's own rule locally, so both backends produce the same string and it joins to the graph (verified over all 885,218 body nodes). It is the same vocabularySliceNode.refuses. Migration: treat it as an opaque handle to pass back toget_source()/describe(), never as a string to parse or build;get_source()still accepts the old callable-half spelling, so a stored id keeps resolving. (#320) -
One completeness protocol on every bounded result.
EdgePage,SliceandFlowPathsall exposecomplete: bool, and it is serialised —model_dump()carries it, so a result written out and read back by another process still says whether it was whole. It replaces two names for one fact:EdgePage.has_moreandSlice.truncated(both unreleased) are gone, not aliased. All three also behave as the list they carry:for edge in page,len(slice),paths[0], andif paths:isFalsewhen empty — pydantic's defaults hadfor x in pageyielding(field, value)tuples andbool(page)alwaysTrue.FlowPathsis now a model (paths,complete) rather than alistsubclass whosetruncatedattributejson.dumpssilently dropped.FlowPath.weakeststays an unserialised property;prov_rankover the dumped hops recovers it. A single probe table (tests/analysis/python/test_completeness_protocol.py) runs over all three so the protocol cannot drift again. -
The predicates and path queries are unbounded by default; only the slices bound themselves.
paths_between,call_paths_between,flows_to_callandflows_to_argumenthad inherited the slices'depth=5, and on a boolean or a path list a hop budget is not a smaller answer but a wrong one:flows_to_call("kwargs", "Website.create", within="Website.configurator_apply")answeredFalsefor a flow that exists, and the matchingpaths_betweenanswered[]with nothing to say a bound had fired. All four now default todepth=None, asreachesalready did, and the rule is stated once onDEFAULT_DEPTHand cited by all five;depth=remains an explicit narrowing. The three slices keepdepth=5, because a bounded slice is a complete answer to a narrower question andtotalsays so. -
paths_betweentakessrc_within=anddst_within=, both required.within=is renamed anddst_withinno longer defaults to it: two values of one callable are joined only through recursion, so the old default made the default call the degenerate case.flows_to_call/flows_to_argumentkeep a singlewithin=deliberately — it scopessrc, and their second endpoint is a callable addressed by name, withargscoped bycalleeitself. -
in_module=also takes the dotted module name ("odoo.tools.mail", or a dotted suffix such as"tools.mail"), alongside the repo-relative path. That is the vocabulary every signature reads in, andin_class=was already a dotted suffix. A resolution that fails because a keyword excluded every match now raisesSelectorNotInGraphnaming that keyword (kind="in_module") instead of claiming the callable does not exist. -
reachesandbackward_coneneed Neo4j 5.9+ on the Neo4j backend, asget_call_graph(roots=)already did: their walks are now quantified path patterns constrained at every hop (see the parity fix below). Every other accessor still runs on any 5.x; the version is checked per accessor, so an older server keeps serving the rest.
| SDK version | Requires codeanalyzer-python |
Graph vocabulary |
|---|---|---|
| <= 1.5.0 | 0.3.x | PyCallSite / PY_HAS_CALLSITE / PySymbol |
| 2.0.0-rc.2 (this) | 1.4.1 pinned; graphs emitted by >= 1.4.0 served identically | PyBodyNode / PY_HAS_BODY_NODE, scope by can:// id prefix |
-
Slices and reachability:
slice_backward(src, within=)/slice_forward(src, within=)/backward_cone(sinks)/reaches(src, dst)/callers_of(name)/callees_of(name), on both backends. Names in, names out — a value is addressed as"invoice_id"scoped by its callable, never as"…@formal_in:1", and nocan://URI appears in an argument or a result. The traversal follows data dependence, control dependence, argument passing, returns and call summaries at once (PY_DDG/PY_CDG/PY_PARAM_IN/PY_PARAM_OUT/PY_SUMMARY— 6,089,420 edges on the measured application) and runs in the database: on the Neo4j backend each slice is one variable-length Cypher match, planned as a pruning breadth-first search, so the largest cone measured — 195,784 nodes — is counted and its first 10,000 nodes described in about 1.5 seconds without the SDK ever holding an edge. The local backend answers the same questions interprocedurally atanalysis_level="system_dependency_graph", building its index fromPyApplication.param_in/param_outandPyCallable.summary, which are the same edges the graph is projected from. Belowprogram_dependency_graphthe slices raise through the guardget_ddgalready uses; the call-graph accessors (reaches,callers_of,callees_of,backward_cone) work fromcall_graphand are not guarded, because the call graph exists there.A slice is capped, not paged — and the cap reports what it dropped (
Slice.total,Slice.truncated,Slice.nodes,Slice.roots,Slice.resolved). This is deliberately the opposite ruling toEdgePageon the sibling accessors above, and the measurement is why. On the same application a backward slice is either about 1 node or about 195,786 (the median over seeds that have callers; p95 195,790, max 196,117 over 200 random seeds), and a forward slice reaches 440,270 at the 95th percentile — of 885,218 body nodes in the whole program. The distribution has no middle, so a page is not a useful unit; and a slice is the traversal, so unlike a keyset cursor over an order the database already maintains, every page would re-run the whole closure.totalcarries E5's "a bound is never silent" instead: the whole slice's size arrives with the first and only call, andtruncatedis derived from it rather than stored, so the two cannot disagree. When a cap fires the way forward isdepth=, which bounds the traversal and so returns a complete answer to a narrower question. Nodes come back ordered byrefand the cap takes a prefix of that order, so the same call twice returns the same subset.Slice.rootis the singular seed for the single-seed accessors and raises on a multi-sink cone rather than silently answering with the first.depthdefaults to 5 onslice_backward,slice_forwardandbackward_cone— a finite bound, notNone. Because the distribution has no middle, an unbounded default gave a connected seed 10,000 arbitrary nodes of a 195,819-node closure: honestly flaggedtruncated, and still an unprincipled 5% of a cone. A hop bound answers a narrower question completely instead, and the number is measured, not chosen: over 120 random connectedformal_inseeds, the backward slice at depth 5 has median 33 / p75 188 / max 1,539 and the forward slice median 24 / p75 63 / max 1,053, with not one seed in either direction reachingmax_nodes; at depth 6 a forward slice first exceeds it (14,260), and at depth 8 three do. So a slice is now normally one wheretruncatedisFalseandtotalis the size ofnodes.depth=Nonestill means the whole closure and is how a caller asks for the fifth of the program — every capability Task 6 shipped remains reachable, one keyword away.reacheskeeps its unbounded default deliberately: a hop budget on a boolean turns "no path" and "no path within 5 hops" into the sameFalse, which is a wrong answer rather than a small one, and the unbounded call measures 20ms mean / 112ms worst over 200 random pairs. The same change made a latent local-backend defect load-bearing and it is fixed here:backward_cone(sinks, depth=n)walked forward over the call graph (nx.ego_graphfollows successors, and was not given the reversed view), so a bounded cone returned the sink's descendants; only the unboundednx.ancestorspath was correct.callees_ofincludes calls out of the project — 38,585 of the application's 370,110 call edges, and usually the ones a caller tracing a sink is after. An external comes backkind="external"with a readable dotted name built from the node's ownmodule/name("odoo.exceptions.ValidationError.__init__", never itscan://id) and withfile=""/line=0, because it was never analysed and there is no position to point at.callers_ofis declared callables only. Neither touches the frozenget_all_callers/get_all_callees. -
Paths, mixed flow queries and hydration:
paths_between(src, dst, src_within=, dst_within=)/call_paths_between(src, dst)/flows_to_call(src, callee, within=)/flows_to_argument(src, callee, arg, within=)/describe(nodes), on both backends. A path is a sequence where a slice is a set (E2):paths_betweenreturns the shortest flows from one value to another asFlowPaths, each an ordered list ofPathHops that say what justified the step —viain the caller's words (data/control/argument/return/summary, neverPY_DDG/PY_PARAM_IN), the variable it flows on, and the provenance it was established with — so a caller can argue a flow rather than assert one.FlowPath.weakestnames the hop that caps the claim, ranked by the newPROV_CERTAINTY/prov_rank:ssa(exact def-use) >reaching-defs(a may-analysis over the CFG) >points-to(alias analysis), least certain first; an unlabelled structural hop claims no approximation and ranks strongest; an unrecognised label ranks weakest. Measured:provis a singleton on every one of the 5,134,655PY_DDGedges, so "the weakest hop" is well defined. Only shortest paths are returned — enumerating every walk does not terminate on a real dependence graph — and at mostmax_pathsof them (DEFAULT_MAX_PATHS = 10, witnesses rather than extent; the extent question isslice_forward), in a stated order (hop_sort_key: length, then(via, var, to.ref)hop by hop) identical on both backends, so a cap takes a prefix rather than whichever paths arrived first. Both areFlowPaths, which says whether the cap fired. Over Neo4j the query isallShortestPaths— a bidirectional BFS that answers the pathological seeds in under 0.1s where a variable-length match never finished; locally it is the same shortest-walk enumeration over the SDG index the slices already build. A path from a node to itself is refused (ValueError, in the caller's vocabulary, advising thereachescall that answers the cycle question) rather than answered[], which a node genuinely on a cycle could not be told apart from.call_paths_betweenis the same shape over the call graph — every hopvia="call"with no variable and no provenance, because a call edge carries neither — and the evidence-carrying form ofreaches.flows_to_callasks whether a value reaches any value entering a call to the callee;flows_to_argumentasks whether it reaches the callee's parameter namedarg(E7, no ordinals). They are different questions with different answers — measured,invoice_idofPaymentPortal.invoice_transactionreaches six of_process_transaction's seven entering values and notkwargs— andflows_to_argument ⟹ flows_to_callholds by construction because both run one reachability predicate over a subset and its superset. Anargnaming no value of the callee raises rather than answeringFalse.describe(nodes)fills inSliceNode.sourcefor the positions a caller chose to read, in one round trip whatever the count, and accepts anything carrying a ref — slice nodes, path-hop endpoints,locate()results (E4: references by default, payloads on request). Afterwardssource=Nonemeans exactly one thing: the position exists and this backend has no text for it — a value vertex on either backend, a statement over Neo4j (the graph carries no text below callable granularity; the local backend slices it out of the module), or an external (never analysed, so no text anywhere). A ref that names nothing raisesKeyError, reported by callable andfile:line, never by ref. -
Per-callable control and data flow:
get_cfg(callable)/get_cdg(callable)/get_ddg(callable), on both backends.DdgEdgecarries the variable that flows (var) and the evidence for it (prov—ssa,reaching-defsorpoints-to), so syntactic and alias-aware dependence are told apart without a second call. Endpoints are body-node ids in the same vocabularyLocateResult.node_iduses andget_source()accepts. Requiresanalysis_level="program_dependency_graph"or deeper: a local backend built shallower raisesCodeanalyzerUsageExceptionnaming the level required and the level in use, rather than returning an empty result that could not be told apart from a callable with no dependence.All three return an
EdgePage, not a list (cldk.analysis.commons.results.EdgePage). Per-callable is a scoping bound, not a size one: measured on a 970k-node graph, one callable (Website.configurator_apply) has 1,386,918 DDG edges — 27% of the whole application's 5,134,655 — which as a list is around half a gigabyte of models in one call. The page bounds the response and discards nothing:page.totalis the size of the whole answer,page.edgesis at mostpage_sizeof them (default 10,000, chosen because 15,520 of the application's 15,549 callables fall under it and so need no loop at all), andpage.next_cursor— opaque, passed back ascursor=— reaches the rest.page.has_moredistinguishes "this is everything" from "there is more" without a second call. Edges come in one canonical order, identical on both backends (source, target, thenkindfor CFG andvar/provfor DDG), so page n is the same page whichever backend answered. Paging is by cursor rather than offset because offsets get slower the deeper you read: for the same query on that callable,SKIPcosts 2.6s / 9.0s / 4.3s for the first / middle / last page while the cursor filter is flat at 3.1s / 2.9s / 2.4s (end to end through the SDK, including name resolution and the count, the cursor form measures 3.5s / 3.6s / 3.0s, median of three runs). CFG and CDG page identically despite topping out at 402 and 314 edges — three sibling accessors answering in two shapes is a defect generator on a surface composed at runtime. -
locate(path, line)/locate_many(positions)— resolve a source position to its enclosing callable (ormodule_scope/file_not_in_graphwhen there isn't one), with the callable's source text attached. Backed by the newLocateResultmodel (cldk.analysis.commons.results). -
get_source(node_id)— the source text named by a callable's signature or byLocateResult.node_id(a body-node id), byte-accurate against non-ASCII source. Below-callable granularity is local-backend-only; the Neo4j backend raisesNotImplementedErrorrather than substituting the enclosing callable's text. -
Entrypoint detection surface:
get_entrypoints()(callable-level),get_entrypoint_classes()(class-level, newPyClassOverviewmodel), andget_entrypoint_coverage()(the detection pass's own coverage/failure record via the newEntrypointCoverageresult, so an empty result can be told apart from a pass that had gaps). -
Repository-artifact / dependency / config layer:
get_artifacts(),get_dependencies(),get_config_keys(),get_config_uses(),get_unresolved_config_reads(), andget_config_readers(key). -
External call resolution:
get_external_symbols()returns every call-graph endpoint outside the analyzed project, keyed by itscan://…/@external/…id.get_callsites_for()and the call graph accessors (get_call_graph()/get_callers()/get_callees()/get_class_call_graph()) now resolve a call landing on such a target to that id instead of leavingcallee_signature/the graph node asNone(previously 602 recorded incidents for the former, 38,585 of 364,752 in-scopePY_CALLSedges on a live graph for the latter). -
has_resolution_edges(property) — whether the current backend can resolve call-site callees at all right now. AlwaysTrueon the local backend (it always attempts Jedi resolution); on the Neo4j backend it is probed once at connection time and isFalseonly when the attached graph carries noPY_RESOLVES_TOedge anywhere, which disambiguates a genuinely unresolved call site from a graph with no resolution data. -
Scoping keywords on the enumerating accessors —
get_symbol_table(paths=…),get_classes(module=…)andget_call_graph(roots=…, depth=…), on both backends. All are keyword-only and default to the whole-application behaviour they had before, so no existing call site changes.paths=/module=name modules by symbol-table key (absolute paths and native separators resolve too);roots=names callables by signature, or by@externalid for an out-of-project target.get_call_graph(roots=…)returns the induced sub-graph over the callables reached — every edge among them, not only those on a path out from a root — sopredecessors()cannot lie about a node in the graph it just returned; a root that calls nothing comes back as a graph of one node, including one that appears in no call edge at all — a root is validated against the callables the application declares (plus the graph's own nodes, which is what makes an@externalid a valid root), never against edge participation, so the two backends judge the same domain.depth=bounds the walk in call hops, requiresroots=, and must be anint>= 1;paths=/roots=take sequences and reject a bare string withTypeErrorrather than iterating its characters;paths=resolution is many-to-one and de-duplicated, so two spellings of one module are one entry. The Neo4j backend pushes all of it into Cypher, including the prefetch scope, so one module or one root costs well under a second instead of a full fetch filtered in Python; the bounded call graph is a quantified path pattern scoped to the application at every hop and therefore requires Neo4j 5.9+ — the server version is read when the backend attaches and enforced by that one accessor, so an older server keeps serving all the others. -
SelectorNotInGraph(cldk.utils.exceptions, aValueError) — a value passed topaths=,module=orroots=that names nothing in the analysed application now raises, naming the values that matched nothing and how many were asked for (1 of 2 paths not in graph: 'gone.py'), including when only some of them missed. Previously each contributed nothing, so a mistyped path and a module that genuinely declares no classes were the same{}, and an unknown root and a callable that calls nothing were the same empty graph. The message lists what missed and stops — no near-miss suggestions, per the leg's E8. An empty sequence (paths=[],roots=[]) raises plainValueErrorinstead: nothing missed, but the argument that means "everything" is the argument omitted. A root that misses says whatroots=takes — full signatures or@externalids, not bare names — and where name-based addressing lives, sinceroots=is an exact filter while every name-taking accessor suffix-matches, so a correct short name and a typo miss identically. -
AmbiguousName(cldk.utils.exceptions, aValueError) — raised when a name the caller wrote matches more than one callable or value. Carries every match as data (candidates, sorted), shows the first few plus a total in the message, and names the keyword that would narrow it — only one the caller has not already used. Nothing is picked and nothing is suggested that did not genuinely match. -
Name-based addressing:
resolve_callable(name, in_class=, in_module=)andresolve_value(name, within=)(both Python backends, and on thePythonAnalysisfacade like every other accessor) — a caller names a callable or one of the values entering it the way it already thinks of it, and gets back aSliceNode; nothing in the surface takes, returns, or requires assembling acan://URI, and nothing takes an ordinal. Names match whole or as a dotted suffix on segment boundaries, with an exact match winning outright; more than one match raisesAmbiguousNamecarrying every candidate rather than picking one, and there is no typo-tolerant matching anywhere — not in the resolver, not in the error path. Both backends route through one policy module (cldk.analysis.commons.resolve), so they cannot drift on what "ambiguous" means. Aresolve_callablerefround-trips throughget_source()on either backend. Aresolve_valuerefdoes not, on either:parameter/global/captureare dataflow vertices with no source span, so there is no text to return (the local backend raisesKeyError: … (no span), the Neo4j backendNotImplementedError). Local variables have no address throughresolve_valueat all — only values that enter a callable do;locate(path, line)addresses positions inside a body.SliceNodegainsdefined_in, the defining module of a captured global. -
GraphSchemaMismatch(cldk.utils.exceptions) — raised by the Neo4j backend's one-time schema probe at attach time; see the breaking-change note above.
get_entrypoint_coverage()over Neo4j reads the projected report. It answered with anentrypoint_report_unavailablediagnostic unconditionally; codeanalyzer-python 1.4.1 (#182) projectsentrypoint_frameworks/entrypoint_report_jsononto:PyApplication, so on such a graph the answer is now the pass's ownPyEntrypointReport, field for field what the local backend returns. The diagnostic survives only for a graph that genuinely lacks the property (one emitted by 1.4.0).- Resolved through the 1.4.1 pin -- python-sdk's upstream reports #176 / #177 / #178, fixed
in codeanalyzer-python as #180 (body nodes and parameters carry their
idinanalysis.json; the SDK's composed body-node id is now pinned equal to the analyzer's own per run), #182 / #185 (entrypoint report projected, Odoo@http.route/http.Controllerdetected: 534 callables and 94 classes on the same checkout that 1.4.0 flagged 0 / 0), #181 (PY_EXTENDSis actually emitted: 1,573 edges where every 2.0.0 graph had none) and #183 (_moduleretired,:PyCanNoderange index onid). - The N+1 timing assertion measured the coverage tracer, not the query. The two timed live
tests (
get_symbol_table/get_classesunder 15 s) ran undersys.settrace, which adds ~5 s to a ~10 s Python-side reconstruction; they passed the ceiling by luck. They now run with coverage paused (pytest-cov'sno_cover), so the ceiling measures the round trips it exists to bound. - Java: a missing JDK cache root is refused, not crashed on.
JCodeanalyzer._get_codeanalyzer_execnow raisesCodeanalyzerExecutionException("no cache directory and no project directory") when neither is available, instead of lettingensure_jdkfail withTypeError: ... not 'NoneType'. Only single-file source mode reaches this path; that mode is won't-fix (#256) and its ten witness tests are skip-gated under #255, so the release test gate (live since #306) is green again (#328). - Java: the cached Temurin JDK is actually reused.
ensure_jdklooked for<cache>/jdk/<release>/bin/java, but the archive extracts nested (<release>/bin, or<release>/Contents/Home/binon macOS), so the cache never hit; without a$JAVA_HOMEcarryingjmodsevery process re-downloaded the JDK and then failed extracting over the read-onlylib/server/*.jsafiles already there. It now finds the extracted JDK the same way the download does. Mocked-analyzer Java tests no longer resolve a JDK at all, andJAVA_HOMEno longer leaks between tests. - Neo4j statements no longer leak across applications in a shared database. The per-parent
child fetches (
get_class()and everything reconstructing one declaration) matched by a bare{signature: $sig}/{file_key: $fk}while their bulk twins were application-scoped, so in a database holding two applicationsget_class()merged another application's members andget_all_classes()did not. All twelve carry the same application-scope predicate now (_module IN $modswhen this landed, thecan://id prefix since leg 1.6), and so do the leg-1.5 call-graph statements behindreaches,backward_cone,call_paths_betweenandflows_to_call, which had matched by signature unscoped too. Statements keyed only by a body-node or ghost id (slice_*,paths_between, the value-reachability predicate) are scoped by construction — the id embeds the application — and stay as they are. A fake two-application driver pins the child fetches, and an audit over every class-level statement pins the rule: matched by signature ⇒ carries the scope; otherwise keyed by id. - The call-graph walks no longer route through an external ghost on the Neo4j backend.
reaches,backward_coneandcall_paths_betweenlabelled only the endpoints of their variable-length match, so an intermediate could be a:PyExternal— and ghosts do have outgoingPY_CALLSedges (5,307 on the measured application, 198 landing on a declared callable). For the two in-applicationcallable → ghost → callablechains there,reachesansweredTruewith no all-callable route andcall_paths_betweenreturned a path with anexternalinterior node, whileget_call_graph— built from declared-origin edges on both backends — has no such route. The walks are now constrained at every hop (a quantified path pattern forreaches/backward_cone, an inlinedall(n IN nodes(p) WHERE n:PyCallable)for the shortest-path query); measured cost 0.25s against 0.03s, 0.27s against 0.08s, and unchanged respectively. - The local call graph has the graph backend's shape. It kept edges originating at an external
ghost and edges landing on a class node, both of which the Neo4j projection filters by label, so
reaches/callers_of/call_paths_between/backward_cone— all of which read from it locally — could differ by backend on the same project.backward_conealso now orders byreflocally, asSlice.nodesdocuments and as the graph'sORDER BY m.iddoes, so a capped cone takes the same prefix on both backends. LocateResult.module.pathis repo-relative on the local backend. It was the absolute path on the analysing machine (/Users/…/src/pay.py) while the Neo4j backend answeredaddons/…/payment.py; themodule_scopediagnostic leaked the same string. Both now use the symbol-table key, the vocabularySliceNode.fileandPyCallableOverview.pathalready share.describe()composes withcallees_of().callees_ofdeliberately returns externals anddescribeaccepts anything with a ref, but the five externals of a typical callee list raisedKeyError: 5 of 7 refs name nothingand printed theircan://ids. An external is found and has no source by definition, so it comes backsource=None; theKeyErrorremains for a ref that genuinely names nothing and now reports positions by callable andfile:line.- No
can://id or ordinal in an error message. The self-question refusal onpaths_betweenrendered both endpoints as…@formal_in:0and advised areaches(...)call with those ids as arguments, which does not run; it now names the value within its callable and advises thereachescall on the enclosing callable, which does. - Argument validation precedes name resolution on both backends.
page_size=0plus a bad name raisedValueErrorover Neo4j andSelectorNotInGraphlocally; the cheap argument check now comes first everywhere, before any round trip. analysis_level=now reaches the analyzer.PythonAnalysis/PyCodeanalyzeraccepted the parameter, stored it, and then builtcodeanalyzer-python'sAnalysisOptionswithout it, so every in-process analysis ran at the analyzer's own default (level 1) whatever the caller asked for:cfg/cdg/ddgwere always empty and noformal_invertex ever existed, which is also exactly what a correct level-1 run looks like. The fourAnalysisLevelnames now map to the analyzer's levels 1–4, soanalysis_level="program_dependency_graph"produces intraprocedural dataflow and"system_dependency_graph"produces the interprocedural verticesresolve_valueaddresses.AnalysisLevelmember-name spellings ("call_graph") are accepted alongside the enum's values ("call graph"); an unrecognised name now raises instead of silently becoming level 1.call_graphis also built at every level at or above"call_graph", not only exactly at it. Deeper levels cost more analysis time — callers who relied on the parameter being inert will now pay for the level they asked for.- Body-node ids no longer double the
@. The local backend composed a body node's id asf"{callable.id}@{body_key}", but the analyzer's synthetic body keys already carry a leading@("@formal_in:1","@entry"), soresolve_valueproduced…charge(self,invoice_id)@@formal_in:1— a string naming nothing in the graph and nothing inPyCallable.body. Composition now follows the emitter's own rule (_global_ordinal), andget_source()accepts both spellings of a body key when splitting one back apart. - A captured global is no longer reported as a parameter. 84% of the
formal_invertices on a real application are module globals the callable reads, and both backends labelled themkind="parameter"with the analyzer's internal"<global>:payment::AccessError"inSliceNode.name. They now come backkind="global"(or"capture"for a closure capture) with the readable identifier innameand the defining module indefined_in, and are addressable as"AccessError"or, where several modules supply that name to one callable, as"payment.AccessError". - Ambiguity messages name only keywords the caller can actually use.
resolve_callableno longer advises narrowing within_class=to someone who already passedin_class=, and an ambiguouswithin=insideresolve_value— which accepts neitherin_class=norin_module=— now advises naming more of the dotted path inwithin=, which resolves. - A signature carried by two analysed callables raises instead of resolving arbitrarily. Both backends keyed their candidate map by signature, silently collapsing a collision. Not reachable on a real application (15,549 distinct signatures), and a duplicate elsewhere does not break unrelated resolutions — the check is on the answer.
resolve_callableno longer hands back a sentinel as an address.PyCallable.iddefaults to""andstart_lineto-1; the local backend copied both intoSliceNode, yieldingref=""andline=-1. It now raises.
- C analysis support (
CLDK.c(),cldk.analysis.c,cldk.models.c) — libclang-backed and syntactic-only, it emitted none of the 2.0 code-property graph the query surface is built on and had no analyzer to grow one.clang,libclang,tree-sitter-c, andtree-sitter-goare dropped as dependencies along with it.CLDK.java(),CLDK.python(), andCLDK.typescript()are unaffected. - BREAKING: fifteen
PythonAnalysisaccessors that only raisedNotImplementedError—get_class_hierarchy,get_service_entry_point_classes,get_service_entry_point_methods,get_entry_point_classes,get_entry_point_methods,get_implemented_interfaces,get_methods_with_decorators,get_test_methods,get_calling_lines,get_call_targets,get_all_crud_operations,get_all_create_operations,get_all_read_operations,get_all_update_operations,get_all_delete_operations. Every one raised unconditionally, so no caller could have depended on behaviour — only on the name existing. This also removes theget_entrypoint_classes(works) /get_entry_point_classes(raised) name collision one space apart;See Alsocross-links pointing at the deleted half are removed too.CLDK.java()andCLDK.typescript()keep their same-named accessors, which do work.
- TypeScript bulk/projected accessors — parity with the Python surface from #180/#181, on
both backends (in-process and Neo4j), verified by a live dual-backend parity suite (#298, #302):
get_callables_overview() -> List[TSCallableOverview]— a lightweight projection of every callable (methods, functions, arrows, accessors, namespace and nested functions) with TS-native facets: the analyzer's 7-valuekind,is_exported/is_async/is_static,accessibility, decorator names, and anowner_signature/owner_kindpair ("class" | "interface"; namespace-owned, nested, and module-level callables are ownerless — the dotted signature carries the namespace path). The overview deliberately excludes source text; drill in viaget_method_bodies.get_method_bodies(signatures) -> Dict[str, str]— source bodies keyed by signature; unknown signatures and code-less callables (implicit constructors) are omitted, so every returned value is a realstr.get_decorated_callables(markers) -> List[TSCallableOverview]— overviews of callables decorated with any of the given marker names.get_callsites_for(signatures) -> Dict[str, List[TSCallsite]]— call sites per callable; every existing signature gets an entry (empty list when it has no call sites).
- The blessed TypeScript test fixture (
tests/resources/typescript/analysis_json/slim/analysis.json) is now tracked, so the test suite runs from a clean clone.
- Getter/setter pairs share one signature in the analyzer's output; the two backends currently
collapse such pairs differently (#300 — documented on
get_callables_overview).
- Published wheels bundle the
codeanalyzer-javaJAR again. Every hatchling-built wheel/sdist since v1.2.0 shipped without the JAR, sopip install cldk+CLDK.java(...)raisedCodeanalyzerExecutionException: codeanalyzer jar not found. The root.gitignore*.jarrule was applied at build time while the nested!codeanalyzer-*.jarnegation (which keeps the JAR in git) was not honored, silently dropping the JAR from any build run inside a git repo — i.e. CI. A[tool.hatch.build] artifactsrule force-includes it, and the release workflow now fails if a built artifact is missing the JAR. (#284)
pip install cldkno longer fails on Python 3.14. Removed the unusedpyarrow==20.0.0dependency (the orphaned companion of the previously removed pandas): nothing in the SDK imports it, and its cp39–cp313-only wheels forced an Arrow C++ source build — and usually a failure — on Python 3.14 installs. (#145)- TypeScript Neo4j backend:
get_external_symbols()works again. Reconstructing External (phantom) nodes raised a Pydanticextra_forbiddenerror for every node, because the reconstructor forwarded graph properties (signature, and a fabricatedkind) that the slimTSExternalSymbolmodel deliberately omits. The reconstructor now conforms to the model and the analyzer's published Neo4j schema; the node'ssignatureremains the map key. (#231)
- Intra-class method calls resolve to the method again. Upgraded
codeanalyzer-python0.3.0 → 0.3.1, which fixes callee resolution forself._method(...)calls: the call graph previously pointed such edges at the class node (truncating call chains at their deepest hop) and could omit receiver-method edges entirely. Call-sitecallee_signaturenow names the method, andPyCallsitegains an additiveargumentsfield. (codellm-devkit/codeanalyzer-python#94, #260)
- Module-level functions are now resolvable through
get_method(Python and TypeScript, both backends). Previouslyget_methodsearched class scope only, so module-level functions were unreachable andget_all_callers/get_all_calleesreturned a silent empty result for them even when the call graph contained the edges. The scope argument now accepts a module name as well as a qualified class name; callers/callees,get_method_parameters, and comment lookups inherit the fix. (#250, #251) - Java lookups no longer crash on a miss.
get_method/get_class/get_java_fileare honestly annotated... | None(ABC, both backends, and the public facade),get_method_parametersreturns an empty list instead of raisingAttributeErrorfor an unknown class or signature, and all internal call-graph/comment lookups are guarded. No behavior change on successful lookups. (#252)
- Upgraded
codeanalyzer-python0.2.0 → 0.3.0, which drops CodeQL and uses PyCG for call-graph construction.
use_codeql(BREAKING). Because codeanalyzer-python 0.3.0 removed CodeQL, theuse_codeqlknob no longer maps to anything and is removed from CLDK's public surface: thePyCodeAnalyzerConfig.use_codeqlfield, the deprecatedCLDK(language).analysis(use_codeql=...)parameter, and thePyCodeanalyzer(use_codeql=...)argument. TheCodeQLDatabaseBuildExceptionandCodeQLQueryExecutionExceptionexception classes are removed as well. Call-graph results may differ (PyCG vs CodeQL-augmented Jedi). See #185.
- Bulk, field-projected accessors on the Python facade (
PythonAnalysis) for enumerating an application set-at-a-time, instead of paying the per-entity reconstructionget_methods()does (tens of thousands of round-trips on the Neo4j backend for a large app):get_callables_overview()(returns the newPyCallableOverviewprojection),get_method_bodies(signatures),get_decorated_callables(markers), andget_callsites_for(signatures). On the read-only Neo4j backend each is a single projected Cypher query; in-process each is one symbol-table walk. NewPyCallableOverviewmodel incldk.models.python. tsc_onlytoggle for the TypeScript backend. NewTSCodeAnalyzerConfigbackend config exposestsc_only, which passes--tsc-only(codeanalyzer-typescript >= 0.4.2) to pin the call graph to the resolver path, replacing reliance on the obsolete--call-graph-provider both.- Synthesized anonymous callables for TypeScript. New
TSSynthesizedCallablemodel andget_synthesized_callables()on the TypeScript analysis surface expose the Jelly-resolved anonymous-callback endpoints the symbol table never names, so anonymous call-graph edges no longer dangle. Empty under thetsc-only resolver. - README Cited By section highlighting papers that cite CLDK (SAINT, ASTER, RECON, PRAXIS, Phaedrus, and others), compiled from Semantic Scholar / OpenAlex citation data.
- The read-only Neo4j Python backend now reuses a single read session across queries instead of opening one per Cypher statement, cutting per-call overhead on the reconstruction path.
- Bumped
codeanalyzer-typescript0.4.0 → 0.4.3.
- Per-language factory methods on
CLDK—CLDK.java(),CLDK.python(),CLDK.typescript(), andCLDK.c()— each with an honest signature exposing only the options that apply to that language. These are the preferred entry points, replacing the stringly-typedCLDK(language).analysis(...). - Typed backend-configuration objects in
cldk.analysis.commons.backend_config. The backend is now selected by the type of thebackend=config passed to a factory:CodeAnalyzerConfig(default; in-process analyzer) /PyCodeAnalyzerConfig(addsuse_codeql,use_ray), orNeo4jConnectionConfig(read-only Neo4j).Neo4jConnectionConfigis hoisted here and re-exported fromcldk.analysis.{python,typescript}.neo4jfor backward compatibility. - Unified, language-keyed cache directory. All backends now share a single
cache_dir(default<project>/.codeanalyzer) and write their artifacts under a per-language subdirectory (<cache_dir>/java,<cache_dir>/python,<cache_dir>/typescript), so a polyglot project analyzed under more than one language no longer overwrites a sharedanalysis.json.
- Caching is on by default for Java/TypeScript. The in-process backend now caches
analysis.jsonto disk (under the language-keyedcache_dir) instead of streaming over a stdout pipe. CLDK(language).analysis(...)is deprecated and retained as a thin compatibility shim that forwards to the new factory methods (emits aDeprecationWarning).
- Java
source_code(single-file) input — passproject_pathinstead.
analysis_backend_pathfrom the public interface. The backend binary ships with the packagedcodeanalyzer-*dependency; for TypeScript,$CODEANALYZER_TS_BINremains as the only out-of-band override.analysis_json_pathfrom the public interface — folded into the unifiedcache_dir.
- The language-keyed cache relocates
analysis.jsonfrom<cache_dir>/analysis.jsonto<cache_dir>/<language>/analysis.json; existing caches are not found at the new path, so the first run after upgrading recomputes the analysis.
- Read-only Neo4j-backed TypeScript analysis backend (
cldk.analysis.typescript.neo4j.TSNeo4jBackend). It is a drop-in alternative to the in-memoryTSCodeanalyzer: it answers the sameget_*query surface (call graph, callers/callees, class hierarchy, call sites, decorators, symbol lookups, ...) by running Cypher over a live Neo4j graph instead of walking the pydantic / NetworkX structures. The graph is the onecodeanalyzer-typescriptemits with--emit neo4j(schemaschema.neo4j.json); it is always populated out of band, and the SDK only polls it (read-only — never writes, needs no binary or project sources). TypeScriptAnalysis/CLDK.analysis(language="typescript")now accept an optionalneo4j_config(Neo4jConnectionConfig) to select the Neo4j backend; without it the in-memory backend is used, unchanged.- Read-only Neo4j-backed Python analysis backend (
cldk.analysis.python.neo4j.PyNeo4jBackend), the analog of the TypeScript one. It answers all 21PythonAnalysisBackendqueries via Cypher over the graphcodeanalyzer-python(>= 0.2.0) emits with--emit neo4j. Verified against a real 57-module project: every node/edge present in the graph reconstructs identically to the in-memoryPyCodeanalyzer(3169/3200 checks; zero weight/provenance mismatches on shared call edges). Known gaps are not in the query layer: projection-lossy fields (comments → docstring,PyVariableDeclaration.value/columns, per-binding import detail), and an upstream emitter bug where calls to a bare module name that is also imported (e.g.os/re/json) are dropped from the emitted call graph.PythonAnalysis/CLDK.analysis(language="python")accept the same optionalneo4j_config. - Read-only Neo4j-backed Java analysis backend (
cldk.analysis.java.neo4j.JNeo4jBackend), completing Neo4j parity across all three languages. It reconstructs the canonicalJApplicationfrom the graphcodeanalyzer-java(>= 2.4.0) emits with--emit neo4jand answers all 36JavaAnalysisBackendqueries with the in-memory backend's logic. Verified against the daytrader8 sample (145 classes): everything the graph actually contains reconstructs identically toJCodeanalyzer(97% of checks). Three projection gaps in thecodeanalyzer-java2.4.0 emitter (fields collapsing to one node, imports reduced to packages, a truncated call graph) are fixed in 2.4.1 (codeanalyzer-java#156/#157/#158, verified on daytrader —J_CALLSwent 287 → 1702), the version the SDK release now bundles.JavaAnalysis/CLDK.java(...)accept aNeo4jConnectionConfigas thebackend=config to select it. - Bumped
codeanalyzer-pythonto0.2.0(adds the Neo4j graph emitter); the bundledcodeanalyzer-javajar is now2.4.1(adds the Neo4j graph emitter + the field/import/call-graph projection fixes). The Java analyzer jar is no longer a pip dependency — the SDK release workflow downloads the latestcodeanalyzer-javajar into the bundledjar/directory. - Optional
neo4jextra (pip install cldk[neo4j]) for the Neo4j Python driver.
- Bundled JDK download for the Java backend.
ensure_jdkresolved the Temurin JVM via the Adoptium/assets/version/{release}endpoint, which now returns 404 for pinned releases (e.g.jdk-21.0.5+11) — so the first Java analysis on a clean machine failed before it started. It now resolves via the/binary/version/...endpoint (following the redirect to the GitHub asset) and reads the checksum from the asset's.sha256.txt.
- Doctest-style Examples across the public API surface of JavaAnalysis, PythonAnalysis, CAnalysis, and core CLDK helpers. Coverage includes Java CRUD operations and comment/docstring query APIs, plus concise inline examples for Python and C where applicable.
- Examples documenting expected NotImplementedError behavior for placeholder APIs (PythonAnalysis and CAnalysis) using doctest flags.
- Converted and standardized docstrings to strict Google style (Args, Returns, Raises, Examples) across edited modules.
- Standardized Examples to use the CLDK facade (e.g.,
CLDK(language="java").analysis(...)) instead of raw constructor calls. - Normalized all doctest Example inputs to single-line strings to ensure reliable mkdocstrings rendering.
- Clarified
CLDK.analysisreturn type with a precise union:JavaAnalysis | PythonAnalysis | CAnalysis. - Updated codeanalyzer version to v2.3.6.
- Fixed README.md logo display on PyPI by updating image URLs to use raw GitHub URLs and maintaining theme-based auto-switching with proper fallback
- mkdocstrings rendering issues caused by multi-line doctest strings and formatting inconsistencies.
- Replaced confusing examples like
JavaAnalysis(None, None, ...)with clear CLDK-based initialization patterns. - Packaging: ensured the built wheel includes the
cldkpackage by addingpackages = [{ include = "cldk" }]to Poetry configuration. - Fixed #141
- Multi-line doctest strings in Examples that broke mkdocstrings rendering; all examples are now single-line.
- Removed pandas dependency (#145)
- Added
argument_exprfield to JCallSite model for capturing actual parameter expressions in method calls - Added Star History section to README.md for tracking project popularity
- Updated codeanalyzer jar to version 2.3.5 with support for call argument expressions and fully qualified parameter types
- Modified codeanalyzer.py to preserve fully qualified parameter types in method signatures instead of simplifying them
- Updated method signature format to use fully qualified type names (e.g.,
java.lang.Stringinstead ofString) - Updated test fixtures with new analysis.json data reflecting the signature format changes
- Fixed method signature handling to maintain fully qualified parameter types for better type resolution
- Updated test cases to use fully qualified method signatures for improved accuracy
v1.0.5 - 2025-06-24
- Fixed issue #135
- Analysis level compatibility checking for analysis.json with passed analysis level
- Updated treesitter analysis to use global declarations of parser and language
v1.0.4 - 2025-06-11
- Added missing callable fields field validator
- Updated test fixture setup to use codeanalyzer jar from cldk/analysis/java/codeanalyzer/jar instead of test resources directory
- Updated analysis.json fixtures (daytrader8 and plantsbywebsphere)
- Removed dangling codeanalyzer jars from test resources
- Removed obsolete analysis.json fixture
v1.0.3 - 2025-06-01
- Added code start line attribute to JCallable (corresponding to added attribute in the java code analyzer model)
v1.0.2 - 2025-05-24
- Added test case and fixture for source analysis
- Added missing attributes in compilation unit model
- Fixed handling of
source_codeoption in Java codeanalyzer - Updated core.py to match python analysis signature
v1.0.1 - 2025-05-07
- Updated treesitter analysis to use global declarations of parser and language
v1.0.0 - 2025-04-29
- First stable release
- Updated contributing guidelines
- Updated README.md
- Updated codeanalyzer jar
- Updated java version in release automation
v0.5.1 - 2025-03-13
- Updated Java model to comply with codeanalyzer v2.3.1
- Updated codeanalyzer jar to the latest from codeanalyzer-java
- Updated get_all_docstrings to return dict
v0.5.0 - 2025-02-21
- Added release automation github actions
- Added Java 11 support in github actions
- Added release_config.json
- Added Comment parsing APIs at file, class, method, and docstring level
- Added support for parsing callable parameters and their location information
- Added Dev container instructions with Python, Java, C, and Rust support
- Added C/C++ analysis support
- Added CRUD operations support for Java JPA applications
- Consolidated analysis_level enums in init.py
- Updated codeanalyzer jar to the latest version
- Changed coverage minimum to 70%
- Updated documentation with mkdocs
- Updated badges and logos in README
- Added Discord community support
- Removed CodeQL dependency and refactored treesitter
- Removed ABCs from analysis
- Removed logic to find LLVM in linux OSes (only appears in Darwin)
- Removed redundant is_entry_point fields from JCallable and JType
- Removed unused parameters and code cleanup
- Fixed various test cases and compatibility issues
- Fixed treesitter superclass identification issues
- Fixed entry point detection code
- Fixed recursive error issues
v0.4.0 - 2024-11-13
- Fixed issue 67 - symbol table is none
- Updated poetry build rules to include codeanalyzer-*.jar
- Added test case to verify jar file exists
v0.3.0 - 2024-11-12
- Support for reading slim JSON from codeanalyzer v1.1.0
- Added more test tools (pylint, flake8, black, pspec, coverage)
- Added test coverage reporting
- Updated README.md to include the arXiv paper
- Removed obsolete test cases for unsupported languages
v0.2.0 - 2024-10-11
- Added GitHub Action to publish manual releases
- Added PyPi badge to README.md
v0.1.4 - 2024-10-21
- Fixed codeanalyzer.jar not being a PosixPath
v0.1.3 - 2024-10-21
- Fixed calling the correct codeanalyzer jar on version 0.1.3
- Removed auto-download of codeanalyzer jar
v0.1.2 - 2024-10-17
- Fixed tree-sitter bug
- Defined self.captures explicitly
0.1.0-dev - 2024-10-07
- Initial development version
- Set version to über json support
- Support for slim JSONs from codeanalyzer
- IBM Copyright added to all source files
- Added code parsing support
- Added support for symbol table call graph
- Added notebook examples for code summarization and test generation
- Basic CLDK framework implementation
- Updated dependencies in pyproject.toml
- Added metadata for PyPi distribution
- Updated README with installation instructions
- Fixed caller method implementation
- Fixed incremental analysis support
- Fixed download jar issues