From d455ba5768a01b3b22b812d31f29aa0e85242b14 Mon Sep 17 00:00:00 2001 From: Rahul Krishna Date: Sun, 6 Sep 2026 00:51:36 -0400 Subject: [PATCH 01/13] =?UTF-8?q?docs(spec,plan):=20leg=201.6=20=E2=80=94?= =?UTF-8?q?=20scope=20by=20can://=20id=20prefix,=20retire=20=5Fmodule=20(#?= =?UTF-8?q?327)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit codeanalyzer-python 1.4.1 drops the _module property; the SDK's Neo4j backend scopes 170 sites on it and its probe cannot tell the two graphs apart. Eight decisions (F1-F8): one id prefix replaces every scope predicate, module keys are derived from ids by verified longest match, the ghost rule survives by label, narrowing stays a separate mechanism from scoping, the audit widens before the code moves, and the probe reads analyzer_version. Dual support: 1.4.0 graphs are served with a one-line warning, older refused. --- .../2026-09-06-leg-1.6-id-prefix-scoping.md | 293 ++++++++++++++++++ .../2026-09-06-leg-1.6-id-prefix-scoping.md | 66 ++++ 2 files changed, 359 insertions(+) create mode 100644 docs/design/plans/2026-09-06-leg-1.6-id-prefix-scoping.md create mode 100644 docs/design/specs/2026-09-06-leg-1.6-id-prefix-scoping.md diff --git a/docs/design/plans/2026-09-06-leg-1.6-id-prefix-scoping.md b/docs/design/plans/2026-09-06-leg-1.6-id-prefix-scoping.md new file mode 100644 index 00000000..27dc9360 --- /dev/null +++ b/docs/design/plans/2026-09-06-leg-1.6-id-prefix-scoping.md @@ -0,0 +1,293 @@ +# Leg 1.6 — id-prefix scoping Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Make `PyNeo4jBackend` serve graphs emitted by codeanalyzer-python 1.4.1 (no `_module` property) by scoping every query on the `can://` id prefix, without moving a single public accessor, and pin 1.4.1. + +**Architecture:** One scope helper replaces 49 predicates and 3 hop-scopes; one derivation helper replaces 13 path projections; the probe learns to read the analyzer version; the audit is widened first so it can see what changes. Both containers (1.4.1 on 7689, 1.4.0 on 7688) gate every task. + +**Tech Stack:** Python 3.11+, pydantic, neo4j driver 5.x, Cypher 5 (quantified path patterns already required at 5.9+), codeanalyzer-python 1.4.1 (`codeanalyzer.schema.ids.application_id` / `module_id`). + +**Spec:** `docs/design/specs/2026-09-06-leg-1.6-id-prefix-scoping.md` (decisions F1–F8). + +## Global Constraints + +- **The public API does not move** (F3). No accessor changes name, signature, or return type. Path values stay repo-relative module keys. +- **Read-only against both graphs.** All SDK and test Cypher is `MATCH`/`RETURN`. Never point anything at port 7687 (an ssh tunnel). `tests/analysis/python/test_python_neo4j_backend.py` emits into a scratch server and is gated on `CLDK_TEST_NEO4J_WRITE_URI`; never set that variable to 7688 or 7689. +- **Work in the leg 1.6 worktree**, not the main checkout (another branch is active there). Stage files by name; never `git add -A`. +- Live env for 1.4.1: `CLDK_TEST_NEO4J_URI=bolt://localhost:7689 CLDK_TEST_NEO4J_USER=neo4j CLDK_TEST_NEO4J_PASSWORD=cldkleg16test`. For 1.4.0: `bolt://localhost:7688`, password `cldkleg1test`. Application name `odoo-slim-19` on both. +- Floors at `b5a18b4` (leg 1.5, 1.4.0 pin): offline `tests/analysis/python/ tests/models/python/` **373 passed / 141 skipped**, coverage **54.35%** (gate 50%); live **516 passed**; contract **8 passed**. +- No suggestions, fuzzy matching or edit distance anywhere (leg 1.5 E8). No `can://` or ordinal in any signature, return field other than `ref`/`node_id`/`next_cursor`, or exception message (E6/E7). +- **Never add Claude/AI attribution** to any commit, comment, doc, or changelog entry. + +## Verification environment + +The 1.4.1 graph on 7689 is being emitted with a fresh cache and takes roughly 50 minutes; Task 0 does not need it, Tasks 1–3 do. Confirm it is complete before any live run against 7689: + +``` +MATCH (a:PyApplication {name:'odoo-slim-19'}) RETURN a.analyzer_version, a.schema_version +MATCH (n) RETURN count(n) -- 1.4.0 image on 7688 has 969,421 +``` + +--- + +## Task 0: Bump the pin and measure what moves + +Pure measurement before any query changes: what does 1.4.1 change for the **local** backend and the test suite, with the Neo4j backend left alone? + +**Files:** +- Modify: `pyproject.toml:38` (`codeanalyzer-python==1.4.1`), `pyproject.toml:93` (`[tool.backend-versions]`), `uv.lock` +- Test: nothing new yet; the point is to watch existing tests + +**Interfaces:** +- Consumes: nothing. +- Produces: a written list of every test whose result or meaning changed under 1.4.1, with the emitter change each one traces to. + +- [ ] **Step 1: Bump the pin and re-sync** + +``` +sed -i '' 's/codeanalyzer-python==1.4.0/codeanalyzer-python==1.4.1/' pyproject.toml +sed -i '' 's/^codeanalyzer-python = "1.4.0"/codeanalyzer-python = "1.4.1"/' pyproject.toml +uv lock && uv sync --all-groups --all-extras +uv run python -c "import importlib.metadata as m; print(m.version('codeanalyzer-python'))" +``` +Expected: `1.4.1`. + +- [ ] **Step 2: Run the offline suite and read every change, not just the count** + +``` +uv run pytest tests/analysis/python/ tests/models/python/ -q --ignore=tests/analysis/python/test_python_neo4j_backend.py -p no:cacheprovider +``` +Against the floor (373 / 141). For every test that newly fails, newly passes, or newly skips, name the 1.4.1 change responsible. Two are expected: + +- **#180 — body nodes and parameters now carry `id` in `analysis.json`.** `cldk/analysis/python/codeanalyzer/codeanalyzer.py:125` `body_node_id()` composes the id from `callable.id` and the body key. Add a test asserting that, for every body node in a real level-4 run over `tests/`' small fixtures, the composed id **equals** the analyzer's own `BodyNode.id` when present. This turns leg 1.5's "verified 885,218 / 885,218 by construction" into a per-run assertion. If any differ, that is a finding to report, not to reconcile silently. +- **#182 — the entrypoint report and Odoo detection.** Any test pinning "zero entrypoints" or `entrypoint_report_unavailable` on a fixture is recording 1.4.0 behaviour; re-read each and update the *expectation*, keeping the test's claim. + +- [ ] **Step 3: Confirm the Neo4j backend is now broken against 7689, on purpose** + +Once 7689 is populated, attach and run one scoped call: + +``` +uv run python -c " +from cldk import CLDK +from cldk.analysis.python.neo4j import Neo4jConnectionConfig +py = CLDK.python(backend=Neo4jConnectionConfig(uri='bolt://localhost:7689', username='neo4j', password='cldkleg16test', application_name='odoo-slim-19')) +print(len(py.get_callables_overview()))" +``` +Expected before Task 1: **0** with no error — record it. This is the defect the leg fixes; Task 1's first test pins that it becomes 15,549. + +- [ ] **Step 4: Commit** + +``` +git add pyproject.toml uv.lock tests/analysis/python/ +git commit -m "chore(python): pin codeanalyzer-python 1.4.1 and pin the body-node id composition to the analyzer's own" +``` + +--- + +## Task 1: Widen the audit, then move the scope predicates + +**Files:** +- Modify: `tests/analysis/python/test_neo4j_multi_application_scope.py` (the rule, the enumeration, the fake server) +- Modify: `cldk/analysis/python/neo4j/neo4j_backend.py` — every **scope-predicate** and **hop-scope** site (52), `_load_module_keys` / `_modules` plumbing, `_probe_schema` +- Test: `tests/analysis/python/test_neo4j_multi_application_scope.py`, `tests/analysis/python/test_scoping_keywords.py:447`, `test_e2e_neo4j_live.py` + +**Interfaces:** +- Consumes: `codeanalyzer.schema.ids.application_id(app) -> "can://python/"`, `module_id(app, file_key)`. +- Produces: `PyNeo4jBackend._scope_prefix: str` (`application_id(app) + "/"`), `PyNeo4jBackend._module_prefixes(keys: Sequence[str]) -> list[str]` for narrowing, `_probe_schema` raising `GraphSchemaMismatch` below 1.4.0 and logging once on 1.4.0. + +- [ ] **Step 1: Widen the audit first (F7) and watch it go red** + +In `test_neo4j_multi_application_scope.py`: add `_MATCHES_BY_PREFIX = re.compile(r"\.id STARTS WITH \$")`; change the rule to "by-signature statements must carry `IN $mods` **or** a prefix predicate"; change `_class_level_statements()` to accept any class-attribute string containing a Cypher keyword at its start (`MATCH`, `OPTIONAL MATCH`, `UNWIND`, `CALL`), not only `startswith("MATCH")`. Then change `_fake_two_app_cypher` so its scope filter is `c["id"].startswith(prefix_param)` for prefix statements and keep the `_module`/`IN $mods` branch only until Step 4 removes it. Add a test that constructs the fake graph with **no `_module` property at all** (that is what a 1.4.1 graph is) and asserts the current backend returns empty — the red that Step 4 turns green. + +Run: `uv run pytest tests/analysis/python/test_neo4j_multi_application_scope.py -q` +Expected: the new no-`_module` test FAILS (returns empty); `test_the_audit_sees_the_dataflow_statements_too` may now list more statements than 11 — record the new count. + +- [ ] **Step 2: Add the scope helpers** + +```python +from codeanalyzer.schema.ids import application_id, module_id + +# in _init_with_driver, after application_name is known: +self._scope_prefix = application_id(self.application_name) + "/" + +def _module_prefixes(self, keys: Sequence[str]) -> list[str]: + """Per-module id prefixes for NARROWING (F6) -- scope is self._scope_prefix alone.""" + return [module_id(self.application_name, k) + "/" for k in keys] +``` +Keep `_load_module_keys` / `self._modules`: they still feed `scope_paths`, `resolve_module_key`, and narrowing. + +- [ ] **Step 3: Replace the 49 scope predicates, one statement family at a time** + +For each site in the inventory (`_BULK_CHILD_QUERIES` ×7, `_probe_resolution_edges`, `_callable_full` ×4, `_class_full` ×3, `_call_rows`, `_bounded_call_rows` ×3, `get_python_file`, `get_all_classes`, `get_class`, `get_callables_overview`, `get_method_bodies`, `get_source`, `_RESOLVE_CALLABLE_QUERY`, `resolve_value`, `_OWN_EDGES`, `_REACHES`, `_CONE`, `_CALLERS`, `_CALLEES`, `_CALL_PATHS` ×2, `_CALLEE_VALUES`, `_SOURCES`, `get_decorated_callables`, `get_entrypoints`, `get_entrypoint_classes`, `get_callsites_for`, `get_config_uses`, `get_config_readers`, `_LOCATE_QUERY:1791`): + +`x._module IN $mods` → `x.id STARTS WITH $prefix`, passing `prefix=self._scope_prefix`. + +Where the statement is a **narrowed bulk fetch** (anything reached through `_prefetch_scope`), the predicate is `any(p IN $prefixes WHERE x.id STARTS WITH p)` with `prefixes=self._module_prefixes(scope)`. Measure both this and the module-join alternative on `get_symbol_table(paths=[one module])` against 7689; keep the faster, say which. + +`_bounded_call_rows`' root anchor `root._module IS NULL OR root._module IN $mods` becomes `root.id STARTS WITH $prefix` — ghosts are inside the prefix now, so the `IS NULL` arm is gone by construction; note it in the docstring. + +Run after each family: `uv run pytest tests/analysis/python/test_neo4j_multi_application_scope.py tests/analysis/python/test_scoping_keywords.py -q`. + +- [ ] **Step 4: Move the 3 hop-scopes, preserving the ghost rule (F5)** + +`_bounded_call_rows:790`, `_REACHES:1373`, `_CONE:1392`: `WHERE a._module IN $mods` → `WHERE a:PyCallable AND a.id STARTS WITH $prefix` (the traversal **source** must be a declared callable; a ghost is reached, never traversed through). Update `test_scoping_keywords.py:447`'s string assertion to the new pattern. + +Write the live test the spec's DoD names: find a `callable → @external → callable` chain on 7689 (leg 1.5 found `IrActionsReport._run_wkhtmltoimage → odoo.tools/parse_version → odoo.tools.parse_version.chk`; re-derive it rather than hard-coding), assert `reaches` is `False` across it and `call_paths_between` returns no path with an `external` interior hop. Skip cleanly if the chain does not exist on this graph. + +Then delete the `_module` branch from the fake server. + +- [ ] **Step 5: The probe (F2)** + +In `_probe_schema`, after the relationship-type check: + +```python +row = self._run("MATCH (a:PyApplication {name: $app}) RETURN a.analyzer_version AS v", app=self.application_name) +version = _parse_version(row[0]["v"]) if row and row[0]["v"] else None +if version is None or version < (1, 4, 0): + raise GraphSchemaMismatch(...) # name what was found and the floor +if version < (1, 4, 1): + logger.warning("graph emitted by codeanalyzer-python %s carries no :PyCanNode index; scoped queries scan rather than seek", ...) +``` +Reuse `GraphSchemaMismatch`; extend it if it cannot carry a version message. Unit-test all three branches with the fake driver; live-test that 7688 warns and 7689 is silent. + +- [ ] **Step 6: Run everything against both graphs** + +``` +uv run pytest tests/analysis/python/ tests/models/python/ -q --ignore=tests/analysis/python/test_python_neo4j_backend.py # offline +CLDK_TEST_NEO4J_URI=bolt://localhost:7689 ... uv run pytest tests/analysis/python/ tests/models/python/ -q --ignore=... # 1.4.1 +CLDK_TEST_NEO4J_URI=bolt://localhost:7688 ... uv run pytest tests/analysis/python/ tests/models/python/ -q --ignore=... # 1.4.0 +``` +Expected: offline ≥ floor; 7688 ≥ 516 minus only tests whose 1.4.1-specific expectations Task 0 changed; 7689 — every failure attributed to a path-projection site (Task 2) or a named 1.4.1 emitter change, nothing else. Report the three summary lines verbatim. + +- [ ] **Step 7: Commit** + +``` +git add cldk/analysis/python/neo4j/neo4j_backend.py tests/analysis/python/test_neo4j_multi_application_scope.py tests/analysis/python/test_scoping_keywords.py tests/analysis/python/test_e2e_neo4j_live.py +git commit -m "fix(python): scope every Neo4j statement on the can:// id prefix, not _module + +codeanalyzer-python 1.4.1 drops the _module property; scope now comes from the id +the analyzer already mints (can://python//...). One prefix replaces 49 +predicates. Ghosts fall inside the prefix, so every walk pins its source to +:PyCallable by label to keep the reached-never-traversed rule. The probe reads +analyzer_version: below 1.4.0 refuses, 1.4.0 warns once about the missing index." +``` + +--- + +## Task 2: Derive module keys from ids + +**Files:** +- Create: nothing — one helper in `cldk/analysis/python/neo4j/neo4j_backend.py` (or `reconstruct.py`, where the projection contract lives) +- Modify: the 13 **path-projection** sites, `_LOCATE_QUERY:1790`, `test_e2e_neo4j_live.py:538,567` +- Test: `tests/analysis/python/test_python_bulk_accessors.py:158` (already pins `o.path == _MODULE_KEY`), `test_locate.py`, new unit tests for the helper + +**Interfaces:** +- Consumes: `self._scope_prefix`, `self._modules` (the verified key set), `module_id`. +- Produces: `module_key_of(node_id: str, prefix: str, known: Collection[str]) -> str` — pure, raises on a non-member. + +- [ ] **Step 1: Write the helper's tests first** + +```python +def test_module_key_is_the_id_segment_up_to_the_first_py_boundary(): + known = {"addons/account/models/account_move.py"} + assert module_key_of("can://python/odoo-slim-19/addons/account/models/account_move.py/AccountMove/write(self,vals)", "can://python/odoo-slim-19/", known) == "addons/account/models/account_move.py" + +def test_a_directory_named_like_a_module_cannot_mis_key(): + known = {"pkg/x.py/real.py"} + assert module_key_of("can://python/app/pkg/x.py/real.py/f()", "can://python/app/", known) == "pkg/x.py/real.py" + +def test_a_key_outside_the_application_raises_rather_than_guesses(): + with pytest.raises(KeyError): + module_key_of("can://python/app/gone.py/f()", "can://python/app/", {"kept.py"}) + +def test_a_ghost_id_has_no_module_key(): + with pytest.raises(KeyError): + module_key_of("can://python/app/@external/os/path", "can://python/app/", {"a.py"}) +``` +Run: expected `ImportError`. + +- [ ] **Step 2: Implement it — longest known-key match, not a split** + +```python +def module_key_of(node_id: str, prefix: str, known: Collection[str]) -> str: + """The repo-relative module key embedded in a can:// id (F4). + + Ids are ``/``; a file key can itself contain ``.py/`` as a directory + name, so the key is recovered by LONGEST match against the application's known keys, never by + splitting. A miss raises: a key we cannot verify is a defect, not a guess. + """ + if not node_id.startswith(prefix): + raise KeyError(node_id) + rest = node_id[len(prefix):] + best = max((k for k in known if rest == k or rest.startswith(k + "/")), key=len, default=None) + if best is None: + raise KeyError(node_id) + return best +``` +If `known` is large (1,626 on Odoo) and this sits on a hot path, index `known` once by first segment; measure before optimising. + +- [ ] **Step 3: Replace the 13 projections** + +Each `c._module AS path` / `AS file` / `AS fk` returns `c.id AS id` instead, and the Python side derives the key with `module_key_of(row["id"], self._scope_prefix, self._module_set)`. Sites: `get_python_file:847b`, `_OVERVIEW_PROJECTION:1078`, `_RESOLVE_CALLABLE_QUERY:1133`, `_SLICE:1321`, `_CONE:1394`, `_CALLERS:1413b`, `_CALLEES:1414b`, `_PATHS:1463` (its `head([...| c._module])` becomes `head([... | c.id])`), `_CALL_PATHS:1483`, `get_entrypoint_classes:1598`, `get_config_readers:1759`. For `:PyClass` rows, the id grammar is the same (`/`). + +- [ ] **Step 4: `_LOCATE_QUERY`'s match key** + +`OPTIONAL MATCH (c:PyCallable {_module: pos.path})` → `OPTIONAL MATCH (c:PyCallable) WHERE c.id STARTS WITH pos.module_prefix`, where the Python side supplies `module_prefix = module_id(app, resolved_key) + "/"` per position. The `module_scope` diagnostic keeps naming the key, never the prefix (E6). + +- [ ] **Step 5: Run the vocabulary tests on both graphs** + +`test_python_bulk_accessors.py::test_overview_path_is_the_repo_relative_module_key`, `test_e2e_neo4j_live.py::test_paths_share_one_vocabulary`, `::test_overview_path_joins_locate_and_class_overview`, `test_locate.py`, and the leg-1.5 `test_locate_node_id_joins_to_the_graph` — on 7689 and 7688. Expected: all pass; path values byte-identical to the 1.4.0 run. + +- [ ] **Step 6: Commit** + +``` +git commit -m "fix(python): derive a callable's module key from its can:// id + +The graph no longer stores _module, so the repo-relative path a caller sees +is recovered from the id by longest match against the application's known +module keys -- verified, never split, raised when it cannot be." +``` + +--- + +## Task 3: Full verification on both graphs, docs, and the audit's final shape + +**Files:** +- Modify: `CHANGELOG.md` `[Unreleased]`, `docs/agent-api-reference.md`, the 38 comment/docstring sites that describe `_module` scoping, `docs/design/specs/2026-09-05-leg-1.5-bounded-queries-and-dataflow.md` (one note under the multi-application decision pointing at F1) +- Test: `tests/analysis/python/test_bounded_enumeration.py` (hub numbers on both graphs) + +- [ ] **Step 1: The three full runs, verbatim** + +Offline, 7689, 7688 — exactly as Task 1 Step 6. Expected: offline ≥ 373/141 and coverage ≥ 54%; 7689 and 7688 both 0 failed. Record the three summary lines. + +- [ ] **Step 2: Hub numbers on both graphs** + +`get_call_graph(roots=[hub], depth=1)` and `depth=2` on 7688 must be 638 / 27,568 and 1,818 / 63,871. On 7689, record the numbers; if they differ, name the 1.4.1 change (e.g. #181 `PY_EXTENDS`, #180 ids) that explains it, or report it as a finding. + +- [ ] **Step 3: `grep -c _module`** + +`grep -n "_module" cldk/analysis/python/neo4j/neo4j_backend.py` — every remaining line is prose explaining history, or a name coincidence (`_module_full`, `in_module`). Rewrite the 38 comment sites so none describes a scoping mechanism that no longer exists. + +- [ ] **Step 4: Docs** + +CHANGELOG `[Unreleased]`: **Changed** — pin 1.4.1; scoping by id prefix (no caller-visible change); probe refuses `< 1.4.0`, warns on 1.4.0. **Fixed** — cite upstream #180/#181/#182 and python-sdk #176/#177/#178 as resolved through the pin. `docs/agent-api-reference.md`: the analyzer version floor and what attaching to an older graph does. Leg-1.5 spec: one sentence under the multi-application scope decision pointing at F1. + +- [ ] **Step 5: Commit** + +``` +git commit -m "docs(python): record id-prefix scoping, the 1.4.1 pin, and the graph version floor" +``` + +--- + +## Definition of done + +- Every item in spec §5. +- The three summary lines from Task 3 Step 1 are in the PR body, with the 7688 run named as the back-compat gate. +- No public accessor changed name, signature or return type (`tests/test_public_surface.py` and the contract tests pass unchanged). +- PR against `release/2.0`, `Closes #327`. + +## Not in this plan + +Retiring `_load_module_keys` entirely (it still serves selector validation); adopting `BodyNode.id` from `analysis.json` in place of the composed id (Task 0 pins them equal, which is enough for this leg); the rc.2 cut itself (finishing-cldk-work, after #328 and this land). diff --git a/docs/design/specs/2026-09-06-leg-1.6-id-prefix-scoping.md b/docs/design/specs/2026-09-06-leg-1.6-id-prefix-scoping.md new file mode 100644 index 00000000..eec74354 --- /dev/null +++ b/docs/design/specs/2026-09-06-leg-1.6-id-prefix-scoping.md @@ -0,0 +1,66 @@ +# Leg 1.6 — scope by `can://` id prefix, retire `_module` + +**Status:** decided 2026-09-06. Tracking: python-sdk#327. Follows leg 1.5 +(`2026-09-05-leg-1.5-bounded-queries-and-dataflow.md`). Blocks the `2.0.0rc2` cut. + +## 1. Why this leg exists + +codeanalyzer-python 1.4.1 removes the `_module` node property from every graph it emits. +Its row builder drops it before the write (`rows.py:133`, `props.pop("_module", None)`), and its +own destructive statements now scope on the `can://` id prefix — `x.id = OR x.id STARTS WITH + + '/'` — with a `:PyCanNode` label and a range index on `id` so the prefix predicate seeks +rather than scans. The reason given upstream is the right one: a label anchor plus `_module` is +application-blind, and the id already carries language, application and file. + +`PyNeo4jBackend` scopes on `_module` structurally. Measured on `b5a18b4`: 170 property-relevant +sites — 49 scope predicates, 13 path projections, 3 hop-scopes, 45 parameter-plumbing sites, 22 +test assertions, 38 comments (a further 97 hits are name coincidences such as `in_module=`). +`_probe_schema` checks `CALL db.relationshipTypes()` only. Against a 1.4.1 graph, attach succeeds +and every scoped query returns empty: the ambiguous-empty defect (leg-1 D7) in its worst form, +with no signal to the caller. + +## 2. Decisions + +| # | Decision | Why | +|---|---|---| +| **F1** | **Scope by id prefix, everywhere.** Every application-scope predicate becomes `n.id STARTS WITH $prefix` with `$prefix = application_id(app) + "/"`. The trailing slash is load-bearing: `odoo-slim-19` must not match `odoo-slim-19-b`. | This is the analyzer's own rule, already used once in the SDK (`get_external_symbols`) and in the test purge. One rule, both sides. | +| **F2** | **Dual support, asymmetrically.** The prefix predicate is correct on 1.4.0 graphs too — their ids have the same grammar — so the backend serves both. The probe reads `:PyApplication.analyzer_version` and `CALL db.propertyKeys()`: `< 1.4.0` or unparsable → refuse with `GraphSchemaMismatch`; 1.4.0 (`_module` present, no `:PyCanNode` index) → serve, and log once that scoping will scan rather than seek. | Correctness is free; only performance differs. Refusing 1.4.0 graphs would strand every graph emitted before today for no gain. | +| **F3** | **The public API does not move.** No accessor changes name, signature or return type. Path values stay repo-relative module keys; `LocateResult.node_id`, `SliceNode.ref` and every other id-shaped value are unchanged. | The rung's Iron Rule. This leg changes how queries are written, not what callers see. | +| **F4** | **Module keys are derived, not stored.** A callable's repo-relative path is recovered from its id: strip `$prefix`, take everything up to and including the first `.py/` boundary, and **verify the result is a member of the application's module-key set** (still loaded from `:PyModule.file_key` at attach). A key that fails membership is a defect, raised, never guessed. `locate`'s match key becomes `c.id STARTS WITH module_id(app, pos.path) + "/"`, using `codeanalyzer.schema.ids.module_id`, the exact inverse. | There is no id→key helper in the SDK today; this defines one with a verification step so a pathological directory named `x.py/` cannot silently mis-key. | +| **F5** | **The ghost rule survives the mechanism change.** Leg 1.5 established that an `@external` ghost is *reached, never traversed through*; that was enforced by `a._module IN $mods`, which ghosts fail because they carry no `_module`. A prefix predicate **includes** ghosts (`can://python//@external/…`). Every hop-scope therefore pins the traversal *source* to `:PyCallable` by label. | A mechanical replacement would silently re-open the leak Fix 3 of leg 1.5 closed (`reaches` true through a ghost with no all-callable route). | +| **F6** | **Narrowing is a separate mechanism from scoping.** `_prefetch_scope` passes a *subset* of module keys to narrow bulk fetches; one application prefix cannot express that. Narrowing uses per-module prefixes — `any(p IN $prefixes WHERE c.id STARTS WITH p)` with `$prefixes = [module_id(app, k) + "/" for k in keys]` — or the module join the prefetch already walks, whichever the measurement favours. `$mods` as a list of file keys survives only for selector validation (`scope_paths`, `resolve_module_key`) and narrowing. | Scope answers "which application"; narrowing answers "which of its modules". Conflating them is how a scoping change breaks `get_symbol_table(paths=)`. | +| **F7** | **The audit widens before the code moves.** `test_neo4j_multi_application_scope.py`'s rule gains an affirmative arm — a statement that matches by signature must carry `IN $mods` **or** `.id STARTS WITH $` — and its statement enumeration stops filtering on `startswith("MATCH")`, which today silently skips `_OVERVIEW_PROJECTION` (`OPTIONAL MATCH`) and `_LOCATE_QUERY` (`UNWIND`). The fake two-application server filters on id prefix, not on a `_module` property that will no longer exist. | An audit that cannot see a statement cannot protect it; a fake that filters on a retired property makes every scoping test vacuous the moment the property goes. | +| **F8** | **Pin 1.4.1.** `codeanalyzer-python==1.4.1` in `dependencies` and `[tool.backend-versions]`. 1.4.1 also lands #180 (body nodes and parameters carry their id in `analysis.json`), #181 (`PY_EXTENDS` emitted) and #182 (entrypoint report projected, Odoo detected) — the fixes for python-sdk's upstream reports #176/#177/#178. | Lockstep, and the 2.0 surface wants those three. | + +## 3. What changes for a caller + +Nothing in the surface. Two things in behaviour, both documented: + +- A graph emitted by codeanalyzer-python **older than 1.4.0** is refused at attach with a message + naming the version found and the floor. Before this leg it was served with silent empties. +- Attaching to a **1.4.0** graph logs one line noting that scoped queries scan rather than seek + because the graph carries no `:PyCanNode` range index. Results are identical. + +Tests that pinned 1.4.0 emitter behaviour will move with 1.4.1 and must be **re-read, not re-run**: +`get_entrypoints` on Odoo was 0 (#177) and will not be; the entrypoint report is now projected; +`PY_EXTENDS` now exists. Each is a fact about the analyzer that a test recorded, not a contract. + +## 4. Verification environment + +Two containers, both read-only to the SDK: + +| container | bolt | emitted by | role | +|---|---|---|---| +| `cldk-leg16-it` | 7689 | codeanalyzer-python **1.4.1** (fresh cache) | primary — every live test | +| `cldk-leg1-it` | 7688 | codeanalyzer-python **1.4.0** | before-image and the F2 back-compat gate | + +Port 7687 on the development machine is an ssh tunnel and is never a target. + +## 5. Definition of done + +- `grep -c "_module" cldk/analysis/python/neo4j/neo4j_backend.py` counts only comments explaining history. +- The full live suite passes on 7689 (leg-1.5 baseline on 1.4.0: 516) **and** on 7688. The bounded-call-graph hub numbers (638 / 27,568 at depth 1; 1,818 / 63,871 at depth 2) are unchanged on 7688 and any change on 7689 is attributed to a named 1.4.1 emitter change. +- The multi-application audit enumerates every Cypher statement on the backend, including those that begin with `OPTIONAL MATCH` or `UNWIND`, under the F7 rule; its fake two-application server filters on id prefix. +- A test asserts the ghost rule directly on the new predicates: a `callable → ghost → callable` chain that exists on the live graph does **not** make `reaches` true. +- Attaching to a graph whose `analyzer_version` is below 1.4.0 raises `GraphSchemaMismatch`; attaching to 7688 succeeds and emits the single warning; attaching to 7689 is silent. +- CHANGELOG records the pin bump, the scoping change, the probe behaviour, and the three upstream fixes 1.4.1 brings; `docs/agent-api-reference.md` records the version floor. From 8b708d9825e6f0cd1ec24418cdc59e5ab2bc3947 Mon Sep 17 00:00:00 2001 From: Rahul Krishna Date: Sun, 6 Sep 2026 00:59:57 -0400 Subject: [PATCH 02/13] chore(python): pin codeanalyzer-python 1.4.1 and pin the body-node id composition to the analyzer's own --- pyproject.toml | 4 +- tests/analysis/python/conftest.py | 11 +- tests/analysis/python/test_body_node_ids.py | 112 ++++++++++++++++++++ tests/analysis/python/test_dataflow.py | 9 +- uv.lock | 8 +- 5 files changed, 130 insertions(+), 14 deletions(-) create mode 100644 tests/analysis/python/test_body_node_ids.py diff --git a/pyproject.toml b/pyproject.toml index 10404cbb..ae4b0b5f 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -35,7 +35,7 @@ dependencies = [ "tree-sitter-java==0.23.5", "tree-sitter-python==0.23.6", "tree-sitter-javascript==0.23.1", - "codeanalyzer-python==1.4.0", + "codeanalyzer-python==1.4.1", "codeanalyzer-typescript==0.4.3", ] @@ -90,7 +90,7 @@ include = [ [tool.backend-versions] codeanalyzer-java = "2.4.1" -codeanalyzer-python = "1.4.0" +codeanalyzer-python = "1.4.1" codeanalyzer-typescript = "0.4.3" ######################################## diff --git a/tests/analysis/python/conftest.py b/tests/analysis/python/conftest.py index b1eb0cb2..c56212d5 100644 --- a/tests/analysis/python/conftest.py +++ b/tests/analysis/python/conftest.py @@ -41,12 +41,11 @@ # that don't care about the schema probe (e.g. a future round-trip-counting test) don't have to set # rel_types themselves. # -# PY_EXTENDS is documented (schema/neo4j_backend.py's module docstring) as the class-inheritance -# edge type, but is NOT observed on a real emitted graph: the live 1.4.0 Odoo application used for -# the e2e suite has 0 PY_EXTENDS edges across 1,656 classes, and `CALL db.relationshipTypes()` -# there doesn't even register the type. Kept here as the *documented* vocabulary the schema probe -# would accept, not as evidence codeanalyzer-python 1.4.0 actually emits it -- if you need a fixture -# graph that matches a real emitted one, drop PY_EXTENDS. +# PY_EXTENDS is the class-inheritance edge type. codeanalyzer-python 1.4.0 never landed one on a +# real graph (the live Odoo application had 0 PY_EXTENDS edges across 1,656 classes: its emitter +# looked bases up by signature while `base_classes` held the written spelling, so every row was +# dropped as dangling); 1.4.1 (#181) resolves bases per module and emits them. A 1.4.0-emitted graph +# still has none, so a fixture meant to match one drops PY_EXTENDS. _V2_RELATIONSHIP_TYPES = frozenset( { "PY_CALLS", diff --git a/tests/analysis/python/test_body_node_ids.py b/tests/analysis/python/test_body_node_ids.py new file mode 100644 index 00000000..c83f7885 --- /dev/null +++ b/tests/analysis/python/test_body_node_ids.py @@ -0,0 +1,112 @@ +################################################################################ +# Copyright IBM Corporation 2026 +# +# Licensed under the Apache License, Version 2.0 (the "License"); +# you may not use this file except in compliance with the License. +# You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, software +# distributed under the License is distributed on an "AS IS" BASIS, +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +# See the License for the specific language governing permissions and +# limitations under the License. +################################################################################ + +"""``body_node_id`` against the analyzer's own ``BodyNode.id`` (codeanalyzer-python 1.4.1, #180). + +Until 1.4.1 ``BodyNode`` carried no id in ``analysis.json``, so the local backend composed one from +``callable.id`` and the body key by copying the emitter's ``_global_ordinal`` rule, and leg 1.5 +could only claim the two agreed by construction. 1.4.1 stamps ``BodyNode.id`` (and +``PyParameter.id``) from the same ``codeanalyzer.schema.ids.global_ordinal``, so agreement is now a +per-run assertion over a real level-4 analysis: every id the analyzer wrote must be the one the SDK +composes. A mismatch here is a finding about one of the two rules, never something to paper over +on the SDK side. + +The fixture is small enough to analyse at level 4 in a test and shaped to produce every vertex +kind leg 1.5 addresses: parameters (``formal_in``), a return (``formal_out``), a call between two +callables (``actual_in`` / ``actual_out``), plus a nested function and a nested class so the walk +reaches callables at every depth ``_iter_callables`` enumerates. +""" + +import textwrap + +import pytest + +from cldk.analysis import AnalysisLevel +from cldk.analysis.python.codeanalyzer.codeanalyzer import PyCodeanalyzer, body_node_id + + +@pytest.fixture(scope="module") +def l4(tmp_path_factory) -> PyCodeanalyzer: + root = tmp_path_factory.mktemp("body-ids") + (root / "src").mkdir() + (root / "src" / "pay.py").write_text( + textwrap.dedent( + """ + LIMIT = 100 + + + def helper(x): + return x + 1 + + + class Portal: + class Meta: + def tag(self, n): + return n + + def charge(self, invoice_id): + def bump(v): + return v + 1 + + total = invoice_id * 2 + amount = helper(total) + if amount > LIMIT: + amount = bump(LIMIT) + return amount + """ + ).lstrip() + ) + return PyCodeanalyzer( + project_dir=root, + analysis_level=AnalysisLevel.system_dependency_graph, + analysis_json_path=None, + eager_analysis=False, + cache_dir=tmp_path_factory.mktemp("cache-body-ids"), + ) + + +def _body_nodes(backend): + for c, *_ in backend._iter_callables(): + for key, node in (c.body or {}).items(): + yield c, key, node + + +def test_the_fixture_produces_the_vertex_kinds_the_assertion_is_about(l4): + kinds = {node.kind for *_, node in _body_nodes(l4)} + assert {"entry", "exit", "formal_in", "formal_out", "actual_in", "actual_out"} <= kinds, kinds + + +def test_every_analyzer_stamped_id_is_the_id_the_sdk_composes(l4): + stamped = [(c, key, node) for c, key, node in _body_nodes(l4) if node.id] + assert stamped + mismatches = [(node.kind, key, node.id, body_node_id(c.id, key)) for c, key, node in stamped if body_node_id(c.id, key) != node.id] + assert mismatches == [] + + +def test_no_body_node_is_left_without_an_id(l4): + """Which kinds, if any, the analyzer leaves unstamped -- ``formal_in`` / ``formal_out`` are the + ones leg 1.5 addresses by id, so an unstamped one there is a finding, not a skip.""" + unstamped = sorted({(node.kind, key) for _, key, node in _body_nodes(l4) if not node.id}) + assert unstamped == [] + + +def test_every_parameter_carries_its_formal_in_id(l4): + seen = 0 + for c, *_ in l4._iter_callables(): + for i, p in enumerate(c.parameters or []): + seen += 1 + assert p.id == body_node_id(c.id, f"@formal_in:{i}"), (c.signature, p.name, p.id) + assert seen diff --git a/tests/analysis/python/test_dataflow.py b/tests/analysis/python/test_dataflow.py index 68db0eb7..82f3cff4 100644 --- a/tests/analysis/python/test_dataflow.py +++ b/tests/analysis/python/test_dataflow.py @@ -49,6 +49,7 @@ from codeanalyzer.neo4j.project import _project_program_graphs from codeanalyzer.neo4j.rows import RowBuilder +from codeanalyzer.schema.ids import application_id from cldk.analysis import AnalysisLevel from cldk.analysis.commons.resolve import value_candidate @@ -207,7 +208,9 @@ def test_the_local_answer_is_edge_for_edge_what_the_graph_would_hold(local_l4): database; the live tests above are what covers that. """ rows = RowBuilder() - _project_program_graphs(rows, local_l4.application, {}, {}) + # 1.4.1 anchors ghost ids on the application's can:// id; the analyzer names the app after + # the project directory (codeanalyzer/core.py: `app_name or project_dir.name`). + _project_program_graphs(rows, local_l4.application, {}, {}, application_id(local_l4.project_dir.name)) edges = rows.finish().edges def emitted(rel, *props): @@ -352,7 +355,9 @@ def test_local_pages_are_the_emitter_rows_sliced_by_the_same_order(local_l4): that same key — the two backends agree page for page. """ rows = RowBuilder() - _project_program_graphs(rows, local_l4.application, {}, {}) + # 1.4.1 anchors ghost ids on the application's can:// id; the analyzer names the app after + # the project directory (codeanalyzer/core.py: `app_name or project_dir.name`). + _project_program_graphs(rows, local_l4.application, {}, {}, application_id(local_l4.project_dir.name)) emitted = sorted((e.from_ref.value, e.to_ref.value, e.props.get("var") or "", list(e.props.get("prov") or [])) for e in rows.finish().edges if e.type == "PY_DDG") walked, cursor, pages = [], None, [] while True: diff --git a/uv.lock b/uv.lock index a3fb66bb..4d84ee6f 100644 --- a/uv.lock +++ b/uv.lock @@ -340,7 +340,7 @@ test = [ [package.metadata] requires-dist = [ - { name = "codeanalyzer-python", specifier = "==1.4.0" }, + { name = "codeanalyzer-python", specifier = "==1.4.1" }, { name = "codeanalyzer-typescript", specifier = "==0.4.3" }, { name = "neo4j", marker = "extra == 'neo4j'", specifier = ">=5.14,<7" }, { name = "networkx", specifier = ">=3.4.2,<4" }, @@ -386,7 +386,7 @@ wheels = [ [[package]] name = "codeanalyzer-python" -version = "1.4.0" +version = "1.4.1" source = { registry = "https://pypi.org/simple" } dependencies = [ { name = "astor" }, @@ -403,9 +403,9 @@ dependencies = [ { name = "typing-extensions" }, { name = "uv" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/b9/e9/8ce468742f0f9bea274bd19a144bdfadcea7c994cf9b0fc95f3eb721a6a6/codeanalyzer_python-1.4.0.tar.gz", hash = "sha256:5cfc367c967f26a764306daf23b4359a106418b8f264833a68b438c5b4ef6b4b", size = 224507, upload-time = "2026-09-02T20:23:02.706Z" } +sdist = { url = "https://files.pythonhosted.org/packages/5b/6f/ee449705998ac7cae481435cae9905da51a2a77e677499b7943bcef9391b/codeanalyzer_python-1.4.1.tar.gz", hash = "sha256:8b9f6c19107bccc2c26a862bab0b51fed9287b4f0329b8d4e9d4d9cb0f029f4b", size = 229087, upload-time = "2026-09-06T01:13:28.613Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/dd/ce/3bba0d0eea0e64a5f9ee41fb1c1f74c7c66618dff6dadbaea2de389c621a/codeanalyzer_python-1.4.0-py3-none-any.whl", hash = "sha256:7ebd20231f5c5f3c4f47c71ac5870634aeccc37f20028f4a52402ed39729ec35", size = 230088, upload-time = "2026-09-02T20:23:01.128Z" }, + { url = "https://files.pythonhosted.org/packages/bc/4d/ce2c1b45ca9854231ec1f04b2fb8d5a864463b1c7fdc3368a8e2412a4f5d/codeanalyzer_python-1.4.1-py3-none-any.whl", hash = "sha256:206efe708e30ab85c76b6725e7bf1c70619feb8c312d5128687a282398d58fc0", size = 235028, upload-time = "2026-09-06T01:13:27.3Z" }, ] [[package]] From 9d09ede231290780f1c9fa2b8ef566caf8564021 Mon Sep 17 00:00:00 2001 From: Rahul Krishna Date: Sun, 6 Sep 2026 01:01:32 -0400 Subject: [PATCH 03/13] =?UTF-8?q?docs(plan):=20leg=201.6=20Task=204=20?= =?UTF-8?q?=E2=80=94=20read=20the=20entrypoint=20report=201.4.1=20projects?= =?UTF-8?q?;=20Task=201=20owns=20the=20e2e=20suite's=20own=20=5Fmodule=20C?= =?UTF-8?q?ypher?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- .../2026-09-06-leg-1.6-id-prefix-scoping.md | 39 ++++++++++++++++++- 1 file changed, 38 insertions(+), 1 deletion(-) diff --git a/docs/design/plans/2026-09-06-leg-1.6-id-prefix-scoping.md b/docs/design/plans/2026-09-06-leg-1.6-id-prefix-scoping.md index 27dc9360..5350d7e6 100644 --- a/docs/design/plans/2026-09-06-leg-1.6-id-prefix-scoping.md +++ b/docs/design/plans/2026-09-06-leg-1.6-id-prefix-scoping.md @@ -90,7 +90,7 @@ git commit -m "chore(python): pin codeanalyzer-python 1.4.1 and pin the body-nod **Files:** - Modify: `tests/analysis/python/test_neo4j_multi_application_scope.py` (the rule, the enumeration, the fake server) - Modify: `cldk/analysis/python/neo4j/neo4j_backend.py` — every **scope-predicate** and **hop-scope** site (52), `_load_module_keys` / `_modules` plumbing, `_probe_schema` -- Test: `tests/analysis/python/test_neo4j_multi_application_scope.py`, `tests/analysis/python/test_scoping_keywords.py:447`, `test_e2e_neo4j_live.py` +- Test: `tests/analysis/python/test_neo4j_multi_application_scope.py`, `tests/analysis/python/test_scoping_keywords.py:447`, `test_e2e_neo4j_live.py` — its **own** raw Cypher matches on `_module` at lines **198, 352, 538, 567, 948–949, 974, 1088** (`{_module: path}`, `cl._module IN $m`, `s._module IN $mods`, `_module:$mod`) and every one returns empty on a 1.4.1 graph; they move to the prefix in this task, not Task 2 **Interfaces:** - Consumes: `codeanalyzer.schema.ids.application_id(app) -> "can://python/"`, `module_id(app, file_key)`. @@ -281,9 +281,46 @@ git commit -m "docs(python): record id-prefix scoping, the 1.4.1 pin, and the gr --- +## Task 4: Read the entrypoint report the graph now carries + +Found by Task 0. `PyNeo4jBackend.get_entrypoint_coverage` (`neo4j_backend.py:~1603`) returns an `entrypoint_report_unavailable` diagnostic whose claim — the projection never carries `PyApplication.entrypoint_report` — was true of 1.4.0 and is false of 1.4.1: #182 writes `entrypoint_frameworks` and `entrypoint_report_json` onto `:PyApplication`. Three tests pin the old claim: `test_entrypoints.py:258` and `test_e2e_neo4j_live.py:784` / `:803`; on 7689 the last fails by design with its own "the projection now carries {...}; the diagnostic is stale" message. This is an accessor **body** change — name, signature and return type do not move (F3). + +**Files:** +- Modify: `cldk/analysis/python/neo4j/neo4j_backend.py` `get_entrypoint_coverage`; `reconstruct.py` if the report is rebuilt there +- Test: `tests/analysis/python/test_entrypoints.py:258`, `tests/analysis/python/test_e2e_neo4j_live.py:784,803` + +**Interfaces:** +- Consumes: `:PyApplication.entrypoint_report_json` (a JSON string) and `entrypoint_frameworks`; the local backend's existing `PyEntrypointReport` model. +- Produces: the same `get_entrypoint_coverage` return type, populated from the graph; the diagnostic survives **only** for a graph that genuinely lacks the property (a 1.4.0 graph on 7688). + +- [ ] **Step 1: Flip the three tests' expectations, watch them fail on 7689** + +Keep each test's *claim* (coverage is answerable / not answerable); change the expectation to "answered from the graph" for a 1.4.1 graph and keep "diagnostic" for 1.4.0. Parametrise on the graph's `analyzer_version` where the test is live. + +- [ ] **Step 2: Read the property** + +`MATCH (a:PyApplication {name: $app}) RETURN a.entrypoint_report_json AS j, a.entrypoint_frameworks AS f` — parse `j` through the same pydantic model the local backend returns (1.4.1's #188 dumps it through the compat shim, so the shape matches). If `j` is null, emit the existing diagnostic unchanged. + +- [ ] **Step 3: Parity** + +On 7689, `get_entrypoint_coverage()` over Neo4j must equal the local backend's answer over the same checkout at level 4 for the fields the projection carries; state any documented lossiness. On 7688 the diagnostic is unchanged. + +- [ ] **Step 4: Commit** + +``` +git commit -m "fix(python): read the entrypoint report the 1.4.1 graph carries + +get_entrypoint_coverage over Neo4j answered with a diagnostic whose premise -- +the projection never carries the report -- 1.4.1 (#182) made false. The +diagnostic now fires only when the property is genuinely absent." +``` + +--- + ## Definition of done - Every item in spec §5. +- `get_entrypoint_coverage` over a 1.4.1 graph is answered from the graph; over a 1.4.0 graph it still says why it cannot be. - The three summary lines from Task 3 Step 1 are in the PR body, with the 7688 run named as the back-compat gate. - No public accessor changed name, signature or return type (`tests/test_public_surface.py` and the contract tests pass unchanged). - PR against `release/2.0`, `Closes #327`. From 6b829cf4e50dde8be9c98111ad742676e13d11fc Mon Sep 17 00:00:00 2001 From: Rahul Krishna Date: Sun, 6 Sep 2026 01:49:11 -0400 Subject: [PATCH 04/13] fix(python): scope every Neo4j statement on the can:// id prefix, not _module codeanalyzer-python 1.4.1 drops the _module property; scope now comes from the id the analyzer already mints (can://python//...). One prefix replaces 49 predicates. Ghosts fall inside the prefix, so every walk pins its source to :PyCallable by label to keep the reached-never-traversed rule. The probe reads analyzer_version: below 1.4.0 refuses, 1.4.0 warns once about the missing index. --- cldk/analysis/python/neo4j/neo4j_backend.py | 247 +++++++++++------- tests/analysis/python/conftest.py | 25 +- tests/analysis/python/test_bodies.py | 2 +- tests/analysis/python/test_dataflow.py | 2 +- tests/analysis/python/test_e2e_neo4j_live.py | 116 ++++++-- tests/analysis/python/test_entrypoints.py | 4 +- tests/analysis/python/test_locate.py | 10 +- .../test_neo4j_multi_application_scope.py | 140 +++++++--- .../analysis/python/test_resolution_probe.py | 8 +- tests/analysis/python/test_schema_probe.py | 38 +++ .../analysis/python/test_scoping_keywords.py | 17 +- 11 files changed, 432 insertions(+), 177 deletions(-) diff --git a/cldk/analysis/python/neo4j/neo4j_backend.py b/cldk/analysis/python/neo4j/neo4j_backend.py index 939392d7..faad6923 100644 --- a/cldk/analysis/python/neo4j/neo4j_backend.py +++ b/cldk/analysis/python/neo4j/neo4j_backend.py @@ -77,13 +77,14 @@ from __future__ import annotations import logging +import re from collections import defaultdict from contextlib import contextmanager from typing import Any, Dict, List, Sequence, Tuple import networkx as nx from codeanalyzer.schema import model_dump_json -from codeanalyzer.schema.ids import application_id +from codeanalyzer.schema.ids import application_id, module_id from cldk.analysis.commons.resolve import CallableCandidate, body_node_kind, resolve_callable_signature, resolve_value_name, resolve_within, value_candidate from cldk.analysis.commons.results import CallableRef, Diagnostic, EdgePage, EntrypointCoverage, FlowPath, FlowPaths, LocateResult, ModuleRef, PathHop, Slice, SliceNode, TypeRef @@ -143,6 +144,13 @@ logger = logging.getLogger(__name__) + +def _semver(raw: Any) -> Tuple[int, int, int] | None: + """``"1.4.1"`` (or ``"1.4.1.post0"``) as ``(1, 4, 1)``; ``None`` for anything that does not + start with three dotted integers, so an unparsable version is *unknown*, never silently zero.""" + m = re.match(r"(\d+)\.(\d+)\.(\d+)", raw) if isinstance(raw, str) else None + return (int(m[1]), int(m[2]), int(m[3])) if m else None + # One statement per parent->child collection, each fetching that whole collection for the *entire* # application in a single round trip and returning the parent's key as ``pk``. These are the bulk # twins of the per-parent statements inlined in ``PyNeo4jBackend._callable_full`` / ``_class_full`` @@ -171,18 +179,18 @@ "RETURN m.file_key AS pk, pkg.name AS module, e.imported_names AS names" ), # class -> its members - "class_methods": "MATCH (c:PyClass)-[:PY_HAS_METHOD]->(m:PyCallable) WHERE c._module IN $mods RETURN c.signature AS pk, properties(m) AS p", - "class_attributes": "MATCH (c:PyClass)-[:PY_HAS_ATTRIBUTE]->(a:PyAttribute) WHERE c._module IN $mods RETURN c.signature AS pk, properties(a) AS p", - "class_inner_classes": "MATCH (c:PyClass)-[:PY_DECLARES]->(ic:PyClass) WHERE c._module IN $mods RETURN c.signature AS pk, properties(ic) AS p", + "class_methods": "MATCH (c:PyClass)-[:PY_HAS_METHOD]->(m:PyCallable) WHERE any(p IN $prefixes WHERE c.id STARTS WITH p) RETURN c.signature AS pk, properties(m) AS p", + "class_attributes": "MATCH (c:PyClass)-[:PY_HAS_ATTRIBUTE]->(a:PyAttribute) WHERE any(p IN $prefixes WHERE c.id STARTS WITH p) RETURN c.signature AS pk, properties(a) AS p", + "class_inner_classes": "MATCH (c:PyClass)-[:PY_DECLARES]->(ic:PyClass) WHERE any(p IN $prefixes WHERE c.id STARTS WITH p) RETURN c.signature AS pk, properties(ic) AS p", # callable -> its body and nested declarations "callable_callsites": ( - "MATCH (f:PyCallable)-[:PY_HAS_BODY_NODE]->(s:PyBodyNode {kind: 'call'}) WHERE f._module IN $mods " + "MATCH (f:PyCallable)-[:PY_HAS_BODY_NODE]->(s:PyBodyNode {kind: 'call'}) WHERE any(p IN $prefixes WHERE f.id STARTS WITH p) " "RETURN f.signature AS pk, properties(s) AS p ORDER BY s.start_line" ), - "callable_inner_callables": "MATCH (f:PyCallable)-[:PY_DECLARES]->(d:PyCallable) WHERE f._module IN $mods RETURN f.signature AS pk, properties(d) AS p", - "callable_inner_classes": "MATCH (f:PyCallable)-[:PY_DECLARES]->(d:PyClass) WHERE f._module IN $mods RETURN f.signature AS pk, properties(d) AS p", + "callable_inner_callables": "MATCH (f:PyCallable)-[:PY_DECLARES]->(d:PyCallable) WHERE any(p IN $prefixes WHERE f.id STARTS WITH p) RETURN f.signature AS pk, properties(d) AS p", + "callable_inner_classes": "MATCH (f:PyCallable)-[:PY_DECLARES]->(d:PyClass) WHERE any(p IN $prefixes WHERE f.id STARTS WITH p) RETURN f.signature AS pk, properties(d) AS p", "callable_variables": ( - "MATCH (f:PyCallable)-[:PY_DECLARES_VAR]->(v:PyVariable) WHERE f._module IN $mods " + "MATCH (f:PyCallable)-[:PY_DECLARES_VAR]->(v:PyVariable) WHERE any(p IN $prefixes WHERE f.id STARTS WITH p) " "RETURN f.signature AS pk, properties(v) AS p ORDER BY v.start_line, v.name" ), } @@ -326,6 +334,13 @@ def _init_with_driver(self, driver: Any, *, application_name: str | None = None, # (:meth:`_bounded_call_rows`), so an older server keeps serving every other accessor. _QUANTIFIED_PATH_MIN_SERVER = (5, 9) + #: The oldest codeanalyzer-python whose graph this backend serves. 1.4.0 introduced the + #: ``can://`` id grammar every statement here scopes on; 1.4.1 dropped the ``_module`` property + #: and added the ``:PyCanNode`` range index on ``id`` that lets a prefix predicate seek. A + #: 1.4.0 graph is therefore served correctly but scanned (see :meth:`_probe_schema`). + _ANALYZER_FLOOR = (1, 4, 0) + _ANALYZER_INDEXED = (1, 4, 1) + def _probe_schema(self) -> None: """Verify the connected graph's vocabulary once, at connection time, and record the server's version while the connection is already open. @@ -347,6 +362,32 @@ def _probe_schema(self) -> None: raise GraphSchemaMismatch(expected=set(self._REQUIRED_RELATIONSHIP_TYPES), found=found, missing=missing) self._server_version = self._read_server_version() + # The analyzer generation that emitted *this application*, from the property it stamps on + # its :PyApplication node. Below the floor the id grammar the scoping relies on does not + # exist and every statement would come back empty; refusing here is what keeps that from + # reading as "no callables". An absent application has no version either, and is refused + # for the same reason. + rows = self._run("MATCH (a:PyApplication {name: $app}) RETURN a.analyzer_version AS v", app=self.application_name) + raw = rows[0].get("v") if rows else None + version = _semver(raw) + floor = ".".join(map(str, self._ANALYZER_FLOOR)) + if version is None or version < self._ANALYZER_FLOOR: + what = f"was emitted by codeanalyzer-python {raw}" if version else (f"reports analyzer_version {raw!r}" if raw else "carries no analyzer_version (no :PyApplication with that name, or one emitted before the property existed)") + raise GraphSchemaMismatch( + expected=set(self._REQUIRED_RELATIONSHIP_TYPES), + found=found, + missing=set(), + message=f"The graph for application {self.application_name!r} {what}; this backend needs a graph emitted by codeanalyzer-python {floor} or newer.", + ) + if version < self._ANALYZER_INDEXED: + logger.warning( + "The graph for application %r was emitted by codeanalyzer-python %s, which carries no :PyCanNode index on id: " + "scoped queries scan rather than seek. Results are identical; re-emit with %s or newer for the index.", + self.application_name, + raw, + ".".join(map(str, self._ANALYZER_INDEXED)), + ) + def _read_server_version(self) -> Tuple[int, ...] | None: """The attached server's version as an int tuple, or ``None`` when it cannot be read. @@ -394,8 +435,8 @@ def _probe_resolution_edges(self) -> bool: This is information, not an error — this never raises the way :meth:`_probe_schema` does. """ rows = self._run( - "MATCH (s:PyBodyNode)-[:PY_RESOLVES_TO]->() WHERE s._module IN $mods RETURN s LIMIT 1", - mods=self._modules, + "MATCH (s:PyBodyNode)-[:PY_RESOLVES_TO]->() WHERE s.id STARTS WITH $prefix RETURN s LIMIT 1", + prefix=self._scope_prefix, ) return bool(rows) @@ -418,6 +459,31 @@ def has_resolution_edges(self) -> bool: #: one module does not prefetch the application's other 77,000 call sites to answer. _prefetch_scope: List[str] | None = None + #: ``_prefetch_scope`` as id prefixes -- what the seven signature-keyed bulk statements scope on + #: (see :meth:`_module_prefixes`). Set alongside it by :meth:`_bulk`. + _prefetch_prefixes: List[str] | None = None + + @property + def _scope_prefix(self) -> str: + """The application scope every statement carries: ``can://python//``. + + Every node the analyzer emits for this application -- module, class, callable, body node + and ``@external`` ghost alike -- has an id under this prefix, and nothing from any other + application does. The trailing slash is load-bearing: without it ``odoo-slim-19`` would + also match ``odoo-slim-19-b``. Derived, not stored, so a backend built through the + ``object.__new__`` seam the unit tests use has it too. + """ + return application_id(self.application_name) + "/" + + def _module_prefixes(self, keys: Sequence[str] | None) -> List[str]: + """Per-module id prefixes for **narrowing** a bulk fetch to a subset of the application's + modules (``get_symbol_table(paths=...)``, ``get_all_classes(module=...)``); ``None`` is the + whole application, i.e. the one prefix :attr:`_scope_prefix`. Scope answers "which + application", narrowing "which of its modules" -- a single application prefix cannot + express the second, and a list of 1,626 module prefixes is a measurably slower way to + express the first.""" + return [self._scope_prefix] if keys is None else [module_id(self.application_name, k) + "/" for k in keys] + def close(self) -> None: """Close the reused session (if any) and the underlying Neo4j driver.""" self._close_session() @@ -502,7 +568,7 @@ def _children(self, bucket: str, key: str, query: str, **params: Any) -> List[Di would trade an N+1 for a much larger constant. """ if self._prefetch is None: - return self._run(query, mods=self._modules, **params) + return self._run(query, mods=self._modules, prefix=self._scope_prefix, **params) index = self._prefetch.get(bucket) if index is None: index = self._prefetch[bucket] = self._collect(bucket) @@ -511,7 +577,7 @@ def _children(self, bucket: str, key: str, query: str, **params: Any) -> List[Di def _collect(self, bucket: str) -> Dict[str, List[Dict[str, Any]]]: """One whole child collection for this application, in one round trip, grouped by ``pk``.""" index: Dict[str, List[Dict[str, Any]]] = defaultdict(list) - for row in self._run(_BULK_CHILD_QUERIES[bucket], mods=self._prefetch_scope): + for row in self._run(_BULK_CHILD_QUERIES[bucket], mods=self._prefetch_scope, prefixes=self._prefetch_prefixes): index[row["pk"]].append(row) return index @@ -527,14 +593,15 @@ def _bulk(self, mods: Sequence[str] | None = None) -> Any: scope* — narrowing inside an already-primed block would serve a half-filled bucket as if it were complete. """ - outer, outer_scope = self._prefetch, self._prefetch_scope + outer, outer_scope, outer_prefixes = self._prefetch, self._prefetch_scope, self._prefetch_prefixes if outer is None: self._prefetch = {} self._prefetch_scope = list(mods) if mods is not None else self._modules + self._prefetch_prefixes = self._module_prefixes(mods) try: yield finally: - self._prefetch, self._prefetch_scope = outer, outer_scope + self._prefetch, self._prefetch_scope, self._prefetch_prefixes = outer, outer_scope, outer_prefixes def _callable_full(self, props: Dict[str, Any]) -> PyCallable: """Rebuild a full :class:`PyCallable` (call sites, inner callables/classes, locals). @@ -556,7 +623,7 @@ def _callable_full(self, props: Dict[str, Any]) -> PyCallable: "callable_callsites", sig, "MATCH (par:PyCallable {signature: $sig})-[:PY_HAS_BODY_NODE]->(s:PyBodyNode {kind: 'call'}) " - "WHERE par._module IN $mods RETURN properties(s) AS p ORDER BY s.start_line", + "WHERE par.id STARTS WITH $prefix RETURN properties(s) AS p ORDER BY s.start_line", sig=sig, ) ] @@ -564,7 +631,7 @@ def _callable_full(self, props: Dict[str, Any]) -> PyCallable: for r in self._children( "callable_inner_callables", sig, - "MATCH (par:PyCallable {signature: $sig})-[:PY_DECLARES]->(d:PyCallable) WHERE par._module IN $mods RETURN properties(d) AS p", + "MATCH (par:PyCallable {signature: $sig})-[:PY_DECLARES]->(d:PyCallable) WHERE par.id STARTS WITH $prefix RETURN properties(d) AS p", sig=sig, ): ic = self._callable_full(r["p"]) @@ -573,7 +640,7 @@ def _callable_full(self, props: Dict[str, Any]) -> PyCallable: for r in self._children( "callable_inner_classes", sig, - "MATCH (par:PyCallable {signature: $sig})-[:PY_DECLARES]->(d:PyClass) WHERE par._module IN $mods RETURN properties(d) AS p", + "MATCH (par:PyCallable {signature: $sig})-[:PY_DECLARES]->(d:PyClass) WHERE par.id STARTS WITH $prefix RETURN properties(d) AS p", sig=sig, ): ic2 = self._class_full(r["p"]) @@ -584,7 +651,7 @@ def _callable_full(self, props: Dict[str, Any]) -> PyCallable: "callable_variables", sig, "MATCH (par:PyCallable {signature: $sig})-[:PY_DECLARES_VAR]->(v:PyVariable) " - "WHERE par._module IN $mods RETURN properties(v) AS p ORDER BY v.start_line, v.name", + "WHERE par.id STARTS WITH $prefix RETURN properties(v) AS p ORDER BY v.start_line, v.name", sig=sig, ) ] @@ -597,7 +664,7 @@ def _class_full(self, props: Dict[str, Any]) -> PyClass: for r in self._children( "class_methods", sig, - "MATCH (par:PyClass {signature: $sig})-[:PY_HAS_METHOD]->(m:PyCallable) WHERE par._module IN $mods RETURN properties(m) AS p", + "MATCH (par:PyClass {signature: $sig})-[:PY_HAS_METHOD]->(m:PyCallable) WHERE par.id STARTS WITH $prefix RETURN properties(m) AS p", sig=sig, ): m = self._callable_full(r["p"]) @@ -606,7 +673,7 @@ def _class_full(self, props: Dict[str, Any]) -> PyClass: for r in self._children( "class_attributes", sig, - "MATCH (par:PyClass {signature: $sig})-[:PY_HAS_ATTRIBUTE]->(a:PyAttribute) WHERE par._module IN $mods RETURN properties(a) AS p", + "MATCH (par:PyClass {signature: $sig})-[:PY_HAS_ATTRIBUTE]->(a:PyAttribute) WHERE par.id STARTS WITH $prefix RETURN properties(a) AS p", sig=sig, ): a = R.attribute(r["p"]) @@ -615,7 +682,7 @@ def _class_full(self, props: Dict[str, Any]) -> PyClass: for r in self._children( "class_inner_classes", sig, - "MATCH (par:PyClass {signature: $sig})-[:PY_DECLARES]->(ic:PyClass) WHERE par._module IN $mods RETURN properties(ic) AS p", + "MATCH (par:PyClass {signature: $sig})-[:PY_DECLARES]->(ic:PyClass) WHERE par.id STARTS WITH $prefix RETURN properties(ic) AS p", sig=sig, ): ic = self._class_full(r["p"]) @@ -686,16 +753,16 @@ def _call_rows(self) -> List[Dict[str, Any]]: ``:PyExternal`` carries no ``signature`` property at all -- only ``id``/``name``/``module``. ``coalesce(t.signature, t.id)`` resolves it to its addressable ``@external`` can-id instead of projecting ``None``, same idiom ``get_callsites_for`` already uses for the identical - situation one screen below. ``s`` is never external here: ``:PyExternal`` carries no - ``_module`` property, so ``s._module IN $mods`` is already false for it -- a call - *originating* at an external ghost exists in the raw graph (5,307 edges on the live Odoo - graph) but is filtered out by this scoping before it ever reaches the RETURN, so ``s`` - needs no coalesce. + situation one screen below. ``s`` is never external here: the pattern pins it to + ``:PyCallable`` by label (a ghost's id sits under the same application prefix, so the + prefix alone would admit it) -- a call *originating* at an external ghost exists in the raw + graph (5,307 edges on the live Odoo graph) but never reaches the RETURN, so ``s`` needs no + coalesce. """ return self._run( - "MATCH (s:PyCallable|PyExternal)-[r:PY_CALLS]->(t:PyCallable|PyExternal) WHERE s._module IN $mods " + "MATCH (s:PyCallable)-[r:PY_CALLS]->(t:PyCallable|PyExternal) WHERE s.id STARTS WITH $prefix " "RETURN s.signature AS src, coalesce(t.signature, t.id) AS tgt, properties(r) AS p", - mods=self._modules, + prefix=self._scope_prefix, ) def _require_quantified_paths(self, accessor: str) -> None: @@ -750,7 +817,8 @@ def _bounded_call_rows(self, roots: List[str], depth: int | None) -> List[Dict[s this one application — a two-hop budget spent walking ghost-to-ghost instead of through the application's own callables — not a hop into a neighbouring application: every ghost id embeds the application name (``can://python/odoo-slim-19/@external/IPython/start_ipython``), - so a ghost is not in fact shared. Requiring ``a._module IN $mods`` of every hop's *source* makes + so a ghost is not in fact shared. Pinning every hop's *source* to ``:PyCallable`` by label + (a ghost's id sits under the same prefix, so the prefix alone would admit it) makes the traversed edge set exactly :meth:`_call_rows`'s, so a ghost is still reached (it is a legitimate callee, and the local backend has it too) but is never traversed *through*, and the two backends agree node-for-node and edge-for-edge. The node labels repeat @@ -766,6 +834,10 @@ def _bounded_call_rows(self, roots: List[str], depth: int | None) -> List[Dict[s :func:`~cldk.analysis.python.backend.call_graph_scope` and re-coerced here, so nothing caller-controlled reaches the statement as text; ``roots`` stays a parameter. + The root anchor is one prefix test: a declared callable and an ``@external`` ghost both + carry an id under this application's prefix, so the ``root._module IS NULL`` arm that once + admitted ghosts (they never had a ``_module``) is gone by construction, not dropped. + The quantifier starts at ``0``, so a root with no outgoing calls still contributes itself. The first ``MATCH`` is therefore also what defines the **domain a root is validated against**: a ``:PyCallable`` this application declares (by signature) or a ``:PyExternal`` @@ -785,15 +857,14 @@ def _bounded_call_rows(self, roots: List[str], depth: int | None) -> List[Dict[s self._require_quantified_paths("get_call_graph(roots=...)") hops = "" if depth is None else str(int(depth)) return self._run( - "MATCH (root:PyCallable|PyExternal) WHERE coalesce(root.signature, root.id) IN $roots " - "AND (root._module IS NULL OR root._module IN $mods) " - f"MATCH (root) ((a:PyCallable|PyExternal)-[:PY_CALLS]->(b:PyCallable|PyExternal) WHERE a._module IN $mods){{0,{hops}}} (n) " + "MATCH (root:PyCallable|PyExternal) WHERE coalesce(root.signature, root.id) IN $roots AND root.id STARTS WITH $prefix " + f"MATCH (root) ((a:PyCallable)-[:PY_CALLS]->(b:PyCallable|PyExternal) WHERE a.id STARTS WITH $prefix){{0,{hops}}} (n) " "WITH collect(DISTINCT n) AS ns " "UNWIND ns AS s " - "OPTIONAL MATCH (s)-[r:PY_CALLS]->(t) WHERE t IN ns AND s._module IN $mods " + "OPTIONAL MATCH (s)-[r:PY_CALLS]->(t) WHERE t IN ns AND s:PyCallable AND s.id STARTS WITH $prefix " "RETURN coalesce(s.signature, s.id) AS src, coalesce(t.signature, t.id) AS tgt, properties(r) AS p", roots=list(roots), - mods=self._modules, + prefix=self._scope_prefix, ) # ===================================================================================== @@ -844,9 +915,9 @@ def get_python_module(self, file_path: str) -> PyModule | None: def get_python_file(self, qualified_class_name: str) -> str | None: # Only top-level classes are in the in-memory _class_to_file map (module.types). rows = self._run( - "MATCH (:PyModule)-[:PY_DECLARES]->(c:PyClass {signature: $sig}) WHERE c._module IN $mods RETURN c._module AS fk LIMIT 1", + "MATCH (:PyModule)-[:PY_DECLARES]->(c:PyClass {signature: $sig}) WHERE c.id STARTS WITH $prefix RETURN c._module AS fk LIMIT 1", sig=qualified_class_name, - mods=self._modules, + prefix=self._scope_prefix, ) return rows[0]["fk"] if rows else None @@ -946,12 +1017,12 @@ def get_all_classes(self, *, module: str | None = None) -> Dict[str, PyClass]: # The statement was already scoped by ``$mods``, so narrowing to one module is just a # shorter list -- and the same list narrows the prefetch, so the seven bulk statements # fetch one module's members instead of the application's. - scope = self._resolve_paths(None if module is None else [module], kind="module") or self._modules + scope = self._resolve_paths(None if module is None else [module], kind="module") result: Dict[str, PyClass] = {} with self._bulk(scope): # every class's members in seven queries, not one per child collection for r in self._run( - "MATCH (:PyModule)-[:PY_DECLARES]->(c:PyClass) WHERE c._module IN $mods RETURN properties(c) AS p", - mods=scope, + "MATCH (:PyModule)-[:PY_DECLARES]->(c:PyClass) WHERE any(p IN $prefixes WHERE c.id STARTS WITH p) RETURN properties(c) AS p", + prefixes=self._module_prefixes(scope), ): c = self._class_full(r["p"]) result[c.signature] = c @@ -960,9 +1031,9 @@ def get_all_classes(self, *, module: str | None = None) -> Dict[str, PyClass]: def get_class(self, qualified_class_name: str) -> PyClass | None: # Top-level classes only, matching get_all_classes().get(...) on the in-memory backend. rows = self._run( - "MATCH (:PyModule)-[:PY_DECLARES]->(c:PyClass {signature: $sig}) WHERE c._module IN $mods RETURN properties(c) AS p LIMIT 1", + "MATCH (:PyModule)-[:PY_DECLARES]->(c:PyClass {signature: $sig}) WHERE c.id STARTS WITH $prefix RETURN properties(c) AS p LIMIT 1", sig=qualified_class_name, - mods=self._modules, + prefix=self._scope_prefix, ) return self._class_full(rows[0]["p"]) if rows else None @@ -1081,16 +1152,16 @@ def get_all_fields(self, qualified_class_name: str) -> List[PyClassAttribute]: def get_callables_overview(self) -> List[PyCallableOverview]: rows = self._run( - "MATCH (c:PyCallable) WHERE c._module IN $mods " + self._OVERVIEW_PROJECTION, - mods=self._modules, + "MATCH (c:PyCallable) WHERE c.id STARTS WITH $prefix " + self._OVERVIEW_PROJECTION, + prefix=self._scope_prefix, ) return [R.overview(r) for r in rows] def get_method_bodies(self, signatures: List[str]) -> Dict[str, str]: rows = self._run( - "MATCH (c:PyCallable) WHERE c._module IN $mods AND c.signature IN $sigs AND c.code IS NOT NULL " + "MATCH (c:PyCallable) WHERE c.id STARTS WITH $prefix AND c.signature IN $sigs AND c.code IS NOT NULL " "RETURN c.signature AS signature, c.code AS code", - mods=self._modules, + prefix=self._scope_prefix, sigs=list(signatures), ) return {r["signature"]: r["code"] for r in rows} @@ -1117,9 +1188,9 @@ def get_source(self, node_id: str) -> str: "for a statement or call site." ) rows = self._run( - "MATCH (c:PyCallable) WHERE c._module IN $mods AND (c.signature = $sig OR c.id = $sig) AND c.code IS NOT NULL " + "MATCH (c:PyCallable) WHERE c.id STARTS WITH $prefix AND (c.signature = $sig OR c.id = $sig) AND c.code IS NOT NULL " "RETURN c.code AS code", - mods=self._modules, + prefix=self._scope_prefix, sig=sig, ) if not rows: @@ -1128,7 +1199,7 @@ def get_source(self, node_id: str) -> str: # -----[ addressing ]----- _RESOLVE_CALLABLE_QUERY = ( - "MATCH (c:PyCallable) WHERE c._module IN $mods AND (c.signature = $name OR c.signature ENDS WITH $dotted) " + "MATCH (c:PyCallable) WHERE c.id STARTS WITH $prefix AND (c.signature = $name OR c.signature ENDS WITH $dotted) " "OPTIONAL MATCH (owner:PyClass)-[:PY_HAS_METHOD]->(c) " "RETURN c.signature AS signature, c.name AS name, c.id AS id, c._module AS path, " "c.start_line AS start_line, owner.signature AS class_signature" @@ -1147,7 +1218,7 @@ def resolve_callable(self, name: str, *, in_class: str | None = None, in_module: against different sets is the defect this construction avoids; running the same predicate twice (once in Cypher, once in the shared policy) is the cheap price of avoiding it. """ - rows = self._run(self._RESOLVE_CALLABLE_QUERY, mods=self._modules, name=name, dotted="." + name) + rows = self._run(self._RESOLVE_CALLABLE_QUERY, prefix=self._scope_prefix, name=name, dotted="." + name) # Two callables sharing a signature would collapse into one entry and resolve arbitrarily; # recorded and raised on only if the name lands on one, so an unrelated duplicate cannot # break every resolution. Not reachable on a real application (15,549 distinct signatures). @@ -1186,10 +1257,10 @@ def resolve_value(self, name: str, *, within: str) -> SliceNode: """ owner = resolve_within(self.resolve_callable, within) rows = self._run( - "MATCH (c:PyCallable) WHERE c._module IN $mods AND c.signature = $sig " + "MATCH (c:PyCallable) WHERE c.id STARTS WITH $prefix AND c.signature = $sig " "MATCH (c)-[:PY_HAS_BODY_NODE]->(b:PyBodyNode) WHERE b.kind = 'formal_in' " "RETURN b.var AS var, b.id AS id", - mods=self._modules, + prefix=self._scope_prefix, sig=owner.callable, ) # A list, not a dict keyed by name: two values resolving to the same name are a genuine @@ -1216,7 +1287,7 @@ def resolve_value(self, name: str, *, within: str) -> SliceNode: #: on odoo-slim-19), but a graph built some other way must not be able to widen the answer #: silently. _OWN_EDGES = ( - "MATCH (c:PyCallable) WHERE c._module IN $mods AND c.signature = $sig " + "MATCH (c:PyCallable) WHERE c.id STARTS WITH $prefix AND c.signature = $sig " "MATCH (c)-[:PY_HAS_BODY_NODE]->(s:PyBodyNode)-[r:{rel}]->(d:PyBodyNode)<-[:PY_HAS_BODY_NODE]-(c) " ) @@ -1251,7 +1322,7 @@ def _own_edges(self, name: str, in_class: str | None, rel: str, projection: str, check_page_size(page_size) sig = self.resolve_callable(name, in_class=in_class).callable match = self._OWN_EDGES.format(rel=rel) - params: Dict[str, Any] = {"mods": self._modules, "sig": sig} + params: Dict[str, Any] = {"prefix": self._scope_prefix, "sig": sig} total = self._run(match + "RETURN count(r) AS total", **params)[0]["total"] where = f"WHERE {keyset_where(order.exprs)} " if cursor is not None else "" rows = self._run( @@ -1330,7 +1401,7 @@ def _slice(self, src: str, within: str, depth: int | None, max_nodes: int, *, ba **Not scoped by ``_module``,** unlike the per-callable accessors. A body-node id is stamped with its application (``can://python//…``) and the emitter only ever links nodes from its own run, so the traversal cannot leave the application it started in; adding - ``m._module IN $mods`` would cost a list membership test on every one of 195,784 reached + ``m.id STARTS WITH $prefix`` would cost a string-prefix test on every one of 195,784 reached nodes to re-establish something the ids already guarantee. The seed is app-scoped by :meth:`resolve_value`, and a live test checks the reached set against this application's module keys rather than taking the argument on trust. @@ -1369,8 +1440,8 @@ def slice_forward(self, src: str, *, within: str, depth: int | None = DEFAULT_DE #: form, and it still plans as a pruning expansion (``WITH DISTINCT m`` keeps it one); needs #: Neo4j 5.9+ like ``roots=`` does. The self-question ``reaches(x, x)`` still terminates. _REACHES = ( - "MATCH (a:PyCallable {{signature:$a}}) WHERE a._module IN $mods " - "MATCH (a) ((x:PyCallable)-[:PY_CALLS]->(y:PyCallable) WHERE x._module IN $mods){{1,{depth}}} (m:PyCallable) " + "MATCH (a:PyCallable {{signature:$a}}) WHERE a.id STARTS WITH $prefix " + "MATCH (a) ((x:PyCallable)-[:PY_CALLS]->(y:PyCallable) WHERE x.id STARTS WITH $prefix){{1,{depth}}} (m:PyCallable) " "WITH DISTINCT m WHERE m.signature = $b RETURN count(m) > 0 AS ok" ) @@ -1380,7 +1451,7 @@ def reaches(self, src: str, dst: str, *, depth: int | None = None) -> bool: self._require_quantified_paths("reaches") a = self.resolve_callable(src).callable b = self.resolve_callable(dst).callable - return bool(self._run(self._REACHES.format(depth="" if depth is None else depth), a=a, b=b, mods=self._modules)[0]["ok"]) + return bool(self._run(self._REACHES.format(depth="" if depth is None else depth), a=a, b=b, prefix=self._scope_prefix)[0]["ok"]) #: ``{0,}`` again, so a sink with no callers is its own cone rather than an empty answer that #: a caller could not tell from "this name is wrong" (D7). Every hop is labelled and scoped @@ -1388,8 +1459,8 @@ def reaches(self, src: str, dst: str, *, depth: int | None = None) -> bool: #: 0.27s against 0.08s. Properties are projected into maps *before* the cap so only ``$cap`` of #: them cross the wire. _CONE = ( - "MATCH (s:PyCallable) WHERE s.signature IN $sigs AND s._module IN $mods " - "MATCH (s) (()<-[:PY_CALLS]-(x:PyCallable) WHERE x._module IN $mods){{0,{depth}}} (m:PyCallable) " + "MATCH (s:PyCallable) WHERE s.signature IN $sigs AND s.id STARTS WITH $prefix " + "MATCH (s) (()<-[:PY_CALLS]-(x:PyCallable) WHERE x.id STARTS WITH $prefix){{0,{depth}}} (m:PyCallable) " "WITH DISTINCT m ORDER BY m.id " "WITH collect({{callable: m.signature, name: m.name, ref: m.id, file: m._module, line: m.start_line}}) AS found " "RETURN size(found) AS total, found[0..$cap] AS page" @@ -1401,27 +1472,27 @@ def backward_cone(self, sinks: Sequence[str], *, depth: int | None = DEFAULT_DEP check_max_nodes(max_nodes) self._require_quantified_paths("backward_cone") roots = cone_sinks(self.resolve_callable, sinks) - row = self._run(self._CONE.format(depth="" if depth is None else depth), sigs=[r.callable for r in roots], cap=max_nodes, mods=self._modules)[0] + row = self._run(self._CONE.format(depth="" if depth is None else depth), sigs=[r.callable for r in roots], cap=max_nodes, prefix=self._scope_prefix)[0] nodes = [SliceNode(file=n["file"], line=n["line"], callable=n["callable"], kind="callable", name=n["name"], source=None, ref=n["ref"]) for n in row["page"]] return Slice(nodes=nodes, roots=roots, resolved=slice_resolved(roots), total=row["total"]) #: ``t`` may be a ``:PyExternal`` ghost, which carries ``module``/``name``/``id`` and no #: ``signature``, ``_module`` or ``start_line`` -- so the projection names each property #: explicitly and :func:`_call_neighbour` decides what a row means from whether ``signature`` - #: came back. ``s._module IN $mods`` scopes the *caller* side; it is already false for an - #: external, which is how a call originating at a ghost stays out of ``callers_of``. - _CALLERS = "MATCH (s:PyCallable)-[:PY_CALLS]->(t:PyCallable {signature: $sig}) WHERE s._module IN $mods RETURN s.signature AS signature, s.name AS name, s.id AS ref, s._module AS file, s.start_line AS line, s.module AS module" - _CALLEES = "MATCH (s:PyCallable {signature: $sig})-[:PY_CALLS]->(t:PyCallable|PyExternal) WHERE s._module IN $mods RETURN t.signature AS signature, t.name AS name, t.id AS ref, t._module AS file, t.start_line AS line, t.module AS module" + #: came back. ``(s:PyCallable)`` pins the *caller* side by label -- a ghost's id sits under + #: the same prefix -- which is how a call originating at a ghost stays out of ``callers_of``. + _CALLERS = "MATCH (s:PyCallable)-[:PY_CALLS]->(t:PyCallable {signature: $sig}) WHERE s.id STARTS WITH $prefix RETURN s.signature AS signature, s.name AS name, s.id AS ref, s._module AS file, s.start_line AS line, s.module AS module" + _CALLEES = "MATCH (s:PyCallable {signature: $sig})-[:PY_CALLS]->(t:PyCallable|PyExternal) WHERE s.id STARTS WITH $prefix RETURN t.signature AS signature, t.name AS name, t.id AS ref, t._module AS file, t.start_line AS line, t.module AS module" def callers_of(self, name: str, *, in_class: str | None = None, in_module: str | None = None) -> List[SliceNode]: """Who calls this (see :meth:`PythonAnalysisBackend.callers_of`).""" sig = self.resolve_callable(name, in_class=in_class, in_module=in_module).callable - return [_call_neighbour(r) for r in self._run(self._CALLERS, sig=sig, mods=self._modules)] + return [_call_neighbour(r) for r in self._run(self._CALLERS, sig=sig, prefix=self._scope_prefix)] def callees_of(self, name: str, *, in_class: str | None = None, in_module: str | None = None) -> List[SliceNode]: """What this calls, externals included (see :meth:`PythonAnalysisBackend.callees_of`).""" sig = self.resolve_callable(name, in_class=in_class, in_module=in_module).callable - return [_call_neighbour(r) for r in self._run(self._CALLEES, sig=sig, mods=self._modules)] + return [_call_neighbour(r) for r in self._run(self._CALLEES, sig=sig, prefix=self._scope_prefix)] # -----[ paths, mixed queries, hydration ]----- #: The caller's word for a hop, computed in Cypher so the ORDER BY below sorts by the same @@ -1476,8 +1547,8 @@ def callees_of(self, name: str, *, in_class: str | None = None, in_module: str | #: the nodes still project to :func:`_call_neighbour`'s row shape so nothing could leak a #: ``can://`` id even if that changed. _CALL_PATHS = ( - "MATCH (a:PyCallable {{signature:$src}}) WHERE a._module IN $mods " - "MATCH (b:PyCallable {{signature:$dst}}) WHERE b._module IN $mods " + "MATCH (a:PyCallable {{signature:$src}}) WHERE a.id STARTS WITH $prefix " + "MATCH (b:PyCallable {{signature:$dst}}) WHERE b.id STARTS WITH $prefix " "MATCH p = allShortestPaths((a)-[:PY_CALLS*1..{depth}]->(b)) WHERE all(n IN nodes(p) WHERE n:PyCallable) " "WITH p, " + _PATH_ORDER + " AS key ORDER BY length(p), key LIMIT $cap " "RETURN [n IN nodes(p) | {{signature: n.signature, name: n.name, ref: n.id, file: n._module, " @@ -1491,7 +1562,7 @@ def _paths(self, query: str, node_of, a: SliceNode, b: SliceNode, *, src: str, d ``a``/``b`` are the resolved endpoints (for the self-question's message); ``src``/``dst`` are the keys the query matches them by.""" check_distinct_endpoints(a, b) - rows = self._run(query.format(rels=SDG_REL_PATTERN, depth="" if depth is None else depth), src=src, dst=dst, cap=max_paths + 1, mods=self._modules) + rows = self._run(query.format(rels=SDG_REL_PATTERN, depth="" if depth is None else depth), src=src, dst=dst, cap=max_paths + 1, prefix=self._scope_prefix) paths = [flow_path([node_of(n) for n in r["ns"]], [(e["via"], e["var"], e["prov"]) for e in r["rs"]]) for r in rows[:max_paths]] return FlowPaths(paths=paths, complete=len(rows) <= max_paths) @@ -1522,7 +1593,7 @@ def call_paths_between(self, src: str, dst: str, *, depth: int | None = None, ma #: Every value that *enters* ``$sig`` -- its parameters, and the globals and captures it reads. #: Scoped, because a signature is not application-stamped the way an id is. - _CALLEE_VALUES = "MATCH (c:PyCallable {signature:$sig})-[:PY_HAS_BODY_NODE]->(b:PyBodyNode {kind:'formal_in'}) WHERE c._module IN $mods RETURN collect(b.id) AS ids" + _CALLEE_VALUES = "MATCH (c:PyCallable {signature:$sig})-[:PY_HAS_BODY_NODE]->(b:PyBodyNode {kind:'formal_in'}) WHERE c.id STARTS WITH $prefix RETURN collect(b.id) AS ids" def _value_reaches(self, src: str, dsts: List[str], depth: int | None) -> bool: """Does the value at ``src`` reach any of ``dsts``? The one predicate both mixed queries @@ -1538,7 +1609,7 @@ def flows_to_call(self, src: str, callee: str, *, within: str, depth: int | None check_depth(depth) root = self.resolve_value(src, within=within) sig = self.resolve_callable(callee).callable - return self._value_reaches(root.ref, self._run(self._CALLEE_VALUES, sig=sig, mods=self._modules)[0]["ids"], depth) + return self._value_reaches(root.ref, self._run(self._CALLEE_VALUES, sig=sig, prefix=self._scope_prefix)[0]["ids"], depth) def flows_to_argument(self, src: str, callee: str, arg: str, *, within: str, depth: int | None = None) -> bool: """Does this value reach ``callee``'s ``arg`` @@ -1556,7 +1627,7 @@ def flows_to_argument(self, src: str, callee: str, arg: str, *, within: str, dep #: the callable arm is ``_module``-scoped: a body-node id and a ghost id both embed the #: application, while a signature does not. _SOURCES = ( - "MATCH (c:PyCallable) WHERE c._module IN $mods AND (c.id IN $refs OR c.signature IN $refs) " + "MATCH (c:PyCallable) WHERE c.id STARTS WITH $prefix AND (c.id IN $refs OR c.signature IN $refs) " "RETURN c.id AS id, c.signature AS sig, c.code AS code " "UNION MATCH (b:PyBodyNode) WHERE b.id IN $refs RETURN b.id AS id, null AS sig, null AS code " "UNION MATCH (e:PyExternal) WHERE e.id IN $refs RETURN e.id AS id, null AS sig, null AS code" @@ -1566,7 +1637,7 @@ def _sources_for(self, refs: Sequence[str]) -> Dict[str, "str | None"]: """Source text for every ref this graph holds (see :meth:`PythonAnalysisBackend._sources_for`).""" wanted = set(refs) found: Dict[str, "str | None"] = {} - for row in self._run(self._SOURCES, mods=self._modules, refs=list(wanted)): + for row in self._run(self._SOURCES, prefix=self._scope_prefix, refs=list(wanted)): # A callable answers to both of its names, exactly as ``get_source`` accepts either -- # a ``SliceNode.ref`` is the ``can://`` id, but a caller holding a signature must not # get "names nothing" for a callable that plainly exists. @@ -1577,26 +1648,26 @@ def _sources_for(self, refs: Sequence[str]) -> Dict[str, "str | None"]: def get_decorated_callables(self, markers: List[str]) -> List[PyCallableOverview]: rows = self._run( - "MATCH (c:PyCallable) WHERE c._module IN $mods " + "MATCH (c:PyCallable) WHERE c.id STARTS WITH $prefix " "AND any(d IN c.decorators WHERE d IN $markers) " + self._OVERVIEW_PROJECTION, - mods=self._modules, + prefix=self._scope_prefix, markers=list(markers), ) return [R.overview(r) for r in rows] def get_entrypoints(self) -> List[PyCallableOverview]: rows = self._run( - "MATCH (c:PyCallable) WHERE c._module IN $mods AND c.is_entrypoint = true " + self._OVERVIEW_PROJECTION, - mods=self._modules, + "MATCH (c:PyCallable) WHERE c.id STARTS WITH $prefix AND c.is_entrypoint = true " + self._OVERVIEW_PROJECTION, + prefix=self._scope_prefix, ) return [R.overview(r) for r in rows] def get_entrypoint_classes(self) -> List[PyClassOverview]: rows = self._run( - "MATCH (cl:PyClass) WHERE cl._module IN $mods AND cl.is_entrypoint = true " + "MATCH (cl:PyClass) WHERE cl.id STARTS WITH $prefix AND cl.is_entrypoint = true " "RETURN cl.signature AS signature, cl.name AS name, cl.decorators AS decorators, " "cl._module AS path, cl.start_line AS start_line, cl.end_line AS end_line", - mods=self._modules, + prefix=self._scope_prefix, ) return [R.class_overview(r) for r in rows] @@ -1634,12 +1705,12 @@ def get_callsites_for(self, signatures: List[str]) -> Dict[str, List[PyCallsite] # unresolved, or when the graph was populated at an analysis level below the one where the # defuse-linker backfill runs (see that same docstring's caveat). rows = self._run( - "MATCH (c:PyCallable) WHERE c._module IN $mods AND c.signature IN $sigs " + "MATCH (c:PyCallable) WHERE c.id STARTS WITH $prefix AND c.signature IN $sigs " "OPTIONAL MATCH (c)-[:PY_HAS_BODY_NODE]->(s:PyBodyNode {kind: 'call'}) " "OPTIONAL MATCH (s)-[:PY_RESOLVES_TO]->(t) " "RETURN c.signature AS owner, properties(s) AS p, coalesce(t.signature, t.id) AS callee " "ORDER BY s.start_line", - mods=self._modules, + prefix=self._scope_prefix, sigs=list(signatures), ) out: Dict[str, List[PyCallsite]] = {} @@ -1712,11 +1783,11 @@ def get_config_uses(self, key: str | None = None) -> List[PyConfigUseEdge]: # PY_USES_CONFIG (the one prefixed edge in this layer, per the leg-1 brief) connects # (:PyBodyNode)-->(:ConfigKey) directly, so its endpoints ARE src/dst -- no reconstruction # helper needed, unlike artifact()/dependency()/config_key() above. Scoped like every other - # body-node query in this file (`bn._module IN $mods`), not via the Artifact/ConfigKey path, + # body-node query in this file (`bn.id STARTS WITH $prefix`), not via the Artifact/ConfigKey path, # since a config key can be read from a module outside this application's declared modules # only if it were mis-scoped -- $mods is the same guard get_method_bodies/_call_rows use. - query = "MATCH (bn:PyBodyNode)-[u:PY_USES_CONFIG]->(ck:ConfigKey) WHERE bn._module IN $mods" - params: Dict[str, Any] = {"mods": self._modules} + query = "MATCH (bn:PyBodyNode)-[u:PY_USES_CONFIG]->(ck:ConfigKey) WHERE bn.id STARTS WITH $prefix" + params: Dict[str, Any] = {"prefix": self._scope_prefix} if key is not None: query += " AND ck.key = $key" params["key"] = key @@ -1752,13 +1823,13 @@ def get_config_readers(self, key: str) -> List[PyCallableOverview]: # PyBodyNode a PY_USES_CONFIG edge points at is guaranteed to already have that edge to its # owner. DISTINCT because one callable can read the same key at several call sites. rows = self._run( - "MATCH (bn:PyBodyNode)-[:PY_USES_CONFIG]->(ck:ConfigKey) WHERE bn._module IN $mods AND ck.key = $key " + "MATCH (bn:PyBodyNode)-[:PY_USES_CONFIG]->(ck:ConfigKey) WHERE bn.id STARTS WITH $prefix AND ck.key = $key " "MATCH (reader:PyCallable)-[:PY_HAS_BODY_NODE]->(bn) " "OPTIONAL MATCH (cls:PyClass)-[:PY_HAS_METHOD]->(reader) " "RETURN DISTINCT reader.signature AS signature, reader.name AS name, reader.decorators AS decorators, " "reader._module AS path, reader.start_line AS start_line, reader.end_line AS end_line, " "cls.signature AS class_signature", - mods=self._modules, + prefix=self._scope_prefix, key=key, ) return [R.overview(r) for r in rows] @@ -1788,7 +1859,7 @@ def get_config_readers(self, key: str) -> List[PyCallableOverview]: "OPTIONAL MATCH (:PyApplication {name: $app})-[:PY_HAS_MODULE]->(m:PyModule {file_key: pos.path}) " "WITH pos, m " "OPTIONAL MATCH (c:PyCallable {_module: pos.path}) " - "WHERE c._module IN $mods " + "WHERE c.id STARTS WITH $prefix " "AND c.start_line IS NOT NULL AND c.end_line IS NOT NULL " "AND c.start_line <= pos.line AND pos.line <= c.end_line " "WITH pos, m, c " @@ -1937,7 +2008,7 @@ def locate_many(self, positions: Sequence[Tuple[str, int]]) -> List[LocateResult rows = self._run( self._LOCATE_QUERY, app=self.application_name, - mods=self._modules, + prefix=self._scope_prefix, positions=[{"idx": i, "path": key, "line": line} for i, (key, (_, line)) in enumerate(zip(keys, positions))], ) by_idx: Dict[int, List[Dict[str, Any]]] = defaultdict(list) diff --git a/tests/analysis/python/conftest.py b/tests/analysis/python/conftest.py index c56212d5..1ff7d24e 100644 --- a/tests/analysis/python/conftest.py +++ b/tests/analysis/python/conftest.py @@ -100,6 +100,8 @@ def run(self, query: str, **params: Any) -> list[_FakeRecord]: self._driver.statements.append(query) if "db.relationshipTypes" in query: return [_FakeRecord({"relationshipType": rt}) for rt in self._driver.rel_types] + if "RETURN a.analyzer_version AS v" in query: + return [] if self._driver.analyzer_version is None else [_FakeRecord({"v": self._driver.analyzer_version})] if self._driver.responder is not None: return [_FakeRecord(d) for d in self._driver.responder(query, params)] return [] @@ -109,8 +111,9 @@ def close(self) -> None: class FakeDriver: - """Stub Neo4j driver with a settable ``rel_types`` set, standing in for a real - ``neo4j.GraphDatabase`` driver in unit tests. + """Stub Neo4j driver with a settable ``rel_types`` set and ``analyzer_version`` (``None`` + means "no :PyApplication of that name"), standing in for a real ``neo4j.GraphDatabase`` + driver in unit tests. ``responder``, when set, answers every statement that isn't the schema probe — a tiny in-memory Cypher stub (query text, params) -> rows, the same shape as the ad hoc @@ -122,8 +125,10 @@ def __init__( self, rel_types: set[str] | frozenset[str] = _V2_RELATIONSHIP_TYPES, responder: Callable[[str, dict], list[dict]] | None = None, + analyzer_version: str | None = "1.4.1", ) -> None: self.rel_types: set[str] = set(rel_types) + self.analyzer_version = analyzer_version self.statements: list[str] = [] self.responder = responder @@ -353,7 +358,7 @@ def _locate_callable_id(signature: str) -> str: """A stand-in for the analyzer's ``can://`` id. Only its shape matters here: a body node's graph ``id`` is ``@``, which is what the Neo4j backend's innermost-node tie break splits on.""" - return f"can://{_LOCATE_MODULE_PATH}#{signature}" + return f"can://python/app/{_LOCATE_MODULE_PATH}/{signature}" # -----[ the local backend's view: a real in-memory PyApplication ]----- @@ -462,17 +467,18 @@ def _locate_responder(query: str, params: dict) -> list[dict]: ``get_source`` for the ``py`` fixture. This evaluates the query's WHERE clauses rather than short-circuiting them, so the assertions it - backs are about the query the backend actually sends: ``$mods`` really gates the callable match - (drop the application's module keys and the callables disappear), a callable row repeats once per - matching body node, and every containment decision is made on lines the way Cypher would. + backs are about the query the backend actually sends: ``$prefix`` really gates the callable + match (attach as another application and the callables disappear), a callable row repeats once + per matching body node, and every containment decision is made on lines the way Cypher would. """ + in_scope = "prefix" in params and _locate_callable_id("").startswith(params["prefix"]) if "RETURN m.file_key AS k" in query: return [{"k": _LOCATE_MODULE_PATH}] if "c.code IS NOT NULL" in query: # get_method_bodies (bulk, "c.signature IN $sigs") and get_source (single, "c.signature = - # $sig") both gate on $mods and on a real code property -- a bodyless callable (has_span + # $sig") both gate on $prefix and on a real code property -- a bodyless callable (has_span # False, e.g. Store.stub) has none, so it is simply absent from the rows, never "". - if _LOCATE_MODULE_PATH not in params.get("mods", []): + if not in_scope: return [] if "IN $sigs" in query: return [ @@ -489,8 +495,7 @@ def _locate_responder(query: str, params: dict) -> list[dict]: if pos["path"] != _LOCATE_MODULE_PATH: rows.append(_locate_row(pos["idx"])) # no :PyModule for this file_key continue - # ``OPTIONAL MATCH (c:PyCallable {_module: pos.path}) WHERE c._module IN $mods AND ...`` - in_scope = _LOCATE_MODULE_PATH in params["mods"] + # ``OPTIONAL MATCH (c:PyCallable {_module: pos.path}) WHERE c.id STARTS WITH $prefix AND ...`` matches = [c for c in _LOCATE_CALLABLE_SPECS if in_scope and c["start_line"] <= pos["line"] <= c["end_line"]] if not matches: rows.append(_locate_row(pos["idx"], _LOCATE_MODULE_PROPS)) diff --git a/tests/analysis/python/test_bodies.py b/tests/analysis/python/test_bodies.py index 79ede2bc..814443b4 100644 --- a/tests/analysis/python/test_bodies.py +++ b/tests/analysis/python/test_bodies.py @@ -86,7 +86,7 @@ def test_get_source_of_a_body_node_local(py_local): ``signature`` produced a string that named nothing in the graph. """ r = py_local.locate("src/app.py", 11) - assert r.node_id == "can://src/app.py#src.app.Store.Meta.tag@11:12" + assert r.node_id == "can://python/app/src/app.py/src.app.Store.Meta.tag@11:12" assert py_local.get_source(r.node_id) == _locate_code(11, 11) assert "café" in py_local.get_source(r.node_id) diff --git a/tests/analysis/python/test_dataflow.py b/tests/analysis/python/test_dataflow.py index 82f3cff4..5880c70d 100644 --- a/tests/analysis/python/test_dataflow.py +++ b/tests/analysis/python/test_dataflow.py @@ -1461,7 +1461,7 @@ def test_local_locate_speaks_the_repo_relative_path(slice_l4): def test_slice_node_kinds_are_the_graphs_vocabulary_translated(live_analysis): """``SliceNode.KINDS`` is pinned against the graph's own ``:PyBodyNode.kind`` values, each translated the way both backends translate it, plus the two call-graph kinds.""" - graph_kinds = {r["k"] for r in live_analysis.backend._run("MATCH (b:PyBodyNode) WHERE b._module IN $mods RETURN DISTINCT b.kind AS k", mods=live_analysis.backend._modules)} + graph_kinds = {r["k"] for r in live_analysis.backend._run("MATCH (b:PyBodyNode) WHERE b.id STARTS WITH $prefix RETURN DISTINCT b.kind AS k", prefix=live_analysis.backend._scope_prefix)} assert graph_kinds, "no body nodes?" translated = set() for kind in graph_kinds: diff --git a/tests/analysis/python/test_e2e_neo4j_live.py b/tests/analysis/python/test_e2e_neo4j_live.py index 9df45142..ac95681e 100644 --- a/tests/analysis/python/test_e2e_neo4j_live.py +++ b/tests/analysis/python/test_e2e_neo4j_live.py @@ -89,6 +89,9 @@ NEO4J_USER = os.environ.get("CLDK_TEST_NEO4J_USER", "neo4j") NEO4J_PASSWORD = os.environ.get("CLDK_TEST_NEO4J_PASSWORD", "neo4j") APP_NAME = os.environ.get("CLDK_TEST_NEO4J_APP", "odoo-slim-19") +#: The application's id prefix -- the scope every SDK statement carries, and the scope this +#: suite's own fixture-derivation Cypher carries too, since a 1.4.1 graph has no ``_module``. +APP_PREFIX = f"can://python/{APP_NAME}/" # The four relationship types PyNeo4jBackend._probe_schema insists on. Duplicated here on purpose: # a test that imports the constant it is checking cannot catch the constant changing. @@ -195,7 +198,7 @@ def sample(cypher) -> Dict[str, Any]: """ MATCH (m:PyModule) WITH m.module_name AS mn, collect(m.file_key) AS ks WHERE size(ks) = 1 WITH mn, ks[0] AS path - MATCH (any:PyCallable {_module: path}) WITH mn, path, min(any.start_line) AS first_def + MATCH (any:PyCallable) WHERE any.id STARTS WITH $prefix + path + '/' WITH mn, path, min(any.start_line) AS first_def WHERE first_def > 2 MATCH (:PyModule {file_key: path})-[:PY_DECLARES]->(c:PyCallable)-[:PY_HAS_BODY_NODE]->(b:PyBodyNode) WHERE c.code IS NOT NULL AND c.end_line - c.start_line >= 3 @@ -205,7 +208,8 @@ def sample(cypher) -> Dict[str, Any]: mn AS module_name, path AS module_path, c.start_line AS start_line, c.end_line AS end_line, call_line, first_def ORDER BY c.signature LIMIT 1 - """ + """, + prefix=APP_PREFIX, ) assert rows, "no top-level module function in the graph satisfies the fixture constraints" picked = dict(rows[0]) @@ -349,7 +353,8 @@ def test_overview_path_joins_locate_and_class_overview(analysis, sample, cypher) # callable overview draws from. (Not a subset check against the callable paths: a module can # declare a class and no callable, and 66 of this application's 1,157 class-bearing modules do.) module_keys = {r["k"] for r in cypher("MATCH (:PyApplication {name: $n})-[:PY_HAS_MODULE]->(m:PyModule) RETURN m.file_key AS k", n=APP_NAME)} - class_paths = {r["p"] for r in cypher("MATCH (cl:PyClass) WHERE cl._module IN $m RETURN DISTINCT cl._module AS p", m=list(module_keys))} + # Task 2: the ``cl._module AS p`` projection becomes a key derived from ``cl.id``. + class_paths = {r["p"] for r in cypher("MATCH (cl:PyClass) WHERE cl.id STARTS WITH $prefix RETURN DISTINCT cl._module AS p", prefix=APP_PREFIX)} assert class_paths, "no classes in the graph" assert row.path in module_keys and class_paths <= module_keys, "class and callable overviews disagree on path spelling" @@ -534,10 +539,11 @@ def test_locate_many_agrees_with_locate_position_by_position(analysis, sample, c """``locate_many`` is not allowed to drift from ``locate``, nor to reorder its results.""" others = cypher( """ - MATCH (c:PyCallable) WHERE c.code IS NOT NULL AND c.start_line IS NOT NULL AND c.end_line > c.start_line + MATCH (c:PyCallable) WHERE c.id STARTS WITH $prefix AND c.code IS NOT NULL AND c.start_line IS NOT NULL AND c.end_line > c.start_line RETURN c._module AS path, c.start_line + 1 AS line, c.signature AS signature ORDER BY c.signature LIMIT 12 - """ + """, + prefix=APP_PREFIX, # Task 2: ``c._module AS path`` becomes a key derived from ``c.id`` ) positions = [(sample["module_path"], sample["inner_line"])] positions += [(r["path"], r["line"]) for r in others] @@ -563,9 +569,10 @@ def test_locate_many_is_a_single_round_trip(analysis, cypher): (r["path"], r["line"]) for r in cypher( """ - MATCH (c:PyCallable) WHERE c.start_line IS NOT NULL AND c.end_line > c.start_line + MATCH (c:PyCallable) WHERE c.id STARTS WITH $prefix AND c.start_line IS NOT NULL AND c.end_line > c.start_line RETURN c._module AS path, c.start_line + 1 AS line ORDER BY c.signature LIMIT 40 - """ + """, + prefix=APP_PREFIX, # Task 2: ``c._module AS path`` becomes a key derived from ``c.id`` ) ] assert len(positions) == 40 @@ -808,6 +815,32 @@ def test_entrypoint_report_is_genuinely_absent_from_the_graph(cypher): assert not any("entrypoint" in p for p in props), f"the projection now carries {props}; the diagnostic is stale" +# ===================================================================================== +# The analyzer-version probe (leg 1.6, F2), on whichever generation this graph is +# ===================================================================================== +def test_attach_warns_on_a_1_4_0_graph_and_is_silent_from_1_4_1(cypher, caplog): + """1.4.0 graphs are served (same id grammar) with one warning that scoped queries scan, since + they carry no ``:PyCanNode`` index; 1.4.1 and newer attach silently. Which branch this graph + exercises is read off the graph, so the same test is the back-compat gate on 7688 and the + silence check on 7689.""" + version = cypher("MATCH (a:PyApplication {name: $n}) RETURN a.analyzer_version AS v", n=APP_NAME)[0]["v"] + with caplog.at_level(logging.WARNING, logger="cldk.analysis.python.neo4j.neo4j_backend"): + facade = CLDK.python(backend=Neo4jConnectionConfig(uri=NEO4J_URI, username=NEO4J_USER, password=NEO4J_PASSWORD, application_name=APP_NAME)) + facade.backend.close() + warnings = [r.getMessage() for r in caplog.records if "scan rather than seek" in r.getMessage()] + if version == "1.4.0": + assert len(warnings) == 1 and "1.4.0" in warnings[0] + else: + assert not warnings, f"a {version} graph should attach silently" + + +def test_attaching_to_an_absent_application_is_refused_not_served_empty(): + """The version probe doubles as the "is this application even here" check: no ``:PyApplication`` + of that name means no version, and unknown is refused rather than served as silent empties.""" + with pytest.raises(GraphSchemaMismatch, match="1.4.0 or newer"): + CLDK.python(backend=Neo4jConnectionConfig(uri=NEO4J_URI, username=NEO4J_USER, password=NEO4J_PASSWORD, application_name="definitely-not-an-application-here")) + + # ===================================================================================== # External symbols and resolution # ===================================================================================== @@ -945,8 +978,8 @@ def busy_callable(cypher, module_keys) -> Dict[str, Any]: """ rows = cypher( """ - MATCH (m:PyModule)-[:PY_DECLARES]->(c:PyCallable) WHERE c._module IN $mods - MATCH (caller:PyCallable)-[:PY_CALLS]->(c) WHERE caller._module IN $mods + MATCH (m:PyModule)-[:PY_DECLARES]->(c:PyCallable) WHERE c.id STARTS WITH $prefix + MATCH (caller:PyCallable)-[:PY_CALLS]->(c) WHERE caller.id STARTS WITH $prefix WITH m, c, count(caller) AS n_callers WHERE n_callers > 0 MATCH (c)-[:PY_CALLS]->(ext:PyExternal) @@ -954,7 +987,7 @@ def busy_callable(cypher, module_keys) -> Dict[str, Any]: RETURN m.module_name AS module_name, c.name AS name, c.signature AS signature, external_callee_id ORDER BY c.signature LIMIT 1 """, - mods=list(module_keys), + prefix=APP_PREFIX, ) assert rows, "no top-level module function here has both a real caller and an external callee" return dict(rows[0]) @@ -971,8 +1004,8 @@ def test_get_call_graph_builds_with_external_targets_resolved(analysis, cypher, graph = analysis.get_call_graph() expected_edges = cypher( - "MATCH (s:PyCallable|PyExternal)-[r:PY_CALLS]->(t:PyCallable|PyExternal) WHERE s._module IN $mods RETURN count(r) AS c", - mods=list(module_keys), + "MATCH (s:PyCallable)-[r:PY_CALLS]->(t:PyCallable|PyExternal) WHERE s.id STARTS WITH $prefix RETURN count(r) AS c", + prefix=APP_PREFIX, )[0]["c"] assert graph.number_of_edges() == expected_edges assert graph.number_of_nodes() > 0 @@ -1085,8 +1118,8 @@ def test_get_classes_returns_every_top_level_class(analysis, cypher): # wrong direction because nothing exercised it. # ===================================================================================== _CROSSING_FLOW = """ -MATCH (a:PyBodyNode {kind:'formal_in', _module:$mod})-[:PY_DDG]->(:PyBodyNode)-[:PY_DDG]->(:PyBodyNode {kind:'actual_in'})-[:PY_PARAM_IN]->(v:PyBodyNode {kind:'formal_in'}) -WHERE a.var IS NOT NULL AND NOT a.var STARTS WITH '<' AND v.var IS NOT NULL AND NOT v.var STARTS WITH '<' +MATCH (a:PyBodyNode {kind:'formal_in'})-[:PY_DDG]->(:PyBodyNode)-[:PY_DDG]->(:PyBodyNode {kind:'actual_in'})-[:PY_PARAM_IN]->(v:PyBodyNode {kind:'formal_in'}) +WHERE a.id STARTS WITH $module_prefix AND a.var IS NOT NULL AND NOT a.var STARTS WITH '<' AND v.var IS NOT NULL AND NOT v.var STARTS WITH '<' WITH a, v ORDER BY a.id, v.id LIMIT 1 MATCH (c1:PyCallable)-[:PY_HAS_BODY_NODE]->(a) MATCH (c2:PyCallable)-[:PY_HAS_BODY_NODE]->(v) RETURN c1.signature AS src_callable, a.var AS src_value, a.id AS src_ref, @@ -1099,9 +1132,9 @@ def crossing_flow(analysis, cypher, module_keys) -> Dict[str, Any]: """A real value-to-value flow that crosses a call boundary, found by asking the graph. Scoped to one module at a time and taken from the first module (in ``file_key`` order) that has - one: the unscoped form is the same query without ``_module``, and it takes 66 seconds because - ordering the whole crossing set is the cost. Module-scoped it is 0.08 s, and the first hit on - this graph is the fifth module. + one: the unscoped form is the same query without the module prefix, and it takes 66 seconds + because ordering the whole crossing set is the cost. Module-scoped it is 0.08 s, and the first + hit on this graph is the fifth module. Both ends are then run back through ``resolve_value``, and a module whose names do not resolve uniquely is skipped rather than worked around: the fixture has to be addressable *the way a @@ -1109,7 +1142,7 @@ def crossing_flow(analysis, cypher, module_keys) -> Dict[str, Any]: real front door. """ for path in sorted(module_keys)[:40]: - rows = cypher(_CROSSING_FLOW, mod=path) + rows = cypher(_CROSSING_FLOW, module_prefix=f"{APP_PREFIX}{path}/") if not rows: continue found = dict(rows[0]) @@ -1236,6 +1269,53 @@ def test_paths_between_explains_the_flow_the_slice_only_asserts(analysis, crossi assert len({tuple((h.via, h.var, h.to.ref) for h in p.hops) for p in paths}) == len(paths), "no path is returned twice" +@pytest.fixture(scope="module") +def ghost_chain(cypher) -> Dict[str, str]: + """A real ``callable -> @external ghost -> callable`` chain inside this application with **no + all-callable route** beside it, or skip. + + Found by asking the graph, not hard-coded: leg 1.5 found two on this graph, both through + ``odoo.tools/parse_version``, and a rebuilt graph may hold different ones or none. The + all-callable check is the same quantified path pattern ``reaches`` compiles to, run here as + ground truth so the assertion below is about the ghost and nothing else. + """ + chains = cypher( + "MATCH (a:PyCallable)-[:PY_CALLS]->(g:PyExternal)-[:PY_CALLS]->(b:PyCallable) " + "WHERE a.id STARTS WITH $prefix AND b.id STARTS WITH $prefix AND a <> b " + "RETURN a.signature AS src, g.id AS ghost, b.signature AS dst ORDER BY src, ghost, dst LIMIT 20", + prefix=APP_PREFIX, + ) + for chain in chains: + routed = cypher( + "MATCH (a:PyCallable {signature:$a}) WHERE a.id STARTS WITH $prefix " + "MATCH (a) ((x:PyCallable)-[:PY_CALLS]->(y:PyCallable) WHERE x.id STARTS WITH $prefix){1,} (m:PyCallable) " + "WITH DISTINCT m WHERE m.signature = $b RETURN count(m) > 0 AS ok", + a=chain["src"], b=chain["dst"], prefix=APP_PREFIX, + )[0]["ok"] + if not routed: + return dict(chain) + pytest.skip("every callable -> ghost -> callable chain on this graph also has an all-callable route, so none isolates the ghost rule") + + +def test_a_ghost_is_reached_but_never_traversed_through(analysis, ghost_chain): + """The ghost rule (leg 1.5, kept by leg 1.6 F5) asserted on the live predicates. + + Every ghost's id sits under the application prefix -- ``can://python//@external/...`` -- + so a prefix predicate alone would admit it as a hop *source* and ``reaches`` would answer + ``True`` through it, contradicting ``get_call_graph``, which both backends build from + declared-callable-origin edges only. The chain here is chosen so the source calls **no** + all-callable route to ``dst`` exists (verified in Cypher by the fixture), so the ghost is the + whole question, not a detour beside a legitimate path. + """ + src, dst = ghost_chain["src"], ghost_chain["dst"] + assert analysis.reaches(src, dst) is False, f"reaches walked through {ghost_chain['ghost']}" + paths = analysis.call_paths_between(src, dst) + assert not [p for p in paths if any(h.to.kind == "external" for h in p.hops[:-1])], "a path routed through a ghost" + assert not paths, "no all-callable route exists, so there is no path to return" + # The ghost itself is still reached: it is a legitimate callee of the source. + assert ghost_chain["ghost"] in {c.ref for c in analysis.callees_of(src)} + + def test_call_paths_between_agrees_with_reaches(analysis, crossing_flow): """``reaches`` says whether; ``call_paths_between`` says how, and the two cannot disagree.""" sig = crossing_flow["src_callable"] diff --git a/tests/analysis/python/test_entrypoints.py b/tests/analysis/python/test_entrypoints.py index dd1c521b..f2d0c19d 100644 --- a/tests/analysis/python/test_entrypoints.py +++ b/tests/analysis/python/test_entrypoints.py @@ -131,7 +131,7 @@ def test_entrypoints_query_filters_on_is_entrypoint_property(): assert [o.signature for o in eps] == ["svc.app.handler"] assert isinstance(eps[0], PyCallableOverview) query = run.call_args.args[0] - assert "c._module IN $mods" in query + assert "c.id STARTS WITH $prefix" in query assert "SET" not in query and "CREATE" not in query and "MERGE" not in query and "DELETE" not in query @@ -209,7 +209,7 @@ def test_entrypoint_classes_query_filters_on_is_entrypoint_property(): assert [c.signature for c in classes] == ["svc.views.AdminView"] assert isinstance(classes[0], PyClassOverview) query = run.call_args.args[0] - assert "cl._module IN $mods" in query + assert "cl.id STARTS WITH $prefix" in query assert "SET" not in query and "CREATE" not in query and "MERGE" not in query and "DELETE" not in query diff --git a/tests/analysis/python/test_locate.py b/tests/analysis/python/test_locate.py index 86b3e9c3..cd1d2c83 100644 --- a/tests/analysis/python/test_locate.py +++ b/tests/analysis/python/test_locate.py @@ -253,16 +253,16 @@ def test_locate_parity_documented_module_source_divergence(py, py_local): # Application scope, path normalisation, and the file_not_in_graph distinction. # ================================================================================================ def test_locate_query_is_scoped_to_the_application(py, fake_driver): - """Every other query in neo4j_backend.py constrains ``_module IN $mods``; so must this one, or - a same-valued file_key from another application in the same database can win.""" + """Every other query in neo4j_backend.py constrains ``.id STARTS WITH $prefix``; so must this + one, or a same-valued file_key from another application in the same database can win.""" py.locate("src/app.py", 21) statement = next(s for s in fake_driver.statements if "UNWIND $positions AS pos" in s) - assert "c._module IN $mods" in statement + assert "c.id STARTS WITH $prefix" in statement def test_locate_scope_is_actually_honoured(py): - """Not just present in the text: drop the application's module keys and no callable matches.""" - py._modules = [] + """Not just present in the text: attach as another application and no callable matches.""" + py.application_name = "some_other_application" r = py.locate("src/app.py", 21) assert r.callable is None assert "module_scope" in [d.code for d in r.diagnostics] diff --git a/tests/analysis/python/test_neo4j_multi_application_scope.py b/tests/analysis/python/test_neo4j_multi_application_scope.py index 2be136fd..78b54ffc 100644 --- a/tests/analysis/python/test_neo4j_multi_application_scope.py +++ b/tests/analysis/python/test_neo4j_multi_application_scope.py @@ -46,20 +46,33 @@ from __future__ import annotations +import inspect import re from typing import Any, Dict, List import pytest +from codeanalyzer.schema.ids import module_id +from cldk.analysis.python.neo4j import neo4j_backend from cldk.analysis.python.neo4j.neo4j_backend import _BULK_CHILD_QUERIES, PyNeo4jBackend +APP_A, APP_B = "app_a", "app_b" APP_A_MODULE = "a/mod.py" APP_B_MODULE = "b/mod.py" +_APP_OF = {APP_A_MODULE: APP_A, APP_B_MODULE: APP_B} #: The colliding parent keys — the same class and method signature in both applications. CLASS_SIG = "shared.Widget" METHOD_SIG = "shared.Widget.render" + +def _node(module: str, key: str, **props: Any) -> Dict[str, Any]: + """A fixture node the way codeanalyzer-python 1.4.1 emits it: the application and the module + live in the ``id`` (``can://python///...``) and nowhere else -- there is no + ``_module`` property, which is exactly what the fake server has to be able to scope without.""" + return {"id": f"{module_id(_APP_OF[module], module)}/{key}", **props} + + #: ``bucket -> (parent key, {application module -> that application's one child's properties})``. #: Every child is named for its application, so ``alpha`` present and ``beta`` absent is the #: assertion; ``beta`` present means the statement that fetched it was not application-scoped. @@ -67,42 +80,42 @@ "class_methods": ( CLASS_SIG, { - APP_A_MODULE: {"signature": METHOD_SIG, "name": "alpha_method", "path": APP_A_MODULE, "_module": APP_A_MODULE}, - APP_B_MODULE: {"signature": METHOD_SIG, "name": "beta_method", "path": APP_B_MODULE, "_module": APP_B_MODULE}, + APP_A_MODULE: _node(APP_A_MODULE, "Widget/render", signature=METHOD_SIG, name="alpha_method", path=APP_A_MODULE), + APP_B_MODULE: _node(APP_B_MODULE, "Widget/render", signature=METHOD_SIG, name="beta_method", path=APP_B_MODULE), }, ), "class_attributes": ( CLASS_SIG, - {APP_A_MODULE: {"name": "alpha_attr"}, APP_B_MODULE: {"name": "beta_attr"}}, + {APP_A_MODULE: _node(APP_A_MODULE, "Widget/alpha_attr", name="alpha_attr"), APP_B_MODULE: _node(APP_B_MODULE, "Widget/beta_attr", name="beta_attr")}, ), "class_inner_classes": ( CLASS_SIG, { - APP_A_MODULE: {"signature": "alpha.Inner", "name": "Inner", "path": APP_A_MODULE, "_module": APP_A_MODULE}, - APP_B_MODULE: {"signature": "beta.Inner", "name": "Inner", "path": APP_B_MODULE, "_module": APP_B_MODULE}, + APP_A_MODULE: _node(APP_A_MODULE, "Widget/Inner", signature="alpha.Inner", name="Inner", path=APP_A_MODULE), + APP_B_MODULE: _node(APP_B_MODULE, "Widget/Inner", signature="beta.Inner", name="Inner", path=APP_B_MODULE), }, ), "callable_callsites": ( METHOD_SIG, - {APP_A_MODULE: {"method_name": "alpha_call"}, APP_B_MODULE: {"method_name": "beta_call"}}, + {APP_A_MODULE: _node(APP_A_MODULE, "Widget/render@1:0", method_name="alpha_call"), APP_B_MODULE: _node(APP_B_MODULE, "Widget/render@1:0", method_name="beta_call")}, ), "callable_inner_callables": ( METHOD_SIG, { - APP_A_MODULE: {"signature": "alpha.inner_fn", "name": "alpha_inner_fn", "path": APP_A_MODULE, "_module": APP_A_MODULE}, - APP_B_MODULE: {"signature": "beta.inner_fn", "name": "beta_inner_fn", "path": APP_B_MODULE, "_module": APP_B_MODULE}, + APP_A_MODULE: _node(APP_A_MODULE, "Widget/render/inner_fn", signature="alpha.inner_fn", name="alpha_inner_fn", path=APP_A_MODULE), + APP_B_MODULE: _node(APP_B_MODULE, "Widget/render/inner_fn", signature="beta.inner_fn", name="beta_inner_fn", path=APP_B_MODULE), }, ), "callable_inner_classes": ( METHOD_SIG, { - APP_A_MODULE: {"signature": "alpha.InnerInFn", "name": "InnerInFn", "path": APP_A_MODULE, "_module": APP_A_MODULE}, - APP_B_MODULE: {"signature": "beta.InnerInFn", "name": "InnerInFn", "path": APP_B_MODULE, "_module": APP_B_MODULE}, + APP_A_MODULE: _node(APP_A_MODULE, "Widget/render/InnerInFn", signature="alpha.InnerInFn", name="InnerInFn", path=APP_A_MODULE), + APP_B_MODULE: _node(APP_B_MODULE, "Widget/render/InnerInFn", signature="beta.InnerInFn", name="InnerInFn", path=APP_B_MODULE), }, ), "callable_variables": ( METHOD_SIG, - {APP_A_MODULE: {"name": "alpha_var"}, APP_B_MODULE: {"name": "beta_var"}}, + {APP_A_MODULE: _node(APP_A_MODULE, "Widget/render/alpha_var", name="alpha_var"), APP_B_MODULE: _node(APP_B_MODULE, "Widget/render/beta_var", name="beta_var")}, ), } @@ -120,10 +133,18 @@ # One :PyClass node per application, same signature, different owning module — what ``get_class`` # and ``get_all_classes`` select on before any child fetch happens. _CLASSES: List[Dict[str, Any]] = [ - {"signature": CLASS_SIG, "name": "Widget", "path": APP_A_MODULE, "_module": APP_A_MODULE}, - {"signature": CLASS_SIG, "name": "Widget", "path": APP_B_MODULE, "_module": APP_B_MODULE}, + _node(APP_A_MODULE, "Widget", signature=CLASS_SIG, name="Widget", path=APP_A_MODULE), + _node(APP_B_MODULE, "Widget", signature=CLASS_SIG, name="Widget", path=APP_B_MODULE), ] +#: The two spellings of the application scope a statement may carry: the whole application +#: (``$prefix``) or, for a narrowed bulk fetch, a list of per-module prefixes (``$prefixes``). +_MATCHES_BY_PREFIX = re.compile(r"\.id STARTS WITH \$prefix\b|any\(p IN \$prefixes WHERE \w+\.id STARTS WITH p\)") + + +def _is_scoped(statement: str) -> bool: + return bool(_MATCHES_BY_PREFIX.search(statement)) or "IN $mods" in statement + def _bucket_of(query: str) -> str | None: """Which child collection a statement is fetching — bulk and per-parent twins alike. @@ -149,25 +170,28 @@ def _bucket_of(query: str) -> str | None: return None -def _fake_two_app_cypher(query: str, **params: Any) -> List[Dict[str, Any]]: - """Answer the statements a class reconstruction issues, honestly. +def _in_scope(query: str, params: Dict[str, Any], props: Dict[str, Any]) -> bool: + """Evaluate the statement's application-scope predicate the way a 1.4.1 server would: on the + node's ``id`` prefix, and **only if the statement asks**. An unscoped statement sees both + applications' rows, which is what makes these tests fail without the predicate. A statement + still asking for ``_module IN $mods`` gets what a 1.4.1 graph gives it -- nothing.""" + if "STARTS WITH $prefix" in query: + return props["id"].startswith(params["prefix"]) + if "$prefixes" in query: + return any(props["id"].startswith(p) for p in params["prefixes"]) + if "IN $mods" in query: + return props.get("_module") in (params.get("mods") or []) + return True - "Honestly" is the whole point: the ``IN $mods`` filter is applied **only when the query - actually asks for it**, exactly as a real server would. An unscoped statement therefore sees - both applications' rows, which is what makes these tests fail without the predicate. - """ - mods = params.get("mods") or [] - - def scope(by_module: Dict[str, Dict[str, Any]]) -> List[Dict[str, Any]]: - items = by_module.items() - return [props for module, props in items if "IN $mods" not in query or module in mods] +def _fake_two_app_cypher(query: str, **params: Any) -> List[Dict[str, Any]]: + """Answer the statements a class reconstruction issues, honestly (see :func:`_in_scope`).""" bucket = _bucket_of(query) if bucket in _CHILDREN: pk, by_module = _CHILDREN[bucket] if params.get("sig", pk) != pk: # a per-parent statement about some other parent return [] - rows = scope(by_module) + rows = [props for props in by_module.values() if _in_scope(query, params, props)] if "AS pk" in query: # the bulk twin returns the parent key alongside the child return [{"pk": pk, "p": props} for props in rows] return [{"p": props} for props in rows] @@ -175,7 +199,7 @@ def scope(by_module: Dict[str, Dict[str, Any]]) -> List[Dict[str, Any]]: # "AS pk" excludes the module_classes bulk bucket, which has no rows in this fixture. if "(c:PyClass" in query and "PY_DECLARES" in query and "AS pk" not in query: rows = [c for c in _CLASSES if "$sig" not in query or c["signature"] == params["sig"]] - return [{"p": c} for c in rows if "IN $mods" not in query or c["_module"] in mods] + return [{"p": c} for c in rows if _in_scope(query, params, c)] return [] # imports and the module-level buckets: none in this fixture @@ -188,7 +212,7 @@ def run(query: str, **params: Any) -> List[Dict[str, Any]]: return _fake_two_app_cypher(query, **params) backend = object.__new__(PyNeo4jBackend) - backend.application_name = "app_a" + backend.application_name = APP_A backend._database = None backend._driver = None backend._session_obj = None @@ -198,10 +222,19 @@ def run(query: str, **params: Any) -> List[Dict[str, Any]]: return backend +def test_the_fake_graph_carries_no_module_property(): + """What a 1.4.1 graph is: scope lives in the id, and there is no ``_module`` to fall back on. + A fixture that grew the property back would let a ``_module``-scoped statement pass here while + returning nothing on a real graph.""" + nodes = [c for _, by_module in _CHILDREN.values() for c in by_module.values()] + _CLASSES + assert nodes and not any("_module" in n for n in nodes) + assert all(n["id"].startswith(("can://python/app_a/", "can://python/app_b/")) for n in nodes) + + def test_per_parent_path_does_not_leak_another_applications_children(): """``get_class`` must not merge application B's methods into application A's class.""" cls = _two_app_backend().get_class(CLASS_SIG) - assert cls is not None + assert cls is not None, "the application's own class came back empty -- the statement scoped on a property the graph does not carry" assert set(cls.callables) == {"alpha_method"}, "per-parent child fetch leaked another application's members" @@ -219,10 +252,10 @@ def test_scoped_and_bulk_paths_agree_across_applications(): def test_every_signature_keyed_child_collection_is_application_scoped(bucket: str, bulk: bool): """Each of the seven collections addressed by a *signature*, on both fetch paths. - ``class_methods`` was the only one with a regression test when the ``$mods`` predicate was - added to the per-parent statements; the other six were fixed in the same change and pinned by - nothing. They are all reachable from one class reconstruction, so one walk covers them and the - parameter names which child is being judged. + ``class_methods`` was the only one with a regression test when the application-scope predicate + was added to the per-parent statements; the other six were fixed in the same change and pinned + by nothing. They are all reachable from one class reconstruction, so one walk covers them and + the parameter names which child is being judged. """ backend = _two_app_backend() props = dict(_CLASSES[0]) @@ -240,8 +273,8 @@ def test_every_signature_keyed_child_collection_is_application_scoped(bucket: st def test_every_child_statement_carries_the_application_scope(): """The four module-keyed collections cannot be caught by the fixture above (see the module docstring), so they are caught here: every statement either path issues to fetch children is - scoped by ``$mods``, and a predicate silently dropped from any of the eleven fails this.""" - assert [b for b, q in _BULK_CHILD_QUERIES.items() if "IN $mods" not in q] == [] + application-scoped, and a predicate silently dropped from any of the eleven fails this.""" + assert [b for b, q in _BULK_CHILD_QUERIES.items() if not _is_scoped(q)] == [] issued: List[str] = [] backend = _two_app_backend(record=issued) @@ -249,7 +282,7 @@ def test_every_child_statement_carries_the_application_scope(): backend._class_full(dict(_CLASSES[0])) assert issued, "the walk issued no statements, so this asserts nothing" - assert [q for q in issued if "IN $mods" not in q] == [] + assert [q for q in issued if not _is_scoped(q)] == [] # ---------------------------------------------------------------------------------------------- @@ -259,23 +292,35 @@ def test_every_child_statement_carries_the_application_scope(): _MATCHES_BY_SIGNATURE = re.compile(r"signature\s*[:=]\s*\$|\.signature IN \$") _MATCHES_BY_ID = re.compile(r"\bid\s*:\s*\$|\.id IN \$") +#: Class-level strings that are Cypher but not a whole statement: appended to a scoped ``MATCH`` +#: at each use, so their scope is judged at the use sites (below), not on the fragment. +_FRAGMENTS = {"_OVERVIEW_PROJECTION"} + def _class_level_statements() -> Dict[str, str]: - return {name: value for name, value in vars(PyNeo4jBackend).items() if isinstance(value, str) and value.lstrip().startswith("MATCH")} + """Every class attribute that starts a Cypher clause -- not only ``MATCH``: filtering on that + alone silently skipped ``_OVERVIEW_PROJECTION`` (``OPTIONAL MATCH``) and ``_LOCATE_QUERY`` + (``UNWIND``), and an audit that cannot see a statement cannot protect it.""" + return { + name: value + for name, value in vars(PyNeo4jBackend).items() + if isinstance(value, str) and re.match(r"(MATCH|OPTIONAL MATCH|UNWIND|CALL)\b", value.lstrip()) + } def test_the_audit_sees_the_dataflow_statements_too(): names = set(_class_level_statements()) - for expected in ("_REACHES", "_CONE", "_PATHS", "_CALL_PATHS", "_VALUE_REACHES", "_CALLEE_VALUES", "_SOURCES", "_SLICE", "_CALLERS", "_CALLEES", "_OWN_EDGES"): + for expected in ("_REACHES", "_CONE", "_PATHS", "_CALL_PATHS", "_VALUE_REACHES", "_CALLEE_VALUES", "_SOURCES", "_SLICE", "_CALLERS", "_CALLEES", "_OWN_EDGES", "_LOCATE_QUERY", "_OVERVIEW_PROJECTION"): assert expected in names, f"{expected} is not a class-level statement any more; move it back or extend the audit" -@pytest.mark.parametrize("name", sorted(_class_level_statements())) +@pytest.mark.parametrize("name", sorted(set(_class_level_statements()) - _FRAGMENTS)) def test_every_statement_is_application_scoped_or_keyed_by_an_application_stamped_id(name): """Two ways a statement stays inside one application, and every statement must use one. A **signature** is not application-stamped: two applications in one database can declare the - same one, so any statement that matches a node by signature must also carry ``IN $mods``. + same one, so any statement that matches a node by signature must also carry the application + scope -- ``.id STARTS WITH $prefix`` (or, on a narrowed bulk fetch, the per-module prefixes). A body-node or ghost **id** embeds the application (``can://python//…``) and the emitter only ever links nodes from its own run, so a statement keyed *only* by id is scoped by construction and may omit the predicate -- ``_SLICE``, ``_PATHS`` and ``_VALUE_REACHES`` do, @@ -283,7 +328,18 @@ def test_every_statement_is_application_scoped_or_keyed_by_an_application_stampe neither would be unscoped and fails here. """ statement = _class_level_statements()[name] - by_signature, by_id = bool(_MATCHES_BY_SIGNATURE.search(statement)), bool(_MATCHES_BY_ID.search(statement)) - assert by_signature or by_id, f"{name} matches its nodes by neither signature nor id" + by_signature, by_id, scoped = bool(_MATCHES_BY_SIGNATURE.search(statement)), bool(_MATCHES_BY_ID.search(statement)), _is_scoped(statement) + assert scoped or by_id, f"{name} carries no application scope and is not keyed by an application-stamped id" if by_signature: - assert "IN $mods" in statement, f"{name} matches by signature without the application scope" + assert scoped, f"{name} matches by signature without the application scope" + + +def test_the_overview_projection_is_only_ever_appended_to_a_scoped_match(): + """``_OVERVIEW_PROJECTION`` binds no node of its own (``(c)`` is whatever the preceding + ``MATCH`` bound), so it is scoped exactly when every statement it is appended to is.""" + source = inspect.getsource(neo4j_backend) + uses = source.split("self._OVERVIEW_PROJECTION")[:-1] + assert len(uses) >= 3, "the projection is used from fewer places than expected; did it move?" + for before in uses: + statement = before[before.rindex("self._run(") :] + assert _is_scoped(statement), f"an unscoped MATCH feeds the overview projection: {statement[-200:]!r}" diff --git a/tests/analysis/python/test_resolution_probe.py b/tests/analysis/python/test_resolution_probe.py index d35465c0..bf842e92 100644 --- a/tests/analysis/python/test_resolution_probe.py +++ b/tests/analysis/python/test_resolution_probe.py @@ -42,8 +42,8 @@ def test_has_resolution_edges_false_on_a_level_1_graph(fake_driver): def test_probe_is_scoped_to_this_applications_modules(fake_driver): - """The probe query filters on ``s._module IN $mods`` -- verify the Cypher actually does, not - just that the boolean comes back right.""" + """The probe query filters on ``s.id STARTS WITH $prefix`` -- verify the Cypher actually does, + not just that the boolean comes back right.""" seen_queries = [] def responder(query, params): @@ -55,8 +55,8 @@ def responder(query, params): probe_calls = [(q, p) for q, p in seen_queries if "PY_RESOLVES_TO" in q] assert len(probe_calls) == 1 query, params = probe_calls[0] - assert "s._module IN $mods" in query - assert "mods" in params + assert "s.id STARTS WITH $prefix" in query + assert params["prefix"] == "can://python/app/" def test_probe_does_not_raise_on_a_legitimate_empty_result(fake_driver): diff --git a/tests/analysis/python/test_schema_probe.py b/tests/analysis/python/test_schema_probe.py index 2bb0e8a9..90ae6195 100644 --- a/tests/analysis/python/test_schema_probe.py +++ b/tests/analysis/python/test_schema_probe.py @@ -21,6 +21,8 @@ the mismatch is instead caught loudly, once, at connection time. """ +import logging + import pytest from cldk.utils.exceptions import GraphSchemaMismatch @@ -48,3 +50,39 @@ def test_probe_raises_on_empty_graph(fake_driver): fake_driver.rel_types = set() with pytest.raises(GraphSchemaMismatch): PyNeo4jBackend._from_driver(fake_driver, application_name="app") + + +# -----[ the analyzer-version floor (leg 1.6, F2) ]----- +def test_probe_refuses_a_graph_below_the_analyzer_floor(fake_driver): + """A 1.3.x graph has none of the ``can://`` id grammar the scoping relies on: every statement + would come back empty, so attach refuses and names what it found and the floor.""" + fake_driver.analyzer_version = "1.3.9" + with pytest.raises(GraphSchemaMismatch, match=r"1\.3\.9.*1\.4\.0 or newer"): + PyNeo4jBackend._from_driver(fake_driver, application_name="app") + + +@pytest.mark.parametrize("raw", [None, "garbage", ""], ids=["no-application", "unparsable", "empty"]) +def test_probe_refuses_when_the_version_cannot_be_read(fake_driver, raw): + """No :PyApplication of that name, or a version that is not one, is *unknown* -- and unknown + is refused, because serving it would be the silent-empty defect with no signal.""" + fake_driver.analyzer_version = raw + with pytest.raises(GraphSchemaMismatch, match="1.4.0 or newer"): + PyNeo4jBackend._from_driver(fake_driver, application_name="app") + + +def test_probe_serves_a_1_4_0_graph_and_warns_once_about_scanning(fake_driver, caplog): + """1.4.0 ids have the same grammar, so results are identical; only the index is missing.""" + fake_driver.analyzer_version = "1.4.0" + with caplog.at_level(logging.WARNING, logger="cldk.analysis.python.neo4j.neo4j_backend"): + PyNeo4jBackend._from_driver(fake_driver, application_name="app") + warnings = [r for r in caplog.records if "scan rather than seek" in r.getMessage()] + assert len(warnings) == 1 + assert "1.4.0" in warnings[0].getMessage() and "1.4.1" in warnings[0].getMessage() + + +@pytest.mark.parametrize("raw", ["1.4.1", "1.5.0", "2.0.0", "1.4.1.post1"]) +def test_probe_is_silent_from_1_4_1_up(fake_driver, caplog, raw): + fake_driver.analyzer_version = raw + with caplog.at_level(logging.WARNING, logger="cldk.analysis.python.neo4j.neo4j_backend"): + PyNeo4jBackend._from_driver(fake_driver, application_name="app") + assert not [r for r in caplog.records if "scan rather than seek" in r.getMessage()] diff --git a/tests/analysis/python/test_scoping_keywords.py b/tests/analysis/python/test_scoping_keywords.py index 81af42be..613bb839 100644 --- a/tests/analysis/python/test_scoping_keywords.py +++ b/tests/analysis/python/test_scoping_keywords.py @@ -384,6 +384,7 @@ def test_neo4j_scoped_prefetch_does_not_fetch_the_whole_application(): with backend._bulk(["pkg/b.py"]): backend._children("class_methods", "sig", "unused") assert [s["mods"] for s in seen] == [["pkg/b.py"]] + assert [s["prefixes"] for s in seen] == [["can://python/app/pkg/b.py/"]], "the signature-keyed buckets narrow on per-module id prefixes" def test_neo4j_bulk_scope_is_the_application_by_default(): @@ -391,6 +392,7 @@ def test_neo4j_bulk_scope_is_the_application_by_default(): with backend._bulk(): backend._children("class_methods", "sig", "unused") assert seen[0]["mods"] == ["pkg/a.py", "pkg/b.py"] + assert seen[0]["prefixes"] == ["can://python/app/"], "the whole application is one prefix, not one per module" def test_neo4j_nested_bulk_keeps_the_outer_scope(): @@ -401,6 +403,7 @@ def test_neo4j_nested_bulk_keeps_the_outer_scope(): with backend._bulk(["pkg/b.py"]): backend._children("class_methods", "sig", "unused") assert seen[0]["mods"] == ["pkg/a.py", "pkg/b.py"] + assert seen[0]["prefixes"] == ["can://python/app/"] def _rows_backend(rows: List[Dict[str, Any]], modules: List[str]) -> tuple[PyNeo4jBackend, List[Dict[str, Any]]]: @@ -435,16 +438,18 @@ def test_neo4j_unbounded_depth_leaves_the_upper_bound_open(): def test_neo4j_walk_is_scoped_to_the_application_at_every_hop(): """Not just at the endpoint. A ``*0..n`` variable-length pattern can only constrain where it - lands, so the walk could step out through a ``:PyExternal`` ghost — which carries no - ``_module`` for a per-hop predicate to test, and has 5,307 outgoing ``PY_CALLS`` edges on the - live graph, 5,108 of them to another ghost — and spend the rest of its hop budget walking the - ghost layer instead of this application's own callables. Not a hop into a *neighbouring* - application: every ghost id embeds the application name, so ghosts are not shared.""" + lands, so the walk could step out through a ``:PyExternal`` ghost — which has 5,307 outgoing + ``PY_CALLS`` edges on the live graph, 5,108 of them to another ghost — and spend the rest of its + hop budget walking the ghost layer instead of this application's own callables. A ghost's id + sits under the same application prefix as a callable's, so the prefix alone would admit it as a + hop *source*: every hop's source is pinned to ``:PyCallable`` by label. Not a hop into a + *neighbouring* application: every ghost id embeds the application name, so ghosts are not + shared.""" backend, seen = _recording_backend(["pkg/a.py"]) with pytest.raises(SelectorNotInGraph): backend.get_call_graph(roots=["pkg.a.go"], depth=2) query = seen[0]["query"] - assert "(a:PyCallable|PyExternal)-[:PY_CALLS]->(b:PyCallable|PyExternal) WHERE a._module IN $mods" in query + assert "(a:PyCallable)-[:PY_CALLS]->(b:PyCallable|PyExternal) WHERE a.id STARTS WITH $prefix" in query assert "*0.." not in query, "a variable-length hop cannot carry a per-hop scope predicate" From d7ce15cffcf2f44f60e73eec0a99db8df043056b Mon Sep 17 00:00:00 2001 From: Rahul Krishna Date: Sun, 6 Sep 2026 02:28:59 -0400 Subject: [PATCH 05/13] fix(python): derive a callable's module key from its can:// id The graph no longer stores _module, so the repo-relative path a caller sees is recovered from the id by longest match against the application's known module keys -- verified, never split, raised when it cannot be. --- cldk/analysis/python/neo4j/neo4j_backend.py | 113 +++++++++++------- cldk/analysis/python/neo4j/reconstruct.py | 32 ++++- tests/analysis/python/conftest.py | 6 +- tests/analysis/python/test_artifacts.py | 2 +- tests/analysis/python/test_e2e_neo4j_live.py | 32 ++--- tests/analysis/python/test_entrypoints.py | 15 +-- tests/analysis/python/test_locate.py | 9 +- tests/analysis/python/test_module_key_of.py | 62 ++++++++++ .../test_neo4j_multi_application_scope.py | 69 ++++++++--- 9 files changed, 246 insertions(+), 94 deletions(-) create mode 100644 tests/analysis/python/test_module_key_of.py diff --git a/cldk/analysis/python/neo4j/neo4j_backend.py b/cldk/analysis/python/neo4j/neo4j_backend.py index faad6923..09b73940 100644 --- a/cldk/analysis/python/neo4j/neo4j_backend.py +++ b/cldk/analysis/python/neo4j/neo4j_backend.py @@ -80,7 +80,8 @@ import re from collections import defaultdict from contextlib import contextmanager -from typing import Any, Dict, List, Sequence, Tuple +from functools import cached_property +from typing import Any, Callable, Dict, FrozenSet, List, Sequence, Tuple import networkx as nx from codeanalyzer.schema import model_dump_json @@ -168,14 +169,14 @@ def _semver(raw: Any) -> Tuple[int, int, int] | None: # merge another application's methods while ``get_all_classes`` would not. _BULK_CHILD_QUERIES: Dict[str, str] = { # module -> its own top-level declarations - "module_classes": "MATCH (m:PyModule)-[:PY_DECLARES]->(c:PyClass) WHERE m.file_key IN $mods RETURN m.file_key AS pk, properties(c) AS p", - "module_functions": "MATCH (m:PyModule)-[:PY_DECLARES]->(f:PyCallable) WHERE m.file_key IN $mods RETURN m.file_key AS pk, properties(f) AS p", + "module_classes": "MATCH (m:PyModule)-[:PY_DECLARES]->(c:PyClass) WHERE m.file_key IN $mods AND m.id STARTS WITH $prefix RETURN m.file_key AS pk, properties(c) AS p", + "module_functions": "MATCH (m:PyModule)-[:PY_DECLARES]->(f:PyCallable) WHERE m.file_key IN $mods AND m.id STARTS WITH $prefix RETURN m.file_key AS pk, properties(f) AS p", "module_variables": ( - "MATCH (m:PyModule)-[:PY_DECLARES_VAR]->(v:PyVariable) WHERE m.file_key IN $mods " + "MATCH (m:PyModule)-[:PY_DECLARES_VAR]->(v:PyVariable) WHERE m.file_key IN $mods AND m.id STARTS WITH $prefix " "RETURN m.file_key AS pk, properties(v) AS p ORDER BY v.start_line, v.name" ), "module_imports": ( - "MATCH (m:PyModule)-[e:PY_IMPORTS]->(pkg:PyPackage) WHERE m.file_key IN $mods " + "MATCH (m:PyModule)-[e:PY_IMPORTS]->(pkg:PyPackage) WHERE m.file_key IN $mods AND m.id STARTS WITH $prefix " "RETURN m.file_key AS pk, pkg.name AS module, e.imported_names AS names" ), # class -> its members @@ -196,9 +197,13 @@ def _semver(raw: Any) -> Tuple[int, int, int] | None: } -def _slice_node(row: Dict[str, Any]) -> SliceNode: +def _slice_node(row: Dict[str, Any], module_key: Callable[[str], str]) -> SliceNode: """One row of the slice query as a :class:`SliceNode`, in the caller's vocabulary. + ``file`` is derived from the body node's own ``ref`` by ``module_key`` (the backend's + :meth:`PyNeo4jBackend._module_key`): a body-node id is its callable's id plus ``@``, so + both embed the same module key, and the graph stores no path to project instead. + ``kind``/``name`` go through :func:`~cldk.analysis.commons.resolve.body_node_kind`, the same translation the local backend uses, so a vertex a caller addressed through ``resolve_value`` as a ``global`` comes back from a slice labelled a ``global`` too. @@ -210,7 +215,7 @@ def _slice_node(row: Dict[str, Any]) -> SliceNode: """ kind, name, defined_in = body_node_kind(row["kind"], row["var"]) return SliceNode( - file=row["file"], + file=module_key(row["ref"]), line=row["line"] if row["line"] is not None else row["c_line"], callable=row["callable"], kind=kind, @@ -221,7 +226,7 @@ def _slice_node(row: Dict[str, Any]) -> SliceNode: ) -def _call_neighbour(row: Dict[str, Any]) -> SliceNode: +def _call_neighbour(row: Dict[str, Any], module_key: Callable[[str], str]) -> SliceNode: """One ``PY_CALLS`` neighbour as a :class:`SliceNode` — declared callable or external ghost. Which it is, is read off the row rather than asked for in a second query: only a @@ -232,7 +237,7 @@ def _call_neighbour(row: Dict[str, Any]) -> SliceNode: discovered as sentinels. """ if row["signature"] is not None: - return SliceNode(file=row["file"], line=row["line"], callable=row["signature"], kind="callable", name=row["name"], source=None, ref=row["ref"]) + return SliceNode(file=module_key(row["ref"]), line=row["line"], callable=row["signature"], kind="callable", name=row["name"], source=None, ref=row["ref"]) qualified = f"{row['module']}.{row['name']}" if row["module"] else row["name"] return SliceNode(file="", line=0, callable=qualified, kind="external", name=row["name"], source=None, ref=row["ref"]) @@ -475,6 +480,24 @@ def _scope_prefix(self) -> str: """ return application_id(self.application_name) + "/" + @cached_property + def _module_set(self) -> FrozenSet[str]: + """:attr:`_modules` as a set -- the membership side of :func:`~cldk.analysis.python.neo4j.reconstruct.module_key_of`. + The list stays the Cypher parameter (the driver does not pack a set); this is the view every + projected row's key is verified against, built once.""" + return frozenset(self._modules) + + def _module_key(self, node_id: str) -> str: + """The repo-relative module key a node's ``can://`` id embeds (F4). The graph stores no + path property to project, so every ``path``/``file`` a caller sees is derived from the id + it came with and verified against the application's module keys -- never split, never + guessed.""" + return R.module_key_of(node_id, self._scope_prefix, self._module_set) + + def _overview(self, row: Dict[str, Any]) -> PyCallableOverview: + """A projected callable row (``_OVERVIEW_PROJECTION``'s shape) with its ``path`` derived.""" + return R.overview({**row, "path": self._module_key(row["id"])}) + def _module_prefixes(self, keys: Sequence[str] | None) -> List[str]: """Per-module id prefixes for **narrowing** a bulk fetch to a subset of the application's modules (``get_symbol_table(paths=...)``, ``get_all_classes(module=...)``); ``None`` is the @@ -577,7 +600,7 @@ def _children(self, bucket: str, key: str, query: str, **params: Any) -> List[Di def _collect(self, bucket: str) -> Dict[str, List[Dict[str, Any]]]: """One whole child collection for this application, in one round trip, grouped by ``pk``.""" index: Dict[str, List[Dict[str, Any]]] = defaultdict(list) - for row in self._run(_BULK_CHILD_QUERIES[bucket], mods=self._prefetch_scope, prefixes=self._prefetch_prefixes): + for row in self._run(_BULK_CHILD_QUERIES[bucket], mods=self._prefetch_scope, prefixes=self._prefetch_prefixes, prefix=self._scope_prefix): index[row["pk"]].append(row) return index @@ -696,7 +719,7 @@ def _module_full(self, props: Dict[str, Any]) -> PyModule: for r in self._children( "module_classes", file_key, - "MATCH (par:PyModule {file_key: $fk})-[:PY_DECLARES]->(c:PyClass) WHERE par.file_key IN $mods RETURN properties(c) AS p", + "MATCH (par:PyModule {file_key: $fk})-[:PY_DECLARES]->(c:PyClass) WHERE par.file_key IN $mods AND par.id STARTS WITH $prefix RETURN properties(c) AS p", fk=file_key, ): c = self._class_full(r["p"]) @@ -705,7 +728,7 @@ def _module_full(self, props: Dict[str, Any]) -> PyModule: for r in self._children( "module_functions", file_key, - "MATCH (par:PyModule {file_key: $fk})-[:PY_DECLARES]->(f:PyCallable) WHERE par.file_key IN $mods RETURN properties(f) AS p", + "MATCH (par:PyModule {file_key: $fk})-[:PY_DECLARES]->(f:PyCallable) WHERE par.file_key IN $mods AND par.id STARTS WITH $prefix RETURN properties(f) AS p", fk=file_key, ): fn = self._callable_full(r["p"]) @@ -716,7 +739,7 @@ def _module_full(self, props: Dict[str, Any]) -> PyModule: "module_variables", file_key, "MATCH (par:PyModule {file_key: $fk})-[:PY_DECLARES_VAR]->(v:PyVariable) " - "WHERE par.file_key IN $mods RETURN properties(v) AS p ORDER BY v.start_line, v.name", + "WHERE par.file_key IN $mods AND par.id STARTS WITH $prefix RETURN properties(v) AS p ORDER BY v.start_line, v.name", fk=file_key, ) ] @@ -730,7 +753,7 @@ def _module_imports(self, file_key: str) -> List[Any]: "module_imports", file_key, "MATCH (par:PyModule {file_key: $fk})-[e:PY_IMPORTS]->(pkg:PyPackage) " - "WHERE par.file_key IN $mods RETURN pkg.name AS module, e.imported_names AS names", + "WHERE par.file_key IN $mods AND par.id STARTS WITH $prefix RETURN pkg.name AS module, e.imported_names AS names", fk=file_key, ): names = r.get("names") or [] @@ -915,11 +938,11 @@ def get_python_module(self, file_path: str) -> PyModule | None: def get_python_file(self, qualified_class_name: str) -> str | None: # Only top-level classes are in the in-memory _class_to_file map (module.types). rows = self._run( - "MATCH (:PyModule)-[:PY_DECLARES]->(c:PyClass {signature: $sig}) WHERE c.id STARTS WITH $prefix RETURN c._module AS fk LIMIT 1", + "MATCH (:PyModule)-[:PY_DECLARES]->(c:PyClass {signature: $sig}) WHERE c.id STARTS WITH $prefix RETURN c.id AS id LIMIT 1", sig=qualified_class_name, prefix=self._scope_prefix, ) - return rows[0]["fk"] if rows else None + return self._module_key(rows[0]["id"]) if rows else None # ===================================================================================== # call graph @@ -1090,9 +1113,10 @@ def _get_module_functions(self, module_name: str) -> Dict[str, PyCallable]: """ rows = self._run( "MATCH (m:PyModule {module_name: $name})-[:PY_DECLARES]->(f:PyCallable) " - "WHERE m.file_key IN $mods RETURN properties(f) AS p", + "WHERE m.file_key IN $mods AND m.id STARTS WITH $prefix RETURN properties(f) AS p", name=module_name, mods=self._modules, + prefix=self._scope_prefix, ) return {fn.name: fn for fn in (self._callable_full(r["p"]) for r in rows)} @@ -1146,7 +1170,7 @@ def get_all_fields(self, qualified_class_name: str) -> List[PyClassAttribute]: _OVERVIEW_PROJECTION = ( "OPTIONAL MATCH (owner:PyClass)-[:PY_HAS_METHOD]->(c) " "RETURN c.signature AS signature, c.name AS name, c.decorators AS decorators, " - "c._module AS path, c.start_line AS start_line, c.end_line AS end_line, " + "c.id AS id, c.start_line AS start_line, c.end_line AS end_line, " "owner.signature AS class_signature" ) @@ -1155,7 +1179,7 @@ def get_callables_overview(self) -> List[PyCallableOverview]: "MATCH (c:PyCallable) WHERE c.id STARTS WITH $prefix " + self._OVERVIEW_PROJECTION, prefix=self._scope_prefix, ) - return [R.overview(r) for r in rows] + return [self._overview(r) for r in rows] def get_method_bodies(self, signatures: List[str]) -> Dict[str, str]: rows = self._run( @@ -1201,7 +1225,7 @@ def get_source(self, node_id: str) -> str: _RESOLVE_CALLABLE_QUERY = ( "MATCH (c:PyCallable) WHERE c.id STARTS WITH $prefix AND (c.signature = $name OR c.signature ENDS WITH $dotted) " "OPTIONAL MATCH (owner:PyClass)-[:PY_HAS_METHOD]->(c) " - "RETURN c.signature AS signature, c.name AS name, c.id AS id, c._module AS path, " + "RETURN c.signature AS signature, c.name AS name, c.id AS id, " "c.start_line AS start_line, owner.signature AS class_signature" ) @@ -1228,13 +1252,13 @@ def resolve_callable(self, name: str, *, in_class: str | None = None, in_module: if r["signature"] in by_sig: collisions.add(r["signature"]) by_sig[r["signature"]] = r - candidates = [CallableCandidate(r["signature"], r["class_signature"], r["path"]) for r in by_sig.values()] + candidates = [CallableCandidate(r["signature"], r["class_signature"], self._module_key(r["id"])) for r in by_sig.values()] sig = resolve_callable_signature(name, candidates, in_class=in_class, in_module=in_module) if sig in collisions: raise ValueError(f"{sig!r} is carried by more than one analysed callable; neither can be addressed unambiguously") row = by_sig[sig] return SliceNode( - file=row["path"], + file=self._module_key(row["id"]), line=row["start_line"], callable=row["signature"], kind="callable", @@ -1389,7 +1413,7 @@ def get_ddg(self, callable: str, *, in_class: str | None = None, page_size: int "UNWIND page AS nid " "MATCH (c:PyCallable)-[:PY_HAS_BODY_NODE]->(b:PyBodyNode {{id:nid}}) " "RETURN total, b.id AS ref, b.kind AS kind, b.var AS var, b.start_line AS line, " - "c.signature AS callable, c._module AS file, c.start_line AS c_line" + "c.signature AS callable, c.start_line AS c_line" ) def _slice(self, src: str, within: str, depth: int | None, max_nodes: int, *, backward: bool) -> Slice: @@ -1416,7 +1440,7 @@ def _slice(self, src: str, within: str, depth: int | None, max_nodes: int, *, ba right="-" if backward else "->", ) rows = self._run(query, id=root.ref, cap=max_nodes) - return Slice(nodes=[_slice_node(r) for r in rows], roots=[root], resolved=slice_resolved([root]), total=rows[0]["total"] if rows else 0) + return Slice(nodes=[_slice_node(r, self._module_key) for r in rows], roots=[root], resolved=slice_resolved([root]), total=rows[0]["total"] if rows else 0) def slice_backward(self, src: str, *, within: str, depth: int | None = DEFAULT_DEPTH, max_nodes: int = DEFAULT_MAX_NODES) -> Slice: """What affects this value (see :meth:`PythonAnalysisBackend.slice_backward`).""" @@ -1462,7 +1486,7 @@ def reaches(self, src: str, dst: str, *, depth: int | None = None) -> bool: "MATCH (s:PyCallable) WHERE s.signature IN $sigs AND s.id STARTS WITH $prefix " "MATCH (s) (()<-[:PY_CALLS]-(x:PyCallable) WHERE x.id STARTS WITH $prefix){{0,{depth}}} (m:PyCallable) " "WITH DISTINCT m ORDER BY m.id " - "WITH collect({{callable: m.signature, name: m.name, ref: m.id, file: m._module, line: m.start_line}}) AS found " + "WITH collect({{callable: m.signature, name: m.name, ref: m.id, line: m.start_line}}) AS found " "RETURN size(found) AS total, found[0..$cap] AS page" ) @@ -1473,7 +1497,7 @@ def backward_cone(self, sinks: Sequence[str], *, depth: int | None = DEFAULT_DEP self._require_quantified_paths("backward_cone") roots = cone_sinks(self.resolve_callable, sinks) row = self._run(self._CONE.format(depth="" if depth is None else depth), sigs=[r.callable for r in roots], cap=max_nodes, prefix=self._scope_prefix)[0] - nodes = [SliceNode(file=n["file"], line=n["line"], callable=n["callable"], kind="callable", name=n["name"], source=None, ref=n["ref"]) for n in row["page"]] + nodes = [SliceNode(file=self._module_key(n["ref"]), line=n["line"], callable=n["callable"], kind="callable", name=n["name"], source=None, ref=n["ref"]) for n in row["page"]] return Slice(nodes=nodes, roots=roots, resolved=slice_resolved(roots), total=row["total"]) #: ``t`` may be a ``:PyExternal`` ghost, which carries ``module``/``name``/``id`` and no @@ -1481,18 +1505,18 @@ def backward_cone(self, sinks: Sequence[str], *, depth: int | None = DEFAULT_DEP #: explicitly and :func:`_call_neighbour` decides what a row means from whether ``signature`` #: came back. ``(s:PyCallable)`` pins the *caller* side by label -- a ghost's id sits under #: the same prefix -- which is how a call originating at a ghost stays out of ``callers_of``. - _CALLERS = "MATCH (s:PyCallable)-[:PY_CALLS]->(t:PyCallable {signature: $sig}) WHERE s.id STARTS WITH $prefix RETURN s.signature AS signature, s.name AS name, s.id AS ref, s._module AS file, s.start_line AS line, s.module AS module" - _CALLEES = "MATCH (s:PyCallable {signature: $sig})-[:PY_CALLS]->(t:PyCallable|PyExternal) WHERE s.id STARTS WITH $prefix RETURN t.signature AS signature, t.name AS name, t.id AS ref, t._module AS file, t.start_line AS line, t.module AS module" + _CALLERS = "MATCH (s:PyCallable)-[:PY_CALLS]->(t:PyCallable {signature: $sig}) WHERE s.id STARTS WITH $prefix RETURN s.signature AS signature, s.name AS name, s.id AS ref, s.start_line AS line, s.module AS module" + _CALLEES = "MATCH (s:PyCallable {signature: $sig})-[:PY_CALLS]->(t:PyCallable|PyExternal) WHERE s.id STARTS WITH $prefix RETURN t.signature AS signature, t.name AS name, t.id AS ref, t.start_line AS line, t.module AS module" def callers_of(self, name: str, *, in_class: str | None = None, in_module: str | None = None) -> List[SliceNode]: """Who calls this (see :meth:`PythonAnalysisBackend.callers_of`).""" sig = self.resolve_callable(name, in_class=in_class, in_module=in_module).callable - return [_call_neighbour(r) for r in self._run(self._CALLERS, sig=sig, prefix=self._scope_prefix)] + return [_call_neighbour(r, self._module_key) for r in self._run(self._CALLERS, sig=sig, prefix=self._scope_prefix)] def callees_of(self, name: str, *, in_class: str | None = None, in_module: str | None = None) -> List[SliceNode]: """What this calls, externals included (see :meth:`PythonAnalysisBackend.callees_of`).""" sig = self.resolve_callable(name, in_class=in_class, in_module=in_module).callable - return [_call_neighbour(r) for r in self._run(self._CALLEES, sig=sig, prefix=self._scope_prefix)] + return [_call_neighbour(r, self._module_key) for r in self._run(self._CALLEES, sig=sig, prefix=self._scope_prefix)] # -----[ paths, mixed queries, hydration ]----- #: The caller's word for a hop, computed in Cypher so the ORDER BY below sorts by the same @@ -1531,7 +1555,6 @@ def callees_of(self, name: str, *, in_class: str | None = None, in_module: str | "WITH p, " + _PATH_ORDER + " AS key ORDER BY length(p), key LIMIT $cap " "RETURN [n IN nodes(p) | {{ref: n.id, kind: n.kind, var: n.var, line: n.start_line, " "callable: head([(c:PyCallable)-[:PY_HAS_BODY_NODE]->(n) | c.signature]), " - "file: head([(c:PyCallable)-[:PY_HAS_BODY_NODE]->(n) | c._module]), " "c_line: head([(c:PyCallable)-[:PY_HAS_BODY_NODE]->(n) | c.start_line])}}] AS ns, " "[r IN relationships(p) | {{via: type(r), var: r.var, prov: r.prov}}] AS rs" ) @@ -1551,7 +1574,7 @@ def callees_of(self, name: str, *, in_class: str | None = None, in_module: str | "MATCH (b:PyCallable {{signature:$dst}}) WHERE b.id STARTS WITH $prefix " "MATCH p = allShortestPaths((a)-[:PY_CALLS*1..{depth}]->(b)) WHERE all(n IN nodes(p) WHERE n:PyCallable) " "WITH p, " + _PATH_ORDER + " AS key ORDER BY length(p), key LIMIT $cap " - "RETURN [n IN nodes(p) | {{signature: n.signature, name: n.name, ref: n.id, file: n._module, " + "RETURN [n IN nodes(p) | {{signature: n.signature, name: n.name, ref: n.id, " "line: n.start_line, module: n.module}}] AS ns, " "[r IN relationships(p) | {{via: type(r), var: null, prov: null}}] AS rs" ) @@ -1563,7 +1586,7 @@ def _paths(self, query: str, node_of, a: SliceNode, b: SliceNode, *, src: str, d are the keys the query matches them by.""" check_distinct_endpoints(a, b) rows = self._run(query.format(rels=SDG_REL_PATTERN, depth="" if depth is None else depth), src=src, dst=dst, cap=max_paths + 1, prefix=self._scope_prefix) - paths = [flow_path([node_of(n) for n in r["ns"]], [(e["via"], e["var"], e["prov"]) for e in r["rs"]]) for r in rows[:max_paths]] + paths = [flow_path([node_of(n, self._module_key) for n in r["ns"]], [(e["via"], e["var"], e["prov"]) for e in r["rs"]]) for r in rows[:max_paths]] return FlowPaths(paths=paths, complete=len(rows) <= max_paths) # Argument validation precedes name resolution on every accessor below, as it does on the @@ -1653,23 +1676,23 @@ def get_decorated_callables(self, markers: List[str]) -> List[PyCallableOverview prefix=self._scope_prefix, markers=list(markers), ) - return [R.overview(r) for r in rows] + return [self._overview(r) for r in rows] def get_entrypoints(self) -> List[PyCallableOverview]: rows = self._run( "MATCH (c:PyCallable) WHERE c.id STARTS WITH $prefix AND c.is_entrypoint = true " + self._OVERVIEW_PROJECTION, prefix=self._scope_prefix, ) - return [R.overview(r) for r in rows] + return [self._overview(r) for r in rows] def get_entrypoint_classes(self) -> List[PyClassOverview]: rows = self._run( "MATCH (cl:PyClass) WHERE cl.id STARTS WITH $prefix AND cl.is_entrypoint = true " "RETURN cl.signature AS signature, cl.name AS name, cl.decorators AS decorators, " - "cl._module AS path, cl.start_line AS start_line, cl.end_line AS end_line", + "cl.id AS id, cl.start_line AS start_line, cl.end_line AS end_line", prefix=self._scope_prefix, ) - return [R.class_overview(r) for r in rows] + return [R.class_overview({**r, "path": self._module_key(r["id"])}) for r in rows] def get_entrypoint_coverage(self) -> EntrypointCoverage: # codeanalyzer-python's neo4j/project.py projects only the derived is_entrypoint / @@ -1827,12 +1850,12 @@ def get_config_readers(self, key: str) -> List[PyCallableOverview]: "MATCH (reader:PyCallable)-[:PY_HAS_BODY_NODE]->(bn) " "OPTIONAL MATCH (cls:PyClass)-[:PY_HAS_METHOD]->(reader) " "RETURN DISTINCT reader.signature AS signature, reader.name AS name, reader.decorators AS decorators, " - "reader._module AS path, reader.start_line AS start_line, reader.end_line AS end_line, " + "reader.id AS id, reader.start_line AS start_line, reader.end_line AS end_line, " "cls.signature AS class_signature", prefix=self._scope_prefix, key=key, ) - return [R.overview(r) for r in rows] + return [self._overview(r) for r in rows] # ===================================================================================== # locate / locate_many — one round trip, UNWIND over the position list @@ -1858,8 +1881,8 @@ def get_config_readers(self, key: str) -> List[PyCallableOverview]: "UNWIND $positions AS pos " "OPTIONAL MATCH (:PyApplication {name: $app})-[:PY_HAS_MODULE]->(m:PyModule {file_key: pos.path}) " "WITH pos, m " - "OPTIONAL MATCH (c:PyCallable {_module: pos.path}) " - "WHERE c.id STARTS WITH $prefix " + "OPTIONAL MATCH (c:PyCallable) " + "WHERE c.id STARTS WITH pos.module_prefix " "AND c.start_line IS NOT NULL AND c.end_line IS NOT NULL " "AND c.start_line <= pos.line AND pos.line <= c.end_line " "WITH pos, m, c " @@ -2005,11 +2028,15 @@ def locate_many(self, positions: Sequence[Tuple[str, int]]) -> List[LocateResult # the graph's file_key before it becomes a Cypher parameter — an unnormalised path would # match no :PyModule and read back as file_not_in_graph. keys = [resolve_module_key(path, self._modules) for path, _ in positions] + # ``module_prefix`` is the exact inverse of ``_module_key``: ``module_id(app, key) + "/"`` + # selects the module's own callables and nothing under a longer key sharing the spelling. rows = self._run( self._LOCATE_QUERY, app=self.application_name, - prefix=self._scope_prefix, - positions=[{"idx": i, "path": key, "line": line} for i, (key, (_, line)) in enumerate(zip(keys, positions))], + positions=[ + {"idx": i, "path": key, "module_prefix": module_id(self.application_name, key) + "/", "line": line} + for i, (key, (_, line)) in enumerate(zip(keys, positions)) + ], ) by_idx: Dict[int, List[Dict[str, Any]]] = defaultdict(list) for r in rows: diff --git a/cldk/analysis/python/neo4j/reconstruct.py b/cldk/analysis/python/neo4j/reconstruct.py index d67bbda7..20b895a9 100644 --- a/cldk/analysis/python/neo4j/reconstruct.py +++ b/cldk/analysis/python/neo4j/reconstruct.py @@ -38,7 +38,7 @@ from __future__ import annotations import json -from typing import Any, Dict, List, Mapping +from typing import Any, Collection, Dict, List, Mapping from cldk.models.python import ( BodyNode, @@ -68,6 +68,26 @@ Props = Mapping[str, Any] +# -----[ ids ]----- +def module_key_of(node_id: str, prefix: str, known: Collection[str]) -> str: + """The repo-relative module key embedded in a ``can://`` id (F4). + + Ids are ``/`` (or exactly ```` for a module), and a + file key can itself contain ``.py/`` as a directory name, so the key is never recovered by + splitting: every ``/``-boundary prefix of the id is tried longest first and the first that is + a member of ``known`` -- the application's verified module keys -- wins. A miss raises: a key + we cannot verify is a defect, not a guess. ``known`` should be a set; this runs once per row. + """ + if not node_id.startswith(prefix): + raise KeyError(node_id) + parts = node_id[len(prefix) :].split("/") + for n in range(len(parts), 0, -1): + candidate = "/".join(parts[:n]) + if candidate in known: + return candidate + raise KeyError(node_id) + + # -----[ helpers ]----- def comments(props: Props) -> List[PyComment]: """Rebuild the (lossy) comment list from the single ``docstring`` property.""" @@ -289,9 +309,9 @@ def overview(row: Props) -> PyCallableOverview: ``decorators``, ``path``, ``start_line``, ``end_line``, and ``class_signature`` (the owning class via ``PY_HAS_METHOD``, or ``None`` for a module-level / nested function). - ``path`` is projected from ``:PyCallable._module`` — the repo-relative module key, the same - vocabulary :func:`class_overview` and ``locate`` speak — *not* ``:PyCallable.path``, which is - the absolute path on the machine that ran the analysis. + ``path`` is the repo-relative module key derived from the node's ``id`` by the backend (see + :func:`module_key_of`), the same vocabulary :func:`class_overview` and ``locate`` speak — *not* + ``:PyCallable.path``, which is the absolute path on the machine that ran the analysis. """ class_sig = row.get("class_signature") return PyCallableOverview( @@ -311,8 +331,8 @@ def class_overview(row: Props) -> PyClassOverview: callable equivalent). ``row`` is a flat ``RETURN`` projection: ``signature``, ``name``, ``decorators``, - ``start_line``, ``end_line``, and ``path`` -- sourced from ``:PyClass``'s ``_module`` property, - since (unlike ``:PyCallable``) the node itself carries no ``path`` property. + ``start_line``, ``end_line``, and ``path`` -- derived from ``:PyClass``'s ``id`` by the backend + (see :func:`module_key_of`), since the node itself carries no ``path`` property. """ return PyClassOverview( signature=row.get("signature", ""), diff --git a/tests/analysis/python/conftest.py b/tests/analysis/python/conftest.py index 1ff7d24e..80854427 100644 --- a/tests/analysis/python/conftest.py +++ b/tests/analysis/python/conftest.py @@ -471,7 +471,7 @@ def _locate_responder(query: str, params: dict) -> list[dict]: match (attach as another application and the callables disappear), a callable row repeats once per matching body node, and every containment decision is made on lines the way Cypher would. """ - in_scope = "prefix" in params and _locate_callable_id("").startswith(params["prefix"]) + in_scope = "prefix" in params and _locate_callable_id("").startswith(params["prefix"]) # get_method_bodies / get_source if "RETURN m.file_key AS k" in query: return [{"k": _LOCATE_MODULE_PATH}] if "c.code IS NOT NULL" in query: @@ -495,8 +495,8 @@ def _locate_responder(query: str, params: dict) -> list[dict]: if pos["path"] != _LOCATE_MODULE_PATH: rows.append(_locate_row(pos["idx"])) # no :PyModule for this file_key continue - # ``OPTIONAL MATCH (c:PyCallable {_module: pos.path}) WHERE c.id STARTS WITH $prefix AND ...`` - matches = [c for c in _LOCATE_CALLABLE_SPECS if in_scope and c["start_line"] <= pos["line"] <= c["end_line"]] + # ``OPTIONAL MATCH (c:PyCallable) WHERE c.id STARTS WITH pos.module_prefix AND ...`` + matches = [c for c in _LOCATE_CALLABLE_SPECS if _locate_callable_id(c["signature"]).startswith(pos["module_prefix"]) and c["start_line"] <= pos["line"] <= c["end_line"]] if not matches: rows.append(_locate_row(pos["idx"], _LOCATE_MODULE_PROPS)) continue diff --git a/tests/analysis/python/test_artifacts.py b/tests/analysis/python/test_artifacts.py index b1b7d5ee..a7c12ecf 100644 --- a/tests/analysis/python/test_artifacts.py +++ b/tests/analysis/python/test_artifacts.py @@ -192,7 +192,7 @@ def py_local() -> PyCodeanalyzer: "signature": "src.db.connect", "name": "connect", "decorators": [], - "path": _MODULE_PATH, + "id": f"can://python/{_APP}/{_MODULE_PATH}/connect", # what get_config_readers projects; path is derived "start_line": 5, "end_line": 11, "class_signature": None, diff --git a/tests/analysis/python/test_e2e_neo4j_live.py b/tests/analysis/python/test_e2e_neo4j_live.py index ac95681e..f3cedf85 100644 --- a/tests/analysis/python/test_e2e_neo4j_live.py +++ b/tests/analysis/python/test_e2e_neo4j_live.py @@ -81,6 +81,7 @@ from cldk import CLDK from cldk.analysis.commons.backend_config import Neo4jConnectionConfig +from cldk.analysis.python.neo4j.reconstruct import module_key_of from cldk.utils.exceptions import GraphSchemaMismatch logging.getLogger("neo4j").setLevel(logging.ERROR) @@ -347,14 +348,15 @@ def test_overview_path_joins_locate_and_class_overview(analysis, sample, cypher) located = analysis.locate(sample["module_path"], sample["inner_line"]) assert located.module.path == row.path - # ...and the third vocabulary. ``PyClassOverview.path`` is projected from ``cl._module``; this - # graph has no entrypoint classes to read one back through, so the check goes to that property - # directly. Both projections now name a module by its ``file_key`` -- the same dictionary the - # callable overview draws from. (Not a subset check against the callable paths: a module can - # declare a class and no callable, and 66 of this application's 1,157 class-bearing modules do.) + # ...and the third vocabulary. ``PyClassOverview.path`` is derived from ``cl.id``; this graph + # has no entrypoint classes to read one back through, so the check derives it from every + # class id directly, the same way. Both projections name a module by its ``file_key`` -- the + # same dictionary the callable overview draws from. (Not a subset check against the callable + # paths: a module can declare a class and no callable, and 66 of this application's 1,157 + # class-bearing modules do.) module_keys = {r["k"] for r in cypher("MATCH (:PyApplication {name: $n})-[:PY_HAS_MODULE]->(m:PyModule) RETURN m.file_key AS k", n=APP_NAME)} - # Task 2: the ``cl._module AS p`` projection becomes a key derived from ``cl.id``. - class_paths = {r["p"] for r in cypher("MATCH (cl:PyClass) WHERE cl.id STARTS WITH $prefix RETURN DISTINCT cl._module AS p", prefix=APP_PREFIX)} + class_ids = [r["i"] for r in cypher("MATCH (cl:PyClass) WHERE cl.id STARTS WITH $prefix RETURN cl.id AS i", prefix=APP_PREFIX)] + class_paths = {module_key_of(i, APP_PREFIX, module_keys) for i in class_ids} assert class_paths, "no classes in the graph" assert row.path in module_keys and class_paths <= module_keys, "class and callable overviews disagree on path spelling" @@ -535,18 +537,18 @@ def test_locate_on_a_file_outside_the_graph_reports_file_not_in_graph(analysis): assert "module_xyzzy.py" in result.diagnostics[0].message -def test_locate_many_agrees_with_locate_position_by_position(analysis, sample, cypher): +def test_locate_many_agrees_with_locate_position_by_position(analysis, sample, cypher, module_keys): """``locate_many`` is not allowed to drift from ``locate``, nor to reorder its results.""" others = cypher( """ MATCH (c:PyCallable) WHERE c.id STARTS WITH $prefix AND c.code IS NOT NULL AND c.start_line IS NOT NULL AND c.end_line > c.start_line - RETURN c._module AS path, c.start_line + 1 AS line, c.signature AS signature + RETURN c.id AS id, c.start_line + 1 AS line, c.signature AS signature ORDER BY c.signature LIMIT 12 """, - prefix=APP_PREFIX, # Task 2: ``c._module AS path`` becomes a key derived from ``c.id`` + prefix=APP_PREFIX, ) positions = [(sample["module_path"], sample["inner_line"])] - positions += [(r["path"], r["line"]) for r in others] + positions += [(module_key_of(r["id"], APP_PREFIX, module_keys), r["line"]) for r in others] positions += [ (sample["module_path"], sample["module_scope_line"]), # module scope ("definitely/not/a/real/module_xyzzy.py", 3), # not in graph @@ -560,19 +562,19 @@ def test_locate_many_agrees_with_locate_position_by_position(analysis, sample, c assert got.model_dump() == one.model_dump(), f"locate_many disagreed with locate at {path}:{line}" -def test_locate_many_is_a_single_round_trip(analysis, cypher): +def test_locate_many_is_a_single_round_trip(analysis, cypher, module_keys): """Many positions must cost **one** Cypher statement, not one per position. Counted by wrapping the backend's ``_run`` seam — the only way to observe round trips at all. """ positions = [ - (r["path"], r["line"]) + (module_key_of(r["id"], APP_PREFIX, module_keys), r["line"]) for r in cypher( """ MATCH (c:PyCallable) WHERE c.id STARTS WITH $prefix AND c.start_line IS NOT NULL AND c.end_line > c.start_line - RETURN c._module AS path, c.start_line + 1 AS line ORDER BY c.signature LIMIT 40 + RETURN c.id AS id, c.start_line + 1 AS line ORDER BY c.signature LIMIT 40 """, - prefix=APP_PREFIX, # Task 2: ``c._module AS path`` becomes a key derived from ``c.id`` + prefix=APP_PREFIX, ) ] assert len(positions) == 40 diff --git a/tests/analysis/python/test_entrypoints.py b/tests/analysis/python/test_entrypoints.py index f2d0c19d..38b34a84 100644 --- a/tests/analysis/python/test_entrypoints.py +++ b/tests/analysis/python/test_entrypoints.py @@ -120,7 +120,7 @@ def test_entrypoints_query_filters_on_is_entrypoint_property(): "signature": "svc.app.handler", "name": "handler", "decorators": [], - "path": "svc/app.py", + "id": "can://python/app/svc/app.py/handler", "start_line": 1, "end_line": 2, "class_signature": None, @@ -145,8 +145,9 @@ def test_entrypoints_parity_between_backends(): """Same fixture, same signatures, on both backends.""" local_sigs = {o.signature for o in _local_backend().get_entrypoints()} - row_handler = {"signature": "svc.app.handler", "name": "handler", "decorators": [], "path": "svc/app.py", "start_line": 1, "end_line": 2, "class_signature": None} - row_run = {"signature": "svc.app.Service.run", "name": "run", "decorators": [], "path": "svc/app.py", "start_line": 3, "end_line": 4, "class_signature": "svc.app.Service"} + common = {"decorators": [], "start_line": 1, "end_line": 2} + row_handler = {"signature": "svc.app.handler", "name": "handler", "id": "can://python/app/svc/app.py/handler", "class_signature": None, **common} + row_run = {"signature": "svc.app.Service.run", "name": "run", "id": "can://python/app/svc/app.py/Service/run", "class_signature": "svc.app.Service", **common} backend = _neo4j_backend() with patch.object(PyNeo4jBackend, "_run", side_effect=_run_keyed({"c.is_entrypoint = true": [row_handler, row_run]})): neo4j_sigs = {o.signature for o in backend.get_entrypoints()} @@ -199,11 +200,11 @@ def test_entrypoint_classes_query_filters_on_is_entrypoint_property(): "signature": "svc.views.AdminView", "name": "AdminView", "decorators": [], - "path": "svc/views.py", + "id": "can://python/app/svc/views.py/AdminView", "start_line": 1, "end_line": 10, } - backend = _neo4j_backend() + backend = _neo4j_backend(modules=("svc/views.py",)) # the row's path is derived from its id and verified against these with patch.object(PyNeo4jBackend, "_run", side_effect=_run_keyed({"cl.is_entrypoint = true": [row]})) as run: classes = backend.get_entrypoint_classes() assert [c.signature for c in classes] == ["svc.views.AdminView"] @@ -226,11 +227,11 @@ def test_entrypoint_classes_parity_between_backends(): "signature": "svc.views.AdminView", "name": "AdminView", "decorators": [], - "path": "svc/views.py", + "id": "can://python/app/svc/views.py/AdminView", "start_line": 1, "end_line": 10, } - backend = _neo4j_backend() + backend = _neo4j_backend(modules=("svc/views.py",)) # the row's path is derived from its id and verified against these with patch.object(PyNeo4jBackend, "_run", side_effect=_run_keyed({"cl.is_entrypoint = true": [row]})): neo4j_sigs = {c.signature for c in backend.get_entrypoint_classes()} diff --git a/tests/analysis/python/test_locate.py b/tests/analysis/python/test_locate.py index cd1d2c83..f11e891d 100644 --- a/tests/analysis/python/test_locate.py +++ b/tests/analysis/python/test_locate.py @@ -253,11 +253,14 @@ def test_locate_parity_documented_module_source_divergence(py, py_local): # Application scope, path normalisation, and the file_not_in_graph distinction. # ================================================================================================ def test_locate_query_is_scoped_to_the_application(py, fake_driver): - """Every other query in neo4j_backend.py constrains ``.id STARTS WITH $prefix``; so must this - one, or a same-valued file_key from another application in the same database can win.""" + """Every other query in neo4j_backend.py constrains ``.id STARTS WITH $prefix``; this one + narrows further, to the position's own module: ``module_id(app, key) + "/"`` per position, so a + same-valued file_key from another application in the same database cannot win, and neither can + a module whose key merely extends this one's spelling.""" py.locate("src/app.py", 21) statement = next(s for s in fake_driver.statements if "UNWIND $positions AS pos" in s) - assert "c.id STARTS WITH $prefix" in statement + assert "c.id STARTS WITH pos.module_prefix" in statement + assert "_module" not in statement, "the graph stores no _module property to match on" def test_locate_scope_is_actually_honoured(py): diff --git a/tests/analysis/python/test_module_key_of.py b/tests/analysis/python/test_module_key_of.py new file mode 100644 index 00000000..e6d4cf47 --- /dev/null +++ b/tests/analysis/python/test_module_key_of.py @@ -0,0 +1,62 @@ +################################################################################ +# Copyright IBM Corporation 2026 +# +# Licensed under the Apache License, Version 2.0 (the "License"); +# you may not use this file except in compliance with the License. +# You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, software +# distributed under the License is distributed on an "AS IS" BASIS, +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +# See the License for the specific language governing permissions and +# limitations under the License. +################################################################################ + +"""``module_key_of``: the repo-relative module key a ``can://`` id embeds, recovered by verified +longest match against the application's known keys -- never by splitting on ``.py/``.""" + +import pytest + +from cldk.analysis.python.neo4j.reconstruct import module_key_of + +APP = "can://python/app/" + + +def test_module_key_is_the_id_segment_up_to_the_first_py_boundary(): + known = {"addons/account/models/account_move.py"} + node_id = "can://python/odoo-slim-19/addons/account/models/account_move.py/AccountMove/write(self,vals)" + assert module_key_of(node_id, "can://python/odoo-slim-19/", known) == "addons/account/models/account_move.py" + + +def test_a_directory_named_like_a_module_cannot_mis_key(): + known = {"pkg/x.py/real.py"} + assert module_key_of("can://python/app/pkg/x.py/real.py/f()", APP, known) == "pkg/x.py/real.py" + + +def test_a_key_outside_the_application_raises_rather_than_guesses(): + with pytest.raises(KeyError): + module_key_of("can://python/app/gone.py/f()", APP, {"kept.py"}) + + +def test_a_ghost_id_has_no_module_key(): + with pytest.raises(KeyError): + module_key_of("can://python/app/@external/os/path", APP, {"a.py"}) + + +def test_a_body_node_id_keys_to_its_callables_module(): + """A body node's id is its callable's id plus ``@``, and the key itself may contain + ``/`` (``@15:2/actual_in:0``); the module key is still the verified prefix, so a slice row can + derive its file from the body node's own ``ref``.""" + known = {"pkg/a.py", "pkg"} + assert module_key_of("can://python/app/pkg/a.py/f(x)@15:2/actual_in:0", APP, known) == "pkg/a.py" + + +def test_the_module_id_itself_keys_to_its_own_key(): + assert module_key_of("can://python/app/pkg/a.py", APP, {"pkg/a.py"}) == "pkg/a.py" + + +def test_an_id_under_another_application_raises(): + with pytest.raises(KeyError): + module_key_of("can://python/app-b/pkg/a.py/f()", APP, {"pkg/a.py"}) diff --git a/tests/analysis/python/test_neo4j_multi_application_scope.py b/tests/analysis/python/test_neo4j_multi_application_scope.py index 78b54ffc..47b1300c 100644 --- a/tests/analysis/python/test_neo4j_multi_application_scope.py +++ b/tests/analysis/python/test_neo4j_multi_application_scope.py @@ -33,15 +33,14 @@ applications' children. Every child below is named for its own application, so a leak is visible by name rather than by count. -**Why the seven buckets covered here are the seven that can leak.** Of the eleven child -collections, four (``module_classes``, ``module_functions``, ``module_variables``, -``module_imports``) are keyed by a module ``file_key`` rather than a signature. Their per-parent -statements pin the parent with ``{file_key: $fk}`` and their bulk rows come back under -``pk = file_key``, so separating two applications there requires their module keys to differ — -and if two applications *shared* a module key, ``$mods`` (a list of module keys) could not tell -them apart either, since it is the same key. Those four are covered by -``test_every_child_statement_carries_the_application_scope`` instead, which is what catches the -predicate being dropped from them. +**The four module-keyed collections** (``module_classes``, ``module_functions``, +``module_variables``, ``module_imports``) are addressed by a module ``file_key`` rather than a +signature, and ``$mods`` (a list of module keys) cannot tell two applications apart when they +*share* a key -- ``src/__init__.py`` in both is the ordinary case, not an exotic one. So those +statements carry the application's id prefix too, and the second fixture below is exactly that +collision: one ``file_key`` declared by both applications, a function named for each. The audit +``test_every_child_statement_carries_the_application_scope`` catches the predicate being dropped +from any of the eleven. """ from __future__ import annotations @@ -130,6 +129,15 @@ def _node(module: str, key: str, **props: Any) -> Dict[str, Any]: "callable_variables": (lambda c: {v.name for v in c.callables["alpha_method"].local_variables}, {"alpha_var"}), } +#: The module-key collision: one ``file_key`` declared by **both** applications, each holding one +#: top-level function named for its application. ``module_id`` puts the application in the id and +#: nowhere else, exactly as for the class fixture above. +SHARED_MODULE = "src/__init__.py" +_SHARED_FUNCTIONS: Dict[str, Dict[str, Any]] = { + app: {"id": f"{module_id(app, SHARED_MODULE)}/{name}", "signature": f"src.{name}", "name": name, "path": SHARED_MODULE} + for app, name in ((APP_A, "alpha_fn"), (APP_B, "beta_fn")) +} + # One :PyClass node per application, same signature, different owning module — what ``get_class`` # and ``get_all_classes`` select on before any child fetch happens. _CLASSES: List[Dict[str, Any]] = [ @@ -137,13 +145,15 @@ def _node(module: str, key: str, **props: Any) -> Dict[str, Any]: _node(APP_B_MODULE, "Widget", signature=CLASS_SIG, name="Widget", path=APP_B_MODULE), ] -#: The two spellings of the application scope a statement may carry: the whole application -#: (``$prefix``) or, for a narrowed bulk fetch, a list of per-module prefixes (``$prefixes``). -_MATCHES_BY_PREFIX = re.compile(r"\.id STARTS WITH \$prefix\b|any\(p IN \$prefixes WHERE \w+\.id STARTS WITH p\)") +#: The three spellings of the application scope a statement may carry: the whole application +#: (``$prefix``); for a narrowed bulk fetch, a list of per-module prefixes (``$prefixes``); and for +#: ``locate``, one module's own prefix per position (``pos.module_prefix``, minted from the same +#: application name). ``file_key IN $mods`` is *not* one: a module key is not application-stamped. +_MATCHES_BY_PREFIX = re.compile(r"\.id STARTS WITH (\$prefix\b|pos\.module_prefix\b)|any\(p IN \$prefixes WHERE \w+\.id STARTS WITH p\)") def _is_scoped(statement: str) -> bool: - return bool(_MATCHES_BY_PREFIX.search(statement)) or "IN $mods" in statement + return bool(_MATCHES_BY_PREFIX.search(statement)) def _bucket_of(query: str) -> str | None: @@ -167,6 +177,8 @@ def _bucket_of(query: str) -> str | None: return "callable_inner_callables" if "(d:PyClass)" in query: return "callable_inner_classes" + if "(f:PyCallable)" in query: + return "module_functions" return None @@ -179,7 +191,7 @@ def _in_scope(query: str, params: Dict[str, Any], props: Dict[str, Any]) -> bool return props["id"].startswith(params["prefix"]) if "$prefixes" in query: return any(props["id"].startswith(p) for p in params["prefixes"]) - if "IN $mods" in query: + if "_module IN $mods" in query: return props.get("_module") in (params.get("mods") or []) return True @@ -187,6 +199,16 @@ def _in_scope(query: str, params: Dict[str, Any], props: Dict[str, Any]) -> bool def _fake_two_app_cypher(query: str, **params: Any) -> List[Dict[str, Any]]: """Answer the statements a class reconstruction issues, honestly (see :func:`_in_scope`).""" bucket = _bucket_of(query) + if bucket == "module_functions": + # Keyed by file_key: the per-parent twin names it (``$fk``), the bulk twin lists the scope + # (``$mods``). Both applications declare SHARED_MODULE, so that filter alone admits both. + wanted = [params["fk"]] if "fk" in params else (params.get("mods") or []) + if SHARED_MODULE not in wanted: + return [] + rows = [props for props in _SHARED_FUNCTIONS.values() if _in_scope(query, params, props)] + if "AS pk" in query: + return [{"pk": SHARED_MODULE, "p": props} for props in rows] + return [{"p": props} for props in rows] if bucket in _CHILDREN: pk, by_module = _CHILDREN[bucket] if params.get("sig", pk) != pk: # a per-parent statement about some other parent @@ -216,7 +238,7 @@ def run(query: str, **params: Any) -> List[Dict[str, Any]]: backend._database = None backend._driver = None backend._session_obj = None - backend._modules = [APP_A_MODULE] + backend._modules = [APP_A_MODULE, SHARED_MODULE] backend._call_graph = None backend._run = run return backend @@ -226,7 +248,7 @@ def test_the_fake_graph_carries_no_module_property(): """What a 1.4.1 graph is: scope lives in the id, and there is no ``_module`` to fall back on. A fixture that grew the property back would let a ``_module``-scoped statement pass here while returning nothing on a real graph.""" - nodes = [c for _, by_module in _CHILDREN.values() for c in by_module.values()] + _CLASSES + nodes = [c for _, by_module in _CHILDREN.values() for c in by_module.values()] + _CLASSES + list(_SHARED_FUNCTIONS.values()) assert nodes and not any("_module" in n for n in nodes) assert all(n["id"].startswith(("can://python/app_a/", "can://python/app_b/")) for n in nodes) @@ -270,6 +292,21 @@ def test_every_signature_keyed_child_collection_is_application_scoped(bucket: st assert extract(cls) == expected, f"{bucket} leaked another application's children" +@pytest.mark.parametrize("bulk", [False, True], ids=["per_parent", "bulk"]) +def test_a_module_key_shared_by_two_applications_does_not_leak(bulk: bool): + """``file_key IN $mods`` cannot separate two applications that both declare ``src/__init__.py``; + only the id prefix can. Application A's module must come back with A's function and not B's, + on both fetch paths.""" + backend = _two_app_backend() + props = {"file_key": SHARED_MODULE, "module_name": "src"} + if bulk: + with backend._bulk(): + module = backend._module_full(props) + else: + module = backend._module_full(props) + assert set(module.functions) == {"alpha_fn"}, "the module-keyed fetch leaked another application's function" + + def test_every_child_statement_carries_the_application_scope(): """The four module-keyed collections cannot be caught by the fixture above (see the module docstring), so they are caught here: every statement either path issues to fetch children is From 2d9f791e16ff44269ec7a6a37ea497fc475c6bdc Mon Sep 17 00:00:00 2001 From: Rahul Krishna Date: Sun, 6 Sep 2026 03:03:31 -0400 Subject: [PATCH 06/13] fix(python): read the entrypoint report the 1.4.1 graph carries get_entrypoint_coverage over Neo4j answered with a diagnostic whose premise -- the projection never carries the report -- 1.4.1 (#182) made false. The diagnostic now fires only when the property is genuinely absent. --- cldk/analysis/commons/results.py | 7 +-- cldk/analysis/python/backend.py | 6 +-- cldk/analysis/python/neo4j/neo4j_backend.py | 28 ++++++++--- cldk/analysis/python/python_analysis.py | 5 +- docs/agent-api-reference.md | 7 +-- tests/analysis/python/test_e2e_neo4j_live.py | 51 +++++++++++++------- tests/analysis/python/test_entrypoints.py | 41 ++++++++++++---- 7 files changed, 100 insertions(+), 45 deletions(-) diff --git a/cldk/analysis/commons/results.py b/cldk/analysis/commons/results.py index f12deb44..ddcc946e 100644 --- a/cldk/analysis/commons/results.py +++ b/cldk/analysis/commons/results.py @@ -263,9 +263,10 @@ class EntrypointCoverage(BaseModel): unresolved: Count of near-misses, keyed by rule/framework, that could not be resolved to a definite entrypoint — non-zero counts are exactly the under-approximation gap. errors: Hard failures the detection pass hit while running. - diagnostics: Non-empty when a backend cannot supply this report at all: the Neo4j - projection does not carry ``PyApplication.entrypoint_report`` on the graph (only the - derived ``is_entrypoint``/``entrypoint_frameworks`` per-node properties), so it returns + diagnostics: Non-empty when a backend cannot supply this report at all: a Neo4j graph + emitted by codeanalyzer-python 1.4.0 does not carry ``PyApplication.entrypoint_report`` + (only the derived ``is_entrypoint``/``entrypoint_frameworks`` per-node properties; + 1.4.1 projects it onto ``:PyApplication``), so the Neo4j backend returns ``entrypoint_report_unavailable`` here instead of fabricating empty-but-clean-looking fields — the same "say so honestly" precedent as ``LocateResult``'s ``module_source_unavailable``. When ``diagnostics`` is non-empty, the other fields are diff --git a/cldk/analysis/python/backend.py b/cldk/analysis/python/backend.py index ca45b810..230b9cef 100644 --- a/cldk/analysis/python/backend.py +++ b/cldk/analysis/python/backend.py @@ -894,9 +894,9 @@ def get_entrypoint_coverage(self) -> EntrypointCoverage: so a caller can tell "the pass ran clean and found nothing" apart from "the pass had gaps" — a distinction :meth:`get_entrypoints`'s empty list alone cannot make. See :class:`~cldk.analysis.commons.results.EntrypointCoverage` for the field-by-field contract, - including the per-backend availability caveat (the Neo4j projection does not carry this - report at all; that backend answers with a ``diagnostics``-only result rather than - fabricating empty-but-clean-looking coverage fields).""" + including the per-backend availability caveat (a Neo4j graph emitted by codeanalyzer-python + 1.4.0 does not carry this report; that backend then answers with a ``diagnostics``-only + result rather than fabricating empty-but-clean-looking coverage fields).""" @property @abstractmethod diff --git a/cldk/analysis/python/neo4j/neo4j_backend.py b/cldk/analysis/python/neo4j/neo4j_backend.py index 09b73940..f7d6b507 100644 --- a/cldk/analysis/python/neo4j/neo4j_backend.py +++ b/cldk/analysis/python/neo4j/neo4j_backend.py @@ -86,6 +86,7 @@ import networkx as nx from codeanalyzer.schema import model_dump_json from codeanalyzer.schema.ids import application_id, module_id +from codeanalyzer.schema.py_schema import PyEntrypointReport from cldk.analysis.commons.resolve import CallableCandidate, body_node_kind, resolve_callable_signature, resolve_value_name, resolve_within, value_candidate from cldk.analysis.commons.results import CallableRef, Diagnostic, EdgePage, EntrypointCoverage, FlowPath, FlowPaths, LocateResult, ModuleRef, PathHop, Slice, SliceNode, TypeRef @@ -1695,12 +1696,27 @@ def get_entrypoint_classes(self) -> List[PyClassOverview]: return [R.class_overview({**r, "path": self._module_key(r["id"])}) for r in rows] def get_entrypoint_coverage(self) -> EntrypointCoverage: - # codeanalyzer-python's neo4j/project.py projects only the derived is_entrypoint / - # entrypoint_frameworks properties onto :PyCallable/:PyClass -- it never projects - # PyApplication.entrypoint_report (frameworks_detected/rulesets/unresolved/errors) onto the - # graph at all (confirmed: no such property or node anywhere in neo4j/project.py). Say so - # rather than fabricate empty-but-clean-looking coverage fields -- same precedent as - # LocateResult's module_source_unavailable for the module-text gap. + """The entrypoint pass's coverage record, read off ``:PyApplication.entrypoint_report_json``. + + codeanalyzer-python 1.4.1 (#182, ``neo4j/project.py``) projects ``PyApplication.entrypoint_report`` + onto the application node as the sorted-key JSON of the whole ``PyEntrypointReport`` -- + ``frameworks_detected``, ``rulesets``, ``unresolved``, ``errors`` -- so this parses the same + model the local backend passes through and there is no lossiness between the two. (The + sibling ``entrypoint_frameworks`` property is that report's ``frameworks_detected`` and is + not read separately.) A 1.4.0 graph has no such property: that absence is reported with + ``entrypoint_report_unavailable`` rather than fabricated as empty-but-clean-looking fields + -- same precedent as ``LocateResult``'s ``module_source_unavailable`` for the module-text gap. + """ + rows = self._run("MATCH (a:PyApplication {name: $app}) RETURN a.entrypoint_report_json AS j", app=self.application_name) + raw = rows[0].get("j") if rows else None + if raw is not None: + r = PyEntrypointReport.model_validate_json(raw) + return EntrypointCoverage( + frameworks_detected=list(r.frameworks_detected), + rulesets=list(r.rulesets), + unresolved=dict(r.unresolved), + errors=list(r.errors), + ) return EntrypointCoverage( diagnostics=[ Diagnostic( diff --git a/cldk/analysis/python/python_analysis.py b/cldk/analysis/python/python_analysis.py index 95a7b228..56b83966 100644 --- a/cldk/analysis/python/python_analysis.py +++ b/cldk/analysis/python/python_analysis.py @@ -664,8 +664,9 @@ def get_entrypoint_coverage(self) -> EntrypointCoverage: Returns: An :class:`~cldk.analysis.commons.results.EntrypointCoverage`. Non-empty - ``diagnostics`` means this backend cannot supply the report at all (the Neo4j - projection does not carry it) rather than the pass having run clean — see the model's + ``diagnostics`` means this backend cannot supply the report at all (a Neo4j graph + emitted by codeanalyzer-python 1.4.0 does not carry it) rather than the pass having + run clean — see the model's own docstring for the field-by-field contract. See Also: diff --git a/docs/agent-api-reference.md b/docs/agent-api-reference.md index e3c1277b..f22284e1 100644 --- a/docs/agent-api-reference.md +++ b/docs/agent-api-reference.md @@ -266,9 +266,10 @@ checkout it detects **zero** entrypoints across 15,549 callables, in a framework from HTTP routes. So an empty `get_entrypoints()` means either "no entrypoints" or "the pass found nothing", and you -cannot tell from the list. `get_entrypoint_coverage()` is how you ask. Over Neo4j it reports -`entrypoint_report_unavailable` — the graph does not carry the report — which is itself the answer: -*you cannot trust the zero*. +cannot tell from the list. `get_entrypoint_coverage()` is how you ask. Over a Neo4j graph emitted +by codeanalyzer-python 1.4.0 it reports `entrypoint_report_unavailable` — that graph does not carry +the report — which is itself the answer: *you cannot trust the zero*. From 1.4.1 the graph carries +it and the answer is the pass's own report, same as the local backend. Concluding "this application has no attack surface" from an empty list is the single worst mistake available in this API. diff --git a/tests/analysis/python/test_e2e_neo4j_live.py b/tests/analysis/python/test_e2e_neo4j_live.py index f3cedf85..f9c473eb 100644 --- a/tests/analysis/python/test_e2e_neo4j_live.py +++ b/tests/analysis/python/test_e2e_neo4j_live.py @@ -72,6 +72,7 @@ from __future__ import annotations +import json import logging import os from pathlib import Path @@ -790,31 +791,45 @@ def test_entrypoints_faithfully_report_what_the_graph_says(analysis, cypher): assert len(analysis.get_entrypoint_classes()) == flagged_classes -def test_entrypoint_coverage_reports_that_it_cannot_tell(analysis): - """The honest "I cannot tell you whether that zero is real". +def test_entrypoint_coverage_is_the_graphs_report_or_says_why_it_cannot_tell(analysis, cypher): + """Whether a caller can trust ``get_entrypoints() == []`` depends on the analyzer generation. - This diagnostic is the only thing standing between a caller and concluding, from - ``get_entrypoints() == []``, that this application has no attack surface. The Neo4j projection - never carries ``PyApplication.entrypoint_report``, so empty coverage fields here mean *unknown*, - not *none* — and the caller has to be able to see the difference. + From 1.4.1 (#182) ``:PyApplication`` carries the pass's own report, and the answer is that + report verbatim — the same fields the local backend returns. A 1.4.0 graph never carried it, + and there the diagnostic is the only thing standing between a caller and concluding, from an + empty list, that this application has no attack surface: empty coverage fields mean *unknown*, + not *none*. Which branch runs is read off the graph, so this is the back-compat gate on 7688 + and the parity check on 7689. """ + row = cypher("MATCH (a:PyApplication {name: $n}) RETURN a.analyzer_version AS v, a.entrypoint_report_json AS j", n=APP_NAME)[0] coverage = analysis.get_entrypoint_coverage() - codes = {d.code for d in coverage.diagnostics} - assert "entrypoint_report_unavailable" in codes - - message = next(d.message for d in coverage.diagnostics if d.code == "entrypoint_report_unavailable") - assert "entrypoint_report" in message - # The clean-looking empties are exactly what the diagnostic is there to qualify. - assert coverage.frameworks_detected == [] - assert coverage.rulesets == [] + if row["v"] == "1.4.0": + assert [d.code for d in coverage.diagnostics] == ["entrypoint_report_unavailable"] + assert "entrypoint_report" in coverage.diagnostics[0].message + # The clean-looking empties are exactly what the diagnostic is there to qualify. + assert coverage.frameworks_detected == [] + assert coverage.rulesets == [] + else: + assert coverage.diagnostics == [], f"a {row['v']} graph carries the report; the diagnostic is stale" + assert coverage.model_dump(exclude={"diagnostics"}) == json.loads(row["j"]) + assert coverage.frameworks_detected, "the 1.4.1 pass detects Odoo; an empty list here is a regression, not a clean run" -def test_entrypoint_report_is_genuinely_absent_from_the_graph(cypher): - """Ground truth for the diagnostic above: no node anywhere carries an entrypoint report.""" +def test_entrypoint_report_presence_matches_the_analyzer_generation(cypher): + """Ground truth for the branch above: 1.4.0 projected no report anywhere; 1.4.1 writes + ``entrypoint_frameworks`` and ``entrypoint_report_json`` onto ``:PyApplication``, the latter + being the analyzer's own ``PyEntrypointReport`` with all four fields — nothing is dropped.""" rows = cypher("MATCH (a:PyApplication {name: $n}) RETURN properties(a) AS p", n=APP_NAME) assert rows, f"application {APP_NAME!r} vanished" - props = set(rows[0]["p"]) - assert not any("entrypoint" in p for p in props), f"the projection now carries {props}; the diagnostic is stale" + props = rows[0]["p"] + entrypoint_props = {p for p in props if "entrypoint" in p} + if props["analyzer_version"] == "1.4.0": + assert not entrypoint_props, f"the 1.4.0 projection now carries {entrypoint_props}; the diagnostic branch is stale" + else: + assert entrypoint_props == {"entrypoint_frameworks", "entrypoint_report_json"} + report = json.loads(props["entrypoint_report_json"]) + assert set(report) == {"frameworks_detected", "rulesets", "unresolved", "errors"} + assert report["frameworks_detected"] == props["entrypoint_frameworks"] # ===================================================================================== diff --git a/tests/analysis/python/test_entrypoints.py b/tests/analysis/python/test_entrypoints.py index 38b34a84..5c3391fa 100644 --- a/tests/analysis/python/test_entrypoints.py +++ b/tests/analysis/python/test_entrypoints.py @@ -34,11 +34,12 @@ * **Finding 1** -- an empty ``get_entrypoints()`` cannot distinguish "ran clean, found none" from "the detection pass had gaps" (its own ``PyEntrypointReport`` docstring: "under-approximates by design, so silence is its failure mode"). ``get_entrypoint_coverage()`` surfaces that report. - Verified against ``codeanalyzer/neo4j/project.py``: the Neo4j projection never emits - ``PyApplication.entrypoint_report`` (frameworks_detected/rulesets/unresolved/errors) onto the - graph at all -- only the derived ``is_entrypoint``/``entrypoint_frameworks`` per-node properties - exist there -- so the Neo4j backend answers with a ``diagnostics``-only - ``entrypoint_report_unavailable`` result rather than fabricating a clean-looking empty report. + Over Neo4j the report is read off ``:PyApplication.entrypoint_report_json``, which + codeanalyzer-python projects from 1.4.1 (#182) as the sorted-key JSON of the very same + ``PyEntrypointReport`` the local backend passes through. A 1.4.0 graph has no such property + (only the derived ``is_entrypoint``/``entrypoint_frameworks`` per-node marks), and there the + Neo4j backend answers with a ``diagnostics``-only ``entrypoint_report_unavailable`` result + rather than fabricating a clean-looking empty report. The local backend is built the same way ``test_python_bulk_accessors.py`` builds its fixture (``object.__new__`` + a hand-assembled ``PyApplication``); the Neo4j backend is built the same way @@ -46,6 +47,7 @@ no live server, no FakeDriver machinery needed for a single filtered projection query. """ +import json from unittest.mock import patch from codeanalyzer.schema.py_schema import PyApplication, PyCallable, PyClass, PyEntrypointReport, PyModule @@ -258,13 +260,32 @@ def test_entrypoint_coverage_surfaces_the_report_locally(): assert coverage.diagnostics == [] -def test_entrypoint_coverage_over_neo4j_says_it_cannot_answer(): - """The Neo4j projection never emits PyApplication.entrypoint_report onto the graph (verified - against codeanalyzer/neo4j/project.py: only the derived is_entrypoint/entrypoint_frameworks - per-node properties exist there) -- say so via a diagnostic rather than fabricate a +def test_entrypoint_coverage_over_neo4j_is_read_from_the_application_node(): + """A 1.4.1 graph carries the report as ``:PyApplication.entrypoint_report_json`` -- the same + ``PyEntrypointReport`` the local backend passes through, dumped as sorted-key JSON + (codeanalyzer/neo4j/project.py, #182) -- so both backends answer identically from one analysis.""" + report = PyEntrypointReport( + frameworks_detected=["flask"], + rulesets=["shipped"], + unresolved={"flask": 2}, + errors=["timeout scanning svc/legacy.py"], + ) + local = object.__new__(PyCodeanalyzer) + local.application = PyApplication(symbol_table={}, entrypoint_report=report) + row = {"j": json.dumps(report.model_dump(mode="json"), sort_keys=True)} + backend = _neo4j_backend() + with patch.object(PyNeo4jBackend, "_run", side_effect=_run_keyed({"a.entrypoint_report_json AS j": [row]})): + coverage = backend.get_entrypoint_coverage() + assert coverage == local.get_entrypoint_coverage() + assert coverage.diagnostics == [] + + +def test_entrypoint_coverage_over_a_graph_without_the_report_says_it_cannot_answer(): + """A 1.4.0 graph never had the property -- say so via a diagnostic rather than fabricate a clean-looking empty report.""" backend = _neo4j_backend() - coverage = backend.get_entrypoint_coverage() + with patch.object(PyNeo4jBackend, "_run", side_effect=_run_keyed({"a.entrypoint_report_json AS j": [{"j": None}]})): + coverage = backend.get_entrypoint_coverage() assert len(coverage.diagnostics) == 1 assert coverage.diagnostics[0].code == "entrypoint_report_unavailable" assert coverage.frameworks_detected == [] From 292bc436001320c75fb6e2724d1b73dde758f89c Mon Sep 17 00:00:00 2001 From: Rahul Krishna Date: Sun, 6 Sep 2026 03:46:55 -0400 Subject: [PATCH 07/13] fix(python): seek locate through :PyCanNode on 1.4.1 graphs; read the entrypoint report by properties() The probe now keeps the analyzer generation on the backend (_analyzer_version), and one statement uses it: locate_many names :PyCanNode on a 1.4.1+ graph so the per-module prefix seeks the range index (40 positions: 399 -> 89 ms on the 1.4.1 odoo graph). Every application-prefix statement stays unlabelled by measurement: the seek then walks all 955,961 nodes under the application before the label and signature filters apply, 2-20x slower than the :PyCallable label scan (_RESOLVE_CALLABLE_QUERY 16.7 -> 198 ms). A 1.4.0 graph has no such label and is never asked for it. get_entrypoint_coverage reads properties(a) rather than a.entrypoint_report_json, so a 1.4.0 graph -- which has no such property key -- no longer draws a server warning on every call. Every remaining `_module` in the backend is a name coincidence or a sentence of history; no comment describes it as a scoping mechanism any more. --- cldk/analysis/python/neo4j/neo4j_backend.py | 112 +++++++++++++------- tests/analysis/python/test_entrypoints.py | 6 +- tests/analysis/python/test_locate.py | 18 ++++ 3 files changed, 95 insertions(+), 41 deletions(-) diff --git a/cldk/analysis/python/neo4j/neo4j_backend.py b/cldk/analysis/python/neo4j/neo4j_backend.py index f7d6b507..39b2253b 100644 --- a/cldk/analysis/python/neo4j/neo4j_backend.py +++ b/cldk/analysis/python/neo4j/neo4j_backend.py @@ -44,9 +44,11 @@ * call-graph edges are ``(:PyCallable|:PyExternal)-[:PY_CALLS {weight, prov}]->(...)`` with a constant ``CALL_DEP`` type; * class inheritance is ``(:PyClass)-[:PY_EXTENDS]->(:PyClass)`` (plus a ``base_classes`` property); -* every project-owned node carries a ``_module`` provenance property, so a single database may hold - several applications — all queries here are scoped to this backend's application, anchored on - ``(:PyApplication {name})-[:PY_HAS_MODULE]->(:PyModule)``. +* every node the analyzer emits for an application — module, class, callable, body node and + ``@external`` ghost — carries an id under ``can://python//``, so a single database may hold + several applications; every statement here is scoped to this backend's application by that id + prefix. (1.4.0 graphs also stamped a ``_module`` provenance property on project-owned nodes; + 1.4.1 retired it, and nothing here reads it.) In-memory dict keys this backend reproduces exactly (the projection stores nodes by ``signature`` only, so the keys are rebuilt from node properties): ``module.types`` / a class's own ``types`` → @@ -159,11 +161,12 @@ def _semver(raw: Any) -> Tuple[int, int, int] | None: # / ``_module_full``, and reproduce those statements' row shapes exactly so either source can feed # the same reconstruction code (see ``PyNeo4jBackend._children``). # -# Scoping: the module-level buckets key on ``m.file_key``, the rest on the parent's ``_module`` -# provenance property -- which the emitter indexes for every module-owned label (see -# ``codeanalyzer/neo4j/schema.py``'s ``INDEXES``). Both confine the result to this backend's -# application. The per-parent statements carry the *same* ``IN $mods`` predicate on the parent -# (``PyNeo4jBackend._children`` supplies ``mods`` to every unprimed run), because a bare +# Scoping: the module-level buckets key on ``m.file_key IN $mods`` plus the application prefix on +# the module's id; the signature-keyed buckets on the parent's id prefix -- the application's when +# the whole application is read, the per-module prefixes when a scoped accessor narrows the fetch +# (``$prefixes``, see ``PyNeo4jBackend._module_prefixes``). Both confine the result to this +# backend's application. The per-parent statements carry the *same* prefix predicate on the parent +# (``PyNeo4jBackend._children`` supplies ``prefix`` to every unprimed run), because a bare # ``{signature: $sig}`` would also match a same-signature node belonging to another application # in a shared database -- and a Unified Knowledge Graph holding several applications is the # expected deployment. Without it the two paths provably disagree there: ``get_class`` would @@ -385,6 +388,7 @@ def _probe_schema(self) -> None: missing=set(), message=f"The graph for application {self.application_name!r} {what}; this backend needs a graph emitted by codeanalyzer-python {floor} or newer.", ) + self._analyzer_version = version if version < self._ANALYZER_INDEXED: logger.warning( "The graph for application %r was emitted by codeanalyzer-python %s, which carries no :PyCanNode index on id: " @@ -481,6 +485,30 @@ def _scope_prefix(self) -> str: """ return application_id(self.application_name) + "/" + #: The codeanalyzer-python generation that emitted this application, set by + #: :meth:`_probe_schema` (which refuses anything below :attr:`_ANALYZER_FLOOR`). The class-level + #: ``None`` is for the ``object.__new__`` seam the unit tests build backends through: an unknown + #: generation names no optional label. + _analyzer_version: Tuple[int, ...] | None = None + + @property + def _can_node(self) -> str: + """``":PyCanNode"`` when the attached graph carries that label's range index on ``id`` + (codeanalyzer-python 1.4.1+), else ``""``. Interpolated into ONE statement, ``_LOCATE_QUERY``. + + Naming the label makes the planner seek ``:PyCanNode(id)`` for a prefix instead of scanning + ``:PyCallable``, and that pays only when the prefix is narrow. Measured on the 1.4.1 odoo + graph: ``locate_many`` over 40 positions (per-module prefixes) 399 -> 89 ms; every statement + whose prefix is the whole application -- ``resolve_callable``, ``get_source``, + ``resolve_value``, ``get_class``, the per-parent child fetches, the callers/callees + neighbourhoods -- 2-20x *slower* (``_RESOLVE_CALLABLE_QUERY`` 16.7 -> 198 ms), because the + seek then walks all 955,961 nodes under ``can://python//`` before the label and + signature filters apply, where the label scan touched 15,549. Those stay unlabelled by + measurement, not oversight. On a 1.4.0 graph the label does not exist and naming it would + match nothing, hence the gate. + """ + return ":PyCanNode" if self._analyzer_version is not None and self._analyzer_version >= self._ANALYZER_INDEXED else "" + @cached_property def _module_set(self) -> FrozenSet[str]: """:attr:`_modules` as a set -- the membership side of :func:`~cldk.analysis.python.neo4j.reconstruct.module_key_of`. @@ -567,11 +595,10 @@ def _load_module_keys(self) -> List[str]: # # **On nesting depth:** there is none to bound. The recursion is real (an inner class has # methods, a nested callable has call sites), but it never happens *in Cypher*: every bulk - # statement is a single flat hop scoped by the parent's ``_module`` provenance property, which - # every projected node carries at every nesting depth — ``codeanalyzer/neo4j/project.py`` - # threads the module's ``file_key`` down through ``_project_class`` / ``_project_callable``'s - # own recursion, so a class nested five levels deep appears in the ``class_inner_classes`` rows - # exactly like a top-level one. The tree is then rebuilt in Python by the same recursive calls + # statement is a single flat hop scoped by the parent's id prefix, and every projected node's + # id embeds its application and module at every nesting depth — ``codeanalyzer/neo4j/project.py`` + # mints a nested declaration's id under its module's, so a class nested five levels deep + # appears in the ``class_inner_classes`` rows exactly like a top-level one. The tree is then rebuilt in Python by the same recursive calls # as before, to whatever depth the graph actually has. No variable-length path, no depth # ceiling, and therefore no depth at which a deeply nested declaration would be silently # truncated. @@ -580,8 +607,8 @@ def _children(self, bucket: str, key: str, query: str, **params: Any) -> List[Di """The child rows of one parent node — from the bulk index when primed, one query when not. Unprimed (the default, and what every single-node accessor pays) runs ``query``: the - statement naming this one parent, application-scoped by the ``mods`` parameter this - method supplies, exactly as its bulk twin is scoped. Primed + statement naming this one parent, application-scoped by the ``prefix`` parameter this + method supplies (plus ``mods`` for the module-keyed ones), exactly as its bulk twin is scoped. Primed (inside :meth:`_bulk`) answers from ``_BULK_CHILD_QUERIES[bucket]``, fetched lazily on first use so an accessor is never charged for a collection it does not read — ``get_all_classes`` never touches the four module-level buckets. Both paths yield the same @@ -835,9 +862,9 @@ def _bounded_call_rows(self, roots: List[str], depth: int | None) -> List[Dict[s The walk is **application-scoped at every hop**, which is why it is a quantified path pattern (Cypher 5.9+) rather than a ``*0..n`` variable-length one: a variable-length pattern can constrain only its endpoint, so the walk could step out through an external - ghost. ``:PyExternal`` carries no ``_module``, so no per-hop module predicate can be - expressed about it, and it has 5,307 outgoing ``PY_CALLS`` edges on this graph, 5,108 of - them landing on another ghost. The leak is traversal **through** the ghost layer *inside* + ghost. A ghost's id sits under the same application prefix as a callable's, so the prefix + alone cannot keep the walk out of it -- the per-hop ``a:PyCallable`` label does -- and it + has 5,307 outgoing ``PY_CALLS`` edges on this graph, 5,108 of them landing on another ghost. The leak is traversal **through** the ghost layer *inside* this one application — a two-hop budget spent walking ghost-to-ghost instead of through the application's own callables — not a hop into a neighbouring application: every ghost id embeds the application name (``can://python/odoo-slim-19/@external/IPython/start_ipython``), @@ -1163,11 +1190,14 @@ def get_all_fields(self, qualified_class_name: str) -> List[PyClassAttribute]: # Field-projected RETURNs that sidestep the per-entity reconstruction fan-out: each is a single # Cypher statement, not the N+1 walk get_symbol_table()/get_all_methods_in_application() pays. # - # ``path`` comes from ``c._module``, not ``c.path``: the latter is the absolute path on the - # machine that ran the analysis (``/Users/…/checkout/addons/…``), which joins to nothing a - # caller holds -- not ``locate().module.path``, not ``PyClassOverview.path`` (already - # ``cl._module``), not ``get_symbol_table()``'s keys, and not any path on another host. - # ``_module`` is the repo-relative module key, i.e. the one vocabulary the whole facade speaks. + # ``path`` is derived from ``c.id`` (:meth:`_module_key`), never read from ``c.path``: the + # latter is the absolute path on the machine that ran the analysis + # (``/Users/…/checkout/addons/…``), which joins to nothing a caller holds -- not + # ``locate().module.path``, not ``PyClassOverview.path`` (derived the same way), not + # ``get_symbol_table()``'s keys, and not any path on another host. The derived key is the + # repo-relative module key, i.e. the one vocabulary the whole facade speaks. (Leg 1 projected + # the 1.4.0 graphs' ``_module`` property for this; 1.4.1 retired the property, and the id + # embeds the same key.) _OVERVIEW_PROJECTION = ( "OPTIONAL MATCH (owner:PyClass)-[:PY_HAS_METHOD]->(c) " "RETURN c.signature AS signature, c.name AS name, c.decorators AS decorators, " @@ -1423,7 +1453,7 @@ def _slice(self, src: str, within: str, depth: int | None, max_nodes: int, *, ba The two differ only in which way the arrows point, so they share a query and a builder -- a second copy would be a second place for the node vocabulary to drift. - **Not scoped by ``_module``,** unlike the per-callable accessors. A body-node id is stamped + **Not scoped by the application prefix,** unlike the per-callable accessors. A body-node id is stamped with its application (``can://python//…``) and the emitter only ever links nodes from its own run, so the traversal cannot leave the application it started in; adding ``m.id STARTS WITH $prefix`` would cost a string-prefix test on every one of 195,784 reached @@ -1502,7 +1532,7 @@ def backward_cone(self, sinks: Sequence[str], *, depth: int | None = DEFAULT_DEP return Slice(nodes=nodes, roots=roots, resolved=slice_resolved(roots), total=row["total"]) #: ``t`` may be a ``:PyExternal`` ghost, which carries ``module``/``name``/``id`` and no - #: ``signature``, ``_module`` or ``start_line`` -- so the projection names each property + #: ``signature`` or ``start_line`` -- so the projection names each property #: explicitly and :func:`_call_neighbour` decides what a row means from whether ``signature`` #: came back. ``(s:PyCallable)`` pins the *caller* side by label -- a ghost's id sits under #: the same prefix -- which is how a call originating at a ghost stays out of ``callers_of``. @@ -1610,7 +1640,7 @@ def call_paths_between(self, src: str, dst: str, *, depth: int | None = None, ma return self._paths(self._CALL_PATHS, _call_neighbour, a, b, src=a.callable, dst=b.callable, depth=depth, max_paths=max_paths) #: ``WITH DISTINCT m`` before the membership test, for :attr:`_REACHES`' measured reason: it is - #: what makes this a pruning BFS instead of a trail enumeration. Not scoped by ``_module``, for + #: what makes this a pruning BFS instead of a trail enumeration. Not scoped by the application prefix, for #: :meth:`_slice`'s reason: body-node ids embed the application, so both the seed and every #: ``$dsts`` id are this application's by construction. _VALUE_REACHES = "MATCH (a:PyBodyNode {{id:$src}})-[:{rels}*1..{depth}]->(m:PyBodyNode) WITH DISTINCT m WHERE m.id IN $dsts RETURN count(m) > 0 AS ok" @@ -1648,7 +1678,7 @@ def flows_to_argument(self, src: str, callee: str, arg: str, *, within: str, dep #: ``null`` code deliberately: the graph carries no text below callable granularity, and a #: ghost was never analysed, so those rows say "found, and there is nothing to read", which is #: what keeps that apart from "not found" (see :meth:`PythonAnalysisBackend.describe`). Only - #: the callable arm is ``_module``-scoped: a body-node id and a ghost id both embed the + #: the callable arm carries the prefix predicate: a body-node id and a ghost id both embed the #: application, while a signature does not. _SOURCES = ( "MATCH (c:PyCallable) WHERE c.id STARTS WITH $prefix AND (c.id IN $refs OR c.signature IN $refs) " @@ -1707,8 +1737,10 @@ def get_entrypoint_coverage(self) -> EntrypointCoverage: ``entrypoint_report_unavailable`` rather than fabricated as empty-but-clean-looking fields -- same precedent as ``LocateResult``'s ``module_source_unavailable`` for the module-text gap. """ - rows = self._run("MATCH (a:PyApplication {name: $app}) RETURN a.entrypoint_report_json AS j", app=self.application_name) - raw = rows[0].get("j") if rows else None + # ``properties(a)`` rather than ``a.entrypoint_report_json``: a 1.4.0 graph has no such + # property key at all, and naming one statically makes the server log a warning per call. + rows = self._run("MATCH (a:PyApplication {name: $app}) RETURN properties(a) AS p", app=self.application_name) + raw = rows[0]["p"].get("entrypoint_report_json") if rows else None if raw is not None: r = PyEntrypointReport.model_validate_json(raw) return EntrypointCoverage( @@ -1760,11 +1792,10 @@ def get_callsites_for(self, signatures: List[str]) -> Dict[str, List[PyCallsite] return out def get_external_symbols(self) -> Dict[str, PyExternalSymbol]: - # :PyExternal carries no `_module` property (it isn't owned by one module -- see this - # file's module docstring on MODULE_OWNED_LABELS), so it can't be scoped the way every - # other query here is. Its id embeds this application's own can:// id by construction - # (`/@external//`), which is app-scoping enough on its own -- a - # second application in the same database mints a disjoint id prefix. + # A ghost is owned by no module, so there is no module key to narrow on; its id embeds this + # application's own can:// id by construction (`/@external//`), and + # that prefix is the whole scope -- a second application in the same database mints a + # disjoint one. (This was the one prefix-scoped statement before leg 1.6 made it the rule.) prefix = f"{application_id(self.application_name)}/@external/" rows = self._run( "MATCH (e:PyExternal) WHERE e.id STARTS WITH $prefix RETURN properties(e) AS p", @@ -1890,9 +1921,12 @@ def get_config_readers(self, key: str) -> List[PyCallableOverview]: # no span, so the emitter prunes their start_line/end_line away entirely — the # ``IS NOT NULL`` guard is what stops a span-less vertex being read as "contains everything". # - # Every other query in this file is scoped to the application with ``_module IN $mods``, and so - # is this one: a database may hold several applications, and a same-valued ``file_key`` from a - # different application would otherwise win the ``{_module: pos.path}`` match. + # Every other query in this file is scoped to the application by id prefix, and this one + # narrows further, to the position's own module: ``pos.module_prefix`` is + # ``module_id(app, key) + "/"``, so a same-valued ``file_key`` from a different application + # cannot win, and neither can a module whose key merely extends this one's spelling. On a + # 1.4.1+ graph ``locate_many`` names ``:PyCanNode`` on the callable so that per-module prefix + # seeks the ``:PyCanNode(id)`` range index (see :attr:`_can_node`). _LOCATE_QUERY = ( "UNWIND $positions AS pos " "OPTIONAL MATCH (:PyApplication {name: $app})-[:PY_HAS_MODULE]->(m:PyModule {file_key: pos.path}) " @@ -2047,7 +2081,9 @@ def locate_many(self, positions: Sequence[Tuple[str, int]]) -> List[LocateResult # ``module_prefix`` is the exact inverse of ``_module_key``: ``module_id(app, key) + "/"`` # selects the module's own callables and nothing under a longer key sharing the spelling. rows = self._run( - self._LOCATE_QUERY, + # One label, swapped in per graph generation (see _can_node); the class-level statement + # keeps the spelling every served graph accepts, and test_locate pins both. + self._LOCATE_QUERY.replace("OPTIONAL MATCH (c:PyCallable) ", f"OPTIONAL MATCH (c:PyCallable{self._can_node}) ", 1), app=self.application_name, positions=[ {"idx": i, "path": key, "module_prefix": module_id(self.application_name, key) + "/", "line": line} diff --git a/tests/analysis/python/test_entrypoints.py b/tests/analysis/python/test_entrypoints.py index 5c3391fa..84ae791e 100644 --- a/tests/analysis/python/test_entrypoints.py +++ b/tests/analysis/python/test_entrypoints.py @@ -272,9 +272,9 @@ def test_entrypoint_coverage_over_neo4j_is_read_from_the_application_node(): ) local = object.__new__(PyCodeanalyzer) local.application = PyApplication(symbol_table={}, entrypoint_report=report) - row = {"j": json.dumps(report.model_dump(mode="json"), sort_keys=True)} + row = {"p": {"analyzer_version": "1.4.1", "entrypoint_report_json": json.dumps(report.model_dump(mode="json"), sort_keys=True)}} backend = _neo4j_backend() - with patch.object(PyNeo4jBackend, "_run", side_effect=_run_keyed({"a.entrypoint_report_json AS j": [row]})): + with patch.object(PyNeo4jBackend, "_run", side_effect=_run_keyed({"RETURN properties(a) AS p": [row]})): coverage = backend.get_entrypoint_coverage() assert coverage == local.get_entrypoint_coverage() assert coverage.diagnostics == [] @@ -284,7 +284,7 @@ def test_entrypoint_coverage_over_a_graph_without_the_report_says_it_cannot_answ """A 1.4.0 graph never had the property -- say so via a diagnostic rather than fabricate a clean-looking empty report.""" backend = _neo4j_backend() - with patch.object(PyNeo4jBackend, "_run", side_effect=_run_keyed({"a.entrypoint_report_json AS j": [{"j": None}]})): + with patch.object(PyNeo4jBackend, "_run", side_effect=_run_keyed({"RETURN properties(a) AS p": [{"p": {"analyzer_version": "1.4.0"}}]})): coverage = backend.get_entrypoint_coverage() assert len(coverage.diagnostics) == 1 assert coverage.diagnostics[0].code == "entrypoint_report_unavailable" diff --git a/tests/analysis/python/test_locate.py b/tests/analysis/python/test_locate.py index f11e891d..870c896b 100644 --- a/tests/analysis/python/test_locate.py +++ b/tests/analysis/python/test_locate.py @@ -21,6 +21,8 @@ ambiguous empty is a defect (see ``cldk/analysis/commons/results.py``). """ +from cldk.analysis.python.neo4j import PyNeo4jBackend + def test_locate_inside_callable(py): r = py.locate("src/app.py", 19) @@ -263,6 +265,22 @@ def test_locate_query_is_scoped_to_the_application(py, fake_driver): assert "_module" not in statement, "the graph stores no _module property to match on" +def test_locate_names_pycannode_only_on_a_graph_that_has_it(py, fake_driver): + """The per-module prefix seeks the ``:PyCanNode(id)`` range index a 1.4.1 graph carries + (measured on odoo: 40 positions 399 -> 89 ms); a 1.4.0 graph has no such label and naming it + would match nothing. Both spellings are pinned because the swap is a string edit on the + statement -- a renamed anchor would silently lose the seek, never the answer.""" + py.locate("src/app.py", 21) + assert "OPTIONAL MATCH (c:PyCallable:PyCanNode) " in next(s for s in fake_driver.statements if "UNWIND $positions AS pos" in s) + + fake_driver.statements.clear() + fake_driver.analyzer_version = "1.4.0" + old = PyNeo4jBackend._from_driver(fake_driver, application_name="app") + assert old.locate("src/app.py", 21).callable is not None, "the 1.4.0 spelling must still answer" + statement = next(s for s in fake_driver.statements if "UNWIND $positions AS pos" in s) + assert "OPTIONAL MATCH (c:PyCallable) " in statement and "PyCanNode" not in statement + + def test_locate_scope_is_actually_honoured(py): """Not just present in the text: attach as another application and no callable matches.""" py.application_name = "some_other_application" From 0a7e5c03438f0ca105a7fd27bf040054befcb210 Mon Sep 17 00:00:00 2001 From: Rahul Krishna Date: Sun, 6 Sep 2026 03:46:55 -0400 Subject: [PATCH 08/13] test(python): time the query with coverage paused; record closure sizes per analyzer generation The two N+1 wall-clock tests measured the coverage tracer, not the round trips: ~10 s under --no-cov, ~15.3 s with sys.settrace on, against a 15 s ceiling. They run under pytest-cov's no_cover marker now (9.6 s / 9.0 s under --cov), and the ceiling still fires on a stalled bulk fetch (75.7 s) while the round-trip ceiling catches a reintroduced N+1 (1,667). The three counts pinned on the 1.4.0 graph -- PY_DDG total, the pathological forward slice, the default-depth seed's unbounded closure -- are recorded per analyzer_version and looked up through a fixture, because the 1.4.0 and 1.4.1 graphs of the same checkout differ by 2,724 level-4 vertices whose existence follows call resolution, which differs between the two emit runs and not between the two analyzer versions. Each test's claim is unchanged. --- tests/analysis/python/conftest.py | 8 ++++ .../python/test_bounded_enumeration.py | 8 ++++ tests/analysis/python/test_dataflow.py | 48 ++++++++++++------- tests/analysis/python/test_e2e_neo4j_live.py | 11 +++-- 4 files changed, 54 insertions(+), 21 deletions(-) diff --git a/tests/analysis/python/conftest.py b/tests/analysis/python/conftest.py index 80854427..5b552e0e 100644 --- a/tests/analysis/python/conftest.py +++ b/tests/analysis/python/conftest.py @@ -571,6 +571,14 @@ def live_analysis(): facade.backend.close() +@pytest.fixture(scope="session") +def live_analyzer_version(live_analysis) -> str: + """The ``analyzer_version`` the live graph's ``:PyApplication`` carries (``"1.4.0"`` on 7688, + ``"1.4.1"`` on 7689) -- the key a count recorded against one emitter run is looked up by.""" + rows = live_analysis.backend._run("MATCH (a:PyApplication {name: $app}) RETURN a.analyzer_version AS v", app=LIVE_NEO4J_APP) + return rows[0]["v"] + + @pytest.fixture def busy_callable() -> str: """A live-graph callable with a non-trivial data dependence, named the way a caller would. diff --git a/tests/analysis/python/test_bounded_enumeration.py b/tests/analysis/python/test_bounded_enumeration.py index d8db756d..aee955b2 100644 --- a/tests/analysis/python/test_bounded_enumeration.py +++ b/tests/analysis/python/test_bounded_enumeration.py @@ -69,6 +69,12 @@ #: Spec §8's wall-clock target for ``get_symbol_table`` / ``get_classes`` on the odoo graph, and #: the ceiling asserted: headroom over the measured ~10s so a slow machine does not flake at the #: boundary, low enough to catch a regression back towards the minutes this leg removed. +#: +#: The two timed tests run with coverage paused (pytest-cov's ``no_cover`` marker) because the +#: clock is meant to measure the query, not the tracer: measured on both live graphs, the same +#: call is 10.2-10.5 s under ``--no-cov`` and ~15.3 s with the coverage tracer on -- the +#: reconstruction is a Python-side walk over ~950k rows, and ``sys.settrace`` costs it half again. +#: Under the tracer the ceiling passed by luck, and only sometimes. _WALL_CLOCK_TARGET = 10 _WALL_CLOCK_CEILING = 15 @@ -104,6 +110,7 @@ def _assert_children_survived(classes: List[PyClass], callables: List[PyCallable assert any(m.callables for m in callables), "no nested callables survived the collapse" +@pytest.mark.no_cover # the clock measures the query, not the tracer -- see _WALL_CLOCK_CEILING def test_symbol_table_is_not_n_plus_one(live_analysis, count_round_trips): n = count_round_trips(live_analysis) started = time.monotonic() @@ -122,6 +129,7 @@ def test_symbol_table_is_not_n_plus_one(live_analysis, count_round_trips): assert any(m.variables for m in table.values()), "no module variables survived the collapse" +@pytest.mark.no_cover # as above def test_classes_is_not_n_plus_one(live_analysis, count_round_trips): n = count_round_trips(live_analysis) started = time.monotonic() diff --git a/tests/analysis/python/test_dataflow.py b/tests/analysis/python/test_dataflow.py index 5880c70d..e0c5e9d6 100644 --- a/tests/analysis/python/test_dataflow.py +++ b/tests/analysis/python/test_dataflow.py @@ -267,7 +267,8 @@ def test_below_level_three_the_backend_refuses_rather_than_returning_empty(local # ---------------------------------------------------------------------------------------------- # Pagination (E5). Per-callable is a *scoping* bound, not a *size* one: measured on odoo-slim-19 -# one callable's DDG is 1,386,918 edges, 27% of the whole application's 5,134,655. The bound that +# one callable's DDG is 1,386,918 edges, 27% of the whole application's 5,134,655 (1.4.0 graph; +# 5,129,295 on 1.4.1). The bound that # was missing is on the response, and the ruling is to paginate rather than truncate — truncation # throws away edges a caller may need, a page keeps every one of them reachable. # @@ -375,7 +376,7 @@ def test_local_pages_are_the_emitter_rows_sliced_by_the_same_order(local_l4): def test_the_local_sort_key_is_total_on_a_real_analyzer_run(local_l4): """Keyset paging resumes *after* a key, so a repeated key would drop its twin. Both backends are safe only while the key is unique; on the live graph it is (measured: 0 duplicate tuples - across all 5,134,655 PY_DDG, 247,906 PY_CFG_NEXT and 139,065 PY_CDG edges, and the emitter's + across all 5,134,655 PY_DDG, 247,906 PY_CFG_NEXT and 139,065 PY_CDG edges of the 1.4.0 graph, and the emitter's MERGE makes it structurally so there), and this is the same claim for the local path, where nothing dedupes.""" for get, key in ((local_l4.get_cfg, cfg_sort_key), (local_l4.get_cdg, cdg_sort_key), (local_l4.get_ddg, ddg_sort_key)): @@ -417,7 +418,7 @@ def test_a_cursor_from_somewhere_else_is_refused_rather_than_answered(local_l4): # # WHAT THE MEASUREMENT DECIDED. The sibling accessors above paginate; these cap. The reason is # the measured shape of a slice on odoo-slim-19, over the five SDG relationship types -# (PY_DDG 5,134,655 / PY_CDG 139,065 / PY_PARAM_IN 229,035 / PY_PARAM_OUT 133,267 / +# (1.4.0 graph: PY_DDG 5,134,655 / PY_CDG 139,065 / PY_PARAM_IN 229,035 / PY_PARAM_OUT 133,267 / # PY_SUMMARY 453,398 -- all five present, verified against ``CALL db.relationshipTypes()``): # # slice_backward over 200 random ``formal_in`` seeds: median 1, p95 195,790, max 196,117 @@ -437,7 +438,20 @@ def test_a_cursor_from_somewhere_else_is_refused_rather_than_answered(local_l4): # two rather than stored, so they cannot disagree. # ============================================================================================== HEAVY_STATEMENT_SLICE = 195_785 # backward slice of any statement in configurator_apply -HEAVY_FORWARD_SLICE = 440_270 # forward slice of its ``kwargs`` -- half the application + +#: Whole-closure sizes recorded per analyzer generation, keyed by the graph's ``analyzer_version`` +#: (the ``live_analyzer_version`` fixture). They are facts about an emitter run, not contracts: +#: the 1.4.0 and 1.4.1 graphs of the same checkout differ by 2,724 body nodes, every one a level-4 +#: ``formal_in``/``formal_out``/``actual_in``/``actual_out`` vertex (the level-1 node set is +#: identical), because which callee a call lands on decides whose globals become formals up the +#: chain, and 869 ``PY_CALLS`` edges resolve differently between the two runs. A closure over +#: those vertices moves with them; the tests' claims -- bounded, true size reported, complete at +#: the default depth -- do not. An unrecorded generation fails on the lookup rather than passing +#: against a stale number. +RECORDED = { + "1.4.0": {"ddg_edges": 5_134_655, "heavy_forward_slice": 440_270, "depth_seed_unbounded": 195_790}, + "1.4.1": {"ddg_edges": 5_129_295, "heavy_forward_slice": 438_017, "depth_seed_unbounded": 195_263}, +} @live_only @@ -510,31 +524,32 @@ def test_depth_bounds_a_slice_without_capping_it(live_analysis, busy_callable): @live_only -def test_the_pathological_slice_is_bounded_and_reports_its_true_size(live_analysis): +def test_the_pathological_slice_is_bounded_and_reports_its_true_size(live_analysis, live_analyzer_version): """The case the cap exists for, on the callable this leg keeps returning to. - ``kwargs`` of ``configurator_apply`` reaches 440,270 nodes — **half** of the application's - 885,218 body nodes — and 27% of its DDG edges live in this one callable. Ten nodes come back, + ``kwargs`` of ``configurator_apply`` reaches 440,270 nodes on the 1.4.0 graph (438,017 on 1.4.1) + — **half** of the application's 885,218 body nodes — and 27% of its DDG edges live in this one + callable. Ten nodes come back, and the result says how many there were, in one call and in about a second. Backward from the same value is one node, because nothing calls it: the two directions of the same seed differ by five orders of magnitude, which is the distribution the cap exists for.""" fwd = live_analysis.slice_forward("kwargs", within=HEAVY_CALLABLE, depth=None, max_nodes=10) - assert len(fwd.nodes) == 10 and fwd.total == HEAVY_FORWARD_SLICE and not fwd.complete + assert len(fwd.nodes) == 10 and fwd.total == RECORDED[live_analyzer_version]["heavy_forward_slice"] and not fwd.complete back = live_analysis.slice_backward("kwargs", within=HEAVY_CALLABLE, depth=None) assert back.total == 1 and back.complete #: A *global* the callable reads, in a callable nothing about this test needs to be heavy: its -#: backward slice is 195,790 nodes unbounded and 76 at the default depth. The callable name is -#: unique in the application, so the seed is addressable the way a caller would say it. +#: backward slice is 195,790 nodes unbounded (1.4.0 graph; 195,263 on 1.4.1 -- see ``RECORDED``) +#: and 76 at the default depth on both. The callable name is unique in the application, so the +#: seed is addressable the way a caller would say it. DEPTH_SEED_CALLABLE = "odoo.tools.mail.email_domain_extract" DEPTH_SEED_VALUE = "found_email" DEPTH_SEED_AT_DEFAULT = 76 -DEPTH_SEED_UNBOUNDED = 195_790 @live_only -def test_the_default_depth_answers_completely_where_unbounded_truncates(live_analysis): +def test_the_default_depth_answers_completely_where_unbounded_truncates(live_analysis, live_analyzer_version): """The reason the default is finite (Task 6.1). Unbounded, this seed's backward slice is a fifth of the application and the caller gets 10,000 arbitrary nodes of it — 5% of a closure, flagged incomplete and useless. At the default depth the same call answers the narrower @@ -545,7 +560,7 @@ def test_the_default_depth_answers_completely_where_unbounded_truncates(live_ana assert len(near.nodes) == near.total and near.complete whole = live_analysis.slice_backward(DEPTH_SEED_VALUE, within=DEPTH_SEED_CALLABLE, depth=None) - assert whole.total == DEPTH_SEED_UNBOUNDED, "depth=None still means the whole closure" + assert whole.total == RECORDED[live_analyzer_version]["depth_seed_unbounded"], "depth=None still means the whole closure" assert len(whole.nodes) == 10_000 and not whole.complete assert {n.ref for n in near.nodes} <= {n.ref for n in whole.nodes} or near.total < whole.total @@ -1074,15 +1089,16 @@ def test_an_empty_describe_costs_nothing(live_analysis, count_round_trips): @live_only -def test_prov_is_a_singleton_on_every_ddg_edge(live_analysis): +def test_prov_is_a_singleton_on_every_ddg_edge(live_analysis, live_analyzer_version): """The measurement ``weakest`` rests on, re-checked against the graph rather than trusted. If a future analyzer generation started emitting several provenances on one edge, "the weakest hop" would become a question about how to *combine* a set, and :func:`prov_rank`'s conservative - choice would start being observable. + choice would start being observable. The one row is the claim; its count is the generation's + whole ``PY_DDG`` (``RECORDED``), so a graph that lost half its edges could not pass either. """ rows = live_analysis.backend._run("MATCH ()-[r:PY_DDG]->() RETURN size(r.prov) AS n, count(*) AS c ORDER BY n") - assert [(r["n"], r["c"]) for r in rows] == [(1, 5_134_655)] + assert [(r["n"], r["c"]) for r in rows] == [(1, RECORDED[live_analyzer_version]["ddg_edges"])] def test_weakest_is_the_most_approximate_hop_not_the_alphabetically_first(): diff --git a/tests/analysis/python/test_e2e_neo4j_live.py b/tests/analysis/python/test_e2e_neo4j_live.py index f9c473eb..b39eafb9 100644 --- a/tests/analysis/python/test_e2e_neo4j_live.py +++ b/tests/analysis/python/test_e2e_neo4j_live.py @@ -778,11 +778,12 @@ def test_get_unresolved_config_reads_returns_real_edges(analysis, cypher): def test_entrypoints_faithfully_report_what_the_graph_says(analysis, cypher): """**Not** an assertion that Odoo has entrypoints. - ``is_entrypoint`` is ``FALSE`` on all 15,549 callables and all 1,656 classes of this graph: the - analyzer's detection pass found nothing on a framework built entirely from HTTP routes. That is - upstream under-detection, not an SDK defect, and this suite must not paper over it by demanding - a non-zero count. What the SDK owes the caller is fidelity — exactly as many entrypoints as the - graph flags, no more and no fewer. + What the graph flags depends on the analyzer generation. The 1.4.0 graph has ``is_entrypoint`` + ``FALSE`` on all 15,549 callables and all 1,656 classes: the pass shipped no Odoo rules, on a + framework built entirely from HTTP routes (upstream under-detection, python-sdk#177). The + 1.4.1 graph flags 534 callables and 94 classes (#182/#185: ``@http.route`` methods and + ``http.Controller`` subclasses). Neither number is this suite's to demand. What the SDK owes + the caller is fidelity — exactly as many entrypoints as the graph flags, no more and no fewer. """ flagged_callables = cypher("MATCH (c:PyCallable) WHERE c.is_entrypoint = true RETURN count(c) AS c")[0]["c"] flagged_classes = cypher("MATCH (c:PyClass) WHERE c.is_entrypoint = true RETURN count(c) AS c")[0]["c"] From 51a3415dabe1f198db8ff61665e07ee459b34b5a Mon Sep 17 00:00:00 2001 From: Rahul Krishna Date: Sun, 6 Sep 2026 03:46:55 -0400 Subject: [PATCH 09/13] docs(python): record id-prefix scoping, the 1.4.1 pin, and the graph version floor --- CHANGELOG.md | 52 +++++++++++++-- docs/agent-api-reference.md | 27 ++++++-- .../2026-09-06-leg-1.6-id-prefix-scoping.md | 65 +++++++++++-------- ...05-leg-1.5-bounded-queries-and-dataflow.md | 4 +- 4 files changed, 107 insertions(+), 41 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index b1a011d5..e74209d6 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -7,11 +7,30 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ## [Unreleased] -Python legs 1 and 1.5 of the CLDK 2.0 agent-facing query facade (see -`docs/design/specs/2026-09-03-agent-facing-query-facade.md` and -`docs/design/specs/2026-09-05-leg-1.5-bounded-queries-and-dataflow.md`). Targets `2.0.0-rc.1`. +Python legs 1, 1.5 and 1.6 of the CLDK 2.0 agent-facing query facade (see +`docs/design/specs/2026-09-03-agent-facing-query-facade.md`, +`docs/design/specs/2026-09-05-leg-1.5-bounded-queries-and-dataflow.md` and +`docs/design/specs/2026-09-06-leg-1.6-id-prefix-scoping.md`). Targets `2.0.0-rc.1`. ### Changed +- **Pinned `codeanalyzer-python` 1.4.0 → 1.4.1** (`pyproject.toml` `dependencies` and + `[tool.backend-versions]`). 1.4.1 removes the `_module` node property from every graph it emits + (upstream #183), so the Neo4j backend no longer scopes on it: **every statement is scoped to the + application by its `can://` id prefix** (`n.id STARTS WITH 'can://python//'`), the rule the + analyzer's own destructive statements use, and a callable's repo-relative `path` is derived from + its id and verified against the application's module keys rather than projected from a property. + No accessor changes name, signature, return type or value; the `_module` scoping is simply gone. + Ghosts (`:PyExternal`) fall inside the prefix, so every call-graph walk pins its traversal source + to `:PyCallable` by label -- an `@external` node is still reached and never traversed through. +- **The schema probe reads the graph's analyzer generation.** Attaching to a graph whose + `:PyApplication.analyzer_version` is below **1.4.0** (or missing) raises `GraphSchemaMismatch` + naming the version found and the floor -- before this it was served with silent empties. A + **1.4.0** graph is served with identical results and one `WARNING` that scoped statements scan + rather than seek (it carries no `:PyCanNode` index). A **1.4.1** graph attaches silently, and + `locate` / `locate_many` name `:PyCanNode` on it so the per-module prefix seeks the range index + (40 positions: 399 → 89 ms). Statements whose prefix is the whole application stay unlabelled: + measured, the seek over 955,961 application nodes is 2–20× slower than the `:PyCallable` label + scan they use. - **BREAKING: Neo4j graph vocabulary migration.** The Python Neo4j backend (`cldk.analysis.python.neo4j.PyNeo4jBackend`) now queries `PyBodyNode` / `PY_HAS_BODY_NODE` instead of the pre-1.4.0 `PyCallSite` / `PY_HAS_CALLSITE` / `PySymbol` vocabulary, matching what @@ -43,8 +62,9 @@ Python legs 1 and 1.5 of the CLDK 2.0 agent-facing query facade (see `PyClassOverview.path` and the keys of `get_symbol_table()`, which it previously could not be joined against. Both backends changed together: the Neo4j projections behind `get_callables_overview()`, `get_decorated_callables()`, `get_entrypoints()` and - `get_config_readers()` now read `:PyCallable._module` rather than `:PyCallable.path`, and the - local backend projects the symbol table key rather than `PyCallable.path`. + `get_config_readers()` now project the module key rather than `:PyCallable.path` (from the + `_module` property on a 1.4.0 graph in leg 1; derived from the node's `can://` id since leg 1.6), + and the local backend projects the symbol table key rather than `PyCallable.path`. **Migration:** a caller that stored or persisted these paths will see different strings for the same callable, and one that stripped a project-root prefix off them must stop. @@ -104,7 +124,7 @@ Python legs 1 and 1.5 of the CLDK 2.0 agent-facing query facade (see | SDK version | Requires `codeanalyzer-python` | Graph vocabulary | | --- | --- | --- | | <= 1.5.0 | 0.3.x | `PyCallSite` / `PY_HAS_CALLSITE` / `PySymbol` | -| 2.0.0-rc.1 (this) | >= 1.4.0 | `PyBodyNode` / `PY_HAS_BODY_NODE` | +| 2.0.0-rc.1 (this) | 1.4.1 pinned; graphs emitted by >= 1.4.0 served (1.4.0 scanned, 1.4.1 indexed) | `PyBodyNode` / `PY_HAS_BODY_NODE`, scope by `can://` id prefix | ### Added - **Slices and reachability: `slice_backward(src, within=)` / `slice_forward(src, within=)` / @@ -312,6 +332,23 @@ Python legs 1 and 1.5 of the CLDK 2.0 agent-facing query facade (see schema probe at attach time; see the breaking-change note above. ### Fixed +- **`get_entrypoint_coverage()` over Neo4j reads the projected report.** It answered with an + `entrypoint_report_unavailable` diagnostic unconditionally; codeanalyzer-python 1.4.1 (#182) + projects `entrypoint_frameworks` / `entrypoint_report_json` onto `:PyApplication`, so on such a + graph the answer is now the pass's own `PyEntrypointReport`, field for field what the local + backend returns. The diagnostic survives only for a graph that genuinely lacks the property (one + emitted by 1.4.0). +- **Resolved through the 1.4.1 pin** -- python-sdk's upstream reports #176 / #177 / #178, fixed + in codeanalyzer-python as #180 (body nodes and parameters carry their `id` in `analysis.json`; + the SDK's composed body-node id is now pinned equal to the analyzer's own per run), #182 / #185 + (entrypoint report projected, Odoo `@http.route` / `http.Controller` detected: 534 callables and + 94 classes on the same checkout that 1.4.0 flagged 0 / 0), #181 (`PY_EXTENDS` is actually emitted: + 1,573 edges where every 2.0.0 graph had none) and #183 (`_module` retired, `:PyCanNode` range + index on `id`, `:PyExternal` ghosts under the application prefix). +- **The N+1 timing assertion measured the coverage tracer, not the query.** The two timed live + tests (`get_symbol_table` / `get_classes` under 15 s) ran under `sys.settrace`, which adds ~5 s + to a ~10 s Python-side reconstruction; they passed the ceiling by luck. They now run with coverage + paused (`pytest-cov`'s `no_cover`), so the ceiling measures the round trips it exists to bound. - **Java: a missing JDK cache root is refused, not crashed on.** `JCodeanalyzer._get_codeanalyzer_exec` now raises `CodeanalyzerExecutionException` ("no cache directory and no project directory") when neither is available, instead of letting `ensure_jdk` fail with `TypeError: ... not @@ -328,7 +365,8 @@ Python legs 1 and 1.5 of the CLDK 2.0 agent-facing query facade (see child fetches (`get_class()` and everything reconstructing one declaration) matched by a bare `{signature: $sig}` / `{file_key: $fk}` while their bulk twins were application-scoped, so in a database holding two applications `get_class()` merged another application's members and - `get_all_classes()` did not. All twelve carry the same `_module IN $mods` predicate now, and so + `get_all_classes()` did not. All twelve carry the same application-scope predicate now (`_module + IN $mods` when this landed, the `can://` id prefix since leg 1.6), and so do the leg-1.5 call-graph statements behind `reaches`, `backward_cone`, `call_paths_between` and `flows_to_call`, which had matched by signature unscoped too. Statements keyed only by a body-node or ghost id (`slice_*`, `paths_between`, the value-reachability predicate) are scoped diff --git a/docs/agent-api-reference.md b/docs/agent-api-reference.md index f22284e1..9b1bcfd8 100644 --- a/docs/agent-api-reference.md +++ b/docs/agent-api-reference.md @@ -3,10 +3,11 @@ What you can ask a CLDK Python analysis, what comes back, and what will mislead you if you don't know it. Written for an agent composing queries at runtime. -Status: legs 1 and 1.5 (`codeanalyzer-python` 1.4.0) — every accessor below is implemented on both -backends and exercised against a live graph. Design records: -`docs/design/specs/2026-09-03-agent-facing-query-facade.md` and -`docs/design/specs/2026-09-05-leg-1.5-bounded-queries-and-dataflow.md`. +Status: legs 1, 1.5 and 1.6 (`codeanalyzer-python` 1.4.1 pinned; graphs emitted by 1.4.0 or newer +served) — every accessor below is implemented on both backends and exercised against live graphs of +both generations. Design records: `docs/design/specs/2026-09-03-agent-facing-query-facade.md`, +`docs/design/specs/2026-09-05-leg-1.5-bounded-queries-and-dataflow.md` and +`docs/design/specs/2026-09-06-leg-1.6-id-prefix-scoping.md`. --- @@ -29,6 +30,15 @@ Attaching raises `GraphSchemaMismatch` if the graph was built by a different ana That is deliberate: the alternative is every query silently returning zero rows. If you see it, the graph needs re-ingesting — it is not a bug in your query. +**Graph version floor: codeanalyzer-python 1.4.0.** The probe reads `:PyApplication.analyzer_version` +and the message names what it found and the floor. What attaching to each generation does: + +| graph emitted by | attach | behaviour | +| --- | --- | --- | +| < 1.4.0, or no `analyzer_version` | refused (`GraphSchemaMismatch`) | the `can://` id grammar every query scopes on does not exist there | +| 1.4.0 | served, one `WARNING` logged | identical results; the graph has no `:PyCanNode` index, so scoped statements scan rather than seek, and `get_entrypoint_coverage()` reports `entrypoint_report_unavailable` (the report was not projected yet) | +| 1.4.1 and newer | served, silent | `locate` seeks the `:PyCanNode(id)` index; the entrypoint report is read off the graph | + --- ## The four moves @@ -262,8 +272,10 @@ class EntrypointCoverage: **This is the most important gotcha in the API.** The analyzer's entrypoint pass *under-approximates by design* — its own docs say "silence is its failure mode". On a real Odoo -checkout it detects **zero** entrypoints across 15,549 callables, in a framework built entirely -from HTTP routes. +checkout, codeanalyzer-python 1.4.0 detected **zero** entrypoints across 15,549 callables, in a +framework built entirely from HTTP routes; 1.4.1 ships Odoo rules and flags 534 callables and 94 +classes on the same checkout. The number is a fact about the analyzer generation and its rules, +never about the application alone. So an empty `get_entrypoints()` means either "no entrypoints" or "the pass found nothing", and you cannot tell from the list. `get_entrypoint_coverage()` is how you ask. Over a Neo4j graph emitted @@ -356,7 +368,8 @@ tells you the `reaches` call to ask instead. **The per-callable graphs return a page, not a list.** Naming a callable says *which* edges you want, not *how many*: on a real application one callable's DDG is 1,386,918 edges — 27% of the whole -application's 5,134,655 — while 15,520 of its 15,549 callables have fewer than 10,000. So all three +application's 5,134,655 (1.4.0 graph; 5,129,295 on 1.4.1) — while 15,520 of its 15,549 callables +have fewer than 10,000. So all three return an `EdgePage`, and nothing is discarded to make it fit: ```python diff --git a/docs/design/plans/2026-09-06-leg-1.6-id-prefix-scoping.md b/docs/design/plans/2026-09-06-leg-1.6-id-prefix-scoping.md index 5350d7e6..37facfd8 100644 --- a/docs/design/plans/2026-09-06-leg-1.6-id-prefix-scoping.md +++ b/docs/design/plans/2026-09-06-leg-1.6-id-prefix-scoping.md @@ -43,7 +43,7 @@ Pure measurement before any query changes: what does 1.4.1 change for the **loca - Consumes: nothing. - Produces: a written list of every test whose result or meaning changed under 1.4.1, with the emitter change each one traces to. -- [ ] **Step 1: Bump the pin and re-sync** +- [x] **Step 1: Bump the pin and re-sync** ``` sed -i '' 's/codeanalyzer-python==1.4.0/codeanalyzer-python==1.4.1/' pyproject.toml @@ -53,7 +53,7 @@ uv run python -c "import importlib.metadata as m; print(m.version('codeanalyzer- ``` Expected: `1.4.1`. -- [ ] **Step 2: Run the offline suite and read every change, not just the count** +- [x] **Step 2: Run the offline suite and read every change, not just the count** ``` uv run pytest tests/analysis/python/ tests/models/python/ -q --ignore=tests/analysis/python/test_python_neo4j_backend.py -p no:cacheprovider @@ -63,7 +63,7 @@ Against the floor (373 / 141). For every test that newly fails, newly passes, or - **#180 — body nodes and parameters now carry `id` in `analysis.json`.** `cldk/analysis/python/codeanalyzer/codeanalyzer.py:125` `body_node_id()` composes the id from `callable.id` and the body key. Add a test asserting that, for every body node in a real level-4 run over `tests/`' small fixtures, the composed id **equals** the analyzer's own `BodyNode.id` when present. This turns leg 1.5's "verified 885,218 / 885,218 by construction" into a per-run assertion. If any differ, that is a finding to report, not to reconcile silently. - **#182 — the entrypoint report and Odoo detection.** Any test pinning "zero entrypoints" or `entrypoint_report_unavailable` on a fixture is recording 1.4.0 behaviour; re-read each and update the *expectation*, keeping the test's claim. -- [ ] **Step 3: Confirm the Neo4j backend is now broken against 7689, on purpose** +- [x] **Step 3: Confirm the Neo4j backend is now broken against 7689, on purpose** Once 7689 is populated, attach and run one scoped call: @@ -76,7 +76,7 @@ print(len(py.get_callables_overview()))" ``` Expected before Task 1: **0** with no error — record it. This is the defect the leg fixes; Task 1's first test pins that it becomes 15,549. -- [ ] **Step 4: Commit** +- [x] **Step 4: Commit** ``` git add pyproject.toml uv.lock tests/analysis/python/ @@ -96,14 +96,14 @@ git commit -m "chore(python): pin codeanalyzer-python 1.4.1 and pin the body-nod - Consumes: `codeanalyzer.schema.ids.application_id(app) -> "can://python/"`, `module_id(app, file_key)`. - Produces: `PyNeo4jBackend._scope_prefix: str` (`application_id(app) + "/"`), `PyNeo4jBackend._module_prefixes(keys: Sequence[str]) -> list[str]` for narrowing, `_probe_schema` raising `GraphSchemaMismatch` below 1.4.0 and logging once on 1.4.0. -- [ ] **Step 1: Widen the audit first (F7) and watch it go red** +- [x] **Step 1: Widen the audit first (F7) and watch it go red** In `test_neo4j_multi_application_scope.py`: add `_MATCHES_BY_PREFIX = re.compile(r"\.id STARTS WITH \$")`; change the rule to "by-signature statements must carry `IN $mods` **or** a prefix predicate"; change `_class_level_statements()` to accept any class-attribute string containing a Cypher keyword at its start (`MATCH`, `OPTIONAL MATCH`, `UNWIND`, `CALL`), not only `startswith("MATCH")`. Then change `_fake_two_app_cypher` so its scope filter is `c["id"].startswith(prefix_param)` for prefix statements and keep the `_module`/`IN $mods` branch only until Step 4 removes it. Add a test that constructs the fake graph with **no `_module` property at all** (that is what a 1.4.1 graph is) and asserts the current backend returns empty — the red that Step 4 turns green. Run: `uv run pytest tests/analysis/python/test_neo4j_multi_application_scope.py -q` Expected: the new no-`_module` test FAILS (returns empty); `test_the_audit_sees_the_dataflow_statements_too` may now list more statements than 11 — record the new count. -- [ ] **Step 2: Add the scope helpers** +- [x] **Step 2: Add the scope helpers** ```python from codeanalyzer.schema.ids import application_id, module_id @@ -117,7 +117,7 @@ def _module_prefixes(self, keys: Sequence[str]) -> list[str]: ``` Keep `_load_module_keys` / `self._modules`: they still feed `scope_paths`, `resolve_module_key`, and narrowing. -- [ ] **Step 3: Replace the 49 scope predicates, one statement family at a time** +- [x] **Step 3: Replace the 49 scope predicates, one statement family at a time** For each site in the inventory (`_BULK_CHILD_QUERIES` ×7, `_probe_resolution_edges`, `_callable_full` ×4, `_class_full` ×3, `_call_rows`, `_bounded_call_rows` ×3, `get_python_file`, `get_all_classes`, `get_class`, `get_callables_overview`, `get_method_bodies`, `get_source`, `_RESOLVE_CALLABLE_QUERY`, `resolve_value`, `_OWN_EDGES`, `_REACHES`, `_CONE`, `_CALLERS`, `_CALLEES`, `_CALL_PATHS` ×2, `_CALLEE_VALUES`, `_SOURCES`, `get_decorated_callables`, `get_entrypoints`, `get_entrypoint_classes`, `get_callsites_for`, `get_config_uses`, `get_config_readers`, `_LOCATE_QUERY:1791`): @@ -129,7 +129,7 @@ Where the statement is a **narrowed bulk fetch** (anything reached through `_pre Run after each family: `uv run pytest tests/analysis/python/test_neo4j_multi_application_scope.py tests/analysis/python/test_scoping_keywords.py -q`. -- [ ] **Step 4: Move the 3 hop-scopes, preserving the ghost rule (F5)** +- [x] **Step 4: Move the 3 hop-scopes, preserving the ghost rule (F5)** `_bounded_call_rows:790`, `_REACHES:1373`, `_CONE:1392`: `WHERE a._module IN $mods` → `WHERE a:PyCallable AND a.id STARTS WITH $prefix` (the traversal **source** must be a declared callable; a ghost is reached, never traversed through). Update `test_scoping_keywords.py:447`'s string assertion to the new pattern. @@ -137,7 +137,7 @@ Write the live test the spec's DoD names: find a `callable → @external → cal Then delete the `_module` branch from the fake server. -- [ ] **Step 5: The probe (F2)** +- [x] **Step 5: The probe (F2)** In `_probe_schema`, after the relationship-type check: @@ -151,7 +151,7 @@ if version < (1, 4, 1): ``` Reuse `GraphSchemaMismatch`; extend it if it cannot carry a version message. Unit-test all three branches with the fake driver; live-test that 7688 warns and 7689 is silent. -- [ ] **Step 6: Run everything against both graphs** +- [x] **Step 6: Run everything against both graphs** ``` uv run pytest tests/analysis/python/ tests/models/python/ -q --ignore=tests/analysis/python/test_python_neo4j_backend.py # offline @@ -160,7 +160,7 @@ CLDK_TEST_NEO4J_URI=bolt://localhost:7688 ... uv run pytest tests/analysis/pytho ``` Expected: offline ≥ floor; 7688 ≥ 516 minus only tests whose 1.4.1-specific expectations Task 0 changed; 7689 — every failure attributed to a path-projection site (Task 2) or a named 1.4.1 emitter change, nothing else. Report the three summary lines verbatim. -- [ ] **Step 7: Commit** +- [x] **Step 7: Commit** ``` git add cldk/analysis/python/neo4j/neo4j_backend.py tests/analysis/python/test_neo4j_multi_application_scope.py tests/analysis/python/test_scoping_keywords.py tests/analysis/python/test_e2e_neo4j_live.py @@ -186,7 +186,7 @@ analyzer_version: below 1.4.0 refuses, 1.4.0 warns once about the missing index. - Consumes: `self._scope_prefix`, `self._modules` (the verified key set), `module_id`. - Produces: `module_key_of(node_id: str, prefix: str, known: Collection[str]) -> str` — pure, raises on a non-member. -- [ ] **Step 1: Write the helper's tests first** +- [x] **Step 1: Write the helper's tests first** ```python def test_module_key_is_the_id_segment_up_to_the_first_py_boundary(): @@ -207,7 +207,7 @@ def test_a_ghost_id_has_no_module_key(): ``` Run: expected `ImportError`. -- [ ] **Step 2: Implement it — longest known-key match, not a split** +- [x] **Step 2: Implement it — longest known-key match, not a split** ```python def module_key_of(node_id: str, prefix: str, known: Collection[str]) -> str: @@ -227,19 +227,19 @@ def module_key_of(node_id: str, prefix: str, known: Collection[str]) -> str: ``` If `known` is large (1,626 on Odoo) and this sits on a hot path, index `known` once by first segment; measure before optimising. -- [ ] **Step 3: Replace the 13 projections** +- [x] **Step 3: Replace the 13 projections** Each `c._module AS path` / `AS file` / `AS fk` returns `c.id AS id` instead, and the Python side derives the key with `module_key_of(row["id"], self._scope_prefix, self._module_set)`. Sites: `get_python_file:847b`, `_OVERVIEW_PROJECTION:1078`, `_RESOLVE_CALLABLE_QUERY:1133`, `_SLICE:1321`, `_CONE:1394`, `_CALLERS:1413b`, `_CALLEES:1414b`, `_PATHS:1463` (its `head([...| c._module])` becomes `head([... | c.id])`), `_CALL_PATHS:1483`, `get_entrypoint_classes:1598`, `get_config_readers:1759`. For `:PyClass` rows, the id grammar is the same (`/`). -- [ ] **Step 4: `_LOCATE_QUERY`'s match key** +- [x] **Step 4: `_LOCATE_QUERY`'s match key** `OPTIONAL MATCH (c:PyCallable {_module: pos.path})` → `OPTIONAL MATCH (c:PyCallable) WHERE c.id STARTS WITH pos.module_prefix`, where the Python side supplies `module_prefix = module_id(app, resolved_key) + "/"` per position. The `module_scope` diagnostic keeps naming the key, never the prefix (E6). -- [ ] **Step 5: Run the vocabulary tests on both graphs** +- [x] **Step 5: Run the vocabulary tests on both graphs** `test_python_bulk_accessors.py::test_overview_path_is_the_repo_relative_module_key`, `test_e2e_neo4j_live.py::test_paths_share_one_vocabulary`, `::test_overview_path_joins_locate_and_class_overview`, `test_locate.py`, and the leg-1.5 `test_locate_node_id_joins_to_the_graph` — on 7689 and 7688. Expected: all pass; path values byte-identical to the 1.4.0 run. -- [ ] **Step 6: Commit** +- [x] **Step 6: Commit** ``` git commit -m "fix(python): derive a callable's module key from its can:// id @@ -257,23 +257,34 @@ module keys -- verified, never split, raised when it cannot be." - Modify: `CHANGELOG.md` `[Unreleased]`, `docs/agent-api-reference.md`, the 38 comment/docstring sites that describe `_module` scoping, `docs/design/specs/2026-09-05-leg-1.5-bounded-queries-and-dataflow.md` (one note under the multi-application decision pointing at F1) - Test: `tests/analysis/python/test_bounded_enumeration.py` (hub numbers on both graphs) -- [ ] **Step 1: The three full runs, verbatim** +- [x] **Step 1: The three full runs, verbatim** Offline, 7689, 7688 — exactly as Task 1 Step 6. Expected: offline ≥ 373/141 and coverage ≥ 54%; 7689 and 7688 both 0 failed. Record the three summary lines. -- [ ] **Step 2: Hub numbers on both graphs** +> **As done (sequentially, never in parallel — parallel runs in one worktree share the coverage data file):** +> - offline: `400 passed, 144 skipped, 8 warnings in 25.06s` — `Total coverage: 55.91%` +> - 7689 (1.4.1): `544 passed, 8 warnings in 410.97s (0:06:50)` +> - 7688 (1.4.0): `544 passed, 8 warnings in 420.67s (0:07:00)` +> +> The two timed N+1 tests now run with coverage paused (`no_cover`): 9.6 s / 9.0 s under `--cov` on 7689, where the tracer had put them at 15.2–15.4 s against a 15 s ceiling. Mutation-checked: a stalled `_collect` fails the clock (75.7 s), a per-class N+1 fails the round-trip ceiling (1,667). + +- [x] **Step 2: Hub numbers on both graphs** `get_call_graph(roots=[hub], depth=1)` and `depth=2` on 7688 must be 638 / 27,568 and 1,818 / 63,871. On 7689, record the numbers; if they differ, name the 1.4.1 change (e.g. #181 `PY_EXTENDS`, #180 ids) that explains it, or report it as a finding. -- [ ] **Step 3: `grep -c _module`** +> **As done.** Hub = `odoo.orm.fields_relational.Many2many.write_real` (out-degree 637) on both. 7688: **638 / 27,568** and **1,818 / 63,871**, unchanged. 7689: **638 / 27,784** and **1,818 / 64,092** — the same node sets, 219 (+3 gone) and 256 (+35 gone) different edges. All 219 at depth 1 leave one source, `addons.web.controllers.export.ExportXlsxWriter.write`, with provenance `defuse`: on the 7689 run its `self.worksheet.write(...)` call site has no Jedi target, so the level-2 def-use linker falls back to every `write` method in the application (the same fan-out gives `AccountMove.write` two new callers). This is **not** a v1.4.0..v1.4.1 emitter change: the range touches only `entrypoints/`, `neo4j/`, `schema/`, both pins carry jedi 0.19.2, and both versions produce byte-identical call targets over a fixture on the same interpreter. It is the emit environment of the 7689 run (interpreter and installed packages seen by Jedi: the same fixture flips `threading.Lock()` from `_thread/allocate_lock` to `_thread.LockType/__init__` — the 7688 and 7689 spellings — merely by running under Python 3.13 instead of 3.12). Recorded as a finding; the tests that pinned closure sizes now key them on the graph's `analyzer_version`. + +- [x] **Step 3: `grep -c _module`** `grep -n "_module" cldk/analysis/python/neo4j/neo4j_backend.py` — every remaining line is prose explaining history, or a name coincidence (`_module_full`, `in_module`). Rewrite the 38 comment sites so none describes a scoping mechanism that no longer exists. -- [ ] **Step 4: Docs** +> **As done:** 55 lines match; 49 are name coincidences (`_modules`, `_module_key`, `_module_full`, `_module_set`, `_module_prefixes`, `_module_imports`, `in_module`, `resolve_module_key`, `_load_module_keys`, `get_python_module`, `_get_module_functions`) and 6 are history — the module docstring, `_ANALYZER_FLOOR`'s note, `_bounded_call_rows`' "the `root._module IS NULL` arm is gone by construction", and `_OVERVIEW_PROJECTION`'s note on where `path` used to come from. None describes a live mechanism. + +- [x] **Step 4: Docs** CHANGELOG `[Unreleased]`: **Changed** — pin 1.4.1; scoping by id prefix (no caller-visible change); probe refuses `< 1.4.0`, warns on 1.4.0. **Fixed** — cite upstream #180/#181/#182 and python-sdk #176/#177/#178 as resolved through the pin. `docs/agent-api-reference.md`: the analyzer version floor and what attaching to an older graph does. Leg-1.5 spec: one sentence under the multi-application scope decision pointing at F1. -- [ ] **Step 5: Commit** +- [x] **Step 5: Commit** ``` git commit -m "docs(python): record id-prefix scoping, the 1.4.1 pin, and the graph version floor" @@ -293,19 +304,21 @@ Found by Task 0. `PyNeo4jBackend.get_entrypoint_coverage` (`neo4j_backend.py:~16 - Consumes: `:PyApplication.entrypoint_report_json` (a JSON string) and `entrypoint_frameworks`; the local backend's existing `PyEntrypointReport` model. - Produces: the same `get_entrypoint_coverage` return type, populated from the graph; the diagnostic survives **only** for a graph that genuinely lacks the property (a 1.4.0 graph on 7688). -- [ ] **Step 1: Flip the three tests' expectations, watch them fail on 7689** +- [x] **Step 1: Flip the three tests' expectations, watch them fail on 7689** Keep each test's *claim* (coverage is answerable / not answerable); change the expectation to "answered from the graph" for a 1.4.1 graph and keep "diagnostic" for 1.4.0. Parametrise on the graph's `analyzer_version` where the test is live. -- [ ] **Step 2: Read the property** +- [x] **Step 2: Read the property** `MATCH (a:PyApplication {name: $app}) RETURN a.entrypoint_report_json AS j, a.entrypoint_frameworks AS f` — parse `j` through the same pydantic model the local backend returns (1.4.1's #188 dumps it through the compat shim, so the shape matches). If `j` is null, emit the existing diagnostic unchanged. -- [ ] **Step 3: Parity** +- [x] **Step 3: Parity** On 7689, `get_entrypoint_coverage()` over Neo4j must equal the local backend's answer over the same checkout at level 4 for the fields the projection carries; state any documented lossiness. On 7688 the diagnostic is unchanged. -- [ ] **Step 4: Commit** +> **As done:** parity was checked against the report the graph itself carries (`entrypoint_report_json`, the analyzer's own `PyEntrypointReport` dumped sorted-key), not a fresh local level-4 run. A local `PyCodeanalyzer` over the checkout re-runs Jedi and the L3/L4 passes whatever `cache_dir` points at (~50 minutes on this checkout), so "the same checkout at level 4" is not a cheap oracle; the projected JSON is the same object the local backend would return, field for field. + +- [x] **Step 4: Commit** ``` git commit -m "fix(python): read the entrypoint report the 1.4.1 graph carries diff --git a/docs/design/specs/2026-09-05-leg-1.5-bounded-queries-and-dataflow.md b/docs/design/specs/2026-09-05-leg-1.5-bounded-queries-and-dataflow.md index 10d68f19..4ac9585f 100644 --- a/docs/design/specs/2026-09-05-leg-1.5-bounded-queries-and-dataflow.md +++ b/docs/design/specs/2026-09-05-leg-1.5-bounded-queries-and-dataflow.md @@ -522,4 +522,6 @@ prefix-parameterised like leg 1's, so it ports, but proving that is not this leg **The three upstream analyzer issues** surfaced by leg 1's live run — zero entrypoints detected in Odoo, half-dedented nested-callable source, and external→app `PY_CALLS` edges invisible to -module scoping — belong to `codeanalyzer-python`. +module scoping — belong to `codeanalyzer-python`. Module scoping itself did not outlive this leg: +leg 1.6 (F1 of `2026-09-06-leg-1.6-id-prefix-scoping.md`) replaced every `_module IN $mods` +predicate with the application's `can://` id prefix, under which the external ghosts fall. From b46cc91ad5cd51bea85fc922e120ca7837fbabcb Mon Sep 17 00:00:00 2001 From: Rahul Krishna Date: Sun, 6 Sep 2026 04:52:53 -0400 Subject: [PATCH 10/13] fix(python): seek locate and resolve through :PySymbol on every served graph The version-gated :PyCanNode label is gone. Both 1.4.0 and 1.4.1 graphs carry the unique pysymbol_id range index on (PySymbol.id), so _LOCATE_QUERY and _RESOLVE_CALLABLE_QUERY anchor on (c:PyCallable:PySymbol) and the planner seeks on either generation: locate x40 381 -> 46 ms (1.4.1), 427 -> 53 ms (1.4.0); resolve 19.2 -> 15.3 ms (1.4.1), a wash on 1.4.0. :PyCanNode was measured and rejected -- its index spans all 955,961 application nodes, so seeking it made resolve 10x slower, and 1.4.0 has no such label. With that, the 1.4.0 attach warning's premise ("scans rather than seeks") is false and the warning is removed; _can_node, _ANALYZER_INDEXED and the per-generation statement interpolation go with it. Spec F2 is amended with the measurements; CHANGELOG and the API reference no longer claim 1.4.0 scans. The vocabulary test that listed PySymbol as "not emitted by 1.4.0" was wrong (22,920 such nodes on the 1.4.0 graph) and now pins the label as a seek anchor only, never a signature-keyed match. Also: the ghost-chain fixture gets an independent shortestPath ground truth; CHANGELOG drops the #183 ghost-prefix attribution (1.4.0 already did that) and targets 2.0.0-rc.2; the plan's Python 3.13 claim is marked as inferred. --- CHANGELOG.md | 19 ++--- cldk/analysis/python/neo4j/neo4j_backend.py | 69 +++++++------------ docs/agent-api-reference.md | 4 +- .../2026-09-06-leg-1.6-id-prefix-scoping.md | 4 +- .../2026-09-06-leg-1.6-id-prefix-scoping.md | 25 +++++-- tests/analysis/python/test_e2e_neo4j_live.py | 31 +++++---- tests/analysis/python/test_locate.py | 19 ++--- .../analysis/python/test_neo4j_vocabulary.py | 15 +++- tests/analysis/python/test_schema_probe.py | 23 +++---- 9 files changed, 111 insertions(+), 98 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index e74209d6..68baf642 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 Python legs 1, 1.5 and 1.6 of the CLDK 2.0 agent-facing query facade (see `docs/design/specs/2026-09-03-agent-facing-query-facade.md`, `docs/design/specs/2026-09-05-leg-1.5-bounded-queries-and-dataflow.md` and -`docs/design/specs/2026-09-06-leg-1.6-id-prefix-scoping.md`). Targets `2.0.0-rc.1`. +`docs/design/specs/2026-09-06-leg-1.6-id-prefix-scoping.md`). Targets `2.0.0-rc.2`. ### Changed - **Pinned `codeanalyzer-python` 1.4.0 → 1.4.1** (`pyproject.toml` `dependencies` and @@ -25,12 +25,13 @@ Python legs 1, 1.5 and 1.6 of the CLDK 2.0 agent-facing query facade (see - **The schema probe reads the graph's analyzer generation.** Attaching to a graph whose `:PyApplication.analyzer_version` is below **1.4.0** (or missing) raises `GraphSchemaMismatch` naming the version found and the floor -- before this it was served with silent empties. A - **1.4.0** graph is served with identical results and one `WARNING` that scoped statements scan - rather than seek (it carries no `:PyCanNode` index). A **1.4.1** graph attaches silently, and - `locate` / `locate_many` name `:PyCanNode` on it so the per-module prefix seeks the range index - (40 positions: 399 → 89 ms). Statements whose prefix is the whole application stay unlabelled: - measured, the seek over 955,961 application nodes is 2–20× slower than the `:PyCallable` label - scan they use. + **1.4.0** graph and a **1.4.1** graph are served identically and silently: both carry the unique + `:PySymbol(id)` range index, so `locate` / `locate_many` and `resolve_callable` anchor on + `(c:PyCallable:PySymbol)` and the prefix *seeks* on either generation (40 positions: 381 → 46 ms + on 1.4.1, 427 → 53 ms on 1.4.0; `resolve_callable` 19.2 → 15.3 ms on 1.4.1, a wash on 1.4.0). + 1.4.1's `:PyCanNode(id)` index was measured and rejected -- it spans all 955,961 application + nodes, so seeking it made `resolve_callable` 10× slower -- and no statement names it. Statements + whose prefix is the whole application stay on the `:PyCallable` label scan, which is faster there. - **BREAKING: Neo4j graph vocabulary migration.** The Python Neo4j backend (`cldk.analysis.python.neo4j.PyNeo4jBackend`) now queries `PyBodyNode` / `PY_HAS_BODY_NODE` instead of the pre-1.4.0 `PyCallSite` / `PY_HAS_CALLSITE` / `PySymbol` vocabulary, matching what @@ -124,7 +125,7 @@ Python legs 1, 1.5 and 1.6 of the CLDK 2.0 agent-facing query facade (see | SDK version | Requires `codeanalyzer-python` | Graph vocabulary | | --- | --- | --- | | <= 1.5.0 | 0.3.x | `PyCallSite` / `PY_HAS_CALLSITE` / `PySymbol` | -| 2.0.0-rc.1 (this) | 1.4.1 pinned; graphs emitted by >= 1.4.0 served (1.4.0 scanned, 1.4.1 indexed) | `PyBodyNode` / `PY_HAS_BODY_NODE`, scope by `can://` id prefix | +| 2.0.0-rc.2 (this) | 1.4.1 pinned; graphs emitted by >= 1.4.0 served identically | `PyBodyNode` / `PY_HAS_BODY_NODE`, scope by `can://` id prefix | ### Added - **Slices and reachability: `slice_backward(src, within=)` / `slice_forward(src, within=)` / @@ -344,7 +345,7 @@ Python legs 1, 1.5 and 1.6 of the CLDK 2.0 agent-facing query facade (see (entrypoint report projected, Odoo `@http.route` / `http.Controller` detected: 534 callables and 94 classes on the same checkout that 1.4.0 flagged 0 / 0), #181 (`PY_EXTENDS` is actually emitted: 1,573 edges where every 2.0.0 graph had none) and #183 (`_module` retired, `:PyCanNode` range - index on `id`, `:PyExternal` ghosts under the application prefix). + index on `id`). - **The N+1 timing assertion measured the coverage tracer, not the query.** The two timed live tests (`get_symbol_table` / `get_classes` under 15 s) ran under `sys.settrace`, which adds ~5 s to a ~10 s Python-side reconstruction; they passed the ceiling by luck. They now run with coverage diff --git a/cldk/analysis/python/neo4j/neo4j_backend.py b/cldk/analysis/python/neo4j/neo4j_backend.py index 39b2253b..1d0861ad 100644 --- a/cldk/analysis/python/neo4j/neo4j_backend.py +++ b/cldk/analysis/python/neo4j/neo4j_backend.py @@ -34,11 +34,13 @@ Identity model (must match the in-memory backend; see ``codeanalyzer/neo4j/project.py``): * a class/callable/external is keyed by ``id``, under its specific label — ``:PyClass`` / - ``:PyCallable`` / ``:PyExternal`` — which is all this backend ever matches on: the producer also - stamps a shared secondary label across all three (a declared symbol is id-keyed there too, not - signature-keyed, unlike 0.3.x — the one exception is an unresolved ``PY_EXTENDS`` base-class - ghost, still merged by ``signature``, irrelevant to every query below), but the specific labels - already uniquely identify these nodes so the shared one goes unqueried; + ``:PyCallable`` / ``:PyExternal`` — which is what this backend matches on. The producer also + stamps a shared secondary label, ``:PySymbol``, across all three (a declared symbol is id-keyed + there too, not signature-keyed, unlike 0.3.x — the one exception is an unresolved ``PY_EXTENDS`` + base-class ghost, still merged by ``signature``, irrelevant to every query below), backed by a + unique range index on ``id`` that every served generation carries. Two point lookups name it + (``_RESOLVE_CALLABLE_QUERY``, ``_LOCATE_QUERY``) so their prefix predicate seeks that index + instead of scanning ``:PyCallable``; see the note above ``_LOCATE_QUERY``; * a module is a ``:PyModule`` keyed by ``file_key`` (which equals the original ``PyModule.file_path`` and the symbol-table key); * call-graph edges are ``(:PyCallable|:PyExternal)-[:PY_CALLS {weight, prov}]->(...)`` with a @@ -345,10 +347,9 @@ def _init_with_driver(self, driver: Any, *, application_name: str | None = None, #: The oldest codeanalyzer-python whose graph this backend serves. 1.4.0 introduced the #: ``can://`` id grammar every statement here scopes on; 1.4.1 dropped the ``_module`` property - #: and added the ``:PyCanNode`` range index on ``id`` that lets a prefix predicate seek. A - #: 1.4.0 graph is therefore served correctly but scanned (see :meth:`_probe_schema`). + #: this backend once read. Both generations carry the ``:PySymbol(id)`` index the point lookups + #: seek, so they are served identically -- results and cost alike (see :meth:`_probe_schema`). _ANALYZER_FLOOR = (1, 4, 0) - _ANALYZER_INDEXED = (1, 4, 1) def _probe_schema(self) -> None: """Verify the connected graph's vocabulary once, at connection time, and record the @@ -389,14 +390,6 @@ def _probe_schema(self) -> None: message=f"The graph for application {self.application_name!r} {what}; this backend needs a graph emitted by codeanalyzer-python {floor} or newer.", ) self._analyzer_version = version - if version < self._ANALYZER_INDEXED: - logger.warning( - "The graph for application %r was emitted by codeanalyzer-python %s, which carries no :PyCanNode index on id: " - "scoped queries scan rather than seek. Results are identical; re-emit with %s or newer for the index.", - self.application_name, - raw, - ".".join(map(str, self._ANALYZER_INDEXED)), - ) def _read_server_version(self) -> Tuple[int, ...] | None: """The attached server's version as an int tuple, or ``None`` when it cannot be read. @@ -487,28 +480,9 @@ def _scope_prefix(self) -> str: #: The codeanalyzer-python generation that emitted this application, set by #: :meth:`_probe_schema` (which refuses anything below :attr:`_ANALYZER_FLOOR`). The class-level - #: ``None`` is for the ``object.__new__`` seam the unit tests build backends through: an unknown - #: generation names no optional label. + #: ``None`` is for the ``object.__new__`` seam the unit tests build backends through. _analyzer_version: Tuple[int, ...] | None = None - @property - def _can_node(self) -> str: - """``":PyCanNode"`` when the attached graph carries that label's range index on ``id`` - (codeanalyzer-python 1.4.1+), else ``""``. Interpolated into ONE statement, ``_LOCATE_QUERY``. - - Naming the label makes the planner seek ``:PyCanNode(id)`` for a prefix instead of scanning - ``:PyCallable``, and that pays only when the prefix is narrow. Measured on the 1.4.1 odoo - graph: ``locate_many`` over 40 positions (per-module prefixes) 399 -> 89 ms; every statement - whose prefix is the whole application -- ``resolve_callable``, ``get_source``, - ``resolve_value``, ``get_class``, the per-parent child fetches, the callers/callees - neighbourhoods -- 2-20x *slower* (``_RESOLVE_CALLABLE_QUERY`` 16.7 -> 198 ms), because the - seek then walks all 955,961 nodes under ``can://python//`` before the label and - signature filters apply, where the label scan touched 15,549. Those stay unlabelled by - measurement, not oversight. On a 1.4.0 graph the label does not exist and naming it would - match nothing, hence the gate. - """ - return ":PyCanNode" if self._analyzer_version is not None and self._analyzer_version >= self._ANALYZER_INDEXED else "" - @cached_property def _module_set(self) -> FrozenSet[str]: """:attr:`_modules` as a set -- the membership side of :func:`~cldk.analysis.python.neo4j.reconstruct.module_key_of`. @@ -1253,8 +1227,10 @@ def get_source(self, node_id: str) -> str: return rows[0]["code"] # -----[ addressing ]----- + #: ``:PySymbol`` is named so the prefix seeks its unique ``id`` range index (see the note above + #: ``_LOCATE_QUERY``): 19.2 -> 15.3 ms on the 1.4.1 graph, 17.4 -> 17.9 ms on 1.4.0 (a wash). _RESOLVE_CALLABLE_QUERY = ( - "MATCH (c:PyCallable) WHERE c.id STARTS WITH $prefix AND (c.signature = $name OR c.signature ENDS WITH $dotted) " + "MATCH (c:PyCallable:PySymbol) WHERE c.id STARTS WITH $prefix AND (c.signature = $name OR c.signature ENDS WITH $dotted) " "OPTIONAL MATCH (owner:PyClass)-[:PY_HAS_METHOD]->(c) " "RETURN c.signature AS signature, c.name AS name, c.id AS id, " "c.start_line AS start_line, owner.signature AS class_signature" @@ -1924,14 +1900,21 @@ def get_config_readers(self, key: str) -> List[PyCallableOverview]: # Every other query in this file is scoped to the application by id prefix, and this one # narrows further, to the position's own module: ``pos.module_prefix`` is # ``module_id(app, key) + "/"``, so a same-valued ``file_key`` from a different application - # cannot win, and neither can a module whose key merely extends this one's spelling. On a - # 1.4.1+ graph ``locate_many`` names ``:PyCanNode`` on the callable so that per-module prefix - # seeks the ``:PyCanNode(id)`` range index (see :attr:`_can_node`). + # cannot win, and neither can a module whose key merely extends this one's spelling. + # + # ``:PySymbol`` is named on the callable so that per-module prefix *seeks*: the producer stamps + # the label on every class, callable and external and backs it with a unique range index on + # ``id`` (``pysymbol_id``) on every served generation -- 1.4.0 and 1.4.1 alike -- so the planner + # emits ``NodeUniqueIndexSeekByRange`` where the bare ``:PyCallable`` anchor scanned the label. + # Measured, 40 positions, median of 5: 381 -> 46 ms on the 1.4.1 graph, 427 -> 53 ms on 1.4.0. + # 1.4.1's own ``:PyCanNode(id)`` range index was measured and rejected: it spans all 955,961 + # application nodes, so seeking it walked a 40x larger range (locate 54 ms, but + # ``_RESOLVE_CALLABLE_QUERY`` 19 -> 210 ms), and 1.4.0 graphs have no such label at all. _LOCATE_QUERY = ( "UNWIND $positions AS pos " "OPTIONAL MATCH (:PyApplication {name: $app})-[:PY_HAS_MODULE]->(m:PyModule {file_key: pos.path}) " "WITH pos, m " - "OPTIONAL MATCH (c:PyCallable) " + "OPTIONAL MATCH (c:PyCallable:PySymbol) " "WHERE c.id STARTS WITH pos.module_prefix " "AND c.start_line IS NOT NULL AND c.end_line IS NOT NULL " "AND c.start_line <= pos.line AND pos.line <= c.end_line " @@ -2081,9 +2064,7 @@ def locate_many(self, positions: Sequence[Tuple[str, int]]) -> List[LocateResult # ``module_prefix`` is the exact inverse of ``_module_key``: ``module_id(app, key) + "/"`` # selects the module's own callables and nothing under a longer key sharing the spelling. rows = self._run( - # One label, swapped in per graph generation (see _can_node); the class-level statement - # keeps the spelling every served graph accepts, and test_locate pins both. - self._LOCATE_QUERY.replace("OPTIONAL MATCH (c:PyCallable) ", f"OPTIONAL MATCH (c:PyCallable{self._can_node}) ", 1), + self._LOCATE_QUERY, app=self.application_name, positions=[ {"idx": i, "path": key, "module_prefix": module_id(self.application_name, key) + "/", "line": line} diff --git a/docs/agent-api-reference.md b/docs/agent-api-reference.md index 9b1bcfd8..b0822467 100644 --- a/docs/agent-api-reference.md +++ b/docs/agent-api-reference.md @@ -36,8 +36,8 @@ and the message names what it found and the floor. What attaching to each genera | graph emitted by | attach | behaviour | | --- | --- | --- | | < 1.4.0, or no `analyzer_version` | refused (`GraphSchemaMismatch`) | the `can://` id grammar every query scopes on does not exist there | -| 1.4.0 | served, one `WARNING` logged | identical results; the graph has no `:PyCanNode` index, so scoped statements scan rather than seek, and `get_entrypoint_coverage()` reports `entrypoint_report_unavailable` (the report was not projected yet) | -| 1.4.1 and newer | served, silent | `locate` seeks the `:PyCanNode(id)` index; the entrypoint report is read off the graph | +| 1.4.0 | served, silent | identical results and plans (`locate` / `resolve_callable` seek the `:PySymbol(id)` index both generations carry); `get_entrypoint_coverage()` reports `entrypoint_report_unavailable` (the report was not projected yet) | +| 1.4.1 and newer | served, silent | as 1.4.0, and the entrypoint report is read off the graph | --- diff --git a/docs/design/plans/2026-09-06-leg-1.6-id-prefix-scoping.md b/docs/design/plans/2026-09-06-leg-1.6-id-prefix-scoping.md index 37facfd8..47781501 100644 --- a/docs/design/plans/2026-09-06-leg-1.6-id-prefix-scoping.md +++ b/docs/design/plans/2026-09-06-leg-1.6-id-prefix-scoping.md @@ -151,6 +151,8 @@ if version < (1, 4, 1): ``` Reuse `GraphSchemaMismatch`; extend it if it cannot carry a version message. Unit-test all three branches with the fake driver; live-test that 7688 warns and 7689 is silent. +> **As done, then amended in the review round:** the 1.4.0 warning was removed. Its premise -- that a 1.4.0 graph scans where 1.4.1 seeks -- was false: both generations carry the unique `:PySymbol(id)` index, and anchoring `_LOCATE_QUERY` / `_RESOLVE_CALLABLE_QUERY` on `(c:PyCallable:PySymbol)` seeks on both (spec F2 as amended, with the measurements). `_can_node` / `_ANALYZER_INDEXED` and the version-gated label swap are gone; 7688 and 7689 both attach silently. + - [x] **Step 6: Run everything against both graphs** ``` @@ -272,7 +274,7 @@ Offline, 7689, 7688 — exactly as Task 1 Step 6. Expected: offline ≥ 373/141 `get_call_graph(roots=[hub], depth=1)` and `depth=2` on 7688 must be 638 / 27,568 and 1,818 / 63,871. On 7689, record the numbers; if they differ, name the 1.4.1 change (e.g. #181 `PY_EXTENDS`, #180 ids) that explains it, or report it as a finding. -> **As done.** Hub = `odoo.orm.fields_relational.Many2many.write_real` (out-degree 637) on both. 7688: **638 / 27,568** and **1,818 / 63,871**, unchanged. 7689: **638 / 27,784** and **1,818 / 64,092** — the same node sets, 219 (+3 gone) and 256 (+35 gone) different edges. All 219 at depth 1 leave one source, `addons.web.controllers.export.ExportXlsxWriter.write`, with provenance `defuse`: on the 7689 run its `self.worksheet.write(...)` call site has no Jedi target, so the level-2 def-use linker falls back to every `write` method in the application (the same fan-out gives `AccountMove.write` two new callers). This is **not** a v1.4.0..v1.4.1 emitter change: the range touches only `entrypoints/`, `neo4j/`, `schema/`, both pins carry jedi 0.19.2, and both versions produce byte-identical call targets over a fixture on the same interpreter. It is the emit environment of the 7689 run (interpreter and installed packages seen by Jedi: the same fixture flips `threading.Lock()` from `_thread/allocate_lock` to `_thread.LockType/__init__` — the 7688 and 7689 spellings — merely by running under Python 3.13 instead of 3.12). Recorded as a finding; the tests that pinned closure sizes now key them on the graph's `analyzer_version`. +> **As done.** Hub = `odoo.orm.fields_relational.Many2many.write_real` (out-degree 637) on both. 7688: **638 / 27,568** and **1,818 / 63,871**, unchanged. 7689: **638 / 27,784** and **1,818 / 64,092** — the same node sets, 219 (+3 gone) and 256 (+35 gone) different edges. All 219 at depth 1 leave one source, `addons.web.controllers.export.ExportXlsxWriter.write`, with provenance `defuse`: on the 7689 run its `self.worksheet.write(...)` call site has no Jedi target, so the level-2 def-use linker falls back to every `write` method in the application (the same fan-out gives `AccountMove.write` two new callers). This is **not** a v1.4.0..v1.4.1 emitter change: the range touches only `entrypoints/`, `neo4j/`, `schema/`, both pins carry jedi 0.19.2, and both versions produce byte-identical call targets over a fixture on the same interpreter. It is the emit environment of the 7689 run (interpreter and installed packages seen by Jedi: the same fixture flips `threading.Lock()` from `_thread/allocate_lock` to `_thread.LockType/__init__` — the 7688 and 7689 spellings — merely by running under a different interpreter -- ≥ 3.13 inferred from the `_thread.LockType` spelling -- instead of 3.12). Recorded as a finding; the tests that pinned closure sizes now key them on the graph's `analyzer_version`. - [x] **Step 3: `grep -c _module`** diff --git a/docs/design/specs/2026-09-06-leg-1.6-id-prefix-scoping.md b/docs/design/specs/2026-09-06-leg-1.6-id-prefix-scoping.md index eec74354..8977b372 100644 --- a/docs/design/specs/2026-09-06-leg-1.6-id-prefix-scoping.md +++ b/docs/design/specs/2026-09-06-leg-1.6-id-prefix-scoping.md @@ -1,7 +1,7 @@ # Leg 1.6 — scope by `can://` id prefix, retire `_module` **Status:** decided 2026-09-06. Tracking: python-sdk#327. Follows leg 1.5 -(`2026-09-05-leg-1.5-bounded-queries-and-dataflow.md`). Blocks the `2.0.0rc2` cut. +(`2026-09-05-leg-1.5-bounded-queries-and-dataflow.md`). Blocks the `2.0.0-rc.2` cut. ## 1. Why this leg exists @@ -24,7 +24,7 @@ with no signal to the caller. | # | Decision | Why | |---|---|---| | **F1** | **Scope by id prefix, everywhere.** Every application-scope predicate becomes `n.id STARTS WITH $prefix` with `$prefix = application_id(app) + "/"`. The trailing slash is load-bearing: `odoo-slim-19` must not match `odoo-slim-19-b`. | This is the analyzer's own rule, already used once in the SDK (`get_external_symbols`) and in the test purge. One rule, both sides. | -| **F2** | **Dual support, asymmetrically.** The prefix predicate is correct on 1.4.0 graphs too — their ids have the same grammar — so the backend serves both. The probe reads `:PyApplication.analyzer_version` and `CALL db.propertyKeys()`: `< 1.4.0` or unparsable → refuse with `GraphSchemaMismatch`; 1.4.0 (`_module` present, no `:PyCanNode` index) → serve, and log once that scoping will scan rather than seek. | Correctness is free; only performance differs. Refusing 1.4.0 graphs would strand every graph emitted before today for no gain. | +| **F2** | **Dual support, symmetrically** *(amended in the review round; the decided text said "asymmetrically" and had 1.4.0 warn once that scoping scans)*. The prefix predicate is correct on 1.4.0 graphs too — their ids have the same grammar — so the backend serves both. The probe reads `:PyApplication.analyzer_version`: `< 1.4.0` or unparsable → refuse with `GraphSchemaMismatch`; anything from 1.4.0 up → serve, silently. **`:PySymbol` is the seek label on both generations**: the producer stamps it on every class, callable and external and backs it with the unique range index `pysymbol_id (PySymbol.id)`, present on 1.4.0 and 1.4.1 graphs alike (verified with `SHOW INDEXES` on both; 15,549 / 15,549 `:PyCallable` carry it on both). The two point lookups whose prefix is narrow — `_LOCATE_QUERY` and `_RESOLVE_CALLABLE_QUERY` — name `(c:PyCallable:PySymbol)` and get `NodeUniqueIndexSeekByRange` where the bare anchor was `NodeByLabelScan`. 1.4.1's own `:PyCanNode(id)` index was measured and **rejected** (table below): it spans all 955,961 application nodes, so the seek walks a 40× larger range than `:PySymbol`'s 23,138, and the label does not exist on 1.4.0. There is no warning: the premise ("1.4.0 scans rather than seeks") is false once `:PySymbol` seeks on both. | Correctness is free, and so is the seek. Refusing 1.4.0 graphs would strand every graph emitted before today for no gain; a version-gated label was an interpolation, a test and a warning for a slower plan. | | **F3** | **The public API does not move.** No accessor changes name, signature or return type. Path values stay repo-relative module keys; `LocateResult.node_id`, `SliceNode.ref` and every other id-shaped value are unchanged. | The rung's Iron Rule. This leg changes how queries are written, not what callers see. | | **F4** | **Module keys are derived, not stored.** A callable's repo-relative path is recovered from its id: strip `$prefix`, take everything up to and including the first `.py/` boundary, and **verify the result is a member of the application's module-key set** (still loaded from `:PyModule.file_key` at attach). A key that fails membership is a defect, raised, never guessed. `locate`'s match key becomes `c.id STARTS WITH module_id(app, pos.path) + "/"`, using `codeanalyzer.schema.ids.module_id`, the exact inverse. | There is no id→key helper in the SDK today; this defines one with a verification step so a pathological directory named `x.py/` cannot silently mis-key. | | **F5** | **The ghost rule survives the mechanism change.** Leg 1.5 established that an `@external` ghost is *reached, never traversed through*; that was enforced by `a._module IN $mods`, which ghosts fail because they carry no `_module`. A prefix predicate **includes** ghosts (`can://python//@external/…`). Every hop-scope therefore pins the traversal *source* to `:PyCallable` by label. | A mechanical replacement would silently re-open the leak Fix 3 of leg 1.5 closed (`reaches` true through a ghost with no all-callable route). | @@ -32,14 +32,29 @@ with no signal to the caller. | **F7** | **The audit widens before the code moves.** `test_neo4j_multi_application_scope.py`'s rule gains an affirmative arm — a statement that matches by signature must carry `IN $mods` **or** `.id STARTS WITH $` — and its statement enumeration stops filtering on `startswith("MATCH")`, which today silently skips `_OVERVIEW_PROJECTION` (`OPTIONAL MATCH`) and `_LOCATE_QUERY` (`UNWIND`). The fake two-application server filters on id prefix, not on a `_module` property that will no longer exist. | An audit that cannot see a statement cannot protect it; a fake that filters on a retired property makes every scoping test vacuous the moment the property goes. | | **F8** | **Pin 1.4.1.** `codeanalyzer-python==1.4.1` in `dependencies` and `[tool.backend-versions]`. 1.4.1 also lands #180 (body nodes and parameters carry their id in `analysis.json`), #181 (`PY_EXTENDS` emitted) and #182 (entrypoint report projected, Odoo detected) — the fixes for python-sdk's upstream reports #176/#177/#178. | Lockstep, and the 2.0 surface wants those three. | +### F2 measurements (medians of 5, odoo-slim-19, ms) + +| statement | shipped (`:PyCallable`) | `:PyCanNode` | `:PySymbol` | +|---|---|---|---| +| 7689 (1.4.1) `_RESOLVE_CALLABLE_QUERY` | 19.2 | 210.2 | **15.3** | +| 7689 (1.4.1) `_LOCATE_QUERY` ×40 positions | 381 | 54 | **46** | +| 7688 (1.4.0) `_LOCATE_QUERY` ×40 positions | 427 | n/a (no label) | **53** | +| 7688 (1.4.0) `_RESOLVE_CALLABLE_QUERY` | 17.4 | n/a | 17.9 | + +Every other statement's prefix is the whole application; there the `:PyCallable` label scan +(15,549 nodes) already beats a seek over 23k `:PySymbol` entries followed by a label filter, so +they stay unlabelled — apply `:PySymbol` elsewhere only on a measured win on both graphs. + ## 3. What changes for a caller Nothing in the surface. Two things in behaviour, both documented: - A graph emitted by codeanalyzer-python **older than 1.4.0** is refused at attach with a message naming the version found and the floor. Before this leg it was served with silent empties. -- Attaching to a **1.4.0** graph logs one line noting that scoped queries scan rather than seek - because the graph carries no `:PyCanNode` range index. Results are identical. +- Attaching to a **1.4.0** graph is indistinguishable from attaching to a 1.4.1 one: same results, + same plans (F2 as amended), nothing logged. The only 1.4.0-specific behaviour a caller can see is + `get_entrypoint_coverage()` reporting `entrypoint_report_unavailable`, because 1.4.0 never + projected the report. Tests that pinned 1.4.0 emitter behaviour will move with 1.4.1 and must be **re-read, not re-run**: `get_entrypoints` on Odoo was 0 (#177) and will not be; the entrypoint report is now projected; @@ -62,5 +77,5 @@ Port 7687 on the development machine is an ssh tunnel and is never a target. - The full live suite passes on 7689 (leg-1.5 baseline on 1.4.0: 516) **and** on 7688. The bounded-call-graph hub numbers (638 / 27,568 at depth 1; 1,818 / 63,871 at depth 2) are unchanged on 7688 and any change on 7689 is attributed to a named 1.4.1 emitter change. - The multi-application audit enumerates every Cypher statement on the backend, including those that begin with `OPTIONAL MATCH` or `UNWIND`, under the F7 rule; its fake two-application server filters on id prefix. - A test asserts the ghost rule directly on the new predicates: a `callable → ghost → callable` chain that exists on the live graph does **not** make `reaches` true. -- Attaching to a graph whose `analyzer_version` is below 1.4.0 raises `GraphSchemaMismatch`; attaching to 7688 succeeds and emits the single warning; attaching to 7689 is silent. +- Attaching to a graph whose `analyzer_version` is below 1.4.0 raises `GraphSchemaMismatch`; attaching to 7688 and to 7689 succeeds silently (F2 as amended). No statement names `:PyCanNode`; `locate_many` over 40 positions seeks on both graphs (< 60 ms where the label scan took ~400 ms). - CHANGELOG records the pin bump, the scoping change, the probe behaviour, and the three upstream fixes 1.4.1 brings; `docs/agent-api-reference.md` records the version floor. diff --git a/tests/analysis/python/test_e2e_neo4j_live.py b/tests/analysis/python/test_e2e_neo4j_live.py index b39eafb9..d74b2387 100644 --- a/tests/analysis/python/test_e2e_neo4j_live.py +++ b/tests/analysis/python/test_e2e_neo4j_live.py @@ -836,20 +836,16 @@ def test_entrypoint_report_presence_matches_the_analyzer_generation(cypher): # ===================================================================================== # The analyzer-version probe (leg 1.6, F2), on whichever generation this graph is # ===================================================================================== -def test_attach_warns_on_a_1_4_0_graph_and_is_silent_from_1_4_1(cypher, caplog): - """1.4.0 graphs are served (same id grammar) with one warning that scoped queries scan, since - they carry no ``:PyCanNode`` index; 1.4.1 and newer attach silently. Which branch this graph - exercises is read off the graph, so the same test is the back-compat gate on 7688 and the - silence check on 7689.""" +def test_attach_is_silent_on_every_served_generation(cypher, caplog): + """A 1.4.0 graph and a 1.4.1 graph are served identically -- same id grammar, same + ``:PySymbol(id)`` index behind the point lookups -- so attaching logs nothing on either. The + version is read off the graph and recorded in the failure, so the same test is the back-compat + gate on 7688 and the check on 7689; the ``:PySymbol`` anchor behind the seek is pinned in ``test_locate.py``.""" version = cypher("MATCH (a:PyApplication {name: $n}) RETURN a.analyzer_version AS v", n=APP_NAME)[0]["v"] - with caplog.at_level(logging.WARNING, logger="cldk.analysis.python.neo4j.neo4j_backend"): + with caplog.at_level(logging.INFO, logger="cldk.analysis.python.neo4j.neo4j_backend"): facade = CLDK.python(backend=Neo4jConnectionConfig(uri=NEO4J_URI, username=NEO4J_USER, password=NEO4J_PASSWORD, application_name=APP_NAME)) facade.backend.close() - warnings = [r.getMessage() for r in caplog.records if "scan rather than seek" in r.getMessage()] - if version == "1.4.0": - assert len(warnings) == 1 and "1.4.0" in warnings[0] - else: - assert not warnings, f"a {version} graph should attach silently" + assert not caplog.records, f"attaching to a {version} graph logged {[r.getMessage() for r in caplog.records]}" def test_attaching_to_an_absent_application_is_refused_not_served_empty(): @@ -1304,12 +1300,23 @@ def ghost_chain(cypher) -> Dict[str, str]: prefix=APP_PREFIX, ) for chain in chains: + args = dict(a=chain["src"], b=chain["dst"], prefix=APP_PREFIX) routed = cypher( "MATCH (a:PyCallable {signature:$a}) WHERE a.id STARTS WITH $prefix " "MATCH (a) ((x:PyCallable)-[:PY_CALLS]->(y:PyCallable) WHERE x.id STARTS WITH $prefix){1,} (m:PyCallable) " "WITH DISTINCT m WHERE m.signature = $b RETURN count(m) > 0 AS ok", - a=chain["src"], b=chain["dst"], prefix=APP_PREFIX, + **args, + )[0]["ok"] + # A second opinion from a pattern ``_REACHES`` does not use: a shortest path whose every + # node is a callable. If the two disagree the fixture is wrong, not the rule. + routed_by_shortest_path = cypher( + "MATCH (a:PyCallable {signature:$a}), (b:PyCallable {signature:$b}) " + "WHERE a.id STARTS WITH $prefix AND b.id STARTS WITH $prefix " + "MATCH p = shortestPath((a)-[:PY_CALLS*1..12]->(b)) WHERE all(n IN nodes(p) WHERE n:PyCallable) " + "RETURN count(p) > 0 AS ok", + **args, )[0]["ok"] + assert routed == routed_by_shortest_path, f"the two all-callable route checks disagree on {chain['src']} -> {chain['dst']}" if not routed: return dict(chain) pytest.skip("every callable -> ghost -> callable chain on this graph also has an all-callable route, so none isolates the ghost rule") diff --git a/tests/analysis/python/test_locate.py b/tests/analysis/python/test_locate.py index 870c896b..86a3452c 100644 --- a/tests/analysis/python/test_locate.py +++ b/tests/analysis/python/test_locate.py @@ -265,20 +265,21 @@ def test_locate_query_is_scoped_to_the_application(py, fake_driver): assert "_module" not in statement, "the graph stores no _module property to match on" -def test_locate_names_pycannode_only_on_a_graph_that_has_it(py, fake_driver): - """The per-module prefix seeks the ``:PyCanNode(id)`` range index a 1.4.1 graph carries - (measured on odoo: 40 positions 399 -> 89 ms); a 1.4.0 graph has no such label and naming it - would match nothing. Both spellings are pinned because the swap is a string edit on the - statement -- a renamed anchor would silently lose the seek, never the answer.""" +def test_locate_and_resolve_seek_the_pysymbol_index(py, fake_driver): + """The per-module prefix seeks the unique ``:PySymbol(id)`` range index every served graph + carries (1.4.0 and 1.4.1 alike; measured on odoo, 40 positions: 381 -> 46 ms and 427 -> 53 ms). + The anchor is pinned as text because dropping the label loses the seek and never the answer -- + a scan is a silent 8x, not a failure. ``resolve_callable`` names it for the same reason; both + are the same statement whatever generation the graph is.""" py.locate("src/app.py", 21) - assert "OPTIONAL MATCH (c:PyCallable:PyCanNode) " in next(s for s in fake_driver.statements if "UNWIND $positions AS pos" in s) + assert "OPTIONAL MATCH (c:PyCallable:PySymbol) " in next(s for s in fake_driver.statements if "UNWIND $positions AS pos" in s) + assert PyNeo4jBackend._RESOLVE_CALLABLE_QUERY.startswith("MATCH (c:PyCallable:PySymbol) WHERE c.id STARTS WITH $prefix") fake_driver.statements.clear() fake_driver.analyzer_version = "1.4.0" old = PyNeo4jBackend._from_driver(fake_driver, application_name="app") - assert old.locate("src/app.py", 21).callable is not None, "the 1.4.0 spelling must still answer" - statement = next(s for s in fake_driver.statements if "UNWIND $positions AS pos" in s) - assert "OPTIONAL MATCH (c:PyCallable) " in statement and "PyCanNode" not in statement + assert old.locate("src/app.py", 21).callable is not None + assert "OPTIONAL MATCH (c:PyCallable:PySymbol) " in next(s for s in fake_driver.statements if "UNWIND $positions AS pos" in s) def test_locate_scope_is_actually_honoured(py): diff --git a/tests/analysis/python/test_neo4j_vocabulary.py b/tests/analysis/python/test_neo4j_vocabulary.py index 5d485b82..b795e460 100644 --- a/tests/analysis/python/test_neo4j_vocabulary.py +++ b/tests/analysis/python/test_neo4j_vocabulary.py @@ -21,7 +21,13 @@ SRC = pathlib.Path("cldk/analysis/python/neo4j/neo4j_backend.py").read_text() -RETIRED = ["PyCallSite", "PY_HAS_CALLSITE", "PySymbol"] +#: 0.3.x's call-site vocabulary. ``PySymbol`` is deliberately *not* here: 0.3.x used it as the +#: primary, signature-keyed label, and 1.4.x kept the name for the shared secondary label it stamps +#: on every class, callable and external, backed by the unique ``pysymbol_id`` range index on ``id`` +#: (22,920 nodes on the 1.4.0 odoo graph, 23,138 on 1.4.1). The backend names it on two point +#: lookups so their prefix predicate seeks that index; what must not come back is matching a +#: symbol *by signature* through it. +RETIRED = ["PyCallSite", "PY_HAS_CALLSITE"] def test_no_retired_labels_in_cypher(): @@ -29,6 +35,13 @@ def test_no_retired_labels_in_cypher(): assert name not in SRC, f"{name} is not emitted by codeanalyzer-python 1.4.0" +def test_pysymbol_is_only_ever_a_seek_anchor_beside_pycallable(): + """Every mention of ``:PySymbol`` in a statement is the compound anchor ``(c:PyCallable:PySymbol)`` + -- never a node matched on its own, and never keyed by signature the 0.3.x way.""" + assert re.findall(r"\(\w+:PySymbol\)|PySymbol \{signature", SRC) == [] + assert SRC.count("(c:PyCallable:PySymbol)") == 2, "the two measured point lookups (_LOCATE_QUERY, _RESOLVE_CALLABLE_QUERY)" + + def test_call_sites_query_uses_body_nodes(): assert "PY_HAS_BODY_NODE" in SRC assert re.search(r"PyBodyNode\s*\{?\s*kind", SRC), "call sites are body nodes with kind='call'" diff --git a/tests/analysis/python/test_schema_probe.py b/tests/analysis/python/test_schema_probe.py index 90ae6195..25655c6d 100644 --- a/tests/analysis/python/test_schema_probe.py +++ b/tests/analysis/python/test_schema_probe.py @@ -70,19 +70,12 @@ def test_probe_refuses_when_the_version_cannot_be_read(fake_driver, raw): PyNeo4jBackend._from_driver(fake_driver, application_name="app") -def test_probe_serves_a_1_4_0_graph_and_warns_once_about_scanning(fake_driver, caplog): - """1.4.0 ids have the same grammar, so results are identical; only the index is missing.""" - fake_driver.analyzer_version = "1.4.0" - with caplog.at_level(logging.WARNING, logger="cldk.analysis.python.neo4j.neo4j_backend"): - PyNeo4jBackend._from_driver(fake_driver, application_name="app") - warnings = [r for r in caplog.records if "scan rather than seek" in r.getMessage()] - assert len(warnings) == 1 - assert "1.4.0" in warnings[0].getMessage() and "1.4.1" in warnings[0].getMessage() - - -@pytest.mark.parametrize("raw", ["1.4.1", "1.5.0", "2.0.0", "1.4.1.post1"]) -def test_probe_is_silent_from_1_4_1_up(fake_driver, caplog, raw): +@pytest.mark.parametrize("raw", ["1.4.0", "1.4.1", "1.5.0", "2.0.0", "1.4.1.post1"]) +def test_probe_serves_every_generation_from_the_floor_up_silently(fake_driver, caplog, raw): + """1.4.0 ids have the same grammar and the same ``:PySymbol(id)`` index as 1.4.1, so a 1.4.0 + graph is served identically -- results and cost -- and there is nothing to warn about.""" fake_driver.analyzer_version = raw - with caplog.at_level(logging.WARNING, logger="cldk.analysis.python.neo4j.neo4j_backend"): - PyNeo4jBackend._from_driver(fake_driver, application_name="app") - assert not [r for r in caplog.records if "scan rather than seek" in r.getMessage()] + with caplog.at_level(logging.INFO, logger="cldk.analysis.python.neo4j.neo4j_backend"): + backend = PyNeo4jBackend._from_driver(fake_driver, application_name="app") + assert backend._analyzer_version == tuple(int(x) for x in raw.split(".")[:3]) + assert not caplog.records, [r.getMessage() for r in caplog.records] From e86b6a03e3fd73a37eae1ff2e3b24099d89bf391 Mon Sep 17 00:00:00 2001 From: Rahul Krishna Date: Sun, 6 Sep 2026 04:52:53 -0400 Subject: [PATCH 11/13] test(python): pause coverage around the timed tests with a marker that survives --no-cov pytest-cov's no_cover hook (through 7.1.0) calls cov_controller.pause() unguarded; under --no-cov the controller is None and both timed N+1 tests died with AttributeError in 0.3 s instead of running. A local `timed` marker and a hookwrapper in conftest pause the tracer only when the _cov plugin is present with a live controller. Verified on the 1.4.1 graph: 9.6 / 8.7 s under --cov, 9.6 / 9.0 s under --no-cov, and a 6 s stall in _collect still fails the clock (75.9 s). --- tests/analysis/python/conftest.py | 24 +++++++++++++++++++ .../python/test_bounded_enumeration.py | 13 ++++++---- 2 files changed, 32 insertions(+), 5 deletions(-) diff --git a/tests/analysis/python/conftest.py b/tests/analysis/python/conftest.py index 5b552e0e..610449c6 100644 --- a/tests/analysis/python/conftest.py +++ b/tests/analysis/python/conftest.py @@ -77,6 +77,30 @@ ) +def pytest_configure(config): + config.addinivalue_line("markers", "timed: the test asserts on a wall clock; coverage is paused around its call") + + +@pytest.hookimpl(hookwrapper=True) +def pytest_runtest_call(item): + """Pause the coverage tracer around a ``timed`` test's call -- when there is one to pause. + + pytest-cov's own ``no_cover`` marker does the same, but its hook (through 7.1.0) dereferences + ``cov_controller`` unguarded, and under ``--no-cov`` that is ``None``: the marked test dies with + ``AttributeError`` before it runs. This checks for the plugin *and* a live controller. + """ + cov = item.config.pluginmanager.get_plugin("_cov") + controller = getattr(cov, "cov_controller", None) + if item.get_closest_marker("timed") and controller is not None: + controller.pause() + try: + yield + finally: + controller.resume() + else: + yield + + class _FakeRecord: """Stands in for ``neo4j.Record``: ``PyNeo4jBackend._run`` calls ``.data()`` on every row.""" diff --git a/tests/analysis/python/test_bounded_enumeration.py b/tests/analysis/python/test_bounded_enumeration.py index aee955b2..3b498880 100644 --- a/tests/analysis/python/test_bounded_enumeration.py +++ b/tests/analysis/python/test_bounded_enumeration.py @@ -70,11 +70,14 @@ #: the ceiling asserted: headroom over the measured ~10s so a slow machine does not flake at the #: boundary, low enough to catch a regression back towards the minutes this leg removed. #: -#: The two timed tests run with coverage paused (pytest-cov's ``no_cover`` marker) because the -#: clock is meant to measure the query, not the tracer: measured on both live graphs, the same +#: The two timed tests run with coverage paused (the ``timed`` marker, see ``conftest.py``) because +#: the clock is meant to measure the query, not the tracer: measured on both live graphs, the same #: call is 10.2-10.5 s under ``--no-cov`` and ~15.3 s with the coverage tracer on -- the #: reconstruction is a Python-side walk over ~950k rows, and ``sys.settrace`` costs it half again. -#: Under the tracer the ceiling passed by luck, and only sometimes. +#: Under the tracer the ceiling passed by luck, and only sometimes. (Not pytest-cov's own +#: ``no_cover``: through 7.1.0 its hook calls ``cov_controller.pause()`` unguarded, and under +#: ``--no-cov`` the controller is ``None`` -- the marked test then fails with ``AttributeError`` +#: instead of running.) _WALL_CLOCK_TARGET = 10 _WALL_CLOCK_CEILING = 15 @@ -110,7 +113,7 @@ def _assert_children_survived(classes: List[PyClass], callables: List[PyCallable assert any(m.callables for m in callables), "no nested callables survived the collapse" -@pytest.mark.no_cover # the clock measures the query, not the tracer -- see _WALL_CLOCK_CEILING +@pytest.mark.timed # the clock measures the query, not the tracer -- see _WALL_CLOCK_CEILING def test_symbol_table_is_not_n_plus_one(live_analysis, count_round_trips): n = count_round_trips(live_analysis) started = time.monotonic() @@ -129,7 +132,7 @@ def test_symbol_table_is_not_n_plus_one(live_analysis, count_round_trips): assert any(m.variables for m in table.values()), "no module variables survived the collapse" -@pytest.mark.no_cover # as above +@pytest.mark.timed # as above def test_classes_is_not_n_plus_one(live_analysis, count_round_trips): n = count_round_trips(live_analysis) started = time.monotonic() From eec105c1ce1b06df27fa7c60494693db9b118939 Mon Sep 17 00:00:00 2001 From: Rahul Krishna Date: Sun, 6 Sep 2026 04:52:53 -0400 Subject: [PATCH 12/13] fix(python): reload the module keys once before calling an unknown module a defect The key set module_key_of verifies against was read once at attach, from a graph this backend does not own; a re-emit that added a module made every answer carrying one of its callables raise a bare KeyError with a can:// id in it -- one new callable took down a 15,549-row overview. _module_key now reloads the keys once on a miss and retries; a second miss raises CodeanalyzerExecutionException naming the key count, never the id. F4's "verified, never guessed" is unchanged. --- cldk/analysis/python/neo4j/neo4j_backend.py | 23 ++++++++++- tests/analysis/python/test_module_key_of.py | 43 +++++++++++++++++++++ 2 files changed, 64 insertions(+), 2 deletions(-) diff --git a/cldk/analysis/python/neo4j/neo4j_backend.py b/cldk/analysis/python/neo4j/neo4j_backend.py index 1d0861ad..7c147e32 100644 --- a/cldk/analysis/python/neo4j/neo4j_backend.py +++ b/cldk/analysis/python/neo4j/neo4j_backend.py @@ -494,8 +494,27 @@ def _module_key(self, node_id: str) -> str: """The repo-relative module key a node's ``can://`` id embeds (F4). The graph stores no path property to project, so every ``path``/``file`` a caller sees is derived from the id it came with and verified against the application's module keys -- never split, never - guessed.""" - return R.module_key_of(node_id, self._scope_prefix, self._module_set) + guessed. + + The key set was read once, at attach, from a graph this backend does not own: a re-emit + since then can add a module, and its callables would then fail membership although they + are this application's. So a miss reloads the module keys **once** and retries; a second + miss is a genuine defect and is raised as such, without the id (E6). + """ + try: + return R.module_key_of(node_id, self._scope_prefix, self._module_set) + except KeyError: + pass + self._modules = self._load_module_keys() + self.__dict__.pop("_module_set", None) # drop the cached frozenset; rebuilt on next read + try: + return R.module_key_of(node_id, self._scope_prefix, self._module_set) + except KeyError: + raise CodeanalyzerExecutionException( + f"A node of application {self.application_name!r} belongs to none of the {len(self._module_set)} module keys the graph " + "holds for it, even after reloading them: the module set changed since attach, or the node is not one of its " + "declared modules'. Re-attach to the graph." + ) from None def _overview(self, row: Dict[str, Any]) -> PyCallableOverview: """A projected callable row (``_OVERVIEW_PROJECTION``'s shape) with its ``path`` derived.""" diff --git a/tests/analysis/python/test_module_key_of.py b/tests/analysis/python/test_module_key_of.py index e6d4cf47..869872e9 100644 --- a/tests/analysis/python/test_module_key_of.py +++ b/tests/analysis/python/test_module_key_of.py @@ -60,3 +60,46 @@ def test_the_module_id_itself_keys_to_its_own_key(): def test_an_id_under_another_application_raises(): with pytest.raises(KeyError): module_key_of("can://python/app-b/pkg/a.py/f()", APP, {"pkg/a.py"}) + + +# ---------------------------------------------------------------------------------------------- +# The backend's use of it: the key set is read once at attach, from a graph the SDK does not own. +# ---------------------------------------------------------------------------------------------- +from cldk.analysis.python.neo4j.neo4j_backend import PyNeo4jBackend # noqa: E402 +from cldk.utils.exceptions import CodeanalyzerExecutionException # noqa: E402 + + +def _overview_row(node_id: str) -> dict: + return {"id": node_id, "signature": node_id.rsplit("/", 1)[1], "name": "f", "decorators": None, "start_line": 1, "end_line": 2, "class_signature": None} + + +def test_a_module_added_since_attach_is_found_after_one_reload_and_a_foreign_id_is_a_typed_error(fake_driver): + """A re-emit can add a module after attach; its callables must not take down a whole-application + answer. One miss reloads the key set once and retries. A second miss is a real defect, raised + as an SDK exception that names the key count and never the ``can://`` id (E6).""" + graph = {"modules": ["a.py"], "callables": ["can://python/app/a.py/f"]} + + def responder(query, params): + if "RETURN m.file_key AS k" in query: + return [{"k": k} for k in graph["modules"]] + if "OPTIONAL MATCH (owner:PyClass)-[:PY_HAS_METHOD]->(c)" in query: + return [_overview_row(i) for i in graph["callables"]] + return [] + + fake_driver.responder = responder + backend = PyNeo4jBackend._from_driver(fake_driver, application_name="app") + loads = lambda: sum("RETURN m.file_key AS k" in s for s in fake_driver.statements) # noqa: E731 + assert loads() == 1 + + graph["modules"].append("b.py") # the graph moved under us + graph["callables"].append("can://python/app/b.py/g") + assert {o.path for o in backend.get_callables_overview()} == {"a.py", "b.py"} + assert loads() == 2, "one reload, on the first miss" + + graph["callables"].append("can://python/app/zzz.py/h") # no such module, before or after reload + with pytest.raises(CodeanalyzerExecutionException) as e: + backend.get_callables_overview() + assert loads() == 3 + assert "changed since attach" in str(e.value) and "2 module keys" in str(e.value) + assert "can://" not in str(e.value) and "zzz" not in str(e.value) + From 63b19aef47da0a747cd53536b4a9c6105f988399 Mon Sep 17 00:00:00 2001 From: Rahul Krishna Date: Sun, 6 Sep 2026 04:52:54 -0400 Subject: [PATCH 13/13] test(python): audit every inline Cypher statement, and scope every arm of _SOURCES The multi-application audit saw class-level statements and the child walk; the ~20 statements written inline at self._run( sites were outside both. The audit now reassembles each of the 44 sites from the class source (constants, f-strings, + chains, .format receivers, self. strings and local assignments) and classifies every statement as prefix-scoped, :PyApplication{name}-anchored, id-keyed or server introspection -- anything else fails, and so does a statement naming :PyCanNode. Stripping the prefix from one inline statement turns the audit red. _SOURCES' body-node and ghost arms carried no prefix, so describe() reported another application's ref as "found, with nothing to read"; both arms now scope on the prefix, and the two-application fake evaluates each arm on its own predicate. --- cldk/analysis/python/neo4j/neo4j_backend.py | 10 +- .../test_neo4j_multi_application_scope.py | 146 ++++++++++++++++-- 2 files changed, 141 insertions(+), 15 deletions(-) diff --git a/cldk/analysis/python/neo4j/neo4j_backend.py b/cldk/analysis/python/neo4j/neo4j_backend.py index 7c147e32..30097ea5 100644 --- a/cldk/analysis/python/neo4j/neo4j_backend.py +++ b/cldk/analysis/python/neo4j/neo4j_backend.py @@ -1672,14 +1672,14 @@ def flows_to_argument(self, src: str, callee: str, arg: str, *, within: str, dep #: are one, its interior the other). The ``:PyBodyNode`` and ``:PyExternal`` arms return #: ``null`` code deliberately: the graph carries no text below callable granularity, and a #: ghost was never analysed, so those rows say "found, and there is nothing to read", which is - #: what keeps that apart from "not found" (see :meth:`PythonAnalysisBackend.describe`). Only - #: the callable arm carries the prefix predicate: a body-node id and a ghost id both embed the - #: application, while a signature does not. + #: what keeps that apart from "not found" (see :meth:`PythonAnalysisBackend.describe`). All + #: three arms carry the prefix predicate: a ref naming another application's node must read as + #: "not found" here, not as "found, with nothing to read". _SOURCES = ( "MATCH (c:PyCallable) WHERE c.id STARTS WITH $prefix AND (c.id IN $refs OR c.signature IN $refs) " "RETURN c.id AS id, c.signature AS sig, c.code AS code " - "UNION MATCH (b:PyBodyNode) WHERE b.id IN $refs RETURN b.id AS id, null AS sig, null AS code " - "UNION MATCH (e:PyExternal) WHERE e.id IN $refs RETURN e.id AS id, null AS sig, null AS code" + "UNION MATCH (b:PyBodyNode) WHERE b.id STARTS WITH $prefix AND b.id IN $refs RETURN b.id AS id, null AS sig, null AS code " + "UNION MATCH (e:PyExternal) WHERE e.id STARTS WITH $prefix AND e.id IN $refs RETURN e.id AS id, null AS sig, null AS code" ) def _sources_for(self, refs: Sequence[str]) -> Dict[str, "str | None"]: diff --git a/tests/analysis/python/test_neo4j_multi_application_scope.py b/tests/analysis/python/test_neo4j_multi_application_scope.py index 47b1300c..6780ab10 100644 --- a/tests/analysis/python/test_neo4j_multi_application_scope.py +++ b/tests/analysis/python/test_neo4j_multi_application_scope.py @@ -45,8 +45,10 @@ from __future__ import annotations +import ast import inspect import re +from collections import defaultdict from typing import Any, Dict, List import pytest @@ -79,8 +81,8 @@ def _node(module: str, key: str, **props: Any) -> Dict[str, Any]: "class_methods": ( CLASS_SIG, { - APP_A_MODULE: _node(APP_A_MODULE, "Widget/render", signature=METHOD_SIG, name="alpha_method", path=APP_A_MODULE), - APP_B_MODULE: _node(APP_B_MODULE, "Widget/render", signature=METHOD_SIG, name="beta_method", path=APP_B_MODULE), + APP_A_MODULE: _node(APP_A_MODULE, "Widget/render", signature=METHOD_SIG, name="alpha_method", path=APP_A_MODULE, code="alpha code"), + APP_B_MODULE: _node(APP_B_MODULE, "Widget/render", signature=METHOD_SIG, name="beta_method", path=APP_B_MODULE, code="beta code"), }, ), "class_attributes": ( @@ -145,6 +147,14 @@ def _node(module: str, key: str, **props: Any) -> Dict[str, Any]: _node(APP_B_MODULE, "Widget", signature=CLASS_SIG, name="Widget", path=APP_B_MODULE), ] +#: What ``_SOURCES`` (``describe``) may be handed, per label and per application: the shared method +#: (by signature), each application's one call-site body node, and one ``@external`` ghost each. +_SOURCE_NODES: Dict[str, List[Dict[str, Any]]] = { + "PyCallable": list(_CHILDREN["class_methods"][1].values()), + "PyBodyNode": list(_CHILDREN["callable_callsites"][1].values()), + "PyExternal": [{"id": f"can://python/{app}/@external/os/path", "name": "path", "module": "os"} for app in (APP_A, APP_B)], +} + #: The three spellings of the application scope a statement may carry: the whole application #: (``$prefix``); for a narrowed bulk fetch, a list of per-module prefixes (``$prefixes``); and for #: ``locate``, one module's own prefix per position (``pos.module_prefix``, minted from the same @@ -198,6 +208,14 @@ def _in_scope(query: str, params: Dict[str, Any], props: Dict[str, Any]) -> bool def _fake_two_app_cypher(query: str, **params: Any) -> List[Dict[str, Any]]: """Answer the statements a class reconstruction issues, honestly (see :func:`_in_scope`).""" + if "UNION MATCH (b:PyBodyNode)" in query: # _SOURCES: three arms, each judged on its own predicate + rows = [] + for arm in query.split(" UNION "): + label = re.search(r"MATCH \(\w+:(\w+)\)", arm)[1] + for props in _SOURCE_NODES[label]: + if (props["id"] in params["refs"] or props.get("signature") in params["refs"]) and _in_scope(arm, params, props): + rows.append({"id": props["id"], "sig": props.get("signature"), "code": props.get("code")}) + return rows bucket = _bucket_of(query) if bucket == "module_functions": # Keyed by file_key: the per-parent twin names it (``$fk``), the bulk twin lists the scope @@ -307,6 +325,18 @@ def test_a_module_key_shared_by_two_applications_does_not_leak(bulk: bool): assert set(module.functions) == {"alpha_fn"}, "the module-keyed fetch leaked another application's function" +def test_describe_does_not_find_another_applications_body_nodes_or_ghosts(): + """``_SOURCES`` answers ``describe`` for callables, body nodes and ghosts in one statement. A + body-node or ghost id embeds its application, but "found, with nothing to read" and "not found" + are different answers, and a ref from another application must get the second -- so every arm + carries the prefix, not only the callable one, and the shared signature resolves to A's code.""" + backend = _two_app_backend() + a_body, b_body = (n["id"] for n in _SOURCE_NODES["PyBodyNode"]) + a_ghost, b_ghost = (n["id"] for n in _SOURCE_NODES["PyExternal"]) + found = backend._sources_for([METHOD_SIG, a_body, b_body, a_ghost, b_ghost]) + assert found == {METHOD_SIG: "alpha code", a_body: None, a_ghost: None} + + def test_every_child_statement_carries_the_application_scope(): """The four module-keyed collections cannot be caught by the fixture above (see the module docstring), so they are caught here: every statement either path issues to fetch children is @@ -345,13 +375,107 @@ def _class_level_statements() -> Dict[str, str]: } +def _inline_statements() -> Dict[str, str]: + """The Cypher written *inline* at every ``self._run(`` site, keyed ``@`` -- the + statements the class-attribute enumeration cannot see (about twenty of the 44 sites). + + Reassembled from the class's own source: string constants verbatim; f-strings and ``+`` chains + joined; ``.format(...)`` unwrapped to its receiver; ``self.`` replaced by the class-level + string it names; a local name replaced by every assignment to it in the enclosing method, in + order (``query += " AND ..."``); anything else becomes ``{…}``. A site whose first argument is + a bare parameter or a lookup resolves to ``{…}`` alone -- ``_children`` (its callers' literals + are judged by the child walk above), ``_collect`` (``_BULK_CHILD_QUERIES``, judged directly) + and ``_paths`` (``_PATHS`` / ``_CALL_PATHS``, class-level) -- and is reported as such. + """ + class_strings = {name: value for name, value in vars(PyNeo4jBackend).items() if isinstance(value, str)} + out: Dict[str, str] = {} + for fn in ast.walk(ast.parse(inspect.getsource(PyNeo4jBackend))): + if not isinstance(fn, ast.FunctionDef): + continue + assigned: Dict[str, List[ast.expr]] = defaultdict(list) + for node in ast.walk(fn): + if isinstance(node, ast.Assign): + for target in node.targets: + if isinstance(target, ast.Name): + assigned[target.id].append(node.value) + elif isinstance(node, ast.AugAssign) and isinstance(node.target, ast.Name): + assigned[node.target.id].append(node.value) + + def text(e: ast.expr, depth: int = 0) -> str: + if depth > 8: + return "{…}" + if isinstance(e, ast.Constant) and isinstance(e.value, str): + return e.value + if isinstance(e, ast.JoinedStr): + return "".join(text(v.value if isinstance(v, ast.FormattedValue) else v, depth + 1) for v in e.values) + if isinstance(e, ast.BinOp) and isinstance(e.op, ast.Add): + return text(e.left, depth + 1) + text(e.right, depth + 1) + if isinstance(e, ast.Call) and isinstance(e.func, ast.Attribute) and e.func.attr == "format": + return text(e.func.value, depth + 1) + if isinstance(e, ast.Attribute) and isinstance(e.value, ast.Name) and e.value.id == "self" and e.attr in class_strings: + return class_strings[e.attr] + if isinstance(e, ast.Name) and e.id in assigned: + return "".join(text(v, depth + 1) for v in assigned[e.id]) + return "{…}" + + for call in ast.walk(fn): + if isinstance(call, ast.Call) and isinstance(call.func, ast.Attribute) and call.func.attr == "_run" and isinstance(call.func.value, ast.Name) and call.func.value.id == "self" and call.args: + out.setdefault(f"{fn.name}@{call.lineno}", text(call.args[0])) + return out + + +#: The four ways a statement stays inside one application. ``introspection`` is the server, not +#: the data (``CALL db.relationshipTypes()``, ``CALL dbms.components()``); ``application`` walks out +#: from ``(:PyApplication {name: $app})`` and cannot leave it; ``prefix`` and ``id`` are the two +#: ``can://``-stamped forms the docstring below explains. Anything else is unscoped. +_INTROSPECTION = re.compile(r"^\s*CALL (db|dbms)\.") +_ANCHORED_ON_THE_APPLICATION = re.compile(r"\(\w*:PyApplication \{name: \$app\}\)") + + +def _scope_kind(statement: str) -> str | None: + if _INTROSPECTION.match(statement): + return "introspection" + if _is_scoped(statement): + return "prefix" + if _ANCHORED_ON_THE_APPLICATION.search(statement): + return "application" + if _MATCHES_BY_ID.search(statement): + return "id" + return None + + +def _every_statement() -> Dict[str, str]: + """Class-level statements (minus fragments) plus every inline one that resolved to Cypher.""" + inline = {name: s for name, s in _inline_statements().items() if s != "{…}"} + return {**{n: s for n, s in _class_level_statements().items() if n not in _FRAGMENTS}, **inline} + + def test_the_audit_sees_the_dataflow_statements_too(): names = set(_class_level_statements()) for expected in ("_REACHES", "_CONE", "_PATHS", "_CALL_PATHS", "_VALUE_REACHES", "_CALLEE_VALUES", "_SOURCES", "_SLICE", "_CALLERS", "_CALLEES", "_OWN_EDGES", "_LOCATE_QUERY", "_OVERVIEW_PROJECTION"): assert expected in names, f"{expected} is not a class-level statement any more; move it back or extend the audit" -@pytest.mark.parametrize("name", sorted(set(_class_level_statements()) - _FRAGMENTS)) +def test_the_audit_sees_every_inline_statement_too(): + """One harvested statement per ``self._run(`` in the source -- a site the harvester cannot + reassemble shows up as ``{…}`` and must be one of the three known indirections, so a new + ``self._run(some_variable)`` cannot slip past unjudged.""" + inline = _inline_statements() + assert len(inline) == inspect.getsource(PyNeo4jBackend).count("self._run("), "a self._run( site the harvester did not see" + indirect = sorted(name.split("@")[0] for name, s in inline.items() if s == "{…}") + assert indirect == ["_children", "_collect", "_paths"], f"unjudged statements passed through a variable: {indirect}" + assert all(re.match(r"(MATCH|OPTIONAL MATCH|UNWIND|CALL)\b", s.lstrip()) for s in inline.values() if s != "{…}"), "an inline statement does not start with a Cypher clause" + for expected in ("get_source", "get_class", "get_python_file", "get_method_bodies", "resolve_value", "_call_rows", "_bounded_call_rows", "_get_module_functions", "get_callsites_for", "get_config_uses", "get_config_readers", "get_decorated_callables", "get_entrypoints", "get_entrypoint_classes", "_probe_resolution_edges", "_own_edges", "get_dependencies"): + assert any(name.startswith(expected + "@") for name in inline), f"{expected}'s statement is not harvested" + + +def test_no_statement_names_pycannode(): + """1.4.1's ``:PyCanNode`` label was measured and rejected as a seek anchor (the spec's F2 table) + and does not exist on 1.4.0 graphs; a statement naming it would match nothing there.""" + assert [name for name, s in _every_statement().items() if "PyCanNode" in s] == [] + + +@pytest.mark.parametrize("name", sorted(_every_statement())) def test_every_statement_is_application_scoped_or_keyed_by_an_application_stamped_id(name): """Two ways a statement stays inside one application, and every statement must use one. @@ -361,14 +485,16 @@ def test_every_statement_is_application_scoped_or_keyed_by_an_application_stampe A body-node or ghost **id** embeds the application (``can://python//…``) and the emitter only ever links nodes from its own run, so a statement keyed *only* by id is scoped by construction and may omit the predicate -- ``_SLICE``, ``_PATHS`` and ``_VALUE_REACHES`` do, - for the measured cost of testing 195,784 reached nodes against a list. A statement keyed by - neither would be unscoped and fails here. + for the measured cost of testing 195,784 reached nodes against a list. A statement anchored + on ``(:PyApplication {name: $app})`` walks out from the application node and cannot leave it + (the module list, artifacts, config keys, the entrypoint report). A statement keyed by none + of these would be unscoped and fails here -- class-level and inline alike. """ - statement = _class_level_statements()[name] - by_signature, by_id, scoped = bool(_MATCHES_BY_SIGNATURE.search(statement)), bool(_MATCHES_BY_ID.search(statement)), _is_scoped(statement) - assert scoped or by_id, f"{name} carries no application scope and is not keyed by an application-stamped id" - if by_signature: - assert scoped, f"{name} matches by signature without the application scope" + statement = _every_statement()[name] + kind = _scope_kind(statement) + assert kind is not None, f"{name} carries no application scope and is not keyed by an application-stamped id: {statement[:160]!r}" + if _MATCHES_BY_SIGNATURE.search(statement): + assert kind == "prefix", f"{name} matches by signature without the application prefix" def test_the_overview_projection_is_only_ever_appended_to_a_scoped_match():