diff --git a/.claude/SCHEMA_DECISIONS.md b/.claude/SCHEMA_DECISIONS.md index 266b152..6ee8952 100644 --- a/.claude/SCHEMA_DECISIONS.md +++ b/.claude/SCHEMA_DECISIONS.md @@ -285,6 +285,51 @@ made output load-dependent (#145). `callee: null→id` remains the single sanctioned L1→L2 refinement. - Neo4j projection unchanged (`PY_CALLS` carries `prov` as data). +## 2026-08-27 — Neutral `Artifact`/`Package` subgraph (issue #157, Task 6) + +Design: `.superpowers/sdd/2026-08-27-artifacts-and-dependencies/task-6-brief.md` +(gitignored working brief — not committed; this entry is the durable record). + +Projects `PyApplication.artifacts`/`dependencies`/`unresolved_imports` (Tasks +1-5) into the graph: new labels `Artifact`, `Package` (merge key `id`); new +rels `HAS_ARTIFACT` (PyApplication→Artifact), `DECLARES_DEPENDENCY` +(Artifact→Package, props `spec`/`kind`/`extras`/`prov`), `LOCKS` +(Artifact→Package, prop `version`), `PY_PROVIDES` (Package→PyExternal), +`PY_UNRESOLVED_IMPORT` (PyApplication→PyExternal, prop `prov`). + +- **`Artifact`/`Package` deliberately break the `Py`-prefix convention** the + Level-3 CPG section above establishes (`PySymbol`, `PyBodyNode`, `PY_CALLS`, + …). Opposite rationale, same namespacing question: a manifest file or a + PyPI package is not a Python-language concept — a TypeScript analyzer + reading `package.json` in the same repo should MERGE onto the same + `Artifact`/`Package` nodes, not create `TSArtifact`/`TSPackage` twins. The + edges that stay this analyzer's own claim (`PY_PROVIDES` — "this analyzer + resolved this import to this package", `PY_UNRESOLVED_IMPORT`) keep the + `PY_` prefix; the nodes they connect to do not. +- **`PY_PROVIDES`/`PY_UNRESOLVED_IMPORT` target ids are minted here, not + looked up.** `app.external_symbols` only homes call-graph endpoints + (`Codeanalyzer._home_external_symbols` walks `app.call_graph` alone), so a + module that is imported but never called — the common case for + `PyDependency.provides_imports`, and the *only* case for an unresolved + import — has no existing `:PyExternal` ghost to MERGE onto. The projection + builds one with the same id shape `_call_endpoint`/`_home_external_symbols` + already use for a dot-less (no `.`) call-graph signature: ` + /@external/`, `module` absent — and the same two-label + `["PySymbol", "PyExternal"]` RowBuilder idiom every other ghost in this file + uses (schema declares `PyExternal`'s merge label as `PySymbol`). If a call + into that same bare name is ever projected too, both rows collapse onto one + node under `RowBuilder`'s MERGE-by-`(label, id)` semantics — correctly, + since they name the same real-world symbol. +- **`LOCKS` fans out to every lock artifact present**, not just the one that + pinned a given dependency: `PyDependency.locked_version` merges all lock + files' pins upstream (`build_dependency_view`) with no per-lock-file + attribution left to project. One lock file is the overwhelmingly common + case; revisit if a project with two conflicting lock files in one repo + turns out to matter in practice. +- Always projected regardless of `-a` — this section is L1 data, identical at + every analysis level (mirrors `analysis.json`), consistent with Neo4j's + existing full-depth-always posture for `--emit neo4j`. + ## 2026-08-27 — Artifacts, dependencies, and the `can://artifact/` namespace Design: `docs/design/specs/2026-08-27-artifacts-and-dependencies-design.md`. @@ -305,7 +350,7 @@ at every level, like entrypoints. `--resolve-installed` flag. - **Neo4j**: neutral labels `:Artifact` / `:Package` (no `Py` prefix, shared MERGE targets across analyzers); `:Package.id` is a purl (`pkg:pypi/`). - `PY_PROVIDES` joins packages to the existing `:PyExternal` ghost ids, wiring + `PY_PROVIDES` joins packages to existing-or-minted module-level `:PyExternal` ghosts, wiring dependencies into the call graph. New edges: `HAS_ARTIFACT`, `DECLARES_DEPENDENCY`, `LOCKS`, `PY_PROVIDES`, `PY_UNRESOLVED_IMPORT`. - Capture broad (config files as nodes with `roles`), extract narrow diff --git a/.github/workflows/release.yml b/.github/workflows/release.yml index f2a0ecf..5f37b77 100644 --- a/.github/workflows/release.yml +++ b/.github/workflows/release.yml @@ -114,12 +114,14 @@ jobs: # Platform-independent, version-locked release assets published alongside the # wheels/sdist: the Neo4j schema contract (so a consumer can validate - # producer/consumer compatibility without installing the package) and the - # cargo-dist-style install script. + # producer/consumer compatibility without installing the package), the + # same contract's DDL as runnable Cypher, and the cargo-dist-style + # install script. - name: Stage release assets (Neo4j schema + installer script) run: | mkdir -p release-assets uv run canpy --emit schema > release-assets/schema.json + uv run python -c "from codeanalyzer.neo4j.schema import uniqueness_constraints, INDEXES; print('\n'.join(s + ';' for s in uniqueness_constraints() + INDEXES))" > release-assets/schema.cypher cp packaging/install/canpy-installer.sh release-assets/canpy-installer.sh ls -lh release-assets diff --git a/.gitignore b/.gitignore index e678fe6..3a62449 100644 --- a/.gitignore +++ b/.gitignore @@ -195,3 +195,6 @@ node_modules/ !.claude/ .claude/* !.claude/SCHEMA_DECISIONS.md + +# Track fixture lock files past the repo-root uv.lock ignore +!test/fixtures/**/uv.lock diff --git a/CHANGELOG.md b/CHANGELOG.md index 022f59a..d9cb398 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -5,6 +5,27 @@ All notable changes to this project will be documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/), and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html). +## [Unreleased] + +### Added + +- Schema v2 now captures non-code artifacts (`application.artifacts`), declared + dependencies with provenance (`application.dependencies`), and undeclared + imports (`application.unresolved_imports`) at every analysis level (#157). + Neo4j gains language-neutral `:Artifact`/`:Package` nodes (purl ids) joined + to the existing `:PyExternal` ghosts. New flag: `--resolve-installed`. + Lockfile-only (transitive) pins are emitted as `direct: false` dependency + records attributed to the lock artifact (#152 reconciliation). +- Artifact discovery never drops a file: every non-`.py` file is now + inventoried, matched or not (unmatched decodable files as `text`/`unknown`, + anything not UTF-8 decodable as `binary` with empty `source`). New + `PyArtifact.text_truncated` field plus `--artifact-text/--no-artifact-text` + and `--artifact-text-max-bytes` flags control verbatim `source` capture; + `sha256`/`size_bytes` always reflect the full file regardless (#157). +- The release workflow now also stages `schema.cypher` (the same Neo4j + schema contract as runnable, `;`-terminated Cypher DDL -- uniqueness + constraints plus indexes) as a GitHub Release asset alongside `schema.json`. + ## [1.2.0] - 2026-08-26 ### Changed diff --git a/README.md b/README.md index 27d1fc1..8cea1d0 100644 --- a/README.md +++ b/README.md @@ -157,167 +157,156 @@ $ canpy --help Static Analysis on Python source code using Jedi and Tree sitter. -╭─ Options ────────────────────────────────────────────────────────────────────╮ -│ --version Show the canpy │ -│ version and │ -│ exit. │ -│ --input -i Path to the │ -│ project root │ -│ directory (not │ -│ required for │ -│ --emit schema). │ -│ --output -o Output directory │ -│ for artifacts. │ -│ --emit json │ -│ (analysis.json, │ -│ default) | neo4j │ -│ (graph.cypher or │ -│ live Bolt push) │ -│ | schema (the │ -│ Neo4j │ -│ schema.json │ -│ contract). │ -│ [default: json] │ -│ --app-name Logical │ -│ application name │ -│ for the graph │ -│ :PyApplication │ -│ anchor (default: │ -│ input dir name). │ -│ --neo4j-uri Push the graph │ -│ to a live Neo4j │ -│ over Bolt │ -│ (incremental); │ -│ omit to write │ -│ graph.cypher. │ -│ [env var: │ -│ NEO4J_URI] │ -│ --neo4j-user Neo4j username. │ -│ [env var: │ -│ NEO4J_USERNAME] │ -│ [default: neo4j] │ -│ --neo4j-password Neo4j password. │ -│ Prefer the env │ -│ var over the │ -│ flag (the flag │ -│ is visible in │ -│ shell history / │ -│ process list). │ -│ [env var: │ -│ NEO4J_PASSWORD] │ -│ [default: neo4j] │ -│ --neo4j-database Neo4j database │ -│ name (default: │ -│ server default). │ -│ [env var: │ -│ NEO4J_DATABASE] │ -│ --analysis-level -a Analysis depth: │ -│ [1<=x<=4] 1=symbol │ -│ table+Jedi call │ -│ graph, │ -│ 2=+defuse-linker │ -│ call graph, │ -│ 3=+native │ -│ intraprocedural │ -│ dataflow │ -│ (CFG/PDG), │ -│ 4=+interprocedu… │ -│ SDG │ -│ (param/summary │ -│ edges, │ -│ alias-aware │ -│ DDG). │ -│ [default: (1)] │ -│ --graphs Level 3+ only: │ -│ comma-separated │ -│ program-graph │ -│ sections to emit │ -│ (cfg, dfg, pdg, │ -│ sdg). Default: │ -│ cfg,dfg,pdg. │ -│ `dfg` emits the │ -│ PDG's data edges │ -│ only; `sdg` │ -│ requires -a 4. │ -│ Incompatible │ -│ with --emit │ -│ neo4j (always │ -│ full-depth). │ -│ [default: │ -│ (cfg,dfg,pdg)] │ -│ --graph-field-de… Level 3 only: │ -│ [x>=1] k-limit on │ -│ access-path │ -│ depth (x.f.g.h │ -│ with k=3 becomes │ -│ x.f.g.*). │ -│ Mandatory bound │ -│ — it is what │ -│ guarantees the │ -│ interprocedural │ -│ fixpoint │ -│ terminates. │ -│ [default: 3] │ -│ --ray --no-ray Enable Ray for │ -│ distributed │ -│ analysis. │ -│ [default: │ -│ no-ray] │ -│ --eager --lazy Enable eager or │ -│ lazy analysis. │ -│ Defaults to │ -│ lazy. │ -│ [default: lazy] │ -│ --skip-tests --include-tests Skip test files │ -│ in analysis. │ -│ [default: │ -│ skip-tests] │ -│ --no-venv --venv Skip virtualenv │ -│ creation and │ -│ dependency │ -│ installation; │ -│ resolve imports │ -│ against the │ -│ ambient Python │ -│ environment │ -│ instead. │ -│ [default: venv] │ -│ --file-name Analyze only the │ -│ specified file │ -│ (relative to │ -│ input │ -│ directory). │ -│ --cache-dir -c Directory to │ -│ store analysis │ -│ cache. Defaults │ -│ to │ -│ '.codeanalyzer' │ -│ in the input │ -│ directory. │ -│ --clear-cache --keep-cache Clear cache │ -│ after analysis. │ -│ By default, │ -│ cache is │ -│ retained. │ -│ [default: │ -│ keep-cache] │ -│ -v Increase │ -│ verbosity: -v, │ -│ -vv, -vvv │ -│ [default: 0] │ -│ --entrypoint-rul… Extra entrypoint │ -│ rules file │ -│ (YAML). │ -│ Repeatable; │ -│ merges with the │ -│ shipped rules. A │ -│ malformed file │ -│ is an error. │ -│ --help Show this │ -│ message and │ -│ exit. │ -╰──────────────────────────────────────────────────────────────────────────────╯ +╭─ Options ────────────────────────────────────────────────────────────────────────────────────────╮ +│ --version Show the canpy version │ +│ and exit. │ +│ --input -i Path to the project │ +│ root directory (not │ +│ required for --emit │ +│ schema). │ +│ --output -o Output directory for │ +│ artifacts. │ +│ --emit Output target: json │ +│ (analysis.json, │ +│ default) | neo4j │ +│ (graph.cypher or live │ +│ Bolt push) | schema │ +│ (the Neo4j schema.json │ +│ contract). │ +│ [default: json] │ +│ --app-name Logical application │ +│ name for the graph │ +│ :PyApplication anchor │ +│ (default: input dir │ +│ name). │ +│ --neo4j-uri Push the graph to a │ +│ live Neo4j over Bolt │ +│ (incremental); omit to │ +│ write graph.cypher. │ +│ [env var: NEO4J_URI] │ +│ --neo4j-user Neo4j username. │ +│ [env var: │ +│ NEO4J_USERNAME] │ +│ [default: neo4j] │ +│ --neo4j-password Neo4j password. Prefer │ +│ the env var over the │ +│ flag (the flag is │ +│ visible in shell │ +│ history / process │ +│ list). │ +│ [env var: │ +│ NEO4J_PASSWORD] │ +│ [default: neo4j] │ +│ --neo4j-database Neo4j database name │ +│ (default: server │ +│ default). │ +│ [env var: │ +│ NEO4J_DATABASE] │ +│ --analysis-level -a [1<=x<=4] Analysis depth: │ +│ 1=symbol table+Jedi │ +│ call graph, │ +│ 2=+defuse-linker call │ +│ graph, 3=+native │ +│ intraprocedural │ +│ dataflow (CFG/PDG), │ +│ 4=+interprocedural SDG │ +│ (param/summary edges, │ +│ alias-aware DDG). │ +│ [default: (1)] │ +│ --graphs Level 3+ only: │ +│ comma-separated │ +│ program-graph sections │ +│ to emit (cfg, dfg, │ +│ pdg, sdg). Default: │ +│ cfg,dfg,pdg. `dfg` │ +│ emits the PDG's data │ +│ edges only; `sdg` │ +│ requires -a 4. │ +│ Incompatible with │ +│ --emit neo4j (always │ +│ full-depth). │ +│ [default: │ +│ (cfg,dfg,pdg)] │ +│ --graph-field-depth [x>=1] Level 3 only: k-limit │ +│ on access-path depth │ +│ (x.f.g.h with k=3 │ +│ becomes x.f.g.*). │ +│ Mandatory bound — it │ +│ is what guarantees the │ +│ interprocedural │ +│ fixpoint terminates. │ +│ [default: 3] │ +│ --ray --no-ray Enable Ray for │ +│ distributed analysis. │ +│ [default: no-ray] │ +│ --eager --lazy Enable eager or lazy │ +│ analysis. Defaults to │ +│ lazy. │ +│ [default: lazy] │ +│ --skip-tests --include-tests Skip test files in │ +│ analysis. │ +│ [default: skip-tests] │ +│ --no-venv --venv Skip virtualenv │ +│ creation and │ +│ dependency │ +│ installation; resolve │ +│ imports against the │ +│ ambient Python │ +│ environment instead. │ +│ [default: venv] │ +│ --resolve-installed Additionally bind │ +│ imports via the │ +│ project venv's │ +│ installed metadata │ +│ (*.dist-info); output │ +│ becomes │ +│ machine-dependent │ +│ (prov: │ +│ installed-metadata). │ +│ --file-name Analyze only the │ +│ specified file │ +│ (relative to input │ +│ directory). │ +│ --cache-dir -c Directory to store │ +│ analysis cache. │ +│ Defaults to │ +│ '.codeanalyzer' in the │ +│ input directory. │ +│ --clear-cache --keep-cache Clear cache after │ +│ analysis. By default, │ +│ cache is retained. │ +│ [default: keep-cache] │ +│ -v Increase verbosity: │ +│ -v, -vv, -vvv │ +│ [default: 0] │ +│ --entrypoint-rules Extra entrypoint rules │ +│ file (YAML). │ +│ Repeatable; merges │ +│ with the shipped │ +│ rules. A malformed │ +│ file is an error. │ +│ --artifact-text --no-artifact-text Capture verbatim │ +│ `source` text on │ +│ discovered artifacts. │ +│ --no-artifact-text │ +│ empties `source` │ +│ everywhere (inventory │ +│ unchanged). │ +│ [default: │ +│ artifact-text] │ +│ --artifact-text-max-by… [x>=1] Per-file byte cap on │ +│ captured artifact │ +│ `source`; a decodable │ +│ file over the cap is │ +│ truncated │ +│ (text_truncated=True). │ +│ sha256/size_bytes │ +│ always reflect the │ +│ full file. │ +│ [default: 262144] │ +│ --help Show this message and │ +│ exit. │ +╰──────────────────────────────────────────────────────────────────────────────────────────────────╯ ``` @@ -485,6 +474,18 @@ A **callable** (function or method) carries its own CPG, keyed by node id: } ``` +The application envelope also contains three substrate sections: + +- **`artifacts`** — discovered non-code files (manifests, configs, Docker files, CI workflows, + packaging files, scripts, docs, and legal files) with extraction status (`none`, `partial`, or + `full`; default `none`), keyed by relative path; each artifact carries the + `can://artifact//` id namespace. +- **`dependencies`** — declared packages with kind (`runtime`/`dev`/`optional`/`build`), spec, + locked version, and provenance (`prov`): where each binding came from (manifest file, lock file, + installed metadata). +- **`unresolved_imports`** — modules imported but not resolvable in the declared dependency set, + one entry per module. + Notable properties: - **Durable `can://` ids** identify every node at callable granularity and above @@ -513,9 +514,13 @@ binary format). ### Neo4j graph -`--emit neo4j` projects the same schema v2.0.0 analysis into a labeled property graph. Every node -label is `Py`-prefixed and every relationship type is `PY_`-prefixed (e.g. `:PyClass`, `PY_CALLS`) -so multiple language analyzers can share one database without label or relationship-type collisions. +`--emit neo4j` projects the same schema v2.0.0 analysis into a labeled property graph. Every +Python-specific node label is `Py`-prefixed and every Python-specific relationship type is +`PY_`-prefixed (e.g. `:PyClass`, `PY_CALLS`) so multiple language analyzers can share one database +without label or relationship-type collisions. The one deliberate exception is the language-neutral +`Artifact`/`Package` subgraph (non-code files and third-party dependencies) — those nodes carry no +`Py` prefix, since they are meant as cross-language merge targets: a sibling-language analyzer over +the same repo should land on the same `Artifact`/`Package` nodes, not a per-language duplicate. Declarations are keyed by their **`can://` id** under a shared `:PySymbol` label; calls, imports, inheritance, decorators, and call sites are relationships. At `-a 3`/`-a 4` the projection gains the **CPG overlay** — `:PyBodyNode` nodes (statements, and at level 4 the parameter vertices) wired by @@ -604,7 +609,7 @@ WHERE f.id STARTS WITH "can://python/myapp/src/api.py" RETURN a.id, f.id ``` -Landing with #157 (schema v2 artifacts + dependencies, 1.3.0) — not in the current release: +Artifact and dependency queries (1.3.0+): ```cypher // all container/orchestration configs in the app diff --git a/codeanalyzer/__main__.py b/codeanalyzer/__main__.py index 3bedc66..053d6c2 100644 --- a/codeanalyzer/__main__.py +++ b/codeanalyzer/__main__.py @@ -193,6 +193,14 @@ def main( "imports against the ambient Python environment instead.", ), ] = False, + resolve_installed: Annotated[ + bool, + typer.Option( + "--resolve-installed", + help="Additionally bind imports via the project venv's installed metadata " + "(*.dist-info); output becomes machine-dependent (prov: installed-metadata).", + ), + ] = False, file_name: Annotated[ Optional[Path], typer.Option( @@ -226,6 +234,24 @@ def main( "the shipped rules. A malformed file is an error.", ), ] = None, + artifact_text: Annotated[ + bool, + typer.Option( + "--artifact-text/--no-artifact-text", + help="Capture verbatim `source` text on discovered artifacts. " + "--no-artifact-text empties `source` everywhere (inventory unchanged).", + ), + ] = True, + artifact_text_max_bytes: Annotated[ + int, + typer.Option( + "--artifact-text-max-bytes", + help="Per-file byte cap on captured artifact `source`; a decodable " + "file over the cap is truncated (text_truncated=True). " + "sha256/size_bytes always reflect the full file.", + min=1, + ), + ] = 262144, ): # Determinism: pin the interpreter hash seed before any analysis (no-op # when PYTHONHASHSEED is already set; --version exits before this). @@ -303,11 +329,14 @@ def main( rebuild_analysis=rebuild_analysis, skip_tests=skip_tests, no_venv=no_venv, + resolve_installed=resolve_installed, file_name=file_name, cache_dir=cache_dir, clear_cache=clear_cache, verbosity=verbosity, entrypoint_rules=tuple(entrypoint_rules or ()), + artifact_text=artifact_text, + artifact_text_max_bytes=artifact_text_max_bytes, ) _set_log_level(options.verbosity) diff --git a/codeanalyzer/artifacts/__init__.py b/codeanalyzer/artifacts/__init__.py new file mode 100644 index 0000000..a7af4c9 --- /dev/null +++ b/codeanalyzer/artifacts/__init__.py @@ -0,0 +1,11 @@ +"""Non-code artifact capture and dependency extraction (spec 2026-08-27). + +Capture never drops a file (every non-`.py` file becomes a +:class:`~codeanalyzer.schema.py_schema.PyArtifact`, rule-matched or not, +text or binary -- issue #157 follow-up); extraction is narrow (only +dependency manifests are parsed for meaning in this unit).""" + +from codeanalyzer.artifacts.dependencies import build_dependency_view +from codeanalyzer.artifacts.discovery import discover_artifacts + +__all__ = ["discover_artifacts", "build_dependency_view"] diff --git a/codeanalyzer/artifacts/dependencies.py b/codeanalyzer/artifacts/dependencies.py new file mode 100644 index 0000000..3207b56 --- /dev/null +++ b/codeanalyzer/artifacts/dependencies.py @@ -0,0 +1,233 @@ +"""PyDependency / PyImportBinding construction from discovered artifacts. + +Deterministic by default: reads only repo files. ``resolve_installed`` adds +filesystem reads of ``/**/site-packages/*.dist-info`` (never runs an +interpreter), tagged ``prov: installed-metadata``. + +``-r``/``-c`` refs in a requirements-format manifest are chased one level +only, by design: a chased target's own refs are not followed further. + +``unresolved_imports`` is byte-identical run-to-run only within one Python +minor version: it is filtered against ``sys.stdlib_module_names``, and that +set's membership varies across minors (a module added to or removed from the +stdlib).""" + +import posixpath +import re +import sys +from pathlib import Path +from typing import Dict, List, Optional, Tuple + +from codeanalyzer.artifacts.parsers import ( + RawDep, _kind_for_requirements, normalize_name, parse_lock_pins, parse_manifest, + parse_requirement_refs, +) +from codeanalyzer.schema.py_schema import ( + PyArtifact, PyDependency, PyImportBinding, PyModule, +) + +_LOCK_BASENAMES = ("poetry.lock", "uv.lock", "Pipfile.lock") + +# Small, non-exhaustive alias table for the worst offenders; everything else +# rides the same-name rule or --resolve-installed. prov: heuristic. Identity +# entries (key == value) do not belong here -- the same-name rule below +# already covers them, and a redundant identity entry only mints a spurious +# "heuristic" prov on what is actually a plain same-name match. +_KNOWN_IMPORT_ALIASES: Dict[str, str] = { + "pyyaml": "yaml", "beautifulsoup4": "bs4", "pillow": "PIL", + "scikit-learn": "sklearn", "opencv-python": "cv2", "python-dateutil": "dateutil", + "msgpack-python": "msgpack", "protobuf": "google.protobuf", "attrs": "attr", +} + + +def _stdlib_names() -> set: + return set(getattr(sys, "stdlib_module_names", ())) | {"__future__"} + + +def _installed_top_levels(venv_dir: Optional[Path]) -> Dict[str, List[str]]: + """{normalized dist name: [top-level import names]} from *.dist-info files.""" + out: Dict[str, List[str]] = {} + if venv_dir is None or not venv_dir.exists(): + return out + for di in sorted(venv_dir.glob("**/site-packages/*.dist-info")): + name = None + meta = di / "METADATA" + if meta.exists(): + m = re.search(r"^Name:\s*(.+)$", meta.read_text(errors="replace"), re.M) + if m: + name = normalize_name(m.group(1).strip()) + if name is None: + continue + tl = di / "top_level.txt" + if tl.exists(): + out[name] = [l.strip() for l in tl.read_text().splitlines() if l.strip()] + return out + + +def _is_requirements_format(path: str) -> bool: + base = path.rsplit("/", 1)[-1] + return base.startswith("requirements") and base.endswith(".txt") + + +def _resolve_ref(manifest_path: str, ref: str) -> Optional[str]: + """POSIX-join a ``-r``/``-c`` ref against its manifest's directory, + normalized and repo-relative. ``None`` if it would escape ``project_dir``.""" + manifest_dir = manifest_path.rsplit("/", 1)[0] if "/" in manifest_path else "" + joined = posixpath.normpath(posixpath.join(manifest_dir, ref) if manifest_dir else ref) + if joined == ".." or joined.startswith("../") or posixpath.isabs(joined): + return None + return joined + + +def _full_text(project_dir: Path, path: str, art: PyArtifact) -> str: + """Manifest/lock extraction must never depend on the stored ``source`` -- + that's capped by ``text_max_bytes`` and emptied by ``capture_text=False`` + (both payload-size controls on the JSON/Neo4j payload, not extraction + controls). Read the real file fresh instead; fall back to ``art.source`` + only if it is gone (e.g. a synthetic artifact in a unit test, or the file + vanished mid-run).""" + try: + return (project_dir / path).read_bytes().decode("utf-8") + except (OSError, UnicodeDecodeError): + return art.source + + +def build_dependency_view( + artifacts: Dict[str, PyArtifact], + modules: Dict[str, PyModule], + project_dir: Path, + venv_dir: Optional[Path], + resolve_installed: bool, +) -> Tuple[List[PyDependency], List[PyImportBinding]]: + deps: List[PyDependency] = [] + + def _emit(raw: List[RawDep], declared_in: str, kind_override: Optional[str] = None) -> None: + for r in raw: + deps.append(PyDependency( + name=r.name, spec=r.spec, + kind=kind_override if kind_override is not None else r.kind, + extras=sorted(r.extras), declared_in=declared_in, prov=["declared"], + )) + + # 1. Declared records from every dependency-manifest artifact (non-lock). + for path in sorted(artifacts): + art = artifacts[path] + if "dependency-manifest" not in art.roles: + continue + if path.rsplit("/", 1)[-1] in _LOCK_BASENAMES: + continue + text = _full_text(project_dir, path, art) + raw, partial = parse_manifest(path, text) + art.extraction = "partial" if partial else "full" + _emit(raw, art.id) + + # 1b. -r/-c refs: a target that is itself a dependency-manifest is + # parsed on its own above; a target with no RULES match for that role + # (e.g. base.txt -- never-drop inventory still captures it, just not + # as a manifest) is chased here and attributed to the referring + # artifact. Gate on the role, not mere presence in `artifacts`: since + # #157 every file is discovered, so presence alone no longer implies + # "already parsed as a manifest above". + if not _is_requirements_format(path): + continue + for ref in parse_requirement_refs(text): + resolved = _resolve_ref(path, ref) + if resolved is None: + continue + target_art = artifacts.get(resolved) + if target_art is not None and "dependency-manifest" in target_art.roles: + continue + target = project_dir / resolved + if not target.is_file(): + continue + try: + ref_text = target.read_bytes().decode("utf-8") + except UnicodeDecodeError: + continue + # Force requirements-format dispatch (chased targets may not be + # named requirements*.txt), but recompute kind from the real + # basename so e.g. `-r dev.txt` still yields kind="dev". + raw_ref, _ = parse_manifest("requirements.txt", ref_text) + real_kind = _kind_for_requirements(resolved.rsplit("/", 1)[-1]) + _emit(raw_ref, art.id, kind_override=real_kind) + + # 2. Lock backfill (locked_version + prov "lockfile"). A pin with no + # manifest declaration is a *transitive* dependency: emitted with + # direct=False, attributed to the lock artifact (#152 reconciliation). + pins: Dict[str, str] = {} + pin_lock_artifact: Dict[str, str] = {} + for path in sorted(artifacts): + if path.rsplit("/", 1)[-1] in _LOCK_BASENAMES: + lock_text = _full_text(project_dir, path, artifacts[path]) + lock_pins = parse_lock_pins(path, lock_text) + pins.update(lock_pins) + for name in lock_pins: + pin_lock_artifact[name] = artifacts[path].id + # A lock with real content that yields zero pins failed to parse + # (corrupt/unrecognized shape) -- don't claim "full" extraction + # for nothing extracted. An empty/whitespace-only lock is not a + # failure (nothing to extract), so it still counts as "full". + artifacts[path].extraction = ( + "full" if lock_pins or not lock_text.strip() else "partial" + ) + for d in deps: + if d.name in pins: + d.locked_version = pins[d.name] + d.prov = sorted(set(d.prov) | {"lockfile"}) + declared_names = {d.name for d in deps} + for name in sorted(set(pins) - declared_names): + deps.append(PyDependency( + name=name, kind="runtime", declared_in=pin_lock_artifact[name], + direct=False, locked_version=pins[name], prov=["lockfile"], + )) + + # 3. Import universe from the symbol table (top-level segments only). + # `module_name` is `py_file.stem` -- the leaf filename only (e.g. "api" + # for "odoo/api.py"), never the package path -- so it alone misses the + # top-level package name itself. Derive that from the symbol-table KEYS + # (repo-relative POSIX paths) too: first path segment when nested, else + # the root file's own stem. Keep the module_name-derived stems as well + # (harmless -- still excludes leaf-name imports the key pass can't see). + local = {m.module_name.split(".")[0] for m in modules.values() if m.module_name} + local |= { + key.split("/", 1)[0] if "/" in key else Path(key).stem for key in modules + } + stdlib = _stdlib_names() + imported: set = set() + for m in modules.values(): + for imp in m.imports or []: + top = (imp.module or imp.name or "").split(".")[0] + if top and top not in stdlib and top not in local: + imported.add(top) + + # 4. provides_imports: same-name rule, alias table, optional installed metadata. + installed = _installed_top_levels(venv_dir) if resolve_installed else {} + for d in deps: + provides: List[str] = [] + same = d.name.replace("-", "_") + for candidate in {d.name, same}: + if candidate in imported: + provides.append(candidate) + alias = _KNOWN_IMPORT_ALIASES.get(d.name) + if alias and alias.split(".")[0] in imported: + provides.append(alias) + d.prov = sorted(set(d.prov) | {"heuristic"}) + if d.name in installed: + for top in installed[d.name]: + if top in imported and top not in provides: + provides.append(top) + d.prov = sorted(set(d.prov) | {"installed-metadata"}) + d.provides_imports = sorted(set(provides)) + + # 5. Unresolved: imported, not stdlib/local, not provided by any dependency. + # Top-level segment only: `imported` (step 3) is already top-level-only, but + # a dotted alias (e.g. protobuf -> "google.protobuf") puts the FULL dotted + # path into provides_imports, so comparing it against `imported` verbatim + # never matches and "google" falsely resurfaces as unresolved even though + # protobuf declares it. + provided = {p.split(".")[0] for d in deps for p in d.provides_imports} + unresolved = [ + PyImportBinding(module=m) for m in sorted(imported - provided) + ] + deps.sort(key=lambda d: (d.name, d.declared_in)) + return deps, unresolved diff --git a/codeanalyzer/artifacts/discovery.py b/codeanalyzer/artifacts/discovery.py new file mode 100644 index 0000000..5bf20c7 --- /dev/null +++ b/codeanalyzer/artifacts/discovery.py @@ -0,0 +1,158 @@ +from __future__ import annotations + +import fnmatch +import hashlib +from pathlib import Path +from typing import Dict, List, Tuple + +from codeanalyzer.schema.ids import artifact_id +from codeanalyzer.schema.py_schema import PyArtifact + +# (glob pattern against the repo-relative POSIX path, format, roles). +# First match wins; patterns are checked in order. +RULES: List[Tuple[str, str, List[str]]] = [ + ("requirements*.txt", "requirements", ["dependency-manifest"]), + ("pyproject.toml", "toml", ["dependency-manifest", "tool-config"]), + ("setup.py", "text", ["dependency-manifest"]), + ("setup.cfg", "ini", ["dependency-manifest", "tool-config"]), + ("Pipfile", "toml", ["dependency-manifest"]), + ("Pipfile.lock", "json", ["dependency-manifest"]), + ("poetry.lock", "toml", ["dependency-manifest"]), + ("uv.lock", "toml", ["dependency-manifest"]), + ("environment.yml", "yaml", ["dependency-manifest"]), + ("environment.yaml", "yaml", ["dependency-manifest"]), + ("Dockerfile", "dockerfile", ["container-image"]), + ("*.dockerfile", "dockerfile", ["container-image"]), + ("docker-compose*.yml", "yaml", ["service-topology"]), + ("docker-compose*.yaml", "yaml", ["service-topology"]), + ("compose.yml", "yaml", ["service-topology"]), + ("compose.yaml", "yaml", ["service-topology"]), + ("k8s/*.yml", "yaml", ["service-topology"]), + ("k8s/*.yaml", "yaml", ["service-topology"]), + ("kind/*.yml", "yaml", ["service-topology"]), + ("kind/*.yaml", "yaml", ["service-topology"]), + ("Chart.yaml", "yaml", ["service-topology"]), + ("values.yaml", "yaml", ["service-topology"]), + (".github/workflows/*.yml", "yaml", ["ci"]), + (".github/workflows/*.yaml", "yaml", ["ci"]), + (".gitlab-ci.yml", "yaml", ["ci"]), + (".env", "text", ["env"]), + (".env.*", "text", ["env"]), + ("tox.ini", "ini", ["tool-config"]), + ("noxfile.py", "text", ["tool-config"]), + ("Makefile", "text", ["tool-config"]), + ("MANIFEST.in", "text", ["packaging"]), + ("LICENSE*", "text", ["legal"]), + ("COPYRIGHT*", "text", ["legal"]), + ("NOTICE*", "text", ["legal"]), + ("*.md", "text", ["docs"]), + ("*.rst", "text", ["docs"]), + ("*.cfg", "ini", ["unknown"]), + ("*.toml", "toml", ["unknown"]), +] + +_IGNORED_DIRS = { + ".git", ".hg", ".svn", "__pycache__", ".venv", "venv", ".tox", ".nox", + "node_modules", ".mypy_cache", ".pytest_cache", ".ruff_cache", ".idea", + "build", "dist", ".eggs", ".codeanalyzer", "virtualenv", "site-packages", +} + + +def _classify(rel_posix: str) -> Tuple[str, List[str]] | None: + name = rel_posix.rsplit("/", 1)[-1] + for pattern, fmt, roles in RULES: + target = rel_posix if ("/" in pattern or pattern.startswith("**")) else name + if fnmatch.fnmatch(target, pattern): + return fmt, roles + return None + + +def _capture_source( + raw: bytes, text: str, capture_text: bool, text_max_bytes: int +) -> Tuple[str, bool]: + """Decide ``(source, text_truncated)`` for a decodable file. + + Slices ``raw`` (not ``text``) for the cap, so it is a true byte cap even + when it lands inside a multi-byte character -- ``errors="ignore"`` drops + the dangling partial char at the cut, so this never raises.""" + if not capture_text: + return "", False + if len(raw) <= text_max_bytes: + return text, False + return raw[:text_max_bytes].decode("utf-8", errors="ignore"), True + + +def discover_artifacts( + project_dir: Path, + app_name: str, + *, + capture_text: bool = True, + text_max_bytes: int = 262144, +) -> Dict[str, PyArtifact]: + """Walk the project and return every file as an artifact, sorted by path. + + Never-drop inventory (issue #157 follow-up): a rule-matched file keeps its + RULES format/roles; everything else falls back to ``text``/``["unknown"]`` + (``source`` captured), or ``binary``/empty ``source`` when it is not UTF-8 + decodable -- rule-matched but undecodable files downgrade to ``binary`` + too, keeping the rule's roles. The one exclusion is a `.py` file no RULES + entry names: the symbol table already owns it. ``setup.py`` is the + deliberate exception -- it IS rule-matched (a dependency-manifest), so it + is captured like any other manifest despite the `.py` suffix. + + ``capture_text=False`` empties ``source`` everywhere (inventory otherwise + identical); a decodable file over ``text_max_bytes`` gets a truncated + ``source`` and ``text_truncated=True`` -- except a ``dependency-manifest`` + role artifact, which is always captured in full when decodable and + ``capture_text`` is on: its source is what ``build_dependency_view`` + parses, not bulk/incidental content, so the byte cap does not apply to + it (``capture_text=False`` still empties it like everything else). + ``sha256``/``size_bytes`` always reflect the full file regardless of + either knob.""" + out: Dict[str, PyArtifact] = {} + for path in sorted(project_dir.rglob("*")): + if not path.is_file(): + continue + rel = path.relative_to(project_dir) + if any(part in _IGNORED_DIRS for part in rel.parts): + continue + rel_posix = rel.as_posix() + name = rel_posix.rsplit("/", 1)[-1] + hit = _classify(rel_posix) + if hit is None and name.endswith(".py"): + continue # symbol table's domain (setup.py is rule-matched above) + + raw = path.read_bytes() + try: + text = raw.decode("utf-8") + decodable = True + except UnicodeDecodeError: + text, decodable = "", False + + if hit is not None: + fmt, roles = hit + else: + fmt, roles = "text", ["unknown"] + # Extensionless shebang script (e.g. odoo-bin): no RULES glob can + # name these (nothing to match on but the shebang itself), so this + # is the one deterministic content-sniff refinement. + if decodable and "." not in name and text.startswith("#!"): + roles = ["script"] + if decodable: + # A dependency-manifest's source IS the extracted meaning (build_ + # dependency_view parses it) -- the byte cap targets bulk/incidental + # assets, never the files extraction depends on, so manifests are + # exempt from it. capture_text=False still empties source (handled + # inside _capture_source); only the byte CAP is bypassed here. + cap = len(raw) if "dependency-manifest" in roles else text_max_bytes + source, text_truncated = _capture_source(raw, text, capture_text, cap) + else: + fmt, source, text_truncated = "binary", "", False + + out[rel_posix] = PyArtifact( + id=artifact_id(app_name, rel_posix), path=rel_posix, format=fmt, + roles=list(roles), size_bytes=len(raw), + sha256=hashlib.sha256(raw).hexdigest(), + source=source, text_truncated=text_truncated, + ) + return out diff --git a/codeanalyzer/artifacts/parsers.py b/codeanalyzer/artifacts/parsers.py new file mode 100644 index 0000000..f6eaf3b --- /dev/null +++ b/codeanalyzer/artifacts/parsers.py @@ -0,0 +1,248 @@ +"""Dependency-manifest readers. Pure text-in/records-out; no execution, no I/O.""" + +import ast +import configparser +import json +import re +import sys +from dataclasses import dataclass +from typing import Dict, List, Optional, Tuple + +if sys.version_info >= (3, 11): + import tomllib +else: # pragma: no cover - exercised on the 3.10 CI leg + import tomli as tomllib + +import yaml + + +@dataclass(frozen=True) +class RawDep: + name: str # PEP 503 normalized + spec: str = "" + kind: str = "runtime" # runtime|dev|optional|build + extras: Tuple[str, ...] = () + + +def normalize_name(raw: str) -> str: + return re.sub(r"[-_.]+", "-", raw).lower() + + +_REQ_LINE = re.compile( + r"^\s*(?P[A-Za-z0-9][A-Za-z0-9._-]*)\s*(?:\[(?P[^\]]+)\])?\s*(?P[^;#]*)" +) + + +def parse_requirement_line(line: str, kind: str = "runtime") -> Optional[RawDep]: + """One PEP 508-ish requirement line -> RawDep (None for options/paths/URLs).""" + line = line.split("#", 1)[0].strip() + if not line or line.startswith(("-", "--")) or line.startswith((".", "/")): + return None + if " @ " in line: + line = line.split(" @ ", 1)[0].strip() # direct ref (PEP 508): keep the name, drop the URL + elif "://" in line: + return None + m = _REQ_LINE.match(line) + if not m: + return None + extras = tuple(e.strip() for e in (m.group("extras") or "").split(",") if e.strip()) + spec = m.group("spec").strip().rstrip(",") + if spec.endswith("\\"): + spec = spec[:-1].rstrip() # pip-compile --generate-hashes line continuation + return RawDep(normalize_name(m.group("name")), spec, kind, extras) + + +_REF_LINE = re.compile(r"^(?:-r|--requirement|-c|--constraint)\s+(\S+)") + + +def parse_requirement_refs(text: str) -> List[str]: + """-r/--requirement/-c/--constraint targets from a requirements file, in order.""" + out = [] + for line in text.splitlines(): + line = line.split("#", 1)[0].strip() + m = _REF_LINE.match(line) + if m: + out.append(m.group(1)) + return out + + +def _kind_for_requirements(basename: str) -> str: + return "dev" if re.search(r"\b(dev|test|lint|doc)\b", basename, re.I) else "runtime" + + +def _parse_requirements(basename: str, text: str) -> List[RawDep]: + kind = _kind_for_requirements(basename) + out = [] + for line in text.splitlines(): + dep = parse_requirement_line(line, kind) + if dep: + out.append(dep) + return out + + +def _spec_and_extras(spec) -> Tuple[str, Tuple[str, ...]]: + """Poetry/Pipfile dep spec (bare string or {version, extras} inline table) -> (version, extras).""" + if isinstance(spec, dict): + return spec.get("version", "") or "", tuple(spec.get("extras", []) or []) + return (spec if isinstance(spec, str) else ""), () + + +def _parse_pyproject(text: str) -> List[RawDep]: + data = tomllib.loads(text) + out: List[RawDep] = [] + for req in (data.get("build-system") or {}).get("requires", []): + d = parse_requirement_line(req, "build") + if d: + out.append(d) + proj = data.get("project") or {} + for req in proj.get("dependencies", []): + d = parse_requirement_line(req) + if d: + out.append(d) + for group in (proj.get("optional-dependencies") or {}).values(): + for req in group: + d = parse_requirement_line(req, "optional") + if d: + out.append(d) + poetry = ((data.get("tool") or {}).get("poetry")) or {} + for name, spec in (poetry.get("dependencies") or {}).items(): + if normalize_name(name) == "python": + continue + v, ex = _spec_and_extras(spec) + out.append(RawDep(normalize_name(name), v, "runtime", ex)) + for gname, group in (poetry.get("group") or {}).items(): + kind = "dev" if gname == "dev" else "optional" + for name, spec in (group.get("dependencies") or {}).items(): + v, ex = _spec_and_extras(spec) + out.append(RawDep(normalize_name(name), v, kind, ex)) + for name, spec in (poetry.get("dev-dependencies") or {}).items(): # legacy poetry + v, ex = _spec_and_extras(spec) + out.append(RawDep(normalize_name(name), v, "dev", ex)) + return out + + +def _parse_setup_py(text: str) -> Tuple[List[RawDep], bool]: + """Static AST only. Literal lists lift; anything computed -> partial=True.""" + try: + tree = ast.parse(text) + except SyntaxError: + return [], True + out: List[RawDep] = [] + partial = False + for node in ast.walk(tree): + if not (isinstance(node, ast.Call) and getattr(node.func, "id", getattr(node.func, "attr", "")) == "setup"): + continue + for kw in node.keywords: + if kw.arg == "install_requires": + lifted = _lift_str_list(kw.value) + if lifted is None: + partial = True + else: + out += [d for d in (parse_requirement_line(s) for s in lifted) if d] + elif kw.arg == "extras_require": + if not isinstance(kw.value, ast.Dict): + partial = True + continue + for v in kw.value.values: + lifted = _lift_str_list(v) + if lifted is None: + partial = True + else: + out += [d for d in (parse_requirement_line(s, "optional") for s in lifted) if d] + return out, partial + + +def _lift_str_list(node: ast.AST) -> Optional[List[str]]: + if isinstance(node, (ast.List, ast.Tuple)) and all( + isinstance(e, ast.Constant) and isinstance(e.value, str) for e in node.elts + ): + return [e.value for e in node.elts] + return None + + +def _parse_setup_cfg(text: str) -> List[RawDep]: + cp = configparser.ConfigParser() + cp.read_string(text) + out: List[RawDep] = [] + if cp.has_option("options", "install_requires"): + for line in cp.get("options", "install_requires").splitlines(): + d = parse_requirement_line(line) + if d: + out.append(d) + if cp.has_section("options.extras_require"): + for _, val in cp.items("options.extras_require"): + for line in val.splitlines(): + d = parse_requirement_line(line, "optional") + if d: + out.append(d) + return out + + +def _parse_pipfile(text: str) -> List[RawDep]: + data = tomllib.loads(text) + out: List[RawDep] = [] + for section, kind in (("packages", "runtime"), ("dev-packages", "dev")): + for name, spec in (data.get(section) or {}).items(): + v, ex = _spec_and_extras(spec) + out.append(RawDep(normalize_name(name), "" if v == "*" else v, kind, ex)) + return out + + +def _parse_environment_yml(text: str) -> List[RawDep]: + data = yaml.safe_load(text) or {} + out: List[RawDep] = [] + for item in data.get("dependencies") or []: + if isinstance(item, str): + d = parse_requirement_line(item) + if d and d.name not in ("pip", "python"): + out.append(d) + elif isinstance(item, dict): + for req in item.get("pip") or []: + d = parse_requirement_line(req) + if d: + out.append(d) + return out + + +def parse_manifest(path: str, text: str) -> Tuple[List[RawDep], bool]: + """Dispatch on basename -> (records, partial). Unknown basenames -> ([], False).""" + base = path.rsplit("/", 1)[-1] + try: + if base.startswith("requirements") and base.endswith(".txt"): + return _parse_requirements(base, text), False + if base == "pyproject.toml": + return _parse_pyproject(text), False + if base == "setup.py": + return _parse_setup_py(text) + if base == "setup.cfg": + return _parse_setup_cfg(text), False + if base == "Pipfile": + return _parse_pipfile(text), False + if base in ("environment.yml", "environment.yaml"): + return _parse_environment_yml(text), False + except Exception: + return [], True # unparseable manifest: keep the artifact, flag extraction + return [], False + + +def parse_lock_pins(path: str, text: str) -> Dict[str, str]: + """Lock file -> {normalized name: pinned version}. Never creates records.""" + base = path.rsplit("/", 1)[-1] + try: + if base in ("poetry.lock", "uv.lock"): + data = tomllib.loads(text) + return { + normalize_name(p["name"]): str(p["version"]) + for p in data.get("package") or [] if "name" in p and "version" in p + } + if base == "Pipfile.lock": + data = json.loads(text) + out = {} + for section in ("default", "develop"): + for name, meta in (data.get(section) or {}).items(): + v = (meta or {}).get("version", "") + out[normalize_name(name)] = v.lstrip("=") + return out + except Exception: + return {} + return {} diff --git a/codeanalyzer/core.py b/codeanalyzer/core.py index 6da3473..f29ba77 100644 --- a/codeanalyzer/core.py +++ b/codeanalyzer/core.py @@ -648,6 +648,23 @@ def analyze(self) -> Analysis: detect_entrypoints(app, self.project_dir, self.options.entrypoint_rules) + # Artifacts + dependencies: L1 data, every level, never varies with -a + # (spec 2026-08-27). Deterministic by default; venv probing is opt-in. + from codeanalyzer.artifacts import build_dependency_view, discover_artifacts + + app.artifacts = discover_artifacts( + self.project_dir, app_name, + capture_text=self.options.artifact_text, + text_max_bytes=self.options.artifact_text_max_bytes, + ) + app.dependencies, app.unresolved_imports = build_dependency_view( + app.artifacts, + app.symbol_table, + self.project_dir, + self.virtualenv if self.options.resolve_installed else None, + self.options.resolve_installed, + ) + # L3: intraprocedural dataflow (CFG/CDG/DDG) emitted onto the v2 tree. if self.analysis_level >= 3: from codeanalyzer.dataflow.builder import ( diff --git a/codeanalyzer/neo4j/project.py b/codeanalyzer/neo4j/project.py index 6001fcd..39ce62b 100644 --- a/codeanalyzer/neo4j/project.py +++ b/codeanalyzer/neo4j/project.py @@ -47,6 +47,7 @@ PyModule, PyVariableDeclaration, ) +from codeanalyzer.schema.ids import application_id, purl_pypi from codeanalyzer.schema.py_schema import PyDecorator @@ -100,6 +101,10 @@ def project(app: PyApplication, app_name: str, sig_to_id: dict, # MERGE — a no-op when no callable carries L3 fields (levels 1/2). _project_program_graphs(b, app, externals, sig_to_id) + # Neutral artifact/dependency subgraph (Task 6). L1 data — always present, + # full-depth-always regardless of -a. + _project_artifacts(b, app, app_name, app_ref) + return b.finish() @@ -248,6 +253,125 @@ def _project_program_graphs( ) +# ---------------------------------------------------------------------------------------------- +# Artifact / dependency subgraph (spec 2026-08-27, Task 6) +# ---------------------------------------------------------------------------------------------- + +_LOCK_BASENAMES = ("poetry.lock", "uv.lock", "Pipfile.lock") + + +def _import_ghost(b: RowBuilder, app_can_id: str, name: str) -> NodeRef: + """A ``:PyExternal`` ghost for a bare imported module name (``PY_PROVIDES``'s + ``provides_imports`` entries, ``PY_UNRESOLVED_IMPORT``'s ``module``). + + ``app.external_symbols`` only homes call-graph endpoints (``_home_external_ + symbols`` walks ``app.call_graph``), so a module that is imported but never + called — the overwhelmingly common case for ``provides_imports`` and the + *only* case for an unresolved import — has no existing ghost to MERGE onto. + This builds one with the same id shape ``_call_endpoint``/``_home_external_ + symbols`` use for a dot-less (no ``.`` in the signature) call target: + ``/@external/``, ``module=None``. Same two-label + ``["PySymbol", "PyExternal"]`` idiom as ``_call_endpoint`` -- the schema + declares :PyExternal's merge label as PySymbol, and RowBuilder MERGEs by + ``(labels[0], value)``, so if a call to that same bare name is ever + projected too, both rows collapse onto this one node — correctly, since + they name the same real-world symbol.""" + return b.node( + ["PySymbol", "PyExternal"], "id", f"{app_can_id}/@external/{name}", {"name": name} + ) + + +def _project_artifacts(b: RowBuilder, app: PyApplication, app_name: str, app_ref: NodeRef) -> None: + """Non-code artifacts, declared dependencies and undeclared imports (Tasks + 1-5) -- neutral ``Artifact``/``Package`` nodes with no ``Py`` prefix + (deliberate: cross-language merge targets, unlike everything else this + module projects). Always emitted regardless of ``-a`` -- this section is + L1 data, identical at every analysis level (mirrors ``analysis.json``).""" + app_can_id = application_id(app_name) + + for path in sorted(app.artifacts or {}): + art = app.artifacts[path] + art_ref = b.node( + ["Artifact"], + "id", + art.id, + prune( + { + "path": art.path, + "format": art.format, + "roles": art.roles, + "size_bytes": art.size_bytes, + "sha256": art.sha256, + "source": art.source, + "text_truncated": art.text_truncated, + "extraction": art.extraction, + } + ), + ) + b.edge("HAS_ARTIFACT", app_ref, art_ref) + + # Every lock artifact present LOCKS every dependency it pinned. The pins + # from all lock files are already merged into one `locked_version` per + # dependency upstream (Task 5 `build_dependency_view`) -- there is no + # per-lock-file attribution to split on, so (like the JSON projection) a + # dependency locked with N lock artifacts present gets N LOCKS edges. + lock_ids = [ + app.artifacts[p].id + for p in sorted(app.artifacts or {}) + if p.rsplit("/", 1)[-1] in _LOCK_BASENAMES + ] + + # app.dependencies has one PyDependency per DECLARING MANIFEST, so a + # package declared in 2+ manifests (e.g. requirements.txt + + # requirements-dev.txt both listing "requests") walks this loop once per + # manifest. DECLARES_DEPENDENCY is correctly one row per declaration (its + # `from_ref` is the manifest, so those rows are already distinct) -- but + # LOCKS/PY_PROVIDES/PY_UNRESOLVED_IMPORT are per-PACKAGE facts, and + # RowBuilder.edge() is append-only (unlike node(), it does not MERGE-dedup) + # -- so without a guard they'd be emitted once per declaring manifest + # instead of once, violating GraphRows' documented deduped-bag contract. + seen: set = set() + + for d in app.dependencies or []: + pkg_id = purl_pypi(d.name) + pkg_ref = b.node(["Package"], "id", pkg_id, {"ecosystem": "pypi", "name": d.name}) + # kind-discriminated: the same manifest may declare one package twice + # under different kinds (e.g. requests in [project.dependencies] AND + # again under [project.optional-dependencies]) -- same endpoint pair, + # so a plain MERGE would collapse the two declarations into one row. + b.edge( + "DECLARES_DEPENDENCY", + NodeRef("Artifact", "id", d.declared_in), + pkg_ref, + prune({"spec": d.spec, "kind": d.kind, "extras": d.extras, "prov": d.prov, "direct": d.direct}), + key=d.kind, + ) + if d.locked_version: + for lock_id in lock_ids: + key = ("LOCKS", lock_id, pkg_id) + if key not in seen: + seen.add(key) + b.edge( + "LOCKS", + NodeRef("Artifact", "id", lock_id), + pkg_ref, + {"version": d.locked_version}, + ) + for top in d.provides_imports: + ghost_ref = _import_ghost(b, app_can_id, top) + key = ("PY_PROVIDES", pkg_id, ghost_ref.value) + if key not in seen: + seen.add(key) + b.edge("PY_PROVIDES", pkg_ref, ghost_ref) + + for u in app.unresolved_imports or []: + ghost_ref = _import_ghost(b, app_can_id, u.module) + key = ("PY_UNRESOLVED_IMPORT", app_ref.value, ghost_ref.value) + if key not in seen: + seen.add(key) + b.edge("PY_UNRESOLVED_IMPORT", app_ref, ghost_ref, prune({"prov": u.prov})) + + def _sym(can_id: str) -> NodeRef: return NodeRef("PySymbol", "id", can_id) diff --git a/codeanalyzer/neo4j/schema.py b/codeanalyzer/neo4j/schema.py index d4ef95c..c45f134 100644 --- a/codeanalyzer/neo4j/schema.py +++ b/codeanalyzer/neo4j/schema.py @@ -201,6 +201,20 @@ class RelType: "_module": "string", }, ), + # Neutral artifact/dependency subgraph (spec 2026-08-27, Task 6). No `Py` + # prefix -- deliberate: `Artifact`/`Package` are cross-language merge + # targets, so a sibling-language analyzer over the same repo lands on the + # same nodes instead of a per-language duplicate. `PY_PROVIDES` / + # `PY_UNRESOLVED_IMPORT` stay PY_-namespaced (this analyzer's own claim + # about what an import resolves to) and target `:PyExternal`. + NodeLabel("Artifact", "Artifact", "id", { + "id": "string", "path": "string", "format": "string", + "roles": "string[]", "size_bytes": "integer", "sha256": "string", + "source": "string", "text_truncated": "boolean", "extraction": "string", + }), + NodeLabel("Package", "Package", "id", { + "id": "string", "ecosystem": "string", "name": "string", + }), ] _DECL_TARGETS = ["PyClass", "PyCallable"] @@ -250,6 +264,19 @@ class RelType: RelType("PY_PARAM_IN", ["PyBodyNode"], ["PyBodyNode"], {"var": "string"}), RelType("PY_PARAM_OUT", ["PyBodyNode"], ["PyBodyNode"], {"var": "string"}), RelType("PY_SUMMARY", ["PyBodyNode"], ["PyBodyNode"]), + # Neutral artifact/dependency subgraph (Task 6). + RelType("HAS_ARTIFACT", ["PyApplication"], ["Artifact"]), + # ``_k`` (merges per ``kind``): the same manifest may declare one package + # twice under different kinds (e.g. a runtime dep re-listed under an + # optional extra) -- same endpoint pair, so without the discriminant the + # plain MERGE collapses the two declarations into one row. + RelType("DECLARES_DEPENDENCY", ["Artifact"], ["Package"], { + "spec": "string", "kind": "string", "extras": "string[]", "prov": "string[]", + "direct": "boolean", "_k": "string", + }), + RelType("LOCKS", ["Artifact"], ["Package"], {"version": "string"}), + RelType("PY_PROVIDES", ["Package"], ["PyExternal"]), + RelType("PY_UNRESOLVED_IMPORT", ["PyApplication"], ["PyExternal"], {"prov": "string[]"}), ] diff --git a/codeanalyzer/options/options.py b/codeanalyzer/options/options.py index b7830d0..0b0014a 100644 --- a/codeanalyzer/options/options.py +++ b/codeanalyzer/options/options.py @@ -37,8 +37,13 @@ class AnalysisOptions: rebuild_analysis: bool = False skip_tests: bool = True no_venv: bool = False + resolve_installed: bool = False file_name: Optional[Path] = None cache_dir: Optional[Path] = None clear_cache: bool = False verbosity: int = 0 entrypoint_rules: Tuple[Path, ...] = () + # Artifact text-capture controls (#157 follow-up): whether to capture + # `source` at all, and the per-file byte cap before it truncates. + artifact_text: bool = True + artifact_text_max_bytes: int = 262144 diff --git a/codeanalyzer/schema/ids.py b/codeanalyzer/schema/ids.py index 75bc706..09626ff 100644 --- a/codeanalyzer/schema/ids.py +++ b/codeanalyzer/schema/ids.py @@ -21,3 +21,17 @@ def callable_sig_segment(name: str, param_names: List[str]) -> str: def ordinal_id(callable_id: str, tag: str) -> str: return f"{callable_id}@{tag}" + + +def artifact_id(app_name: str, rel_path: str) -> str: + """Language-neutral artifact id: ``can://artifact//``. + + The first segment is a namespace (a language for code nodes, the literal + ``artifact`` for files), so sibling analyzers over the same repo emit the + same id for the same file.""" + return f"can://artifact/{app_name}/{rel_path}" + + +def purl_pypi(name: str) -> str: + """Package URL for a (PEP 503 normalized) PyPI distribution name.""" + return f"pkg:pypi/{name}" diff --git a/codeanalyzer/schema/py_schema.py b/codeanalyzer/schema/py_schema.py index 18cdd19..26db9d7 100644 --- a/codeanalyzer/schema/py_schema.py +++ b/codeanalyzer/schema/py_schema.py @@ -470,6 +470,52 @@ class PyExternalSymbol(BaseModel): module: Optional[str] = None # best-effort owning module, e.g. "requests" +@builder +class PyArtifact(BaseModel): + """Any non-`.py` project file (config, manifest, CI, container spec, or + plain data/binary) -- never dropped from the walk. Captured broadly (node + + verbatim ``source``); *meaning* is extracted narrowly -- only + ``dependency-manifest`` roles feed ``dependencies`` today. ``id`` is + language-neutral (``can://artifact//``).""" + + id: str = "" + kind: str = "artifact" + path: str # repo-relative POSIX path (also the map key) + format: str # toml|yaml|json|ini|requirements|dockerfile|text|binary + roles: List[str] = [] + size_bytes: int = 0 + sha256: str = "" # always the full file's hash, even when source is truncated/empty + source: str = "" # verbatim by default; "" for binary or when capture is disabled + text_truncated: bool = False # True when `source` is a prefix, not the full file + extraction: str = "none" # none|partial|full + + +@builder +class PyDependency(BaseModel): + """One declared third-party dependency, evidence-tagged via ``prov``.""" + + name: str # PEP 503 normalized + spec: str = "" + kind: str = "runtime" # runtime|dev|optional|build + extras: List[str] = [] + declared_in: str = "" # PyArtifact id + # False for lockfile-only (transitive) dependencies -- pinned in a lock + # with no manifest declaration (#152 reconciliation). + direct: bool = True + locked_version: Optional[str] = None + provides_imports: List[str] = [] + prov: List[str] = [] # declared|lockfile|installed-metadata|heuristic + + +@builder +class PyImportBinding(BaseModel): + """A top-level import no declared dependency accounts for.""" + + module: str + bound_to: Optional[str] = None # best-effort distribution name + prov: List[str] = [] + + @builder class PyRepositoryInfo(BaseModel): """Where the analyzed source came from: git provenance captured at analysis time.""" @@ -502,6 +548,11 @@ class PyApplication(BaseModel): # builtin members), keyed by signature. Populated by the analyzer so every # backend (JSON and Neo4j) shares one authoritative external-symbol set. external_symbols: Dict[str, PyExternalSymbol] = {} + # Non-code artifacts, declared dependencies, and undeclared imports + # (spec 2026-08-27). L1 data: identical at every analysis level. + artifacts: Dict[str, PyArtifact] = {} + dependencies: List[PyDependency] = [] + unresolved_imports: List[PyImportBinding] = [] # Coverage/failure record for the entrypoint pass; see PyEntrypointReport (#27). entrypoint_report: PyEntrypointReport = PyEntrypointReport() # Git provenance of the analyzed checkout, captured at analysis time. diff --git a/docs/design/plans/2026-08-27-artifacts-and-dependencies.md b/docs/design/plans/2026-08-27-artifacts-and-dependencies.md new file mode 100644 index 0000000..20c0faa --- /dev/null +++ b/docs/design/plans/2026-08-27-artifacts-and-dependencies.md @@ -0,0 +1,1305 @@ +# Artifacts and Dependencies Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Emit non-code artifacts, evidence-tagged dependency records, and unresolved imports in both schema v2 projections (closes #157). + +**Architecture:** A new `codeanalyzer/artifacts/` package scans the project once (after the symbol table, at every level): `discovery.py` turns rule-matched files into `PyArtifact` nodes, `parsers.py` reads dependency manifests into raw records, `dependencies.py` builds `PyDependency`/`PyImportBinding` lists. `core.py` attaches all three to `PyApplication`; `neo4j/` projects them as language-neutral `:Artifact`/`:Package` nodes joined to the existing `:PyExternal` ghosts. + +**Tech Stack:** pydantic models (v1/v2 compat helpers in `codeanalyzer/schema/__init__.py`), stdlib `tomllib` (`tomli` backport below 3.11), `yaml` (already a dependency via entrypoints), stdlib `ast`/`configparser`/`hashlib`. + +**Spec:** `docs/design/specs/2026-08-27-artifacts-and-dependencies-design.md` + +## Global Constraints + +- Never add AI/Claude attribution anywhere (commits, code, docs). Conventional Commits. +- Artifact ids are language-neutral: `can://artifact//`. +- `source` is verbatim and unbounded; artifacts are text-only (rule-matched formats). +- Dependency `prov` vocabulary exactly: `declared`, `lockfile`, `installed-metadata`, `heuristic`. +- Default run reads only repo files (byte-identical across machines); venv probing only behind `--resolve-installed`. +- Lock files never create records — they only backfill `locked_version`. +- All three new sections are L1 data: identical at every `-a`; never gated by level. +- `setup.py` is parsed by static AST only — never executed, never imported. +- Sorted iteration everywhere (walks, dict builds) — determinism by construction. +- Pydantic v1 must keep working: no `model_dump`/`model_copy` calls on models outside the compat helpers. +- New runtime dependency allowed: `tomli>=2.0; python_version < '3.11'` only. + +--- + +### Task 1: Schema models and ids + +**Files:** +- Modify: `codeanalyzer/schema/ids.py` (append) +- Modify: `codeanalyzer/schema/py_schema.py` (new models near `PyExternalSymbol`; three new fields on `PyApplication`) +- Test: `test/test_artifact_models.py` (create) + +**Interfaces:** +- Produces: `artifact_id(app_name: str, rel_path: str) -> str`; `purl_pypi(name: str) -> str`; models `PyArtifact`, `PyDependency`, `PyImportBinding`; `PyApplication.artifacts: Dict[str, PyArtifact]`, `.dependencies: List[PyDependency]`, `.unresolved_imports: List[PyImportBinding]`. + +- [ ] **Step 1: Write the failing test** + +```python +# test/test_artifact_models.py +"""Task 1: artifact/dependency schema models and id constructors.""" +from codeanalyzer.schema import model_validate_json, model_dump_json +from codeanalyzer.schema.ids import artifact_id, purl_pypi +from codeanalyzer.schema.py_schema import ( + PyApplication, PyArtifact, PyDependency, PyImportBinding, +) + + +def test_artifact_id_is_language_neutral(): + assert artifact_id("myapp", "deploy/docker-compose.yml") == \ + "can://artifact/myapp/deploy/docker-compose.yml" + + +def test_purl_pypi(): + assert purl_pypi("pyyaml") == "pkg:pypi/pyyaml" + + +def test_models_round_trip(): + art = PyArtifact( + id=artifact_id("a", "pyproject.toml"), path="pyproject.toml", + format="toml", roles=["dependency-manifest"], size_bytes=10, + sha256="ab" * 32, source="[project]\n", + ) + dep = PyDependency( + name="requests", spec=">=2.31", kind="runtime", + declared_in=art.id, provides_imports=["requests"], prov=["declared"], + ) + imp = PyImportBinding(module="yaml", bound_to="pyyaml", prov=["heuristic"]) + app = PyApplication.builder().symbol_table({}).call_graph([]).build() + app.artifacts = {art.path: art} + app.dependencies = [dep] + app.unresolved_imports = [imp] + back = model_validate_json(PyApplication, model_dump_json(app)) + assert back.artifacts["pyproject.toml"].kind == "artifact" + assert back.artifacts["pyproject.toml"].extraction == "none" + assert back.dependencies[0].locked_version is None + assert back.unresolved_imports[0].bound_to == "pyyaml" + + +def test_defaults_empty_on_old_payload(): + app = PyApplication.builder().symbol_table({}).call_graph([]).build() + back = model_validate_json(PyApplication, model_dump_json(app)) + assert back.artifacts == {} and back.dependencies == [] and back.unresolved_imports == [] +``` + +- [ ] **Step 2: Run test to verify it fails** + +Run: `uv run pytest test/test_artifact_models.py -v` +Expected: FAIL with `ImportError` (`artifact_id` / `PyArtifact` not defined). + +- [ ] **Step 3: Implement** + +Append to `codeanalyzer/schema/ids.py`: + +```python +def artifact_id(app_name: str, rel_path: str) -> str: + """Language-neutral artifact id: ``can://artifact//``. + + The first segment is a namespace (a language for code nodes, the literal + ``artifact`` for files), so sibling analyzers over the same repo emit the + same id for the same file.""" + return f"can://artifact/{app_name}/{rel_path}" + + +def purl_pypi(name: str) -> str: + """Package URL for a (PEP 503 normalized) PyPI distribution name.""" + return f"pkg:pypi/{name}" +``` + +In `codeanalyzer/schema/py_schema.py`, after `PyExternalSymbol` (follow its style; use `@builder` like neighbors): + +```python +@builder +class PyArtifact(BaseModel): + """A recognized non-code file (config, manifest, CI, container spec). + + Captured broadly (node + verbatim ``source``); *meaning* is extracted + narrowly — only ``dependency-manifest`` roles feed ``dependencies`` today. + ``id`` is language-neutral (``can://artifact//``).""" + + id: str = "" + kind: str = "artifact" + path: str # repo-relative POSIX path (also the map key) + format: str # toml|yaml|json|ini|requirements|dockerfile|text + roles: List[str] = [] + size_bytes: int = 0 + sha256: str = "" + source: str = "" # verbatim, unbounded by decision (spec §3) + extraction: str = "none" # none|partial|full + + +@builder +class PyDependency(BaseModel): + """One declared third-party dependency, evidence-tagged via ``prov``.""" + + name: str # PEP 503 normalized + spec: str = "" + kind: str = "runtime" # runtime|dev|optional|build + extras: List[str] = [] + declared_in: str = "" # PyArtifact id + locked_version: Optional[str] = None + provides_imports: List[str] = [] + prov: List[str] = [] # declared|lockfile|installed-metadata|heuristic + + +@builder +class PyImportBinding(BaseModel): + """A top-level import no declared dependency accounts for.""" + + module: str + bound_to: Optional[str] = None # best-effort distribution name + prov: List[str] = [] +``` + +On `PyApplication`, after `external_symbols`: + +```python + # Non-code artifacts, declared dependencies, and undeclared imports + # (spec 2026-08-27). L1 data: identical at every analysis level. + artifacts: Dict[str, PyArtifact] = {} + dependencies: List[PyDependency] = [] + unresolved_imports: List[PyImportBinding] = [] +``` + +- [ ] **Step 4: Run tests** + +Run: `uv run pytest test/test_artifact_models.py test/test_v2_keystone.py -v` +Expected: PASS (keystone suite proves no regression to the envelope). + +- [ ] **Step 5: Commit** + +```bash +git add codeanalyzer/schema/ids.py codeanalyzer/schema/py_schema.py test/test_artifact_models.py +git commit -m "feat(schema): PyArtifact/PyDependency/PyImportBinding models and can://artifact ids" +``` + +--- + +### Task 2: Artifact discovery + +**Files:** +- Create: `codeanalyzer/artifacts/__init__.py`, `codeanalyzer/artifacts/discovery.py` +- Test: `test/test_artifact_discovery.py` (create) + +**Interfaces:** +- Consumes: `artifact_id`, `PyArtifact` (Task 1). +- Produces: `discover_artifacts(project_dir: Path, app_name: str) -> Dict[str, PyArtifact]` (sorted keys, repo-relative POSIX paths). `RULES: List[Tuple[str, str, List[str]]]` (glob pattern, format, roles). + +- [ ] **Step 1: Write the failing test** + +```python +# test/test_artifact_discovery.py +"""Task 2: rule-matched files become PyArtifact nodes; nothing else does.""" +import hashlib +from pathlib import Path +from codeanalyzer.artifacts.discovery import discover_artifacts + + +def _mk(tmp_path: Path, rel: str, text: str = "x: 1\n") -> Path: + p = tmp_path / rel + p.parent.mkdir(parents=True, exist_ok=True) + p.write_text(text) + return p + + +def test_discovers_known_shapes(tmp_path): + _mk(tmp_path, "pyproject.toml", "[project]\nname='a'\n") + _mk(tmp_path, "requirements-dev.txt", "pytest\n") + _mk(tmp_path, "deploy/docker-compose.yml") + _mk(tmp_path, "Dockerfile", "FROM python:3.12\n") + _mk(tmp_path, ".github/workflows/ci.yml") + _mk(tmp_path, "src/app.py", "x = 1\n") # code: never an artifact + _mk(tmp_path, "notes.md", "hi\n") # unmatched: no node + arts = discover_artifacts(tmp_path, "myapp") + assert sorted(arts) == [ + ".github/workflows/ci.yml", "Dockerfile", "deploy/docker-compose.yml", + "pyproject.toml", "requirements-dev.txt", + ] + py = arts["pyproject.toml"] + assert py.id == "can://artifact/myapp/pyproject.toml" + assert py.format == "toml" and "dependency-manifest" in py.roles + assert arts["Dockerfile"].roles == ["container-image"] + assert arts["deploy/docker-compose.yml"].roles == ["service-topology"] + assert arts[".github/workflows/ci.yml"].roles == ["ci"] + + +def test_source_hash_and_ignores(tmp_path): + _mk(tmp_path, "pyproject.toml", "content-here\n") + _mk(tmp_path, ".venv/pyvenv.cfg", "home = /x\n") + _mk(tmp_path, ".git/config", "[core]\n") + _mk(tmp_path, "node_modules/a/package.json", "{}") + arts = discover_artifacts(tmp_path, "a") + assert list(arts) == ["pyproject.toml"] + a = arts["pyproject.toml"] + assert a.source == "content-here\n" + assert a.sha256 == hashlib.sha256(b"content-here\n").hexdigest() + assert a.size_bytes == len(b"content-here\n") + + +def test_unreadable_binary_is_skipped(tmp_path): + (tmp_path / "settings.json").write_bytes(b"\xff\xfe\x00bad") + arts = discover_artifacts(tmp_path, "a") + assert arts == {} +``` + +- [ ] **Step 2: Run to verify it fails** + +Run: `uv run pytest test/test_artifact_discovery.py -v` +Expected: FAIL with `ModuleNotFoundError: codeanalyzer.artifacts`. + +- [ ] **Step 3: Implement** + +`codeanalyzer/artifacts/__init__.py`: + +```python +"""Non-code artifact capture and dependency extraction (spec 2026-08-27). + +Capture is broad (every rule-matched config-shaped file becomes a +:class:`~codeanalyzer.schema.py_schema.PyArtifact`); extraction is narrow +(only dependency manifests are parsed for meaning in this unit).""" + +from codeanalyzer.artifacts.discovery import discover_artifacts + +__all__ = ["discover_artifacts"] +``` + +`codeanalyzer/artifacts/discovery.py`: + +```python +import fnmatch +import hashlib +from pathlib import Path +from typing import Dict, List, Tuple + +from codeanalyzer.schema.ids import artifact_id +from codeanalyzer.schema.py_schema import PyArtifact + +# (glob pattern against the repo-relative POSIX path, format, roles). +# First match wins; patterns are checked in order. +RULES: List[Tuple[str, str, List[str]]] = [ + ("requirements*.txt", "requirements", ["dependency-manifest"]), + ("**/requirements*.txt", "requirements", ["dependency-manifest"]), + ("pyproject.toml", "toml", ["dependency-manifest", "tool-config"]), + ("setup.py", "text", ["dependency-manifest"]), + ("setup.cfg", "ini", ["dependency-manifest", "tool-config"]), + ("Pipfile", "toml", ["dependency-manifest"]), + ("Pipfile.lock", "json", ["dependency-manifest"]), + ("poetry.lock", "toml", ["dependency-manifest"]), + ("uv.lock", "toml", ["dependency-manifest"]), + ("environment.yml", "yaml", ["dependency-manifest"]), + ("environment.yaml", "yaml", ["dependency-manifest"]), + ("Dockerfile", "dockerfile", ["container-image"]), + ("**/Dockerfile", "dockerfile", ["container-image"]), + ("*.dockerfile", "dockerfile", ["container-image"]), + ("docker-compose*.yml", "yaml", ["service-topology"]), + ("docker-compose*.yaml", "yaml", ["service-topology"]), + ("**/docker-compose*.yml", "yaml", ["service-topology"]), + ("**/docker-compose*.yaml", "yaml", ["service-topology"]), + ("compose.yml", "yaml", ["service-topology"]), + ("compose.yaml", "yaml", ["service-topology"]), + ("k8s/**/*.yml", "yaml", ["service-topology"]), + ("k8s/**/*.yaml", "yaml", ["service-topology"]), + ("**/Chart.yaml", "yaml", ["service-topology"]), + ("**/values.yaml", "yaml", ["service-topology"]), + (".github/workflows/*.yml", "yaml", ["ci"]), + (".github/workflows/*.yaml", "yaml", ["ci"]), + (".gitlab-ci.yml", "yaml", ["ci"]), + (".env", "text", ["env"]), + (".env.*", "text", ["env"]), + ("tox.ini", "ini", ["tool-config"]), + ("noxfile.py", "text", ["tool-config"]), + ("Makefile", "text", ["tool-config"]), + ("*.cfg", "ini", ["unknown"]), + ("*.toml", "toml", ["unknown"]), +] + +_IGNORED_DIRS = { + ".git", ".hg", ".svn", "__pycache__", ".venv", "venv", ".tox", ".nox", + "node_modules", ".mypy_cache", ".pytest_cache", ".ruff_cache", ".idea", + "build", "dist", ".eggs", +} + + +def _classify(rel_posix: str) -> Tuple[str, List[str]] | None: + name = rel_posix.rsplit("/", 1)[-1] + for pattern, fmt, roles in RULES: + target = rel_posix if ("/" in pattern or pattern.startswith("**")) else name + if fnmatch.fnmatch(target, pattern): + return fmt, roles + return None + + +def discover_artifacts(project_dir: Path, app_name: str) -> Dict[str, PyArtifact]: + """Walk the project and return rule-matched files as artifacts, sorted by path.""" + out: Dict[str, PyArtifact] = {} + for path in sorted(project_dir.rglob("*")): + if not path.is_file(): + continue + rel = path.relative_to(project_dir) + if any(part in _IGNORED_DIRS for part in rel.parts): + continue + rel_posix = rel.as_posix() + hit = _classify(rel_posix) + if hit is None: + continue + fmt, roles = hit + raw = path.read_bytes() + try: + text = raw.decode("utf-8") + except UnicodeDecodeError: + continue # text-only by spec; binaries never become artifacts + out[rel_posix] = PyArtifact( + id=artifact_id(app_name, rel_posix), path=rel_posix, format=fmt, + roles=list(roles), size_bytes=len(raw), + sha256=hashlib.sha256(raw).hexdigest(), source=text, + ) + return out +``` + +- [ ] **Step 4: Run tests** — `uv run pytest test/test_artifact_discovery.py -v` — PASS. + +- [ ] **Step 5: Commit** + +```bash +git add codeanalyzer/artifacts test/test_artifact_discovery.py +git commit -m "feat(artifacts): rule-table discovery walk producing PyArtifact nodes" +``` + +--- + +### Task 3: Manifest parsers + +**Files:** +- Create: `codeanalyzer/artifacts/parsers.py` +- Modify: `pyproject.toml` (add `"tomli>=2.0; python_version < '3.11'"` to `dependencies`) +- Test: `test/test_manifest_parsers.py` (create) + +**Interfaces:** +- Produces: `RawDep` dataclass `(name, spec, kind, extras)` (name PEP 503 normalized); `normalize_name(raw: str) -> str`; `parse_manifest(fmt_path: str, text: str) -> Tuple[List[RawDep], bool]` returning `(deps, partial)` and dispatching on basename: requirements/pyproject/setup.py/setup.cfg/Pipfile/environment.yml; `parse_lock_pins(basename: str, text: str) -> Dict[str, str]` for poetry.lock, uv.lock, Pipfile.lock; `parse_requirement_line(line: str) -> RawDep | None` (exported for reuse). + +- [ ] **Step 1: Write the failing test** + +```python +# test/test_manifest_parsers.py +"""Task 3: every spec §6 manifest format parses into RawDep records.""" +import textwrap +from codeanalyzer.artifacts.parsers import ( + RawDep, normalize_name, parse_lock_pins, parse_manifest, +) + + +def test_normalize_name(): + assert normalize_name("PyYAML") == "pyyaml" + assert normalize_name("ruamel.yaml") == "ruamel-yaml" + assert normalize_name("typing_extensions") == "typing-extensions" + + +def test_requirements_txt(): + text = textwrap.dedent("""\ + # comment + requests>=2.31,<3 + pyyaml + celery[redis]==5.3.* + -e ./local-pkg + --index-url https://example.invalid + """) + deps, partial = parse_manifest("requirements.txt", text) + assert not partial + assert [(d.name, d.spec) for d in deps] == [ + ("requests", ">=2.31,<3"), ("pyyaml", ""), ("celery", "==5.3.*"), + ] + assert deps[2].extras == ["redis"] + assert all(d.kind == "runtime" for d in deps) + + +def test_requirements_dev_kind(): + deps, _ = parse_manifest("requirements-dev.txt", "pytest\n") + assert deps[0].kind == "dev" + + +def test_pyproject_pep621_poetry_and_build(): + text = textwrap.dedent("""\ + [build-system] + requires = ["setuptools>=68"] + [project] + dependencies = ["requests>=2.31"] + [project.optional-dependencies] + docs = ["sphinx"] + [tool.poetry.dependencies] + python = "^3.10" + rich = "^13.0" + [tool.poetry.group.dev.dependencies] + mypy = "*" + """) + deps, partial = parse_manifest("pyproject.toml", text) + assert not partial + by = {(d.name, d.kind) for d in deps} + assert ("setuptools", "build") in by + assert ("requests", "runtime") in by + assert ("sphinx", "optional") in by + assert ("rich", "runtime") in by and ("mypy", "dev") in by + assert ("python", "runtime") not in {(d.name, d.kind) for d in deps} # interpreter, not a dep + + +def test_setup_py_static_literals(): + text = 'from setuptools import setup\nsetup(install_requires=["flask>=2"], extras_require={"test": ["pytest"]})\n' + deps, partial = parse_manifest("setup.py", text) + assert not partial + assert {(d.name, d.kind) for d in deps} == {("flask", "runtime"), ("pytest", "optional")} + + +def test_setup_py_dynamic_is_partial(): + text = "from setuptools import setup\nreqs = compute()\nsetup(install_requires=reqs)\n" + deps, partial = parse_manifest("setup.py", text) + assert partial and deps == [] + + +def test_setup_cfg(): + text = "[options]\ninstall_requires =\n numpy>=1.24\n pandas\n" + deps, _ = parse_manifest("setup.cfg", text) + assert [(d.name, d.spec) for d in deps] == [("numpy", ">=1.24"), ("pandas", "")] + + +def test_pipfile_and_environment_yml(): + pip = '[packages]\nrequests = ">=2.31"\n[dev-packages]\nblack = "*"\n' + deps, _ = parse_manifest("Pipfile", pip) + assert {(d.name, d.kind, d.spec) for d in deps} == { + ("requests", "runtime", ">=2.31"), ("black", "dev", ""), + } + env = "dependencies:\n - numpy=1.26\n - pip\n - pip:\n - fastapi>=0.100\n" + deps, _ = parse_manifest("environment.yml", env) + assert {(d.name, d.spec) for d in deps} == {("numpy", "=1.26"), ("fastapi", ">=0.100")} + + +def test_lock_pins(): + poetry = '[[package]]\nname = "requests"\nversion = "2.31.0"\n[[package]]\nname = "PyYAML"\nversion = "6.0.1"\n' + assert parse_lock_pins("poetry.lock", poetry) == {"requests": "2.31.0", "pyyaml": "6.0.1"} + uv = '[[package]]\nname = "requests"\nversion = "2.32.0"\n' + assert parse_lock_pins("uv.lock", uv) == {"requests": "2.32.0"} + pipf = '{"default": {"requests": {"version": "==2.31.0"}}, "develop": {}}' + assert parse_lock_pins("Pipfile.lock", pipf) == {"requests": "2.31.0"} +``` + +- [ ] **Step 2: Run to verify it fails** — `uv run pytest test/test_manifest_parsers.py -v` — FAIL (`No module named codeanalyzer.artifacts.parsers`). + +- [ ] **Step 3: Implement `codeanalyzer/artifacts/parsers.py`** + +```python +"""Dependency-manifest readers. Pure text-in/records-out; no execution, no I/O.""" + +import ast +import configparser +import json +import re +import sys +from dataclasses import dataclass, field +from typing import Dict, List, Optional, Tuple + +if sys.version_info >= (3, 11): + import tomllib +else: # pragma: no cover - exercised on the 3.10 CI leg + import tomli as tomllib + +import yaml + + +@dataclass(frozen=True) +class RawDep: + name: str # PEP 503 normalized + spec: str = "" + kind: str = "runtime" # runtime|dev|optional|build + extras: Tuple[str, ...] = () + + +def normalize_name(raw: str) -> str: + return re.sub(r"[-_.]+", "-", raw).lower() + + +_REQ_LINE = re.compile( + r"^\s*(?P[A-Za-z0-9][A-Za-z0-9._-]*)\s*(?:\[(?P[^\]]+)\])?\s*(?P[^;#]*)" +) + + +def parse_requirement_line(line: str, kind: str = "runtime") -> Optional[RawDep]: + """One PEP 508-ish requirement line -> RawDep (None for options/paths/URLs).""" + line = line.split("#", 1)[0].strip() + if not line or line.startswith(("-", "--")) or "://" in line or line.startswith((".", "/")): + return None + m = _REQ_LINE.match(line) + if not m: + return None + extras = tuple(e.strip() for e in (m.group("extras") or "").split(",") if e.strip()) + return RawDep(normalize_name(m.group("name")), m.group("spec").strip().rstrip(","), kind, extras) + + +def _kind_for_requirements(basename: str) -> str: + return "dev" if re.search(r"(dev|test|lint|doc)", basename, re.I) else "runtime" + + +def _parse_requirements(basename: str, text: str) -> List[RawDep]: + kind = _kind_for_requirements(basename) + out = [] + for line in text.splitlines(): + dep = parse_requirement_line(line, kind) + if dep: + out.append(dep) + return out + + +def _parse_pyproject(text: str) -> List[RawDep]: + data = tomllib.loads(text) + out: List[RawDep] = [] + for req in (data.get("build-system") or {}).get("requires", []): + d = parse_requirement_line(req, "build") + if d: + out.append(d) + proj = data.get("project") or {} + for req in proj.get("dependencies", []): + d = parse_requirement_line(req) + if d: + out.append(d) + for group in (proj.get("optional-dependencies") or {}).values(): + for req in group: + d = parse_requirement_line(req, "optional") + if d: + out.append(d) + poetry = ((data.get("tool") or {}).get("poetry")) or {} + for name, spec in (poetry.get("dependencies") or {}).items(): + if normalize_name(name) == "python": + continue + out.append(RawDep(normalize_name(name), spec if isinstance(spec, str) else "", "runtime")) + for gname, group in (poetry.get("group") or {}).items(): + kind = "dev" if gname == "dev" else "optional" + for name, spec in (group.get("dependencies") or {}).items(): + out.append(RawDep(normalize_name(name), spec if isinstance(spec, str) else "", kind)) + for name, spec in (poetry.get("dev-dependencies") or {}).items(): # legacy poetry + out.append(RawDep(normalize_name(name), spec if isinstance(spec, str) else "", "dev")) + return out + + +def _parse_setup_py(text: str) -> Tuple[List[RawDep], bool]: + """Static AST only. Literal lists lift; anything computed -> partial=True.""" + try: + tree = ast.parse(text) + except SyntaxError: + return [], True + out: List[RawDep] = [] + partial = False + for node in ast.walk(tree): + if not (isinstance(node, ast.Call) and getattr(node.func, "id", getattr(node.func, "attr", "")) == "setup"): + continue + for kw in node.keywords: + if kw.arg == "install_requires": + lifted = _lift_str_list(kw.value) + if lifted is None: + partial = True + else: + out += [d for d in (parse_requirement_line(s) for s in lifted) if d] + elif kw.arg == "extras_require": + if not isinstance(kw.value, ast.Dict): + partial = True + continue + for v in kw.value.values: + lifted = _lift_str_list(v) + if lifted is None: + partial = True + else: + out += [d for d in (parse_requirement_line(s, "optional") for s in lifted) if d] + return out, partial + + +def _lift_str_list(node: ast.AST) -> Optional[List[str]]: + if isinstance(node, (ast.List, ast.Tuple)) and all( + isinstance(e, ast.Constant) and isinstance(e.value, str) for e in node.elts + ): + return [e.value for e in node.elts] + return None + + +def _parse_setup_cfg(text: str) -> List[RawDep]: + cp = configparser.ConfigParser() + cp.read_string(text) + out: List[RawDep] = [] + if cp.has_option("options", "install_requires"): + for line in cp.get("options", "install_requires").splitlines(): + d = parse_requirement_line(line) + if d: + out.append(d) + if cp.has_section("options.extras_require"): + for _, val in cp.items("options.extras_require"): + for line in val.splitlines(): + d = parse_requirement_line(line, "optional") + if d: + out.append(d) + return out + + +def _parse_pipfile(text: str) -> List[RawDep]: + data = tomllib.loads(text) + out: List[RawDep] = [] + for section, kind in (("packages", "runtime"), ("dev-packages", "dev")): + for name, spec in (data.get(section) or {}).items(): + s = spec if isinstance(spec, str) else (spec.get("version", "") if isinstance(spec, dict) else "") + out.append(RawDep(normalize_name(name), "" if s == "*" else s, kind)) + return out + + +def _parse_environment_yml(text: str) -> List[RawDep]: + data = yaml.safe_load(text) or {} + out: List[RawDep] = [] + for item in data.get("dependencies") or []: + if isinstance(item, str): + name, _, spec = item.partition("=") + if normalize_name(name) in ("pip", "python"): + continue + out.append(RawDep(normalize_name(name), f"={spec}" if spec else "")) + elif isinstance(item, dict): + for req in item.get("pip") or []: + d = parse_requirement_line(req) + if d: + out.append(d) + return out + + +def parse_manifest(path: str, text: str) -> Tuple[List[RawDep], bool]: + """Dispatch on basename -> (records, partial). Unknown basenames -> ([], False).""" + base = path.rsplit("/", 1)[-1] + try: + if base.startswith("requirements") and base.endswith(".txt"): + return _parse_requirements(base, text), False + if base == "pyproject.toml": + return _parse_pyproject(text), False + if base == "setup.py": + return _parse_setup_py(text) + if base == "setup.cfg": + return _parse_setup_cfg(text), False + if base == "Pipfile": + return _parse_pipfile(text), False + if base in ("environment.yml", "environment.yaml"): + return _parse_environment_yml(text), False + except Exception: + return [], True # unparseable manifest: keep the artifact, flag extraction + return [], False + + +def parse_lock_pins(path: str, text: str) -> Dict[str, str]: + """Lock file -> {normalized name: pinned version}. Never creates records.""" + base = path.rsplit("/", 1)[-1] + try: + if base in ("poetry.lock", "uv.lock"): + data = tomllib.loads(text) + return { + normalize_name(p["name"]): str(p["version"]) + for p in data.get("package") or [] if "name" in p and "version" in p + } + if base == "Pipfile.lock": + data = json.loads(text) + out = {} + for section in ("default", "develop"): + for name, meta in (data.get(section) or {}).items(): + v = (meta or {}).get("version", "") + out[normalize_name(name)] = v.lstrip("=") + return out + except Exception: + return {} + return {} +``` + +Add to `pyproject.toml` `[project] dependencies`: `"tomli>=2.0; python_version < '3.11'",` + +- [ ] **Step 4: Run tests** — `uv run pytest test/test_manifest_parsers.py -v` — PASS. + +- [ ] **Step 5: Commit** + +```bash +git add codeanalyzer/artifacts/parsers.py test/test_manifest_parsers.py pyproject.toml uv.lock +git commit -m "feat(artifacts): manifest parsers for all spec formats (static, no execution)" +``` + +(Run `uv lock` before committing so `uv.lock` picks up the tomli marker dep.) + +--- + +### Task 4: Dependency view — records, lock backfill, import binding + +**Files:** +- Create: `codeanalyzer/artifacts/dependencies.py` +- Test: `test/test_dependency_view.py` (create) + +**Interfaces:** +- Consumes: Tasks 1–3 (`PyArtifact`, `RawDep`, `parse_manifest`, `parse_lock_pins`, `normalize_name`, `purl_pypi`). +- Produces: `build_dependency_view(artifacts: Dict[str, PyArtifact], modules: Dict[str, "PyModule"], venv_dir: Optional[Path], resolve_installed: bool) -> Tuple[List[PyDependency], List[PyImportBinding]]`. Mutates `artifacts` only to set `extraction` (`"full"`/`"partial"` on manifests). Dependencies sorted by `(name, declared_in)`; unresolved sorted by `module`. + +- [ ] **Step 1: Write the failing test** + +```python +# test/test_dependency_view.py +"""Task 4: declared records, lock backfill, provides_imports, unresolved imports.""" +from pathlib import Path +from codeanalyzer.artifacts.discovery import discover_artifacts +from codeanalyzer.artifacts.dependencies import build_dependency_view +from codeanalyzer.schema.py_schema import PyImport, PyModule + + +def _module(name, imports): + return PyModule.builder().file_path(f"/tmp/{name}.py").module_name(name).imports( + [PyImport(module=m, name="*") for m in imports] + ).build() + + +def _setup(tmp_path): + (tmp_path / "pyproject.toml").write_text( + '[project]\ndependencies = ["requests>=2.31", "PyYAML"]\n' + ) + (tmp_path / "uv.lock").write_text( + '[[package]]\nname = "requests"\nversion = "2.32.3"\n' + ) + arts = discover_artifacts(tmp_path, "app") + mods = { + "app.py": _module("app", ["requests", "yaml", "colorama", "os", "app.util"]), + "app/util.py": _module("app.util", []), + } + return arts, mods + + +def test_records_lock_and_binding(tmp_path): + arts, mods = _setup(tmp_path) + deps, unresolved = build_dependency_view(arts, mods, None, False) + by = {d.name: d for d in deps} + assert by["requests"].prov == ["declared", "lockfile"] + assert by["requests"].locked_version == "2.32.3" + assert by["requests"].declared_in == "can://artifact/app/pyproject.toml" + assert by["pyyaml"].locked_version is None and by["pyyaml"].prov == ["declared"] + # provides_imports: requests trivially; pyyaml has no same-name import -> [] + assert by["requests"].provides_imports == ["requests"] + assert by["pyyaml"].provides_imports == [] + assert arts["pyproject.toml"].extraction == "full" + + +def test_unresolved_imports(tmp_path): + arts, mods = _setup(tmp_path) + _, unresolved = build_dependency_view(arts, mods, None, False) + u = {b.module: b for b in unresolved} + # yaml: known-alias heuristic binds it to declared pyyaml -> NOT unresolved + # colorama: imported, never declared -> unresolved, unbound + # os: stdlib; app.util: local module -> neither appears + assert set(u) == {"colorama"} + assert u["colorama"].bound_to is None and u["colorama"].prov == [] + + +def test_known_alias_binding(tmp_path): + arts, mods = _setup(tmp_path) + deps, unresolved = build_dependency_view(arts, mods, None, False) + yaml_dep = next(d for d in deps if d.name == "pyyaml") + assert "yaml" in yaml_dep.provides_imports and "heuristic" in yaml_dep.prov + + +def test_installed_metadata_binding(tmp_path): + arts, mods = _setup(tmp_path) + venv = tmp_path / ".venv" + di = venv / "lib" / "python3.12" / "site-packages" / "PyYAML-6.0.1.dist-info" + di.mkdir(parents=True) + (di / "METADATA").write_text("Metadata-Version: 2.1\nName: PyYAML\nVersion: 6.0.1\n") + (di / "top_level.txt").write_text("yaml\n_yaml\n") + deps, _ = build_dependency_view(arts, mods, venv, True) + yaml_dep = next(d for d in deps if d.name == "pyyaml") + assert "yaml" in yaml_dep.provides_imports + assert "installed-metadata" in yaml_dep.prov +``` + +- [ ] **Step 2: Run to verify it fails** — `uv run pytest test/test_dependency_view.py -v` — FAIL (module missing). + +- [ ] **Step 3: Implement `codeanalyzer/artifacts/dependencies.py`** + +```python +"""PyDependency / PyImportBinding construction from discovered artifacts. + +Deterministic by default: reads only repo files. ``resolve_installed`` adds +filesystem reads of ``/**/site-packages/*.dist-info`` (never runs an +interpreter), tagged ``prov: installed-metadata``.""" + +import re +import sys +from pathlib import Path +from typing import Dict, List, Optional, Tuple + +from codeanalyzer.artifacts.parsers import ( + RawDep, normalize_name, parse_lock_pins, parse_manifest, +) +from codeanalyzer.schema.py_schema import ( + PyArtifact, PyDependency, PyImport, PyImportBinding, PyModule, +) + +_LOCK_BASENAMES = ("poetry.lock", "uv.lock", "Pipfile.lock") + +# Small, non-exhaustive alias table for the worst offenders; everything else +# rides the same-name rule or --resolve-installed. prov: heuristic. +_KNOWN_IMPORT_ALIASES: Dict[str, str] = { + "pyyaml": "yaml", "beautifulsoup4": "bs4", "pillow": "PIL", + "scikit-learn": "sklearn", "opencv-python": "cv2", "python-dateutil": "dateutil", + "msgpack-python": "msgpack", "protobuf": "google.protobuf", + "setuptools": "setuptools", "attrs": "attr", "pymongo": "pymongo", +} + + +def _stdlib_names() -> set: + return set(getattr(sys, "stdlib_module_names", ())) | {"__future__"} + + +def _installed_top_levels(venv_dir: Optional[Path]) -> Dict[str, List[str]]: + """{normalized dist name: [top-level import names]} from *.dist-info files.""" + out: Dict[str, List[str]] = {} + if venv_dir is None or not venv_dir.exists(): + return out + for di in sorted(venv_dir.glob("lib/python*/site-packages/*.dist-info")): + name = None + meta = di / "METADATA" + if meta.exists(): + m = re.search(r"^Name:\s*(.+)$", meta.read_text(errors="replace"), re.M) + if m: + name = normalize_name(m.group(1).strip()) + if name is None: + continue + tl = di / "top_level.txt" + if tl.exists(): + out[name] = [l.strip() for l in tl.read_text().splitlines() if l.strip()] + return out + + +def build_dependency_view( + artifacts: Dict[str, PyArtifact], + modules: Dict[str, PyModule], + venv_dir: Optional[Path], + resolve_installed: bool, +) -> Tuple[List[PyDependency], List[PyImportBinding]]: + # 1. Declared records from every dependency-manifest artifact (non-lock). + deps: List[PyDependency] = [] + for path in sorted(artifacts): + art = artifacts[path] + if "dependency-manifest" not in art.roles: + continue + if path.rsplit("/", 1)[-1] in _LOCK_BASENAMES: + continue + raw, partial = parse_manifest(path, art.source) + art.extraction = "partial" if partial else "full" + for r in raw: + deps.append(PyDependency( + name=r.name, spec=r.spec, kind=r.kind, extras=sorted(r.extras), + declared_in=art.id, prov=["declared"], + )) + + # 2. Lock backfill (locked_version + prov "lockfile"); locks never create records. + pins: Dict[str, str] = {} + for path in sorted(artifacts): + if path.rsplit("/", 1)[-1] in _LOCK_BASENAMES: + pins.update(parse_lock_pins(path, artifacts[path].source)) + artifacts[path].extraction = "full" + for d in deps: + if d.name in pins: + d.locked_version = pins[d.name] + d.prov = sorted(set(d.prov) | {"lockfile"}) + + # 3. Import universe from the symbol table (top-level segments only). + local = {m.module_name.split(".")[0] for m in modules.values() if m.module_name} + stdlib = _stdlib_names() + imported: set = set() + for m in modules.values(): + for imp in m.imports or []: + top = (imp.module or imp.name or "").split(".")[0] + if top and top not in stdlib and top not in local: + imported.add(top) + + # 4. provides_imports: same-name rule, alias table, optional installed metadata. + installed = _installed_top_levels(venv_dir) if resolve_installed else {} + for d in deps: + provides: List[str] = [] + same = d.name.replace("-", "_") + for candidate in {d.name, same}: + if candidate in imported: + provides.append(candidate) + alias = _KNOWN_IMPORT_ALIASES.get(d.name) + if alias and alias.split(".")[0] in imported: + provides.append(alias) + d.prov = sorted(set(d.prov) | {"heuristic"}) + if d.name in installed: + for top in installed[d.name]: + if top in imported and top not in provides: + provides.append(top) + d.prov = sorted(set(d.prov) | {"installed-metadata"}) + d.provides_imports = sorted(set(provides)) + + # 5. Unresolved: imported, not stdlib/local, not provided by any dependency. + provided = {p for d in deps for p in d.provides_imports} + unresolved = [ + PyImportBinding(module=m) for m in sorted(imported - provided) + ] + deps.sort(key=lambda d: (d.name, d.declared_in)) + return deps, unresolved +``` + +Also export from `codeanalyzer/artifacts/__init__.py`: add `from codeanalyzer.artifacts.dependencies import build_dependency_view` and extend `__all__`. + +- [ ] **Step 4: Run tests** — `uv run pytest test/test_dependency_view.py test/test_manifest_parsers.py -v` — PASS. + +- [ ] **Step 5: Commit** + +```bash +git add codeanalyzer/artifacts test/test_dependency_view.py +git commit -m "feat(artifacts): dependency records with lock backfill and import binding" +``` + +--- + +### Task 5: Core wiring + CLI flag + +**Files:** +- Modify: `codeanalyzer/core.py` (in `analyze`, directly after the `detect_entrypoints(app, self.project_dir, self.options.entrypoint_rules)` call) +- Modify: `codeanalyzer/options/options.py` (`AnalysisOptions`: add `resolve_installed: bool = False`) +- Modify: `codeanalyzer/__main__.py` (add the flag; keep option count/order consistent with neighbors) +- Test: `test/test_artifact_pipeline.py` (create) + +**Interfaces:** +- Consumes: `discover_artifacts`, `build_dependency_view` (Tasks 2/4). +- Produces: populated `app.artifacts` / `app.dependencies` / `app.unresolved_imports` at every level; CLI `--resolve-installed` (default off). + +- [ ] **Step 1: Write the failing test** + +```python +# test/test_artifact_pipeline.py +"""Task 5: sections populated at every level, identically (L1-data posture).""" +import json +from pathlib import Path +from codeanalyzer.core import Codeanalyzer +from codeanalyzer.options import AnalysisOptions +from codeanalyzer.schema import model_dump + + +def _run(tmp_path, project, level): + out = tmp_path / f"out{level}" + opts = AnalysisOptions( + input=project, output=out, analysis_level=level, + no_venv=True, cache_dir=tmp_path / f"cache{level}", + ) + artifacts = Codeanalyzer(opts).analyze() + return artifacts.application + + +def _fixture(tmp_path) -> Path: + proj = tmp_path / "proj" + proj.mkdir() + (proj / "pyproject.toml").write_text('[project]\ndependencies = ["requests"]\n') + (proj / "Dockerfile").write_text("FROM python:3.12\n") + (proj / "app.py").write_text("import requests\nimport colorama\n") + return proj + + +def test_sections_identical_across_levels(tmp_path): + proj = _fixture(tmp_path) + a1 = _run(tmp_path, proj, 1) + a2 = _run(tmp_path, proj, 2) + d1 = json.loads(model_dump_json_app(a1)) + d2 = json.loads(model_dump_json_app(a2)) + for field in ("artifacts", "dependencies", "unresolved_imports"): + assert d1[field] == d2[field] + assert sorted(d1["artifacts"]) == ["Dockerfile", "pyproject.toml"] + assert [d["name"] for d in d1["dependencies"]] == ["requests"] + assert [u["module"] for u in d1["unresolved_imports"]] == ["colorama"] + + +def model_dump_json_app(app): + from codeanalyzer.schema import model_dump_json + return model_dump_json(app) + + +def test_resolve_installed_flag_default_off(): + from codeanalyzer.options import AnalysisOptions + assert AnalysisOptions.__fields__ if hasattr(AnalysisOptions, "__fields__") else True + assert AnalysisOptions(input=Path(".")).resolve_installed is False +``` + +- [ ] **Step 2: Run to verify it fails** — `uv run pytest test/test_artifact_pipeline.py -v` — FAIL (`resolve_installed` unknown / sections empty). + +- [ ] **Step 3: Implement** + +`options.py`: add `resolve_installed: bool = False` beside `no_venv`. + +`core.py`, after the `detect_entrypoints(...)` call (comment style matches neighbors): + +```python + # Artifacts + dependencies: L1 data, every level, never varies with -a + # (spec 2026-08-27). Deterministic by default; venv probing is opt-in. + from codeanalyzer.artifacts import build_dependency_view, discover_artifacts + + app.artifacts = discover_artifacts(self.project_dir, self.app_name) + app.dependencies, app.unresolved_imports = build_dependency_view( + app.artifacts, + app.symbol_table, + self.virtualenv if self.options.resolve_installed else None, + self.options.resolve_installed, + ) +``` + +(`self.app_name` and `self.virtualenv` already exist on `Codeanalyzer`; check their exact attribute names in `__init__` and use those.) + +`__main__.py`: add beside `--no-venv`: + +```python + resolve_installed: bool = typer.Option( + False, "--resolve-installed", + help="Additionally bind imports via the project venv's installed metadata " + "(*.dist-info); output becomes machine-dependent (prov: installed-metadata).", + ), +``` + +and pass `resolve_installed=resolve_installed` into `AnalysisOptions(...)`. + +- [ ] **Step 4: Run tests** — `uv run pytest test/test_artifact_pipeline.py test/test_cli.py -v` — PASS (CLI test asserts existing flags unaffected). Then `uv run python scripts/update_readme.py` to refresh the README `--help` block. + +- [ ] **Step 5: Commit** + +```bash +git add codeanalyzer/core.py codeanalyzer/options/options.py codeanalyzer/__main__.py README.md test/test_artifact_pipeline.py +git commit -m "feat(cli): emit artifacts/dependencies at every level; add --resolve-installed" +``` + +--- + +### Task 6: Neo4j projection + +**Files:** +- Modify: `codeanalyzer/neo4j/schema.py` (catalog: 2 node labels, 5 rel types, 2 constraints) +- Modify: `codeanalyzer/neo4j/project.py` (new `_project_artifacts(...)` called from `project(...)`) +- Modify: `schema.neo4j.json` (regenerate) +- Test: `test/test_neo4j_artifacts.py` (create) + +**Interfaces:** +- Consumes: populated app sections (Task 5); `purl_pypi` (Task 1); existing `RowBuilder`, external-ghost id scheme. +- Produces: catalog entries — `NodeLabel("Artifact", "Artifact", "id", {...})`, `NodeLabel("Package", "Package", "id", {"id": "string", "ecosystem": "string", "name": "string"})`; rels `HAS_ARTIFACT` (PyApplication→Artifact), `DECLARES_DEPENDENCY` (Artifact→Package, props `{spec, kind, extras: "string[]", prov: "string[]"}`), `LOCKS` (Artifact→Package, `{version}`), `PY_PROVIDES` (Package→PyExternal), `PY_UNRESOLVED_IMPORT` (PyApplication→PyExternal, `{prov: "string[]"}`). + +- [ ] **Step 1: Write the failing test** + +```python +# test/test_neo4j_artifacts.py +"""Task 6: artifact/dependency rows in the Neo4j projection.""" +from pathlib import Path +from codeanalyzer.neo4j.schema import NODE_LABELS, REL_TYPES + + +def test_catalog_has_neutral_vocabulary(): + labels = {n.label: n for n in NODE_LABELS} + assert labels["Artifact"].key == "id" and labels["Package"].key == "id" + rels = {r.type for r in REL_TYPES} + assert {"HAS_ARTIFACT", "DECLARES_DEPENDENCY", "LOCKS", + "PY_PROVIDES", "PY_UNRESOLVED_IMPORT"} <= rels + + +def test_rows_projected(tmp_path): + proj = tmp_path / "p" + proj.mkdir() + (proj / "pyproject.toml").write_text('[project]\ndependencies = ["requests"]\n') + (proj / "app.py").write_text("import requests\nimport colorama\nrequests.get('u')\n") + from codeanalyzer.core import Codeanalyzer + from codeanalyzer.options import AnalysisOptions + app = Codeanalyzer(AnalysisOptions( + input=proj, analysis_level=2, no_venv=True, cache_dir=tmp_path / "c", + )).analyze().application + from codeanalyzer.neo4j.project import project + from codeanalyzer.schema.assign_ids import build_sig_to_id + rows = project(app, "p", build_sig_to_id(app), full_depth=True) + nodes = {(r.label, r.key_value) for r in rows.nodes} + assert ("Artifact", "can://artifact/p/pyproject.toml") in nodes + assert ("Package", "pkg:pypi/requests") in nodes + rel_types = {r.type for r in rows.rels} + assert {"HAS_ARTIFACT", "DECLARES_DEPENDENCY", "PY_PROVIDES"} <= rel_types +``` + +(Adapt the two helper imports — `project(...)` signature and the sig-to-id builder name — to what `codeanalyzer/neo4j/project.py` and `codeanalyzer/schema/assign_ids.py` actually export; the existing `test/test_neo4j_*.py` files show the working invocation to copy. Row objects: use the same accessors those tests use.) + +- [ ] **Step 2: Run to verify it fails** — `uv run pytest test/test_neo4j_artifacts.py -v` — FAIL (labels missing). + +- [ ] **Step 3: Implement** + +`schema.py` — append to `NODE_LABELS`: + +```python + NodeLabel("Artifact", "Artifact", "id", { + "id": "string", "path": "string", "format": "string", + "roles": "string[]", "size_bytes": "integer", "sha256": "string", + "source": "string", "extraction": "string", + }), + NodeLabel("Package", "Package", "id", { + "id": "string", "ecosystem": "string", "name": "string", + }), +``` + +Append to `REL_TYPES`: + +```python + RelType("HAS_ARTIFACT", ["PyApplication"], ["Artifact"]), + RelType("DECLARES_DEPENDENCY", ["Artifact"], ["Package"], { + "spec": "string", "kind": "string", "extras": "string[]", "prov": "string[]", + }), + RelType("LOCKS", ["Artifact"], ["Package"], {"version": "string"}), + RelType("PY_PROVIDES", ["Package"], ["PyExternal"]), + RelType("PY_UNRESOLVED_IMPORT", ["PyApplication"], ["PyExternal"], {"prov": "string[]"}), +``` + +`project.py` — new step called at the end of `project(...)`, next to `_project_program_graphs`: + +```python +def _project_artifacts(b, app, app_name: str) -> None: + """Neutral artifact/package subgraph (spec 2026-08-27). Package prov of an + import binding lands on the edge; PY_PROVIDES targets the same @external + ghost ids the call graph merges on, so dependencies join it.""" + from codeanalyzer.schema.ids import purl_pypi + + app_key = app_name # PyApplication merges on name + for path in sorted(app.artifacts or {}): + art = app.artifacts[path] + b.node("Artifact", art.id, { + "path": art.path, "format": art.format, "roles": art.roles, + "size_bytes": art.size_bytes, "sha256": art.sha256, + "source": art.source, "extraction": art.extraction, + }) + b.rel("PyApplication", app_key, "HAS_ARTIFACT", "Artifact", art.id) + seen_pkgs = set() + for d in app.dependencies or []: + pkg = purl_pypi(d.name) + if pkg not in seen_pkgs: + b.node("Package", pkg, {"ecosystem": "pypi", "name": d.name}) + seen_pkgs.add(pkg) + b.rel("Artifact", d.declared_in, "DECLARES_DEPENDENCY", "Package", pkg, { + "spec": d.spec, "kind": d.kind, "extras": d.extras, "prov": d.prov, + }) + if d.locked_version: + # every lock artifact that pinned it: point from each lock file present + for lpath in sorted(app.artifacts or {}): + if lpath.rsplit("/", 1)[-1] in ("poetry.lock", "uv.lock", "Pipfile.lock"): + b.rel("Artifact", app.artifacts[lpath].id, "LOCKS", "Package", pkg, + {"version": d.locked_version}) + for top in d.provides_imports: + ghost = f"can://python/{app_name}/@external/{top}" + b.merge_external(ghost, module=top) + b.rel("Package", pkg, "PY_PROVIDES", "PyExternal", ghost) + for u in app.unresolved_imports or []: + ghost = f"can://python/{app_name}/@external/{u.module}" + b.merge_external(ghost, module=u.module) + b.rel("PyApplication", app_key, "PY_UNRESOLVED_IMPORT", "PyExternal", ghost, + {"prov": u.prov}) +``` + +**Adapt the `b.node`/`b.rel`/ghost-merge calls to the actual `RowBuilder` API** in `codeanalyzer/neo4j/rows.py` (read it first; `_project_program_graphs` shows the working idiom, and external ghosts are merged in `project(...)` — reuse that helper rather than inventing `merge_external` if one already exists; match the exact `can://…/@external/…` id shape used there, which may include the member name segment). + +Regenerate the contract: `uv run canpy --emit schema > schema.neo4j.json`. + +- [ ] **Step 4: Run tests** — `uv run pytest test/test_neo4j_artifacts.py test/ -k "neo4j" -v` — PASS (schema-conformance test must accept the regenerated contract). + +- [ ] **Step 5: Commit** + +```bash +git add codeanalyzer/neo4j/schema.py codeanalyzer/neo4j/project.py schema.neo4j.json test/test_neo4j_artifacts.py +git commit -m "feat(neo4j): neutral Artifact/Package subgraph joined to external ghosts" +``` + +--- + +### Task 7: End-to-end fixture, determinism, docs + +**Files:** +- Create: `test/fixtures/whole_applications/manifests_app/` (pyproject.toml, requirements-dev.txt, setup.py with a computed `install_requires`, uv.lock, environment.yml, Dockerfile, docker-compose.yml, `.github/workflows/ci.yml`, `pkg/__init__.py`, `pkg/main.py` importing one declared + one undeclared package) +- Test: `test/test_artifacts_end_to_end.py` (create) +- Modify: `CHANGELOG.md` (Unreleased → Added), `README.md` (Output shape section: three new sections, one paragraph; cookbook: drop the "Landing with #157" marker) + +**Interfaces:** +- Consumes: everything above. + +- [ ] **Step 1: Create the fixture** — files exactly: + +`pkg/main.py`: + +```python +import requests +import colorama # deliberately undeclared + +def fetch(url): + return requests.get(url) +``` + +`pyproject.toml`: `[project]\nname = "manifests-app"\ndependencies = ["requests>=2.31"]\n` +`requirements-dev.txt`: `pytest>=8\n` +`setup.py`: `from setuptools import setup\nextra = compute_extras()\nsetup(install_requires=extra)\n` +`uv.lock`: `[[package]]\nname = "requests"\nversion = "2.32.3"\n` +`environment.yml`: `dependencies:\n - numpy=1.26\n` +`Dockerfile`: `FROM python:3.12-slim\n` +`docker-compose.yml`: `services:\n web:\n build: .\n` +`.github/workflows/ci.yml`: `on: push\njobs: {}\n` +`pkg/__init__.py`: empty. + +- [ ] **Step 2: Write the failing test** + +```python +# test/test_artifacts_end_to_end.py +"""Task 7: full pipeline over the manifests_app fixture + determinism.""" +import json +from pathlib import Path +from codeanalyzer.core import Codeanalyzer +from codeanalyzer.options import AnalysisOptions +from codeanalyzer.schema import model_dump_json + +FIXTURE = Path(__file__).parent / "fixtures" / "whole_applications" / "manifests_app" + + +def _app(tmp_path, tag): + return Codeanalyzer(AnalysisOptions( + input=FIXTURE, analysis_level=1, no_venv=True, cache_dir=tmp_path / tag, + )).analyze().application + + +def test_full_surface(tmp_path): + app = _app(tmp_path, "a") + arts = app.artifacts + assert {"pyproject.toml", "requirements-dev.txt", "setup.py", "uv.lock", + "environment.yml", "Dockerfile", "docker-compose.yml", + ".github/workflows/ci.yml"} <= set(arts) + assert arts["setup.py"].extraction == "partial" # computed install_requires + assert arts["pyproject.toml"].extraction == "full" + deps = {d.name: d for d in app.dependencies} + assert deps["requests"].locked_version == "2.32.3" + assert deps["requests"].prov == ["declared", "lockfile"] + assert deps["pytest"].kind == "dev" + assert deps["numpy"].spec == "=1.26" + assert [u.module for u in app.unresolved_imports] == ["colorama"] + + +def test_determinism_two_runs(tmp_path): + a = model_dump_json(_app(tmp_path, "r1")) + b = model_dump_json(_app(tmp_path, "r2")) + assert a == b +``` + +- [ ] **Step 3: Run to verify current state** — first run FAILs only if earlier tasks missed something; otherwise both PASS immediately (this task's value is the fixture + gate). + +- [ ] **Step 4: Full suite + docs** + +Run: `uv run pytest test/ -q` — all green. +`CHANGELOG.md` under a new `## [Unreleased]` → `### Added`: + +```markdown +- Schema v2 now captures non-code artifacts (`application.artifacts`), declared + dependencies with provenance (`application.dependencies`), and undeclared + imports (`application.unresolved_imports`) at every analysis level (#157). + Neo4j gains language-neutral `:Artifact`/`:Package` nodes (purl ids) joined + to the existing `:PyExternal` ghosts. New flag: `--resolve-installed`. +``` + +README: in **Output shape**, add a short paragraph naming the three sections and the `can://artifact/` namespace; in the cookbook, delete the "Landing with #157 … not in the current release" sentence (queries now real). + +- [ ] **Step 5: Commit** + +```bash +git add test/fixtures/whole_applications/manifests_app test/test_artifacts_end_to_end.py CHANGELOG.md README.md +git commit -m "test(artifacts): manifests_app fixture, end-to-end + determinism gates; docs" +``` diff --git a/pyproject.toml b/pyproject.toml index 5127985..27d1859 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -47,6 +47,9 @@ dependencies = [ # pyyaml: the entrypoint rules pack (#27) is YAML. Declared explicitly -- # it was previously only reachable as a transitive dep of ray. "pyyaml>=6.0,<7.0", + # tomli: stdlib tomllib only ships from 3.11; artifacts/parsers.py needs + # TOML parsing (pyproject.toml, Pipfile, lock files) below that floor. + "tomli>=2.0; python_version < '3.11'", ] [project.optional-dependencies] diff --git a/schema.neo4j.json b/schema.neo4j.json index 5f9a353..ffda3f1 100644 --- a/schema.neo4j.json +++ b/schema.neo4j.json @@ -151,6 +151,32 @@ "end_line": "integer", "_module": "string" } + }, + { + "label": "Artifact", + "merge_label": "Artifact", + "key": "id", + "properties": { + "id": "string", + "path": "string", + "format": "string", + "roles": "string[]", + "size_bytes": "integer", + "sha256": "string", + "source": "string", + "text_truncated": "boolean", + "extraction": "string" + } + }, + { + "label": "Package", + "merge_label": "Package", + "key": "id", + "properties": { + "id": "string", + "ecosystem": "string", + "name": "string" + } } ], "relationship_types": [ @@ -354,6 +380,67 @@ "PyBodyNode" ], "properties": {} + }, + { + "type": "HAS_ARTIFACT", + "from": [ + "PyApplication" + ], + "to": [ + "Artifact" + ], + "properties": {} + }, + { + "type": "DECLARES_DEPENDENCY", + "from": [ + "Artifact" + ], + "to": [ + "Package" + ], + "properties": { + "spec": "string", + "kind": "string", + "extras": "string[]", + "prov": "string[]", + "direct": "boolean", + "_k": "string" + } + }, + { + "type": "LOCKS", + "from": [ + "Artifact" + ], + "to": [ + "Package" + ], + "properties": { + "version": "string" + } + }, + { + "type": "PY_PROVIDES", + "from": [ + "Package" + ], + "to": [ + "PyExternal" + ], + "properties": {} + }, + { + "type": "PY_UNRESOLVED_IMPORT", + "from": [ + "PyApplication" + ], + "to": [ + "PyExternal" + ], + "properties": { + "prov": "string[]" + } } ], "constraints": [ @@ -364,7 +451,9 @@ "CREATE CONSTRAINT pydecorator_name IF NOT EXISTS FOR (x:PyDecorator) REQUIRE x.name IS UNIQUE", "CREATE CONSTRAINT pyattribute_id IF NOT EXISTS FOR (x:PyAttribute) REQUIRE x.id IS UNIQUE", "CREATE CONSTRAINT pyvariable_id IF NOT EXISTS FOR (x:PyVariable) REQUIRE x.id IS UNIQUE", - "CREATE CONSTRAINT pybodynode_id IF NOT EXISTS FOR (x:PyBodyNode) REQUIRE x.id IS UNIQUE" + "CREATE CONSTRAINT pybodynode_id IF NOT EXISTS FOR (x:PyBodyNode) REQUIRE x.id IS UNIQUE", + "CREATE CONSTRAINT artifact_id IF NOT EXISTS FOR (x:Artifact) REQUIRE x.id IS UNIQUE", + "CREATE CONSTRAINT package_id IF NOT EXISTS FOR (x:Package) REQUIRE x.id IS UNIQUE" ], "indexes": [ "CREATE INDEX py_callable_name IF NOT EXISTS FOR (c:PyCallable) ON (c.name)", diff --git a/test/fixtures/whole_applications/manifests_app/.github/workflows/ci.yml b/test/fixtures/whole_applications/manifests_app/.github/workflows/ci.yml new file mode 100644 index 0000000..9aa4fbd --- /dev/null +++ b/test/fixtures/whole_applications/manifests_app/.github/workflows/ci.yml @@ -0,0 +1,2 @@ +on: push +jobs: {} diff --git a/test/fixtures/whole_applications/manifests_app/Dockerfile b/test/fixtures/whole_applications/manifests_app/Dockerfile new file mode 100644 index 0000000..bee3c16 --- /dev/null +++ b/test/fixtures/whole_applications/manifests_app/Dockerfile @@ -0,0 +1 @@ +FROM python:3.12-slim diff --git a/test/fixtures/whole_applications/manifests_app/data.csv b/test/fixtures/whole_applications/manifests_app/data.csv new file mode 100644 index 0000000..bed40b0 --- /dev/null +++ b/test/fixtures/whole_applications/manifests_app/data.csv @@ -0,0 +1,3 @@ +name,value +alpha,1 +beta,2 diff --git a/test/fixtures/whole_applications/manifests_app/docker-compose.yml b/test/fixtures/whole_applications/manifests_app/docker-compose.yml new file mode 100644 index 0000000..3dad310 --- /dev/null +++ b/test/fixtures/whole_applications/manifests_app/docker-compose.yml @@ -0,0 +1,3 @@ +services: + web: + build: . diff --git a/test/fixtures/whole_applications/manifests_app/environment.yml b/test/fixtures/whole_applications/manifests_app/environment.yml new file mode 100644 index 0000000..1b6fdb6 --- /dev/null +++ b/test/fixtures/whole_applications/manifests_app/environment.yml @@ -0,0 +1,2 @@ +dependencies: + - numpy=1.26 diff --git a/test/fixtures/whole_applications/manifests_app/logo.png b/test/fixtures/whole_applications/manifests_app/logo.png new file mode 100644 index 0000000..846d29a Binary files /dev/null and b/test/fixtures/whole_applications/manifests_app/logo.png differ diff --git a/test/fixtures/whole_applications/manifests_app/pkg/__init__.py b/test/fixtures/whole_applications/manifests_app/pkg/__init__.py new file mode 100644 index 0000000..e69de29 diff --git a/test/fixtures/whole_applications/manifests_app/pkg/main.py b/test/fixtures/whole_applications/manifests_app/pkg/main.py new file mode 100644 index 0000000..1656432 --- /dev/null +++ b/test/fixtures/whole_applications/manifests_app/pkg/main.py @@ -0,0 +1,5 @@ +import requests +import colorama # deliberately undeclared + +def fetch(url): + return requests.get(url) diff --git a/test/fixtures/whole_applications/manifests_app/pyproject.toml b/test/fixtures/whole_applications/manifests_app/pyproject.toml new file mode 100644 index 0000000..c610b67 --- /dev/null +++ b/test/fixtures/whole_applications/manifests_app/pyproject.toml @@ -0,0 +1,3 @@ +[project] +name = "manifests-app" +dependencies = ["requests>=2.31"] diff --git a/test/fixtures/whole_applications/manifests_app/requirements-dev.txt b/test/fixtures/whole_applications/manifests_app/requirements-dev.txt new file mode 100644 index 0000000..cc52e8c --- /dev/null +++ b/test/fixtures/whole_applications/manifests_app/requirements-dev.txt @@ -0,0 +1 @@ +pytest>=8 diff --git a/test/fixtures/whole_applications/manifests_app/setup.py b/test/fixtures/whole_applications/manifests_app/setup.py new file mode 100644 index 0000000..e7abd8c --- /dev/null +++ b/test/fixtures/whole_applications/manifests_app/setup.py @@ -0,0 +1,3 @@ +from setuptools import setup +extra = compute_extras() +setup(install_requires=extra) diff --git a/test/fixtures/whole_applications/manifests_app/uv.lock b/test/fixtures/whole_applications/manifests_app/uv.lock new file mode 100644 index 0000000..744029b --- /dev/null +++ b/test/fixtures/whole_applications/manifests_app/uv.lock @@ -0,0 +1,3 @@ +[[package]] +name = "requests" +version = "2.32.3" diff --git a/test/sample_graph_app.py b/test/sample_graph_app.py index a45b604..40eb511 100644 --- a/test/sample_graph_app.py +++ b/test/sample_graph_app.py @@ -5,7 +5,8 @@ graph with a resolved edge and a ghost edge, each callable's CPG ``body``/``cfg``/``cdg``/``ddg`` (level 3), and — new at level 4 — the interprocedural ``param_in``/``param_out``/``summary`` param-passing overlay plus -the points-to ``ddg`` delta. +the points-to ``ddg`` delta. Also carries the artifact/dependency subgraph +(Task 6): a manifest + lock artifact and a locked, import-providing dependency. The symbol table is built from a real (temporary) source file so ``build_function_pdgs`` can recover each callable's AST; ``assign_ids`` + @@ -37,9 +38,12 @@ from codeanalyzer.dataflow.syntactic import SyntacticOracle from codeanalyzer.schema import PyApplication, PyExternalSymbol from codeanalyzer.schema.assign_ids import assign_ids +from codeanalyzer.schema.ids import artifact_id from codeanalyzer.schema.l1_body import populate_l1_body from codeanalyzer.schema.l2_callees import backfill_callees -from codeanalyzer.schema.py_schema import PyCallEdge +from codeanalyzer.schema.py_schema import ( + PyArtifact, PyCallEdge, PyDependency, PyImportBinding, +) from codeanalyzer.semantic_analysis.call_graph import ( iter_callables_in_symbol_table, ) @@ -143,6 +147,38 @@ def make_sample_app() -> Tuple[PyApplication, Dict[str, str]]: ), ] + # Artifact/dependency subgraph (Task 6): a manifest + a lock artifact, one + # locked runtime dependency it declares (provides an import), and one + # import nothing declared -- exercises every Artifact/Package label and + # HAS_ARTIFACT/DECLARES_DEPENDENCY/LOCKS/PY_PROVIDES/PY_UNRESOLVED_IMPORT + # relationship in the catalog (guarded by + # test_all_catalog_node_kinds_and_relationships_are_exercised). + manifest_id = artifact_id("sample-app", "pyproject.toml") + lock_id = artifact_id("sample-app", "poetry.lock") + app.artifacts = { + "pyproject.toml": PyArtifact( + id=manifest_id, path="pyproject.toml", format="toml", + roles=["dependency-manifest", "tool-config"], size_bytes=42, + sha256="a" * 64, source="[project]\ndependencies = [\"acme\"]\n", + extraction="full", + ), + "poetry.lock": PyArtifact( + id=lock_id, path="poetry.lock", format="toml", + roles=["dependency-manifest"], size_bytes=7, sha256="b" * 64, + source="", extraction="full", + ), + } + app.dependencies = [ + PyDependency( + name="acme", spec=">=1.0", kind="runtime", extras=[], + declared_in=manifest_id, locked_version="1.2.3", + provides_imports=["acme"], prov=["declared", "lockfile"], + ), + ] + app.unresolved_imports = [ + PyImportBinding(module="colorama", prov=["heuristic"]), + ] + # Identity + L1 bodies, then the intraprocedural (syntactic) L3 overlay. sig_to_id = assign_ids(app, "sample-app") ext_id = "can://python/sample-app/@external/os/getcwd" diff --git a/test/test_artifact_discovery.py b/test/test_artifact_discovery.py new file mode 100644 index 0000000..50938f4 --- /dev/null +++ b/test/test_artifact_discovery.py @@ -0,0 +1,245 @@ +"""Task 2: every file becomes a PyArtifact node except `.py` files and ignored +dirs (never-drop inventory, issue #157 follow-up). Rule-matched files keep +their format/roles; unmatched files fall back to text/unknown (or binary). +Also covers the text-capture controls (`capture_text`/`text_max_bytes`).""" +import hashlib +from pathlib import Path +from codeanalyzer.schema import model_dump +from codeanalyzer.artifacts.discovery import discover_artifacts + + +def _mk(tmp_path: Path, rel: str, text: str = "x: 1\n") -> Path: + p = tmp_path / rel + p.parent.mkdir(parents=True, exist_ok=True) + p.write_text(text) + return p + + +def test_discovers_known_shapes(tmp_path): + _mk(tmp_path, "pyproject.toml", "[project]\nname='a'\n") + _mk(tmp_path, "requirements-dev.txt", "pytest\n") + _mk(tmp_path, "deploy/docker-compose.yml") + _mk(tmp_path, "Dockerfile", "FROM python:3.12\n") + _mk(tmp_path, "svc/Dockerfile", "FROM alpine:latest\n") + _mk(tmp_path, ".github/workflows/ci.yml") + _mk(tmp_path, "k8s/deploy.yaml") + _mk(tmp_path, "src/app.py", "x = 1\n") # code: never an artifact + _mk(tmp_path, "notes.md", "hi\n") # *.md -> docs + _mk(tmp_path, "data.bin", "hi\n") # unmatched, decodable: text/unknown + arts = discover_artifacts(tmp_path, "myapp") + assert sorted(arts) == [ + ".github/workflows/ci.yml", "Dockerfile", "data.bin", "deploy/docker-compose.yml", + "k8s/deploy.yaml", "notes.md", "pyproject.toml", "requirements-dev.txt", + "svc/Dockerfile", + ] + assert arts["notes.md"].roles == ["docs"] + py = arts["pyproject.toml"] + assert py.id == "can://artifact/myapp/pyproject.toml" + assert py.format == "toml" and "dependency-manifest" in py.roles + assert arts["Dockerfile"].roles == ["container-image"] + assert arts["svc/Dockerfile"].roles == ["container-image"] + assert arts["deploy/docker-compose.yml"].roles == ["service-topology"] + assert arts["k8s/deploy.yaml"].roles == ["service-topology"] + assert arts[".github/workflows/ci.yml"].roles == ["ci"] + assert arts["data.bin"].format == "text" and arts["data.bin"].roles == ["unknown"] + + +def test_source_hash_and_ignores(tmp_path): + _mk(tmp_path, "pyproject.toml", "content-here\n") + _mk(tmp_path, ".venv/pyvenv.cfg", "home = /x\n") + _mk(tmp_path, ".git/config", "[core]\n") + _mk(tmp_path, "node_modules/a/package.json", "{}") + arts = discover_artifacts(tmp_path, "a") + assert list(arts) == ["pyproject.toml"] + a = arts["pyproject.toml"] + assert a.source == "content-here\n" + assert a.sha256 == hashlib.sha256(b"content-here\n").hexdigest() + assert a.size_bytes == len(b"content-here\n") + + +def test_matched_non_utf8_file_becomes_binary_artifact(tmp_path): + """Rule-matched (pyproject.toml) but not UTF-8 decodable: never dropped -- + format downgrades to binary, the rule's roles survive, source is empty.""" + raw = b"\xff\xfe\x00bad" + (tmp_path / "pyproject.toml").write_bytes(raw) + arts = discover_artifacts(tmp_path, "a") + assert list(arts) == ["pyproject.toml"] + a = arts["pyproject.toml"] + assert a.format == "binary" + assert a.roles == ["dependency-manifest", "tool-config"] + assert a.source == "" + assert a.sha256 == hashlib.sha256(raw).hexdigest() + assert a.size_bytes == len(raw) + + +def test_ignores_own_codeanalyzer_venv(tmp_path): + """Regression: a default (non --no-venv) run provisions its own analysis + virtualenv under /.codeanalyzer//virtualenv/ before + discovery runs (Codeanalyzer.__enter__). Without an ignore, discovery + walks straight into it and picks up pyvenv.cfg (machine-specific `home=` + -> breaks determinism) and site-packages *.toml files.""" + _mk(tmp_path, ".codeanalyzer/x/virtualenv/pyvenv.cfg", "home = /machine/specific\n") + _mk(tmp_path, "pyproject.toml", "[project]\nname='a'\n") + arts = discover_artifacts(tmp_path, "a") + assert list(arts) == ["pyproject.toml"] + + +def test_discovers_kind_yaml(tmp_path): + """kind/*.yml|yaml -> service-topology, alongside the existing k8s/ rules + (user-surfaced miss on a real corpus).""" + _mk(tmp_path, "kind/cluster.yml") + arts = discover_artifacts(tmp_path, "a") + assert list(arts) == ["kind/cluster.yml"] + assert arts["kind/cluster.yml"].roles == ["service-topology"] + + +def test_discovers_packaging_docs_legal_files(tmp_path): + """Round-2 role vocabulary growth: packaging/docs/legal.""" + _mk(tmp_path, "MANIFEST.in", "include *.txt\n") + _mk(tmp_path, "LICENSE", "MIT\n") + _mk(tmp_path, "LICENSE.md", "MIT\n") # legal-prefixed rule wins over *.md + _mk(tmp_path, "COPYRIGHT.txt", "(c) 2026\n") + _mk(tmp_path, "NOTICE", "third-party notices\n") + _mk(tmp_path, "CONTRIBUTING.rst", "how to contribute\n") + arts = discover_artifacts(tmp_path, "a") + assert arts["MANIFEST.in"].roles == ["packaging"] + assert arts["LICENSE"].roles == ["legal"] + assert arts["LICENSE.md"].roles == ["legal"] + assert arts["COPYRIGHT.txt"].roles == ["legal"] + assert arts["NOTICE"].roles == ["legal"] + assert arts["CONTRIBUTING.rst"].roles == ["docs"] + + +def test_extensionless_shebang_script_is_captured(tmp_path): + """odoo-bin-style entrypoint: no extension, so no RULES glob can name it -- + the shebang fallback refines its roles to ["script"].""" + _mk(tmp_path, "odoo-bin", "#!/usr/bin/env python3\nimport sys\n") + arts = discover_artifacts(tmp_path, "a") + assert list(arts) == ["odoo-bin"] + assert arts["odoo-bin"].format == "text" and arts["odoo-bin"].roles == ["script"] + + +def test_extensionless_binary_captured_as_binary_artifact(tmp_path): + raw = b"\xff\xfe\x00#!bad" + (tmp_path / "odoo-bin").write_bytes(raw) + arts = discover_artifacts(tmp_path, "a") + assert list(arts) == ["odoo-bin"] + a = arts["odoo-bin"] + assert a.format == "binary" and a.roles == ["unknown"] and a.source == "" + assert a.sha256 == hashlib.sha256(raw).hexdigest() + + +def test_extensionless_text_without_shebang_captured_as_unknown(tmp_path): + _mk(tmp_path, "README", "just some notes, no shebang\n") + arts = discover_artifacts(tmp_path, "a") + assert list(arts) == ["README"] + assert arts["README"].format == "text" and arts["README"].roles == ["unknown"] + assert arts["README"].source == "just some notes, no shebang\n" + + +def test_unmatched_text_file_captured_as_unknown(tmp_path): + """No RULES glob names *.csv: still captured, never dropped.""" + _mk(tmp_path, "data.csv", "a,b\n1,2\n") + arts = discover_artifacts(tmp_path, "a") + assert list(arts) == ["data.csv"] + assert arts["data.csv"].format == "text" + assert arts["data.csv"].roles == ["unknown"] + assert arts["data.csv"].source == "a,b\n1,2\n" + + +def test_unmatched_binary_file_captured_with_empty_source(tmp_path): + raw = b"\x89PNG\r\n\x1a\n\x00\x01\x02\x03" + (tmp_path / "logo.png").write_bytes(raw) + arts = discover_artifacts(tmp_path, "a") + assert list(arts) == ["logo.png"] + a = arts["logo.png"] + assert a.format == "binary" + assert a.roles == ["unknown"] + assert a.source == "" + assert a.sha256 == hashlib.sha256(raw).hexdigest() + assert a.size_bytes == len(raw) + + +def test_py_files_never_become_artifacts(tmp_path): + """`.py` is the symbol table's domain -- an unmatched `.py` file never + becomes an artifact (unlike every other unmatched extension). `setup.py` + is the deliberate, pre-existing exception: it is rule-matched (as a + dependency-manifest) despite the `.py` suffix, so "rule-matched: as + today" still applies to it -- only the *unmatched* fallback excludes + `.py`.""" + _mk(tmp_path, "src/app.py", "x = 1\n") + _mk(tmp_path, "pkg/__init__.py", "") + _mk(tmp_path, "setup.py", "from setuptools import setup\n") + arts = discover_artifacts(tmp_path, "a") + assert "src/app.py" not in arts and "pkg/__init__.py" not in arts + assert "setup.py" in arts and arts["setup.py"].roles == ["dependency-manifest"] + + +def test_text_max_bytes_caps_source_and_flags_truncation(tmp_path): + content = "0123456789abcdefGHIJ" # 21 bytes, past a 16-byte cap + _mk(tmp_path, "notes.md", content) + raw = content.encode("utf-8") + arts = discover_artifacts(tmp_path, "a", text_max_bytes=16) + art = arts["notes.md"] + assert art.source == content[:16] + assert art.text_truncated is True + assert art.sha256 == hashlib.sha256(raw).hexdigest() # sha256 always full-file + assert art.size_bytes == len(raw) + + +def test_text_max_bytes_never_raises_on_split_multibyte_char(tmp_path): + """A cap that lands mid-codepoint must decode cleanly, never raise -- + back off to the last clean char boundary (errors='ignore' on the prefix).""" + raw = ("a" * 15 + "é" + "extra-tail").encode("utf-8") # 'é' straddles byte 16 + (tmp_path / "notes.md").write_bytes(raw) + arts = discover_artifacts(tmp_path, "a", text_max_bytes=16) + art = arts["notes.md"] + assert art.text_truncated is True + assert art.source == "a" * 15 + assert len(art.source.encode("utf-8")) <= 16 + assert art.sha256 == hashlib.sha256(raw).hexdigest() + + +def test_file_under_cap_is_not_truncated(tmp_path): + _mk(tmp_path, "notes.md", "short\n") + arts = discover_artifacts(tmp_path, "a", text_max_bytes=16) + assert arts["notes.md"].source == "short\n" + assert arts["notes.md"].text_truncated is False + + +def test_capture_text_false_empties_source_everywhere_else_identical(tmp_path): + _mk(tmp_path, "pyproject.toml", "[project]\nname='a'\n") + _mk(tmp_path, "data.csv", "a,b\n1,2\n") + (tmp_path / "logo.png").write_bytes(b"\x89PNG\r\n\x1a\n\x00\x01") + with_text = discover_artifacts(tmp_path, "a") + without_text = discover_artifacts(tmp_path, "a", capture_text=False) + assert set(with_text) == set(without_text) + for path in with_text: + b = without_text[path] + assert b.source == "" and b.text_truncated is False + a_dict = model_dump(with_text[path]) + b_dict = model_dump(b) + a_dict["source"] = b_dict["source"] = "" + a_dict["text_truncated"] = b_dict["text_truncated"] = False + assert a_dict == b_dict + + +def test_dependency_manifest_exempt_from_text_max_bytes(tmp_path): + """#157 review fix: a dependency-manifest's source IS the extraction + input -- the byte cap targets bulk/incidental assets, never manifests. + A cap far below the file's real size must not truncate it.""" + content = '[project]\ndependencies = ["requests"]\n' # > 16 bytes + _mk(tmp_path, "pyproject.toml", content) + arts = discover_artifacts(tmp_path, "a", text_max_bytes=16) + art = arts["pyproject.toml"] + assert art.source == content + assert art.text_truncated is False + + +def test_dependency_manifest_still_empty_source_with_capture_text_false(tmp_path): + """The manifest exemption is from the byte CAP only -- capture_text=False + still empties source for manifests exactly like everything else.""" + _mk(tmp_path, "pyproject.toml", '[project]\ndependencies = ["requests"]\n') + arts = discover_artifacts(tmp_path, "a", capture_text=False, text_max_bytes=16) + art = arts["pyproject.toml"] + assert art.source == "" and art.text_truncated is False diff --git a/test/test_artifact_models.py b/test/test_artifact_models.py new file mode 100644 index 0000000..116ef5c --- /dev/null +++ b/test/test_artifact_models.py @@ -0,0 +1,44 @@ +"""Task 1: artifact/dependency schema models and id constructors.""" +from codeanalyzer.schema import model_validate_json, model_dump_json +from codeanalyzer.schema.ids import artifact_id, purl_pypi +from codeanalyzer.schema.py_schema import ( + PyApplication, PyArtifact, PyDependency, PyImportBinding, +) + + +def test_artifact_id_is_language_neutral(): + assert artifact_id("myapp", "deploy/docker-compose.yml") == \ + "can://artifact/myapp/deploy/docker-compose.yml" + + +def test_purl_pypi(): + assert purl_pypi("pyyaml") == "pkg:pypi/pyyaml" + + +def test_models_round_trip(): + art = PyArtifact( + id=artifact_id("a", "pyproject.toml"), path="pyproject.toml", + format="toml", roles=["dependency-manifest"], size_bytes=10, + sha256="ab" * 32, source="[project]\n", + ) + dep = PyDependency( + name="requests", spec=">=2.31", kind="runtime", + declared_in=art.id, provides_imports=["requests"], prov=["declared"], + ) + imp = PyImportBinding(module="yaml", bound_to="pyyaml", prov=["heuristic"]) + app = PyApplication.builder().symbol_table({}).call_graph([]).build() + app.artifacts = {art.path: art} + app.dependencies = [dep] + app.unresolved_imports = [imp] + back = model_validate_json(PyApplication, model_dump_json(app)) + assert back.artifacts["pyproject.toml"].kind == "artifact" + assert back.artifacts["pyproject.toml"].extraction == "none" + assert back.artifacts["pyproject.toml"].text_truncated is False + assert back.dependencies[0].locked_version is None + assert back.unresolved_imports[0].bound_to == "pyyaml" + + +def test_defaults_empty_on_old_payload(): + app = PyApplication.builder().symbol_table({}).call_graph([]).build() + back = model_validate_json(PyApplication, model_dump_json(app)) + assert back.artifacts == {} and back.dependencies == [] and back.unresolved_imports == [] diff --git a/test/test_artifact_pipeline.py b/test/test_artifact_pipeline.py new file mode 100644 index 0000000..a6716b8 --- /dev/null +++ b/test/test_artifact_pipeline.py @@ -0,0 +1,71 @@ +"""Task 5: sections populated at every level, identically (L1-data posture).""" +import json +from pathlib import Path +from codeanalyzer.core import Codeanalyzer +from codeanalyzer.options import AnalysisOptions +from codeanalyzer.schema import model_dump_json + + +def _run(tmp_path, project, level): + out = tmp_path / f"out{level}" + opts = AnalysisOptions( + input=project, output=out, analysis_level=level, + no_venv=True, cache_dir=tmp_path / f"cache{level}", + ) + artifacts = Codeanalyzer(opts).analyze() + return artifacts.application + + +def _fixture(tmp_path) -> Path: + proj = tmp_path / "proj" + proj.mkdir() + (proj / "pyproject.toml").write_text('[project]\ndependencies = ["requests"]\n') + (proj / "Dockerfile").write_text("FROM python:3.12\n") + (proj / "app.py").write_text("import requests\nimport colorama\n") + return proj + + +def test_sections_identical_across_levels(tmp_path): + proj = _fixture(tmp_path) + a1 = _run(tmp_path, proj, 1) + a2 = _run(tmp_path, proj, 2) + d1 = json.loads(model_dump_json(a1)) + d2 = json.loads(model_dump_json(a2)) + for field in ("artifacts", "dependencies", "unresolved_imports"): + assert d1[field] == d2[field] + assert sorted(d1["artifacts"]) == ["Dockerfile", "pyproject.toml"] + assert [d["name"] for d in d1["dependencies"]] == ["requests"] + assert [u["module"] for u in d1["unresolved_imports"]] == ["colorama"] + + +def test_resolve_installed_flag_default_off(): + assert AnalysisOptions(input=Path(".")).resolve_installed is False + + +def test_artifact_text_flags_defaults(): + opts = AnalysisOptions(input=Path(".")) + assert opts.artifact_text is True + assert opts.artifact_text_max_bytes == 262144 + + +def test_artifact_text_options_thread_through_core(tmp_path): + """core.py must pass artifact_text/artifact_text_max_bytes to + discover_artifacts -- verified end to end, not just at the discovery unit.""" + proj = tmp_path / "proj" + proj.mkdir() + (proj / "notes.md").write_text("0123456789abcdefGHIJ") # 21 bytes + + capped = Codeanalyzer(AnalysisOptions( + input=proj, analysis_level=1, no_venv=True, cache_dir=tmp_path / "cache-capped", + artifact_text_max_bytes=16, + )).analyze().application + art = capped.artifacts["notes.md"] + assert art.text_truncated is True + assert len(art.source.encode("utf-8")) <= 16 + + no_text = Codeanalyzer(AnalysisOptions( + input=proj, analysis_level=1, no_venv=True, cache_dir=tmp_path / "cache-no-text", + artifact_text=False, + )).analyze().application + assert no_text.artifacts["notes.md"].source == "" + assert no_text.artifacts["notes.md"].text_truncated is False diff --git a/test/test_artifacts_end_to_end.py b/test/test_artifacts_end_to_end.py new file mode 100644 index 0000000..de37909 --- /dev/null +++ b/test/test_artifacts_end_to_end.py @@ -0,0 +1,73 @@ +"""Task 7: full pipeline over the manifests_app fixture + determinism.""" +from pathlib import Path +from codeanalyzer.core import Codeanalyzer +from codeanalyzer.options import AnalysisOptions +from codeanalyzer.schema import model_dump_json + +FIXTURE = Path(__file__).parent / "fixtures" / "whole_applications" / "manifests_app" + + +def _app(tmp_path, tag): + return Codeanalyzer(AnalysisOptions( + input=FIXTURE, analysis_level=1, no_venv=True, cache_dir=tmp_path / tag, + )).analyze().application + + +def test_full_surface(tmp_path): + app = _app(tmp_path, "a") + arts = app.artifacts + # never-drop inventory (#157) makes the artifact set deterministic -- + # exact equality, not a subset check: every non-.py fixture file, and + # nothing else, is inventoried (pkg/__init__.py, pkg/main.py excluded). + assert set(arts) == { + "pyproject.toml", "requirements-dev.txt", "setup.py", "uv.lock", + "environment.yml", "Dockerfile", "docker-compose.yml", + ".github/workflows/ci.yml", "data.csv", "logo.png", + } + assert arts["setup.py"].extraction == "partial" # computed install_requires + assert arts["pyproject.toml"].extraction == "full" + # never-drop inventory (#157 follow-up): unmatched files are captured too. + assert arts["data.csv"].format == "text" and arts["data.csv"].roles == ["unknown"] + assert arts["logo.png"].format == "binary" and arts["logo.png"].roles == ["unknown"] + assert arts["logo.png"].source == "" and arts["logo.png"].sha256 != "" + deps = {d.name: d for d in app.dependencies} + assert deps["requests"].locked_version == "2.32.3" + assert deps["requests"].prov == ["declared", "lockfile"] + assert deps["pytest"].kind == "dev" + assert deps["numpy"].spec == "=1.26" + # setup.py is repo code to the symbol table (statically parsed, never + # executed); its setuptools import is undeclared here by design. + assert [u.module for u in app.unresolved_imports] == ["colorama", "setuptools"] + + +def test_determinism_two_runs(tmp_path): + a = model_dump_json(_app(tmp_path, "r1")) + b = model_dump_json(_app(tmp_path, "r2")) + assert a == b + + +def test_level_invariance_artifacts_and_deps(tmp_path): + """Artifacts, dependencies, and unresolved_imports must be identical at L1 and L4.""" + app_l1 = Codeanalyzer(AnalysisOptions( + input=FIXTURE, analysis_level=1, no_venv=True, cache_dir=tmp_path / "l1", + )).analyze().application + app_l4 = Codeanalyzer(AnalysisOptions( + input=FIXTURE, analysis_level=4, no_venv=True, cache_dir=tmp_path / "l4", + )).analyze().application + + # artifacts section + assert set(app_l1.artifacts.keys()) == set(app_l4.artifacts.keys()) + for name in app_l1.artifacts: + assert app_l1.artifacts[name] == app_l4.artifacts[name] + + # dependencies section + deps_l1 = {d.name: d for d in app_l1.dependencies} + deps_l4 = {d.name: d for d in app_l4.dependencies} + assert set(deps_l1.keys()) == set(deps_l4.keys()) + for name in deps_l1: + assert deps_l1[name] == deps_l4[name] + + # unresolved_imports section + unres_l1 = sorted([u.module for u in app_l1.unresolved_imports]) + unres_l4 = sorted([u.module for u in app_l4.unresolved_imports]) + assert unres_l1 == unres_l4 diff --git a/test/test_dependency_view.py b/test/test_dependency_view.py new file mode 100644 index 0000000..e5ed419 --- /dev/null +++ b/test/test_dependency_view.py @@ -0,0 +1,236 @@ +"""Task 4: declared records, lock backfill, provides_imports, unresolved imports.""" +from pathlib import Path +from codeanalyzer.artifacts.discovery import discover_artifacts +from codeanalyzer.artifacts.dependencies import build_dependency_view +from codeanalyzer.schema.py_schema import PyImport, PyModule + + +def _module(name, imports): + return PyModule.builder().file_path(f"/tmp/{name}.py").module_name(name).imports( + [PyImport(module=m, name="*") for m in imports] + ).build() + + +def _setup(tmp_path): + (tmp_path / "pyproject.toml").write_text( + '[project]\ndependencies = ["requests>=2.31", "PyYAML"]\n' + ) + (tmp_path / "uv.lock").write_text( + '[[package]]\nname = "requests"\nversion = "2.32.3"\n' + ) + arts = discover_artifacts(tmp_path, "app") + mods = { + "app.py": _module("app", ["requests", "yaml", "colorama", "os", "app.util"]), + "app/util.py": _module("app.util", []), + } + return arts, mods + + +def test_records_lock_and_binding(tmp_path): + arts, mods = _setup(tmp_path) + deps, unresolved = build_dependency_view(arts, mods, tmp_path, None, False) + by = {d.name: d for d in deps} + assert by["requests"].prov == ["declared", "lockfile"] + assert by["requests"].locked_version == "2.32.3" + assert by["requests"].declared_in == "can://artifact/app/pyproject.toml" + # pyyaml's prov gains "heuristic" via the alias-table binding below (same + # cause as the dropped provides_imports==[] assert); see test_known_alias_binding. + assert by["pyyaml"].locked_version is None + # provides_imports: requests trivially; pyyaml binds via the alias table + # (see test_known_alias_binding — yaml->pyyaml, not a same-name match). + assert by["requests"].provides_imports == ["requests"] + assert arts["pyproject.toml"].extraction == "full" + + +def test_unresolved_imports(tmp_path): + arts, mods = _setup(tmp_path) + _, unresolved = build_dependency_view(arts, mods, tmp_path, None, False) + u = {b.module: b for b in unresolved} + # yaml: known-alias heuristic binds it to declared pyyaml -> NOT unresolved + # colorama: imported, never declared -> unresolved, unbound + # os: stdlib; app.util: local module -> neither appears + assert set(u) == {"colorama"} + assert u["colorama"].bound_to is None and u["colorama"].prov == [] + + +def test_known_alias_binding(tmp_path): + arts, mods = _setup(tmp_path) + deps, unresolved = build_dependency_view(arts, mods, tmp_path, None, False) + yaml_dep = next(d for d in deps if d.name == "pyyaml") + assert "yaml" in yaml_dep.provides_imports and "heuristic" in yaml_dep.prov + + +def test_installed_metadata_binding(tmp_path): + arts, mods = _setup(tmp_path) + venv = tmp_path / ".venv" + di = venv / "lib" / "python3.12" / "site-packages" / "PyYAML-6.0.1.dist-info" + di.mkdir(parents=True) + (di / "METADATA").write_text("Metadata-Version: 2.1\nName: PyYAML\nVersion: 6.0.1\n") + (di / "top_level.txt").write_text("yaml\n_yaml\n") + deps, _ = build_dependency_view(arts, mods, tmp_path, venv, True) + yaml_dep = next(d for d in deps if d.name == "pyyaml") + assert "yaml" in yaml_dep.provides_imports + assert "installed-metadata" in yaml_dep.prov + + +def test_requirement_refs_chased(tmp_path): + (tmp_path / "reqs").mkdir() + (tmp_path / "reqs" / "requirements.txt").write_text( + "-r base.txt\n-r ../../etc/passwd\nrequests\n" + ) + (tmp_path / "reqs" / "base.txt").write_text("flask>=2\n") + arts = discover_artifacts(tmp_path, "app") + # never-drop inventory (#157): captured, but not independently a manifest. + assert arts["reqs/base.txt"].roles == ["unknown"] + deps, _ = build_dependency_view(arts, {}, tmp_path, None, False) + manifest_id = arts["reqs/requirements.txt"].id + by = {d.name: d for d in deps} + assert set(by) == {"requests", "flask"} # ../../etc/passwd ref ignored + assert by["flask"].declared_in == manifest_id + assert by["requests"].declared_in == manifest_id + + +def test_requirement_ref_chased_kind_from_real_basename(tmp_path): + (tmp_path / "reqs").mkdir() + (tmp_path / "reqs" / "requirements.txt").write_text("-r dev.txt\n") + (tmp_path / "reqs" / "dev.txt").write_text("mypy\n") + arts = discover_artifacts(tmp_path, "app") + # never-drop inventory (#157): captured, but not independently a manifest. + assert arts["reqs/dev.txt"].roles == ["unknown"] + deps, _ = build_dependency_view(arts, {}, tmp_path, None, False) + mypy = next(d for d in deps if d.name == "mypy") + assert mypy.kind == "dev" # from dev.txt's real basename, not the forced "requirements.txt" + + +def test_requirement_ref_to_discovered_target_not_duplicated(tmp_path): + (tmp_path / "requirements.txt").write_text("-r requirements-extra.txt\n") + (tmp_path / "requirements-extra.txt").write_text("rich\n") + arts = discover_artifacts(tmp_path, "app") + assert "requirements-extra.txt" in arts # discovered and parsed on its own + deps, _ = build_dependency_view(arts, {}, tmp_path, None, False) + rich_deps = [d for d in deps if d.name == "rich"] + assert len(rich_deps) == 1 # not duplicated by the chase + assert rich_deps[0].declared_in == arts["requirements-extra.txt"].id + + +def test_resolve_ref_does_not_over_reject_dotdot_prefixed_name(tmp_path): + (tmp_path / "requirements.txt").write_text("-r ..bak.txt\n") + (tmp_path / "..bak.txt").write_text("click\n") + arts = discover_artifacts(tmp_path, "app") + deps, _ = build_dependency_view(arts, {}, tmp_path, None, False) + assert {d.name for d in deps} == {"click"} # "..bak.txt" != escaping ".."/"../..." + + +def test_dotted_alias_does_not_falsely_unresolve_top_level(tmp_path): + """Regression: protobuf's alias table entry maps to the dotted + "google.protobuf". provides_imports keeps that full dotted string, but + the unresolved check must compare TOP-LEVEL segments -- else "google" + falsely resurfaces as unresolved even though protobuf declares it.""" + (tmp_path / "pyproject.toml").write_text('[project]\ndependencies = ["protobuf"]\n') + arts = discover_artifacts(tmp_path, "app") + mods = {"app.py": _module("app", ["google.protobuf"])} + deps, unresolved = build_dependency_view(arts, mods, tmp_path, None, False) + assert "google" not in {u.module for u in unresolved} + protobuf_dep = next(d for d in deps if d.name == "protobuf") + assert "google.protobuf" in protobuf_dep.provides_imports + + +def test_same_name_match_carries_no_heuristic_prov(tmp_path): + """Regression: identity entries ("setuptools": "setuptools", "pymongo": + "pymongo") in the alias table minted a spurious "heuristic" prov on a + plain same-name match. A same-name match must be prov == ["declared"].""" + (tmp_path / "pyproject.toml").write_text('[project]\ndependencies = ["setuptools"]\n') + arts = discover_artifacts(tmp_path, "app") + mods = {"app.py": _module("app", ["setuptools"])} + deps, _ = build_dependency_view(arts, mods, tmp_path, None, False) + setuptools_dep = next(d for d in deps if d.name == "setuptools") + assert setuptools_dep.prov == ["declared"] + assert setuptools_dep.provides_imports == ["setuptools"] + + +def test_local_package_top_level_not_falsely_unresolved(tmp_path): + """Regression (confirmed on odoo-slim): module_name is py_file.stem -- + the leaf filename only ("api" for "odoo/api.py") -- so a local set built + from module_name alone never contains the top-level PACKAGE name itself. + A sibling module doing `import odoo` then falsely lands in + unresolved_imports. Fix: also derive local tops from the symbol-table + keys (first path segment when nested).""" + mods = { + "odoo/api.py": _module("api", []), + "pkg/main.py": _module("main", ["odoo"]), + } + _, unresolved = build_dependency_view({}, mods, tmp_path, None, False) + assert "odoo" not in {u.module for u in unresolved} + + +def test_lock_only_transitive_dep_emitted_indirect(tmp_path): + (tmp_path / "pyproject.toml").write_text('[project]\ndependencies = ["requests>=2.31"]\n') + (tmp_path / "uv.lock").write_text( + '[[package]]\nname = "requests"\nversion = "2.32.3"\n' + '[[package]]\nname = "urllib3"\nversion = "2.2.1"\n' + ) + from codeanalyzer.artifacts.discovery import discover_artifacts + from codeanalyzer.artifacts.dependencies import build_dependency_view + arts = discover_artifacts(tmp_path, "app") + deps, _ = build_dependency_view(arts, {}, tmp_path, None, False) + by = {d.name: d for d in deps} + assert by["requests"].direct is True and by["requests"].locked_version == "2.32.3" + u = by["urllib3"] + assert u.direct is False + assert u.prov == ["lockfile"] and u.locked_version == "2.2.1" + assert u.declared_in == arts["uv.lock"].id + assert u.spec == "" + + +def _big_lock_text(min_bytes: int) -> str: + """Enough valid ``[[package]]`` entries to exceed ``min_bytes``.""" + parts = [] + size = 0 + i = 0 + while size < min_bytes: + entry = f'[[package]]\nname = "pkg{i}"\nversion = "1.0.{i}"\n\n' + parts.append(entry) + size += len(entry.encode("utf-8")) + i += 1 + return "".join(parts) + + +def test_large_lock_parses_all_pins_under_default_text_cap(tmp_path): + """#157 review fix: a lock bigger than the default 262144-byte cap must + still parse in full -- dependency-manifest artifacts are exempt from + text_max_bytes at discovery, and extraction reads the file fresh besides.""" + text = _big_lock_text(300_000) + assert len(text.encode("utf-8")) > 262144 + (tmp_path / "uv.lock").write_text(text) + package_count = text.count("[[package]]") + arts = discover_artifacts(tmp_path, "app") # default text_max_bytes + assert arts["uv.lock"].text_truncated is False + deps, _ = build_dependency_view(arts, {}, tmp_path, None, False) + pinned = {d.name for d in deps if d.locked_version} + assert len(pinned) == package_count + assert arts["uv.lock"].extraction == "full" # set by build_dependency_view + + +def test_large_lock_parses_all_pins_even_with_capture_text_false(tmp_path): + """Extraction must not depend on the stored `source` at all: with + --no-artifact-text (capture_text=False) the stored source is empty, but + build_dependency_view reads the real file fresh, so deps are unaffected.""" + text = _big_lock_text(300_000) + (tmp_path / "uv.lock").write_text(text) + package_count = text.count("[[package]]") + arts = discover_artifacts(tmp_path, "app", capture_text=False) + assert arts["uv.lock"].source == "" + deps, _ = build_dependency_view(arts, {}, tmp_path, None, False) + pinned = {d.name for d in deps if d.locked_version} + assert len(pinned) == package_count + assert arts["uv.lock"].extraction == "full" + + +def test_corrupted_lock_extraction_is_partial_not_full(tmp_path): + """A lock with real (non-empty) content that parse_lock_pins can't make + sense of must not claim extraction="full" for zero pins extracted.""" + (tmp_path / "uv.lock").write_text("this is not valid toml at all {{{\n") + arts = discover_artifacts(tmp_path, "app") + deps, _ = build_dependency_view(arts, {}, tmp_path, None, False) + assert arts["uv.lock"].extraction == "partial" + assert deps == [] diff --git a/test/test_manifest_parsers.py b/test/test_manifest_parsers.py new file mode 100644 index 0000000..fd918ce --- /dev/null +++ b/test/test_manifest_parsers.py @@ -0,0 +1,157 @@ +"""Task 3: every spec §6 manifest format parses into RawDep records.""" +import textwrap +from codeanalyzer.artifacts.parsers import ( + RawDep, normalize_name, parse_lock_pins, parse_manifest, + parse_requirement_line, parse_requirement_refs, +) + + +def test_normalize_name(): + assert normalize_name("PyYAML") == "pyyaml" + assert normalize_name("ruamel.yaml") == "ruamel-yaml" + assert normalize_name("typing_extensions") == "typing-extensions" + + +def test_requirements_txt(): + text = textwrap.dedent("""\ + # comment + requests>=2.31,<3 + pyyaml + celery[redis]==5.3.* + -e ./local-pkg + --index-url https://example.invalid + """) + deps, partial = parse_manifest("requirements.txt", text) + assert not partial + assert [(d.name, d.spec) for d in deps] == [ + ("requests", ">=2.31,<3"), ("pyyaml", ""), ("celery", "==5.3.*"), + ] + assert deps[2].extras == ("redis",) + assert all(d.kind == "runtime" for d in deps) + + +def test_requirements_dev_kind(): + deps, _ = parse_manifest("requirements-dev.txt", "pytest\n") + assert deps[0].kind == "dev" + + +def test_pyproject_pep621_poetry_and_build(): + text = textwrap.dedent("""\ + [build-system] + requires = ["setuptools>=68"] + [project] + dependencies = ["requests>=2.31"] + [project.optional-dependencies] + docs = ["sphinx"] + [tool.poetry.dependencies] + python = "^3.10" + rich = "^13.0" + [tool.poetry.group.dev.dependencies] + mypy = "*" + """) + deps, partial = parse_manifest("pyproject.toml", text) + assert not partial + by = {(d.name, d.kind) for d in deps} + assert ("setuptools", "build") in by + assert ("requests", "runtime") in by + assert ("sphinx", "optional") in by + assert ("rich", "runtime") in by and ("mypy", "dev") in by + assert ("python", "runtime") not in {(d.name, d.kind) for d in deps} # interpreter, not a dep + + +def test_setup_py_static_literals(): + text = 'from setuptools import setup\nsetup(install_requires=["flask>=2"], extras_require={"test": ["pytest"]})\n' + deps, partial = parse_manifest("setup.py", text) + assert not partial + assert {(d.name, d.kind) for d in deps} == {("flask", "runtime"), ("pytest", "optional")} + + +def test_setup_py_dynamic_is_partial(): + text = "from setuptools import setup\nreqs = compute()\nsetup(install_requires=reqs)\n" + deps, partial = parse_manifest("setup.py", text) + assert partial and deps == [] + + +def test_setup_cfg(): + text = "[options]\ninstall_requires =\n numpy>=1.24\n pandas\n" + deps, _ = parse_manifest("setup.cfg", text) + assert [(d.name, d.spec) for d in deps] == [("numpy", ">=1.24"), ("pandas", "")] + + +def test_pipfile_and_environment_yml(): + pip = '[packages]\nrequests = ">=2.31"\n[dev-packages]\nblack = "*"\n' + deps, _ = parse_manifest("Pipfile", pip) + assert {(d.name, d.kind, d.spec) for d in deps} == { + ("requests", "runtime", ">=2.31"), ("black", "dev", ""), + } + env = "dependencies:\n - numpy=1.26\n - scipy>=1.10\n - pip\n - pip:\n - fastapi>=0.100\n" + deps, _ = parse_manifest("environment.yml", env) + assert {(d.name, d.spec) for d in deps} == { + ("numpy", "=1.26"), ("scipy", ">=1.10"), ("fastapi", ">=0.100"), + } + + +def test_lock_pins(): + poetry = '[[package]]\nname = "requests"\nversion = "2.31.0"\n[[package]]\nname = "PyYAML"\nversion = "6.0.1"\n' + assert parse_lock_pins("poetry.lock", poetry) == {"requests": "2.31.0", "pyyaml": "6.0.1"} + uv = '[[package]]\nname = "requests"\nversion = "2.32.0"\n' + assert parse_lock_pins("uv.lock", uv) == {"requests": "2.32.0"} + pipf = '{"default": {"requests": {"version": "==2.31.0"}}, "develop": {}}' + assert parse_lock_pins("Pipfile.lock", pipf) == {"requests": "2.31.0"} + + +def test_requirements_kind_word_boundary(): + deps, _ = parse_manifest("requirements-docker.txt", "requests\n") + assert deps[0].kind == "runtime" + deps, _ = parse_manifest("requirements-latest.txt", "requests\n") + assert deps[0].kind == "runtime" + deps, _ = parse_manifest("requirements-dev.txt", "requests\n") + assert deps[0].kind == "dev" + + +def test_malformed_manifests_are_partial(): + assert parse_manifest("pyproject.toml", "not [ valid toml") == ([], True) + assert parse_manifest("Pipfile", "not [ valid toml") == ([], True) + assert parse_manifest("setup.cfg", "[options\nbroken") == ([], True) + assert parse_manifest("environment.yml", "dependencies: [unclosed") == ([], True) + + +def test_malformed_lock_files_return_empty(): + assert parse_lock_pins("poetry.lock", "not [ valid toml") == {} + assert parse_lock_pins("uv.lock", "not [ valid toml") == {} + assert parse_lock_pins("Pipfile.lock", "not valid json") == {} + + +def test_poetry_inline_table_dep(): + text = textwrap.dedent("""\ + [tool.poetry.dependencies] + python = "^3.10" + requests = {version = "^2.28", extras = ["socks"]} + """) + deps, partial = parse_manifest("pyproject.toml", text) + assert not partial + assert deps == [RawDep("requests", "^2.28", "runtime", ("socks",))] + + +def test_pipfile_inline_table_extras(): + text = '[packages]\nrequests = {version = "*", extras = ["security"]}\n' + deps, _ = parse_manifest("Pipfile", text) + assert deps == [RawDep("requests", "", "runtime", ("security",))] + + +def test_direct_ref_and_line_continuation(): + assert parse_requirement_line("mylib @ git+https://x/y.git") == RawDep("mylib", "", "runtime", ()) + assert parse_requirement_line("pkg==1.4.2 \\") == RawDep("pkg", "==1.4.2", "runtime", ()) + + +def test_parse_requirement_refs(): + text = "requests\n-r base.txt\n-c constraints/prod.txt\npytest\n" + assert parse_requirement_refs(text) == ["base.txt", "constraints/prod.txt"] + long_form = "--requirement base.txt\n--constraint constraints/prod.txt\n" + assert parse_requirement_refs(long_form) == ["base.txt", "constraints/prod.txt"] + + +def test_setup_py_syntax_error(): + deps, partial = parse_manifest("setup.py", "def foo(:\n pass") + assert deps == [] + assert partial is True diff --git a/test/test_neo4j_artifacts.py b/test/test_neo4j_artifacts.py new file mode 100644 index 0000000..42452ee --- /dev/null +++ b/test/test_neo4j_artifacts.py @@ -0,0 +1,106 @@ +"""Task 6: artifact/dependency rows in the Neo4j projection.""" +from codeanalyzer.neo4j.schema import NODE_LABELS, REL_TYPES + + +def test_catalog_has_neutral_vocabulary(): + labels = {n.label: n for n in NODE_LABELS} + assert labels["Artifact"].key == "id" and labels["Package"].key == "id" + assert labels["Artifact"].properties["text_truncated"] == "boolean" + rels = {r.type for r in REL_TYPES} + assert {"HAS_ARTIFACT", "DECLARES_DEPENDENCY", "LOCKS", + "PY_PROVIDES", "PY_UNRESOLVED_IMPORT"} <= rels + + +def test_rows_projected(tmp_path): + proj = tmp_path / "p" + proj.mkdir() + (proj / "pyproject.toml").write_text('[project]\ndependencies = ["requests"]\n') + (proj / "app.py").write_text("import requests\nimport colorama\nrequests.get('u')\n") + from codeanalyzer.core import Codeanalyzer + from codeanalyzer.options import AnalysisOptions + app = Codeanalyzer(AnalysisOptions( + input=proj, analysis_level=2, no_venv=True, cache_dir=tmp_path / "c", + )).analyze().application + + from codeanalyzer.neo4j.project import project + from codeanalyzer.schema.assign_ids import assign_ids + rows = project(app, "p", assign_ids(app, "p")) + + nodes = {(n.labels[0], n.value) for n in rows.nodes} + assert ("Artifact", "can://artifact/p/pyproject.toml") in nodes + assert ("Package", "pkg:pypi/requests") in nodes + pyproject_node = next( + n for n in rows.nodes + if n.labels[0] == "Artifact" and n.value == "can://artifact/p/pyproject.toml" + ) + assert pyproject_node.props["text_truncated"] is False + rel_types = {e.type for e in rows.edges} + assert {"HAS_ARTIFACT", "DECLARES_DEPENDENCY", "PY_PROVIDES"} <= rel_types + + +def test_locks_and_provides_dedup_across_multi_manifest_declarations(tmp_path): + """A package declared in 2+ manifests yields one PyDependency record per + manifest (Task 5's `build_dependency_view`), so DECLARES_DEPENDENCY -- + correctly -- fires once per manifest. LOCKS/PY_PROVIDES are per-PACKAGE + facts, not per-declaration, and must not duplicate just because the + package happens to be declared twice.""" + proj = tmp_path / "p" + proj.mkdir() + (proj / "requirements.txt").write_text("requests==2.31.0\n") + (proj / "requirements-dev.txt").write_text("requests==2.31.0\n") + (proj / "poetry.lock").write_text('[[package]]\nname = "requests"\nversion = "2.31.0"\n') + (proj / "app.py").write_text("import requests\n") + from codeanalyzer.core import Codeanalyzer + from codeanalyzer.options import AnalysisOptions + app = Codeanalyzer(AnalysisOptions( + input=proj, analysis_level=2, no_venv=True, cache_dir=tmp_path / "c", + )).analyze().application + assert len(app.dependencies) == 2, "fixture must produce two declarations of one package" + + from codeanalyzer.neo4j.project import project + from codeanalyzer.schema.assign_ids import assign_ids + rows = project(app, "p", assign_ids(app, "p")) + + locks = [e for e in rows.edges if e.type == "LOCKS"] + provides = [e for e in rows.edges if e.type == "PY_PROVIDES"] + declares = [e for e in rows.edges if e.type == "DECLARES_DEPENDENCY"] + + assert len(locks) == 1, f"expected exactly one LOCKS row, got {len(locks)}" + assert locks[0].from_ref.value == "can://artifact/p/poetry.lock" + assert locks[0].to_ref.value == "pkg:pypi/requests" + + assert len(provides) == 1, f"expected exactly one PY_PROVIDES row, got {len(provides)}" + assert provides[0].from_ref.value == "pkg:pypi/requests" + + assert len(declares) == 2, "one DECLARES_DEPENDENCY row per declaring manifest" + + +def test_declares_dependency_distinct_by_kind_within_one_manifest(tmp_path): + """One manifest re-declaring the same package under two kinds (e.g. + requests in [project.dependencies] AND again under + [project.optional-dependencies]) yields two PyDependency records with + identical DECLARES_DEPENDENCY endpoints (same manifest, same package) and + no discriminant -- without `key=d.kind` the Cypher/Bolt MERGE collapses + them into one relationship.""" + proj = tmp_path / "p" + proj.mkdir() + (proj / "pyproject.toml").write_text( + '[project]\ndependencies = ["requests"]\n' + '[project.optional-dependencies]\nextra = ["requests"]\n' + ) + (proj / "app.py").write_text("import requests\n") + from codeanalyzer.core import Codeanalyzer + from codeanalyzer.options import AnalysisOptions + app = Codeanalyzer(AnalysisOptions( + input=proj, analysis_level=2, no_venv=True, cache_dir=tmp_path / "c", + )).analyze().application + assert len(app.dependencies) == 2, "fixture must produce two declarations of one package" + + from codeanalyzer.neo4j.project import project + from codeanalyzer.schema.assign_ids import assign_ids + rows = project(app, "p", assign_ids(app, "p")) + + declares = [e for e in rows.edges if e.type == "DECLARES_DEPENDENCY"] + assert len(declares) == 2, "one DECLARES_DEPENDENCY row per declaration, not collapsed" + keys = {e.key for e in declares} + assert keys == {"runtime", "optional"}, f"expected two distinct kind keys, got {keys}"