Is your feature request related to a problem? Please describe.
Two byte-identical invocations of the analyzer produce different call graphs. The variance is small but real, and it is not the sharding defect from #145 — it survives with --pycg-shard sequential, with --ray, and at any --pycg-max-iter.
Measured on test/fixtures/whole_applications/flask (82 files), -a 2 --no-venv --pycg-shard --pycg-shard-ceiling 25 --clear-cache, two runs per configuration:
| config |
edges A / B |
differing edges |
sequential, --pycg-max-iter 50 |
1263 / 1263 |
1 |
sequential, --pycg-max-iter 3 |
1263 / 1263 |
2 |
--ray, --pycg-max-iter 50 |
1168 / 1168 |
3 |
--ray, --pycg-max-iter 3 |
1168 / 1167 |
2 |
0.1–0.3% of edges. Ray adds no meaningful variance, so this is not a scheduling artifact.
Almost every differing edge involves file-IO types — _TextIOBase, _BufferedIOBase, FileIO. The reproducible instance is src/flask/cli.py:1004:
with open(startup) as f:
eval(compile(f.read(), startup, "exec"), ctx)
open() is overloaded in typeshed (text mode → TextIOWrapper, binary → BufferedReader). Jedi's overload resolution returns a different candidate set on different runs, so f.read() resolves to either:
can://python/flask/src/flask/cli.py/shell_command() -> …/@external/builtins._TextIOBase/read
can://python/flask/src/flask/cli.py/shell_command() -> …/@external/builtins._BufferedIOBase/read
SymbolTableBuilder._first_definition is already a deterministic picker — min(definitions, key=lambda d: (d.full_name or "", d.name or "")). Given both candidates it would always choose _BufferedIOBase (B < T). The instability is in the candidate set Jedi hands us, not in our selection.
PYTHONHASHSEED=0 is already pinned for the driver (#99) and propagated into Ray workers via runtime_env.env_vars (core.py:48), so hash randomisation is not the cause.
Describe the solution you'd like
Not stated in the original issue.
Describe alternatives you've considered
Not stated in the original issue.
Additional context
Caveats and known risks
Definition of done
Is your feature request related to a problem? Please describe.
Two byte-identical invocations of the analyzer produce different call graphs. The variance is small but real, and it is not the sharding defect from #145 — it survives with
--pycg-shardsequential, with--ray, and at any--pycg-max-iter.Measured on
test/fixtures/whole_applications/flask(82 files),-a 2 --no-venv --pycg-shard --pycg-shard-ceiling 25 --clear-cache, two runs per configuration:--pycg-max-iter 50--pycg-max-iter 3--ray,--pycg-max-iter 50--ray,--pycg-max-iter 30.1–0.3% of edges. Ray adds no meaningful variance, so this is not a scheduling artifact.
Almost every differing edge involves file-IO types —
_TextIOBase,_BufferedIOBase,FileIO. The reproducible instance issrc/flask/cli.py:1004:open()is overloaded in typeshed (text mode →TextIOWrapper, binary →BufferedReader). Jedi's overload resolution returns a different candidate set on different runs, sof.read()resolves to either:SymbolTableBuilder._first_definitionis already a deterministic picker —min(definitions, key=lambda d: (d.full_name or "", d.name or "")). Given both candidates it would always choose_BufferedIOBase(B<T). The instability is in the candidate set Jedi hands us, not in our selection.PYTHONHASHSEED=0is already pinned for the driver (#99) and propagated into Ray workers viaruntime_env.env_vars(core.py:48), so hash randomisation is not the cause.Describe the solution you'd like
Not stated in the original issue.
Describe alternatives you've considered
Not stated in the original issue.
Additional context
Caveats and known risks
test_l2_runs_are_byte_identical) runs onrequests— 35 files, below the shard ceiling, so it never reaches the sharded path and cannot observe this.only_stubs/prefer_stubs), disabling a Jedi cache (cost: speed), or post-filtering ambiguous IO overloads to a canonical choice. Each trades precision or runtime; none is obviously right.Definition of done
(src, dst, weight, prov)--pycg-shard-ceiling 25yields ~4 shards), which L2 call graph is nondeterministic run-to-run (PyCG fixpoint under --pycg-max-iter) #99'srequestsgate cannot