Skip to content

Python impact/test-impact loop: type test parameters from their fixtures (+ iterations) - #1873

Closed
swapnilpaliwal-sd wants to merge 8 commits into
apps/integration-0.1.9from
fix/py-impact-loop
Closed

swapnilpaliwal-sd wants to merge 8 commits into
apps/integration-0.1.9from
fix/py-impact-loop

Conversation

@swapnilpaliwal-sd

@swapnilpaliwal-sd swapnilpaliwal-sd commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Iterative work on Python impact / test-impact, measured on 16 public repositories (10 tuned on, 6 held out) with three behavioural oracles built for it:

  • mutation — each of 60 sampled functions per repo broken in turn, the test files that then fail (per-test baseline subtraction, collection errors and a suite that never starts both counted, -B with a swept __pycache__, per-mutant timeout at 4x the clean suite);
  • runtime trace — the call graph the suite actually executes (nearest repository frame, with or without a library frame between), for impact's callers and path's chains;
  • issue -> fix PR — the source files a later fix touched, for context.

Commit 1 — a test parameter typed by its fixture

pytest fills def test_commit(session) with what the session fixture returned or yielded; no call site spells that call, so the parameter was untyped and every call on it — and everything derived from it — was ambiguous_unknown. 1,529 such calls across ten suites.

value-flow.dl now treats the runner's call as one more argument reaching a parameter (param_arg_type from the fixture's value type). Which fixture serves the parameter comes from a syntax-only copy of the runner's lookup in framework-behavior/dispatch.dl (§2a): the existing py_fixture_injection reads module_member_method for star-imported fixtures, so reading it from resolution makes the nearest-conftest max a cyclic aggregate. Contract note: resolution now reads two relations defined in the framework-behavior file (§2a, §2b); nothing there writes a resolution relation. Star-imported / pytest_plugins fixtures keep reaching their tests through the injection edge and only stay untyped.

on 10 held-in repos base this
impact --tests recall 0.668 0.698
every failing file selected 46.7% 50.5%
precision 0.403 0.408
callers / path / context — unchanged

Three repos carrying the pattern move; seven without it are byte-identical. Engine suite 43/43 with identical case results; query cases 310/310; new case fixture-value-types-the-parameter fails on the base engine on its derived-local check.

swapnilpaliwal-sd and others added 3 commits October 9, 2026 01:17
pytest fills a test's parameter with whatever the fixture of that name returned or
yielded, and no call site spells that call, so the parameter stayed untyped: every
method called on it was ambiguous_unknown, and everything derived from it (a local
assigned from one of those calls, a with block) went with it. Across ten public
suites 1,529 such calls were unresolved; a suite whose fixtures are annotated
already resolved its own.

value-flow.dl now treats the runner's call as one more argument reaching a
parameter: param_arg_type from the fixture's value type (its return, its yield, or
its declared return). The fixture serving a parameter comes from a syntax-only copy
of the runner's lookup (class, own module, nearest conftest): the existing
py_fixture_injection reads module_member_method for star-imported fixtures, which
is resolution, and its nearest-conftest MAX would then be a cyclic aggregate.
Star-imported and pytest_plugins fixtures still reach their tests through the
injection edge; they only stay untyped.

Measured with a behavioural oracle (break each of 600 sampled functions, record
which test files fail) on ten held-in repositories: impact --tests recall
0.668 -> 0.698, every-failing-file-selected 46.7% -> 50.5%, precision unchanged;
the three repositories carrying the pattern move, the seven without it are
byte-identical. Engine suite 43/43 with identical case results; query cases
310/310. New case fixture-value-types-the-parameter fails on the base engine on its
derived-local check.

Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
…(T, wrapper) returns wrapper

functools.update_wrapper(wrapper, wrapped) hands back `wrapper`, and typing.cast(T, x)
hands back `x`; both are documented, neither module is staged in a client-only run. A
decorator written `return update_wrapper(wrapper, f)`, or typed as
`return t.cast(F, update_wrapper(wrapper, f))`, was therefore opaque: every method it
decorates was decorator_replaced_target and no call through the attribute reached
anything. That shape occurs in five of sixteen public repositories measured.

A catalogue, py_returns_arg(DottedName, Position) in builtins.dl, lists the two
functions and the argument each returns; call_chain.dl follows such calls (to any depth)
to the returned name, matched through the import that binds the callee: a module
import, its alias (`import typing as t` is MODULE_IMPORT_ALIAS, which the copy.copy
rule beside it also misses), or a from-import.

On ten held-in repositories, path's found rate on runtime-proven chains goes
0.720 -> 0.744 (the repository with the shape: 0.52 -> 0.76); test selection and the
other verbs unchanged. Engine suite 43/43 with identical case results; query cases
313/313; case decorator-returning-update-wrapper fails on the base engine on both
decorator checks.

Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
…y the import walk

impact's `at import` rung names every test file importing a module whose BODY reaches
the change through calls that run, because a module that raises while being imported
fails every file importing it. A module body's calls inside `if __name__ == "__main__":`
are real edges, but they run only when the file is executed as a script, never on import.
Walking them named every importer of a module that ends in a demo block.

The fact export now records guard_only(module, callee) where every call site the module
body has to that callee lies inside a top-level main guard, and up_running does not take
those edges. Ordinary reachability is unchanged; only the import hop is.

Measured with the mutation oracle on ten held-in repositories, the at-import rung was
the second-noisiest (precision 0.16 over 720 selected test files), and 637 of those 720
were for targets that never run on import. On the repository carrying the shape its
at-import selections drop 713 -> 305 and test-selection precision 0.335 -> 0.373 with
recall unchanged; the control repository whose at-import selections were all correct is
unchanged. Query cases 315/315 python, 264/264 typescript, 333/333 java; case
main-guard-does-not-run-on-import fails on the base engine on the guarded check.

Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
@swapnilpaliwal-sd

Copy link
Copy Markdown
Contributor Author

Two more commits on the same loop (same oracles, same ten held-in repositories).

Commit 2 — update_wrapper / typing.cast hand back an argument. A decorator written return update_wrapper(wrapper, f), or typed return t.cast(F, update_wrapper(wrapper, f)), was opaque: every method it decorates was decorator_replaced_target. A catalogue py_returns_arg(DottedName, Position) in builtins.dl lists the two (documented behaviour), and call_chain.dl follows such calls to the returned name through the binding import — module, alias (import typing as t is MODULE_IMPORT_ALIAS; the copy.copy rule next to it has the same blind spot), or from-import. The shape is in 5 of 16 repositories. path found-rate on runtime-proven chains 0.720 → 0.744 (0.52 → 0.76 where it occurs); nothing else moves.

Commit 3 — the import walk skips if __name__ == "__main__":. The at import rung was the second-noisiest (precision 0.16 over 720 selected test files; 637 of them for targets that never run on import). Module-body calls under the main guard are real edges but run only as a script; the fact export now records guard_only and up_running does not take those edges. Where the shape occurs: at-import selections 713 → 305, precision 0.335 → 0.373, recall unchanged; the control repository (all its at-import selections correct) unchanged.

Each commit has a case that fails on the base engine; query cases 315/315 python, 264/264 typescript, 333/333 java; engine suite 43/43 with identical case results.

Held-out evaluation (six untouched repositories) is running and will be reported here.

swapnilpaliwal-sd and others added 3 commits October 9, 2026 01:48
…_"`) is a guard too

`if sys.platform != "win32" and __name__ == "__main__":` is false on import whatever the
other conjunct is, so the block cannot run then; an `or` could, and is not read as a guard.
The first guard rule missed the conjunction, which on the repository carrying the shape
was where almost all of the remaining false at-import routes came from.

There: at-import selections 305 -> 55, of which 53 fail under the mutation oracle; test-
selection precision 0.373 -> 0.446. Recall 0.810 -> 0.794: a few failing files had been
selected only through an at-import route that cannot run, right by accident; their real
route is one the graph does not see. Control repository unchanged. IMPACT_VERSION 65, so
fact caches written under 64 are not reused. Case extended with the conjunctive spelling;
query cases green.

Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
…hat builds the object, not from every constructor

impact takes a hop from a protocol member (a dunder no call site names) to whoever
constructs its type: you do not build a context manager without entering it. That holds
for __enter__/__exit__, __bool__, __hash__, __call__ (precision 0.61-0.80 under the
mutation oracle), and not for showing a value (__repr__, __str__, __format__), pickling or
copying it (__getstate__, __setstate__, __reduce__, ...) or deleting from it (__delitem__,
__delattr__, __del__): those run only when some code asks, and a constructor deep inside
other code says nothing about that. Measured on ten held-in repositories, __repr__ alone
was 124 selected test files at precision 0.05, across eight of them.

For those members the hop is now taken only from test code that constructs the type: a
test that builds the object is the one that asks for its repr. Dropping the hop outright
removed 403 false files but left 13 targets with a failing test and an empty answer; this
form keeps every answer non-empty.

Ten held-in repositories (this commit with the two main-guard ones): test-selection
precision 0.408 -> 0.454, recall 0.698 -> 0.694, empty answers unchanged at 84/535. Case
on-demand-dunder-hop-from-tests fails on the base engine on the __repr__ check; its control
(__eq__ keeps every constructor) passes on both. Query cases 318/318 python, 264/264
typescript, 333/333 java.

Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
…setter and deleter

The engine emitted a PROPERTY_READ edge for reading a @Property and nothing for
writing or deleting one, though `obj.x = v` runs `@x.setter` and `del obj.x` runs
`@x.deleter` exactly as a read runs the getter. A setter therefore had no caller: an
impact on it named no test, and everything the setter calls was cut off from whoever
assigns. Setters occur in five of sixteen public repositories measured; the torture
fixture documented the gap as a known miss.

type_property_accessor finds the setter / deleter on the class that wins the name in
the receiver's MRO (they share the getter's binding, so mro_lookup's single answer is
not enough); property_write_edge keys on the attribute access's STORE / DEL context;
the edge is exported as call kind PROPERTY_WRITE (schema vocabulary for Python, the
kind TypeScript already uses for its accessors), mapped to the `property` rung.

Torture: the CPython oracle confirms the new edge (agree 587 -> 588, missing 47 -> 46,
extra unchanged); f43's known-miss marker is removed with its docstring rewritten at the
same line count; goldens updated. Engine suite 43/43. Query cases 322/322 python,
264/264 typescript, 333/333 java; case property-setter-is-a-call fails on the base
engine on all three write/delete checks.

All six commits together, ten held-in repositories: test-selection recall
0.668 -> 0.700, precision 0.403 -> 0.452, impact source-caller recall 0.673 -> 0.686,
path found 0.720 -> 0.752, context recall@5 0.789 -> 0.799; no repository regresses
beyond 0.003 on any of them.

Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
@swapnilpaliwal-sd

Copy link
Copy Markdown
Contributor Author

Two more commits, plus the conjunctive-guard follow-up.

Commit 4b — cond and __name__ == "__main__" is a guard too. The first guard rule missed the conjunction; with it, the at-import rung on the repository carrying the shape goes 713 → 55 selected files, 53 of which fail (precision 0.96).

Commit 5 — on-demand protocol members are reached from test code that builds the object. The construction hop (whoever builds a T runs T's protocol) is right for __enter__/__exit__, __bool__, __hash__, __call__ (0.61–0.80) and wrong for showing (__repr__, __str__, __format__), pickling/copying and deleting (__repr__ alone: 124 files at 0.05 across eight repositories). For those the hop is now taken only from test code that constructs the type. Dropping it outright removed 403 false files but left 13 targets with an empty answer; this form keeps every answer non-empty.

Commit 6 — property writes are calls. obj.x = v runs the setter and del obj.x the deleter; the engine emitted nothing for either, so a setter had no caller and no tests. New PROPERTY_WRITE edges (schema vocabulary for Python). The torture fixture had documented this as a known miss; the CPython oracle confirms the new edge (agree 587 → 588, extra unchanged).

Cumulative, ten held-in repositories (base → now):

base now
test-selection recall 0.668 0.700
test-selection precision 0.403 0.452
every failing file selected 46.7% 48.6%
impact source-caller recall 0.673 0.686
path found (runtime-proven chains) 0.720 0.752
context recall@5 0.789 0.799

No repository regresses by more than 0.003 on any metric. Held-out evaluation on six untouched repositories is still generating truth.

swapnilpaliwal-sd and others added 2 commits October 9, 2026 02:44
…operand of a 3+ way union counts

Two gaps in "Optional[X] IS X", the rule that types a value annotated Optional[X] or
X | None:

- It existed for parameters, returns and fields but not for a LOCAL variable's
  annotation, so `ctx: Optional[Context] = ...` and `x: Foo | None = ...` left every
  call on the local unresolved while the same annotation on a parameter resolved.
- `|` is left-associative: `A | B | None` is `(A | B) | None`, so A and B sit two
  levels below the annotation, under a nested union, and the depth-1 element rule saw
  only that nested union (which names no type) and None. Every union of three or more
  operands lost its members, on parameters, returns, fields and locals alike.

binding_declared_nominal gains the Optional/union clause; union_operand walks nested
PEP 604 unions and feeds their operands to type_ref_element, resolved and by name.
Containers are untouched (union_operand is rooted at an Optional/union kind only).

Optional/union locals occur in 15 of 16 public repositories measured (96 written with
`|`, 50 with Optional[]/Union[]). Ten held-in repositories, against the previous
commit: path found 0.752 -> 0.756, impact source-caller recall 0.686 -> 0.687 (one
repository 0.795 -> 0.807); the new callers are a union's dispatch set and are labelled
`one of a set`, resolved-caller precision unchanged at 0.979. Engine suite 43/43,
torture unchanged; query cases 325/325; case optional-and-union-annotations fails on
the base engine on both the local and the three-operand checks.

Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
…twin

A package that ships `.pyi` stubs beside its `.py` modules declares every function
twice under one qualified name. Name resolution can reach only the stub's declaration:
the parser resolves a relative import such as `from .impl import *` in pkg/__init__.py
to impl.pyi rather than impl.py, so `import pkg; pkg.f()` lands on the stub side. The
.pyi gate (#223) then correctly refuses the stub as an edge target, and the site was left
with no client target at all: every test call into such a package ended at the library
boundary.

method_stub_twin pairs a stub declaration with the non-stub declaration of the same
qualified name, and expr_call_candidate adds the twin wherever the stub is a candidate.
The stub stays a candidate and a type source, and stays off the edge, exactly as the
gate's design intends; the twin is what CPython runs. (The parser's relative-import
resolution preferring the stub is the upstream cause and is left as it is here.)

Two of sixteen public repositories ship stubs this way. On the one held out from tuning,
where this was found, test-selection recall goes 0.015 -> 0.809 (empty answers 59/60 ->
12/60), so it no longer counts as held-out evidence; on the held-in one it confirms
independently: recall 0.512 -> 0.541, precision 0.629 -> 0.670, source-caller recall
0.610 -> 0.636. Every other repository is unchanged. Engine suite 43/43, torture
unchanged; query cases 327/327 python, 264/264 typescript, 333/333 java; case
stub-beside-source-resolves-to-source fails on the base engine on the twin check.

Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
@swapnilpaliwal-sd

Copy link
Copy Markdown
Contributor Author

Held-out evaluation (six repositories never used for tuning; base engine vs this branch, every verb):

base this branch
test-selection recall 0.444 0.453
test-selection precision 0.429 0.434
every failing file selected 38.1% 53.6%
targets with an empty answer 39.4% 23.2%
impact source-caller recall 0.631 0.637
impact caller precision (callers that ran) 0.924 0.936
path found / context recall@5 0.600 / 0.894 unchanged

Honest reading: most of the empty-answer and safe-selection gain comes from one repository whose .pyi stubs sent every test call to the library boundary (recall 0.015 → 0.809); that fix was found on it, so it is not independent evidence — the independent confirmation is the held-in stub repository (recall 0.512 → 0.541). On the other five the gains are small (structlog recall 0.478 → 0.507, source-callers 0.386 → 0.432; tenacity precision +0.9pt) and two dip: one repository loses 2 true files through the on-demand __repr__ rule (its tests repr an object a factory built), another loses 14 through the main-guard rule — its tests run tutorial modules as scripts (subprocess … mod.__file__), so the guarded app() really runs; the old import walk was right by accident there, and the real route (a test that runs a file as a program by that spelling) is not recognised yet.

@swapnilpaliwal-sd

Copy link
Copy Markdown
Contributor Author

Superseded: rebuilt on the integration tip and landed through #1882 / #1885.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant