Skip to content

Performance: cache ClassDef.ancestors() transitive walk - #3048

Open
Pierre-Sassoulas wants to merge 3 commits into
mainfrom
perf/cache-classdef-ancestors
Open

Performance: cache ClassDef.ancestors() transitive walk#3048
Pierre-Sassoulas wants to merge 3 commits into
mainfrom
perf/cache-classdef-ancestors

Conversation

@Pierre-Sassoulas

@Pierre-Sassoulas Pierre-Sassoulas commented May 10, 2026

Copy link
Copy Markdown
Member

Type of Changes

Type
🔨 Refactoring

Description

Split out of the original four-optimization PR, as requested — this one now contains only the ClassDef.ancestors() cache, in a single commit with its tests.

The recursive walk in ancestors(recurs=True) re-resolved shared base classes on every call, amplifying cost on deep MRO chains. Cache the materialized tuple on the instance so each ClassDef pays for its ancestors once, and the cache dies with the instance when the manager drops the AST.

context is intentionally not part of the key: the result is path-independent and the walk's own yielded set handles cycle prevention.

TestAncestorsCaching covers the corner cases: a cache hit reuses the materialized tuple, recurs=False bypasses the cache entirely, a cyclic hierarchy (string.Template = A) unwinds through the _COMPUTING_ANCESTORS sentinel instead of recursing forever, and an exception mid-walk clears the sentinel so a transient failure cannot poison the cache.

Measured on pandas/core/frame.py (interleaved A/B, n=4):

baseline 21.34s ± 0.18  ->  patched 20.48s ± 0.13   (-4.0%)

Cache hit rate on the same run: 98% (66k hits / 1.2k misses on 1k distinct ClassDefs). Larger speedups expected on codebases with deeper MROs (SQLAlchemy/Pydantic-shaped projects).

The other three optimizations are now standalone PRs, each reviewable independently:

Note that CodSpeed's earlier +9.65% on this PR was the combined figure for all four; taken alone this change is the -4.0% above.

Closes #1115

@Pierre-Sassoulas Pierre-Sassoulas added this to the 4.2.0 milestone May 10, 2026
@Pierre-Sassoulas
Pierre-Sassoulas force-pushed the perf/cache-classdef-ancestors branch 4 times, most recently from 4160de2 to eba3439 Compare May 18, 2026 20:09
@codecov

codecov Bot commented May 18, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 93.70%. Comparing base (da4a8cf) to head (6089a01).

Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main    #3048      +/-   ##
==========================================
+ Coverage   93.67%   93.70%   +0.03%     
==========================================
  Files          93       93              
  Lines       11645    11675      +30     
==========================================
+ Hits        10908    10940      +32     
+ Misses        737      735       -2     
Flag Coverage Δ
linux 93.55% <100.00%> (+0.02%) ⬆️
pypy 93.70% <100.00%> (+0.03%) ⬆️
windows 93.67% <100.00%> (+0.03%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
astroid/nodes/scoped_nodes/scoped_nodes.py 93.58% <100.00%> (+0.22%) ⬆️

... and 3 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@Pierre-Sassoulas
Pierre-Sassoulas force-pushed the perf/cache-classdef-ancestors branch from eba3439 to d7a22a7 Compare May 24, 2026 12:06
@Pierre-Sassoulas
Pierre-Sassoulas changed the base branch from main to codspeed-wizard-1774989768270 May 24, 2026 12:06
@codspeed-hq

codspeed-hq Bot commented May 24, 2026

Copy link
Copy Markdown

Merging this PR will improve performance by 2.7%

⚠️ Different runtime environments detected

Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.

Open the report in CodSpeed to investigate

⚡ 2 improved benchmarks
✅ 1 untouched benchmark
⏩ 1 skipped benchmark1

Performance Changes

Mode Benchmark BASE HEAD Efficiency
Simulation test_bench_endtoend_walk_infer_black 35.7 s 34.6 s +3%
Simulation test_bench_endtoend_walk_infer_flask 23.1 s 22.6 s +2.4%

Tip

Curious why performance improved? Comment @codspeedbot explain why performance improved on this PR, or directly use the CodSpeed MCP with your agent.


Comparing perf/cache-classdef-ancestors (6089a01) with main (da4a8cf)

Open in CodSpeed

Footnotes

  1. 1 benchmark was skipped, so the baseline result was used instead. If it was deleted from the codebase, click here and archive it to remove it from the performance reports.

@Pierre-Sassoulas
Pierre-Sassoulas force-pushed the codspeed-wizard-1774989768270 branch 2 times, most recently from 2d391a6 to 2ef6d50 Compare May 24, 2026 13:23
@Pierre-Sassoulas
Pierre-Sassoulas force-pushed the perf/cache-classdef-ancestors branch from d7a22a7 to 2718c3e Compare May 24, 2026 14:34
Pierre-Sassoulas added a commit that referenced this pull request May 24, 2026
…) bypass

Add targeted regression coverage for the four perf commits in #3048
so the corner-case branches don't silently regress.

scoped_nodes.py — ClassDef.ancestors() cache (TestAncestorsCaching):
  * cache hit reuses the materialized tuple across calls
  * recurs=False bypasses the cache entirely
  * cyclic class hierarchy (string.Template = A) unwinds via the
    _COMPUTING_ANCESTORS sentinel without infinite recursion
  * exception during the walk clears the sentinel so the cache is
    not poisoned (covers the except BaseException cleanup path)

scoped_nodes.py — ClassDef._find_metaclass() cache (TestFindMetaclassCaching):
  * cached result reused on second no-context call
  * None result is cached so the MRO walk runs once
  * explicit context= argument bypasses the cache
  * re-entry through the _COMPUTING_METACLASS sentinel returns None
    (covers the cycle-break branch)

brain_builtin_inference.py — single Call dispatcher (TestBuiltinDispatcher):
  * known builtin Name dispatches (bool, dict.fromkeys via Attribute)
  * unknown Name, non-dict Attribute, dynamic call target are skipped
  * re.Pattern = type(...) / re.Match = type(...) still excluded so
    brain_re keeps owning their inference
  * register_builtin_transform populates _BUILTIN_INFERENCE_FUNCS

context.py — InferenceContext.clone() bypass (new tests/test_context.py):
  * path is an independent deep copy; mutations don't leak back
  * lookupname is reset to None on the clone
  * _nodes_inferred is shared (mutable counter preserved across the
    clone family — required for the max_inferred budget)
  * callcontext / boundnode / extra_context propagated by identity
  * constraints is shallow-copied
  * clone() bypasses __init__ (tracked via temporary patch)
  * end-to-end inference still resolves through a cloned context

Closes the two PR #3048 coverage gaps reported by Codecov
(scoped_nodes.py:2218-2220 and :2757).
@Pierre-Sassoulas
Pierre-Sassoulas force-pushed the codspeed-wizard-1774989768270 branch 2 times, most recently from 2fb0a52 to 7e482f9 Compare June 6, 2026 05:51
Base automatically changed from codspeed-wizard-1774989768270 to main June 6, 2026 21:41
Pierre-Sassoulas added a commit that referenced this pull request Jun 6, 2026
…) bypass

Add targeted regression coverage for the four perf commits in #3048
so the corner-case branches don't silently regress.

scoped_nodes.py — ClassDef.ancestors() cache (TestAncestorsCaching):
  * cache hit reuses the materialized tuple across calls
  * recurs=False bypasses the cache entirely
  * cyclic class hierarchy (string.Template = A) unwinds via the
    _COMPUTING_ANCESTORS sentinel without infinite recursion
  * exception during the walk clears the sentinel so the cache is
    not poisoned (covers the except BaseException cleanup path)

scoped_nodes.py — ClassDef._find_metaclass() cache (TestFindMetaclassCaching):
  * cached result reused on second no-context call
  * None result is cached so the MRO walk runs once
  * explicit context= argument bypasses the cache
  * re-entry through the _COMPUTING_METACLASS sentinel returns None
    (covers the cycle-break branch)

brain_builtin_inference.py — single Call dispatcher (TestBuiltinDispatcher):
  * known builtin Name dispatches (bool, dict.fromkeys via Attribute)
  * unknown Name, non-dict Attribute, dynamic call target are skipped
  * re.Pattern = type(...) / re.Match = type(...) still excluded so
    brain_re keeps owning their inference
  * register_builtin_transform populates _BUILTIN_INFERENCE_FUNCS

context.py — InferenceContext.clone() bypass (new tests/test_context.py):
  * path is an independent deep copy; mutations don't leak back
  * lookupname is reset to None on the clone
  * _nodes_inferred is shared (mutable counter preserved across the
    clone family — required for the max_inferred budget)
  * callcontext / boundnode / extra_context propagated by identity
  * constraints is shallow-copied
  * clone() bypasses __init__ (tracked via temporary patch)
  * end-to-end inference still resolves through a cloned context

Closes the two PR #3048 coverage gaps reported by Codecov
(scoped_nodes.py:2218-2220 and :2757).
@Pierre-Sassoulas
Pierre-Sassoulas force-pushed the perf/cache-classdef-ancestors branch from ea6b8c5 to 4f1ff78 Compare June 6, 2026 21:50
@Pierre-Sassoulas
Pierre-Sassoulas force-pushed the perf/cache-classdef-ancestors branch from 4f1ff78 to c009ecd Compare July 4, 2026 12:42
Pierre-Sassoulas added a commit that referenced this pull request Jul 4, 2026
…) bypass

Add targeted regression coverage for the four perf commits in #3048
so the corner-case branches don't silently regress.

scoped_nodes.py — ClassDef.ancestors() cache (TestAncestorsCaching):
  * cache hit reuses the materialized tuple across calls
  * recurs=False bypasses the cache entirely
  * cyclic class hierarchy (string.Template = A) unwinds via the
    _COMPUTING_ANCESTORS sentinel without infinite recursion
  * exception during the walk clears the sentinel so the cache is
    not poisoned (covers the except BaseException cleanup path)

scoped_nodes.py — ClassDef._find_metaclass() cache (TestFindMetaclassCaching):
  * cached result reused on second no-context call
  * None result is cached so the MRO walk runs once
  * explicit context= argument bypasses the cache
  * re-entry through the _COMPUTING_METACLASS sentinel returns None
    (covers the cycle-break branch)

brain_builtin_inference.py — single Call dispatcher (TestBuiltinDispatcher):
  * known builtin Name dispatches (bool, dict.fromkeys via Attribute)
  * unknown Name, non-dict Attribute, dynamic call target are skipped
  * re.Pattern = type(...) / re.Match = type(...) still excluded so
    brain_re keeps owning their inference
  * register_builtin_transform populates _BUILTIN_INFERENCE_FUNCS

context.py — InferenceContext.clone() bypass (new tests/test_context.py):
  * path is an independent deep copy; mutations don't leak back
  * lookupname is reset to None on the clone
  * _nodes_inferred is shared (mutable counter preserved across the
    clone family — required for the max_inferred budget)
  * callcontext / boundnode / extra_context propagated by identity
  * constraints is shallow-copied
  * clone() bypasses __init__ (tracked via temporary patch)
  * end-to-end inference still resolves through a cloned context

Closes the two PR #3048 coverage gaps reported by Codecov
(scoped_nodes.py:2218-2220 and :2757).
Pierre-Sassoulas added a commit that referenced this pull request Jul 4, 2026
The clone() fast path no longer goes through __init__, which left
the explicit nodes_inferred branch (context.py:56) uncovered — the
one genuine coverage loss Codecov reports on #3048. Pin the
__init__ contract for external callers: a passed cell is adopted
by identity, the default is a fresh zeroed cell per context.
@Pierre-Sassoulas
Pierre-Sassoulas force-pushed the perf/cache-classdef-ancestors branch from c009ecd to 8116a9e Compare July 5, 2026 07:10
Pierre-Sassoulas added a commit that referenced this pull request Jul 5, 2026
…) bypass

Add targeted regression coverage for the four perf commits in #3048
so the corner-case branches don't silently regress.

scoped_nodes.py — ClassDef.ancestors() cache (TestAncestorsCaching):
  * cache hit reuses the materialized tuple across calls
  * recurs=False bypasses the cache entirely
  * cyclic class hierarchy (string.Template = A) unwinds via the
    _COMPUTING_ANCESTORS sentinel without infinite recursion
  * exception during the walk clears the sentinel so the cache is
    not poisoned (covers the except BaseException cleanup path)

scoped_nodes.py — ClassDef._find_metaclass() cache (TestFindMetaclassCaching):
  * cached result reused on second no-context call
  * None result is cached so the MRO walk runs once
  * explicit context= argument bypasses the cache
  * re-entry through the _COMPUTING_METACLASS sentinel returns None
    (covers the cycle-break branch)

brain_builtin_inference.py — single Call dispatcher (TestBuiltinDispatcher):
  * known builtin Name dispatches (bool, dict.fromkeys via Attribute)
  * unknown Name, non-dict Attribute, dynamic call target are skipped
  * re.Pattern = type(...) / re.Match = type(...) still excluded so
    brain_re keeps owning their inference
  * register_builtin_transform populates _BUILTIN_INFERENCE_FUNCS

context.py — InferenceContext.clone() bypass (new tests/test_context.py):
  * path is an independent deep copy; mutations don't leak back
  * lookupname is reset to None on the clone
  * _nodes_inferred is shared (mutable counter preserved across the
    clone family — required for the max_inferred budget)
  * callcontext / boundnode / extra_context propagated by identity
  * constraints is shallow-copied
  * clone() bypasses __init__ (tracked via temporary patch)
  * end-to-end inference still resolves through a cloned context

Closes the two PR #3048 coverage gaps reported by Codecov
(scoped_nodes.py:2218-2220 and :2757).
Pierre-Sassoulas added a commit that referenced this pull request Jul 5, 2026
The clone() fast path no longer goes through __init__, which left
the explicit nodes_inferred branch (context.py:56) uncovered — the
one genuine coverage loss Codecov reports on #3048. Pin the
__init__ contract for external callers: a passed cell is adopted
by identity, the default is a fresh zeroed cell per context.
@Pierre-Sassoulas
Pierre-Sassoulas force-pushed the perf/cache-classdef-ancestors branch from 8116a9e to a064244 Compare July 29, 2026 21:34
Pierre-Sassoulas added a commit that referenced this pull request Jul 29, 2026
…) bypass

Add targeted regression coverage for the four perf commits in #3048
so the corner-case branches don't silently regress.

scoped_nodes.py — ClassDef.ancestors() cache (TestAncestorsCaching):
  * cache hit reuses the materialized tuple across calls
  * recurs=False bypasses the cache entirely
  * cyclic class hierarchy (string.Template = A) unwinds via the
    _COMPUTING_ANCESTORS sentinel without infinite recursion
  * exception during the walk clears the sentinel so the cache is
    not poisoned (covers the except BaseException cleanup path)

scoped_nodes.py — ClassDef._find_metaclass() cache (TestFindMetaclassCaching):
  * cached result reused on second no-context call
  * None result is cached so the MRO walk runs once
  * explicit context= argument bypasses the cache
  * re-entry through the _COMPUTING_METACLASS sentinel returns None
    (covers the cycle-break branch)

brain_builtin_inference.py — single Call dispatcher (TestBuiltinDispatcher):
  * known builtin Name dispatches (bool, dict.fromkeys via Attribute)
  * unknown Name, non-dict Attribute, dynamic call target are skipped
  * re.Pattern = type(...) / re.Match = type(...) still excluded so
    brain_re keeps owning their inference
  * register_builtin_transform populates _BUILTIN_INFERENCE_FUNCS

context.py — InferenceContext.clone() bypass (new tests/test_context.py):
  * path is an independent deep copy; mutations don't leak back
  * lookupname is reset to None on the clone
  * _nodes_inferred is shared (mutable counter preserved across the
    clone family — required for the max_inferred budget)
  * callcontext / boundnode / extra_context propagated by identity
  * constraints is shallow-copied
  * clone() bypasses __init__ (tracked via temporary patch)
  * end-to-end inference still resolves through a cloned context

Closes the two PR #3048 coverage gaps reported by Codecov
(scoped_nodes.py:2218-2220 and :2757).
Pierre-Sassoulas added a commit that referenced this pull request Jul 29, 2026
The clone() fast path no longer goes through __init__, which left
the explicit nodes_inferred branch (context.py:56) uncovered — the
one genuine coverage loss Codecov reports on #3048. Pin the
__init__ contract for external callers: a passed cell is adopted
by identity, the default is a fresh zeroed cell per context.
@Pierre-Sassoulas

Copy link
Copy Markdown
Member Author

@DanielNoord I'd like to merge this before release, great performance improvements, confirmed in the original issue #1115 (comment), and it's been ready for a long time.

@DanielNoord DanielNoord left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would it be possible to split out the individual performance improvements into separate PRs? This seems to include 4.
With an upcoming long holiday weekend I don't have a lot of time for reviews. Context switching to a bigger PR is probably something I wouldn't get to before the end of next week. I can probably make some time for smaller PRs in between meetings tomorrow :)

The recursive walk in ``ancestors(recurs=True)`` re-resolved shared base
classes on every call, amplifying cost on deep MRO chains. Cache the
materialized tuple as a ``cached_property`` so each ClassDef pays for
its ancestors once, and the cache dies with the instance when the
manager drops the AST.

``context`` is intentionally not part of the key — the result is
path-independent and the walk's own ``yielded`` set handles cycle
prevention.

``TestAncestorsCaching`` covers the corner cases: a cache hit reuses the
materialized tuple, ``recurs=False`` bypasses the cache entirely, a
cyclic hierarchy (``string.Template = A``) unwinds through the
``_COMPUTING_ANCESTORS`` sentinel instead of recursing forever, and an
exception mid-walk clears the sentinel so a transient failure cannot
poison the cache.

Measured on pandas/core/frame.py (interleaved A/B, n=4):
  baseline 21.34s ± 0.18  ->  patched 20.48s ± 0.13   (-4.0%)

Cache hit rate on the same run: 98% (66k hits / 1.2k misses on 1k
distinct ClassDefs). Larger speedups expected on codebases with
deeper MROs (SQLAlchemy/Pydantic-shaped projects).

Closes #1115
@Pierre-Sassoulas

Copy link
Copy Markdown
Member Author

Done, there's this one still, and now #3166, #3167 and #3168

@DanielNoord DanielNoord left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why do we go via __dict__ and not just set this as an attribute? Pehaps with weak links

@Pierre-Sassoulas Pierre-Sassoulas modified the milestones: 4.2.0, 4.4.0 Aug 3, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Astroid calls to ancestors are uncached and slow for templates and generics in ClassDef.ancestors

2 participants