Remove remaining array and Python overhead - #52
Merged
Conversation
📊 Benchmark Comparison Results
➖ 27 benchmarks within noise (click to expand)
Legend: ✅ Faster (>5%) · ⛔ This PR has severe performance regressions (>20% slower). Please investigate before merging. |
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: e7b02886-a2c5-43cd-bb97-bfa037d3b625
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: e7b02886-a2c5-43cd-bb97-bfa037d3b625
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: e7b02886-a2c5-43cd-bb97-bfa037d3b625
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: e7b02886-a2c5-43cd-bb97-bfa037d3b625
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: e7b02886-a2c5-43cd-bb97-bfa037d3b625
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: e7b02886-a2c5-43cd-bb97-bfa037d3b625
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: e7b02886-a2c5-43cd-bb97-bfa037d3b625
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: e7b02886-a2c5-43cd-bb97-bfa037d3b625
mrecachinas
force-pushed
the
perf/remaining-opportunities
branch
from
July 10, 2026 13:54
4c372f0 to
4de9156
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
The first performance pass made the individual Hamming-distance kernels very fast. Follow-up profiling showed that batch scans and Python wrappers were still spending meaningful time on repeated setup, thread scheduling, and buffer handling.
This PR removes that remaining overhead without changing the public API or results.
What changed
bytearray,memoryview, and other contiguous buffers use a stack-pinned buffer guard instead of boxedPyBuffersetup.Representative results
These comparisons use a same-state
mainbaseline to avoid CPU-frequency differences between benchmark sessions.mainallfirst, match lastbestallallbestallAdditional isolated measurements show 64-byte
bytearraycalls improving 38.2%, 1 MiB hex calls improving 26.0% by avoiding copies, and the 4 KiB byte boundary improving 26.5%.Safety and compatibility
Py_bufferis never moved after acquisition.Testing
cargo test --no-default-features(74 unit tests, 2 doctests)cargo fmt --checkruff check .ruff format --check .python -m pytest -q .(140 passed, 3 skipped)Follow-up