Optimise Patches::filter to better handle sparse masks and sparse patches - #10149
robert3005 wants to merge 1 commit into
Performance Regression: -8.03%
⚠️ Unknown Walltime execution environment detected
Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.
For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.
⚠️ Different runtime environments detected
Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.
⚡ 2 improved benchmarks
❌ 3 regressed benchmarks
✅ 2057 untouched benchmarks
⏩ 503 skipped benchmarks1
Warning
Please fix the performance issues or acknowledge them on CodSpeed.
Performance Changes
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | WallTime | mul_u64_nonnull_neon |
29.6 µs | 40.1 µs | -26.2% |
| ❌ | WallTime | multiply_shapes_neon[(32768, PerRowPerRow)] |
33.3 µs | 39.1 µs | -15.02% |
| ❌ | WallTime | mul_i64_nonnull_neon |
33.5 µs | 39.1 µs | -14.52% |
| ⚡ | WallTime | dict_canonicalize_gt_u8_neon[1000000] |
547 µs | 493.4 µs | +10.86% |
| ⚡ | WallTime | scalar_subtract_neon |
13.9 µs | 12.6 µs | +10.73% |
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing rk/patchesfilter (1d66d12) with develop (a96e0a8)
Footnotes
-
503 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩