Optimise Patches::filter to better handle sparse masks and sparse patches - #10149
robert3005 wants to merge 1 commit into
Conversation
Merging this PR will degrade performance by 8.03%
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | WallTime | mul_u64_nonnull_neon |
29.6 µs | 40.1 µs | -26.2% |
| ❌ | WallTime | multiply_shapes_neon[(32768, PerRowPerRow)] |
33.3 µs | 39.1 µs | -15.02% |
| ❌ | WallTime | mul_i64_nonnull_neon |
33.5 µs | 39.1 µs | -14.52% |
| ⚡ | WallTime | dict_canonicalize_gt_u8_neon[1000000] |
547 µs | 493.4 µs | +10.86% |
| ⚡ | WallTime | scalar_subtract_neon |
13.9 µs | 12.6 µs | +10.73% |
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing rk/patchesfilter (1d66d12) with develop (a96e0a8)
Footnotes
-
503 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
| /// The search compares values in the patch index type, so the patches do not need conversion. | ||
| struct PatchGallop<'a, T> { | ||
| patch_indices: &'a [T], | ||
| offset: usize, | ||
| patch_position: usize, | ||
| mask_position: usize, | ||
| new_patch_indices: &'a mut BufferMut<u64>, | ||
| kept_patches: &'a mut Vec<usize>, | ||
| } | ||
|
|
||
| impl<T: IntegerPType> PatchGallop<'_, T> { | ||
| fn next_mask_index(&mut self, mask_index: usize) -> Result<(), NoMoreMatches> { | ||
| // Mask indices are sorted, so no later mask index fits in `T` either. | ||
| let needle = mask_index | ||
| .checked_add(self.offset) | ||
| .and_then(<T as NumCast>::from) | ||
| .ok_or(NoMoreMatches)?; | ||
|
|
||
| self.patch_position = gallop_lower_bound(self.patch_indices, self.patch_position, &needle); | ||
| if self.patch_position == self.patch_indices.len() { |
We assumed dense case of mask over patches. Most of the time the mask or the
patches will be sparse