Add list_contains membership benchmarks - #10060
robert3005 wants to merge 5 commits into
Conversation
Merging this PR will improve performance by 14.33%
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ⚡ | WallTime | compare_u64_avx512 |
3.7 µs | 3.3 µs | +14.33% |
| 🆕 | Simulation | i64_random_chunked[256] |
N/A | 20.8 ms | N/A |
| 🆕 | Simulation | i64_random[256] |
N/A | 5.7 ms | N/A |
| 🆕 | Simulation | nested_list_random[32] |
N/A | 7.4 ms | N/A |
| 🆕 | Simulation | utf8_random[256] |
N/A | 14.4 ms | N/A |
Tip
Curious why performance improved? Comment @codspeedbot explain why performance improved on this PR, or directly use the CodSpeed MCP with your agent.
Comparing rk/list-contains-benchmarks (27f94c1) with develop (0852739)
Footnotes
-
503 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
8d837fb to
e71e6ad
Compare
e71e6ad to
cbc910b
Compare
cbc910b to
5b1960d
Compare
|
less than 26 benchmarks please can we have only 1 or 2? |
|
I think we can get it down to |
9522a70 to
7cae7b2
Compare
Benchmark constant-set membership before introducing prepared-set probes. Cover dense and random integers, UTF-8, decimals, nested lists, large sets, and flat and chunked needles using the same workloads for comparisons. Signed-off-by: Robert Kruszewski <github@robertk.io>
Signed-off-by: Robert Kruszewski <github@robertk.io>
Keep one benchmark per probe kind, integers, UTF-8 strings, and nested lists, each at a small and a large set, plus integer needles split into chunks, so that every benchmark stays under 1ms of CodSpeed simulation once constant sets are prepared as probes. Nested sets compare whole rows and stay smaller. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Robert Kruszewski <github@robertk.io>
7cae7b2 to
89feff4
Compare
Keep one set size per benchmark, four cases in all, so CodSpeed runs the sizes that show the probe cost rather than the setup cost. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Robert Kruszewski <github@robertk.io>
Summary
Add constant-set
list_containsbenchmarks before the prepared-set optimizations, so the same workloads can measure the existing comparison/OR implementation and the optimized implementation.Stack order: #10057 (SQL null semantics) → this PR → #10052 (membership optimizations) → #10053 (Java/JNI and Python bindings) → #10054 (DuckDB integration).
Changes
vortex-array/benches/list_contains_set.rsand its Cargo benchmark registration out of Optimize list_contains with prepared constant sets #10052.Validation:
cargo clippy -p vortex-array --bench list_contains_set --all-features -- -D warningspassed on this branch. Runtime tests and benchmarks were not executed.