Latest Results
Measure numeric arithmetic and comparison benchmarks per CPU feature (#9599)
## Summary
Sweeps `#[cpu_features]` across the microbenchmarks that clearly earn
it: the binary numeric arithmetic and comparison kernels. Those are
portable lane loops ā the source is identical on every target and the
vector width the compiler picks comes from the build flags ā which is
the case the attribute exists for. Measuring them in simulation under
one fixed `+avx2` build hides the only variable that matters.
Middle of a three-PR stack: #9598, then this PR, then #9587. It carries
no kernel changes of its own ā everything here is a benchmark attribute
ā so it can be reordered or rebased onto `develop` without touching the
other two.
## Changes
Tagged:
- `binary_ops`: the primitive arithmetic cases (`add_*`, `subtract_*`,
`multiply_*`, `mul_*`, `div_i64_*`, `sub_i64_constant`, and the three
`*_shapes` matrices) and the two primitive comparison cases
(`eq_i64_constant`, `lt_i64_nullable`).
- `compare`: `compare_int`, `compare_int_nullable`,
`compare_int_constant`, `compare_int_eq`, `compare_float`.
- `scalar_subtract`.
- `lane_kernels`: `lanezip_checked_add_u32` and its
`arrow_checked_add_u32` baseline. The baseline is tagged too ā comparing
the two is only meaningful under the same build flags.
Left in simulation, with the reasoning recorded in each file's module
docs:
- Decimal arithmetic and comparison: `i128` widening and per-lane
rescaling, not something a wider vector register decides.
- Boolean `and`/`or`: already word-at-a-time over a bitmap.
- String and struct comparison: dominated by view chasing and per-field
dispatch.
- Casts in `lane_kernels`: vectorization-sensitive, but out of scope
here.
Also left alone: the `between` benchmarks in `vortex-fastlanes`
(`new_raw_prim_test_between` is a raw-primitive comparison kernel and
does qualify) and the bit-packed comparison matrices. Both are `types
=`/`consts =` parameterized, which `#[cpu_features]` has no coverage for
yet, and both would fan out to dozens of walltime series. Worth a
follow-up rather than a guess in this PR.
Note that tagging moves a benchmark out of the sharded simulation job,
so these series restart on the walltime legs instead of continuing their
simulation history.
Signed-off-by: Connor Tsui <connor.tsui20@gmail.com> Benchmark primitive comparison across CPU features (#9598)
## Summary
The measurement half of #9587, split out so the numbers land before the
kernels do.
Bottom of a three-PR stack: this PR, then #9599, then #9587. Tagging
these with `#[cpu_features]` first means the walltime legs record the
portable lane kernel's throughput on `avx2`, `avx512`, and `neon` metal
as a baseline series. The kernel PR then reports against it rather than
introducing both the benchmark and the thing it measures in one diff.
## Changes
- Adds the primitive comparison cases a hand-written SIMD kernel would
have to beat: constant on the left, `u8`, `u64`, and `f32`.
- Tags those four with `#[cpu_features]`, so each walltime leg measures
them under its own build flags instead of in simulation. The existing
cases are untouched and keep their simulation series.
- `bench_compare` now carries an `ItemsCount`, so the report reads as
throughput rather than a time that only means something next to another
run over the same array length.
- Adds module docs recording why these four are tagged and the rest are
not.
Signed-off-by: Connor Tsui <connor.tsui20@gmail.com> Latest Branches
-12%
renovate/anthropics-claude-code-action-digest +31%
ct/primitive-comparison-simd -12%
ct/cpu-features-numeric-benches Ā© 2026 CodSpeed Technology