Avatar for the Eventual-Inc user
Eventual-Inc
Daft
BlogDocsChangelog

Performance History

Latest Results

perf(optimizer): rewrite starts_with filters into range predicates Rework per review: instead of teaching the stats layer about starts_with, add a RewriteStartsWith optimizer rule that rewrites starts_with(c, P) to c >= P AND c < incr(P) inside Filter predicates, running right after expression simplification and before PushDownFilter. - The scan expression rewriter classifies any ScalarFn as a UDF and strands a lone starts_with filter in a residual Filter op; plain comparisons are data predicates, so they land in pushdowns.filters and every source can push them down natively (Parquet, CSV, Lance, Iceberg, ...). - Existing Utf8 min/max statistics pruning then skips row groups and scan tasks for free, with no starts_with knowledge needed in daft-stats. - Only Filter predicates are rewritten: projections keep the single, cheaper kernel call. - Arguments are bound by name (keyword-arg calls in any order are safe); an all-max-scalar prefix falls back to a lower-bound-only rewrite; empty-prefix and non-literal patterns are left untouched. Replaces the daft-stats/daft-parquet approach with optimizer plan-shape tests and a parquet e2e test asserting the pushed-down range predicate.
hello-peter-tang:startswith-stats-pruning
4 hours ago
fix(iceberg): apply partition predicates to count pushdown (#7421) ## Changes Made Fixes #7420: `count_rows()` on an identity-partitioned Iceberg table returned the whole table's row count when the only predicate was on the partition column. ```python # dt='a' has 3 rows, dt='b' has 5 rows daft.read_iceberg(table).where(col("dt") == "a").count_rows() # was 8, now 3 ``` ### Why it happened `PushDownFilter` splits a predicate into three groups (`PredicateGroups` in `src/daft-scan/src/expr_rewriter.rs`). A pure partition predicate with an identity transform lands in `partition_only_filter`, which is "applied directly to partition values and can be dropped from the data-level filter" — so it moves into `pushdowns.partition_filters` and `pushdowns.filters` becomes `None`. The count-pushdown guard in `push_down_aggregation.rs` only inspects `filters`: ```rust external_info.pushdowns.filters.is_none() ``` so it reads that as "no filter at all" and pushes the count into the source. `IcebergDataSource._create_count_tasks` then summed `record_count` over every data file, applying neither filter — unlike `_create_regular_tasks` in the same class, which passes `row_filter=convert_row_filter(...)` and prunes each file with `pspec.filter(ExpressionsProjection([pushdowns.partition_filters]))`. ### The fix `_create_count_tasks` now prunes files against `partition_filters` the same way the regular path does, and falls back to a regular scan when `pushdowns.filters` is set, since a row-level predicate cannot be answered from record counts — a file surviving metadata pruning may still hold rows that do not match. This keeps the optimization for the query shape it is most valuable for. The count stays metadata-only; it is simply computed over the surviving partitions: ``` INFO daft.io.iceberg.iceberg_scan: Using Iceberg count pushdown optimization for count mode: All INFO daft.io.iceberg.iceberg_scan: Created Iceberg count pushdown task with total_count=3 for field=x ``` Partition pruning is exact for the predicates that reach this path: they are identity-transform predicates the optimizer already proved resolvable from partition values alone, so every row of a surviving file matches. I considered tightening the Rust guard to also require `partition_filters.is_none()`. It fixes the wrong result, but it disables count pushdown for partition-filtered counts, which is the case the metadata-only path exists to serve. Instead, `DataSource.supports_count_pushdown` now documents the contract: a source that absorbs a count must honor the rest of the pushdowns, and `partition_filters` can be set while `filters` is `None`. ### Scope `GlobScanOperator` also reports `supports_count_pushdown()` for Parquet, but it has no separate count path — its count comes from scan tasks that were already pruned by `partition_filters`, so hive-partitioned Parquet was never affected. Verified: `read_parquet` with `hive_partitioning=True` returns the correct count before and after this change. Iceberg was the only source with its own count path. ## Testing New tests in `tests/io/iceberg/test_iceberg_reads.py` against the local SqlCatalog: - a partition predicate on an identity-partitioned column, checked against the materialized row count - no predicate, counting every partition - a data-column predicate, which must not be answered from record counts - partition and data predicates combined - a partition predicate matching no partition The first and last fail on `main` (8 instead of 3, 8 instead of 0) and pass with this change. `tests/io/` and `tests/catalog` pass (1325 passed, 81 skipped). One pre-existing environment error in `tests/io/test_s3_credentials_refresh.py` from a fixture that cannot start its local mock server; unrelated to this change. ## Related Issues Closes #7420
main
9 hours ago
chore: retrigger CI (codspeed runner apt-get network failure)
hello-peter-tang:perf/6340-lm-pipelined-predicate-eval
22 hours ago
Merge branch 'main' into perf/dense-integer-argsort
GENG-CHOVYYYY:perf/dense-integer-argsort
1 day ago

Latest Branches

CodSpeed Performance Gauge
0%
perf(optimizer): rewrite starts_with filters into range predicates for pushdown pruning#7441
17 hours ago
bae4e5c
hello-peter-tang:startswith-stats-pruning
CodSpeed Performance Gauge
0%
22 hours ago
6004ace
hello-peter-tang:perf/6340-lm-pipelined-predicate-eval
CodSpeed Performance Gauge
0%
1 day ago
d7351d9
GENG-CHOVYYYY:perf/dense-integer-argsort
© 2026 CodSpeed Technology
Home Terms Privacy Docs