Avatar for the Eventual-Inc user
Eventual-Inc
Daft
BlogDocsChangelog

Performance History

Latest Results

fix(iceberg): apply partition predicates to count pushdown (#7421) ## Changes Made Fixes #7420: `count_rows()` on an identity-partitioned Iceberg table returned the whole table's row count when the only predicate was on the partition column. ```python # dt='a' has 3 rows, dt='b' has 5 rows daft.read_iceberg(table).where(col("dt") == "a").count_rows() # was 8, now 3 ``` ### Why it happened `PushDownFilter` splits a predicate into three groups (`PredicateGroups` in `src/daft-scan/src/expr_rewriter.rs`). A pure partition predicate with an identity transform lands in `partition_only_filter`, which is "applied directly to partition values and can be dropped from the data-level filter" — so it moves into `pushdowns.partition_filters` and `pushdowns.filters` becomes `None`. The count-pushdown guard in `push_down_aggregation.rs` only inspects `filters`: ```rust external_info.pushdowns.filters.is_none() ``` so it reads that as "no filter at all" and pushes the count into the source. `IcebergDataSource._create_count_tasks` then summed `record_count` over every data file, applying neither filter — unlike `_create_regular_tasks` in the same class, which passes `row_filter=convert_row_filter(...)` and prunes each file with `pspec.filter(ExpressionsProjection([pushdowns.partition_filters]))`. ### The fix `_create_count_tasks` now prunes files against `partition_filters` the same way the regular path does, and falls back to a regular scan when `pushdowns.filters` is set, since a row-level predicate cannot be answered from record counts — a file surviving metadata pruning may still hold rows that do not match. This keeps the optimization for the query shape it is most valuable for. The count stays metadata-only; it is simply computed over the surviving partitions: ``` INFO daft.io.iceberg.iceberg_scan: Using Iceberg count pushdown optimization for count mode: All INFO daft.io.iceberg.iceberg_scan: Created Iceberg count pushdown task with total_count=3 for field=x ``` Partition pruning is exact for the predicates that reach this path: they are identity-transform predicates the optimizer already proved resolvable from partition values alone, so every row of a surviving file matches. I considered tightening the Rust guard to also require `partition_filters.is_none()`. It fixes the wrong result, but it disables count pushdown for partition-filtered counts, which is the case the metadata-only path exists to serve. Instead, `DataSource.supports_count_pushdown` now documents the contract: a source that absorbs a count must honor the rest of the pushdowns, and `partition_filters` can be set while `filters` is `None`. ### Scope `GlobScanOperator` also reports `supports_count_pushdown()` for Parquet, but it has no separate count path — its count comes from scan tasks that were already pruned by `partition_filters`, so hive-partitioned Parquet was never affected. Verified: `read_parquet` with `hive_partitioning=True` returns the correct count before and after this change. Iceberg was the only source with its own count path. ## Testing New tests in `tests/io/iceberg/test_iceberg_reads.py` against the local SqlCatalog: - a partition predicate on an identity-partitioned column, checked against the materialized row count - no predicate, counting every partition - a data-column predicate, which must not be answered from record counts - partition and data predicates combined - a partition predicate matching no partition The first and last fail on `main` (8 instead of 3, 8 instead of 0) and pass with this change. `tests/io/` and `tests/catalog` pass (1325 passed, 81 skipped). One pre-existing environment error in `tests/io/test_s3_credentials_refresh.py` from a fixture that cannot start its local mock server; unrelated to this change. ## Related Issues Closes #7420
main
2 hours ago
chore: retrigger CI (codspeed runner apt-get network failure)
hello-peter-tang:perf/6340-lm-pipelined-predicate-eval
14 hours ago
Merge branch 'main' into perf/dense-integer-argsort
GENG-CHOVYYYY:perf/dense-integer-argsort
19 hours ago

Latest Branches

CodSpeed Performance Gauge
0%
perf(parquet): LM-pipelined predicate evaluation with stats-based group ordering#7488
14 hours ago
6004ace
hello-peter-tang:perf/6340-lm-pipelined-predicate-eval
CodSpeed Performance Gauge
0%
20 hours ago
d7351d9
GENG-CHOVYYYY:perf/dense-integer-argsort
CodSpeed Performance Gauge
0%
22 hours ago
c12c0e4
djouallah:fix/azure-uri-fabric-host
© 2026 CodSpeed Technology
Home Terms Privacy Docs