Latest Results
fix(iceberg): apply partition predicates to count pushdown (#7421)
## Changes Made
Fixes #7420: `count_rows()` on an identity-partitioned Iceberg table
returned the whole table's row count when the only predicate was on the
partition column.
```python
# dt='a' has 3 rows, dt='b' has 5 rows
daft.read_iceberg(table).where(col("dt") == "a").count_rows() # was 8, now 3
```
### Why it happened
`PushDownFilter` splits a predicate into three groups (`PredicateGroups`
in `src/daft-scan/src/expr_rewriter.rs`). A pure partition predicate
with an identity transform lands in `partition_only_filter`, which is
"applied directly to partition values and can be dropped from the
data-level filter" — so it moves into `pushdowns.partition_filters` and
`pushdowns.filters` becomes `None`.
The count-pushdown guard in `push_down_aggregation.rs` only inspects
`filters`:
```rust
external_info.pushdowns.filters.is_none()
```
so it reads that as "no filter at all" and pushes the count into the
source. `IcebergDataSource._create_count_tasks` then summed
`record_count` over every data file, applying neither filter — unlike
`_create_regular_tasks` in the same class, which passes
`row_filter=convert_row_filter(...)` and prunes each file with
`pspec.filter(ExpressionsProjection([pushdowns.partition_filters]))`.
### The fix
`_create_count_tasks` now prunes files against `partition_filters` the
same way the regular path does, and falls back to a regular scan when
`pushdowns.filters` is set, since a row-level predicate cannot be
answered from record counts — a file surviving metadata pruning may
still hold rows that do not match.
This keeps the optimization for the query shape it is most valuable for.
The count stays metadata-only; it is simply computed over the surviving
partitions:
```
INFO daft.io.iceberg.iceberg_scan: Using Iceberg count pushdown optimization for count mode: All
INFO daft.io.iceberg.iceberg_scan: Created Iceberg count pushdown task with total_count=3 for field=x
```
Partition pruning is exact for the predicates that reach this path: they
are identity-transform predicates the optimizer already proved
resolvable from partition values alone, so every row of a surviving file
matches.
I considered tightening the Rust guard to also require
`partition_filters.is_none()`. It fixes the wrong result, but it
disables count pushdown for partition-filtered counts, which is the case
the metadata-only path exists to serve. Instead,
`DataSource.supports_count_pushdown` now documents the contract: a
source that absorbs a count must honor the rest of the pushdowns, and
`partition_filters` can be set while `filters` is `None`.
### Scope
`GlobScanOperator` also reports `supports_count_pushdown()` for Parquet,
but it has no separate count path — its count comes from scan tasks that
were already pruned by `partition_filters`, so hive-partitioned Parquet
was never affected. Verified: `read_parquet` with
`hive_partitioning=True` returns the correct count before and after this
change. Iceberg was the only source with its own count path.
## Testing
New tests in `tests/io/iceberg/test_iceberg_reads.py` against the local
SqlCatalog:
- a partition predicate on an identity-partitioned column, checked
against the materialized row count
- no predicate, counting every partition
- a data-column predicate, which must not be answered from record counts
- partition and data predicates combined
- a partition predicate matching no partition
The first and last fail on `main` (8 instead of 3, 8 instead of 0) and
pass with this change.
`tests/io/` and `tests/catalog` pass (1325 passed, 81 skipped). One
pre-existing environment error in
`tests/io/test_s3_credentials_refresh.py` from a fixture that cannot
start its local mock server; unrelated to this change.
## Related Issues
Closes #7420 Latest Branches
0%
hello-peter-tang:perf/6340-lm-pipelined-predicate-eval 0%
GENG-CHOVYYYY:perf/dense-integer-argsort 0%
djouallah:fix/azure-uri-fabric-host © 2026 CodSpeed Technology