Avatar for the Eventual-Inc user
Eventual-Inc
Daft
BlogDocsChangelog

Performance History

Latest Results

fix(parquet): apply positional deletes in count-only fast path (#7560) A read that projects zero file columns takes the count_only_stream fast path, which summed physical row-group sizes and ignored delete_rows, start_offset, and num_rows. On Iceberg merge-on-read tables this over-counted by including positionally-deleted rows. Derive the count-only total from build_base_selections -- the same selection machinery the decode path uses -- so offset, positional deletes, and the limit are applied consistently across both paths. Adds a regression test (200 rows over 4 row groups, 40 positional deletes -> 160 live rows) that failed 200 != 160 before the fix.
hello-peter-tang:fix/7560-count-only-positional-deletes
3 hours ago
fix(sql): use signed 64-bit Iceberg snapshot IDs
Solomon-mithra:fix/sql-iceberg-signed-snapshot-id
8 hours ago
feat(io): add basic ORC file reading support (#7576) ## Changes Made Add `daft.read_orc()` for reading raw ORC files into a Daft DataFrame, allowing ORC inputs to be processed directly without a separate format conversion step. The reader uses the existing Python `DataSource` / `DataSourceTask` APIs, Daft file I/O, and PyArrow's ORC Dataset scanner. - Support file paths, directories, glob patterns, and path lists. Search directories recursively for `*.orc` and deduplicate overlapping inputs. - Reuse `IOConfig`, including the planning context's default configuration, for local and remote file access. - Read rows lazily in execution tasks, using one task per file and configurable batch sizes. - Infer the schema from the first matched file. Align subsequent files by field name, fill missing fields with nulls, exclude extra fields, and apply Daft's existing conversion rules. - Support column projection while retaining fields required by filters. Filters and limits use Daft's existing execution operators. - Add public exports, unit and integration tests, API documentation, and an ORC usage guide. Full schema merging, stripe-level task splitting, and ORC-native predicate pruning are outside this PR. ### Example ```python import daft df = daft.read_orc("./data/*.orc", batch_size=65536) df.select("id", "label").show() ``` ### Testing Tests cover path handling, schema alignment, projection/filter interactions, batching, serialization, and resource cleanup. - Native ORC suite: **76 passed**. - Ray ORC suite with xdist and coverage: **76 passed**. - HTTP/S3-compatible integration suite: **9 passed**. - PyArrow 16.0.0 compatibility suite: **33 passed**. - Project pre-commit checks and MkDocs build passed. ### AI assistance Codex and Claude assisted with development and review. ## Related Issues Closes #7575 --------- Signed-off-by: jiangxt2 <jiangxt2@vip.qq.com>
main
11 hours ago
Merge branch 'main' into fix/bisect-round-robin-bundles
DogerW666:fix/bisect-round-robin-bundles
21 hours ago
fix(optimizer): preserve input order observed by batch UDFs
hello-peter-tang:eliminate-redundant-sort
21 hours ago

Latest Branches

CodSpeed Performance Gauge
0%
fix(parquet): apply positional deletes in count-only fast path#7591
18 hours ago
bf2e4c0
hello-peter-tang:fix/7560-count-only-positional-deletes
CodSpeed Performance Gauge
0%
9 hours ago
7206f26
Solomon-mithra:fix/sql-iceberg-signed-snapshot-id
CodSpeed Performance Gauge
0%
20 hours ago
d23836e
hello-peter-tang:perf/6340-lm-pipelined-predicate-eval
© 2026 CodSpeed Technology
Home Terms Privacy Docs