Avatar for the Eventual-Inc user
Eventual-Inc
Daft
BlogDocsChangelog

Performance History

Latest Results

feat(io): support ignoring corrupt ORC files Add opt-in schema fallback and conservative ORC error filtering. Report skipped files through existing execution statistics while preserving emitted batches and unrelated error propagation. Signed-off-by: jiangxt2 <jiangxt2@vip.qq.com>
jiangxt2:orc-ignore-corrupt-files
1 hour ago
fix(sql): use signed 64-bit Iceberg snapshot IDs
Solomon-mithra:fix/sql-iceberg-signed-snapshot-id
15 hours ago
feat(io): add basic ORC file reading support (#7576) ## Changes Made Add `daft.read_orc()` for reading raw ORC files into a Daft DataFrame, allowing ORC inputs to be processed directly without a separate format conversion step. The reader uses the existing Python `DataSource` / `DataSourceTask` APIs, Daft file I/O, and PyArrow's ORC Dataset scanner. - Support file paths, directories, glob patterns, and path lists. Search directories recursively for `*.orc` and deduplicate overlapping inputs. - Reuse `IOConfig`, including the planning context's default configuration, for local and remote file access. - Read rows lazily in execution tasks, using one task per file and configurable batch sizes. - Infer the schema from the first matched file. Align subsequent files by field name, fill missing fields with nulls, exclude extra fields, and apply Daft's existing conversion rules. - Support column projection while retaining fields required by filters. Filters and limits use Daft's existing execution operators. - Add public exports, unit and integration tests, API documentation, and an ORC usage guide. Full schema merging, stripe-level task splitting, and ORC-native predicate pruning are outside this PR. ### Example ```python import daft df = daft.read_orc("./data/*.orc", batch_size=65536) df.select("id", "label").show() ``` ### Testing Tests cover path handling, schema alignment, projection/filter interactions, batching, serialization, and resource cleanup. - Native ORC suite: **76 passed**. - Ray ORC suite with xdist and coverage: **76 passed**. - HTTP/S3-compatible integration suite: **9 passed**. - PyArrow 16.0.0 compatibility suite: **33 passed**. - Project pre-commit checks and MkDocs build passed. ### AI assistance Codex and Claude assisted with development and review. ## Related Issues Closes #7575 --------- Signed-off-by: jiangxt2 <jiangxt2@vip.qq.com>
main
18 hours ago

Latest Branches

CodSpeed Performance Gauge
0%
feat(io): support ignoring corrupt ORC files#7600
2 hours ago
5a50930
jiangxt2:orc-ignore-corrupt-files
CodSpeed Performance Gauge
0%
2 hours ago
989096d
jiangxt2:orc-hive-partitioning
CodSpeed Performance Gauge
0%
3 hours ago
40a24d2
jiangxt2:orc-file-path-column
© 2026 CodSpeed Technology
Home Terms Privacy Docs