Conversation
- Integrate dev's scan cache/preload: file_sizes, metadata_byte_sizes, footer_offsets, rebind(), get_file_size(), get_metadata_byte_size(), get_footer_offset() - Byte ranges in compute_task include header and footer for zero file I/O on preload - host_parquet_representation: add file_size parameter, keep filter_expression and pure_filter_ids for filter pushdown - Partition reserved_compressed_bytes includes metadata size for cache - Keep row_group_pruning partition structure (claim_next_rg_partition, rg_indices) Made-with: Cursor
- Do not set filter on _reader_options (materialization reads all rows in
selected row groups)
- Pass empty {} to make_selected_column_indices so no pure filter columns
- Use temporary pruning_options with filter only for filter_row_groups_with_stats
- Do not pass filter_expression or pure_filter_ids to host_parquet_representation
(defaults nullptr and {})
Made-with: Cursor
joosthooz
force-pushed
the
row_group_pruning
branch
from
March 13, 2026 10:14
62c8842 to
140b4b4
Compare
- pixi.toml: switch to rapidsai channel and libcudf 26.02.* - host_parquet_representation_converters: add 26.02 path for materialize_all_columns (vector<device_buffer>, 4-arg API) alongside 26.04+ path (host_span<device_span>, 5-arg with mr_ref) Made-with: Cursor
- Include filter column names in pruning_options so libcudf can resolve the filter to column statistics when projecting (reader_options had only projected columns). - parquet_scan_task: store _filter_column_names and _pruning_column_names; use set_columns/set_column_names for 26.02/26.04 in reader and pruning. - Guard parquet_io_utils.hpp include for 26.04+. - Tests: adjust for pruning-only (validate partition count, relax decimal pruning expectation, fix filter-on-non-projected-column test). Made-with: Cursor
Collaborator
|
#363 is finally ready. But I don't have access to a good machine for performance evaluation. |
joosthooz
marked this pull request as draft
March 16, 2026 10:58
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Shamelessly building on Kevin's PR #363 but using libcudf nightly build and cucascade with this patch NVIDIA/cuCascade#88