[Parquet] Add writer option to skip bloom filters for column chunks whose data pages are all dictionary encoded - #10963
Open
ranflarion wants to merge 2 commits into
Conversation
…a pages are all dictionary encoded
etseidl
approved these changes
Sep 2, 2026
etseidl
left a comment
Contributor
There was a problem hiding this comment.
This seems reasonable, assuming one has access to the dictionary (which we may be adding here soon #10420).
It would be nice if we could also skip populating the bloom filter while a dictionary is in use. On fallback to non-dictionary, we could then add all the dict keys to the filter.
Comment on lines
+810
to
+811
| /// The dictionary page of such a chunk already lists every distinct value, so parquet-java | ||
| /// skips the bloom filter for it; set this to `false` to write files the same way. |
Contributor
There was a problem hiding this comment.
I don't think we need the comparison to parquet-java here. I'd instead say that since the dictionary contains all distinct values, the bloom filter is redundant and users might wish to save the space by setting this to false.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Which issue does this PR close?
Rationale for this change
A column chunk whose data pages are all dictionary encoded carries its exact set of distinct values in the dictionary page, so a bloom filter for it adds nothing a reader cannot already get exactly, while every value is still hashed into the filter during the write and the filter is serialized after the chunk. parquet-java stopped writing these in PARQUET-2251 (apache/parquet-java#1033, 1.13.0), so files from Spark, Hive and Iceberg never have a bloom filter on a dictionary-only chunk, and there was no way to get the same output from this crate. Details in #10962.
What changes are included in this PR?
WriterProperties::bloom_filter_for_dictionary_encoded_chunkswithWriterPropertiesBuilder::set_bloom_filter_for_dictionary_encoded_chunksandDEFAULT_BLOOM_FILTER_FOR_DICTIONARY_ENCODED_CHUNKS = true, so the default output is unchanged.GenericColumnWriter::close, when the option isfalse, the bloom filter is dropped unlessencoding_statsrecords at least oneDATA_PAGE/DATA_PAGE_V2whose encoding is notPLAIN_DICTIONARYorRLE_DICTIONARY, the same testParquetFileWriter.writeColumnChunkapplies in parquet-java.flush_bloom_filteris still called so the encoder state is reset as before.Are these changes tested?
Yes,
test_bloom_filter_for_dictionary_encoded_chunkswrites a small dictionary-friendly Int32 column across the dictionary on/off × option on/off matrix and asserts a filter is present in every case except dictionary on with the option off.Are there any user-facing changes?
One new writer property, opt-in, documented on the setter. No breaking changes.