Skip to content

[Parquet] Add writer option to skip bloom filters for column chunks whose data pages are all dictionary encoded - #10963

Open
ranflarion wants to merge 2 commits into
apache:mainfrom
ranflarion:bloom-filter-dictionary-encoded-chunks
Open

[Parquet] Add writer option to skip bloom filters for column chunks whose data pages are all dictionary encoded#10963
ranflarion wants to merge 2 commits into
apache:mainfrom
ranflarion:bloom-filter-dictionary-encoded-chunks

Conversation

@ranflarion

Copy link
Copy Markdown
Contributor

Which issue does this PR close?

Rationale for this change

A column chunk whose data pages are all dictionary encoded carries its exact set of distinct values in the dictionary page, so a bloom filter for it adds nothing a reader cannot already get exactly, while every value is still hashed into the filter during the write and the filter is serialized after the chunk. parquet-java stopped writing these in PARQUET-2251 (apache/parquet-java#1033, 1.13.0), so files from Spark, Hive and Iceberg never have a bloom filter on a dictionary-only chunk, and there was no way to get the same output from this crate. Details in #10962.

What changes are included in this PR?

  • WriterProperties::bloom_filter_for_dictionary_encoded_chunks with WriterPropertiesBuilder::set_bloom_filter_for_dictionary_encoded_chunks and DEFAULT_BLOOM_FILTER_FOR_DICTIONARY_ENCODED_CHUNKS = true, so the default output is unchanged.
  • In GenericColumnWriter::close, when the option is false, the bloom filter is dropped unless encoding_stats records at least one DATA_PAGE/DATA_PAGE_V2 whose encoding is not PLAIN_DICTIONARY or RLE_DICTIONARY, the same test ParquetFileWriter.writeColumnChunk applies in parquet-java. flush_bloom_filter is still called so the encoder state is reset as before.

Are these changes tested?

Yes, test_bloom_filter_for_dictionary_encoded_chunks writes a small dictionary-friendly Int32 column across the dictionary on/off × option on/off matrix and asserts a filter is present in every case except dictionary on with the option off.

Are there any user-facing changes?

One new writer property, opt-in, documented on the setter. No breaking changes.

@github-actions github-actions Bot added the parquet Changes to the parquet crate label Sep 2, 2026

@etseidl etseidl left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This seems reasonable, assuming one has access to the dictionary (which we may be adding here soon #10420).

It would be nice if we could also skip populating the bloom filter while a dictionary is in use. On fallback to non-dictionary, we could then add all the dict keys to the filter.

Comment thread parquet/src/file/properties.rs Outdated
Comment on lines +810 to +811
/// The dictionary page of such a chunk already lists every distinct value, so parquet-java
/// skips the bloom filter for it; set this to `false` to write files the same way.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think we need the comparison to parquet-java here. I'd instead say that since the dictionary contains all distinct values, the bloom filter is redundant and users might wish to save the space by setting this to false.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

parquet Changes to the parquet crate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Parquet] Option to skip bloom filters for column chunks whose data pages are all dictionary encoded

2 participants