Skip to content

Avoid unnecessary Parquet Bloom filter reads through incremental pruning #25797

Description

@haohuaijin

Is your feature request related to a problem or challenge?

Parquet Bloom filter pruning currently loads filters for the relevant predicate columns before evaluating whether a row group can be excluded.

For queries with multiple filtering conditions, some of these reads may be unnecessary. For example:

SELECT COUNT(*)
FROM traces
WHERE trace_id = 'abc123'
  AND service_name = 'checkout'
  AND host = 'host-42';

If a row group survives statistics pruning, but its trace_id Bloom filter proves that 'abc123' is absent, the entire row group can already be excluded. Reading the service_name and host Bloom filters cannot change that decision.

This is particularly relevant to observability workloads that filter across several columns, especially when Parquet files reside in remote object storage.

Describe the solution you'd like

Evaluate Bloom filters incrementally and stop reading additional filters for a row group as soon as a necessary condition proves that the group cannot match.

The desired behavior is:

  1. Read a relevant Bloom filter.
  2. Check whether it rules out the row group.
  3. If it does, skip the remaining Bloom filter reads for that group.
  4. Otherwise, continue with the remaining relevant columns.

Existing literal guarantees could provide the necessary conditions. For example, an IN condition can exclude a group only when all candidate values are definitely absent. A possible Bloom hit must remain inconclusive.

This would also allow filters to be released after evaluation instead of retaining all filters until a separate pruning pass. The implementation should preserve conservative handling of compound predicates, NULLs, missing filters, and read failures.

Describe alternatives you've considered

Optimizing evaluation after loading all filters could reduce CPU overhead, but would not avoid unnecessary Bloom filter reads.

Additional context

The relevant loading and pruning flow is in datafusion/datasource-parquet/src/opener/mod.rs, particularly load_bloom_filters() and prune_bloom_filters().

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

enhancementNew feature or request

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions