Is your feature request related to a problem or challenge?
Parquet Bloom filter pruning currently loads filters for the relevant predicate columns before evaluating whether a row group can be excluded.
For queries with multiple filtering conditions, some of these reads may be unnecessary. For example:
SELECT COUNT(*)
FROM traces
WHERE trace_id = 'abc123'
AND service_name = 'checkout'
AND host = 'host-42';
If a row group survives statistics pruning, but its trace_id Bloom filter proves that 'abc123' is absent, the entire row group can already be excluded. Reading the service_name and host Bloom filters cannot change that decision.
This is particularly relevant to observability workloads that filter across several columns, especially when Parquet files reside in remote object storage.
Describe the solution you'd like
Evaluate Bloom filters incrementally and stop reading additional filters for a row group as soon as a necessary condition proves that the group cannot match.
The desired behavior is:
- Read a relevant Bloom filter.
- Check whether it rules out the row group.
- If it does, skip the remaining Bloom filter reads for that group.
- Otherwise, continue with the remaining relevant columns.
Existing literal guarantees could provide the necessary conditions. For example, an IN condition can exclude a group only when all candidate values are definitely absent. A possible Bloom hit must remain inconclusive.
This would also allow filters to be released after evaluation instead of retaining all filters until a separate pruning pass. The implementation should preserve conservative handling of compound predicates, NULLs, missing filters, and read failures.
Describe alternatives you've considered
Optimizing evaluation after loading all filters could reduce CPU overhead, but would not avoid unnecessary Bloom filter reads.
Additional context
The relevant loading and pruning flow is in datafusion/datasource-parquet/src/opener/mod.rs, particularly load_bloom_filters() and prune_bloom_filters().
Is your feature request related to a problem or challenge?
Parquet Bloom filter pruning currently loads filters for the relevant predicate columns before evaluating whether a row group can be excluded.
For queries with multiple filtering conditions, some of these reads may be unnecessary. For example:
If a row group survives statistics pruning, but its
trace_idBloom filter proves that'abc123'is absent, the entire row group can already be excluded. Reading theservice_nameandhostBloom filters cannot change that decision.This is particularly relevant to observability workloads that filter across several columns, especially when Parquet files reside in remote object storage.
Describe the solution you'd like
Evaluate Bloom filters incrementally and stop reading additional filters for a row group as soon as a necessary condition proves that the group cannot match.
The desired behavior is:
Existing literal guarantees could provide the necessary conditions. For example, an
INcondition can exclude a group only when all candidate values are definitely absent. A possible Bloom hit must remain inconclusive.This would also allow filters to be released after evaluation instead of retaining all filters until a separate pruning pass. The implementation should preserve conservative handling of compound predicates, NULLs, missing filters, and read failures.
Describe alternatives you've considered
Optimizing evaluation after loading all filters could reduce CPU overhead, but would not avoid unnecessary Bloom filter reads.
Additional context
The relevant loading and pruning flow is in
datafusion/datasource-parquet/src/opener/mod.rs, particularlyload_bloom_filters()andprune_bloom_filters().