Skip to content

docs(concurrency): note that a large single-group backlog slows claims - #69

Open
psteinroe wants to merge 1 commit into
mainfrom
perf/group-backlog-claim
Open

psteinroe wants to merge 1 commit into
mainfrom
perf/group-backlog-claim

Conversation

@psteinroe

@psteinroe psteinroe commented Oct 2, 2026 •

Copy link
Copy Markdown
Owner

Documents a known limitation of groupConcurrency. A claim skips the pending executions of full groups one by one in claim order, so a very large backlog in one group slows every claim of that task until it drains. With 100k pending rows in one full group, every claim reads all of them to find 10 claimable rows in other groups (about 48 ms instead of 0.5 ms). pg-boss #842 hit the same issue. The note recommends spreading bulk work across groups or moving it to a separate task or queue.

What was tried and why it was dropped

The first version of this PR fixed the claim. When full groups held a large backlog ahead of the batch, the claim walked the task's groups and took the head of each open group. It was rejected as too costly:

  • It added a partial index over all available rows.
  • It added about 180 lines of SQL and three hard-coded 1000-row/1000-group heuristics.
  • Common grouped claims got slower: many small groups went from 0.52 to 1.27 ms, single-row groups from 0.62 to 0.94 ms, and tasks with more than 1000 groups from 48 to 63 ms.
  • With groupConcurrency: 1 under load, the checks would run on nearly every grouped claim.

No simpler approach meets the constraints (no regression in the common scenarios, no index covering ungrouped rows, no thresholds, a small diff):

  • Exact merge of group heads. A btree ordered by claim order cannot skip the rows of a full group. The alternative is a merge of group heads, which costs one index lookup per group: about 8 buffers and ~15 µs per group, measured on Postgres 15. Doing that on every grouped claim makes tasks with many groups much slower than today: 10k groups take about 250 ms.
  • Choosing between the scan and the merge per claim. This needs a cost cutoff, which is a threshold.
  • Adding "group" to get_task_executions. This would let full-group rows be rejected inside the index, but it changes an index over every row. The scan also stays proportional to the backlog and still fetches each heap row, because queue and is_available are evaluated as heap filters.

A real fix probably needs per-group state, for example keeping only each group's head claimable. That is a larger design change than this PR.

@psteinroe psteinroe changed the title perf(claims): claim open group heads behind a saturated group backlog docs(concurrency): note that a large single-group backlog slows claims Oct 2, 2026
@psteinroe
psteinroe force-pushed the perf/group-backlog-claim branch from e045a8a to a9bec83 Compare October 2, 2026 23:04

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant