Skip to content

fix(orchestrator): keep heartbeating until workers have drained - #76

Merged
psteinroe merged 1 commit into
mainfrom
fix/heartbeat-during-drain
Oct 2, 2026
Merged

psteinroe merged 1 commit into
mainfrom
fix/heartbeat-during-drain

Conversation

@psteinroe

Copy link
Copy Markdown
Owner

The orchestrator now keeps heartbeating until its workers have stopped. Before, shutdown aborted the same controller that ran the heartbeat loop, so heartbeats stopped as soon as stop() was called. Since #68, workers wait for running handlers, so a handler that runs for more than STALE_ORCHESTRATOR_MAX_AGE_MS (5 minutes) after shutdown let another orchestrator recover its executions as stale while it was still running, and the handler ran twice. A shutdown signal from a heartbeat (the version-mismatch path) had the same problem.

Separate heartbeat abort

The heartbeat loop now has its own AbortController. cleanup() aborts it and awaits the loop after the workers have stopped, right before deleting the orchestrator row. Shutdown signalling to workers is unchanged: stop(), process signals and a heartbeat shutdown signal still abort the orchestrator's controller. When a heartbeat reports that the row was recovered as stale, the loop still aborts the orchestrator and stops heartbeating, as in #67. #67's guarantees still hold: cleanup awaits the loop before it removes the row, so no heartbeat runs after cleanup and the row is not re-created.

Verification

New parameterized test in orchestrator-cleanup.test.ts, for both stop() and a heartbeat shutdown signal. A handler keeps running after shutdown, DB time moves forward 10 minutes and the heartbeat timer is advanced. Then the orchestrator row has a fresh last_heartbeat_at, recoverStaleOrchestrators leaves the execution locked, and after the handler finishes the execution is completed and the row is removed. Both cases failed before the fix. just lint, just format, bun run typecheck and the full bun test (352 pass) pass.

Follow-ups (out of scope)

  • Once an orchestrator has been recovered as stale it still stops heartbeating right away. Executions it claimed between recovery and the abort can therefore be recovered again if its handlers keep running for another 5 minutes. Those results are still fenced by locked_by.

@psteinroe psteinroe added the ready label Oct 2, 2026
@psteinroe
psteinroe force-pushed the fix/heartbeat-during-drain branch from 31b5aa7 to 3447b11 Compare October 2, 2026 22:50
@psteinroe
psteinroe merged commit 454c43a into main Oct 2, 2026
9 checks passed
@psteinroe
psteinroe deleted the fix/heartbeat-during-drain branch October 2, 2026 22:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant