Repository navigation
fix(orchestrator): keep heartbeating until workers have drained - #76
Merged
Merged
Conversation
psteinroe
force-pushed
the
fix/heartbeat-during-drain
branch
from
October 2, 2026 22:50
31b5aa7 to
3447b11
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The orchestrator now keeps heartbeating until its workers have stopped. Before, shutdown aborted the same controller that ran the heartbeat loop, so heartbeats stopped as soon as
stop()was called. Since #68, workers wait for running handlers, so a handler that runs for more thanSTALE_ORCHESTRATOR_MAX_AGE_MS(5 minutes) after shutdown let another orchestrator recover its executions as stale while it was still running, and the handler ran twice. A shutdown signal from a heartbeat (the version-mismatch path) had the same problem.Separate heartbeat abort
The heartbeat loop now has its own
AbortController.cleanup()aborts it and awaits the loop after the workers have stopped, right before deleting the orchestrator row. Shutdown signalling to workers is unchanged:stop(), process signals and a heartbeatshutdownsignal still abort the orchestrator's controller. When a heartbeat reports that the row was recovered as stale, the loop still aborts the orchestrator and stops heartbeating, as in #67. #67's guarantees still hold: cleanup awaits the loop before it removes the row, so no heartbeat runs after cleanup and the row is not re-created.Verification
New parameterized test in
orchestrator-cleanup.test.ts, for bothstop()and a heartbeat shutdown signal. A handler keeps running after shutdown, DB time moves forward 10 minutes and the heartbeat timer is advanced. Then the orchestrator row has a freshlast_heartbeat_at,recoverStaleOrchestratorsleaves the execution locked, and after the handler finishes the execution is completed and the row is removed. Both cases failed before the fix.just lint,just format,bun run typecheckand the fullbun test(352 pass) pass.Follow-ups (out of scope)
locked_by.