fix(core): resume sessions after restart#35820
Closed
kitlangton wants to merge 3 commits into
Closed
Conversation
Contributor
Author
|
Closing this draft because startup event scanning and a long-held recovery lock are the wrong shape. #35826 establishes single managed-daemon ownership first; the follow-up recovery change will use a bounded pending-recoveries projection instead of scanning history. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Addresses #35646
Summary
EffectFlockHow it works
A recovering process acquires the global Session recovery lock before reading lifecycle state. It holds that lock while all interrupted Sessions resume. Other daemon candidates wait, then re-read lifecycle state after the first process has recorded
session.execution.startedor a terminal outcome, so they have nothing to resume. No new recovery table or migration is needed.User interruptions, completed executions, failures, and unmatched execution starts are not resumed. Unexpected hard-process death still needs an explicit ownership/lease design because blindly resuming an unmatched start can duplicate provider work while another process remains alive.
Verification
bun run test -- test/session-execution-local.test.ts test/util/effect-flock.test.tsbun typecheckinpackages/core