Preserve Jepsen Driver state through cleanup death races - #26
Open
jeregrine wants to merge 2 commits into
Open
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Jepsen cluster cleanup checked whether each Owner was alive before making unguarded calls. A registry-conflict death between those steps crashed the Driver, whose replacement forgot unrelated unlinked Owners. Snapshot refresh had the same race.
Fix
Cleanup and snapshot refresh now share a monitored-owner call path. A confirmed DOWN is processed through the ordinary lifecycle handler, removing only that Owner and retaining every other Owner, monitor, and cached state. A timeout without a matching DOWN remains an explicit error and does not discard state.
Owner cleanup returns its refreshed snapshot in the same turn, removing the second unguarded snapshot pass. Public cleanup propagates failures instead of ignoring them.
Supporting information
The concurrency regressions suspend real harness Owners, queue cleanup or snapshots, and deliver conflict-shaped deaths while the Driver is waiting. They also cover death before snapshot handling, unexpected death evidence, live-owner timeout, and preservation of unrelated public memberships. The policy deciding which conflict deaths are legitimate is unchanged.