Skip to content

Preserve Jepsen Driver state through cleanup death races - #26

Open
jeregrine wants to merge 2 commits into
mainfrom
fix-jepsen-cleanup-deaths
Open

Preserve Jepsen Driver state through cleanup death races#26
jeregrine wants to merge 2 commits into
mainfrom
fix-jepsen-cleanup-deaths

Conversation

@jeregrine

Copy link
Copy Markdown
Member

Problem

Jepsen cluster cleanup checked whether each Owner was alive before making unguarded calls. A registry-conflict death between those steps crashed the Driver, whose replacement forgot unrelated unlinked Owners. Snapshot refresh had the same race.

Fix

Cleanup and snapshot refresh now share a monitored-owner call path. A confirmed DOWN is processed through the ordinary lifecycle handler, removing only that Owner and retaining every other Owner, monitor, and cached state. A timeout without a matching DOWN remains an explicit error and does not discard state.

Owner cleanup returns its refreshed snapshot in the same turn, removing the second unguarded snapshot pass. Public cleanup propagates failures instead of ignoring them.

Supporting information

The concurrency regressions suspend real harness Owners, queue cleanup or snapshots, and deliver conflict-shaped deaths while the Driver is waiting. They also cover death before snapshot handling, unexpected death evidence, live-owner timeout, and preservation of unrelated public memberships. The policy deciding which conflict deaths are legitimate is unchanged.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant