Status: implemented and validated; unrelated repository-baseline gate failures are recorded below
Owner: agentOS native sidecar
Baseline: the request-concurrency stack through
refactor(native-sidecar): separate ownership from conflict policy
This document is the implementation contract and completion checklist. A box may be checked only after the implementation exists and its corresponding test passes. The implementing agent must record the commands and results in the final validation section before declaring the work complete.
Keep the concurrency behavior fixed by request-concurrency-fix-prompt.md, but
replace the overlapping request, ownership, conflict, and VM lifecycle
coordination machinery with the smallest explicit model that provides the same
safety.
The completed design has:
- one bounded operation table for every inbound request on a connection, including ordinary and progress requests;
- one short-held ownership/membership state lock, never held across
.await; - one small lifecycle gate per VM;
- ACP-owned per-route exclusion, without generic core extension-ordering machinery; and
- no duplicate per-connection, per-session, and per-VM operation maps.
The implementation must be a simplification. Production code in the ownership, request-admission, and lifecycle subsystem should be materially smaller after the change. Adding behavioral tests is expected and does not count against that goal.
The current implementation correctly prevents a long ACP prompt from holding the protocol ingress loop, but represents one request in several overlapping places:
RequestOperationRegistryorProgressRequestRegistry;- connection ownership operations;
- session ownership operations;
- VM ownership operations;
- extension conflict registrations;
- task/completion supervision; and
- terminal/progress publication guards.
The ownership coordinator also combines four different concerns:
- entity membership;
- disposal state;
- operation tracking;
- VM lifecycle exclusion.
That makes cancellation and disposal hard to audit and caused ordinary and
progress request IDs to be checked in separate registries. For example, an
ordinary request 42 and progress request 42 can currently be admitted on
the same connection even though request IDs are connection-scoped.
The lifecycle exclusion itself is necessary. The hierarchy around it is not.
- Production ingress only decodes, validates, reserves, registers, and starts work. It never awaits business execution, VM activity, output capacity, adapter I/O, or event drainage.
- Ordinary work is concurrent by default.
- Different VMs never share a lifecycle gate.
- A lifecycle mutation runs only after conflicting work already admitted for that VM has drained.
- Once lifecycle admission becomes pending, new ordinary work for that VM is rejected with a typed, retryable conflict; it is not placed in a hidden waiter queue.
- Progress requests, registered sidecar responses, cancellation, permission responses, shutdown, and transport failure do not require ordinary request admission.
- Progress-critical internal events needed to settle already-admitted work can continue during lifecycle drain, remain bounded, and are included in the drain count.
- No internal event mutates a VM while its lifecycle gate is active.
- A long ACP prompt does not hold the VM lifecycle gate. Only its short VM service operations enter the gate.
- Each
(connection_id, request_id)identifies at most one live inbound request, regardless of whether it is ordinary or progress-classified. - Every admitted request retains exactly one response-publication right and releases all count/byte/gate accounting on every terminal path.
- Entity disposal closes admission before signalling work, waits without holding a state lock, and cannot remove state underneath an admitted operation.
- Every queue, table, retained byte collection, and waiter set remains bounded with a typed error naming the limit and configuration path.
- No mutex or
RefCellborrow crosses.await. - No global
Arc<Mutex<NativeSidecar>>or equivalent is introduced. - Native-sidecar remains opaque to ACP payloads.
fd 0 / fd 3 readers
|
v
+------------------------+
| Protocol router |
| - classify |
| - reserve output |
| - admit operation |
| - start task |
+-----------+------------+
|
v
+------------------------+
| Operation table | one short metadata mutex
| - global request IDs |
| - ordinary/progress |
| - count/byte budgets |
| - ownership + phase |
| - cancellation |
| - drain notification |
+-----------+------------+
|
| VM-scoped work only
v
+------------------------+
| VmLifecycleGate | one independent gate per VM
| - ordinary permits |
| - internal permits |
| - lifecycle permit |
+------------------------+
ACP route state remains inside the ACP extension. Output reservations and the physical output broker remain separate from the operation table.
Replace the separate ordinary and progress registries with one operation table keyed by:
struct OperationKey {
connection_id: String,
request_id: RequestId,
}The table may retain separate configured budgets for ordinary and progress classes, but duplicate-ID detection is global across both classes.
The conceptual record is:
struct OperationRecord {
generation: u64,
class: OperationClass, // Ordinary | Progress
ownership: OwnershipScope,
operation: String,
request_bytes: usize,
cancellation: OperationCancellation,
publication: ResponsePublicationGuard,
}Requirements:
- Replace
RequestOperationRegistryandProgressRequestRegistrywith one authoritative map, or make one a thin class-specific view over a single authoritative map. - Ordinary and progress budgets remain independently bounded.
- Duplicate detection happens before either class reserves or starts work.
- The same numeric request ID remains valid on different connections.
- Replace
TerminalResponseGuardandProgressAcknowledgementGuardwith one response-publication primitive parameterized only by response class. - Keep exactly-once takeover for bounded shutdown and transport failure.
- Remove operation lifecycle states that are used only to mirror task execution. Retain only state that affects admission, cancellation, publication correctness, or externally useful diagnostics.
- Scope cancellation and drain scan the single bounded table. Do not add secondary per-entity operation maps merely to avoid scanning at most the configured in-flight operation bound.
- Ordinary shutdown may close ordinary admission while progress admission remains open for the bounded drain.
- Table closure and entity closure use generations so an old disposal cannot cancel a later entity reusing the same textual ID.
Use one short-held metadata lock for connection, session, and VM membership. This lock is allowed because it protects only bounded in-memory transitions and is always released before external work or waiting.
The simplest recommended representation places entity phases and the bounded
operation map under the same metadata mutex. If the implementation keeps
membership and operation storage behind separate mutexes, it must document one
lock order, make admission/disposal one linearizable transaction through a
generation-bearing lease, and add a deterministic test for the race. Merely
checking Open, releasing the membership lock, and registering an operation
later is not safe.
The ownership state must not contain a second copy of every active operation. The operation table is authoritative for active work and cancellation.
Requirements:
- Entity records have an explicit generation and
Open | Closingphase. - Request admission validates the complete connection/session/VM ownership path against one coherent snapshot.
- Entity creation and disposal are linearized against admission.
- Disposal marks the entity
Closingbefore signalling matching operation records. - Disposal waits on operation-table and lifecycle-gate drain notification without holding the membership lock.
- Entity state is removed only after matching operations and gate permits have drained.
- A stale permit or completion from an earlier entity generation cannot mutate a newly created entity with the same textual ID.
- Delete
ConnectionOperationRegistration,SessionOperationRegistration,VmOperationRegistration, and their independent operation-ID counters. - Delete nested lock sequences used solely to register the same operation at connection, session, and VM scopes.
- Do not reuse
runtime.protocol.maxInFlightRequestsas a session membership limit. Reuse an existing semantically correct bound, remove a redundant retained collection, or add a dedicated documented limit.
Each live VM owns one independent VmLifecycleGate. The gate contains only the
state required to exclude lifecycle mutations from conflicting VM work.
Conceptual API:
impl VmLifecycleGate {
fn try_enter_ordinary(
&self,
cancellation: OperationCancellation,
) -> Result<VmOrdinaryPermit, VmGateError>;
fn try_enter_internal(
&self,
cancellation: OperationCancellation,
) -> Result<InternalAdmission, VmGateError>;
async fn begin_lifecycle(
&self,
cancellation: OperationCancellation,
) -> Result<VmLifecyclePermit, VmGateError>;
}Conceptual state:
struct VmGateState {
phase: VmGatePhase,
closing: bool,
ordinary_active: usize,
internal_active: usize,
next_generation: u64,
}
enum VmGatePhase {
Idle,
Pending { generation: u64 },
Active { generation: u64 },
}Exact state semantics:
Idle
ordinary -> increment ordinary_active
internal -> increment internal_active
lifecycle -> Pending and close new ordinary admission
Pending
new ordinary -> typed retryable conflict
bounded internal settlement work -> admitted or durably deferred
ordinary_active == 0 && internal_active == 0 -> Active
Active
ordinary -> typed retryable conflict
internal -> remains in its durable source; never mutates the VM
matching lifecycle permit drops -> Idle
closing == true (orthogonal to Idle/Pending/Active)
all new gate admission -> typed shutdown/disposal error
existing permits drain and notify disposal
cancellation/permit drop must never reopen admission
Requirements:
-
Idle -> Pendingis the linearization point that closes ordinary VM admission. - Only one lifecycle request may be pending or active. A second receives a typed conflict rather than waiting.
- The lifecycle waiter registers notification before checking drain state, preventing lost wakeups.
- Lifecycle cancellation while pending restores
Idleonly when the cancellation belongs to the current gate generation and the gate is not closing. A closing gate remains unavailable. - Dropping an active lifecycle permit restores
Idleexactly once when the gate is open. It never clears the orthogonal closing state. - Dropping an ordinary or internal permit decrements exactly one counter and wakes a pending lifecycle waiter when the gate may advance.
- Counter underflow, stale generation, and poisoned-state recovery are hard logged invariant failures.
- Ordinary and internal admission remains bounded independently.
- Internal work admitted during
Pendingis limited to the existing settlement/event path. Public requests cannot label themselves internal. - Durable internal events are not removed from their source unless the gate or a tracked deferred permit owns them.
- A pending lifecycle operation cannot spin while internal capacity is unavailable.
- A lifecycle gate for VM A never reads, writes, or notifies VM B's gate.
The production doc comment on VmLifecycleGate must explain all of the
following:
- a standard-library lock cannot cross
.awaitwithout blocking a runtime worker; - ordinary admission must reject after lifecycle becomes pending rather than join an implicit waiter queue;
- progress-critical internal settlement work must remain admissible during
Pendingand must be counted in the lifecycle drain; - disposal needs explicit cancellation, generations, and active counts; and
- the mutex inside the gate protects only non-suspending state transitions and is released before waiting.
The comment must also state that a short global metadata mutex is acceptable: the forbidden design is a lock held across request execution, not a lock around bounded map/counter transitions.
Replace generalized conflict policy with one narrow VM concurrency class:
enum VmConcurrencyClass {
None,
Ordinary,
ExclusiveLifecycle,
}The names may differ, but the model must have no extension-specific conflict variant.
Classification requirements:
- Connection- and session-owned operations use
None. - Extension envelope requests use
None; ACP enforces its own route-level response-loop exclusion. - Short VM service calls made by extensions use
Ordinary. - Internal VM settlement/event work uses the dedicated internal gate API,
not
Ordinaryand not a public request class. - Ordinary core VM work uses
Ordinary. - The following remain
ExclusiveLifecycleunless an operation owner provides a narrower, tested safety argument: - dispose VM; - bootstrap root filesystem; - configure VM; - create or seal layer; - import or export snapshot; - create overlay; - snapshot root filesystem; - link package. - Classification is implemented in one auditable location or on the request type itself. Do not introduce a pairwise conflict matrix.
- A newly added lifecycle request must require an explicit classification in a compile-time exhaustive match or architecture test.
ACP is the only production extension currently returning an ordering key, and
it returns ExtensionManaged; the core coordinator therefore performs no
exclusion for that key. Retain ACP's route state and remove the unused generic
layer around it.
- Delete
ExtensionOrderingPolicy. - Delete
Extension::request_ordering_key. - Delete
Extension::request_ordering_policy. - Delete
ConflictPolicy::Extensionor its replacement equivalent. - Delete connection-level
extension_conflictsstate and registration guards. - Delete tests and architecture guards that require the removed hooks.
- Preserve
Extension::request_classso an opaque extension can identify progress requests without native-sidecar decoding its payload. - Preserve ACP's per-route state machine and typed
session_busyresponse. - Preserve concurrent prompts on different ACP routes.
The following source audit must return no production matches:
rg -n \
'ExtensionOrderingPolicy|request_ordering_key|request_ordering_policy|ConflictPolicy::Extension|extension_conflicts' \
crates/native-sidecar/src crates/agentos-sidecar/srcThis refactor must not simplify away the reserved progress path. It must also verify that progress is not reserved only at ingress/output while silently joining an ordinary internal service queue.
- ACP cancel and permission response bypass ordinary request count/byte admission and the VM ordinary gate.
- Registered
SidecarResponseFramerouting remains direct to its waiter. - Shutdown and terminal transport failure remain directly routable.
- If ACP cancellation must issue
WriteStdin, the service scheduler has bounded progress-reserved admission or another direct bounded path. - Saturating ordinary extension-service work cannot prevent an admitted cancellation from reaching the adapter and receiving its acknowledgement.
- Progress saturation returns a typed limit error without consuming the target operation's terminal response right.
- Continuous progress traffic cannot starve already-retained terminal responses. Use bounded fair scheduling rather than unlimited strict priority if the existing output test exposes starvation.
- Multiple ordinary permits coexist.
- Lifecycle becomes pending and waits for all earlier ordinary permits.
- New ordinary admission is rejected while lifecycle is pending.
- Lifecycle becomes active only after ordinary and internal counts reach zero.
- A second lifecycle request is rejected in both pending and active phases.
- Cancelling a pending lifecycle operation reopens ordinary admission.
- Dropping the active lifecycle permit reopens ordinary admission.
- A stale lifecycle permit cannot reopen or mutate a later generation.
- Internal settlement work can run during pending and is included in drain.
- Internal work is not admitted while lifecycle is active.
- Gate closure cancels/rejects admission and drains all permits.
- Closing during pending and closing during active lifecycle work cannot reopen admission when the lifecycle future or permit later drops.
- Near-limit warnings and typed count-limit errors name the configuration path.
- Ordinary request
42followed by progress request42on the same connection rejects the second request. - Progress request
42followed by ordinary request42rejects the second request. - Request
42on two different connections remains valid. - Ordinary and progress class budgets are independent despite sharing the ID table.
- Cancellation versus natural completion publishes exactly one response.
- Forced shutdown takeover publishes at most one response and releases all accounting.
- Closing a connection rejects later admission and cancels only matching operations.
- Scope drain reaches zero count and zero retained request bytes.
- Same-VM read and write requests execute concurrently.
- VM-A lifecycle drain does not delay VM-B ordinary work.
- Configure/dispose waits for already-admitted same-VM work and rejects new same-VM work while pending.
- A sleeping ACP prompt does not retain the VM gate; same-VM file read/write completes during the prompt.
- Two ACP prompts on different routes run concurrently.
- A second prompt on the same route receives ACP's typed
session_busy. - ACP cancellation progresses while ordinary request admission is full.
- ACP cancellation reaches adapter stdin while ordinary extension-service capacity is saturated.
- Disposal during an active prompt cancels, drains, and removes all route, operation, gate, permission, and event-waiter state.
- Repeated lifecycle/cancel/process-exit races produce no duplicate terminal response, leaked permit, lost event, or hot spin.
- Add a deterministic model test that compares randomized gate operations against a small reference state machine. Cover admit/drop/cancel/close and stale-generation sequences.
- Extend protocol frame fuzzing with mixed ordinary/progress duplicate IDs, lifecycle transitions, cancellation, and shutdown.
- Add a bounded load test with multiple VMs, concurrent ordinary work, periodic lifecycle requests, ACP progress, and deliberately delayed completions.
- The load test asserts a finite completion deadline, per-VM independence, exactly-once responses, and zero final accounting rather than relying on sleeps or log inspection.
- Safeguard-firing saturation tests remain cheap and active in PR CI; tests attempting to prove absence of a resource bound remain explicitly ignored.
- Remove obsolete types, aliases, compatibility wrappers, error variants, metrics, comments, and tests instead of leaving deprecated paths.
- Do not retain both the old coordinator and new gate behind a feature flag.
- Do not introduce a second request scheduler, runtime, or worker pool.
- Do not add a general ordinary request FIFO.
- Remove duplicated admission preflight where one atomic admission method can reserve all required state safely.
- Keep output reservations outside the VM gate; no gate permit may be held while waiting for output capacity.
- Keep public and Rust client behavior identical if any wire/config surface changes.
- Update
request-concurrency-fix-prompt.mdto describe the final simplified ownership/gate model and remove the obsolete ordering-key description. - Update native-sidecar architecture comments and guards to enforce behavior and forbidden dependencies, not incidental private symbol names.
-
cargo fmtand Clippy complete without new allows for dead code, complex types, or too many arguments in the new subsystem. - The production ownership/request/gate implementation is net smaller than the baseline, or the final validation record explains every net-new production abstraction.
- Add failing tests for cross-class duplicate IDs and the standalone VM gate state machine.
- Introduce the single operation table and shared publication guard.
- Introduce
VmLifecycleGatewith ordinary, internal, and lifecycle RAII permits. - Rewire core VM requests and extension service commands to the gate.
- Move disposal cancellation/drain to the authoritative operation table.
- Remove per-entity operation maps and registration guards.
- Remove generic extension ordering hooks and conflict state.
- Run the real-loop ACP/filesystem/lifecycle regressions.
- Saturate the extension service path and close any progress-reservation gap exposed by the test.
- Run model/fuzz/load coverage and the final source/architecture audits.
- Update the original concurrency design document and record final validation below.
The implementing agent may reorder mechanical steps, but must not temporarily weaken the production invariants on the final revision.
Run focused checks first:
cargo fmt --check
cargo test -p agentos-native-sidecar request_operations
cargo test -p agentos-native-sidecar ownership_coordinator
cargo test -p agentos-native-sidecar request_concurrency
cargo test -p agentos-native-sidecar --test architecture_guards
cargo test -p agentos-sidecar acp
pnpm --dir packages/core test:prThen run repository gates proportional to the final diff:
cargo check --workspace
cargo clippy --workspace --all-targets -- -D warnings
pnpm check-types
pnpm buildIf exact test filters change as obsolete modules are removed, replace them with the new target names and record the actual commands below. Do not silently skip a behavioral category because its old filter no longer matches tests.
Required source audits:
rg -n \
'ExtensionOrderingPolicy|request_ordering_key|request_ordering_policy|ConflictPolicy::Extension|extension_conflicts|ProgressRequestRegistry' \
crates/native-sidecar/src crates/agentos-sidecar/src
rg -n \
'ConnectionOperationRegistration|SessionOperationRegistration|VmOperationRegistration' \
crates/native-sidecar/srcBoth commands must return no production matches. Test-only reference-model names should also be renamed so the audit remains unambiguous.
This work is complete only when:
- Every checkbox in sections 1 through 8 is complete.
- The original prompt/filesystem/cancel reproduction passes.
- All required focused validation passes; repository-wide gates either pass or have a reproduced, unrelated baseline failure recorded below.
- Source audits show no obsolete ownership or ordering implementation.
- A reviewer can identify one authoritative operation table and one lifecycle gate per VM without tracing duplicate shadow state.
- The final diff contains no unrelated refactor or generated artifacts.
- The working copy contains no accidental empty jj revision in the stack.
- The implementation revision has a plain conventional-commit description with no coding-agent attribution.
Implementing agent fills this section in before handoff.
- Implementation revision:
xulmxrrl—refactor(native-sidecar): simplify VM lifecycle coordination - Final reviewer revision:
xulmxrrl - Focused Rust tests:
cargo fmt --check; operation-table tests 20/20; lifecycle-coordinator tests 11/11; request-concurrency tests 25/25; architecture guards 41/41; ACP tests 52/52. - ACP/public-client regression: the delayed response after 256 updates and the two-prompts-plus-filesystem regression pass; all four native-sidecar migration parity scenarios pass (6/6 tests across the two files).
- Model/fuzz/load tests: the deterministic randomized gate model, mixed ordinary/progress protocol framing, duplicate-ID cases, lifecycle/cancel races, bounded progress saturation, and multi-VM load deadline all pass in the focused Rust suites.
- Workspace checks:
cargo check --workspace, package-scoped Core and runtime-core builds/typechecks, fixed-version verification, and publish helper typechecks/tests pass. The repository-wide Clippy gate still reaches unchangedlarge_enum_variantand disabled browser-target failures; rootpnpm check-typesstill reaches an unmaterialized example dependency; rootpnpm buildstill requires the generated Codex WASI release artifact; andpackages/core test:prstill has the unchanged top-levelpreadexport assertion (98/99 unit tests). These baseline failures are recorded in the agentOS friction log; the focused changed-package and real-loop gates pass. - Source audits: both required
rgcommands return no production matches.OperationTableis the sole inbound request table and eachVmRecordowns one independentVmLifecycleGate. - Production lines removed/added in the simplified subsystem: before test modules, ownership coordination changed from 1,843 to 1,686 lines (-157) and request operations from 1,484 to 1,524 (+40), for a net reduction of 117 production lines. The request-operation increase is the shared progress projection and cross-class identity/publication coverage replacing the deleted second registry.
- Remaining known risks: no known change-specific correctness gap. The unrelated repository-baseline gates above remain cleanup debt; the release workflow remains the authoritative generated-artifact build.