Status: Accepted · Audience: runtime contributors, advanced users Answers: how does a flow actually run, and how is it still correct after a process dies?
FlowX.Compiler turns a Define(...) body into static data. Nothing about the
shape of a flow is decided at run time.
// GENERATED — obj/generated/FlowX/PlaceOrderFlow.Plan.g.cs
internal static class PlaceOrderFlowPlan
{
public static readonly ExecutionPlan Instance = new(
FlowId: "order.place",
Version: new SemVer(1, 2, 0),
Profile: ExecutionProfile.Durable,
Steps: new StepPlan[]
{
new(Id: 0, Capability: CapabilitySlot.OrderValidate,
Compensation: CapabilitySlot.None, Policies: PolicyChain.Empty),
new(Id: 1, Capability: CapabilitySlot.InventoryReserve,
Compensation: CapabilitySlot.InventoryRelease, Policies: PolicyChain.Retry3),
new(Id: 2, Capability: CapabilitySlot.PaymentCapture,
Compensation: CapabilitySlot.None, Policies: PolicyChain.PaymentGateway),
new(Id: 3, Kind: StepKind.Emit, EventType: EventSlot.OrderPlaced),
},
Deadline: TimeSpan.FromSeconds(30));
}The Flow Engine's inner loop is therefore an index walk over a
ReadOnlySpan<StepPlan> — no dictionary lookups, no delegate allocation, no
virtual dispatch beyond the capability call itself.
flowchart LR
A["Define() body"] --> B["Compiler: symbol graph"]
B --> C["Validation:<br/>contracts, cycles, determinism"]
C --> D["ExecutionPlan<br/><i>static readonly</i>"]
C --> E["Dispatch switch<br/><i>zero reflection</i>"]
C --> F["Trigger bindings"]
C --> G["flowx.manifest.json"]
D --> H["Flow Engine at run time"]
E --> H
| Engine | Owns | Does not own |
|---|---|---|
| Trigger | envelope normalisation, admission, dedup, input binding | business logic, retries |
| Flow | plan walking, state machine, compensation, deadline | policy semantics, I/O |
| Policy | stage chain execution, retry/breaker/cache/authz | which policies apply (compile-time) |
| Capability | resolution + invocation, scope management | orchestration |
| Event | outbox write, publication, schema check | transport specifics (plugin does that) |
| Scheduler | timers, cron, delayed signals, leader election | flow semantics |
| Stream | offsets, windows, watermarks, backpressure | per-record business logic |
| Observability | spans, metrics, log scopes, replay feed | exporting (OTel SDK does that) |
| Plugin Host | plugin lifecycle, capability negotiation, health | plugin behaviour |
Each engine is a separate namespace with an interface, a single production
implementation, and its own test class. "One reason to change" is verified by
the NoCyclicDependencies and layering fitness functions.
// Simplified from FlowX.Runtime/Flow/FlowEngine.cs — the shape is normative.
public async ValueTask<FlowResult> ExecuteAsync(
ExecutionPlan plan, FlowContext ctx, CancellationToken ct)
{
using var flowSpan = _obs.StartFlow(plan, ctx); // no-op when unobserved
var completed = new StepCursor(plan.Steps.Length); // stack-allocated bitmap
for (var i = ctx.ResumeFromStep; i < plan.Steps.Length; i++)
{
ref readonly var step = ref plan.Steps[i];
ctx.Deadline.ThrowIfExceeded();
var result = await _policy.RunAsync(in step, ctx, _capabilities, ct);
if (result.IsFailure)
return await CompensateAsync(plan, ctx, completed, result.Error, ct);
completed.Mark(i);
if (plan.Profile is ExecutionProfile.Durable)
await _journal.CommitStepAsync(ctx.InstanceId, i, result, ctx.State, ct);
}
return FlowResult.Success(plan.Project(ctx));
}Three properties this shape guarantees:
- Resumability — re-entering this loop is the only difference between a fresh
run and a recovery. There is no separate recovery code path to rot. The sketch
shows a scalar
ctx.ResumeFromStep; the shipped seam does not use one. ADR-0015 commitment 2 derives the position by asking the committed journal rows whether a(scope, step)is done, because a scalar cursor cannot describe a half-completedParallelfork — a shape the DSL shipped in P1. The loop still enters at index 0 and skips what has committed. - Deadline propagation — checked at every step boundary and passed into every policy; a flow cannot outlive its budget by more than one step.
- Compensation symmetry —
completedis the exact set to unwind, in reverse order.
A conditional does not change that shape; it changes only how i moves.
When/Otherwise compiles into the same flat array as everything else: a
Branch step carrying the index to continue at when its predicate is false, and
a Jump step closing the then block so a taken branch skips the alternative.
The loop therefore becomes a while whose index advances either by one or to a
target — the engine holds no branch stack, never recurses, and allocates nothing
to take a branch. Targets are validated forward and in range when the graph is
built, which is the only reason the loop is guaranteed to terminate: the DSL
cannot express a loop, so a backward target is always a layout bug rather than
something an author asked for.
Switch is the same shape with more destinations: one Switch step carrying a
target per case plus a default target, and a Jump closing each case block. The
loop still advances by one or to a target, still holds no branch stack, and still
allocates nothing — taking a case is one bounds check, one read out of the node's
CaseTargets, and one assignment to i. Every case target is validated forward
and in range alongside the default, because a termination proof that covered only
the arm nobody takes would be the wrong way round.
Predicates are evaluated through IStepDispatcher.Evaluate, which is
synchronous and returns bool; a switch's selector runs through
IStepDispatcher.Select, which is synchronous and returns the matching case's
position, or -1 for none. An awaitable predicate would put a state machine on
the hot path and would invite exactly the IO FLOWX1011 forbids; a signature
that cannot express IO is cheaper to enforce than a diagnostic that reports it.
Select returns an int rather than the value it selected for a related reason:
returning the value would mean returning it as object, which boxes an enum on
every switch a flow takes.
flowchart TD
Q1{"Must the flow survive<br/>a process crash?"}
Q1 -- no --> E["Ephemeral<br/>~1 µs/step, 0 alloc<br/>state in memory"]
Q1 -- yes --> Q2{"Continuous input<br/>from a stream?"}
Q2 -- yes --> S["Streaming<br/>checkpointed offsets<br/>windows + backpressure"]
Q2 -- no --> Q3{"Does it wait for<br/>external signals or timers?"}
Q3 -- yes --> D2["Durable + Suspendable<br/>zero resources while waiting"]
Q3 -- no --> D1["Durable<br/>~1 ms/step<br/>journaled per step"]
style E fill:#2e7d32,color:#fff
style D1 fill:#1168bd,color:#fff
style D2 fill:#1168bd,color:#fff
style S fill:#6a1b9a,color:#fff
Ephemeral |
Durable |
Streaming |
|
|---|---|---|---|
| State | pooled in memory | journal, per step | checkpoint, per batch |
| Crash | lost; caller retries | resumes on another node | resumes from checkpoint |
| Overhead/step | ~1 µs | ~1 ms (group-committed) | amortised per batch |
| Suspend/resume | no | yes (timers, signals) | n/a |
| Determinism required | no | yes | yes within a batch |
| Typical use | queries, CRUD, validation APIs | payments, sagas, onboarding | ingestion, enrichment, aggregation |
Default is
Ephemeral. Durability is a cost you opt into, per flow. This is the main architectural difference from Temporal-style engines, which make everything durable and charge everything for it. See ADR-0003.
Warning
The Durable column now describes something that runs, one cell excepted; the
Streaming column has no engine at all. Until WP-52 (2026-07-31) this box said
only the Ephemeral column described something that runs: FlowX.Runtime read
ExecutionProfile nowhere, so declaring Durable got you the ephemeral engine
with a different word in the manifest. It then said the seam existed but that
nothing acquires a lease, nothing scans for an abandoned instance, and the
only IFlowJournal in the repository was an in-memory reference — "no store, no
second node". Both of those states have expired.
What runs today: the runtime reads the profile; a Durable flow takes a lease
before its instance is opened, commits one journal row per step boundary under
that lease's fencing token, and gives the lease back (WP-55); a host that can scan
sweeps for instances a dead node left running and finishes them through the same
step loop; and plugins/FlowX.Postgres (WP-53) persists all of it in PostgreSQL.
So "resumes on another node" in the Crash row has a mechanism, and it is
covered by DurableHostTests — with two hosts in one process, against a shared
store. No node has been killed and nothing crosses a process boundary; the chaos
rig that would show this under a SIGKILL is WP-50 and has not started
(21 §8).
"Suspend/resume — yes (timers, signals)" is a mechanism. This box said the whole
cell was design, refused at build time; then, for a day, that half of it was. Both
halves run since WP-63 (2026-08-02). A Durable flow that reaches
.AwaitSignal<T>(timeout) or .Delay(duration) suspends — the invocation returns,
the instance is sealed Suspended at its resume frontier holding no thread, no
pooled context and no lease, and it records which wait it is parked at and when it
is due. FlowHost.SignalAsync re-enters the same step loop FlowRecoveryScan
uses; so does FlowTimerScan, which is the sweep that finds the instances whose
instant has passed. §6 is the
account, including what a sweep interval means for how precise a wait is.
A Durable flow with no journal is refused, not run ephemerally:
flow.durability_not_configured, before its first step. Since WP-55 that is a
statement about a host that registered no stores rather than about the platform —
a host with a journal and a lease store runs durable flows; one with neither
refuses them, which is the right answer for an unconfigured deployment. See
ADR-0015#what-wp-52-landed-and-what-it-did-not)
for the in/out list, ADR-0016 for what
the schema looked like against a real database,
§5 for what it means for replay, and
20-Roadmap for the phases.
The compiler still says so for Streaming. Declaring Streaming is reported
as FLOWX1028, which was narrowed to that one
profile at WP-52 rather than deleted — deleting it outright would have handed
Streaming the silence Durable had. It is a warning rather than an error on
purpose: the declaration is what P7 has to find and honour, and an error would
push every author to delete it to buy back a build.
Durable flows are replayed. Replay is only correct if flow logic is deterministic. FlowX draws the boundary explicitly:
flowchart LR
subgraph det["Deterministic zone — replayed"]
F["Flow Define() body"]
R["Routing / branching decisions"]
P["Projection / Return expression"]
end
subgraph nondet["Non-deterministic zone — journaled, never replayed"]
C["Capability bodies (I/O)"]
T["ctx.UtcNow"]
I["ctx.NewId()"]
RND["ctx.Random"]
end
det -->|"invokes"| nondet
nondet -->|"result recorded in journal"| det
| Rule | Diagnostic | Severity in Durable |
Built? |
|---|---|---|---|
No DateTime.Now/UtcNow, DateTimeOffset.Now/UtcNow in flows or capabilities |
FLOWX1007 |
Error | yes — WP-58 |
No Guid.NewGuid(), Random.Shared |
FLOWX1008 |
Error | yes — WP-58 |
| No mutable static state reachable from a flow | FLOWX1009 |
Error | yes — WP-58 |
Flow branching, step input mappings and the Return projection may only read ctx.State and step results |
FLOWX1011 |
Error | yes (Warning in Ephemeral) |
Anything in ctx.State must be serialisable by a generated STJ context |
FLOWX1006 |
Error | yes — WP-59 |
Outside a durable flow they are Warnings, not Info. This paragraph said they
"drop to Info — there is no replay, so there is no determinism obligation", which is
ADR-0003's clause repeated, and it is the clause
the set was re-decided against. The decision is Warning by default, Error where the
compilation can prove the code is on a durable flow's replay path, never
informational — for a flow that is its own Profile; for a capability, which has no
profile, it is a Durable flow in this compilation naming it as a step, directly or
through a sub-flow. The reasoning is written once, on
the diagnostics index, and
ADR-0003's determinism bullet records that its own
wording is superseded. Repeating it here is what let the two disagree for two phases, so
it is not repeated again.
All four no rows are now yes, and three of them turned on the severity question
this section asked. They were blocked on severity, not on analysis: Ephemeral
was the only profile the runtime executed, ADR-0003 makes FLOWX1007–FLOWX1009
informational there, and an Info diagnostic never reaches a build log — so they would
have shipped doing nothing anywhere. WP-52 removed that premise and WP-58 wrote the
rules: DeterminismAnalyzer and AmbientReads in src/FlowX.Compiler/Analysis/, with
both directions pinned by DeterminismAnalyzerTests. FLOWX1006 was the row that stayed
no longest, and it was never blocked on severity: it checks membership in the generated
System.Text.Json context that
ADR-0015 commitment 5 requires
payloads to be written through, and until WP-59 emitted that writer there was no
membership to check. It is an error uniformly rather than Warning-then-escalate,
because it reports only on a Durable flow and so cannot fire where the set would warn.
FLOWX1012 was the last id in this family and shipped at
WP-60, once the fix it recommends — Profile = Durable, plus a journal and a lease store
registered on the host — stopped being a lie. Its severity is a Warning and is
deliberately not the set's rule restated: it reports precisely because the flow is not
durable, so the escalation above can never apply to it, and the one escalation that looks
plausible — a durable parent composing the flow — is the case rule 4 and WP-57 below say the
engine does not honour.
Its own section on the diagnostics index
carries that argument, kept apart from this one.
FLOWX1011 stopped being an exception, and nothing about it changed to do so. It
shipped as a Warning outside Durable rather than Info, deviating from ADR-0003, because
Info would have made it invisible in every build anyone could run. The stance for the
whole table was then re-decided as a set at WP-58 — the revisit this paragraph asked
for, rather than the one row at a time it was written to prevent — and the conclusion was
that the deviation had been right all along: Warning is the rule and FLOWX1011 is an
instance of it. See
FLOWX1011.
FLOWX1011 also carries more weight than it did. ADR-0015 originally said the
journal must record the branch a Switch took; it has no field for one, and
the amendment#amendments-the-first-implementation-forced-wp-52)
resolved it by replaying the selector against the restored state bag instead. Replay
of control flow now depends on this rule's purity guarantee — and the rule is a
Warning, and says nothing about capability bodies.
Its row covers the whole deterministic zone drawn above, and not only the "routing /
branching decisions" box: the Return projection named in the diagram, the Switch and
ForEach selectors, the Emit and EmitOnFailure payload maps, and the
Step<TCapability, TStepIn> and SubFlow input mappings the replay contract below
requires to be byte-identical. The full list, and what the rule provably cannot see, is on
its page.
Replay contract: replaying a completed durable instance must produce byte-identical step inputs and identical control flow.
Something verifies it now — WP-61, on 2026-07-31. This box said "Nothing verifies it.
ReplayDeterminismTestdoes not exist"; it does, asReplayDeterminismTestsintests/FlowX.Runtime.Tests, and the corpus it is named for exists with it. Eight shapes — linear,When/Otherwise,Switch,ForEach,Parallel,SubFlowinline andDetached, a failure with its unwind, and a flow that reads all three ambient sources — are each run twice: once against the world, once against the journal the first run wrote. What is compared is every action the flow took and every row the journal holds, including the wholeNondeterminismCapture. The replay's clock is deliberately a hundred days from the original's, so a value that was not replayed cannot be mistaken for one that was, and the corpus cannot pass by comparing nothing — which is the failure mode the exit criterion named.Two earlier versions of this box are now wrong, and both are recorded rather than deleted. It said there was no journal type in the solution; WP-51 declared
IFlowJournal. It then said nothing wrote to one; WP-52 (2026-07-31) made the runtime readExecutionProfile, and aDurableflow now journals a step boundary, capturesctx.UtcNow,ctx.NewId()andRandom's seed per step, and resumes by replaying its committed rows into this same loop.The capture is now read as well as written, which is what WP-61 had to buy before it could compare anything:
FlowExecutionContext.ReplayNondeterminismhands a step the instant, the ids and the seed its row records, andctx.UtcNow,ctx.NewId()andctx.Randomanswer from them. This paragraph said "the capture is written and never read, which is exactly the state in which a determinism leak leaves no trace". That state is over. What still does not happen is the engine calling it: the step loop skips a committed step rather than re-running it, so no resumed execution reaches a step it holds a capture for, and a replay has to be driven from outside the loop.flowx replay(WP-64) is the driver that ships.Three fidelity limits are now measured rather than predicted, each pinned by a test that goes red the day it is fixed.
- A
Parallelwhose branches genuinely overlap does not replay. One pooled context is shared by every branch, so a capture can land on a sibling's row — ADR-0015#what-wp-52-landed-and-what-it-did-not) called this best-effort attribution and said WP-61 needed a per-branch context before it was safe. WP-61 did not buy one. It pinned the exact interleaving with a rendezvous instead, so the misattribution is reproduced on every run rather than one in twenty, and it measured the consequence: the branch whose id was taken by its sibling mints a fresh one on replay. A fork whose branches do not overlap replays exactly, and that is the whole of what "aParallelreplays" currently means.- A compensation's ambient reads are captured by nothing. The forward path takes an envelope at every step boundary;
CommitCompensationAsynctakes none, so an undo that readsctx.UtcNowleaves no record of what it read and no replay can reproduce it. This was named in no document before WP-61 measured it.- The engine's own deadline check is not replayed. It reads the clock before any hook outside the loop is reached, so it sees the replaying node's time. Everything the flow reads is replayed; the engine's own read is not.
A fourth gap closed on 2026-08-01, and it is worth naming here because this paragraph used to be the place that said so. It read: "
AwaitSignalis refused at build time under every profile … so a durable flow still runs to completion inside one invocation." It does not. A durable flow suspends at its suspension point and is resumed by a signal or by a clock, and what a resumed invocation replays is what every other resume replays — the committed rows, the state-bag snapshot and the control-flow delegates. The delivered signal itself is journaled as theAwaitSignalstep's own row, so it is replayed by the same mechanism as a step's output and adds no fourth fidelity limit. A timer adds none either: the instant is on the instance row rather than in the flow's own state, so nothing about it is replayed — it is read once, by the engine, when it arrives back at the node it is parked at.FLOWX1017still refusesAwaitSignalon a flow that journals nothing, and now refusesDelaythere too; the rule that used to refuse both atDurableis deleted.
Note
All of this section runs. This box read "nothing in this section runs" until WP-63 landed its signal half on 2026-08-01, and "the clock half does not" until it landed the rest on 2026-08-02.
| Construct | What the compiler and the engine do with it |
|---|---|
.AwaitSignal<T>(timeout) |
Honoured. The plan carries the author's declared duration and FlowEngine stops at the step. A delivered signal satisfies it; the declared duration expiring takes the .OnTimeout block, or ends the flow with flow.signal_not_received when there is none |
.OnTimeout(block) |
Honoured. The block is laid out immediately after the wait, and the wait carries the index a delivered signal jumps to. Both paths rejoin there, so an escalation that should end the flow says .Fail(...) — exactly as it would inside an .Otherwise(...) |
.Delay(duration) |
Honoured. One step of its own, carrying the author's expression. The instance parks and a sweep brings it back |
.PollUntil<T>(until, interval, timeout) |
Honoured. The same parking, re-entered once per attempt: the body is the capability at the next index and the attempt number is its journal scope, so the loop needs no backward target and the step loop's termination argument is untouched. The budget is measured from the instant the first attempt read, which is on its row |
What a wait costs. The instance is sealed Suspended with its committed prefix
intact; the lease is released rather than renewed for the length of the wait; the
pooled context goes back to the pool. What it gains is three values on its own row —
which step it is parked at, in which iteration, and when it is due. That is the
whole of a durable timer, and it is why there is no timer table: a wait is a
property of the instance that is waiting, and a separate table would need its own
transaction with the write that parks the instance, which is not optional — an
instance parked with nothing scheduled to wake it is a wait that never ends.
What brings it back. FlowHost.SignalAsync for a signal; FlowTimerScan for a
clock. Both call the same FlowHost.ResumeAsync a recovery scan calls, which enters
the same FlowEngine.ExecuteAsync a fresh instance enters. The sweep decides which
instances to look at again; the engine, standing on the node the instance is parked
at, decides whether the wait is over — so a sweep never has to know what a suspension
point means. There is also no signal table: a delivered signal is journaled as
the AwaitSignal step's own row, through the same CommitAsync every step boundary
uses.
A wait is a lower bound, never an upper one. The sweep runs on
FlowXOptions.TimerScanInterval (ten seconds by default, jittered), so a
.Delay(TimeSpan.FromSeconds(1)) elapses somewhere between one and eleven. That is
the same promise a scheduled trigger makes and the only one a sweep can keep. A host
whose journal implements no ITimerIndex makes no promise at all: it parks
instances correctly and never wakes them, which FlowTimerScan.IsEnabled reports.
samples/workflow's offer.accept is the running version of the flow below —
.AwaitSignal, an .OnTimeout that withdraws the offer, and a .Delay before
onboarding — and tests/Workflow.Tests/SuspensionTests measures it against the real
host, the real journal and the real sweep.
One limit worth knowing before you write one, and one that has gone. An inline
composed child may not suspend — a parent records a composition as one row written
when the child finishes, so a parent resumed past a waiting child would compose a
second child instance and repeat its effects; it is refused as
flow.suspension_inside_composition, and a Detached child is allowed to wait.
The second limit read "a flow with an [HttpTrigger] should not suspend … and
plugins/FlowX.Http has no path for it yet". It has one at WP-64: the run endpoint
answers 202 with the instance and where to continue it, and the compiler generates
one delivery route per signal — POST {flow route}/{instanceId}/signals/{identity},
which is the arrow drawn in the diagram below
(ADR-0022).
protected override void Define(IFlowBuilder<OnboardCustomer, OnboardResult> flow) => flow
.Step<CreateAccount>()
.Step<SendVerificationEmail>()
.AwaitSignal<EmailVerified>(timeout: TimeSpan.FromDays(3))
.OnTimeout(f => f.Step<ExpireOnboarding>())
.Step<ActivateAccount>()
.Return(ctx => new OnboardResult(ctx.Get<AccountId>()));sequenceDiagram
autonumber
participant FE as Flow Engine
participant J as Journal
participant SE as Timer sweep
participant EXT as External caller
FE->>J: state=Suspended, awaiting EmailVerified, deadline +3d
FE->>SE: the same write records wake_at = +3d, wake_step_id = the wait
FE->>FE: release lease, free context to pool
Note over FE: zero threads, zero memory held anywhere
EXT->>FE: POST /flows/42/signals/EmailVerified
FE->>J: commit the AwaitSignal step's own row, carrying the payload
FE->>FE: the same ExecuteAsync a recovery scan uses resumes at the next step
alt no signal within 3 days
SE->>FE: a sweep finds the row due
FE->>FE: run OnTimeout branch → ExpireOnboarding
end
Every arrow in that diagram is built, and two of them are drawn loosely. "register
timer" is not a call to a scheduler engine: it is the same write that records
state=Suspended, carrying three more columns, which is what makes parking an instance
and scheduling its wake one transaction rather than two. And "timer fires" is a sweep
finding the row rather than a callback the process was holding — FlowTimerScan runs on
an interval and asks the journal which parked instances are due. The arrow that used to
read "append signal, state=Pending" is corrected rather than aspirational: a signal is
not appended to a table of its own, it is the suspension point's journal row.
A suspended instance costs one row. A million suspended onboardings cost a
million rows and zero compute — this is what makes long-running business
processes affordable. This paragraph used to end "that is the argument for building
it, not a description of what it does", and then described only the signal path. It
now describes both: offer.accept in samples/workflow sends an offer out and
suspends holding no lease; a countersignature arriving on another request resumes it
into a settling period, where it parks again on a clock; and a sweep wakes it to start
onboarding. An offer nobody signs is picked up by the same sweep when its window
closes, and the escalation withdraws it — against PostgreSQL, with the instance
Suspended and one committed row in between each pair.
flowchart TD
S1["step 1 ✅"] --> S2["step 2 ✅"] --> S3["step 3 ❌ fails"]
S3 --> C2["compensate step 2<br/>(own retry policy)"]
C2 --> C1["compensate step 1"]
C1 --> END["state = Compensated"]
C2 -- "compensation exhausted" --> DLQ["state = CompensationFailed<br/>dead-letter + alert + manual replay"]
Rules:
- Compensation runs in strict reverse order of successfully completed
steps only. A failed step is never compensated (it did not take effect — or if
it did, that step is not idempotent and that is a capability bug). A step a
Fallbackanswered for is not compensated either, and it is the one case where "successfully completed" and "its capability ran" come apart: neither the step's capability nor the fallback's may declare a side effect (FLOWX1053), so there is nothing an undo could reverse. Since WP-80 a resumed instance agrees — the degraded row names the capability that answered, so the stack a new node rebuilds is the stack the forward path built; a constant fallback writes no such name and is the exception ADR-0079 §2.5 records. - Each compensation carries its own policy chain —
StepNode.CompensationPolicies, resolved against the compensating capability, soFLOWX1014's idempotency rule is checked against the thing that would actually be run twice. Compensation retry is more aggressive than forward retry by default (5 attempts vs 3):PolicySet.CompensationDefault. Since WP-57 the runtime executes it — the only policy it executes anywhere, at theConsistencystage (10 §2). A policy applies because it was declared: an undo with no declared chain is still attempted exactly once.PolicySet.CompensationDefaultitself did not reach a plan untilPolicySetReaderlearned to resolve a set declared in a referenced assembly: it lives inFlowX.Abstractions, its symbol carries no syntax in a consuming compilation, and the emitter wrote no chain for it. So for two releases the set this rule names was the one way to retry an undo that could not work, while a hand-written.CompensationRetry(attempts: 5)did. Two attempt counts do not work and are diagnosed rather than silent:attempts: 1retries nothing, becauseIsRetryingisAttempts > 1(FLOWX1035), and a second.WithPolicy(...)on one step discards the first (FLOWX1034) — so a compensation default applied beside a step's own set deletes that set rather than adding to it. - Compensation is best-effort but loud: exhaustion produces
flowx_flow_compensation_failed_total, a dead-letter record, and a documented operator recovery path. Half of this is now true. WP-57 shipsICompensationAlertSink, raised exactly once per exhausted compensation and carrying the flow, the instance, the step, the compensating capability, the attempt count and the last error — everythingflowx replay --instance <id> --from <step>needs. The metric and the dead-letter record are not built: FlowX ships no metrics infrastructure yet, and the sink is the seam the observability package attaches to rather than a counter invented inside the engine. - Compensation is itself journaled, so a crash during compensation resumes compensation — never re-runs forward steps.
Ephemeralflows may declare compensation, but the guarantee is weaker: a process crash loses the pending unwind along with the instance, because the stack is a field of an in-memory context and nothing writes it down.FLOWX1012says so at build time since WP-60, as a warning — this rule was specified alongsideFLOWX1017in ADR-0003, only one of the pair was built, and for two phases a compensableEphemeralflow compiled in silence. It fires on the default profile as well as on a declaredEphemeralone, which is the whole reason it is not an error; its page carries that argument and the reference sample's answer to it. Rules 1–3 above are implemented and covered byCompensationStackTests,FlowEngineTestsandCompensationPolicyTests. Rule 4 is implemented for a flow's own steps and still open across a composition. Since WP-57CompensateAsynccommits one row per undo attempt —JournalOutcome.Compensatedwhen it worked,Failurewhen it did not, keyed past the forward row it reverses, carrying the compensating capability's id and moving the instance toCompensating— and a resumed instance does not repeat an undo whose row already committed. Two gaps remain, both named rather than papered over: an undo whose row never landed re-runs, which is the same honest limit ADR-0006 states for a forward effect that landed before its commit; and a composed child that already succeeded records nothing, because its instance was sealedCompletedand a journal correctly refuses a write to a finished instance.- A compensation runs under its own identity.
ctx.CapabilityIdis the compensating capability —payment.refund, not thepayment.capturebeing reversed — and the step being undone isctx.CompensatingFor,nullon the forward path. This rule is here because the engine got it wrong. It entered a compensation with the forward step node, so an undo deriving an idempotency key fromctx.CapabilityId— which is what 19 §1 shows and what the reference sample did — produced the forward step's key byte for byte. Any store honouring that key deduplicated the contra write away, the capability returned success, and the engine recordedCompensationOutcome.Succeededover an effect that never happened. The journal row was right throughout, which is exactly what made it invisible: the audit trail named the compensation while the code that ran named the step. Both facts are onCapabilityContextnow, because both have readers and neither is derivable from the other at the point of use — a compensator needs its own identity to key a write, an operator reading a trace needs to know what is being reversed. The two ids come fromStepNode.IdentityandStepNode.CompensationIdentity, one expression each, so the row and the context cannot drift apart again.
flowchart TD
A["Capability returns Result.Fail(Error)"] --> B{"ErrorCategory"}
B -- "Validation / NotFound / Forbidden" --> C["Terminal — no retry"]
B -- "Conflict / Unavailable / Internal" --> D["Retry policy applies"]
D --> E{"Retries exhausted<br/>or deadline hit?"}
E -- no --> A
E -- yes --> C
C --> F{"Prior compensable steps?"}
F -- yes --> G["Compensate"]
F -- no --> H["FlowResult.Failure"]
G --> H
H --> I["Trigger maps to transport:<br/>HTTP 4xx/5xx · gRPC status · DLQ"]
X["Capability throws Exception"] --> Y["Runtime catches at the capability boundary"]
Y --> Z["Error{Internal, code=capability.unhandled}<br/>+ exception logged with trace id<br/>+ flowx_capability_unhandled_total"]
Z --> D
An unhandled exception is never silently converted into a business error. It is recorded as a defect signal with its own metric, because a capability throwing is a bug in that capability.
flow.Parallel(p => p
.Branch<CheckCredit>()
.Branch<CheckFraud>()
.Branch<CheckSanctions>(),
merge: MergeStrategy.AllMustSucceed)
.Step<ApproveApplication>();| Strategy | Semantics | Failure behaviour |
|---|---|---|
AllMustSucceed |
wait for all branches | first failure cancels siblings via linked token |
AllSettled |
wait for all, collect results | flow continues; branch errors available in context |
FirstSuccess |
race | remaining branches cancelled |
Quorum(n) |
first n successes | remainder cancelled |
MergeStrategy is a struct, not an enum, because Quorum(n) carries a number. The
three parameterless strategies are static properties, so merge: MergeStrategy.AllMustSucceed
reads exactly as it does above; MergeStrategy.Quorum(2) is a factory call.
default(MergeStrategy) is AllMustSucceed — the strictest of the four, because silence
should not buy leniency.
AllSettled is the only strategy that writes something for the next step to read: the
engine puts a ParallelOutcome into the context at the join, carrying each branch's error
or null in declaration order. Read it with ctx.Get<ParallelOutcome>(). A flow with
two AllSettled forks keeps the later one, because the state bag is keyed by type — the
StepIndex on the outcome is there so a reader can tell which fork it is holding.
Quorum(n) with n greater than the branch count fails the fork rather than waiting for a
branch that does not exist. MergeStrategy cannot check that itself; it does not know how
many branches there are.
A fork uses the same flat layout as a switch: one node carrying a target per block, and
the blocks laid out after it. The only structural difference is that a fork's blocks need
no closing jumps — a branch runs as a bounded range [BranchTargets[k], BranchTargets[k+1]),
so the next branch's target is the bound, and a jump would only restate it.
0 validate
1 parallel(→2,3,4 join 5) ← one node, three branch targets, one join
2 check credit ← branch 0 = [2,3)
3 check fraud ← branch 1 = [3,4)
4 check sanctions ← branch 2 = [4,5)
5 decide ← the join
Everything the flat model bought survives: no branch stack, no nested plan objects, and the
termination proof is unchanged — every target still points strictly forward, StepGraph
still rejects one that is out of range, and each branch is a forward-only walk over a
disjoint sub-range.
What does not survive is the claim that "the loop index" describes execution. Between a fork and its join there are several indices, on several threads. The engine recurses exactly once per fork, into the same range-walking method with different bounds. A flow that does not fork never reaches that code.
Branches are as concurrent as their steps are. Each branch is started eagerly on the
calling thread and runs until its first incomplete await; the rest interleave on the
thread pool. So a branch whose every step completes synchronously finishes before its
sibling starts — which is correct, and is why a fork is for I/O rather than for CPU work.
Forcing branches onto Task.Run would buy a thread-pool dispatch per branch to make
CPU-bound work contend.
Every branch is drained before the fork returns, cancelled or not. The context is pooled, so a branch still writing after the engine had moved on would eventually write into the next flow's context — possibly another tenant's. It is also what makes compensation sound: nothing is still running when the unwind starts.
A fork allocates. A linked token source, a Task per branch and their awaiters — about
240 B per branch on ~70 B fixed, measured. Budget B2 stays a hard zero for the linear,
conditional and switch paths, which never reach this code; see
14 §1.1.
AllMustSucceed cancels its siblings on the first failure, and a cancelled sibling may
already have completed compensable work. That work is still compensated: branch steps
record onto the flow's single compensation stack as they complete, and CompensateAsync
runs under CancellationToken.None precisely so that the token which stopped the work
cannot also stop the undoing of it.
The unwind order needs one caveat stated plainly. "Strict reverse order" is exact within a branch, and between a branch and everything sequential around it — the fork and the join establish that ordering. Two steps in two different branches have no such ordering between them, so their compensations may unwind in either order. That is not a defect being hidden: branches that write disjoint slots have nothing to order against each other, and a saga that needed one undone before the other was not expressing concurrency in the first place.
Parallel branches share the flow's deadline and its context, and write to disjoint
context slots — enforced at compile time (FLOWX1013) so
parallel steps cannot race on shared state.
The runtime's half of that bargain is narrower than the compiler's, and the difference
matters. FlowExecutionContext holds a plain Dictionary<Type, object>; it is pooled, and
a ConcurrentDictionary would have cost every linear flow an allocation per write to solve
a problem only forks have. So the state bag and the compensation stack are guarded by a
lock, and only when more than one thread can reach them — ExecutionPlan.HasParallel, a
fact precomputed at type initialisation. A flow that never forks takes one always-false
branch and costs exactly what it did before.
A ForEach whose MaxDegreeOfParallelism is above one counts as forking, for the obvious
reason and by exactly the same mechanism: it reuses the fork's linked token source, its
WhenAny drain and this lock, rather than growing a second answer to the same problem. A
loop bounded at one does not, and keeps the unguarded path — so the question the flag
answers is not "does the flow branch" but "can two threads reach the context".
That makes concurrent writes safe: the dictionary cannot be corrupted and a torn read is
impossible. It does not make them meaningful. Two branches writing the same contract type
still race, and the winner is whichever finished last — which is a modelling mistake, and
FLOWX1013's job. Read that page's stated limits before treating a green build as a
proof: it compares declared output contracts, so a capability calling ctx.Set<T>()
from inside its own body writes a slot the rule never sees.
One fidelity limit, stated rather than papered over: ctx.CapabilityId is a single field on
the shared context, so inside a parallel branch a capability may read whichever step most
recently started, including a sibling's. Error attribution does not depend on it — the
engine takes the identity from the step node it is executing — but a log line written by a
capability does. Per-branch identity needs a per-branch context, which is a larger change
than this shape.
In Durable flows, each branch's steps commit their own journal rows. This sentence went
on to say "the merge point is a single checkpoint", and there is no such checkpoint:
neither the fork nor the join writes a row of its own — FlowEngine runs the branches and
jumps to the join index. Nothing is lost by that, and it is the design rather than an
omission: the branches share the fork's cursor, their spans are disjoint, so
(scope, step_id) stays unique without a per-branch scope and the frontier reconstructs
"branch A done, branch B stopped at step 12" purely from which rows exist. A merge
checkpoint would be a second place the same fact is written, and the stored copy would be
the one nothing checks — the same argument
ADR-0015 commitment 2 makes for
deriving the resume position instead of remembering it.
flowchart LR
src["Source<br/>Kafka partition"] --> ch["Bounded channel<br/>capacity = N"]
ch --> proc["Flow instances<br/>degree = D"]
proc --> sink["Sink / commit offset"]
ch -. "channel full" .-> pause["Pause partition consumption"]
pause -. "below low-water mark" .-> src
The Stream Engine never grows an unbounded queue. When the channel is full it
pauses consumption at the source (Kafka Pause, AMQP prefetch, MQTT flow
control) rather than buffering in memory.
Running since 2026-08-02, and narrower than the diagram. This box said there was no Stream Engine and no bounded channel anywhere in
src/.FlowStreamScancreates one per subscription and — this is the part the diagram understates — issues no read at all when it is full, so the backlog stays in the source rather than in a batch already in flight.StreamBackpressureTestsis what holds it to that: a deliberately slow flow, and an assertion on how many records the source was asked for beyond what the flows consumed. It is written so that it fails if the bound stops working, by running the same workload at two capacities and asserting the peak follows the capacity."Pause partition consumption" is not what happens, because there are no partitions. One subscription is read by one node, in order, under a lease; a stream is paused by not being read. Kafka
Pause, AMQP prefetch and MQTT flow control are still design — there is noFlowX.Kafka, and the oneIStreamSourcethat ships is Redis.Two bounds, both configurable and both refusals rather than buffers:
FlowXOptions.StreamChannelCapacityandStreamMaxResidentRecords. Exceeding the second stops the subscription withstream.window_overflow(ADR-0055). B13 is still unmeasured — the subsystem exists, the harness does not.
Offset commit is after flow completion, giving at-least-once semantics; what makes it
effectively-once is not capability idempotency but the instance id a closed window derives —
a rebuilt window meets flow_instance's primary key
(ADR-0055). And the checkpoint is not
"the last record read": it is the last record of the longest prefix whose windows have closed,
which is what lets a restart rebuild an open window rather than lose it.
| Source | Effect |
|---|---|
| Client disconnect (HTTP) | linked token cancelled; ephemeral flow aborts, durable flow continues (the business operation was accepted) |
| Flow deadline exceeded | TimedOut; compensation runs if compensable |
| Shutdown (SIGTERM) | stop admission → finish in-flight up to grace period → release leases → exit |
| Operator cancel | flowx cancel --instance 42 → Compensating |
Deadlines are absolute, set at trigger time, and never reset by a retry. A
retry that would exceed the deadline is not attempted — the policy engine plans
the backoff first and refuses the attempt when the wait alone would run past the
budget — and a step's Timeout is clamped to whatever is left of the deadline
whenever the deadline is shorter. This prevents the classic
"3 retries × 30 s timeout inside a 10 s SLA" failure, and prevents it in both
directions: PolicyExecutionTests.ARetryNeverOutlivesTheDeadline pins the first
half and StepPolicy.EffectiveTimeout the second.
| Not done | Why | Alternative |
|---|---|---|
| Distributed transactions (2PC) | availability cost, operational fragility | sagas + compensation |
| Automatic capability parallelisation | implicit concurrency is unreviewable | explicit Parallel |
| Dynamic flow modification at run time | breaks P4 and replay determinism | deploy a new flow version |
| Retrying non-idempotent capabilities by default | data corruption risk | [Capability(Idempotent = true)] opt-in gates retry |
| Cross-flow shared mutable state | coupling, unanalysable | events, or an explicit store inside a capability |
Next: 07 — Capability Model