This runbook covers the M11-12/R24 release-candidate objectives. It complements the
dynamic Kind/K3d fault evidence; dashboards and alerts cannot replace a
dynamic_execution=true, non-fixture evidence package for the exact candidate commit.
The versioned source of truth is
distribution/config/architecture-production-readiness-policy.json. Metric names,
units, owners, query labels, thresholds, dashboard panels, alert routes, readiness
dependencies, and collector-outage limits must change together with that contract.
- CommitLog remains the authoritative WAL. Never delete or rewrite WAL, PVCs, queue offsets, or message payloads during rollback.
- A production candidate needs at least six hours of soak evidence sampled every minute. Missing samples must not exceed one percent.
- The evidence package must bind the candidate commit, candidate image digests, M11-11 fault run, SLO policy, dashboard, alerts, runbook, and every generated artifact by SHA-256.
- Any failed objective, unresolved alert, unresolved fault, mismatched commit, or incomplete rollback blocks promotion.
- Release identity uses only
service,release_commit, andrelease_nonce. Credentials, configuration objects, message bodies, topic names, consumer groups, hostnames, and request identifiers are not release-identity labels. /readyzis the canonical service-readiness endpoint. It returns success only after the service-specific security, storage or control-plane dependencies and release identity have completed.- Acknowledgement RPO/RTO is evaluated with
distribution/config/ack-failover-evidence-schema.jsonand the acknowledgement and failover contract. An asynchronous acknowledgement never inherits a zero-loss claim from a synchronous profile. - The project does not claim built-in regional disaster recovery. The exact boundary is recorded in the regional disaster recovery ADR.
Alert: RocketMQReleaseIdentityConflict
- Diagnose: Compare
rocketmq_release_infobyservice,release_commit, andrelease_noncewith the activeReleaseState. Confirm that the conflict is not a stale Prometheus target. - Contain: Stop promotion and prevent another rollout until every workload references one commit, nonce, config digest, Secret version, and storage generation.
- Recover: Remove only the stale workload or scrape target, wait for its
series to age out, and verify all five
/readyzendpoints against the intended release identity. - Escalate: Page the release-engineering and owning service teams when more
than one identity remains active for five minutes or the running identity does
not match the complete
ReleaseState.
Readiness is service-specific but governed by one policy:
- Broker requires validated security, a writable MessageStore, started request processors and listeners, NameServer registration, and release identity.
- NameServer requires validated security, route processing, the optional embedded Controller when configured, and release identity.
- Controller requires validated security, recovered storage and Raft state, cluster recovery, and release identity.
- Proxy requires validated security, a healthy metadata route in cluster mode, every configured listener, and release identity.
- MCP requires validated security, initialized application state, the selected transport listener, and release identity.
A bound socket alone is never readiness. When /readyz returns 503, stop the
rollout, inspect the failed dependency, restore that dependency, and wait for a
fresh 200 response. Escalate to the owning service team if the dependency is
healthy but readiness remains unpublished.
Alert: RocketMQMessageDeliveryRatioBurn
- Diagnose: Confirm the alert is not caused by an idle input stream, then check Broker readiness, NameServer route visibility, Controller quorum, Proxy readiness, and the acknowledged send/query probe.
- Contain: Stop promotion when the delivery ratio is below
0.999; stop new writes if the configured acknowledgement contract cannot be preserved. - Recover: Roll back all five workloads to their recorded baseline digests and verify the acknowledged message, queue offset, CommitLog offset, and PVC UID set again.
- Escalate: Page Broker, storage, and control-plane owners if the baseline still misses the objective or the acknowledged message cannot be recovered.
Alert: RocketMQSendMessageP99High
- Diagnose: Compare the functional-probe p99 with
rocketmq_send_message_latency; the release objective is at most1000 ms. Inspect flush/dispatch lag, HA replication lag, disk pressure, CPU throttling, and collector-outage timing before attributing the regression. - Contain: Stop promotion; do not raise the threshold or shorten the soak window to make a candidate pass.
- Recover: Roll back the candidate images when the objective remains exceeded for ten minutes, then repeat the same probe against the baseline.
- Escalate: Page Broker and storage owners when the baseline remains above the objective or durable acknowledgements are delayed.
Alert: RocketMQConsumerLagHigh
- Diagnose: Identify the affected topic and consumer group without exporting
message bodies. Check
rocketmq_consumer_lag_messages, in-flight work, route freshness, consumer connectivity, and dispatch-behind bytes. - Contain: Block promotion when the six-hour maximum exceeds
10000messages and avoid destructive offset changes. - Recover: If the lag began with the candidate, restore baseline images and confirm that the lag is decreasing before closing the incident.
- Escalate: Page Broker and consumer owners if lag does not decrease after baseline recovery or route and storage health disagree.
Alert: RocketMQStoreFlushBehindHigh
- Diagnose: Check
rocketmq_storage_flush_behind_bytes,rocketmq_storage_dispatch_behind_bytes, disk pressure, flush latency, and storage errors. - Contain: Block promotion when flush-behind exceeds
67108864bytes. Keep the Broker available for reads when safe, but do not acknowledge new writes if the configured durability contract cannot be met. - Recover: Roll back executable images and chart revision only. Do not delete CommitLog, ConsumeQueue, Index, RocksDB state, or PVCs.
- Escalate: Page the storage owner for persistent backlog, disk errors, or any mismatch between reported and durable offsets.
Alert: RocketMQHaReplicationLagHigh
- Diagnose: Check
rocketmq_store_ha_replication_lag_bytes, replica connectivity, Controller quorum, confirm offset, and disk/flush health. - Contain: Block promotion when replication lag exceeds
67108864bytes. Do not force a role change that can lose acknowledged data. - Recover: Restore baseline images and verify the same message ID and offsets after quorum recovery.
- Escalate: Page storage and Controller owners if quorum cannot recover or replica and leader offsets remain inconsistent.
The M11-11 policy requires all 16 ordered scenarios in one non-fixture run: rolling upgrade, node eviction, both NameServer availability boundaries, collector outage, disk pressure, disk-full admission, synchronous-write contention, Controller leader loss, Controller quorum loss with duplicate-leader detection, latency/loss/half-open network impairment, HA lag and promotion, interrupted Raft snapshot installation, Proxy long-poll/slow-Broker overload, secret rotation, and acknowledged-message recovery.
Each record must satisfy its declared RPO/RTO and retain the injection,
observable, abort, cleanup, and cleanup-verification evidence. A model test may
drive an otherwise unsafe fault (for example disk-full admission) only when it
uses the production state transition and is executed from the current checkout;
fixture output is never promotion evidence. The run must finish with no netem
qdisc, cordon, disk-pressure taint, synthetic disk-contention process, scaled-down
quorum, rotated credential, candidate image, or unresolved fault.
The M11-11 collector_outage fault scenario must demonstrate a successful
send/query round trip in under 30 seconds while the collector is unavailable. The
telemetry queue remains bounded and the data plane must not wait for export.
The production boundary admits at most 2048 records, 8 MiB in total, and 64 KiB
per record. Admission uses try_enqueue, rejects the newest record when full, and
reports drops through rocketmq.exporter.drop; shutdown reports final queue and
drop totals through rocketmq.exporter.shutdown.
- Diagnose: Inspect accepted, drained, dropped-by-reason, queued-item, and queued-byte measurements together with the send/query probe.
- Contain: Stop promotion if the 30-second budget is exceeded; do not enlarge the queue until resource usage and outage duration are understood.
- Recover: Restore the collector, allow the bounded queue to drain, and repeat the same functional probe.
- Escalate: Page the observability owner if the data plane blocks, queue limits are exceeded, drops are not reported, or recovery remains above 30 seconds.
Never disable the bounded queue or privacy/cardinality guard as a workaround.
The candidate must report zero FailedPreStopHook events and restore:
- all five baseline image digests;
- the baseline Helm chart revision;
- the collector and Controller quorum;
- the original PVC UID set and acknowledged message;
- queue and CommitLog offsets;
- an empty unresolved-fault and unresolved-alert list.
Rollback is complete only after the baseline workloads are ready and the
acknowledged-message query succeeds. Preserve both failed and successful evidence
directories. A failed run must not contain a production run.json that can be
mistaken for a pass.
Pull requests use the affected checks described in the CI validation policy. Record the relevant command results in the PR; no separate candidate JSON, fixed commit, fingerprint, or empty historical-failure ledger is required for routine development.
The historical accepted code/system record is
d88a973131ce4f57d01a65def8ecb7944a45ba21,
with its machine-readable companion.
It records a 93.5 / 100 code/system assessment and is not production certification;
the six-hour soak, target-hardware comparison, complete disaster recovery, Docker
images, and real external adapters remain deferred V1 evidence.
Automated six-hour SLO and Kubernetes fault workflows are retired. Their local runners and evidence guards remain available for an explicitly provisioned, disposable environment with enough capacity for the complete run. Scheduled or manually selected root integration runs still own the broader contract test suites; these tests do not execute a live fault cluster or six-hour soak.
Before a live run, verify published images with
scripts/verify_service_image_publication.py, generate cryptographically random
test credentials with scripts/new-m11-evidence-secrets.ps1, and authenticate any
required registry pulls. Run scripts/kind-architecture-refactor-e2e.ps1 -Mode Run
with -KeepCluster to retain the fault cluster. The caller owns candidate promotion,
cluster deletion, credential cleanup, and registry logout; never upload credential
manifests with the evidence.
The retained fault cluster is promoted to the candidate digests before the soak.
run-architecture-slo-cluster.ps1 then deploys a digest-pinned private Prometheus
that scrapes all five metrics Services and a bounded message send/consume probe.
Prometheus is reached only through a loopback kubectl port-forward. The wrapper
keeps the port-forward owned for the full sampler lifetime and terminates it during
cleanup. Production credentials and an externally reachable metrics endpoint are
not inputs to this isolated evidence environment.
From the repository root, validate policy and synthetic evidence during the later production-validation tier:
python scripts/architecture_slo_guard.py --policy-only
python -m unittest scripts.tests.test_architecture_slo_guard -v
python -m unittest scripts.tests.test_m11_dynamic_evidence -v
cargo test -p rocketmq-observability --test production_readiness_contractThe SLO tests generate synthetic evidence in a temporary repository using the
current policy and release assets, then inject deliberate violations. Updating
the runbook or metric registry does not require refreshing committed fixture
hashes. These samples remain marked fixture=true and status=not-run.
For production evidence, omit --allow-fixture. The guard then requires the
candidate commit to equal the checked-out commit, verifies the embedded M11-11
dynamic fault evidence, compares every objective with the committed threshold, and
checks every SHA-256 entry.