This guide covers testing patterns, rules, and conventions for the MEDRE project. It is the authoritative reference for how tests are written, what each test tier proves, and how to run the suite.
The test suite has 14k+ tests (~125 deselected by the
live/docker/hardware marker policy). Every transport has a
fake adapter that exercises the full pipeline. The standard pytest -q run
requires a generous timeout—the full suite can exceed 600 s on typical
hardware. See the README for project context and the
Operator Workflows for bridge-specific test
commands.
Agent responsibility: Before adding tests to any file, check its current line count (
wc -l <file>). If the file is anywhere near 1,500 lines, create a new file instead. Anything over 1,500 lines causes CI to fail (test_no_file_exceeds_1500_lines). When in doubt, start a new file.
Target: < 1,200 lines per test file. Hard ceiling: 1,500 lines unless explicitly justified in the file and in this guide.
Split files by behavioral domain, not by "coverage" or "misc". When a domain file approaches the target, split it by subdomain following the procedure in the Splitting procedure section below.
There is no oversized-test allowlist. Every test_*.py file stays at
or below 1,500 lines (MAX_LINES). The target remains below 1,200 lines.
If a file approaches the hard cap, split it by behavioral domain following
the procedure in the Splitting procedure section.
Completed splits are listed in the Completed Splits
table as historical record, not as active allowlist entries.
These files are approaching the 1,500-line hard cap and should be split opportunistically by behavioral domain:
test_runtime_snapshot.py(1,301 lines)test_trace.py(1,425 lines)test_replay_recover.py(1,454 lines)- Any other 1,000+ line file should be split when convenient
- Identify distinct subdomains in the file (e.g., delivery vs. failure taxonomy vs. fanout for pipeline tests).
- Create new
test_<module>_<subdomain>.pyfiles. - Move the relevant test functions, fixtures, and imports. Keep shared
helpers in
tests/helpers/if needed by multiple files. - Run
PYTHONPATH=src pytest tests/test_<module>_<subdomain>.py -qto confirm the split tests pass. - Delete the moved code from the original file.
- Run the full suite to confirm nothing is broken:
PYTHONPATH=src pytest -q.
New tests use pytest function style (module-level async def or def),
not unittest.TestCase. Existing TestCase classes are acceptable as-is
but should not be extended with new test methods.
# Preferred: pytest async function
async def test_adapter_delivers_event(fake_adapter, canonical_event):
result = await fake_adapter.deliver(canonical_event)
assert result.status == "success"
# Existing pattern: TestCase classes (do not extend with new methods)
class TestLegacyStorage(unittest.TestCase):
def setUp(self):
self.store = InMemoryStorage()Prefer pytest fixtures for test setup and teardown. Fixtures are composable, support async natively, and have explicit scoping.
@pytest.fixture
async def storage(tmp_path):
store = SqliteStorage(db_path=tmp_path / "test.db")
await store.start()
yield store
await store.stop()
async def test_event_round_trip(storage):
event = make_test_event()
await storage.put(event)
retrieved = await storage.get(event.event_id)
assert retrieved is not NoneThe project uses asyncio_mode = "auto" in pyproject.toml. Most async test
functions do not need an explicit @pytest.mark.asyncio decorator. Add the
decorator only when the auto-detection fails (rare, usually with parametrized
generators).
Never use fixed sleeps directly in tests; use wait_until() or deterministic
hooks (asyncio.Event, mock callbacks).
wait_until() (from tests/helpers/async_utils.py) polls a condition with a
short interval and a bounded timeout. It fails loudly if the condition is not
met within the timeout, rather than silently passing.
from tests.helpers.async_utils import wait_until
# CORRECT: deterministic polling
async def test_delivery_propagates(adapter, event):
await adapter.simulate_inbound(event)
await wait_until(lambda: len(adapter.delivered_payloads) >= 1)
assert adapter.delivered_payloads[0].text == "expected"
# WRONG: fixed sleep
async def test_delivery_propagates_bad(adapter, event):
await adapter.simulate_inbound(event)
await asyncio.sleep(0.3) # nondeterministic, slow, fragile
assert len(adapter.delivered_payloads) >= 1For conditions that can be triggered by side effects, prefer mocking the side
effect or using asyncio.Event over polling:
async def test_pipeline_processes_event(pipeline, event):
processed = asyncio.Event()
original_handle = pipeline.handle_ingress
async def tracking_handle(*args, **kwargs):
result = await original_handle(*args, **kwargs)
processed.set()
return result
pipeline.handle_ingress = tracking_handle
await pipeline.submit(event)
await asyncio.wait_for(processed.wait(), timeout=2.0)Using the wrong mock type is the most common source of RuntimeWarning: coroutine was never awaited and ResourceWarning noise in the test suite.
| Production call | Mock type | Example |
|---|---|---|
await client.close() |
AsyncMock |
client.close = AsyncMock() |
client.add_event_callback(fn) |
MagicMock (never awaited) |
client.add_event_callback = MagicMock() |
await session.start() |
AsyncMock |
session.start = AsyncMock() |
session.config (attribute access) |
Plain attribute or PropertyMock |
session.config = test_config |
If production code awaits the callable, use AsyncMock. For everything
else, use Mock or MagicMock.
When faking scheduler submission helpers (_submit_coro,
run_coroutine_threadsafe, etc.), close passed coroutines before returning.
This prevents "coroutine was never awaited" warnings:
from concurrent.futures import Future
import asyncio
def _submit_done(coro, loop=None):
"""Fake scheduler submit that closes the coroutine immediately."""
if asyncio.iscoroutine(coro):
coro.close()
fut = Future()
fut.set_result(None)
return futAsync fakes that simulate cancellation raise asyncio.CancelledError,
not return a value or raise a different exception. Tests that exercise
cancellation paths catch CancelledError explicitly:
async def _cancel_immediately(*args, **kwargs):
raise asyncio.CancelledError()Tests are classified into seven tiers based on what they honestly prove. Never overclaim the evidence level of a test. If a test uses fake adapters, call it "fake pipeline", not "docker" or "live".
| Tier | Label | What it proves | How to test |
|---|---|---|---|
| 1 | fake_pipeline |
PipelineRunner.handle_ingress() works with direct CanonicalEvent injection |
Import CanonicalEvent, construct it, call runner.handle_ingress() directly |
| 2 | fake_adapter_callback |
adapter.simulate_inbound() produces the same results as direct injection |
Use FakeMatrixAdapter.simulate_inbound(), compare output with direct injection |
| 3 | wrapper_callback |
Real adapter SDK callback (e.g., _on_room_message) bridges to fake target |
Mock the SDK, test the wrapper callback through the pipeline to a fake target adapter |
| 4 | sdk_contract |
Exact pinned optional SDK exposes the constructors, enums and callbacks MEDRE consumes | Dedicated *_sdk marker/job with the adapter extra installed |
| 5 | local_integration |
Real pinned SDK and MEDRE session run together against a deterministic local endpoint | @pytest.mark.local_integration plus the adapter's *_sdk marker |
| 6 | docker_sdk_boundary |
Real SDK code paths work against containerized services (Synapse, meshtasticd) | Docker Compose tests, gated by @pytest.mark.docker |
| 7 | live_network |
Real adapter against real endpoint or hardware | @pytest.mark.live; physical radios additionally use @pytest.mark.hardware |
- Installed-SDK contract tests validate package surfaces without claiming a network/service integration result.
- If a test uses Docker but not real hardware, label it docker_sdk_boundary (tier 6), not "live".
- If a test uses fake adapters but a real pipeline, label it fake_pipeline (tier 1) or fake_adapter_callback (tier 2), not "docker".
- If a test uses a real SDK but routes to a fake outbound target, label it sdk_contract bridge smoke (tier 4), not "live bridge". Reserve local_integration (tier 5) for a deterministic local endpoint and docker_sdk_boundary (tier 6) for tests that actually use a containerized service.
The medre smoke --json report includes an evidence_level field set to
fake_bridge. This is intentional. It does not overclaim.
soak is intentionally not another evidence tier. It describes duration and
repetition at one of the tiers above. A soak test therefore needs to carry the
marker for the layer it exercises as well as pytest.mark.soak. Examples:
local_integration + meshcore_sdk + soakmeans repeated real-SDK/local-node lifecycle and traffic without hardware.hardware + live + meshtastic_sdk + soakmeans repeated MEDRE lifecycle against a physical radio.- A fake-only soak remains synthetic evidence regardless of duration.
The default suite excludes local_integration and soak. Pull requests run the
deterministic local-integration layer separately; local-integration soak jobs are
manual (workflow_dispatch) so they can be extended without making the normal
feedback loop unbounded.
Tests verify that SQLite and in-memory storage backends behave consistently for core operations (put, get, list, delete) and that SQLite persists across restarts while in-memory does not.
async def test_sqlite_persists_across_restarts(tmp_path):
db_path = tmp_path / "test.db"
store = SqliteStorage(db_path=db_path)
await store.start()
await store.put(event)
await store.stop()
store2 = SqliteStorage(db_path=db_path)
await store2.start()
retrieved = await store2.get(event.event_id)
assert retrieved is not None
await store2.stop()medre inspect commands are read-only. They open the database, query data,
print it, and close. They never modify the database. Tests for inspect
subcommands assert that the database content is unchanged after the
command runs.
The schema version (_EXPECTED_SCHEMA_VERSION) stays at 1 during the entire prerelease period. It will not be bumped until storage compatibility becomes release-tracked. Tests assert schema_version == 1 to guard against accidental bumps.
Because the version is frozen, column-shape validation is the primary guard against stale prerelease databases. Tests verify that:
- A fresh database starts with
schema_version == 1in_medre_schema_meta. - A database that reports the correct version but is missing required columns (stale prerelease shape) raises
PreReleaseSchemaMismatchError, not a genericStorageInitializationError(tested intest_prerelease_storage_reset.py). - Column-shape validation reports the affected table and all missing columns in the error (tested in
test_prerelease_storage_reset.py).
Do not add schema version bump tests or migration tests during prerelease. The version stays at 1 and there is no migration code path to test.
Smoke, drill, evidence, trace, and recover commands produce structured JSON
output. Tests assert the normalized JSON shape using shared assertion
helpers from tests/helpers/assertions.py.
from tests.helpers.assertions import assert_report_shape
def test_smoke_report_json_shape(smoke_report):
assert_report_shape(smoke_report)
assert smoke_report["status"] == "passed"
assert smoke_report["evidence_level"] == "fake_bridge"CLI tests use module-level functions, not unittest.TestCase classes. This
allows proper pytest fixture injection and parametrization.
async def test_config_check_valid_config(tmp_path):
config_path = write_test_config(tmp_path)
result = await runner.invoke(["config", "check", "--config", str(config_path)])
assert result.exit_code == 0Patch the canonical module where the object is looked up, not where it is defined. This ensures the patch intercepts the actual import path used by the code under test.
# CORRECT: patch at the lookup site
@patch("medre.adapters.matrix.adapter.HAS_NIO")
async def test_matrix_adapter_without_nio(mock_has_nio):
mock_has_nio.return_value = False
...
# WRONG: patch at the definition site (may not intercept the import)
@patch("medre.adapters.matrix.HAS_NIO")
async def test_matrix_adapter_without_nio_bad(mock_has_nio):
...Adapter package root __init__.py files are lightweight package markers and
should not be used as patch targets. Patch the concrete definition/use site
instead:
# Correct: patch the actual module where the symbol is defined/used
from unittest.mock import patch
with patch("medre.adapters.matrix.compat.HAS_NIO", False):
... # tests that need HAS_NIO=False run here
# Wrong: adapter package roots are docstring-only with no re-exports
# from medre.adapters.matrix import HAS_NIO # This will not workTests must not contain compatibility shims, version detection branches, or
environment-specific workarounds (e.g., if os.getenv("MEDRE_TESTING")).
Test and production code paths are identical.
Treat warnings as bugs where practical. ResourceWarning and
RuntimeWarning about unawaited coroutines indicate real issues (leaked
coroutines, unclosed resources) that will cause problems in production.
For CI hardening, use:
PYTHONPATH=src pytest -W error::ResourceWarning -qThis is not enforced by default (some third-party libraries produce noisy warnings), but failures from this flag should be fixed, not suppressed.
Docker tests exercise real SDK code paths against containerized services (Synapse, meshtasticd). They are opt-in and excluded from default runs.
All Docker tests use the @pytest.mark.docker decorator:
import pytest
@pytest.mark.docker
async def test_synapse_connectivity():
"""Test real Matrix SDK against containerized Synapse."""
...pyproject.toml excludes Docker and live tests from default runs:
[tool.pytest.ini_options]
addopts = "-m 'not live and not docker and not hardware and not local_integration and not soak and not matrix_sdk and not lxmf_sdk and not meshtastic_sdk and not meshcore_sdk'"The excluded markers gate the following tiers (see pyproject.toml markers):
| Marker | What it gates |
|---|---|
live |
Tests connecting to a real service or hardware (skipped by default). |
docker |
Tests requiring Docker services such as Synapse or meshtasticd. |
hardware |
Tests requiring physical hardware (serial/BLE Meshtastic radios, etc.). |
local_integration |
Real pinned SDK plus deterministic local endpoint (no external service). |
soak |
Extended-duration or repeated-cycle transport endurance tests. |
matrix_sdk |
Tests requiring the pinned mindroom-nio Matrix SDK. |
lxmf_sdk |
Tests requiring the pinned LXMF/RNS SDKs. |
meshtastic_sdk |
Tests requiring the pinned mtjk Meshtastic SDK. |
meshcore_sdk |
Tests requiring the pinned meshcore SDK. |
# Prerequisites: Docker daemon running, SDK extras installed
pip install -e ".[matrix,meshtastic,dev]"
# All Docker integration tests
PYTHONPATH=src pytest tests/integration/ -m docker -v
# Matrix (Synapse) only
PYTHONPATH=src pytest tests/integration/test_synapse_connectivity.py -m docker -v
# Meshtastic (meshtasticd) only
PYTHONPATH=src pytest tests/integration/test_meshtasticd_connectivity.py -m docker -vWhen a type checker reports a false positive in a test, fix it by improving the mock or fake to have the correct type, not by suppressing the warning.
Broad # type: ignore or # pyright: ignore comments at module level are not
acceptable in tests. Specific line-level ignores are acceptable only with a
comment explaining why:
adapter._client.send_response = MagicMock() # type: ignore[assignment] # fake does not implement full protocolWhen editing a test file that already has type-ignores, remove any that the edit makes unnecessary.
PYTHONPATH=src pytest -q
# Expected: 14k+ collected, ~125 deselected (live/docker/hardware).
# The full suite takes 600–900 s on typical hardware; use prefix slices
# during development (see slow-suite partition strategy below).This is the primary development command. It runs all unit and fake-pipeline tests. No network, no hardware, no optional SDK dependencies required.
PYTHONPATH=src pytest -m docker -v
# Requires: Docker daemon running, SDK extras installedPYTHONPATH=src pytest -m live -v --tb=short
# Requires: hardware, credentials, environment variablespython -m compileall -q src tests
# Expected: no output (all files compile cleanly)Run all three test tiers plus the compile check before merging. Because the default suite exceeds 600 s, use the partition strategy from Slow-suite partition strategy to cover the full suite in slices:
# 1. Compile check (fast, catches syntax/import errors)
python -m compileall -q src tests
# 2. Collection check (fast, catches fixture/conftest errors)
python -m pytest --collect-only -q
# 3. Default suite in prefix slices (see partition commands above)
PYTHONPATH=src pytest -q tests/test_trace*.py
PYTHONPATH=src pytest -q tests/test_startup*.py
# ... continue through all prefix groups ...
# 4. Docker integration tests (requires Docker)
PYTHONPATH=src pytest -m docker -v
# 5. Live network tests (requires hardware/credentials)
PYTHONPATH=src pytest -m live -v --tb=short# Single test file
PYTHONPATH=src pytest tests/test_pipeline_delivery.py -v
# Files matching a prefix
PYTHONPATH=src pytest tests/test_matrix_session*.py -v
# Keyword match
PYTHONPATH=src pytest -k "test_delivery" -v
# Operator smoke command (Docker-free bridge validation)
PYTHONPATH=src medre smoke
PYTHONPATH=src medre smoke --jsonWhen the full suite hangs, use this per-file timeout loop to isolate which test file is blocking:
# Run from the repository root
set -o pipefail
find tests -type f -name 'test_*.py' | sort | while read -r f; do
echo -n "$(basename "$f"): "
out="$(PYTHONPATH=src timeout 90 python -m pytest -q "$f" 2>&1)"
status=$?
printf '%s\n' "$out" | head -3
if [ "$status" -eq 124 ]; then
echo "TIMEOUT (124)"
fi
doneEach file gets a 90-second timeout. Files that hang will report timeout exit code
124. Note that ordering/pollution across files can mask the real hang, so
always re-run suspect files in isolation to confirm.
Agent responsibility: These rules prevent timeout-tuning loops, output truncation, and retry storms. Violating them wastes time and hides real failures. Follow them exactly.
-
No
timeoutwrappers for routine pytest runs. Runpytestdirectly. The shelltimeoutcommand is reserved for the diagnostic loop above, not for routine validation or failure collection. -
No
tail,head, grep-piping, or output truncation. Capture full pytest output. Piping throughtail,head, orgrephides failure context (tracebacks, fixture teardown errors, import failures). If the output is large, save it to a file or scroll it — never truncate it. -
No broad suite after scoped validation passes. Once a targeted file or keyword run passes, do not escalate to the full suite "just to be sure" unless explicitly requested. The full suite is for pre-merge verification, not iterative debugging.
-
If a test hangs or times out once, stop. Do not rerun it with a longer or different timeout. A hang is the test telling you something is wrong (deadlock, missing cleanup, infinite loop). Longer timeouts do not fix the underlying issue — they just waste time before the same hang.
-
Capture full pytest output once for a failure. When a test fails, the first run's output is the evidence. Read it completely before taking any other action.
-
Static-read the failing test and its source before any rerun. Before rerunning a failing test, use the Read tool to examine both the test file and the source module it exercises. Identify the likely cause from the code, not from repeated execution.
-
Rerun at most once after a concrete suspected fix. Only rerun after making an edit that addresses a specific, identified cause. If the rerun still fails, stop and investigate further — do not loop on runs.
-
Report blockers instead of looping. If the test still hangs or fails after one fix-attempt rerun, stop and report the issue with: the full pytest output, the test file path, the source path, and what was tried. Do not continue running the test with variations.
-
Do not repeatedly run hanging tests with varying timeouts. This is worth stating twice: a test that hangs at 30 s will also hang at 60 s, 90 s, and 300 s. Varying the timeout does not diagnose or fix the hang.
-
Do not use shell pipes to filter pytest output during debugging. Pipes hide information. If you need specific lines, use the Grep or Read tool on saved output, not shell-level truncation.
Agent responsibility: When verifying changes, follow this sequence. Do not skip steps or escalate to broader runs before scoped validation passes.
-
Compile check — catches syntax errors and import-time failures:
python -m compileall -q src tests
Expected: no output. Any output is a blocker.
-
Collection check — catches collection errors (broken fixtures, bad imports, conftest issues) without running any tests:
python -m pytest --collect-only -q
Expected: 14k+ collected, ~125 deselected. Collection takes ~11 s.
-
Targeted files for changed modules — run only the test files that exercise the code you changed. Use file paths, not
-kkeyword filters, for precision:# Example: changed src/medre/core/routing/ PYTHONPATH=src pytest tests/test_route*.py tests/test_routing.py tests/test_routes.py -q
-
Do not run the full suite during scoped validation. See rule 3 in the Test-execution discipline section.
The full suite at 14,000+ tests cannot run within typical agent timeouts (300–600 s). Use directory/prefix slicing to partition the work. The groups below are ordered roughly from slowest to fastest per test; time your slices and stop after one hang.
| Prefix group | Files | Collected | Deselected | Estimated time |
|---|---|---|---|---|
test_meshtastic*.py |
47 | 1,063 | 16 | ~44 s |
test_runtime*.py |
55 | 952 | 12 | ~34 s |
test_matrix*.py |
38 | 832 | 29 | ~29 s |
test_docs*.py |
18 | 729 | 0 | unmeasured |
test_meshcore*.py |
25 | 572 | 11 | ~30 s |
test_lxmf*.py |
25 | 585 | 30 | ~21 s |
test_cli*.py |
25 | 524 | 0 | ~60 s |
test_adapter*.py |
11 | 448 | 0 | unmeasured |
test_replay*.py |
23 | 413 | 0 | ~53 s |
test_capability*.py |
6 | 391 | 0 | unmeasured |
test_evidence*.py |
14 | 373 | 0 | unmeasured |
test_storage*.py |
16 | 357 | 0 | ~40 s |
test_delivery*.py |
11 | 344 | 0 | unmeasured |
test_config*.py |
7 | 308 | 0 | unmeasured |
test_architecture*.py |
11 | 285 | 0 | unmeasured |
test_cross*.py |
6 | 253 | 0 | unmeasured |
test_retry*.py |
14 | 248 | 0 | ~27 s |
test_pipeline*.py |
17 | 247 | 0 | ~30 s |
test_route*.py |
10 | 242 | 0 | unmeasured |
test_soak*, test_longrun*, test_extended* |
8 | 149 | 6 | ~60 s |
conformance/ |
8 | 153 | 0 | unmeasured |
lifecycle/ |
9 | 113 | 0 | unmeasured |
operational/ |
4 | 57 | 0 | unmeasured |
| Other (boundary, canonical, rendering, drill, snapshot, etc.) | ~80 | ~1,900 | varies | unmeasured |
| Total | ~460 | ~14k+ | ~125 | ~600–900 s est. |
The test_soak*.py, test_longrun*.py, and test_extended_longrun*.py files
contain 149 tests that average ~0.4 s/test (60 s total). This is 10× slower
than the runtime or matrix groups (~0.03 s/test). Run this group separately
and only when soak/longrun stability is explicitly in scope.
Run prefix groups one at a time. Use timeout only for diagnostics, not for
routine validation (see rule 1 in the discipline section above).
# Fast groups (under 30 s each)
PYTHONPATH=src pytest -q tests/test_trace*.py
PYTHONPATH=src pytest -q tests/test_startup*.py
PYTHONPATH=src pytest -q tests/test_shutdown*.py
PYTHONPATH=src pytest -q tests/test_retry*.py
# Medium groups (30–60 s each)
PYTHONPATH=src pytest -q tests/test_pipeline*.py
PYTHONPATH=src pytest -q tests/test_storage*.py
PYTHONPATH=src pytest -q tests/test_matrix*.py
PYTHONPATH=src pytest -q tests/test_meshcore*.py
PYTHONPATH=src pytest -q tests/test_lxmf*.py
PYTHONPATH=src pytest -q tests/test_runtime*.py
PYTHONPATH=src pytest -q tests/test_meshtastic*.py
PYTHONPATH=src pytest -q tests/test_replay*.py
PYTHONPATH=src pytest -q tests/test_cli*.py
# Slow group (60 s)
PYTHONPATH=src pytest -q tests/test_soak*.py tests/test_longrun*.py tests/test_extended_longrun*.py
# Subdirectories
PYTHONPATH=src pytest -q tests/conformance/
PYTHONPATH=src pytest -q tests/lifecycle/
PYTHONPATH=src pytest -q tests/operational/
# Remaining "other" files (boundary, architecture, docs, etc.)
PYTHONPATH=src pytest -q tests/test_architecture*.py tests/test_boundary*.py \
tests/test_docs*.py tests/test_capability*.py tests/test_evidence*.py \
tests/test_delivery*.py tests/test_config*.py tests/test_canonical*.py \
tests/test_rendering*.py tests/test_drill*.py tests/test_snapshot*.py \
tests/test_example*.py tests/test_scope*.py tests/test_failure*.pyThe per-file timeout loop in Identifying hanging tests
above isolates individual hanging files. For broader diagnostics, use
collection-based slicing and pytest-timeout per-test timeouts:
# Step 1: Confirm collection works (no test runs, ~11 s)
python -m pytest --collect-only -q
# Step 2: Run prefix groups with a per-test timeout to surface slow individual
# tests without hanging the entire suite. Requires pytest-timeout (in dev deps).
PYTHONPATH=src pytest -q --timeout=30 tests/test_soak*.py tests/test_longrun*.py
# Step 3: If a prefix group hangs, bisect it. Run half the files:
PYTHONPATH=src pytest -q tests/test_longrun_soak.py tests/test_longrun_stability_v3.py
# If that passes, the hang is in the other half.When the full suite hangs at low progress (e.g., 6 % at 600 s), the cause is typically one of:
- A single test that deadlocks or loops — isolated with the per-file
timeout loop or
--timeout=30per-test cap. - Cumulative resource leaks (unclosed sqlite connections, dangling event loops) that slow later tests — detected by running groups in reverse alphabetical order and comparing timings.
- Fixture ordering pollution — a test that leaves global state (env vars, sys.modules patches, working directory) that poisons a later test. Run the suspect file in isolation to confirm.
Do not mask these with longer timeouts, filterwarnings suppressions, or
-x (fail-fast). Diagnose the root cause.
| Symptom | Likely cause | Action |
|---|---|---|
| Docker tests skip with "Docker not available" | Docker daemon not running | docker info |
| Docker tests skip with "mtjk not installed" | Meshtastic SDK not installed | pip install -e ".[meshtastic]" |
| Docker tests skip with "mindroom-nio not installed" | Matrix SDK not installed | pip install -e ".[matrix]" |
| Live tests skip | Missing environment variables | Set required MATRIX_* or MESHTASTIC_* env vars |
| Compile check produces output | Syntax error or import issue | Fix the reported file |
ResourceWarning in test output |
Unclosed resource or leaked coroutine | Fix the mock or add cleanup (see Async Mocking Rules) |
These files have been split by behavioral domain following the procedure above.
| Original file | Result | Domain files |
|---|---|---|
Former tests/test_adapter_callback_bridge.py |
Split | 6 domain files |
Former tests/test_longrun_callback_bridge.py |
Split | 4 domain files |
Former tests/test_operator_workflows.py |
Split | 7 domain files |
Former tests/test_pipeline.py |
Split | 5 domain files (delivery, failure taxonomy, fanout, native refs, capacity) |
Former tests/test_replay.py |
Split | 5 domain files (engine, policy, accounting, capacity, traceability) |
Former tests/test_cli.py |
Split | 9 domain files: test_cli_command_help_hints, test_cli_config_workflows, test_cli_diagnostics_workflows, test_cli_install_metadata, test_cli_replay_surface, test_cli_route_workflows, test_cli_run_workflows, test_cli_scenario_crosscheck, test_cli_smoke_run_session. Helper: helpers/cli.py. |
| Former CLI walkthrough monolith | Split | 4 domain files: test_cli_config_and_smoke, test_cli_inspect_flow, test_cli_replay_flow, test_cli_error_paths. Helper: helpers/walkthrough.py. |
Former tests/test_docker_bridge_artifacts.py |
Split | 4 domain files: test_docker_artifact_core, test_docker_artifact_plan, test_docker_artifact_metadata, test_docker_artifact_honesty. Helper: helpers/docker_artifacts.py. |
tests/test_matrix_session.py |
Split | 3 domain files: test_matrix_session_config (encryption config), test_matrix_session_e2ee (Megolm, encrypted rooms, E2EE diagnostics), test_matrix_session_recovery (sync failure, reconnect, crypto store continuity, sync state resilience). Original retained at 460 lines (lifecycle, diagnostics, start behavior). |
tests/test_storage.py |
Split | 7 domain files: test_storage_durability, test_storage_integrity, test_storage_invariants, test_storage_native_refs, test_storage_path_cli, test_storage_path_validation, test_storage_receipts. Original retained at 231 lines. |
tests/test_replay_routing.py |
Split | 3 domain files: test_replay_routing_controls, test_replay_routing_durability, test_replay_routing_isolation. Original retained at 422 lines. |
tests/test_runtime_builder.py |
Split | 3 domain files: test_runtime_builder_ordering (build ordering, adapter ID propagation), test_runtime_builder_paths (Matrix store path derivation, ensure-dirs), test_runtime_builder_routes (degraded route validation). Original retained at 520 lines (construction, config, fakes). |
tests/test_meshtastic_adapter.py |
Split | 1 domain file: test_meshtastic_adapter_delivery (send semantics, session boundary, session unit). Original retained at 755 lines (connection modes, queue ownership, lifecycle). |
tests/test_meshtastic_fake_bridge.py |
Split | 2 domain files: test_meshtastic_fake_bridge_errors, test_meshtastic_fake_bridge_session. Original retained at 938 lines. |
tests/test_storage_outbox.py |
Deleted | 5 domain files: test_storage_outbox_crud (create, get, idempotent create, list, count, persistence), test_storage_outbox_claim (claim due, release claim, claim clears next_attempt_at), test_storage_outbox_status (status transitions, transition guards, queued lease semantics), test_storage_outbox_atomic_create (atomic create, no-steal guarantees), test_storage_outbox_concurrency (write lock serialisation, transaction rollback, stale queued reclaim, is_claimable property). Original deleted. |
tests/test_fake_runtime_smoke.py |
Split | 2 domain files: test_fake_runtime_soak (diagnostics snapshots, replay delivery, happy path), test_fake_runtime_startup_snapshot (startup/shutdown integration, snapshot integration). Original retained at 931 lines. |
tests/test_operator_recovery.py |
Split | 3 domain files: test_config_repair (malformed config, storage path, config repair workflows), test_startup_recovery (startup failure, degraded runtime, adapter disable/enable), test_deterministic_messaging (no-traceback assertions, deterministic boot/supervision shape). Helper: helpers/operator_recovery.py. Original retained at 295 lines (route validation recovery, replay after restart). |
Former test_cli.py has been split into domain files (all under 1,500 lines). The
monolith has been deleted. test_cli is listed in DELETED_MONOLITHS in
test_test_suite_structure.py.
- Adapter authoring guide -- writing a new transport adapter and its fake
- Source audits -- audit evidence for transport SDK assumptions
- Operator workflows -- operator commands for bridge testing