Celerity's first guiding principle is correctness first — "a fast collection that returns wrong answers is worthless." This document describes how that principle is enforced: the layers of tests, the property-based and fuzz harnesses that hunt for the bugs example-based tests miss, how code coverage is measured and gated, and how to run each layer locally.
| Layer | Project / file | What it proves | Run locally |
|---|---|---|---|
| Behavioural unit tests | Celerity.Tests |
Each public method does the right thing on hand-picked inputs, including collisions, resizes, and the out-of-band default/zero/null key. | dotnet test |
| Edge-case coverage | alongside each type's tests (*Tests.cs, *EnumerationTests.cs, *CollisionTests.cs) |
The corners example tests skip: non-generic IEnumerable/IEnumerator paths, Reset(), indexer misses, Clear() on empty, wrap-around backward-shift. |
dotnet test |
| Property-based tests | Celerity.Tests/Properties/ |
Across thousands of randomized operation sequences, every collection stays observably equal to its BCL oracle. | dotnet test |
| Differential fuzzer | Celerity.Fuzz |
A long random walk finds no divergence from the BCL; failures replay deterministically from a seed. | dotnet run -c Release |
| Native AOT smoke test | Celerity.AotSmokeTest |
Every collection/hasher works in a trimmed, AOT-compiled native binary. | see aot.md |
All of these run in CI. Coverage is measured on all six shipping assemblies and gated at 100% line and branch; the rendered report is published to the coverage dashboard.
Example-based unit tests are necessary but not sufficient for a data-structure library. They prove the cases the author thought of. The bugs that actually ship in open-addressed hash tables live in the cases nobody enumerated: a particular interleaving of inserts and deletes that leaves a tombstone in the wrong slot, a resize triggered mid-probe-chain, a backward-shift that wraps the table boundary and orphans a colliding key.
Celerity attacks those with two adversarial layers that don't rely on the author's imagination:
- Property-based testing generates random operation sequences and checks an invariant — here, equivalence to a known-correct BCL collection — rather than a fixed expected output.
- Differential fuzzing runs the same idea as an unbounded soak: keep generating sequences until something diverges, and when it does, hand back a seed that reproduces it.
Both compare against a BCL oracle (Dictionary<,>, HashSet<>, or a Dictionary<TKey, List<TValue>> model for the multi-map). The oracle is the specification: Celerity claims drop-in parity, so any observable difference is a bug in Celerity.
The bulk of the suite lives in Celerity.Tests, mirroring the library's folder layout. Test names follow Method_ShouldExpectedBehavior_WhenCondition. Notable categories:
- Collision tests (
*CollisionTests.cs) — force every key down one probe chain with a constant hasher, then verify lookups, removals, and backward-shift deletion keep every entry findable. - Enumeration tests (
*EnumerationTests.cs) — the struct enumerators,Keys/Valuesviews, mid-enumeration mutation detection, and the non-generic interface surface (IEnumerable.GetEnumerator(),object IEnumerator.Current,IEnumerator.Reset()). - Load-factor / constructor validation — boundary resizes and argument checking.
- Edge cases live next to the type they exercise rather than in a catch-all file: indexer misses on the out-of-band key and
Clear()on an empty collection sit in*Tests.cs; the wrap-around cluster that exercises thebypassesGapbranch of backward-shift deletion sits in*CollisionTests.cs.
Run them with:
dotnet testCelerity.Tests/Properties/CollectionModelPropertyTests.cs uses CsCheck to make parity the explicit contract. Each test:
- Generates a randomized list of mutating operations (
Set,Remove,TryAdd,Clearfor dictionaries;Add/RemoveValue/RemoveAllfor the multi-map; etc.). - Applies the identical sequence to a Celerity collection and a BCL oracle.
- Asserts the two are observably equal —
Count, per-key lookups across the whole key domain, and full enumeration.
The key domains are deliberately tiny (and include 0 and negatives) so collisions, resizes, the special zero/default/null-key slot, and backward-shift deletion all fire densely. Each test runs 2 000 sequences per invocation.
When a property fails, CsCheck shrinks the failing sequence to a minimal reproduction and prints a seed. Replay it by setting the seed:
# PowerShell
$env:CsCheck_Seed = '0000LASTpRINTED'; dotnet test --filter CollectionModelPropertyTests# bash
CsCheck_Seed='0000LASTpRINTED' dotnet test --filter CollectionModelPropertyTestsCoverage targets: CelerityDictionary, IntDictionary, LongDictionary, CeleritySet, IntSet, LongSet, CelerityMultiMap, and FrozenCelerityDictionary.
The property tests are bounded — a fixed number of sequences per CI run. The fuzzer is the unbounded soak counterpart: it keeps generating cases until you stop it or it finds a divergence. It lives in src/Celerity.Fuzz and shares the same differential idea (drive Celerity and a BCL oracle in lock-step, fail on the first observable difference).
Every case is a pure function of a single 32-bit seed, so a failure is perfectly reproducible. Run it locally:
cd src/Celerity.Fuzz
# 100k random cases across all collections
dotnet run -c Release -- --iterations 100000
# soak for 60 seconds
dotnet run -c Release -- --time 60
# focus one collection
dotnet run -c Release -- --target CelerityMultiMap --iterations 200000
# list the targets
dotnet run -c Release -- --listOn a failure it prints the target, the caseSeed, and a ready-to-paste replay command:
================ FUZZ FAILURE ================
target : CelerityDictionary
caseSeed : 1734023
replay : dotnet run -c Release -- --seed 1734023 --iterations 1
detail : DivergenceException: value[3] 70 != 71
==============================================
Reproduce it with exactly that command. (Note: --target changes the RNG stream, so when replaying a reported caseSeed, omit --target — the seed already determines which collection ran.)
In CI the fuzzer runs as a nightly job (.github/workflows/fuzz.yml) with a wall-clock budget, plus on-demand via workflow_dispatch (where you can pass a seed, time, or target). It is intentionally not a per-PR gate — a soak job belongs on a schedule, while the bounded property tests cover the per-PR signal.
Add an entry to Differential.All in src/Celerity.Fuzz/Differential.cs and write a method that drives your collection against a BCL oracle, calling Check(condition, message) on every observable. The driver discovers it automatically (including in --list).
Coverage is collected with coverlet and scoped to all six shipping assemblies — Celerity, Celerity.Hashing, Celerity.Primitives, Celerity.Ring, Celerity.Sentinel, Celerity.Cardinality — via src/coverage.runsettings. The test, benchmark, fuzz, and AOT-smoke assemblies are tooling, not the subject under measurement.
Four test projects contribute: Celerity.Tests for the three core packages, plus Celerity.Ring.Tests / Celerity.Sentinel.Tests / Celerity.Cardinality.Tests for the showcase tier. Their Cobertura reports are merged on (source file, line number), so a line covered by any run counts as covered — which matters because the showcase projects also exercise Celerity.Collections transitively.
The suite covers 100% of lines and 100% of branches across all six. A small number of guards are excluded at the source with [ExcludeFromCodeCoverage(Justification = "…")], and only where no test could ever reach them:
| Guard | Why no test can reach it |
|---|---|
Deque.ClampToArrayMaxLength |
Needs a backing array above 2³⁰ elements. Pinned by a real [MemoryIntensiveFact(3100)] test in DequeGrowthTests, which allocates ~3 GiB and skips on memory-capped runners — excluded so the gate does not depend on whether the runner had the headroom. |
IndexedPriorityQueue.ClampGrowth |
Needs 2³⁰ live entries with distinct elements, pushing the backing array past the 2 GiB single-object limit. EnsureCapacity calls Resize directly, so capacity cannot be pre-inflated into it. |
FrozenCelerityDictionary.ThrowIfKeyCountExceedsCeiling, FrozenCeleritySet.ThrowIfElementCountExceedsCeiling |
The count is taken after materializing the source into a List<string>, so reaching 2³⁰ needs an 8.6 GB string[] — past the 2 GiB array limit. A source that merely reports a huge ICollection.Count cannot reach it; that count is only a capacity hint. |
CuckooFilter.AtLeastOne, XorFilter.AtLeastOne |
Dead by construction: the constructors' own argument validation already forces both sizing expressions above the floor. |
XorFilter.BuildOrThrow, XorFilter.TryBuild |
The peel retry schedule is independent of the element set, so no hasher can stall all MaxConstructionAttempts seeds. Individual attempts do stall and retry — that path lives in TryPeel, which stays measured. |
Hash64Source.CreateNative |
Its null arm is unobservable. Native is read only by Hash64, and every caller guards that on IsNative64 being true, so the class is never initialized for a 32-bit-only THasher — the arm is evaluated only if the runtime runs the beforefieldinit initializer eagerly, which is its option and not a contract. |
That table is the complete set; grep -rn "ExcludeFromCodeCoverage" src/Celerity*/ should return nothing beyond it.
The rule for new code: exclusions are for genuinely unreachable code and must carry a Justification that says why. Anything a test can reach gets a test.
Adding a shipping package? Add its assembly to the
<Include>list insrc/coverage.runsettings, and its test project to.github/workflows/coverage.yml. Coverlet's assembly filter is exact-match, not a prefix —[Celerity]*compiles to^Celerity$and matches only theCelerity.Collectionsassembly. That is how the 2.0.0 package split left five of six packages silently outside the gate until #314. An unlisted package is unmeasured, and the gate stays green no matter what its coverage is.
Collect and render a report locally:
# 1. collect Cobertura coverage for the three core packages.
# Clear stale results first: the four reports are merged by source-file path, and
# SourceLink resolves those paths from the build's git state — so mixing reports
# from different commits makes the same file appear twice under two spellings and
# the merged totals come out roughly halved.
cd src
rm -rf ./TestResults/coverage ./TestResults/showcase
dotnet test Celerity.Tests/Celerity.Tests.csproj \
--collect:"XPlat Code Coverage" \
--settings coverage.runsettings \
--results-directory ./TestResults/coverage
# 2. and for the three showcase packages
for project in Ring Sentinel Cardinality; do
dotnet test "Celerity.${project}.Tests/Celerity.${project}.Tests.csproj" \
--collect:"XPlat Code Coverage" \
--settings coverage.runsettings \
--results-directory "./TestResults/showcase/${project}"
done
# 3. render the HTML report + badge (pure Python, no extra tooling).
# --input is repeatable; the reports are merged.
python3 ../scripts/coverage_report.py \
--input "./TestResults/coverage/**/coverage.cobertura.xml" \
--input "./TestResults/showcase/*/**/coverage.cobertura.xml" \
--outdir ../coveragereport --min-line 100 --min-branch 100
# 4. open coveragereport/index.htmlThe report is rendered by scripts/coverage_report.py — a small generator that reads the Cobertura XML coverlet produces and emits an index.html styled like the rest of the Celerity site, a badge.svg, and a summary.md. It exists so the report carries the project's own look and no third-party "sponsors only" upsell; there is no dependency on ReportGenerator.
The coverage workflow (.github/workflows/coverage.yml) runs on every PR and on main:
- Collects coverage, renders the report + badge with
scripts/coverage_report.py, and uploads it as a build artifact. - Fails the build if line coverage drops below
MIN_LINE_COVERAGE(100%) or branch coverage belowMIN_BRANCH_COVERAGE(100%). The floor is deliberately a hair-trigger: new code arrives with its tests, or the gate goes red. If you hit a genuinely unreachable branch, exclude it at the source with a justification rather than lowering the floor. - Posts a coverage summary comment on the PR.
- On
main, publishes the HTML report togh-pagesunder/coverageand refreshes the README badge.
| Workflow | Trigger | What it does |
|---|---|---|
ci.yml |
push / PR | dotnet build + dotnet test on Linux, Windows, macOS; Native AOT publish + smoke run. |
coverage.yml |
push / PR | Collect + gate coverage, comment on PRs, publish report on main. |
fuzz.yml |
nightly / manual | Differential fuzz soak with a time budget. |
benchmarks.yml |
push / PR | Same-runner A/B benchmark comparison vs main. |
When fixing a bug, add a test that fails on main and passes on your branch (see CONTRIBUTING.md). For a new collection, the expectation is parity coverage at every layer: behavioural tests, a property-based parity test against the closest BCL oracle, and a fuzz target. Reusing the existing helpers in each file is the fastest way to get there.