SIMD-accelerated Random Linear Network Coding over GF(2⁸)
Author: arryboom · MSRV 1.89 · License Apache-2.0 · v1.2.0
Crate package: rlnc-simdx · Rust import: rlnc_simdx
no_std-friendly · zero mandatory dependencies · safe public API · automatic CPU dispatch
(GFNI / AVX-512 / AVX2 / SSSE3 / NEON / WASM SIMD128)
~55–62 GB/s AXPY on AMD Ryzen 7 5800X3D (dual DDR4, AVX2 path) — about 25–29× faster than scalar.
# Cargo.toml
[dependencies]
rlnc-simdx = "1.2"
# From this repository's workspace root instead:
# rlnc-simdx = { path = "rlnc-simdx" }The default features are std and alloc. Use default-features = false for the
zero-allocation core, or enable only alloc for heap-backed APIs without std.
use rlnc_simdx::{Decoder, Encoder, RlncError};
fn main() -> Result<(), RlncError> {
let k = 4; // generation size (number of source symbols)
let n = 128; // bytes per symbol
let source: Vec<Vec<u8>> = (0..k).map(|i| vec![i as u8; n]).collect();
let source_refs: Vec<&[u8]> = source.iter().map(Vec::as_slice).collect();
let encoder = Encoder::new(k, n)?;
let mut decoder = Decoder::new(k, n)?;
// Systematic packets make this compact example deterministic. Applications
// can use Encoder::encode_random with SimpleRng or their own coefficients.
for index in 0..k {
let packet = encoder.encode_systematic(&source_refs, index)?;
decoder.receive(packet)?;
}
let decoded = decoder.decode()?.expect("decoder has full rank");
assert_eq!(decoded, source);
Ok(())
}use rlnc_simdx::kernel;
let x = [1u8, 2, 3, 4];
let mut y = [0u8; 4];
kernel::axpy(0x03, &x, &mut y); // y[i] ^= c * x[i]; buffers must not overlap
kernel::scale(0x03, &x, &mut y); // y[i] = c * x[i]; buffers must not overlap
kernel::scale_inplace(0x03, &mut y); // y[i] = c * y[i]Check which SIMD path is active:
println!("kernel: {}", rlnc_simdx::active_kernel());
// e.g. "avx2+ssse3 (tier5)" or "gfni+avx512 (tier1)"Heap API (alloc) |
Encoder · Decoder · Recoder · CodedPacket · GfMatrix · AlignedBuffer |
| Zero-allocation API | Gf8 · RlncError · safe kernel operations · active_kernel() |
| Field | GF(2⁸), AES polynomial 0x11B (same field as Intel GFNI hardware) |
| SIMD | Picked automatically at runtime with std; compile-time selection without std |
| Safety | Public kernels check length and non-overlap in release builds and panic on misuse |
| Dependencies | None |
| Tests | 91 unit tests + 3 doctests |
SVE is experimental and intentionally not part of production dispatch. AArch64 uses the production NEON kernel. The SVE module remains crate-private until a correct vector-length-agnostic implementation is hardware-tested.
Reference machine:
| CPU | AMD Ryzen 7 5800X3D |
| RAM | Dual-channel DDR4 |
| Kernel | avx2+ssse3 (tier5) (no GFNI / no AVX-512 on this SKU) |
| Tool | bench_standalone, SI GB/s, aligned buffers |
| Size | Scalar | SIMD dispatch | Speedup |
|---|---|---|---|
| 1 KiB | ~2.2 GB/s | ~41 GB/s | ~19× |
| 16 KiB | ~2.2 GB/s | ~62 GB/s | ~29× |
| 64 KiB | ~2.2 GB/s | ~60 GB/s | ~28× |
| 1 MiB | ~2.2 GB/s | ~55 GB/s | ~25× |
| Size | SIMD dispatch (typical) |
|---|---|
| 16 KiB | ~70 GB/s (~28× scalar) |
| 1 MiB | ~55 GB/s (~22× scalar) |
Mid sizes can vary run-to-run due to clocks and cache state. Prefer several runs for published measurements.
| Symbol size | ≈ Time / packet | Payload throughput |
|---|---|---|
| 1 KiB | ~500 ns | ~2.0 GB/s |
| 16 KiB | ~5 µs | ~3.2 GB/s |
| 64 KiB | ~24 µs | ~2.7 GB/s |
CPUs with GFNI (Ice Lake+, Zen 4+) usually report gfni+avx2 or
gfni+avx512 and higher throughput.
The focused multithreaded benchmark compares scalar and runtime-dispatched SIMD encode/decode throughput using the same RLNC pipeline and Rayon parallelism:
cargo run --release -p rlnc-simdx-mt-bench
# Short run, considering at most two worker threads
cargo run --release -p rlnc-simdx-mt-bench -- --quick --max-threads 2It covers k = 8, 16, 32 and symbol sizes 64 B, 1 KiB, 4 KiB, 16 KiB, 64 KiB. Scalar and SIMD worker counts are autotuned independently for each
workload and operation. A portable ASCII metadata panel reports the active SIMD
kernel, host platform, logical CPUs, mode, tuning range, and measurement scope.
Encode and Decode use separate tables; rows stream as workloads complete, and
each table ends with peak-throughput and best-speedup summaries.
Thread-pool construction, fixture creation, and reusable worker-state allocation
are outside timed samples. See the
rlnc-simdx-mt-bench guide
for the measurement model.
The portable standalone kernel benchmark remains available:
cargo build --release -p rlnc-simdx-bench --bin bench_standalone
# Windows
target\release\bench_standalone.exe --quick
# Linux / macOS
./target/release/bench_standalone --quickYou can copy either release binary to another machine with the same OS and architecture without installing Rust.
Criterion HTML reports are written under target/criterion/:
cargo bench -p rlnc-simdx-bench --bench axpy
cargo bench -p rlnc-simdx-bench --bench decodeThe workspace benchmark package enables the explicitly unstable
bench-internals feature to compare scalar internals with the supported safe
kernel API. Applications should not enable that feature.
Your application (safe Rust)
→ Encoder / Decoder / kernel::axpy | scale | scale_inplace
→ crate-private dispatch and SIMD implementations
- Do: use equal-length, non-overlapping slices for
axpyandscale. - Do: use
scale_inplacewhen multiplying a buffer by a scalar in place. - Don't: treat this crate as encryption or constant-time cryptography.
- Don't: use
SimpleRngas a CSPRNG; it is a small LFSR for coding coefficients.
cargo test --workspace
cargo test -p rlnc-simdx --all-features
cargo check -p rlnc-simdx --no-default-features # zero-allocation core
cargo check -p rlnc-simdx --no-default-features --features alloc # no_std + alloc| Feature | Default | Meaning |
|---|---|---|
alloc |
✓ | Heap APIs: aligned buffers, matrix, encoder, decoder, and recoder |
std |
✓ | Runtime CPU dispatch; implies alloc |
bench-internals |
Unstable scalar internals for this workspace's benchmarks only |
With no features, field, error, kernel, and active_kernel() remain
available and do not require a global allocator.
Maintainers create and push a signed semantic-version tag after updating the workspace version, lockfile, changelog, and synchronized READMEs:
git tag -s v1.2.0
git push origin v1.2.0The tag push runs the release workflow. A manual workflow dispatch may rerun an existing tag idempotently; it does not create the tag. The workflow requires the tag version to match the inherited Cargo workspace version and publishes:
rlnc-simdx-1.2.0.craterlnc-simdx-bench-1.2.0-x86_64-unknown-linux-gnu.tar.gzrlnc-simdx-bench-1.2.0-x86_64-pc-windows-msvc.ziprlnc-simdx-bench-1.2.0-x86_64-apple-darwin.tar.gzSHA256SUMS
Each native archive contains bench_standalone and rlnc-simdx-mt-bench (with
.exe on Windows), plus README.md and LICENSE. Download all assets into one
directory and verify them against the sorted manifest:
# Linux (or macOS with GNU coreutils)
sha256sum -c SHA256SUMS
# macOS fallback
shasum -a 256 -c SHA256SUMS# PowerShell: compare these hashes with SHA256SUMS
Get-ChildItem rlnc-simdx-* | Get-FileHash -Algorithm SHA256Release notes also contain the commit SHA and a per-asset SHA-256 table. See the changelog for release details.
rlnc-simdx/ # publishable library package (import: rlnc_simdx)
rlnc-simdx-bench/ # Criterion suites and portable bench_standalone binary
rlnc-simdx-mt-bench/ # Rayon-autotuned scalar vs SIMD encode/decode benchmark
plans/ # architecture, review, and historical design notes
Deeper design and release notes: architecture · changelog
arryboom — https://github.com/arryboom
Licensed under the Apache License, Version 2.0 (SPDX: Apache-2.0).