The validated system has two RTX 3090 24 GB cards and 128 GB of system memory,
plus NVMe-backed swap. Configure at least 32 GiB of swap before loading the
checkpoint; 48–64 GiB is recommended. The launch profile assumes two visible
CUDA GPUs and reserves approximately 4.13 GiB of KV cache on each GPU for one
262,144-token sequence. See memory.md before changing the cache
or concurrency settings.
Required software:
- Linux x86_64 with a working NVIDIA Container Toolkit;
- Docker with Compose v2 (Compose is optional if using
make serve); - Python 3.10+ for assembly scripts;
hffromhuggingface_hubwithhf_xet;- enough local storage for both upstream sources, the 120 GiB release tree, and a temporary 4 GiB MTP build.
Run make preflight before the first image build. It is read-only. The runtime
container receives SYS_PTRACE for cross-process PLE CUDA IPC. On hosts where
that is insufficient, kernel.yama.ptrace_scope=0 may be required temporarily;
do not make that security relaxation persistent without reviewing it.
The launcher passes --max-parallel-loading-workers 1, but the pinned vLLM
runtime warns that it ignores this option. Do not rely on it to limit loading
memory. The target, PLE worker, pinned expert storage, and conversion buffers
overlap during startup, so swap is still needed.
repro.lock.json is authoritative. It pins:
- Intel's AutoRound checkpoint by full Hub commit;
- RadixArk's PLE/MTP source by full Hub commit;
- the vLLM day-0 container by OCI digest;
- Humming kernels by package version;
- target/MTP tensor counts and payload byte counts;
- the validated hardware and serving profile.
Never replace these with branch names such as main in a tagged release.
Install the current Hugging Face CLI, authenticate, and download the inputs:
python3 -m pip install --upgrade 'huggingface_hub[hf_xet]'
hf auth login
./scripts/download_sources.sh /models/qwen38-sourcesThe recommended builder is the same digest-pinned container used for serving, which already provides the matching PyTorch, safetensors, and compressed-tensors stack:
make build-image
./scripts/assemble_with_docker.sh /models/qwen38-sources uploadThis creates /models/qwen38-sources/upload and keeps every large path under a
single container mount, so hard links remain available. Advanced users can run
assemble_hf_repo.sh directly in a compatible host Python environment.
The target builder does not requantize or repack Intel tensors. It omits the Intel BF16 PLE-only shard and bundled BF16 MTP file, inserts the published FP8 PLE shards, and writes a new auditable index. The separate MTP builder quantizes only routed draft experts to symmetric INT4 group-32; target verification still determines emitted tokens.
Set KEEP_WORK=1 if you want to retain the large symlinked MTP intermediate for
debugging. Otherwise the script deletes only the uniquely named temporary
directory it created beneath the supplied work directory.
./scripts/upload_hf.sh albucino/Qwen3.8-Flash-Next-W4A16-FP8PLE \
/models/qwen38-sources/uploadThe script validates first, uses hf upload/Xet so interrupted uploads can be
resumed, and prints the resulting Hub commit. Before tagging GitHub:
Keep the source directories in place until the upload finishes because the assembled target uses hard links to unchanged upstream shards.
- put the Hugging Face repo and full commit in
repro.lock.json; - put the GitHub repository and tag/commit in the Hugging Face model card;
- rerun
make validate; - commit, tag
v1.0.0, and create a GitHub release; - verify a fresh machine can download by the pinned Hub revision and start the endpoint without relying on an untracked file.
Do not upload model weights to GitHub or GitHub Releases. The Hub is designed for resumable, content-addressed model uploads and keeps the model card next to the tensors.
export MODEL_DIR=/models/qwen38-upload
make build-image
make preflight
make serveEach release from v0.3.0 on is also published to GHCR by
.github/workflows/publish-image.yml, built from the release tag with the same
docker/Dockerfile. The model weights are not in the image. Pull by digest,
not by tag; the digest for each release is in its release notes and the
workflow run summary:
docker pull ghcr.io/dominikbucko/qwen38-flash-next-2x3090:v0.3.0
IMAGE=ghcr.io/dominikbucko/qwen38-flash-next-2x3090@sha256:<digest from the release notes> make serveThe image carries OCI labels for its source commit
(org.opencontainers.image.revision) and base image
(org.opencontainers.image.base.name), and a signed build-provenance
attestation you can check with
gh attestation verify oci://ghcr.io/dominikbucko/qwen38-flash-next-2x3090:v0.3.0 --owner DominikBucko.
It is rebuilt in CI, so it is not byte-identical to the local image the
September 25 numbers were measured on; the build inputs are. That run's exact
settings, host and driver are in
benchmarks/2026-09-25/environment.json.
Equivalent Compose launch:
export MODEL_DIR=/models/qwen38-upload
docker compose -f docker/compose.yaml up --buildThe single-request default is hot84 with the full 262,144-token limit. Existing
.env files override it: change VLLM_WNA16_STATIC_HOT_CACHE_SIZE to 84 when
upgrading. Hot88 remains an optional tighter profile; test long prefill before
using it. Neither setting removes experts or changes their precision.
The approximate QSA selector is default. To opt into exact selection:
export VLLM_QSA_EXACT_TOPK=1
make serveThe exact setting is a precision experiment, not the recommended performance profile. A single AgentBench A/B did not improve hidden-functional passes or mean score.
GitHub CI checks Python syntax, overlay checksums, lock/Docker consistency,
accidental model blobs, symlinks, cache files, and common token formats. It
cannot validate CUDA kernels, 256K allocation, or throughput. Those checks need
the target hardware and the procedure in benchmarks.md.
The September 16 fresh Docker check passed both the two-client 128K profile and a full-256K request as the first inference on hot84 defaults. It also verified all installed overlay hashes and bidirectional P2P copies inside the image. See the test record for exact versions and limits. This does not replace testing on a different host.