Skip to content

Latest commit

 

History

History
167 lines (128 loc) · 6.64 KB

File metadata and controls

167 lines (128 loc) · 6.64 KB

Reproduction and release guide

1. Hardware and host prerequisites

The validated system has two RTX 3090 24 GB cards and 128 GB of system memory, plus NVMe-backed swap. Configure at least 32 GiB of swap before loading the checkpoint; 48–64 GiB is recommended. The launch profile assumes two visible CUDA GPUs and reserves approximately 4.13 GiB of KV cache on each GPU for one 262,144-token sequence. See memory.md before changing the cache or concurrency settings.

Required software:

  • Linux x86_64 with a working NVIDIA Container Toolkit;
  • Docker with Compose v2 (Compose is optional if using make serve);
  • Python 3.10+ for assembly scripts;
  • hf from huggingface_hub with hf_xet;
  • enough local storage for both upstream sources, the 120 GiB release tree, and a temporary 4 GiB MTP build.

Run make preflight before the first image build. It is read-only. The runtime container receives SYS_PTRACE for cross-process PLE CUDA IPC. On hosts where that is insufficient, kernel.yama.ptrace_scope=0 may be required temporarily; do not make that security relaxation persistent without reviewing it.

The launcher passes --max-parallel-loading-workers 1, but the pinned vLLM runtime warns that it ignores this option. Do not rely on it to limit loading memory. The target, PLE worker, pinned expert storage, and conversion buffers overlap during startup, so swap is still needed.

2. Immutable inputs

repro.lock.json is authoritative. It pins:

  • Intel's AutoRound checkpoint by full Hub commit;
  • RadixArk's PLE/MTP source by full Hub commit;
  • the vLLM day-0 container by OCI digest;
  • Humming kernels by package version;
  • target/MTP tensor counts and payload byte counts;
  • the validated hardware and serving profile.

Never replace these with branch names such as main in a tagged release.

3. Build the Hugging Face upload tree

Install the current Hugging Face CLI, authenticate, and download the inputs:

python3 -m pip install --upgrade 'huggingface_hub[hf_xet]'
hf auth login
./scripts/download_sources.sh /models/qwen38-sources

The recommended builder is the same digest-pinned container used for serving, which already provides the matching PyTorch, safetensors, and compressed-tensors stack:

make build-image
./scripts/assemble_with_docker.sh /models/qwen38-sources upload

This creates /models/qwen38-sources/upload and keeps every large path under a single container mount, so hard links remain available. Advanced users can run assemble_hf_repo.sh directly in a compatible host Python environment.

The target builder does not requantize or repack Intel tensors. It omits the Intel BF16 PLE-only shard and bundled BF16 MTP file, inserts the published FP8 PLE shards, and writes a new auditable index. The separate MTP builder quantizes only routed draft experts to symmetric INT4 group-32; target verification still determines emitted tokens.

Set KEEP_WORK=1 if you want to retain the large symlinked MTP intermediate for debugging. Otherwise the script deletes only the uniquely named temporary directory it created beneath the supplied work directory.

4. Upload and cross-pin releases

./scripts/upload_hf.sh albucino/Qwen3.8-Flash-Next-W4A16-FP8PLE \
  /models/qwen38-sources/upload

The script validates first, uses hf upload/Xet so interrupted uploads can be resumed, and prints the resulting Hub commit. Before tagging GitHub:

Keep the source directories in place until the upload finishes because the assembled target uses hard links to unchanged upstream shards.

  1. put the Hugging Face repo and full commit in repro.lock.json;
  2. put the GitHub repository and tag/commit in the Hugging Face model card;
  3. rerun make validate;
  4. commit, tag v1.0.0, and create a GitHub release;
  5. verify a fresh machine can download by the pinned Hub revision and start the endpoint without relying on an untracked file.

Do not upload model weights to GitHub or GitHub Releases. The Hub is designed for resumable, content-addressed model uploads and keeps the model card next to the tensors.

5. Serve

export MODEL_DIR=/models/qwen38-upload
make build-image
make preflight
make serve

Prebuilt image

Each release from v0.3.0 on is also published to GHCR by .github/workflows/publish-image.yml, built from the release tag with the same docker/Dockerfile. The model weights are not in the image. Pull by digest, not by tag; the digest for each release is in its release notes and the workflow run summary:

docker pull ghcr.io/dominikbucko/qwen38-flash-next-2x3090:v0.3.0
IMAGE=ghcr.io/dominikbucko/qwen38-flash-next-2x3090@sha256:<digest from the release notes> make serve

The image carries OCI labels for its source commit (org.opencontainers.image.revision) and base image (org.opencontainers.image.base.name), and a signed build-provenance attestation you can check with gh attestation verify oci://ghcr.io/dominikbucko/qwen38-flash-next-2x3090:v0.3.0 --owner DominikBucko. It is rebuilt in CI, so it is not byte-identical to the local image the September 25 numbers were measured on; the build inputs are. That run's exact settings, host and driver are in benchmarks/2026-09-25/environment.json.

Equivalent Compose launch:

export MODEL_DIR=/models/qwen38-upload
docker compose -f docker/compose.yaml up --build

The single-request default is hot84 with the full 262,144-token limit. Existing .env files override it: change VLLM_WNA16_STATIC_HOT_CACHE_SIZE to 84 when upgrading. Hot88 remains an optional tighter profile; test long prefill before using it. Neither setting removes experts or changes their precision.

The approximate QSA selector is default. To opt into exact selection:

export VLLM_QSA_EXACT_TOPK=1
make serve

The exact setting is a precision experiment, not the recommended performance profile. A single AgentBench A/B did not improve hidden-functional passes or mean score.

6. What CI can and cannot prove

GitHub CI checks Python syntax, overlay checksums, lock/Docker consistency, accidental model blobs, symlinks, cache files, and common token formats. It cannot validate CUDA kernels, 256K allocation, or throughput. Those checks need the target hardware and the procedure in benchmarks.md.

The September 16 fresh Docker check passed both the two-client 128K profile and a full-256K request as the first inference on hot84 defaults. It also verified all installed overlay hashes and bidirectional P2P copies inside the image. See the test record for exact versions and limits. This does not replace testing on a different host.