-
Notifications
You must be signed in to change notification settings - Fork 11
Expand file tree
/
Copy pathfast-256k.env
More file actions
24 lines (23 loc) · 1.15 KB
/
Copy pathfast-256k.env
File metadata and controls
24 lines (23 loc) · 1.15 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
# Fast full-context profile for 2x RTX 3090 + 128 GiB RAM (September 25 runtime).
# Copy these values into .env. Single request, 262,144 tokens, BF16 KV, hot84.
# Requires working bidirectional CUDA P2P (scripts/check_custom_all_reduce.py)
# and enough RAM to keep the whole 51.2 GB PLE table resident: stop other
# memory-heavy jobs while serving.
# Stream each layer's cold experts to the GPU ahead of large prefill chunks.
QWEN38_STREAM_STAGE=1
QWEN38_STREAM_STAGE_MIN_TOKENS=1536
# Overlap down-projection expert copies with the gate/up GEMM in decode.
QWEN38_STAGE_OVERLAP=1
# Overlap host input preparation with GPU work.
QWEN38_ASYNC_SCHEDULING=1
# Tensor-core skinny GEMM for decode-sized dense projections.
QWEN38_TRITON_SKINNY=1
# Draft-only restricted MTP vocabulary; target verification is unchanged.
VLLM_MTP_DRAFT_VOCAB_RANGES=[[0,65536],[248044,248320]]
# Custom P2P all-reduce needs the non-expandable allocator.
DISABLE_CUSTOM_ALL_REDUCE=0
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:False
# Re-read the PLE table into RAM after the checkpoint load evicts it.
QWEN38_PLE_PREFAULT=1
# Persistent kernel caches (created if missing).
JIT_CACHE_DIR=./jit-cache