-
Notifications
You must be signed in to change notification settings - Fork 11
Expand file tree
/
Copy pathagent-128k.env
More file actions
23 lines (21 loc) · 969 Bytes
/
Copy pathagent-128k.env
File metadata and controls
23 lines (21 loc) · 969 Bytes
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
# Agent profile for 2x RTX 3090 + 128 GiB RAM: 128K context, hot100.
# Copy these values into .env. Same runtime switches as configs/fast-256k.env;
# the context window is about half, and the KV memory it frees holds 16 more
# hot experts per GPU, which removes decode expert misses. Single request,
# up to 131,072 prompt tokens plus 4,096 generated tokens.
MAX_MODEL_LEN=135168
# 110 KV blocks: a full 135,168-token request plus the extra recurrent-state
# block per Mamba group that async prefill holds (reported concurrency 1.04x).
KV_CACHE_MEMORY_BYTES=2500000000
VLLM_WNA16_STATIC_HOT_CACHE_SIZE=100
# Runtime switches, as in configs/fast-256k.env.
QWEN38_STREAM_STAGE=1
QWEN38_STREAM_STAGE_MIN_TOKENS=1536
QWEN38_STAGE_OVERLAP=1
QWEN38_ASYNC_SCHEDULING=1
QWEN38_TRITON_SKINNY=1
VLLM_MTP_DRAFT_VOCAB_RANGES=[[0,65536],[248044,248320]]
DISABLE_CUSTOM_ALL_REDUCE=0
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:False
QWEN38_PLE_PREFAULT=1
JIT_CACHE_DIR=./jit-cache