Skip to content

[Bug]: qwen3.8-flash-next: assert numerator % denominator == 0, "{} is not divisible by {}".format #55517

Description

@Alireza3242

Your current environment

The output of python collect_env.py
Collecting environment information...
==============================
        System Info
==============================
OS                           : Ubuntu 24.04.2 LTS (x86_64)
GCC version                  : (Ubuntu 13.3.0-6ubuntu2~24.04.1) 13.3.0
Clang version                : Could not collect
CMake version                : Could not collect
Libc version                 : glibc-2.39

==============================
       PyTorch Info
==============================
PyTorch version              : 2.13.0+cu129
Is debug build               : False
CUDA used to build PyTorch   : 12.9
ROCM used to build PyTorch   : N/A
XPU used to build PyTorch    : N/A

==============================
      Python Environment
==============================
Python version               : 3.12.3 (main, Jul 15 2026, 23:46:41) [GCC 13.3.0] (64-bit runtime)
Python platform              : Linux-5.15.0-141-generic-x86_64-with-glibc2.39

==============================
       CUDA / GPU Info
==============================
Is CUDA available            : True
CUDA runtime version         : 12.9.86
CUDA_MODULE_LOADING set to   :
GPU models and configuration :
GPU 0: NVIDIA A100-SXM4-80GB
GPU 1: NVIDIA A100-SXM4-80GB
GPU 2: NVIDIA A100-SXM4-80GB

Nvidia driver version        : 550.163.01
cuDNN version                : Could not collect
HIP runtime version          : N/A
MIOpen runtime version       : N/A
Is XNNPACK available         : False

==============================
          CPU Info
==============================
Architecture:                         x86_64
CPU op-mode(s):                       32-bit, 64-bit
Address sizes:                        45 bits physical, 48 bits virtual
Byte Order:                           Little Endian
CPU(s):                               24
On-line CPU(s) list:                  0-23
Vendor ID:                            AuthenticAMD
Model name:                           AMD EPYC 7763 64-Core Processor
CPU family:                           25
Model:                                1
Thread(s) per core:                   1
Core(s) per socket:                   12
Socket(s):                            2
Stepping:                             1
BogoMIPS:                             4899.99
Flags:                                fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush mmx fxsr sse sse2 ht syscall nx mmxext fxsr_opt pdpe1gb rdtscp lm constant_tsc rep_good nopl tsc_reliable nonstop_tsc cpuid extd_apicid tsc_known_freq pni pclmulqdq ssse3 fma cx16 pcid sse4_1 sse4_2 x2apic movbe popcnt aes xsave avx f16c rdrand hypervisor lahf_lm cmp_legacy extapic cr8_legacy abm sse4a misalignsse 3dnowprefetch osvw topoext invpcid_single ibpb vmmcall fsgsbase bmi1 avx2 smep bmi2 erms invpcid rdseed adx smap clflushopt clwb sha_ni xsaveopt xsavec xgetbv1 xsaves clzero wbnoinvd arat umip pku ospke vaes vpclmulqdq rdpid overflow_recov succor fsrm
Hypervisor vendor:                    VMware
Virtualization type:                  full
L1d cache:                            768 KiB (24 instances)
L1i cache:                            768 KiB (24 instances)
L2 cache:                             12 MiB (24 instances)
L3 cache:                             64 MiB (2 instances)
NUMA node(s):                         1
NUMA node0 CPU(s):                    0-23
Vulnerability Gather data sampling:   Not affected
Vulnerability Itlb multihit:          Not affected
Vulnerability L1tf:                   Not affected
Vulnerability Mds:                    Not affected
Vulnerability Meltdown:               Not affected
Vulnerability Mmio stale data:        Not affected
Vulnerability Reg file data sampling: Not affected
Vulnerability Retbleed:               Not affected
Vulnerability Spec rstack overflow:   Mitigation; safe RET
Vulnerability Spec store bypass:      Vulnerable
Vulnerability Spectre v1:             Mitigation; usercopy/swapgs barriers and __user pointer sanitization
Vulnerability Spectre v2:             Mitigation; Retpolines; IBPB conditional; STIBP disabled; RSB filling; PBRSB-eIBRS Not affected; BHI Not affected
Vulnerability Srbds:                  Not affected
Vulnerability Tsx async abort:        Not affected

==============================
Versions of relevant libraries
==============================
[pip3] flashinfer-python==0.6.18
[pip3] nccl4py==0.5.0
[pip3] numpy==2.2.6
[pip3] nvidia-cublas-cu12==12.9.1.4
[pip3] nvidia-cuda-cccl-cu12==12.9.27
[pip3] nvidia-cuda-cupti-cu12==12.9.79
[pip3] nvidia-cuda-nvcc-cu12==12.9.86
[pip3] nvidia-cuda-nvdisasm==13.3.73
[pip3] nvidia-cuda-nvrtc-cu12==12.9.86
[pip3] nvidia-cuda-runtime-cu12==12.9.79
[pip3] nvidia-cudnn-cu12==9.20.0.48
[pip3] nvidia-cudnn-frontend==1.28.0
[pip3] nvidia-cufft-cu12==11.4.1.4
[pip3] nvidia-cufile-cu12==1.14.1.1
[pip3] nvidia-curand-cu12==10.3.10.19
[pip3] nvidia-cusolver-cu12==11.7.5.82
[pip3] nvidia-cusparse-cu12==12.5.10.65
[pip3] nvidia-cusparselt-cu12==0.8.1
[pip3] nvidia-cutlass-dsl==4.6.2
[pip3] nvidia-cutlass-dsl-libs-base==4.6.2
[pip3] nvidia-cutlass-dsl-libs-core==4.6.2
[pip3] nvidia-cutlass-dsl-libs-cu12==4.6.2
[pip3] nvidia-ml-py==13.610.43
[pip3] nvidia-nccl-cu12==2.30.7
[pip3] nvidia-nvjitlink-cu12==12.9.86
[pip3] nvidia-nvshmem-cu12==3.4.5
[pip3] nvidia-nvtx-cu12==12.9.79
[pip3] pyzmq==27.2.0
[pip3] tokenspeed-triton==3.8.10.post20260721
[pip3] torch==2.13.0+cu129
[pip3] torch_c_dlpack_ext==0.1.5
[pip3] torchaudio==2.11.0+cu129
[pip3] torchcodec==0.16.0+cu129
[pip3] torchvision==0.28.0+cu129
[pip3] transformers==5.16.1
[pip3] triton==3.7.1
[conda] Could not collect

==============================
         vLLM Info
==============================
ROCM Version                 : Could not collect
vLLM Version                 : 0.28.1rc1.dev437+ge962733e0 (git sha: e962733e0)
vLLM Build Flags:
  CUDA Archs: 7.5 8.0 8.6 8.9 9.0 10.0 12.0; ROCm: Disabled; XPU: Disabled
GPU Topology:
        GPU0    GPU1    GPU2    CPU Affinity    NUMA Affinity   GPU NUMA ID
GPU0     X      PHB     PHB     0-23    0               N/A
GPU1    PHB      X      PHB     0-23    0               N/A
GPU2    PHB     PHB      X      0-23    0               N/A

Legend:

  X    = Self
  SYS  = Connection traversing PCIe as well as the SMP interconnect between NUMA nodes (e.g., QPI/UPI)
  NODE = Connection traversing PCIe as well as the interconnect between PCIe Host Bridges within a NUMA node
  PHB  = Connection traversing PCIe as well as a PCIe Host Bridge (typically the CPU)
  PXB  = Connection traversing multiple PCIe bridges (without traversing the PCIe Host Bridge)
  PIX  = Connection traversing at most a single PCIe bridge
  NV#  = Connection traversing a bonded set of # NVLinks

==============================
     Environment Variables
==============================
NVIDIA_VISIBLE_DEVICES=all
VLLM_BUILD_URL=https://buildkite.com/vllm/release-v2/builds/6149
NVIDIA_REQUIRE_CUDA=cuda>=12.9 brand=unknown,driver>=535,driver<536 brand=grid,driver>=535,driver<536 brand=tesla,driver>=535,driver<536 brand=nvidia,driver>=535,driver<536 brand=quadro,driver>=535,driver<536 brand=quadrortx,driver>=535,driver<536 brand=nvidiartx,driver>=535,driver<536 brand=vapps,driver>=535,driver<536 brand=vpc,driver>=535,driver<536 brand=vcs,driver>=535,driver<536 brand=vws,driver>=535,driver<536 brand=cloudgaming,driver>=535,driver<536 brand=unknown,driver>=550,driver<551 brand=grid,driver>=550,driver<551 brand=tesla,driver>=550,driver<551 brand=nvidia,driver>=550,driver<551 brand=quadro,driver>=550,driver<551 brand=quadrortx,driver>=550,driver<551 brand=nvidiartx,driver>=550,driver<551 brand=vapps,driver>=550,driver<551 brand=vpc,driver>=550,driver<551 brand=vcs,driver>=550,driver<551 brand=vws,driver>=550,driver<551 brand=cloudgaming,driver>=550,driver<551 brand=unknown,driver>=560,driver<561 brand=grid,driver>=560,driver<561 brand=tesla,driver>=560,driver<561 brand=nvidia,driver>=560,driver<561 brand=quadro,driver>=560,driver<561 brand=quadrortx,driver>=560,driver<561 brand=nvidiartx,driver>=560,driver<561 brand=vapps,driver>=560,driver<561 brand=vpc,driver>=560,driver<561 brand=vcs,driver>=560,driver<561 brand=vws,driver>=560,driver<561 brand=cloudgaming,driver>=560,driver<561 brand=unknown,driver>=565,driver<566 brand=grid,driver>=565,driver<566 brand=tesla,driver>=565,driver<566 brand=nvidia,driver>=565,driver<566 brand=quadro,driver>=565,driver<566 brand=quadrortx,driver>=565,driver<566 brand=nvidiartx,driver>=565,driver<566 brand=vapps,driver>=565,driver<566 brand=vpc,driver>=565,driver<566 brand=vcs,driver>=565,driver<566 brand=vws,driver>=565,driver<566 brand=cloudgaming,driver>=565,driver<566 brand=unknown,driver>=570,driver<571 brand=grid,driver>=570,driver<571 brand=tesla,driver>=570,driver<571 brand=nvidia,driver>=570,driver<571 brand=quadro,driver>=570,driver<571 brand=quadrortx,driver>=570,driver<571 brand=nvidiartx,driver>=570,driver<571 brand=vapps,driver>=570,driver<571 brand=vpc,driver>=570,driver<571 brand=vcs,driver>=570,driver<571 brand=vws,driver>=570,driver<571 brand=cloudgaming,driver>=570,driver<571
TORCH_CUDA_ARCH_LIST=7.5 8.0 8.6 8.9 9.0 10.0 12.0
NVIDIA_DRIVER_CAPABILITIES=compute,utility
VLLM_IMAGE_TAG=vllm/vllm-openai:cu129-nightly-e962733e08d10f7ca65dac4df99e116460b8b174
VLLM_USAGE_SOURCE=production-docker-image
CUDA_VERSION=12.9.1
VLLM_ENABLE_CUDA_COMPATIBILITY=0
VLLM_BUILD_PIPELINE=019d130e-464e-4ff7-b84b-492992c0c06b
LD_LIBRARY_PATH=/usr/local/nvidia/lib64:/usr/local/cuda/lib64:/usr/local/cuda/lib64
VLLM_BUILD_COMMIT=e962733e08d10f7ca65dac4df99e116460b8b174
PYTORCH_NVML_BASED_CUDA_CHECK=1
TORCHINDUCTOR_COMPILE_THREADS=1
TORCHINDUCTOR_CACHE_DIR=/tmp/torchinductor_root

🐛 Describe the bug

export CUDA_VISIBLE_DEVICES=0,1,2
export PYTHONPATH="/app"
export VLLM_LOGGING_CONFIG_PATH=/app/logging_config.json
export VLLM_PLE_CPU_OFFLOAD=1

vllm serve /app/data/qwen3.8-flash-next-fp8/model \
  --gpu-memory-utilization 0.8 \
  --tensor-parallel-size 3 --enable-expert-parallel \
  --served-model-name qwen3.8-flash-next \
  --enable-prefix-caching \
  --no-enable-flashinfer-autotune \
  --default-chat-template-kwargs '{"enable_thinking": false, "temperature": 0.7, "top_p":0.8, "top_k":20, "min_p":0.0, "presence_penalty":1.5, "repetition_penalty":1.0}' \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_xml \
  --reasoning-parser qwen3 \
  --middleware src.timing_middleware.log_timing_middleware

error:

qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) INFO 09-06 00:38:49 model_runner.py:390] Loading model from scratch...
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) INFO 09-06 00:38:49 cuda.py:581] Using backend AttentionBackendEnum.FLASH_ATTN for vit attention
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944] WorkerProc failed to start.
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944] Traceback (most recent call last):
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/v1/executor/multiproc_executor.py", line 911, in worker_main
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]     worker = WorkerProc(*args, **kwargs)
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]              ^^^^^^^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/tracing/otel.py", line 178, in sync_wrapper
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]     return func(*args, **kwargs)
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]            ^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/v1/executor/multiproc_executor.py", line 680, in __init__
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]     self.worker.load_model()
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/v1/worker/gpu_worker.py", line 502, in load_model
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]     self.model_runner.load_model(load_dummy_weights=load_dummy_weights)
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/v1/worker/gpu/model_runner.py", line 394, in load_model
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]     self.model = model_loader.load_model(
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]                  ^^^^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/tracing/otel.py", line 178, in sync_wrapper
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]     return func(*args, **kwargs)
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]            ^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/model_executor/model_loader/base_loader.py", line 55, in load_model
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]     model = initialize_model(
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]             ^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/tracing/otel.py", line 178, in sync_wrapper
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]     return func(*args, **kwargs)
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]            ^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/model_executor/model_loader/utils.py", line 60, in initialize_model
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]     model = model_class(vllm_config=vllm_config, prefix=prefix)
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]             ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/models/qwen4_exp/nvidia/model.py", line 898, in __init__
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]     self.visual = Qwen3_VisionTransformer(
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]                   ^^^^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/model_executor/models/qwen3_vl.py", line 656, in __init__
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]     Qwen3_VisionBlock(
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/compilation/decorators.py", line 385, in __init__
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]     old_init(self, *args, **kwargs)
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/model_executor/models/qwen3_vl.py", line 465, in __init__
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]     self.attn = Qwen2_5_VisionAttention(
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]                 ^^^^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/model_executor/models/qwen2_5_vl.py", line 369, in __init__
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]     self.num_attention_heads_per_partition = dist_utils.divide(
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]                                              ^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/distributed/utils.py", line 63, in divide
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]     ensure_divisibility(numerator, denominator)
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/distributed/utils.py", line 55, in ensure_divisibility
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]     assert numerator % denominator == 0, "{} is not divisible by {}".format(
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944]            ^^^^^^^^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (Worker_TP0_EP0 pid=331) ERROR 09-06 00:38:50 multiproc_executor.py:944] AssertionError: 16 is not divisible by 3
qwen3.8-flash-next  | (EngineCore pid=261) INFO 09-06 00:38:50 multiproc_executor.py:472] [shutdown] Executor: waiting for worker exit count=3
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944] WorkerProc failed to start.
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944] Traceback (most recent call last):
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/v1/executor/multiproc_executor.py", line 911, in worker_main
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]     worker = WorkerProc(*args, **kwargs)
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]              ^^^^^^^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/tracing/otel.py", line 178, in sync_wrapper
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]     return func(*args, **kwargs)
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]            ^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/v1/executor/multiproc_executor.py", line 680, in __init__
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]     self.worker.load_model()
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/v1/worker/gpu_worker.py", line 502, in load_model
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]     self.model_runner.load_model(load_dummy_weights=load_dummy_weights)
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/v1/worker/gpu/model_runner.py", line 394, in load_model
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]     self.model = model_loader.load_model(
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]                  ^^^^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/tracing/otel.py", line 178, in sync_wrapper
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]     return func(*args, **kwargs)
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]            ^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/model_executor/model_loader/base_loader.py", line 55, in load_model
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]     model = initialize_model(
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]             ^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/tracing/otel.py", line 178, in sync_wrapper
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]     return func(*args, **kwargs)
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]            ^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/model_executor/model_loader/utils.py", line 60, in initialize_model
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]     model = model_class(vllm_config=vllm_config, prefix=prefix)
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]             ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/models/qwen4_exp/nvidia/model.py", line 898, in __init__
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]     self.visual = Qwen3_VisionTransformer(
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]                   ^^^^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/model_executor/models/qwen3_vl.py", line 656, in __init__
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]     Qwen3_VisionBlock(
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/compilation/decorators.py", line 385, in __init__
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]     old_init(self, *args, **kwargs)
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/model_executor/models/qwen3_vl.py", line 465, in __init__
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]     self.attn = Qwen2_5_VisionAttention(
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]                 ^^^^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/model_executor/models/qwen2_5_vl.py", line 369, in __init__
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]     self.num_attention_heads_per_partition = dist_utils.divide(
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]                                              ^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/distributed/utils.py", line 63, in divide
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]     ensure_divisibility(numerator, denominator)
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/distributed/utils.py", line 55, in ensure_divisibility
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]     assert numerator % denominator == 0, "{} is not divisible by {}".format(
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944]            ^^^^^^^^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (Worker_TP2_EP2 pid=449) ERROR 09-06 00:38:50 multiproc_executor.py:944] AssertionError: 16 is not divisible by 3
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944] WorkerProc failed to start.
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944] Traceback (most recent call last):
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/v1/executor/multiproc_executor.py", line 911, in worker_main
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]     worker = WorkerProc(*args, **kwargs)
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]              ^^^^^^^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/tracing/otel.py", line 178, in sync_wrapper
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]     return func(*args, **kwargs)
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]            ^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/v1/executor/multiproc_executor.py", line 680, in __init__
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]     self.worker.load_model()
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/v1/worker/gpu_worker.py", line 502, in load_model
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]     self.model_runner.load_model(load_dummy_weights=load_dummy_weights)
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/v1/worker/gpu/model_runner.py", line 394, in load_model
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]     self.model = model_loader.load_model(
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]                  ^^^^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/tracing/otel.py", line 178, in sync_wrapper
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]     return func(*args, **kwargs)
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]            ^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/model_executor/model_loader/base_loader.py", line 55, in load_model
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]     model = initialize_model(
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]             ^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/tracing/otel.py", line 178, in sync_wrapper
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]     return func(*args, **kwargs)
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]            ^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/model_executor/model_loader/utils.py", line 60, in initialize_model
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]     model = model_class(vllm_config=vllm_config, prefix=prefix)
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]             ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/models/qwen4_exp/nvidia/model.py", line 898, in __init__
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]     self.visual = Qwen3_VisionTransformer(
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]                   ^^^^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/model_executor/models/qwen3_vl.py", line 656, in __init__
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]     Qwen3_VisionBlock(
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/compilation/decorators.py", line 385, in __init__
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]     old_init(self, *args, **kwargs)
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/model_executor/models/qwen3_vl.py", line 465, in __init__
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]     self.attn = Qwen2_5_VisionAttention(
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]                 ^^^^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/model_executor/models/qwen2_5_vl.py", line 369, in __init__
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]     self.num_attention_heads_per_partition = dist_utils.divide(
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]                                              ^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/distributed/utils.py", line 63, in divide
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]     ensure_divisibility(numerator, denominator)
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]   File "/usr/local/lib/python3.12/dist-packages/vllm/distributed/utils.py", line 55, in ensure_divisibility
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]     assert numerator % denominator == 0, "{} is not divisible by {}".format(
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944]            ^^^^^^^^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (Worker_TP1_EP1 pid=399) ERROR 09-06 00:38:50 multiproc_executor.py:944] AssertionError: 16 is not divisible by 3
qwen3.8-flash-next  | [rank0]:[W906 00:38:51.794319004 ProcessGroupNCCL.cpp:1624] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
qwen3.8-flash-next  | (EngineCore pid=261) INFO 09-06 00:38:52 multiproc_executor.py:479] [shutdown] Executor: all workers exited gracefully
qwen3.8-flash-next  | (EngineCore pid=261) ERROR 09-06 00:38:52 core.py:1385] EngineCore failed to start.
qwen3.8-flash-next  | (EngineCore pid=261) ERROR 09-06 00:38:52 core.py:1385] Traceback (most recent call last):
qwen3.8-flash-next  | (EngineCore pid=261) ERROR 09-06 00:38:52 core.py:1385]   File "/usr/local/lib/python3.12/dist-packages/vllm/v1/engine/core.py", line 1347, in run_engine_core
qwen3.8-flash-next  | (EngineCore pid=261) ERROR 09-06 00:38:52 core.py:1385]     engine_core = EngineCoreProc(*args, engine_index=dp_rank, **kwargs)
qwen3.8-flash-next  | (EngineCore pid=261) ERROR 09-06 00:38:52 core.py:1385]                   ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (EngineCore pid=261) ERROR 09-06 00:38:52 core.py:1385]   File "/usr/local/lib/python3.12/dist-packages/vllm/tracing/otel.py", line 178, in sync_wrapper
qwen3.8-flash-next  | (EngineCore pid=261) ERROR 09-06 00:38:52 core.py:1385]     return func(*args, **kwargs)
qwen3.8-flash-next  | (EngineCore pid=261) ERROR 09-06 00:38:52 core.py:1385]            ^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (EngineCore pid=261) ERROR 09-06 00:38:52 core.py:1385]   File "/usr/local/lib/python3.12/dist-packages/vllm/v1/engine/core.py", line 1104, in __init__
qwen3.8-flash-next  | (EngineCore pid=261) ERROR 09-06 00:38:52 core.py:1385]     super().__init__(
qwen3.8-flash-next  | (EngineCore pid=261) ERROR 09-06 00:38:52 core.py:1385]   File "/usr/local/lib/python3.12/dist-packages/vllm/v1/engine/core.py", line 134, in __init__
qwen3.8-flash-next  | (EngineCore pid=261) ERROR 09-06 00:38:52 core.py:1385]     self.model_executor = executor_class(vllm_config)
qwen3.8-flash-next  | (EngineCore pid=261) ERROR 09-06 00:38:52 core.py:1385]                           ^^^^^^^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (EngineCore pid=261) ERROR 09-06 00:38:52 core.py:1385]   File "/usr/local/lib/python3.12/dist-packages/vllm/v1/executor/multiproc_executor.py", line 116, in __init__
qwen3.8-flash-next  | (EngineCore pid=261) ERROR 09-06 00:38:52 core.py:1385]     super().__init__(vllm_config)
qwen3.8-flash-next  | (EngineCore pid=261) ERROR 09-06 00:38:52 core.py:1385]   File "/usr/local/lib/python3.12/dist-packages/vllm/tracing/otel.py", line 178, in sync_wrapper
qwen3.8-flash-next  | (EngineCore pid=261) ERROR 09-06 00:38:52 core.py:1385]     return func(*args, **kwargs)
qwen3.8-flash-next  | (EngineCore pid=261) ERROR 09-06 00:38:52 core.py:1385]            ^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (EngineCore pid=261) ERROR 09-06 00:38:52 core.py:1385]   File "/usr/local/lib/python3.12/dist-packages/vllm/v1/executor/abstract.py", line 110, in __init__
qwen3.8-flash-next  | (EngineCore pid=261) ERROR 09-06 00:38:52 core.py:1385]     self._init_executor()
qwen3.8-flash-next  | (EngineCore pid=261) ERROR 09-06 00:38:52 core.py:1385]   File "/usr/local/lib/python3.12/dist-packages/vllm/v1/executor/multiproc_executor.py", line 214, in _init_executor
qwen3.8-flash-next  | (EngineCore pid=261) ERROR 09-06 00:38:52 core.py:1385]     self.workers = WorkerProc.wait_for_ready(unready_workers)
qwen3.8-flash-next  | (EngineCore pid=261) ERROR 09-06 00:38:52 core.py:1385]                    ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (EngineCore pid=261) ERROR 09-06 00:38:52 core.py:1385]   File "/usr/local/lib/python3.12/dist-packages/vllm/v1/executor/multiproc_executor.py", line 808, in wait_for_ready
qwen3.8-flash-next  | (EngineCore pid=261) ERROR 09-06 00:38:52 core.py:1385]     raise e from None
qwen3.8-flash-next  | (EngineCore pid=261) ERROR 09-06 00:38:52 core.py:1385] Exception: WorkerProc initialization failed due to an exception in a background process. See stack trace for root cause.
qwen3.8-flash-next  | (EngineCore pid=261) Process EngineCore:
qwen3.8-flash-next  | (EngineCore pid=261) Traceback (most recent call last):
qwen3.8-flash-next  | (EngineCore pid=261)   File "/usr/lib/python3.12/multiprocessing/process.py", line 314, in _bootstrap
qwen3.8-flash-next  | (EngineCore pid=261)     self.run()
qwen3.8-flash-next  | (EngineCore pid=261)   File "/usr/lib/python3.12/multiprocessing/process.py", line 108, in run
qwen3.8-flash-next  | (EngineCore pid=261)     self._target(*self._args, **self._kwargs)
qwen3.8-flash-next  | (EngineCore pid=261)   File "/usr/local/lib/python3.12/dist-packages/vllm/v1/engine/core.py", line 1389, in run_engine_core
qwen3.8-flash-next  | (EngineCore pid=261)     raise e
qwen3.8-flash-next  | (EngineCore pid=261)   File "/usr/local/lib/python3.12/dist-packages/vllm/v1/engine/core.py", line 1347, in run_engine_core
qwen3.8-flash-next  | (EngineCore pid=261)     engine_core = EngineCoreProc(*args, engine_index=dp_rank, **kwargs)
qwen3.8-flash-next  | (EngineCore pid=261)                   ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (EngineCore pid=261)   File "/usr/local/lib/python3.12/dist-packages/vllm/tracing/otel.py", line 178, in sync_wrapper
qwen3.8-flash-next  | (EngineCore pid=261)     return func(*args, **kwargs)
qwen3.8-flash-next  | (EngineCore pid=261)            ^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (EngineCore pid=261)   File "/usr/local/lib/python3.12/dist-packages/vllm/v1/engine/core.py", line 1104, in __init__
qwen3.8-flash-next  | (EngineCore pid=261)     super().__init__(
qwen3.8-flash-next  | (EngineCore pid=261)   File "/usr/local/lib/python3.12/dist-packages/vllm/v1/engine/core.py", line 134, in __init__
qwen3.8-flash-next  | (EngineCore pid=261)     self.model_executor = executor_class(vllm_config)
qwen3.8-flash-next  | (EngineCore pid=261)                           ^^^^^^^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (EngineCore pid=261)   File "/usr/local/lib/python3.12/dist-packages/vllm/v1/executor/multiproc_executor.py", line 116, in __init__
qwen3.8-flash-next  | (EngineCore pid=261)     super().__init__(vllm_config)
qwen3.8-flash-next  | (EngineCore pid=261)   File "/usr/local/lib/python3.12/dist-packages/vllm/tracing/otel.py", line 178, in sync_wrapper
qwen3.8-flash-next  | (EngineCore pid=261)     return func(*args, **kwargs)
qwen3.8-flash-next  | (EngineCore pid=261)            ^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (EngineCore pid=261)   File "/usr/local/lib/python3.12/dist-packages/vllm/v1/executor/abstract.py", line 110, in __init__
qwen3.8-flash-next  | (EngineCore pid=261)     self._init_executor()
qwen3.8-flash-next  | (EngineCore pid=261)   File "/usr/local/lib/python3.12/dist-packages/vllm/v1/executor/multiproc_executor.py", line 214, in _init_executor
qwen3.8-flash-next  | (EngineCore pid=261)     self.workers = WorkerProc.wait_for_ready(unready_workers)
qwen3.8-flash-next  | (EngineCore pid=261)                    ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (EngineCore pid=261)   File "/usr/local/lib/python3.12/dist-packages/vllm/v1/executor/multiproc_executor.py", line 808, in wait_for_ready
qwen3.8-flash-next  | (EngineCore pid=261)     raise e from None
qwen3.8-flash-next  | (EngineCore pid=261) Exception: WorkerProc initialization failed due to an exception in a background process. See stack trace for root cause.
qwen3.8-flash-next  | (APIServer pid=8) INFO 09-06 00:38:54 utils.py:620] [shutdown] Process manager: send sigterm to process EngineCore
qwen3.8-flash-next  | (APIServer pid=8) Traceback (most recent call last):
qwen3.8-flash-next  | (APIServer pid=8)   File "/usr/local/bin/vllm", line 10, in <module>
qwen3.8-flash-next  | (APIServer pid=8)     sys.exit(main())
qwen3.8-flash-next  | (APIServer pid=8)              ^^^^^^
qwen3.8-flash-next  | (APIServer pid=8)   File "/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/cli/main.py", line 97, in main
qwen3.8-flash-next  | (APIServer pid=8)     args.dispatch_function(args)
qwen3.8-flash-next  | (APIServer pid=8)   File "/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/cli/serve.py", line 153, in cmd
qwen3.8-flash-next  | (APIServer pid=8)     uvloop.run(run_server(args))
qwen3.8-flash-next  | (APIServer pid=8)   File "/usr/local/lib/python3.12/dist-packages/uvloop/__init__.py", line 96, in run
qwen3.8-flash-next  | (APIServer pid=8)     return __asyncio.run(
qwen3.8-flash-next  | (APIServer pid=8)            ^^^^^^^^^^^^^^
qwen3.8-flash-next  | (APIServer pid=8)   File "/usr/lib/python3.12/asyncio/runners.py", line 194, in run
qwen3.8-flash-next  | (APIServer pid=8)     return runner.run(main)
qwen3.8-flash-next  | (APIServer pid=8)            ^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (APIServer pid=8)   File "/usr/lib/python3.12/asyncio/runners.py", line 118, in run
qwen3.8-flash-next  | (APIServer pid=8)     return self._loop.run_until_complete(task)
qwen3.8-flash-next  | (APIServer pid=8)            ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (APIServer pid=8)   File "uvloop/loop.pyx", line 1518, in uvloop.loop.Loop.run_until_complete
qwen3.8-flash-next  | (APIServer pid=8)   File "/usr/local/lib/python3.12/dist-packages/uvloop/__init__.py", line 48, in wrapper
qwen3.8-flash-next  | (APIServer pid=8)     return await main
qwen3.8-flash-next  | (APIServer pid=8)            ^^^^^^^^^^
qwen3.8-flash-next  | (APIServer pid=8)   File "/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/launchers/api_server/entry.py", line 176, in run_server
qwen3.8-flash-next  | (APIServer pid=8)     await run_server_worker(listen_address, sock, args, **uvicorn_kwargs)
qwen3.8-flash-next  | (APIServer pid=8)   File "/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/launchers/api_server/entry.py", line 190, in run_server_worker
qwen3.8-flash-next  | (APIServer pid=8)     async with build_async_engine_client(
qwen3.8-flash-next  | (APIServer pid=8)   File "/usr/lib/python3.12/contextlib.py", line 210, in __aenter__
qwen3.8-flash-next  | (APIServer pid=8)     return await anext(self.gen)
qwen3.8-flash-next  | (APIServer pid=8)            ^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (APIServer pid=8)   File "/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/launchers/api_server/entry.py", line 58, in build_async_engine_client
qwen3.8-flash-next  | (APIServer pid=8)     async with build_async_engine_client_from_engine_args(
qwen3.8-flash-next  | (APIServer pid=8)   File "/usr/lib/python3.12/contextlib.py", line 210, in __aenter__
qwen3.8-flash-next  | (APIServer pid=8)     return await anext(self.gen)
qwen3.8-flash-next  | (APIServer pid=8)            ^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (APIServer pid=8)   File "/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/launchers/api_server/entry.py", line 94, in build_async_engine_client_from_engine_args
qwen3.8-flash-next  | (APIServer pid=8)     async_llm = AsyncLLM.from_vllm_config(
qwen3.8-flash-next  | (APIServer pid=8)                 ^^^^^^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (APIServer pid=8)   File "/usr/local/lib/python3.12/dist-packages/vllm/v1/engine/async_llm.py", line 225, in from_vllm_config
qwen3.8-flash-next  | (APIServer pid=8)     return cls(
qwen3.8-flash-next  | (APIServer pid=8)            ^^^^
qwen3.8-flash-next  | (APIServer pid=8)   File "/usr/local/lib/python3.12/dist-packages/vllm/v1/engine/async_llm.py", line 160, in __init__
qwen3.8-flash-next  | (APIServer pid=8)     self.engine_core = EngineCoreClient.make_async_mp_client(
qwen3.8-flash-next  | (APIServer pid=8)                        ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (APIServer pid=8)   File "/usr/local/lib/python3.12/dist-packages/vllm/tracing/otel.py", line 178, in sync_wrapper
qwen3.8-flash-next  | (APIServer pid=8)     return func(*args, **kwargs)
qwen3.8-flash-next  | (APIServer pid=8)            ^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (APIServer pid=8)   File "/usr/local/lib/python3.12/dist-packages/vllm/v1/engine/core_client.py", line 163, in make_async_mp_client
qwen3.8-flash-next  | (APIServer pid=8)     return AsyncMPClient(
qwen3.8-flash-next  | (APIServer pid=8)            ^^^^^^^^^^^^^^
qwen3.8-flash-next  | (APIServer pid=8)   File "/usr/local/lib/python3.12/dist-packages/vllm/tracing/otel.py", line 178, in sync_wrapper
qwen3.8-flash-next  | (APIServer pid=8)     return func(*args, **kwargs)
qwen3.8-flash-next  | (APIServer pid=8)            ^^^^^^^^^^^^^^^^^^^^^
qwen3.8-flash-next  | (APIServer pid=8)   File "/usr/local/lib/python3.12/dist-packages/vllm/v1/engine/core_client.py", line 1060, in __init__
qwen3.8-flash-next  | (APIServer pid=8)     super().__init__(
qwen3.8-flash-next  | (APIServer pid=8)   File "/usr/local/lib/python3.12/dist-packages/vllm/v1/engine/core_client.py", line 656, in __init__
qwen3.8-flash-next  | (APIServer pid=8)     with launch_core_engines(
qwen3.8-flash-next  | (APIServer pid=8)   File "/usr/lib/python3.12/contextlib.py", line 144, in __exit__
qwen3.8-flash-next  | (APIServer pid=8)     next(self.gen)
qwen3.8-flash-next  | (APIServer pid=8)   File "/usr/local/lib/python3.12/dist-packages/vllm/v1/engine/utils.py", line 1240, in launch_core_engines
qwen3.8-flash-next  | (APIServer pid=8)     wait_for_engine_startup(
qwen3.8-flash-next  | (APIServer pid=8)   File "/usr/local/lib/python3.12/dist-packages/vllm/v1/engine/utils.py", line 1320, in wait_for_engine_startup
qwen3.8-flash-next  | (APIServer pid=8)     raise RuntimeError(
qwen3.8-flash-next  | (APIServer pid=8) RuntimeError: Engine core initialization failed. See root cause above. Failed core proc(s): {}

Before submitting a new issue...

  • Make sure you already searched for relevant issues, and asked the chatbot living at the bottom right corner of the documentation page, which can answer lots of frequently asked questions.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions