The unified attention autotuning pruner results in no valid autotuner configs for many testing configurations in the new vLLM test-suite vllm_tdesc that was added in #7413. The vllm_tdesc test-suite first applies the local unified_attention.patch patch before running the unified attention tests to verify the changes in the patch.
Why is the autotuning pruner pruning the original hardcoded config that passes in the vllm_triton_attn test-suite?
Occurs on PVC and BMG.
Error:
triton.runtime.errors.AutotunerError: Autotuner error: No valid autotuner configs after pruning. `early_config_prune` should return at least one config.
Example failing configurations (see skiplist for the full list):
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-32768-None-dtype0-None-16-128-num_heads2-seq_lens0]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-32768-None-dtype0-None-16-128-num_heads2-seq_lens1]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-32768-None-dtype0-None-16-256-num_heads2-seq_lens0]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-32768-None-dtype0-None-16-256-num_heads2-seq_lens1]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-32768-None-dtype0-64-16-128-num_heads2-seq_lens0]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-32768-None-dtype0-64-16-128-num_heads2-seq_lens1]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-32768-None-dtype0-64-16-256-num_heads2-seq_lens0]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-32768-None-dtype0-64-16-256-num_heads2-seq_lens1]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-32768-None-dtype0-128-16-128-num_heads2-seq_lens0]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-32768-None-dtype0-128-16-128-num_heads2-seq_lens1]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-32768-None-dtype0-128-16-256-num_heads2-seq_lens0]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-32768-None-dtype0-128-16-256-num_heads2-seq_lens1]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-32768-None-dtype0-256-16-128-num_heads2-seq_lens0]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-32768-None-dtype0-256-16-128-num_heads2-seq_lens1]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-32768-None-dtype0-256-16-256-num_heads2-seq_lens0]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-32768-None-dtype0-256-16-256-num_heads2-seq_lens1]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-32768-50.0-dtype0-None-16-128-num_heads2-seq_lens0]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-32768-50.0-dtype0-None-16-128-num_heads2-seq_lens1]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-32768-50.0-dtype0-None-16-256-num_heads2-seq_lens0]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-32768-50.0-dtype0-None-16-256-num_heads2-seq_lens1]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-32768-50.0-dtype0-64-16-128-num_heads2-seq_lens0]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-32768-50.0-dtype0-64-16-128-num_heads2-seq_lens1]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-32768-50.0-dtype0-64-16-256-num_heads2-seq_lens0]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-32768-50.0-dtype0-64-16-256-num_heads2-seq_lens1]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-32768-50.0-dtype0-128-16-128-num_heads2-seq_lens0]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-32768-50.0-dtype0-128-16-128-num_heads2-seq_lens1]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-32768-50.0-dtype0-128-16-256-num_heads2-seq_lens0]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-32768-50.0-dtype0-128-16-256-num_heads2-seq_lens1]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-32768-50.0-dtype0-256-16-128-num_heads2-seq_lens0]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-32768-50.0-dtype0-256-16-128-num_heads2-seq_lens1]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-32768-50.0-dtype0-256-16-256-num_heads2-seq_lens0]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-32768-50.0-dtype0-256-16-256-num_heads2-seq_lens1]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-2048-None-dtype0-None-16-128-num_heads2-seq_lens0]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-2048-None-dtype0-None-16-128-num_heads2-seq_lens1]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-2048-None-dtype0-None-16-256-num_heads2-seq_lens0]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-2048-None-dtype0-None-16-256-num_heads2-seq_lens1]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-2048-None-dtype0-64-16-128-num_heads2-seq_lens0]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-2048-None-dtype0-64-16-128-num_heads2-seq_lens1]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-2048-None-dtype0-64-16-256-num_heads2-seq_lens0]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-2048-None-dtype0-64-16-256-num_heads2-seq_lens1]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-2048-None-dtype0-128-16-128-num_heads2-seq_lens0]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-2048-None-dtype0-128-16-128-num_heads2-seq_lens1]
tests/kernels/attention/test_triton_unified_attention.py::test_triton_unified_attn[0-None-2048-None-dtype0-128-16-256-num_heads2-seq_lens0]
...
The unified attention autotuning pruner results in no valid autotuner configs for many testing configurations in the new vLLM test-suite
vllm_tdescthat was added in #7413. Thevllm_tdesctest-suite first applies the localunified_attention.patchpatch before running the unified attention tests to verify the changes in the patch.Why is the autotuning pruner pruning the original hardcoded config that passes in the
vllm_triton_attntest-suite?Occurs on PVC and BMG.
Error:
Example failing configurations (see skiplist for the full list):