Prebuilt FlashAttention wheel for modern NVIDIA GPUs (Blackwell / Ada / Hopper).
๐ pip install --no-deps https://github.com/CozzyTyler/flash-attn-cu130-wheel/releases/download/v2.8.4-cu130-torch2.12-triton3.7/flash_attn-2.8.4+cu130.torch2.12.triton3.7-cp310-cp310-linux_x86_64.whl
- CUDA: 13.0
- PyTorch: 2.12.0.dev+cu130
- Triton: 3.7.0
- Python: 3.10
- ABI: CXX11 (True)
Built on RTX 5090 Laptop GPU (Blackwell).
This wheel is NOT universal.
It is compiled for a specific stack:
- Torch 2.12 (nightly)
- CUDA 13.0
- Triton 3.7
๐ May NOT work on:
- stable PyTorch
- older CUDA
- different Triton
Download .whl from Releases and install:
pip install --no-deps https://github.com/CozzyTyler/flash-attn-cu130-wheel/releases/download/v2.8.4-cu130-torch2.12-triton3.7/flash_attn-2.8.4+cu130.torch2.12.triton3.7-cp310-cp310-linux_x86_64.whl