[CUDA] Support FP8 (E4M3) KV Cache for Group Query Attention #10520
Triggered via pull request
February 14, 2026 04:12
Status
Success
Total duration
20m 52s
Artifacts
–
build_x64_release_ep_generic_interface
18m 26s
Annotations
6 warnings
|
build_x64_release_ep_generic_interface:
onnxruntime/core/mlas/lib/amd64/QgemmU8X8KernelAvx2.asm#L1234
epilog offset from end of function exceeds 4095
|
|
build_x64_release_ep_generic_interface:
onnxruntime/core/mlas/lib/amd64/QgemmU8X8KernelAvx2.asm#L1227
epilog offset from end of function exceeds 4095
|
|
build_x64_release_ep_generic_interface:
onnxruntime/core/mlas/lib/amd64/QgemmU8X8KernelAvx2.asm#L1220
epilog offset from end of function exceeds 4095
|
|
build_x64_release_ep_generic_interface:
onnxruntime/core/mlas/lib/amd64/QgemmU8X8KernelAvx2.asm#L1213
epilog offset from end of function exceeds 4095
|
|
build_x64_release_ep_generic_interface:
onnxruntime/core/mlas/lib/amd64/QgemmU8X8KernelAvx2.asm#L1206
epilog offset from end of function exceeds 4095
|
|
build_x64_release_ep_generic_interface:
onnxruntime/core/mlas/lib/amd64/QgemmU8X8KernelAvx2.asm#L1199
epilog offset from end of function exceeds 4095
|