Skip to content

feat(quantization): support MOE_REQUANTIZE_WEIGHT_DTYPE and MOE_REQUANTIZE_BLOCK_SIZE in MXFP4 MoE - #3494

Open
aashishrampal-lab wants to merge 1 commit into
vllm-project:mainfrom
aashishrampal-lab:pr/moe-requantize-mxfp4
Open

feat(quantization): support MOE_REQUANTIZE_WEIGHT_DTYPE and MOE_REQUANTIZE_BLOCK_SIZE in MXFP4 MoE#3494
aashishrampal-lab wants to merge 1 commit into
vllm-project:mainfrom
aashishrampal-lab:pr/moe-requantize-mxfp4

Conversation

@aashishrampal-lab

Copy link
Copy Markdown
Contributor

Supports configurable requantization dtype (e.g. FP8) and block size for MXFP4 MoE weights.

Part 1 of the DeepSeek-V4 Pipeline Parallelism & Layerwise Loading stack.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant