## Parent Issue Part of #206 — Megatron Roadmap, **Phase 4 (Advanced Features / Stretch): Task 4.2** ## Description Implement expert parallelism communication patterns for Mixture-of-Experts models. ## Requirements - [ ] All-to-all dispatch for token routing to experts - [ ] All-to-all combine for gathering expert outputs - [ ] Capacity factor and load balancing support - [ ] Compatible with Megatron's MoE layer implementation - [ ] Works with EP process groups ## Blocked By - #260 (DeviceMesh beyond 2D) - #253 (Differentiable collectives) ## Blocks - MoE model training on candle
Parent Issue
Part of #206 — Megatron Roadmap, Phase 4 (Advanced Features / Stretch): Task 4.2
Description
Implement expert parallelism communication patterns for Mixture-of-Experts models.
Requirements
Blocked By
Blocks