Motivation
Online softmax and in-place computation could significantly save memory, especially for OPSD teacher TP sharding.
Keypoints
- Support the sharding LM heads with
gather_output=False of ColumnParallelLinear
- Enabling parallel CE loss feature (integrating Liger kernel)
- Profiling the GPU footprint to evaluate
- Unify the untied and tied path for LM heads
Motivation
Online softmax and in-place computation could significantly save memory, especially for OPSD teacher TP sharding.
Keypoints
gather_output=Falseof ColumnParallelLinear