Commit f3db7b0
authored
Cast the lm_head weight to the compute dtype in the chunked projections
Under FSDP2 mixed precision the patched forward reads the fp32 sharded master
weight directly while the hidden states come out bf16, so the tensor-core
projection failed with a dtype mismatch (tests/distributed test_sft_peft[fsdp2]).
Cast the weight to the hidden-states dtype, which is what lm_head's own forward
computes in. Same fix in all three copies: SFT, distillation, async distillation.1 parent a1ab3db commit f3db7b0
3 files changed
Lines changed: 14 additions & 9 deletions
File tree
- trl
- experimental/async_distillation
- trainer
Lines changed: 3 additions & 2 deletions
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
225 | 225 | | |
226 | 226 | | |
227 | 227 | | |
228 | | - | |
229 | | - | |
| 228 | + | |
| 229 | + | |
| 230 | + | |
230 | 231 | | |
231 | 232 | | |
232 | 233 | | |
| |||
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
112 | 112 | | |
113 | 113 | | |
114 | 114 | | |
115 | | - | |
116 | | - | |
117 | | - | |
| 115 | + | |
| 116 | + | |
| 117 | + | |
| 118 | + | |
| 119 | + | |
118 | 120 | | |
119 | 121 | | |
120 | 122 | | |
| |||
125 | 127 | | |
126 | 128 | | |
127 | 129 | | |
128 | | - | |
| 130 | + | |
129 | 131 | | |
130 | 132 | | |
131 | 133 | | |
| |||
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
99 | 99 | | |
100 | 100 | | |
101 | 101 | | |
102 | | - | |
103 | | - | |
104 | | - | |
| 102 | + | |
| 103 | + | |
| 104 | + | |
| 105 | + | |
| 106 | + | |
105 | 107 | | |
106 | 108 | | |
107 | 109 | | |
| |||
0 commit comments