When training with DAPO, combining overlong punishment with default reward results in negative returns. I believe there should be a better mechanism to combine these rewards/punishments (not just simply adding them together, but rather an alpha or tunable parameter applied to penalties)

When training with DAPO, combining overlong punishment with default reward results in negative returns. I believe there should be a better mechanism to combine these rewards/punishments (not just simply adding them together, but rather an alpha or tunable parameter applied to penalties)