If the model is optimized using the loss defined as the negative log-likelihood of the ground-truth preference from the paper, does this mean we only leverage the relative quality between the two edited images without using their exact numerical scores at all?
If the model is optimized using the loss defined as the negative log-likelihood of the ground-truth preference from the paper, does this mean we only leverage the relative quality between the two edited images without using their exact numerical scores at all?