Logging overall results here for reference. Aggregated by task, significantly better on all task (cross training contrast average, cross label average, cross eval contrast average):
Once we fix the HD95 aggregation (we shouldn't just raw average like this), we will most likely also beat all other methods significantly there too.
Logging overall results here for reference. Aggregated by task, significantly better on all task (cross training contrast average, cross label average, cross eval contrast average):
Once we fix the HD95 aggregation (we shouldn't just raw average like this), we will most likely also beat all other methods significantly there too.