Just noticed in the README that you plan to measure perplexity... I don't expect much / any improvement in this metric – the hypothesis is that weight / activation kurtosis will be reduced, thus facilitating model compressibility rather than performance.
Perplexity of heavily quantized models may be interesting though.
Just noticed in the README that you plan to measure perplexity... I don't expect much / any improvement in this metric – the hypothesis is that weight / activation kurtosis will be reduced, thus facilitating model compressibility rather than performance.
Perplexity of heavily quantized models may be interesting though.