v0.8.0

awni released this 21 Mar 21:00

· 1157 commits to main since this release

44390bd

Highlights

More perf!
mx.fast.rms_norm and mx.fast.layer_norm
Switch to Nanobind substantially reduces overhead
Up to 4x faster __setitem__ (e.g. a[...] = b)

Core

mx.inverse, CPU only
vmap over mx.matmul and mx.addmm
Switch to nanobind from pybind11
Faster setitem indexing
- Benchmarks
mx.fast.rms_norm, token generation benchmark
mx.fast.layer_norm, token generation benchmark
vmap for inverse and svd
Faster non-overlapping pooling

Optimizers

Set minimum value in cosine decay scheduler

Bugfixes

Fix bug in multi-dimensional reduction

Assets 2