Commit a2bdd22
committed
chore(release): v0.2.0 — AWQ Option B + 8-bit embed_tokens
Bump lumen-app + frontend + tauri to 0.2.0.
What ships:
- `hsng95/gemma-4-26b-a4b-mlx-imatrix3plus-awq` as the default Gemma 4
catalog entry (12 GB, 3.916 bpw). Multi-seed mean PPL 16.2 ppl
better than the no-AWQ imatrix3plus baseline AND decode is faster
than uniform 4-bit on M3 Max (73 vs 71 tok/s).
- bf16 / mixed-precision embed_tokens support in the Gemma 4 native
loader. Future skip-list quant builds will no longer trip the
"missing embed_tokens.scales" load error.
- Full AWQ tooling under scripts/quant/ (awq_search, awq_apply,
awq_filter_scales, test_awq_apply_parity, test_long_context_extended).
Upgrade path: lumen-app's HF Hub revision tracking detects the new
commit on the same repo id, prompts Update on the MODELS card. The
old v0.1.3 download (still at the bf16-embed revision) will be
replaced in-place by the 8-bit-embed ship build.1 parent 4b88613 commit a2bdd22
4 files changed
Lines changed: 4 additions & 4 deletions
Some generated files are not rendered by default. Learn more about customizing how changed files appear on GitHub.
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
1 | 1 | | |
2 | 2 | | |
3 | | - | |
| 3 | + | |
4 | 4 | | |
5 | 5 | | |
6 | 6 | | |
| |||
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
1 | 1 | | |
2 | 2 | | |
3 | 3 | | |
4 | | - | |
| 4 | + | |
5 | 5 | | |
6 | 6 | | |
7 | 7 | | |
| |||
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
1 | 1 | | |
2 | 2 | | |
3 | 3 | | |
4 | | - | |
| 4 | + | |
5 | 5 | | |
6 | 6 | | |
7 | 7 | | |
| |||
0 commit comments