Skip to content

Milestones

List view

  • Ambitious feature-grade work: larger design-space and hardware-gated features — shape schemes for tensor functions, CDNA MFMA tensor cores, PoPE, CUDA pinned and constant memory.

    Due by October 28, 2026
    0/7 issues closed
  • Robustness pulled forward ahead of the feature milestones, while the issues still describe the code they were filed against: the engineering-hygiene, dedup, and test-seam follow-ups the v1.0 and v1.0.1 review cycles filed; scan, doc, and tooling consolidation; and soundness items (accumulator-width unification, the Metal pooled-accumulation bug, unenforced invariants, benchmark-trust fixes).

    Due by September 16, 2026
    134/202 issues closed
  • Consumers and explorations: models, reproductions, demos, integrations, and the training experience (checkpointing, tracking, plots). Paced by interest.

    Due by October 11, 2026
    0/15 issues closed
  • Performance-chasing in the approximate profile, demonstrated on benchmarks. The numerics-changing tier behind a third `approximate` preset (fused attention via online softmax, Winograd, tf32/fp16 arithmetic), the exact-numerics performance residue (footprint-scoped materialization, register-tile geometry, narrow-accumulator renderings, cost-model fidelity, async staging, memory pressure), and the benchmark legs that expose where OCANNL wins and where it loses (sequence-length and batch scaling, 3x3 convs, roofline and memory columns, Gemma 3 on real weights). An issue here closes with a before/after benchmark cell in a report.

    Due by October 3, 2026
    9/40 issues closed