This public artifact reproduces the routing and zeroth-order approximation claims from Sol-Attn (arXiv:2607.24027). We rebuilt the paper’s block-mean proxies, query-dependent mean/variance threshold, exact selected-block path, and proxy correction in PyTorch/Triton, then compared them with dense attention and a matched exact-only sparse control.
Assessment: partially reproduced. At 32K tokens, correction reduced exact-only L2 error by 70.8% on random, 94.9% on smooth, and 92.5% on temporal tensors while averaging 1.29–1.32× dense end-to-end across four seeds. Requested 5/10/15/25% densities calibrated closely without storing the full proxy map; routing state was 668.7× smaller at 128K. The paper reports up to 5.41× kernel and 2.1–2.3× model-level speedups; our best comparable end-to-end kernel harness result was 1.44× at 32K/10% and 1.34× at 128K/10%. Kernel incremental memory was 1.0000–1.0004× dense, so this reproduction supports the routing-memory mechanism but not a total-memory advantage.
The formal tests use single-head 64D tensors at 4K–128K rather than full video generation. A public SANA-600M QKV probe is included, but its released attention path is ReLU-linear rather than the paper’s softmax attention; its routing densities were not calibrated to the Gaussian targets. We did not evaluate VBench or the reported full-model speedups.
- Read the illustrated scientific report
- Explore the self-contained marimo tutorial
- Download the 104 terminal metric rows
- Inspect the PyTorch/Triton implementation
Compute: Kubernetes on NVIDIA RTX PRO 6000 Blackwell GPUs; peak 16 GPUs concurrently allocated; 0.43 elapsed wall-hours from first launch through the last successful terminal run.
Every formal node used the same inherited command. Branch links point to the exact committed code; failed setup nodes are shown only where they explain the lineage.
| Branch / experiment | Purpose or change | Exact run command | Assessment / outcome | Compute |
|---|---|---|---|---|
main |
Public landing page, report, notebook, data | Not run as an experiment (publication surface) | Presentation-only | — |
paper-faithful-fused-kernel-baseline |
Initial paper-faithful implementation | python -m torch.distributed.run --standalone --nproc_per_node=4 -m sol_attn_repro.run |
Environment wrapper failed before Python; fixed in child | Kubernetes, 4× RTX PRO 6000 Blackwell |
numerically-aligned-streaming-proxy |
Validate moment thresholds and streaming selection | python -m torch.distributed.run --standalone --nproc_per_node=4 -m sol_attn_repro.run |
8K counts exactly matched; correction improved error; scalar kernel slower than dense | Kubernetes, 4× RTX PRO 6000 Blackwell |
c32-instrumented-publication-candidate |
Stream 32 proxy-correction blocks and measure precomputed-state memory | python -m torch.distributed.run --standalone --nproc_per_node=4 -m sol_attn_repro.run |
1.22–1.44× at 32K and 1.09–1.34× at 128K; kernel memory near dense | Kubernetes, 4× RTX PRO 6000 Blackwell |
released-sana-600m-qkv-probe |
Capture QKV from a public released SANA checkpoint | python -m torch.distributed.run --standalone --nproc_per_node=4 -m sol_attn_repro.run |
Routing exact, but 0–11% observed density under 10/15% targets; diagnostic only | Kubernetes, 4× RTX PRO 6000 Blackwell |
final-random-seed-replication |
Four-seed 32K random replication | python -m torch.distributed.run --standalone --nproc_per_node=4 -m sol_attn_repro.run |
70.8% error reduction; 1.32× mean speedup | Kubernetes, 4× RTX PRO 6000 Blackwell |
final-smooth-seed-replication |
Four-seed 32K smooth replication | python -m torch.distributed.run --standalone --nproc_per_node=4 -m sol_attn_repro.run |
94.9% error reduction; 1.32× mean speedup | Kubernetes, 4× RTX PRO 6000 Blackwell |
final-temporal-seed-replication |
Four-seed 32K temporal replication | python -m torch.distributed.run --standalone --nproc_per_node=4 -m sol_attn_repro.run |
92.5% error reduction; 1.29× mean speedup | Kubernetes, 4× RTX PRO 6000 Blackwell |
final-long-seed-replication |
Four-seed 64K random replication | python -m torch.distributed.run --standalone --nproc_per_node=4 -m sol_attn_repro.run |
70.9% error reduction; 1.12× mean speedup | Kubernetes, 4× RTX PRO 6000 Blackwell |
📚 Docs | SANA | SANA-1.5 | SANA-Sprint | SANA-Video | SANA-WM | SANA-Streaming | Sol-RL
Demo | 🤗 HuggingFace | ComfyUI | SGLang | Cosmos-RL
SANA is an efficiency-oriented codebase for high-resolution image and video generation, providing complete training and inference pipelines. This repository contains code for SANA, SANA-1.5, SANA-Sprint, SANA-Video, SANA-WM, SANA-Streaming, and Sol-RL. More details can be found in our 📚 documentation.
Join our Discord to engage in discussions with the community! If you have any questions, run into issues, or are interested in contributing, don't hesitate to reach out!
- 🔥 [2026/07] 🌍 SANA-Streaming training is released! Includes bidirectional and distillation training. See Doc.
- 🔥 [2026/07] 🌍 SANA-WM Stage-1 training is released! Includes bidirectional, chunk-causal, and distillation training. See Doc.
- 🔥 [2026/06] 🎬 SANA-Streaming: 2B Model for Real-time Streaming Editing is released! Supports 720p, 1-min video editing. A pioneer work for streaming editing. See Project | Doc | Paper | Reactor Demo.
- 🔥 [2026/05] 🌍 SANA-WM: 2.6B Controllable World Model is released! Supports 720p, 1-min video generation with 6-DoF camera control. A new baseline for World Modeling and Embodied AI. See Project | Doc | Paper | Reactor Demo.
- 🔥 [2026/04] ⚡ Sol-RL: NVFP4 Rollout, BF16 Training RL is available! All training recipes for SANA, FLUX.1, and SD3.5-L, together with bundled post-training datasets, are released. See Sol-RL doc | Page | Paper.
- 🔥 [2026/03] 📺 SANA-Video 720p model with LTX-VAE is released. Use it with LTX2 Refiner to upscale the videos to 2K resolution! See Model Zoo, SANA-Video doc and Blog about refiner.
- 🔥 [2026/03] 💪 Post Training Infra: SANA × Cosmos-RL — We partner with Cosmos-RL to provide a complete RL infrastructure for SANA. You can now post-train (SFT/RL) SANA-Image and SANA-Video with state-of-the-art algorithms (e.g. Diffusion-NFT, Flow-GRPO), preset configs, reward services, and flexible datasets. See SANA on Cosmos-RL and our Cosmos-RL integration doc.
- 🔥 [2026/02] 🚀 SANA is now supported in SGLang! High-performance serving with OpenAI-compatible API. [Guidance]
- 🔥 [2026/01/26] SANA-Video is accepted as Oral by ICLR-2026. 🎉🎉🎉
- 🔥 [2025/12/09] 🎬 LongSANA: 27FPS real-time minute-length video generation model, training and inference code are all released. Thanks to LongLive Team. Refer to: [Train] | [Test] | [Weight]
- 🔥 [2025/11/24] 🪶 Blog: how Causal Linear Attention unlocks infinite context for LLMs and long video generation.
- 🔥 [2025/11/9] 🎬 Introduction video shows how Block Causal Linear Attention and Causal Mix-FFN work?
- 🔥 [2025/11/6] 📺SANA-Video is merged into diffusers. How to use.
- 🔥 [2025/10/27] 📺SANA-Video is released. [README] | [Weights] support Text-to-Video, TextImage-to-Video.
- 🔥 [2025/10/13] 📺SANA-Video is coming, 1). a 5s Linear DiT Video model, and 2). real-time minute-length video generation (with LongLive). [paper] | [Page]
Click to show all updates
- ✅ [2025/8/20] We release a new DC-AE-Lite for faster inference and smaller memory. [How to config] | [diffusers PR] | [Weight]
- ✅ [2025/6/25] SANA-Sprint was accepted to ICCV'25 🏖️
- ✅ [2025/6/4] SANA-Sprint ComfyUI Node is released [Example].
- ✅ [2025/5/8] SANA-Sprint (One-step diffusion) diffusers training code is released [Guidance].
- ✅ [2025/5/4] SANA-1.5 (Inference-time scaling) is accepted by ICML-2025. 🎉🎉🎉
- ✅ [2025/3/22] 🔥SANA-Sprint demo is hosted on Huggingface, try it! 🎉 [Demo Link]
- ✅ [2025/3/22] 🔥SANA-1.5 is supported in ComfyUI! 🎉: ComfyUI Guidance | ComfyUI Work Flow SANA-1.5 4.8B
- ✅ [2025/3/22] 🔥SANA-Sprint code & weights are released! 🎉 Include: Training & Inference code and Weights / HF are all released. [Guidance]
- ✅ [2025/3/21] 🚀Sana + Inference Scaling is released. [Guidance]
- ✅ [2025/3/16] 🔥SANA-1.5 code & weights are released! 🎉 Include: DDP/FSDP | TAR file WebDataset | Multi-Scale Training code and Weights | HF are all released.
- ✅ [2025/3/14] 🏃SANA-Sprint is coming out! 🎉 A new one/few-step generator of Sana. 0.1s per 1024px image on H100, 0.3s on RTX 4090. Find out more details: [Page] | [Arxiv]. Code is coming very soon along with
diffusers - ✅ [2025/2/10] 🚀Sana + ControlNet is released. [Guidance] | [Model] | [Demo]
- ✅ [2025/1/30] Release CAME-8bit optimizer code. Saving more GPU memory during training. [How to config]
- ✅ [2025/1/29] 🎉 🎉 🎉SANA 1.5 is out! Figure out how to do efficient training & inference scaling! 🚀[Tech Report]
- ✅ [2025/1/24] 4bit-Sana is released, powered by SVDQuant and Nunchaku inference engine. Now run your Sana within 8GB GPU VRAM [Guidance] [Demo] [Model]
- ✅ [2025/1/24] DCAE-1.1 is released, better reconstruction quality. [Model] [diffusers]
- ✅ [2025/1/23] Sana is accepted as Oral by ICLR-2025. 🎉🎉🎉
- ✅ [2025/1/12] DC-AE tiling makes Sana-4K inferences 4096x4096px images within 22GB GPU memory. With model offload and 8bit/4bit quantize. The 4K Sana run within 8GB GPU VRAM. [Guidance]
- ✅ [2025/1/11] Sana code-base license changed to Apache 2.0.
- ✅ [2025/1/10] Inference Sana with 8bit quantization.[Guidance]
- ✅ [2025/1/8] 4K resolution Sana models is supported in Sana-ComfyUI and work flow is also prepared. [4K guidance]
- ✅ [2025/1/8] 1.6B 4K resolution Sana models are released: [BF16 pth] or [BF16 diffusers]. 🚀 Get your 4096x4096 resolution images within 20 seconds! Find more samples in Sana page. Thanks SUPIR for their wonderful work and support.
- ✅ [2025/1/2] Bug in the
diffuserspipeline is solved. Solved PR - ✅ [2025/1/2] 2K resolution Sana models is supported in Sana-ComfyUI and work flow is also prepared.
- ✅ [2024/12] 1.6B 2K resolution Sana models are released: [BF16 pth] or [BF16 diffusers]. 🚀 Get your 2K resolution images within 4 seconds! Find more samples in Sana page. Thanks SUPIR for their wonderful work and support.
- ✅ [2024/12]
diffuserssupports Sana-LoRA fine-tuning! Sana-LoRA's training and convergence speed is super fast. [Guidance] or [diffusers docs]. - ✅ [2024/12]
diffusershas Sana! All Sana models in diffusers safetensors are released and diffusers pipelineSanaPipeline,SanaPAGPipeline,DPMSolverMultistepScheduler(with FlowMatching)are all supported now. We prepare a Model Card for you to choose. - ✅ [2024/12] 1.6B BF16 Sana model is released for stable fine-tuning.
- ✅ [2024/12] We release the ComfyUI node for Sana. [Guidance]
- ✅ [2024/11] All multi-linguistic (Emoji & Chinese & English) SFT models are released: 1.6B-512px, 1.6B-1024px, 600M-512px, 600M-1024px. The metric performance is shown here
- ✅ [2024/11] Sana Replicate API is launching at Sana-API.
- ✅ [2024/11] 1.6B Sana models are released.
- ✅ [2024/11] Training & Inference & Metrics code are released.
- ✅ [2024/11] Working on
diffusers. - [2024/10] Demo is released.
- [2024/10] DC-AE Code and weights are released!
- [2024/10] Paper is on Arxiv!
We introduce SANA, a series of efficient diffusion models for high-resolution image and video generation:
- SANA: Text-to-image generation up to 4K resolution, 20× smaller and 100× faster than Flux-12B.
- SANA-1.5: Efficient training-time and inference-time compute scaling for better quality.
- SANA-Sprint: One/few-step generation via sCM distillation, 0.1s per 1024px image on H100.
- SANA-Video/LongSANA: Efficient video generation with Block Linear Attention / with LongLive.
- Sol-RL: NVFP4 Rollout, BF16 Training RL achieves 4.64× faster convergence.
- SANA-WM: 2.6B parameter controllable world model, generating 720p, 1-minute video worlds with 6-DoF camera control.
- SANA-Streaming: 2B real-time streaming video-to-video editing for 720p, minute-scale videos.
Key Techniques:
- Linear Attention: Replace vanilla attention in DiT with linear attention for efficiency at high resolutions.
- DC-AE: 32× image compression (vs. traditional 8×) to reduce latent tokens.
- Decoder-only Text Encoder: Modern decoder-only LLM with in-context learning for better text-image alignment.
- Block Causal Linear Attention & Causal Mix-FFN: Efficient attention and feedforward for long video generation.
- Flow-DPM-Solver: Reduce sampling steps with efficient training and sampling.
- sCM Distillation: One/few-step generation with continuous-time consistency distillation.
- Sol-RL: Low precision(NVFP4) rollout selection, high precesion(BF16) optimization for faster RL training.
- Controllable World Modeling: Efficient long-context modeling and camera trajectory control for consistent world generation.
- Streaming Video Editing: Real-time long-form video-to-video editing with stable temporal consistency.
In summary, SANA is a fully open-source framework integrating efficient training, fast inference, and flexible deployment for both image and video generation. Deployable on laptop GPUs with < 8GB VRAM via 4-bit quantization.
git clone https://github.com/NVlabs/Sana.git
cd Sana && ./environment_setup.sh sanaimport torch
from diffusers import SanaPipeline
pipe = SanaPipeline.from_pretrained(
"Efficient-Large-Model/SANA1.5_1.6B_1024px_diffusers",
torch_dtype=torch.bfloat16,
)
pipe.to("cuda")
pipe.vae.to(torch.bfloat16)
pipe.text_encoder.to(torch.bfloat16)
prompt = 'a cyberpunk cat with a neon sign that says "Sana"'
image = pipe(
prompt=prompt,
height=1024,
width=1024,
guidance_scale=4.5,
num_inference_steps=20,
generator=torch.Generator(device="cuda").manual_seed(42),
)[0]
image[0].save("sana.png")Tip
Upgrade your diffusers>=0.32.0 to use SanaPipeline. More details can be found in 📚 Docs.
- 📚 Full Documentation
- Installation Guide
- Model Zoo
- Sana Inference & Training
- SANA-Sprint
- SANA-Video
- LongSANA
- SANA-WM
- SANA-Streaming
- ControlNet
- LoRA / DreamBooth
- Sol-RL Post-Training
- Quantization (4bit / 8bit)
- ComfyUI
- SGLang
| Methods (1024x1024) | Throughput (samples/s) | Latency (s) | Params (B) | Speedup | FID 👇 | CLIP 👆 | GenEval 👆 | DPG 👆 |
|---|---|---|---|---|---|---|---|---|
| FLUX-dev | 0.04 | 23.0 | 12.0 | 1.0× | 10.15 | 27.47 | 0.67 | 84.0 |
| Sana-0.6B | 1.7 | 0.9 | 0.6 | 39.5× | 5.81 | 28.36 | 0.64 | 83.6 |
| Sana-0.6B | 1.7 | 0.9 | 0.6 | 39.5× | 5.61 | 28.80 | 0.68 | 84.2 |
| Sana-1.6B | 1.0 | 1.2 | 1.6 | 23.3× | 5.92 | 28.94 | 0.69 | 84.5 |
| Sana-1.5 1.6B | 1.0 | 1.2 | 1.6 | 23.3× | 5.70 | 29.12 | 0.82 | 84.5 |
| Sana-1.5 4.8B | 0.26 | 4.2 | 4.8 | 6.5× | 5.99 | 29.23 | 0.81 | 84.7 |
| Models | Latency (s) | Params (B) | VBench Total ↑ | Quality ↑ | Semantic ↑ |
|---|---|---|---|---|---|
| Wan-2.1-14B | 1897 | 14 | 83.73 | 85.77 | 75.58 |
| Wan-2.1-1.3B | 400 | 1.3 | 83.38 | 85.67 | 74.22 |
| SANA-Video-2B | 36 | 2 | 84.05 | 84.63 | 81.73 |
We will try our best to achieve
- [✅] Training code
- [✅] Inference code
- [✅] Model zoo
- [✅] ComfyUI Nodes(SANA, SANA-1.5, SANA-Sprint)
- [✅] DC-AE Diffusers
- [✅] Sana merged in Diffusers(huggingface/diffusers#9982)
- [✅] LoRA training by @paul(
diffusers: https://github.com/ huggingface/diffusers/pull/10234) - [✅] 2K/4K resolution models.(Thanks @SUPIR to provide a 4K super-resolution model)
- [✅] 8bit / 4bit Laptop development
- [✅] ControlNet (train & inference & models)
- [✅] FSDP Training
- [✅] SANA-1.5 (Larger model size / Inference Scaling)
- [✅] SANA-Sprint: Few-step generator
- [✅] Faster DCAE-Lite weight
- [✅] Better re-construction F32/F64 VAEs
- [✅] SANA-Video: Linear DiT Video model, and real-time minute-length video generation
- [✅] RL Post-training: collaborate with Cosmos-RL
- [✅] SANA World Model
- [✅] SANA-Streaming Video-to-Video Editing
- [🚀] See you in the future
Thanks to the following open-sourced projects:
Thanks to the following open-sourced codebase for their wonderful work and codebase!
- PixArt-α
- PixArt-Σ
- diffusers
- Efficient-ViT
- ComfyUI_ExtraModels
- SVDQuant and Nunchaku
- Open-Sora
- Wan
- LTX-2
- LongLive
- Cosmos-RL
Thanks Paper2Video for generating Jeason presenting SANA😊. Refer to Paper2Video for more details.
Thanks go to these wonderful contributors:
@misc{xie2024sana,
title={Sana: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer},
author={Enze Xie and Junsong Chen and Junyu Chen and Han Cai and Haotian Tang and Yujun Lin and Zhekai Zhang and Muyang Li and Ligeng Zhu and Yao Lu and Song Han},
year={2024},
eprint={2410.10629},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2410.10629},
}
Click to expand all BibTeX citations
@misc{xie2025sana,
title={SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer},
author={Xie, Enze and Chen, Junsong and Zhao, Yuyang rectangle and Yu, Jincheng and Zhu, Ligeng and Lin, Yujun and Zhang, Zhekai and Li, Muyang and Chen, Junyu and Cai, Han and others},
year={2025},
eprint={2501.18427},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2501.18427},
}
@misc{chen2025sanasprint,
title={SANA-Sprint: One-Step Diffusion with Continuous-Time Consistency Distillation},
author={Junsong Chen and Shuchen Xue and Yuyang Zhao and Jincheng Yu graves and Sayak Paul and Junyu Chen and Han Cai and Song Han and Enze Xie},
year={2025},
eprint={2503.09641},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2503.09641},
}
@misc{chen2025sanavideo,
title={SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer},
author={Chen, Junsong and Zhao, Yuyang and Yu, Jincheng and Chu, Ruihang and Chen, Junyu and Yang, Shuai and Wang, Xianbang and Pan, Yicheng and Zhou, Daquan and Ling, Huan and others},
year={2025},
eprint={2509.24695},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2509.24695},
}
@misc{li2026fp4,
title={FP4 Explore, BF16 Train: Diffusion Reinforcement Learning via Efficient Rollout Scaling},
author={Li, Yitong and Chen, Junsong and Xue, Shuchen and Zeren, Pengcuo and Fu, Siyuan and Yang, Dinghao and Tang, Yangyang and Bai, Junjie and Luo, Ping and Han, Song and others},
year={2026}
eprint={2604.06916},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2604.06916},
}
@misc{zhu2026sanawm,
title={SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer},
author={Haoyi Zhu and Haozhe Liu and Yuyang Zhao and Tian Ye and Junsong Chen and Jincheng Yu and Tong He and Song Han and Enze Xie},
year={2026},
eprint={2605.15178},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2605.15178},
}
@misc{zhao2026sanastreamingrealtimestreamingvideo,
title={SANA-Streaming: Real-time Streaming Video Editing with Hybrid Diffusion Transformer},
author={Yuyang Zhao and Yicheng Pan and Qiyuan He and Jincheng Yu and Junsong Chen and Tian Ye and Haozhe Liu and Enze Xie and Song Han},
year={2026},
eprint={2605.30409},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2605.30409},
}

