feat(groot): GR00T N1.6 × Jetson Thor (SM110) — HF parity, 12 bug fixes, 130→28.5 ms - #177
Open
DXICM wants to merge 4 commits into
Open
feat(groot): GR00T N1.6 × Jetson Thor (SM110) — HF parity, 12 bug fixes, 130→28.5 ms#177DXICM wants to merge 4 commits into
DXICM wants to merge 4 commits into
Conversation
Root-cause and fix 12 real bugs where the upstream N1.6 frontend
inherited openpi-family (Pi0/Pi0.5) vision/kernel assumptions that
do not hold for GR00T N1.6's HF behaviour:
1. Tokenization: reproduce Eagle chat template (system/user headers,
formalize, per-view image blocks) instead of bare encode()
2. Resolution: HF eval chain outputs 252x252, not 224
3. SigLIP attention scope: HF(sdpa) does cross-view full attention
on the packed 648-token sequence, not per-view
4. Patch flatten order: HF NaFlex uses (ph,pw,C), not (C,ph,pw)
5. Strided FMHA divergence on non-power-of-2 seq with real data:
parity mode routes SigLIP attention through torch sdpa
6. CKernelQwen3 diverges from HF on real sequences: parity mode
runs HF-native Qwen3Model (bf16, sdpa, graph-captured)
7. Wild pointer after re-capture: Qwen3 graph-captured LN referenced
local tensors; promote to persistent attributes + finiteness guard
8. adaLN chunk order reversed: HF proj_out_1 is (shift, scale)
9. Single-frame FP8 calibration too narrow: multi-frame calibrate
(current + 7 synthetic frames, percentile=99.9)
10. Prompt switch rejected after graph bake: detect change, reset
graph runtime, re-set prompt, re-capture
11. Idle-first-frame garbage: Thor GPU idle reset invalidates captured
graphs; add replay finiteness self-check + re-capture retry
12. Prompt-switch re-capture device-side assert: stale DiT static
buffers/indices not rebuilt; add to stale list
Precision vs HF eager: cos 0.999933 / maxd 0.059 (denormalized action).
No inference hyperparameters changed (4-step, 252x252, T=50, bf16).
Also adds tools/convert_groot_n16_hf_checkpoint.py for HF safetensors
to FlashRT layout conversion (Qwen3 16-layer truncation, DiT repack,
SigLIP mlp1 layout).
New CUDA kernels for the N1.6 Thor NVFP4 pipeline: - fused_fp4/silu_mul_fp4_sfa_bf16: SiLU(gate)*up (bf16) direct to NVFP4+SFA, bit-exact vs torch two-step chain - fused_fp4/dit_norm_fp4_sfa: AdaLN / no-affine LN / weighted RMSNorm direct to NVFP4+SFA (bf16 input variants) - gemm/fp4/cutlass_fp4_gemm_bias_bf16_sm100: bias / bias+residual / bias+tanh-GELU+fp4out epilogue variants - quantize/quantize_fp4_sfa_bf16: vectorized bf16 dynamic quantize - kernels/qk_norm_rope_rotate_half_bf16: fused per-head RMSNorm + rotate-half RoPE (bf16, in-place, one launch per Q/K) Performance rounds (no hyperparameter changes): - DiT NVFP4 fused epilogue: 36.6 -> 15.7 ms (8 kernels/layer) - Qwen3 fused norm/rope/GQA: 12.7 -> 5.0 ms (cos 0.999986) - SigLIP FA4 + fp4 encoder: 10.3 -> 6.9 ms (cos 0.999988) - SigLIP embeddings in-graph: 34 -> 28.5 ms (bit-exact) - E2E total: 130 -> 28.5 ms (4-step, 2-camera, 252x252, T=50) Bandwidth ceiling: Thor measured 252-255 GB/s (~93% of 273 spec); DiT 15.2 ms is weight-bandwidth-bound floor for this config. Tier switches (all default ON, independently fall back): FLASHRT_N16_DIT_FP4, FLASHRT_N16_QWEN3_FP4, FLASHRT_N16_SIGLIP_FP4, FLASHRT_N16_FA4
- docs/groot_n16_thor_sm110.md: single authoritative document covering architecture facts, 12-bug root-cause table, falsified hypotheses, full optimization record (130 -> 28.5 ms), roofline/bandwidth ceiling analysis (252-255 GB/s, ~93% of spec), precision tier switches, and verification methodology. - docs/groot_transformers5_weight_corruption.md: transformers>=5 silent weight corruption via _initialize_missing_keys re-randomizing SigLIP2 vision tower (282 tensors). One-line fix + integrity guard. - docs/thor_gpu_idle_reset_workaround.md: Thor GPU idle reset defect and three-layer CUDA Graph protection (keepalive, idle reinit, finiteness).
cutlass-dsl caches the device arch at import time. The previous code imported cutlass to check its version, then set CUTE_DSL_ARCH=sm_101a — too late; NVVM already cached sm_110a and ICEs on the hd256 2CTA kernel (introduced in flashrt-project#164, commit 7fd75d2). Fix: set CUTE_DSL_ARCH=sm_101a unconditionally before any cutlass import. Also revert the hd256 2CTA dispatch to SM100-only (the dedicated kernel was never validated on SM110) and restore the _fa4_trimmed lazy loader for BlackwellFusedMultiHeadAttentionForward. Verified: all-tier E2E on Thor — median 27.7 ms, p95 28.5 ms, actions finite, cos 0.999933 vs HF eager.
DXICM
force-pushed
the
feat/groot-n16-thor
branch
from
August 18, 2026 03:52
817cbf7 to
4cae6c0
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
GR00T N1.6-3B full adaptation for Jetson AGX Thor (SM110): HF numerical parity, 12 real bug fixes, and FA4/NVFP4 full-kernelization bringing E2E inference from ~130 ms to 28.5 ms (median 27.7 ms, p95 28.5 ms on Thor).
No inference hyperparameters changed (4-step flow-matching, 252×252, T=50, bf16 math). All speedup comes from kernelization / quantization / graph fusion.
Precision vs HF eager (denormalized action space): cos 0.999933 / maxd 0.059.
Commits
fix(groot)perf(groot)docs(groot)fix(fa4)The 12 Bug Fixes
Upstream N1.6 frontend inherited openpi-family (Pi0/Pi0.5) vision/kernel assumptions that do not hold for GR00T N1.6's HF behaviour:
encode()vs Eagle chat templateimage_size=252default(C,ph,pw)vs HF NaFlex(ph,pw,C)(scale,shift)vs HF(shift,scale)Performance (all tiers ON, Thor SM110)
Bandwidth ceiling: Thor measured 252–255 GB/s (~93% of 273 GB/s spec). DiT 15.2 ms is weight-bandwidth-bound floor for this config.
Precision Tier Switches
FLASHRT_N16_DIT_FP4FLASHRT_N16_QWEN3_FP4FLASHRT_N16_SIGLIP_FP4FLASHRT_N16_FA4FLASHRT_N16_DIT_STEPSAll tiers independently fall back to bf16/torch when disabled or unavailable.
New Kernels
csrc/fused_fp4/silu_mul_fp4_sfa_bf16.{cu,cuh}— SiLU(gate)×up (bf16) → NVFP4+SFA, bit-exact vs torchcsrc/fused_fp4/dit_norm_fp4_sfa.cu— AdaLN / no-affine LN / weighted RMSNorm → NVFP4+SFA (bf16 variants)csrc/kernels/qk_norm_rope_rotate_half_bf16.{cu,cuh}— fused per-head RMSNorm + rotate-half RoPEcsrc/gemm/fp4/cutlass_fp4_gemm_bias_bf16_sm100.cu— bias / bias+res / bias+GELU+fp4out epiloguescsrc/quantize/quantize_fp4_sfa_bf16.cu— vectorized bf16 dynamic quantizeUpstream Bug Fixed (fix(fa4) commit)
fa4_backend.pyimported cutlass-dsl (which caches device arch assm_110a) BEFORE settingCUTE_DSL_ARCH=sm_101a. Combined with #164 (7fd75d20) extending the hd256 2CTA dispatch to SM110 without validation, this triggers an NVVM ICE on Thor. Fix: set env var before any cutlass import + restrict hd256 2CTA to SM100 + restore_fa4_trimmedlazy loader. Requiresnvidia-cutlass-dsl >= 4.5.Files Changed (17 files, +2588 / -134)
flash_rt/frontends/torch/groot_thor.py(core, +1461)flash_rt/hardware/thor/attn_backend_groot.pyflash_rt/models/groot/pipeline_thor.pyflash_rt/hardware/thor/fa4_backend.pycsrc/attention/flash_attn_4_src/flashrt_fa4/cute/interface_fwd_sm100.pycsrc/fused_fp4/,csrc/gemm/fp4/,csrc/quantize/,csrc/kernels/CMakeLists.txt,csrc/bindings.cpp,csrc/fp4_bindings.cppdocs/groot_n16_thor_sm110.md,docs/groot_transformers5_weight_corruption.md,docs/thor_gpu_idle_reset_workaround.mdTest Plan
Notes
pipeline_thor.pycontains pre-existing upstream TODO/FIXME in the legacy kernel path (parity=False); this PR does not modify that path.serving/groot_n16/) is maintained separately and not included in this PR.