[CuTe] Tune B200 forward scheduling - #2741
Draft
drisspg wants to merge 1 commit into
Draft
Conversation
drisspg
added a commit
that referenced
this pull request
Jul 27, 2026
Promote two SM100 BF16 output-only regions after explicit-config discovery, boundary, independent holdout, and config=None confirmation on an NVIDIA GB200 (152 SMs, CUTLASS DSL 4.6.0.dev0): policy cells geomean [paired-round 95% CI] weighted min max W/N/L D64 direct O 60 1.204x [1.203x, 1.206x] 1.191x 1.171x 1.264x 60/0/0 D128 1CTA 72 1.206x [1.203x, 1.208x] 1.158x 1.105x 1.560x 72/0/0 Each phase used seven alternating paired rounds with 31 fixed-pointer CUDA-graph samples per arm. D64 and D128 holdouts contained 20 and 24 cells, respectively; every retained discovery, boundary, and holdout cell won. The D64 causal control was rejected at 0.993x geomean with three paired intervals wholly below 1.0. A fresh old-explicit-baseline versus config=None confirmation covered 12 cells per rule (four discovery, four boundary, and four holdout): policy cells geomean [paired-round 95% CI] weighted min max W/N/L D64 direct O 12 1.208x [1.203x, 1.212x] 1.193x 1.173x 1.261x 12/0/0 D128 1CTA 12 1.211x [1.201x, 1.220x] 1.162x 1.112x 1.500x 12/0/0 The rules require exact SM100, dense noncausal equal-head BF16, at least 64 M-blocks, K >= 2048, packed Q > 256, MHA or packed GQA ratios 4/8/16, 128x128 tiles, and no SplitKV or optional feature paths. FP16, SM103, causal/local, LSE-producing and training/Flex, paged, sparse, QV/MLA, gather-KV, modifiers, sinks, and unmeasured layouts retain their previous defaults. The intervals propagate paired-round timing variance within the frozen cells; they measure replay repeatability rather than treating the curated grid as a random workload sample. Full raw samples, manifests, excluded controls, and validation logs are archived at evidence commit 0a9f4dc. stack-info: PR: #2741, branch: drisspg/stack/57
drisspg
force-pushed
the
drisspg/stack/57
branch
from
July 27, 2026 22:54
8f859c8 to
f771372
Compare
This was referenced Jul 27, 2026
drisspg
marked this pull request as draft
July 27, 2026 23:48
drisspg
force-pushed
the
drisspg/stack/57
branch
from
July 29, 2026 02:40
f771372 to
63fe68f
Compare
drisspg
force-pushed
the
drisspg/stack/57
branch
from
July 29, 2026 03:30
63fe68f to
6dcdbba
Compare
drisspg
force-pushed
the
drisspg/stack/57
branch
from
July 29, 2026 06:37
6dcdbba to
e70f08c
Compare
drisspg
force-pushed
the
drisspg/stack/57
branch
from
July 29, 2026 16:41
e70f08c to
9460987
Compare
drisspg
force-pushed
the
drisspg/stack/57
branch
from
July 29, 2026 21:11
9460987 to
5a0b1c7
Compare
drisspg
force-pushed
the
drisspg/stack/57
branch
from
July 30, 2026 00:55
5a0b1c7 to
9069402
Compare
drisspg
force-pushed
the
drisspg/stack/56
branch
from
July 30, 2026 00:55
ad2e8f9 to
c6b3100
Compare
drisspg
force-pushed
the
drisspg/stack/57
branch
from
July 30, 2026 00:58
9069402 to
c977015
Compare
drisspg
force-pushed
the
drisspg/stack/56
branch
2 times, most recently
from
July 30, 2026 01:16
32a98fd to
f58455d
Compare
drisspg
force-pushed
the
drisspg/stack/57
branch
from
July 30, 2026 01:16
c977015 to
2e26120
Compare
drisspg
force-pushed
the
drisspg/stack/57
branch
from
July 30, 2026 01:37
2e26120 to
e3dda2e
Compare
drisspg
force-pushed
the
drisspg/stack/57
branch
from
July 30, 2026 01:52
e3dda2e to
8c21f1a
Compare
drisspg
force-pushed
the
drisspg/stack/57
branch
from
July 30, 2026 02:04
8c21f1a to
b97daee
Compare
drisspg
force-pushed
the
drisspg/stack/57
branch
from
July 30, 2026 02:48
b97daee to
6ad9fc9
Compare
drisspg
force-pushed
the
drisspg/stack/57
branch
from
July 30, 2026 03:48
6ad9fc9 to
146cf56
Compare
drisspg
force-pushed
the
drisspg/stack/57
branch
from
July 30, 2026 05:16
146cf56 to
82595ec
Compare
Promote two SM100 BF16 output-only regions after explicit-config discovery, boundary, independent holdout, and config=None confirmation on an NVIDIA GB200 (152 SMs, CUTLASS DSL 4.6.0.dev0): policy cells geomean [paired-round 95% CI] weighted min max W/N/L D64 direct O 60 1.204x [1.203x, 1.206x] 1.191x 1.171x 1.264x 60/0/0 D128 1CTA 72 1.205x [1.203x, 1.208x] 1.158x 1.105x 1.560x 72/0/0 Each phase used seven alternating paired rounds with 31 fixed-pointer CUDA-graph samples per arm. D64 and D128 holdouts contained 20 and 24 cells, respectively; every retained discovery, boundary, and holdout cell won. The D64 causal control was rejected at 0.993x geomean with three paired intervals wholly below 1.0. A fresh old-explicit-baseline versus config=None confirmation covered 12 cells per rule (four discovery, four boundary, and four holdout): policy cells geomean [paired-round 95% CI] weighted min max W/N/L D64 direct O 12 1.208x [1.203x, 1.212x] 1.193x 1.173x 1.261x 12/0/0 D128 1CTA 12 1.211x [1.201x, 1.220x] 1.162x 1.112x 1.500x 12/0/0 The rules require exact SM100, dense noncausal equal-head BF16, at least 64 M-blocks, K >= 2048, packed Q > 256, MHA or packed GQA ratios 4/8/16, 128x128 tiles, and no SplitKV or optional feature paths. FP16, SM103, causal/local, LSE-producing and training/Flex, paged, sparse, QV/MLA, gather-KV, modifiers, sinks, and unmeasured layouts retain their previous defaults. The intervals propagate paired-round timing variance within the frozen cells; they measure replay repeatability rather than treating the curated grid as a random workload sample. Full raw samples, manifests, excluded controls, and validation logs are archived at evidence commit 0a9f4dc. stack-info: PR: #2741, branch: drisspg/stack/57
drisspg
force-pushed
the
drisspg/stack/57
branch
from
July 30, 2026 16:01
82595ec to
111c47b
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked PRs:
Human Note
Agent Note
Summary
This adds the same measured-policy treatment for B200. For long, sufficiently occupied BF16 output-only attention, D64 is better with direct output instead of TMA O, and D128 is better with 1CTA instead of 2CTA. The causal D64 control was basically flat/slightly worse, so causal stays on the old policy.
This is a Seaborn strip plot over the archived result rows: every circle is one actually timed workload cell, color is the campaign phase, the outlined diamond/error bar is the geomean and paired-round 95% timing interval, and
Xis the time-weighted result. The causal controls are real measurements, not an illustrative baseline.A fresh old-explicit-baseline versus
config=Nonererun then confirmed D64 at 1.208× and D128 at 1.211× geomean across 12 cells per rule, with 24/24 selector and direct-correctness checks passing.The D128 campaign now spells out the resolved register allocation alongside CTA topology: the measured 2CTA baseline is
(softmax=176, correction=88, other=72)and the measured 1CTA candidate is(192, 80, 48). These are not new measurements or a changed policy; they are exactly the values the old kernel-side table derived during the archived run, now made explicit so the checked campaign still reconstructs the same two binaries.The rule is exact SM100 BF16, dense noncausal, D=DV in {64,128}, MHA or packed GQA ratios 4/8/16, packed Q > 256, K≥2048, at least 64 M-blocks, 128×128 tiles, no LSE, and no SplitKV or optional feature path. SM103, FP16/FP8, causal/local, training/Flex, paged, sparse, QV/MLA, gather, modifier, sink, packed-MHA, and underfilled cases keep their previous defaults.
Reproduction and validation
Produced on the B200 machine with:
2.14.0.dev20260708+cu132, CUDA 13.2, CUTLASS DSL4.6.0.dev0, driver580.126.20.0a9f4dc.