These flags control how TT-Lang compiles operations. Pass them on the command line,
or print the list with --ttl-help:
python my_kernel.py --ttl-help
python my_kernel.py --no-ttl-maximize-dst| Flag | Default | Description |
|---|---|---|
--ttl-maximize-dst / --no-ttl-maximize-dst |
enabled | Partition compute iteration spaces into subblocks that maximize DST register utilization, and reorder tile operations within sync regions to group by kind. Disabling falls back to per-tile synchronization. |
--ttl-fpu-binary-ops / --no-ttl-fpu-binary-ops |
enabled | Allow FPU strategy selection for binary add, subtract, and multiply when their operands permit it. Disabling selects SFPU. |
--ttl-block-matmul / --no-ttl-block-matmul |
enabled | Emit matmul_block (processes the full tile block atomically) instead of per-tile matmul loops. Disabling this option is not yet supported. |
--ttl-subblock-sync / --no-ttl-subblock-sync |
disabled | Refine DFB reserve/push to per-subblock granularity, enabling pack_tile_block for contiguous subblocks. When disabled, user-placed reserve/push is preserved as written. |
--ttl-combine-pack-tiles / --no-ttl-combine-pack-tiles |
enabled | Combine consecutive pack_tile ops on the same DFB with contiguous DST and DFB indices into a single pack_tile_block call. |
--ttl-reduce-full-fp32 / --no-ttl-reduce-full-fp32 |
enabled | Prefer full-fp32 accumulation for reduce operations when supported by the target and the complete kernel configuration. |
--ttl-matmul-full-fp32 / --no-ttl-matmul-full-fp32 |
enabled | Prefer full-fp32 accumulation for matmul operations when supported by the target and the complete kernel configuration. |
--ttl-strict-f32-acc / --no-ttl-strict-f32-acc |
disabled | Error at compile time if a += accumulation loop's output block exceeds f32 DST capacity (4 tiles with double-buffering). When enabled, guarantees each accumulation step fits in a single DST section without subblocking. |
--ttl-compiler-dfbs / --no-ttl-compiler-dfbs |
enabled | Insert compiler-allocated intermediate DFBs when an operation requires DFB-attached inputs, fusion would read a source after its DFB is released, or a computed value is stored by operations in multiple MLIR basic blocks. When disabled, the compiler emits an error if materialization is required. |
--ttl-pipe-computed-addresses / --no-ttl-pipe-computed-addresses |
enabled | Use computed receiver DFB addresses for eligible PipeNet transfers. When disabled, transfers use receiver-published destination addresses; multicast still requires proven equal runtime receiver addresses. |
--ttl-pipe-capacity-sync / --no-ttl-pipe-capacity-sync |
enabled | Use capacity-counter synchronization when the receiver wait and pop execute on the receiver NOC thread and the computed-address transfer passes the DFB ownership and count proofs. When disabled, computed-address transfers use receiver-post synchronization. |
--ttl-pipe-global-semaphores-only / --no-ttl-pipe-global-semaphores-only |
disabled | Allocate all compiler-managed PipeNet synchronization counters in GlobalSemaphore storage, leaving local hardware semaphore ids available to the application. |
--ttl-pipe-batch-tiles N |
0 (auto) |
Limit the logical transfers in one PipeTransport group. 0 selects automatically and 1 disables grouping. |
--ttl-l1-budget N |
target-dependent | Override the L1 allocation budget used by DFB validation and PipeTransport selection. |
--ttl-reuse-user-dfbs / --no-ttl-reuse-user-dfbs |
enabled | Reuse physical DFB indices when concurrent-kernel liveness proves that compatible logical DFB lifetimes do not overlap. Disabling compacts provisional user indices without introducing new user-DFB sharing. |
--ttl-dfb-exact-coloring-search-limit N |
1000000 |
Examine at most N states during deterministic exact DFB allocation when order-dependent first-fit does not satisfy a DFB index or L1 limit. This bounds compile time; reaching the limit reports an inconclusive result, not a capacity proof. |
--ttl-specialize-cores / --no-ttl-specialize-cores |
disabled | Clone each TTKernel function whose control flow branches on a core coordinate once per launch coordinate (ttkernel-specialize-cores), replacing my_logical_x_ / my_logical_y_ with constants and tagging clones with ttl.core_coord for per-core dispatch. Opt-in. |
Besides the command line, the same flags can be set through three other mechanisms. When the same flag is set in multiple places, higher-priority sources win and unmentioned flags fall through from lower levels:
| Priority | Mechanism | Example |
|---|---|---|
| 1 (lowest) | CompilerOptions class defaults |
— |
| 2 | @ttl.operation decorator options= parameter |
@ttl.operation(grid=(2,2), options="--no-ttl-maximize-dst") |
| 3 | TTLANG_COMPILER_OPTIONS environment variable |
export TTLANG_COMPILER_OPTIONS="--no-ttl-fpu-binary-ops" |
| 4 (highest) | Command-line arguments (sys.argv) |
python my_kernel.py --no-ttl-maximize-dst |
The options keyword can also be passed at call time to override the decorator
for a single invocation:
my_kernel(tensor_a, tensor_b, options="--no-ttl-fpu-binary-ops")These parameters are set on the @ttl.operation decorator (not via command-line
flags) and control the TTNN compute kernel hardware configuration:
| Parameter | Type | Default | Description |
|---|---|---|---|
fp32_dest_acc_en |
bool or None |
None |
Constrain the Wormhole B0/Blackhole DST register-file element width: true selects 32-bit elements and false selects 16-bit elements. When None, resolve the width from target capabilities and tile-operation requirements. |
dst_full_sync_en |
bool or None |
None |
Enable full DST synchronization (single-buffering mode). Doubles DST capacity (32-bit elements: 8, 16-bit elements: 16) at the cost of a full sync between math and pack threads. |
math_fidelity |
str or None |
None |
Set the compute math fidelity to LoFi, HiFi2, HiFi3, or HiFi4. When None, retain the TTNN default. |
@ttl.operation(
grid=(2, 2),
fp32_dest_acc_en=True,
dst_full_sync_en=False,
math_fidelity="HiFi4",
)
def my_kernel(a, b): ...These environment variables control compilation behavior and diagnostic output. They are independent of the code generation flags above.
| Variable | Type | Default | Description |
|---|---|---|---|
TTLANG_COMPILE_ONLY |
0/1 |
0 |
Compile kernels but do not execute on hardware. |
TTLANG_INITIAL_MLIR |
file path | (unset) | Write the pre-optimization MLIR module to this file. |
TTLANG_FINAL_MLIR |
file path | (unset) | Write the post-optimization MLIR module to this file. |
TTLANG_VERBOSE_PASSES |
any value | (unset) | Print the IR after every pass in the pipeline. Output is very large; redirect to a file. |
TTLANG_DEBUG_LOCATIONS |
0/1 |
0 |
Include source locations in printed MLIR (locations are always tracked internally for error messages). |
TTLANG_VERBOSE_ERRORS |
0/1 |
0 |
Include raw MLIR diagnostics in error output. |
TTLANG_SIM_ONLY |
0/1 |
0 |
Force import ttl to skip loading the compiled MLIR extension. Used when running the simulator from a source tree without an installed tt-lang-sim wheel (which ships the same signal as a marker module). |
Profiling-related environment variables (TTLANG_AUTO_PROFILE,
TTLANG_PERF_DUMP, TTLANG_PERF_SERV, TTLANG_SIGNPOST_PROFILE,
TTLANG_PROFILE_CSV) are documented in the
Performance Tools reference.
The @ttl.operation decorator also accepts these parameters for operation structure
and layout:
| Parameter | Type | Default | Description |
|---|---|---|---|
grid |
tuple or Callable |
(required) | Compute grid dimensions, e.g., (2, 2) |
indexing_maps |
list[Callable] |
None |
Lambda functions for tile indexing |
iterator_types |
list[str] |
None |
"parallel" or "reduction" per dimension |
num_outs |
int |
1 |
Number of output tensor arguments |
memory_space |
str |
"L1" |
Memory space for dataflow buffers: "L1" or "DRAM" |
tiled |
bool |
True |
Use tiled tensor layout |
ttlang-opt is the standalone MLIR optimizer driver for the TTL dialect, used
primarily for compiler development and testing. It accepts all standard
mlir-opt flags (run ttlang-opt --help for the full list) plus the
TTL-specific passes and pipeline documented below.
The main compilation pipeline, equivalent to what the Python API runs internally.
ttlang-opt input.mlir -p 'ttl-to-ttkernel-pipeline{maximize-dst=true lower-to-emitc=true}'| Option | Type | Default | Description |
|---|---|---|---|
maximize-dst |
bool | true |
Enable DST maximization via subblock compute and scheduling. |
enable-fpu-binary-ops |
bool | true |
Allow FPU strategy selection for binary add/sub/mul. |
use-block-matmul |
bool | true |
Lower matmul to block-level hardware calls (matmul_block). |
subblock-sync |
bool | false |
Refine DFB reserve/push to per-subblock granularity. |
combine-pack-tiles |
bool | true |
Combine consecutive pack_tile ops into pack_tile_block. |
reduce-full-fp32 |
bool | true |
Prefer full-fp32 reduce accumulation when supported. |
matmul-full-fp32 |
bool | true |
Prefer full-fp32 matmul accumulation when supported. |
strict-f32-acc |
bool | false |
Error if a += accumulation loop's output block exceeds f32 DST capacity. |
compiler-dfbs |
bool | true |
Insert compiler-allocated intermediate DFBs for DFB-only operands, source-lifetime preservation, and computed values stored by operations in multiple MLIR basic blocks. Error if disabled and any operation requires one. |
pipe-computed-addresses |
bool | true |
Use computed receiver DFB addresses for eligible PipeNet transfers. When disabled, transfers use receiver-published destination addresses; multicast still requires proven equal runtime receiver addresses. |
pipe-capacity-sync |
bool | true |
Use capacity-counter synchronization when the receiver wait and pop execute on the receiver NOC thread and the computed-address transfer passes the DFB ownership and count proofs. When disabled, computed-address transfers use receiver-post synchronization. |
pipe-global-semaphores-only |
bool | false |
Allocate all compiler-managed PipeNet synchronization counters in GlobalSemaphore storage, leaving local hardware semaphore ids available to the application. |
pipe-batch-tiles |
int64_t | 0 (auto) |
Limit logical transfers per PipeTransport group. 0 selects automatically and 1 disables grouping. |
l1-budget-override |
uint32_t | 0 (target default) |
Override the L1 allocation budget used by DFB validation and PipeTransport selection. |
reuse-user-dfbs |
bool | true |
Reuse physical DFB indices for compatible logical DFBs with proven non-overlapping concurrent lifetimes. |
exact-coloring-search-limit |
uint64 | 1000000 |
Maximum states examined during deterministic exact DFB allocation before reporting an inconclusive result. |
specialize-cores |
bool | false |
Clone TTKernel functions that branch on a core coordinate once per launch coordinate (ttkernel-specialize-cores), then run canonicalize / cse. Maps from --ttl-specialize-cores. |
lower-to-emitc |
bool | false |
Run the TTKernel-to-EmitC backend (produces C++ source). |
The pipeline runs these passes in order:
ttl-materialize-loop-state-- replace ranked-tensor loop-carried values with compiler-created DFBsttl-insert-copy-wait-- insert missingttl.waitafterttl.copyops whose transfer handle has no wait userttl-annotate-l1-acc-loops-- detect+=accumulation loops and annotate for L1 packer accumulationttl-create-producer-compute-- create producerttl.computeoperations before intermediate materializationttl-insert-intermediate-dfbs-- materialize DFB-only operands, values that must be preserved before source release, and computed values stored by operations in multiple MLIR basic blocks; verify and error whencompiler-dfbs=falseconvert-ttl-to-compute-- lower TTL elementwise tensor ops tottl.computewith tile opsttl-insert-cb-sync-- insert missing DFB synchronizationttl-verify-pipenet-guards, thenttl-verify-pipenet-schedule-- verify PipeNet launch domains and event ordering while logical DFB identities remain distinct and before physical DFB allocationttl-form-pipe-transports-- group eligible repeated PipeNet transfers and select bounded receiver storagettl-coalesce-dfb-acquires-- coalesce compatible DFB acquiresttl-finalize-dfb-indices-- assign logical DFBs to physical indices, validate capacity, and emit runtime metadata;reuse-user-dfbscontrols user-DFB reuse andexact-coloring-search-limitbounds exhaustive fixed-limit and minimum physical-index-count queriesttl-set-compute-kernel-config-- select tile execution strategies and resolve kernel-wide DST and per-DFB unpack configurationttl-assign-dst-- DST register allocation (linear scan with copy insertion)ttl-subblock-compute-for-dst-- tilettl.computeinto DST-sized subblocks (only ifmaximize-dst=true); optionally refine reserve/push to per-subblock granularity (only ifsubblock-sync=true)ttl-lower-to-loops-- lowerttl.computetoscf.forloops; matmul computes are expanded inline viagenerateMatmulComputettl-schedule-operations-- reorder tile ops by dependency depth and kind (only ifmaximize-dst=true)ttl-annotate-cb-associations-- annotate block args with DFB indicesttl-verify-dfb-spsc-- verify per-node DFB producer/consumer uniqueness after finalizationttl-erase-pipenet-scopes-- remove verified PipeNet structural markersttl-validate-cb-budget-- verify static DFB storage fits the per-core L1 budgetconvert-ttl-to-ttkernel-- lower TTL DMA and PipeNet operations to TTKernel, selecting destination addressing, synchronization protocol, and synchronization-counter storagettkernel-insert-inits-- insert hardware init ops before compute opsttkernel-insert-l1-accumulation-- insertpack_reconfig_l1_accguards for+=and reduction loopsttkernel-combine-pack-tiles-- combine consecutivepack_tileintopack_tile_block(only ifcombine-pack-tiles=true)- Canonicalization and CSE cleanup
ttkernel-specialize-cores, thencanonicalize,cse-- per-core clone and const-fold of coordinate branches; tags clones withttl.core_coord(only ifspecialize-cores=true)- (if
lower-to-emitc=true)lower-affine,convert-ttkernel-to-emitc,emitc-form-expressions
Each pass can also be run standalone for testing. Only passes with configurable options are listed; the remaining passes have no options.
Insert compiler-allocated intermediate DFBs where tensor SSA values require concrete DFB storage.
| Option | Type | Default | Description |
|---|---|---|---|
enable |
bool | true |
Insert compiler-allocated DFBs. When false, emit an error if any operation requires one. |
ttlang-opt input.mlir -p 'func.func(ttl-insert-intermediate-dfbs{enable=false})'Assign physical indices to logical DFBs and emit the complete runtime allocation table.
| Option | Type | Default | Description |
|---|---|---|---|
reuse-user-dfbs |
bool | true |
Reuse a physical index when concurrent-kernel liveness proves that two compatible logical DFB lifetimes cannot overlap. When false, compact provisional user indices without introducing new user-DFB sharing and apply the same lifetime proof only to compiler-created DFBs. |
exact-coloring-search-limit |
uint64 | 1000000 |
Examine at most this many states during deterministic exact DFB allocation. Exhaustive search is used only when order-dependent first-fit prevents acceptance. Reaching the limit fails with an inconclusive-search diagnostic rather than a false capacity diagnostic. |
ttlang-opt input.mlir -p 'builtin.module(ttl-finalize-dfb-indices{reuse-user-dfbs=false exact-coloring-search-limit=1000000})'Resolve tile execution strategies and shared compute-kernel configuration. See Compute Kernel Configuration for the algorithm and invariants.
| Option | Type | Default | Description |
|---|---|---|---|
fp32-dest-acc-en |
string | auto |
Select 32-bit destination elements through the Wormhole B0/Blackhole fp32_dest_acc_en setting: auto, enabled, or disabled. |
dst-full-sync-en |
string | auto |
Select full DST synchronization: auto, enabled, or disabled. |
reduce-full-fp32 |
bool | true |
Prefer full-fp32 reduce accumulation when supported. |
matmul-full-fp32 |
bool | true |
Prefer full-fp32 matmul accumulation when supported. |
enable-fpu-binary-ops |
bool | true |
Allow eligible add/sub/mul operations to select FPU. |
ttlang-opt input.mlir -p 'func.func(ttl-set-compute-kernel-config{fp32-dest-acc-en=enabled})'DST register allocator using linear scan allocation with in-place operation merging.
| Option | Type | Default | Description |
|---|---|---|---|
dst-capacity |
uint32_t | 0 (auto) |
Override DST register capacity. Auto-computed from fp32_dest_acc_en and dst_full_sync_en by default. Single-buffering (dst_full_sync_en=true): 32-bit elements=8, 16-bit elements=16. Double-buffering (default): 32-bit elements=4, 16-bit elements=8. |
separate-output-region |
bool | false |
Allocate outputs in a separate DST region (needed for reductions and some loop optimizations). |
ttlang-opt input.mlir -p 'func.func(ttl-assign-dst{dst-capacity=16})'Partition ttl.compute into DST-sized subblocks.
| Option | Type | Default | Description |
|---|---|---|---|
subblock-sync |
bool | false |
Refine DFB reserve/push to per-subblock granularity, enabling pack_tile_block for contiguous subblocks. When disabled, user-placed reserve/push is preserved. |
strict-f32-acc |
bool | false |
Error if a += accumulation loop with non-f32 output requires subblocking. Subblocking reduces accumulation precision because bf16 L1 intermediates truncate f32 DST values. |
ttlang-opt input.mlir -p 'func.func(ttl-subblock-compute-for-dst{subblock-sync=true})'Group eligible repeated PipeNet transfers and select bounded receiver storage. Later PipeTransport planning replaces proven-private grouped DFB lifecycles with transport-owned scratch; scalar residuals retain the original lifecycle. Selection accounts for DFB allocation, a conservative receiver-published address table, and transport scratch.
| Option | Type | Default | Description |
|---|---|---|---|
group-size |
int64_t | 0 (auto) |
Limit logical transfers per group. 0 selects automatically and 1 disables grouping. |
l1-budget-override |
uint32_t | 0 (target default) |
Override the combined DFB and pipe scratch budget used during grouping selection. |
ttlang-opt input.mlir --ttl-form-pipe-transports='group-size=8'Lower TTL data movement and PipeNet operations to TTKernel.
| Option | Type | Default | Description |
|---|---|---|---|
reduce-full-fp32 |
bool | true |
Enable FP32 accumulation for reduce operations. |
pipe-computed-addresses |
bool | true |
Use computed receiver DFB addresses for eligible PipeNet transfers. When false, transfers use receiver-published destination addresses; multicast still requires proven equal runtime receiver addresses. |
pipe-capacity-sync |
bool | true |
Use capacity-counter synchronization when the receiver wait and pop execute on the receiver NOC thread and the computed-address transfer passes the DFB ownership and count proofs. When false, computed-address transfers use receiver-post synchronization. |
pipe-global-semaphores-only |
bool | false |
Allocate all compiler-managed PipeNet synchronization counters in GlobalSemaphore storage. |
ttlang-opt input.mlir -p 'builtin.module(convert-ttl-to-ttkernel{pipe-computed-addresses=true pipe-capacity-sync=false pipe-global-semaphores-only=true})'Analyze dataflow buffer producer/consumer relationships and dump the flow graph.
| Option | Type | Default | Description |
|---|---|---|---|
output |
string | "" |
Path to write JSON output. Empty string prints to stderr only. |
ttlang-opt input.mlir -p 'ttl-dump-cb-flow-graph{output="/tmp/cb_graph.json"}'Clone TTKernel functions that branch on a core coordinate once per launch
coordinate. Requires a module-level ttl.launch_grid attribute (an i64 array
of length 2 with positive entries). Missing or malformed ttl.launch_grid is
a hard error. A valid single-core grid (product <= 1) skips specialization.
Only scf.if conditions derived from ttkernel.my_logical_x_ /
ttkernel.my_logical_y_ trigger cloning. Functions with symbol uses (for
example func.call targets) are left unspecialized with a warning so erasing
the original does not leave dangling SymbolRefAttrs; unrelated functions in
the module are still specialized. Each clone replaces coordinate reads with
arith.constants and is tagged with ttl.core_coord for runtime dispatch.
Downstream canonicalize / cse fold the now-constant branches.
This pass is off by default. Enable it through the pipeline option
specialize-cores (Python: --ttl-specialize-cores):
ttlang-opt input.mlir -p 'ttl-to-ttkernel-pipeline{specialize-cores=true lower-to-emitc=true}'
# Or stand-alone:
ttlang-opt input.mlir -p 'builtin.module(ttkernel-specialize-cores,canonicalize,cse)'