AIE kernel optimization¶
This skill makes one compiled kernel faster. It assumes the kernel is
already correct on hardware (aie-hw-bringup) and worth the work
(aie-dataflow-opt ranks kernels by ablation).
Evidence rule. The static report tells you what to try and predicts the result. Only two arms measured back to back on the NPU, as traced cycles per call of the kernel itself, show a speedup. Only a raw output-word diff shows bit-exactness. Without a device, report a static change as "candidate, HW unconfirmed", never as faster: an II33 → 18 change measured 594 → 594 (it edited a symbol no factory calls), and an II37 → 31 change measured slower.
Setup¶
source /opt/xilinx/xrt/setup.sh # XRT first, or pyxrt is missing
source <venv>/bin/activate
source utils/env_setup.sh install # from the repository root (docs/Building.md)
W=$(mktemp -d); K=<factory>; CASE=<case name from kernel_cases.py>; CPUS=<fixed host CPU list>
Remarks takes --target aie2p for npu2 (Strix, Krackan) or --target aie2
for npu1 (Phoenix); xrt-smi examine shows which device you have. What an
intrinsic lowers to is in third_party/aie_api/include/aie_api/detail/<arch>/.
Workflow¶
1. Resolve what runs¶
grep -n "def $K\b" -A40 python/iron/kernels/*.py # source, -D flags, extern "C" symbol, trace=
grep -n "Case(\"$K" test/python/npu/kernel_cases.py # cases, calls=, kwargs
python -m aie.utils.compile.remarks --target aie2p --only "^$K" --out $W/rows.json --meta $W/meta.json
- Remarks prints
[OK] <build>: <symbol> from <source>per build. Follow that symbol to the function it calls and edit only that. A.ccoften holds siblings no factory selects (in-place,_scalar, another dtype). - Remarks builds each factory at its defaults and
.dtypes; if a case sets another-Dflag, the rows describe the default build. - Read the contract's
trace=. OnlyTrace.whole_call()gets a cycles row.Trace.none(reason)andTrace.partial(reason)leave you with wall clock.
2. Gate, and prove the gate¶
pytest test/python/test_kernel_contracts.py -k "$K" -q # host only
pytest test/python/npu/test_kernels_e2e.py -m extensive -k "$CASE" --seeds 3
-m extensiveruns every case, every data case (random plus the contract's edge cases) and every seed.-kis a substring match (-k addalso runsmul_add); check the selection with--collect-only -q.- Mutation-prove it. Break the kernel on purpose (drop a term, skip the tail, swap two loads), watch the gate fail, revert. Report a mutation that survives, with the reason; don't swap in a louder one.
- Derive every tolerance from the arithmetic, over the longest case (one kernel's error was 0.07 at 4 calls and 0.25 at 256).
- The reference must not replay the kernel. Write a single-pass fp32 reference, and judge against the mathematics (see the exp2 and tanh traps below).
- Keep outputs poisoned (
kd.upload(..., poison=True)) and guarded (kd.design(..., guard=True)), so an unwritten output or a write past a tile's end fails. Gate every entry point alone, and the composition once.
3. Build the base arm¶
An arm is a directory holding aie_kernels/ and aie_runtime_lib/.
Remarks and the JIT compile from MLIR_AIE_KERNEL_SOURCES=<arm> when it is
set, and both take --baseline-sources <arm> for the comparison.
BASE=$W/base; mkdir -p $BASE
git archive HEAD aie_kernels aie_runtime_lib | tar -x -C $BASE
# Shared tree, or your file includes a sibling (mha.cc includes mm, softmax, zero):
cp -r aie_kernels aie_runtime_lib $BASE/ && git show <rev>:<file> > $BASE/<file>
diff -rq aie_kernels $BASE/aie_kernels # must list only your files
The factories, cases and -D flags come from the installed Python in both
arms, so snapshot Python-side parameters (stack sizes, geometry tables) you
change along with the arm.
4. Read the static report¶
python -m aie.utils.compile.remarks --target aie2p --only "^$K" \
--out $W/rows.json --meta $W/meta.json --keep $W/objs --baseline-sources $BASE
- Per build:
[OK], then anydropped pragma:,calls the runtime library: <symbols>, and astack:warning over the contract's budget. It exits 3 if a build fails to compile; that is a finding. With--baseline-sourcesit ends withN rows differand onename: before -> afterline per row: that is the static A/B. - Rows:
unpipelined_loops,non_zol_loops,missing_bank_loads,pass_failed_warnings,pm_bytes,libcalls,stack_bytes, and per looploop/<fn>/<bb>/IIandnot_zol. They cover only functions the entry symbol reaches. - Meta per loop:
ii,ns(stage count),pipelined,zol,bundle_count,byte_count,file,line; per buildpass_failedandschedule_notes(MII,SwpMaxMii). Meta covers this tree only; for the base arm's loops run again withMLIR_AIE_KERNEL_SOURCES=$BASE. - Spills:
[sp, #inside a loop body in the kept object. - Size: use
pm_bytes/pm_bytes_by_function, not an object's.text, which counts siblings the link drops. - A
noinlinebody reports its caller's II; count its bundles yourself.
5. Diagnose, then change one thing¶
| Report or object shows | Lever |
|---|---|
libcalls names a helper called in a loop |
L02; L11 for __divsi3 |
vector<float> multiply, min or max in a loop |
L03; L04 on AIE2P |
Array of accumulators or vectors indexed by a loop counter; [sp traffic |
L01 |
| The loop you care about isn't innermost or single-block; unpipelined parent | L10 |
| Loop body branches on its counter | L13 |
Address arithmetic in the body; no __restrict |
L06 |
Pipelined, many empty bundles; byte_count flat from ×1 to ×4 |
L07 |
| Stepping 16 lanes on 8/16-bit, or on bf16 on AIE2P | L08 |
| One long mac chain, or broadcasts spilled per block | L09 |
Pipelined, ns 3, one long latency chain with no dominant step |
L17 |
Several mmul<8,8,8> accumulators for Y += S·V; high II, big frame |
L18 |
A reduce_add/reduce_max per row that the row loop never overlaps |
L20 |
bfp16 stream state (sfl/sfh, FIFO) spilled around vst.push/vldb.pop |
L19 |
(tanh+1)/2, or separate constant multiplies |
L05 |
lda.s8/st.s8 byte loop |
L12 |
| Scalar int8 requantize tail | L14 |
| Scalar gather building the mmul A operand | L15 |
| Body is only a rearrangement | L16 |
| II already at a resource or latency bound | stop (§When to stop) |
- One change per candidate. The lever's Check must move in the
--baseline-sourcesdiff; if no static metric moved, revert. - Before an unroll, screen it: two scratch arms that differ only in the
pragma, remarks on
x4with--baseline-sources $W/x1. Loopbyte_countflat at ×4 means latency-bound: go. Roughly ×4, a new[sp, #, or morestack_bytes: stop. Try 2 and 8 when 4 is borderline. - If the change creates a path no case reaches (an unroll remainder, a
non-square shape, an odd block count), add a
Casetotest/python/npu/kernel_cases.pyand its name totest/python/npu/perf_series.txt(test_perf_series_names.pychecks). - A change that hides work from LLVM (an opaque pointer bump,
volatile, a LICM blocker) to get a lower II has measured slower (prefillfv: static II37 → 31, HW 1374-1576 against 1265). - Write the per-call prediction before measuring:
bundles_outside_loop + 5 + (trips - 1) × II, per loop, trips from the loop-count setup in the object. Compare II per unit of work (per row, value or tile), not per loop: a raw II that rises can still win. Never turn an II ratio into a speedup (prefillfv's said -70%; HW gave -35%). - Re-run
test_kernel_contracts.pyand the gate.
6. Measure both arms back to back¶
taskset -c $CPUS pytest test/python/npu/test_kernels_perf.py -m perf -k "[$CASE] or [<unchanged case>]" \
--baseline-sources $BASE --perf-out $W/perf.json --perf-meta $W/meta.json
- Each case runs from this tree, then from
$BASE, with the same inputs. The summary printscycles base -> cur,npu_us min base -> curandsameorN differraw output words per case;--perf-metaholds the same underbaseline.cases.<case>. -k "[$CASE]", brackets included, selects exactly one case.-krejects=; select a prefix and check with--collect-only -q.- To screen a candidate kept in its own copy, set
MLIR_AIE_KERNEL_SOURCES=$W/cand-<name>; the JIT cache key includes it, so no cache wipe is needed.
Verdict:
| You see | Verdict |
|---|---|
| Candidate's max below base's min; the unchanged kernel reproduces to the cycle | confirmed |
| Ranges overlap, or the delta sits inside an unchanged kernel's spread | no change: revert |
| Candidate higher | rejected: revert |
| The unchanged kernel moved | void: the arms differ in more than your change; fix and rerun |
Check the result against the prediction (zero predicted ~75 / ~130,
measured 78 / 134); if they disagree, find out which is wrong first.
7. Exactness and record¶
- "Bit-identical" means every case reads
samein step 6's run. A tolerance pass isn't proof:fused_mmpassed 27 of 27 cases with a dropped term, and a mutation that changed 45,159 words passed rtol 0.04. Mutation-prove the diff too: a broken candidate must show words that differ. If the change reorders accumulation, say so and give the max |Δ|. - One commit per kernel and change, with base → after cycles per case, n, the arm sources and the gate line. List each rejected variant with the number that killed it. NO-CHANGE with a measured bound is a valid result.
- Record provenance with every number: the row's
extra(commit, Peano,kernelsdigest),git status --shortover the include closure, and what$BASEcame from. When two numbers conflict, compare digests first. - Several agents on one NPU: wrap every hardware command in
flockon one shared lock file, and build base arms by copying the live tree. - Keep every
extern "C"name, signature and buffer layout. Leave*_scalarvariants alone.
Reading hardware rows¶
<case>/cycles is the min of the kernel's own intervals, split from its
initializers' by position (kd.cycles_per_call). Its range gives
median X max Y n=N, init[i] min M per initializer, and truncated.
| You see | Do |
|---|---|
No <case>/cycles row |
The contract is none/partial (in the library, only set_rounding). Use wall clock and say so, or move one marker pair to bracket the whole call and declare Trace.whole_call() |
expected E trace intervals, got M |
A marker the contract doesn't declare. pytest test/python/test_kernel_trace_markers.py names the build |
truncated, or none of the kernel's |
Compare n= with calls; flag n < calls/2. Run the marker audit |
| A row moved but its file didn't | Something in its include closure changed: git diff <rev> -- <closure>. A 33% "matmul" gain was all zero.cc |
| A stream you traced yourself, K kernels per call | Population j is intervals[j::K]. Never quote min or median over the mixed stream |
- A cycle row names the design, not the kernel. Before crediting a delta,
confirm the
[OK]line names the entry you changed. core_fraction = cycles × calls ÷ 1.76e9 ÷ npu_seconds(1.76 GHz measured on AIE2P). Below 0.5 the case is dispatch-bound: a 16-call dispatch costs 60-90 µs whatever the core does. Judge the kernel on cycles; if production is dispatch-bound too, hand off toaie-dataflow-opt.- Wall clock, only when the kernel can't be traced: quote the pinned
min (
npu_us_min), never the mean (it showed 7.5% between byte-identical builds). For the per-call cost, difference two cases that differ only incalls:(t20 − t4) / 16. - Ablation prices a region: replace it with a cheap wrong computation and trace. The bf16 GEMV went 1154 → 634, pricing its reduction at 520. A price is where the time goes, not a bound: six local tweaks lost, then restructuring the reduction (L20) won 1154 → 788.
Levers that measured faster¶
Each gives the change, what must move in the static report where it isn't obvious, and one hardware result (cycles per call unless noted; "model" means wall clock in a model outside this repository).
L01 Fully unroll loops that index register arrays. A short loop
indexes acc[i] or calls insert/extract with its counter; [sp around
the accumulators. AIE_LOOP_UNROLL_FULL makes each index constant. Check:
stack traffic gone, the enclosing loop gets an II. conv2dk1_i8 7623 → 504.
L02 No libcalls in hot loops. The libcalls row names a helper (see
the traps). Use aie::max/min for compares, aie::inv times a multiply
for divides, aie::to_float for int → float, integers for index math;
reduce a per-chunk max element-wise and horizontally once at the end.
Check: helper gone, loop gets an II. rms_norm 1597 → 438.
L03 Avoid the f32 vector multiply (AIE2P). Skip multiplies by a known 1
or 0. If one operand is exact in bf16 (a tanh result, a chosen constant),
split only the f32 operand into three bf16 limbs (residuals via
aie::msc; two limbs change bits) and mac each against it. Clamp the
rounded bf16 result against bounds rounded the same way instead of an f32
min/max. Matmul silu epilogue 8098 → 3213, bit-identical.
L04 Run emulated f32 chains 32 lanes wide (AIE2P). A 16-lane emulated
f32 multiply pays for 32 lanes. Step 32 lanes; replace a vector
to_fixed/to_float floor with the magic-number floor x + 1.5*2^23.
Check: bundles per element drop. bf16_exp 32482 → 15425 (with other
changes).
L05 Fold constants into one mac. Write (tanh+1)/2 as one mac into an
accumulator preloaded with 0.5; fold constant scales together
(power-of-two scaling is bit-exact). sigmoid 498 → 118.
L06 Walking __restrict cursors. Peano doesn't strength-reduce
base[i*stride]. Advance a T *__restrict cursor, mark non-aliasing
in/out pointers __restrict (not for in-place kernels), and use
add_2d_byte/add_3d_byte for multi-dimensional walks
(aie_kernels/quant/q4nx_dequant.cc). Check: II and body [sp drop.
zero<float,4096> 519 → 262.
L07 AIE_LOOP_UNROLL(4) on latency-bound bodies, after the unroll
screen passes. Not on load-, resource- or register-bound bodies. Check: II
per element drops, no new [sp. leaky_relu 298 → 86, add 390 → 150.
UNROLL(2) on fused_mm's epilogue chunk loop: gelu 2582 → 1919.
L08 Full register width. Step 64 lanes on 8/16-bit data and
AIE_BF16_LANES (aie_kernels/aie_arch.h) on bf16 arithmetic. Check: trip
count halves at a similar II. rope 1932 → 298 (with L06).
L09 Split dependency chains. Two independent accumulators added at the
end, or one broadcast feeding several row blocks; stop before spills.
int16 mv/32x32 291 → 127.
L10 Make the hot loop innermost and single-block, the only loops Peano
pipelines: fully unroll a short K reduction into its parent, fold nested
tile loops into one counter, split a per-tile if into straight loops. Opt
in per dtype. int8 mm 993 → 737. UNROLL(2) on an unpipelined outer loop
overlaps two trips' loads and stores: fused_mm 1084.5 → 991.5.
L11 Unsigned counted trip. Compute the trip count up front as unsigned;
a signed divide or shift by 2^k is __divsi3. AIE_LOOP_MIN_ITERATION_COUNT
can help but can cost the zero-overhead loop, so re-check non_zol_loops.
axpy 317 → 178.
L12 Wide stores, never byte loops. Peano won't pipeline an
lda.s8/st.s8 loop. Copy with uint64_t/uint32_t or 32 B vector
stores, both ends aligned. bfp16 shuffle 10.4x per call.
L13 UNROLL_FULL, not RANGE, on a loop that switches on its counter.
AIE_LOOP_RANGE is only a trip-count hint and leaves the branch. 5.31 →
2.84 ms (model).
L14 Vector int8 epilogue (AIE2P). Add the bias on an int32 vector, then
acc.to_vector<int8>(shift) under aie::rounding_mode::conv_even, which
is bit-exact with scalar banker's-rounding SRS. -42% on one block (model).
L15 Producer writes mmul-A order (AIE2P). aie::concat won't join
vectors under 128 bits, so a strided A operand becomes a scalar byte copy.
When you own both ends, have the producer store to_vector<int8>(shift)
(already mmul-A order) and the consumer do one aligned load, or
aie::shuffle_down(aie::concat(lo, hi), shift) over two aligned blocks.
-23.6% on one block (model).
L16 Pure rearrangement → DMA. A body that is only a deinterleave or
transpose belongs in a memtile dims_to_stream transform (element ≥ 512 B;
int8 vector loads need a 32 B-aligned start). Hand off to
aie-dataflow-opt. programming_examples/basic/transposes/transposes.py
shows --strategy dma and --strategy combined.
L17 Raise the pipeliner stage cap on one long chain. The loop is
pipelined at ns 3 and no single step dominates. Add
["-mllvm", "--aie-pipeliner-max-stagecount=5"] to that factory's
compile_flags only (as python/iron/kernels/quant.py does). Check: ns
rises; the final II may barely move, so measure anyway. q4nx_dequant
2245 → 2115.
L18 Two 8x8 tiles on one 64-lane accumulator (AIE2P, bf16). For Y +=
S·V, keep two neighbouring output tiles on one accfloat accumulator so
one vmac.f advances both; build the S operand once per row block; load
the next pair's y before this pair's store (clamp the last pointer to
itself). Prefill fv 2034 →
1265, bit-identical.
L19 One bfp16 stream per operand (AIE2P). Output streams share one sf
register and inputs two lf FIFOs, so extra live streams spill every step.
Write output through one contiguous stream; hop one A and one B stream
between rows with pop_seek. mm_bfp 4243 → 951; q4nx_dequant 2.10x.
L20 Batch horizontal reductions. Mac four rows per load of b, pack the
four accumulators so one interleave_unzip + add tree folds all rows,
and finish group g's tree after group g+1's macs issue; keep the single-row
code for the tail. Rows per group are bounded by the stack (8 overflowed
1 KB in mv). bf16 mv/32x256 1154 → 788; mha partial_softmax 8171 →
1696, both bit-identical.
AIE2 (npu1). Gate AIE2-only code with AIE_TUNED_AIE2 or a capability
from aie_kernels/aie_arch.h; the other target's .text must not change.
- The AIE2 pipeliner needs
__restrictto overlap a streaming loop; library kernels spell itAIE2_RESTRICTso AIE2P code stays as tuned. - An unrolled body above MII 27 gets no overlap. One chain under
AIE_LOOP_NO_UNROLLwith__restrictpipelined at II1:add342 → 78. AIE_LOOP_MIN_ITERATION_COUNT(n)on a runtime-count loop let it overlap, with a plain loop kept for shorter rows:axpy269 → 87.- A LUT loop runs in series when its reads are ordered against its stores:
rotate it one vector ahead, or use
lut_map_bf16. - There is no native bf16
sliding_mul; usemac_elem_16_2. - Replace
aie::transposein a loop with two-registershufflestages.
Traps¶
Code that compiles cleanly and then does nothing, runs slowly or produces wrong data. None of these is a compiler bug. A Peano crash or miscompile is: reduce it to a repro, file it against llvm-aie, and name the issue next to any workaround rather than recording it here.
- Libcalls (AIE2 and AIE2P). The scalar unit has no float multiplier
or divider:
float * float(__mulsf3),float / float(__divsf3),(float)int(__floatsisf), float compares (__ltsf2) and anydoubleare runtime calls, as are 64-bit multiply and divide and any integer divide or modulo that isn't by a power-of-two constant. A libcall in a loop also blocks pipelining. Float add/sub and bf16 ↔ float casts are native. The remarks report lists every libcall a kernel makes. - No f32 vector multiplier.
aie::mul/maconvector<float,N>is a bf16 emulation (224 B of code where one bf16 mac is 4 B). - Pragmas. Under Peano
AIE_PREPARE_FOR_PIPELINING,AIE_LOOP_FLATTENand the other Chess hints inaie_kernels/aie_kernel_utils.hexpand to nothing, andAIE_PREPARE_FOR_POSTPIPELININGdisables pipelining. The ones that act areAIE_LOOP_UNROLL(n)/_FULL/NO_UNROLLand the trip-count hints (AIE_LOOP_RANGEis only a hint). Put each immediately before thefor;pass_failedlists any the compiler dropped. - Stack overflow is silent. It corrupts the neighbouring buffer:
suspect it when errors are small, scattered and row-local. The IRON
Worker default is 1024 B; aiecc errors with
this core needs N byteswhen the declaration is short, and the remarksstack_bytesrow warns over the contract'sstack_bytes(it leaves out libcall frames). Anything that grows the frame (unroll, accumulators, markers) needs a new declaration in the same change. to_vector<int32>is a raw accumulator dump;to_vector<int8>(shift)applies the row-major permutation. Indexing one as the other got 61454 of 65536 values wrong.- Software pipelining stops at MII 27 (
SwpMaxMiiinschedule_notes); above it only the postpipeliner runs. A loop over the cap can still win if it does more work per trip. aie::exp2(AIE2P) is an interpolant that overshoots trueexp2by up to 6.15%. Keepnp.exp2in references with a derived envelope (test/python/npu/test_mha_e2e.py).- Native
vtanhreturns x for |x| ≤ 0.5, and tanh wants its f32 argument (a bf16-rounded one raised silu's error 1.35x). The native build is latency-bound, the LUT build (-DACTIVATIONS_TANH_LUT=1) load-bound: gate unrolls onACTIVATIONS_NATIVE_TANHand measure both.
Changes that measured worse¶
- Fewer lanes, 8 rows per group, or hoisting operands in bf16
mv: slower or spilled (register pressure). UNROLL_FULLonfused_mm's i loop: accumulators spilled, k_step 216 → 287, and three builds failed their stack check.mmul<8,8,8>fed by a scalar A gather in a 1x1 int8 conv: 3.34 → 3.58 ms (model).- An
ifor ternary inside a mac loop: 7% slower (model). - Unrolling a load-, resource- or register-bound body (tanh stayed at II4;
swiglu ×2 doubled its stack), and
AIE_LOOP_MIN_ITERATION_COUNTon AIE2P loops that then lost their zero-overhead loop. The compiler reports these before hardware does; trust it.
When to stop¶
Report NO-CHANGE with the bound when the II equals a resource or latency
bound. mm_bfp_mixed's k loop sits at II6, the vmac.f acc → acc latency,
with no fifth accumulator register for another chain. bf16 mm's k loop
needs 34 mv-slot ops per step, so II35 is 97% slot efficiency. Before
calling a loop port-bound, count slot users per iteration in the object; if
the II sits above that count, the bound is something else.
Background¶
programming_guide/section-4/section-4c/README.md (how Peano schedules a
loop, the pragma reference) and section-4d/README.md (this workflow for
humans, with a worked add example).