Python Kernel Library¶
Pre-built AIE kernel wrappers for common operations. These provide ready-to-use
Worker-compatible callables backed by optimized native AIE code. For the C++
kernel sources these wrap, see C++ AIE kernels.
Element-wise operations¶
Element-wise kernel factories: passthrough, scale, add, mul, relu.
add_ref ¶
mul_ref ¶
scale_ref ¶
Numpy reference for scale: x * factor[0] in int64.
factor is the 1-element int32 buffer the kernel reads the multiplier
from; the harness casts the int64 product back to x.dtype, wrapping
on overflow like the C++ store does.
Source code in python/iron/kernels/eltwise.py
passthrough ¶
Element-wise passthrough kernel: copies input tile to output tile.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tile_size
|
int
|
Number of elements per tile, a positive whole number of 64-byte vectors. Compiled into the kernel's loop bound. |
4096
|
dtype
|
type
|
Element data type ( |
int32
|
Returns:
| Type | Description |
|---|---|
ExternalFunction
|
ExternalFunction configured for |
Raises:
| Type | Description |
|---|---|
ValueError
|
When |
Source code in python/iron/kernels/eltwise.py
scale ¶
scale(
tile_size: int = 1024,
dtype: type = int32,
vectorized: bool = True,
use_chess: bool = False,
) -> ExternalFunction
Scalar-multiply kernel: multiplies each element of an input tile by a factor.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tile_size
|
int
|
Number of elements per tile. |
1024
|
dtype
|
type
|
Element data type. Must be |
int32
|
vectorized
|
bool
|
If |
True
|
use_chess
|
bool
|
When |
False
|
Returns:
| Type | Description |
|---|---|
ExternalFunction
|
ExternalFunction configured for the scale kernel. |
Raises:
| Type | Description |
|---|---|
ValueError
|
When |
Source code in python/iron/kernels/eltwise.py
add ¶
Element-wise bf16 addition (tile_size must be 1024, hard-coded in C++).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tile_size
|
int
|
Elements per tile (must be 1024). |
1024
|
dtype
|
type
|
Element data type (only |
bfloat16
|
vectorized
|
bool
|
If |
True
|
Returns:
| Type | Description |
|---|---|
ExternalFunction
|
ExternalFunction for eltwise_add_bf16. |
Raises:
| Type | Description |
|---|---|
ValueError
|
When |
Source code in python/iron/kernels/eltwise.py
mul ¶
Element-wise bf16 multiplication (tile_size must be 1024, hard-coded in C++).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tile_size
|
int
|
Elements per tile (must be 1024). |
1024
|
dtype
|
type
|
Element data type (only |
bfloat16
|
vectorized
|
bool
|
If |
True
|
Returns:
| Type | Description |
|---|---|
ExternalFunction
|
ExternalFunction for eltwise_mul_bf16. |
Raises:
| Type | Description |
|---|---|
ValueError
|
When |
Source code in python/iron/kernels/eltwise.py
mul_add ¶
c = a * b or c = a + b on bf16 tiles, chosen per call by is_mul.
One kernel for a two-phase runtime-parameter design:
programming_examples/ml/scale_shift computes A * B with is_mul = 1
and then + C with is_mul = 0 on the same workers.
aie_kernels/eltwise/scale_shift.cc fixes the tile at 1024 elements.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tile_size
|
int
|
Elements per tile (must be 1024). |
1024
|
Raises:
| Type | Description |
|---|---|
ValueError
|
When |
Source code in python/iron/kernels/eltwise.py
mul_add_ref ¶
Numpy reference for mul_add: a * b if is_mul else a + b.
Source code in python/iron/kernels/eltwise.py
relu ¶
Element-wise bf16 ReLU (tile_size must be 1024, hard-coded in C++).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tile_size
|
int
|
Elements per tile (must be 1024). |
1024
|
Returns:
| Type | Description |
|---|---|
ExternalFunction
|
ExternalFunction for bf16_relu. |
Raises:
| Type | Description |
|---|---|
ValueError
|
When |
Source code in python/iron/kernels/eltwise.py
add_sized ¶
Element-wise bf16 addition, with a compiled-in element count.
Runtime-size sibling of add; design passes
(a, b, c, size) for ABI compatibility. Scalar tails are supported.
Source code in python/iron/kernels/eltwise.py
mul_sized ¶
Element-wise bf16 multiplication, with a compiled-in element count.
Runtime-size sibling of mul; design passes
(a, b, c, size) for ABI compatibility. Scalar tails are supported.
Source code in python/iron/kernels/eltwise.py
relu_sized ¶
Element-wise bf16 ReLU, with a compiled-in element count.
Runtime-size sibling of relu; design passes
(in, out, size) for ABI compatibility. Not LUT-based. Positive
multiples of 32 elements are supported.
Source code in python/iron/kernels/eltwise.py
Data movement¶
Data-movement / conversion kernel factories: axpy, convert_copy, expand, transpose.
Each wraps one source under aie_kernels/datamovement/ — plain aie_api
vector code with no LUT dependency. convert_copy binds
cast_f32_bf16.cc, the f32->bf16 cast with host-matching conv_even
rounding.
axpy_ref ¶
convert_copy_ref ¶
Numpy reference for convert_copy.
ml_dtypes rounds f32 -> bf16 half-to-even, exactly as the kernel's
conv_even does, so the cast is the reference and the match is
bit-for-bit.
Source code in python/iron/kernels/datamovement.py
expand_ref ¶
Numpy reference for expand.
payload holds, per tile, tile_size packed uint4 values
(tile_size // 2 bytes, low nibble first) followed by one bf16 scale per
group_size elements; the result is nibble * scale-of-its-group.
Source code in python/iron/kernels/datamovement.py
expand_sample ¶
Random expand payloads: packed uint4 values then bf16 scales in [0.1, 1).
Source code in python/iron/kernels/datamovement.py
transpose_ref ¶
Numpy reference for transpose.
Transposes each subtile x subtile block of the dim_n x dim_m
matrix in place -- the blocks move, the matrix does not.
Source code in python/iron/kernels/datamovement.py
axpy ¶
SAXPY kernel: z = a * x + y over bf16 tiles.
The scalar a and element count are passed to the kernel at runtime, so a
design supplies (x, y, a, z, size). The vectorized path processes 64
elements per iteration; tile_size must therefore be a multiple of 64.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tile_size
|
int
|
Elements per tile (multiple of 64 for the vectorized path). |
1024
|
vectorized
|
bool
|
If |
True
|
Returns:
| Type | Description |
|---|---|
ExternalFunction
|
ExternalFunction for the saxpy kernel. |
Raises:
| Type | Description |
|---|---|
ValueError
|
When |
Source code in python/iron/kernels/datamovement.py
convert_copy ¶
Convert-copy kernel: element-preserving float32 -> bfloat16.
Reads a length-tile_size f32 tile and writes the same number of bf16
elements (halving the byte footprint). Element count is a runtime arg; the
kernel processes 16 elements per iteration, so tile_size must be a
multiple of 16.
Backed by aie_kernels/datamovement/cast_f32_bf16.cc (symbol
cast_f32_bf16_row), which rounds with conv_even — bit-for-bit
agreeing with a host AVX512-BF16 pack — and restores the core's rounding
mode on exit. The same source builds for aie2.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tile_size
|
int
|
Elements per tile (multiple of 16). |
1024
|
Returns:
| Type | Description |
|---|---|
ExternalFunction
|
ExternalFunction for |
Raises:
| Type | Description |
|---|---|
ValueError
|
When |
Source code in python/iron/kernels/datamovement.py
expand ¶
Dequantize kernel: uint4 -> bfloat16 with per-group scale factors.
Each tile holds tile_size packed unsigned int4 values followed by one
bf16 scale factor per group_size-element group; the kernel zero-extends
and scales into tile_size bf16 outputs (no zero point). tile_size
and group_size are baked in at compile time via -DTILE_SIZE /
-DGROUP_SIZE (group_size must be a multiple of 32, matching the C++
static_assert).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tile_size
|
int
|
Number of uint4 elements per tile. |
1024
|
group_size
|
int
|
Elements sharing one scale factor (multiple of 32). |
32
|
Returns:
| Type | Description |
|---|---|
ExternalFunction
|
ExternalFunction for |
Raises:
| Type | Description |
|---|---|
ValueError
|
When |
Source code in python/iron/kernels/datamovement.py
rope ¶
rope(
tile_size: int = 1024,
two_halves: bool = False,
*,
cols: int | None = None
) -> ExternalFunction
RoPE positional rotation over bf16 tiles; dims read at runtime.
Design passes (in, lut, out, dims). two_halves selects the
HuggingFace-style rope_two_halves over the Llama-paper interleave
rope. cols aliases tile_size. Both architectures use the generic
source; rows must be positive multiples of 16 (interleaved) or 32
(two halves, keeping each half 32-byte aligned). Each input row has its own
streamed (cos, sin) LUT.
Source code in python/iron/kernels/datamovement.py
rope_ref ¶
Rotate bf16 pairs by an interleaved (cos, sin) LUT, in either RoPE layout.
Source code in python/iron/kernels/datamovement.py
transpose ¶
transpose(
dim_m: int = 32,
dim_n: int = 32,
subtile: int = 4,
dtype: type = bfloat16,
) -> ExternalFunction
Blocked transpose through aie::transpose.
Transposes each subtile x subtile block of a dim_n x dim_m
matrix in place: the blocks stay put, the elements inside them move.
dim_m / dim_n are compile-time (-DDIM_m / -DDIM_n). The
kernel only moves bytes, so any 1-, 2- or 4-byte dtype works and
selects -DBIT_WIDTH as the other generic kernels do; bf16 is the
default. programming_examples/basic/transposes uses this kernel for its
combined strategy.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
dim_m
|
int
|
Inner (contiguous) dimension. |
32
|
dim_n
|
int
|
Outer dimension. |
32
|
subtile
|
int
|
Block size to transpose, 4 ( |
4
|
dtype
|
type
|
Element type, 1, 2 or 4 bytes wide. |
bfloat16
|
Returns:
| Type | Description |
|---|---|
ExternalFunction
|
ExternalFunction for the selected transpose variant. |
Raises:
| Type | Description |
|---|---|
ValueError
|
When |
Source code in python/iron/kernels/datamovement.py
Independent, target-native zero-fill kernel.
zero ¶
zero(
tile_size: int | tuple[int, ...] = 1024,
dtype: type | dtype = int32,
*,
vectorized: bool = True,
use_chess: bool = False
) -> ExternalFunction
Fill one tile with zeros, independently of any compute kernel.
tile_size is an element count or shape. For v8bfp16ebs8 it
counts eight-value blocks, matching the ndarray ABI; all nine bytes
of every block (exponent and mantissas) are cleared. Vector stores
use the target's native width, with a scalar tail for smaller tiles.
Source code in python/iron/kernels/zero.py
Quantization¶
Packed quantization kernel factories and byte-exact host references.
q4nx_dequant_ref ¶
Dequantize packed q4nx to the kernel's bfp16ebs8 output bytes.
Input has shape (..., m_tile*k_tile//2 + 4*m_tile*k_tile//group).
It contains little-endian bf16 scales, then bf16 minima, both indexed
[k_group, n], followed by unsigned nibbles (low nibble first) indexed
[n//16, k, n%16]. Values are min + scale * nibble.
Accumulator results narrow to bf16 with floor rounding before BFP
conversion. Output is uint8 with the same leading dimensions, indexed
[k//ct_k, n//8, (k%ct_k)//8, n%8, 9]: each nine-byte block is an
exponent followed by eight signed mantissas for consecutive k values.
Finite scales/minima and finite dequantized bf16 results are required.
Source code in python/iron/kernels/quant.py
q4nx_dequant ¶
AIE2P-only q4nx dequantization into GEMM B-operand BFP storage.
m_tile counts n rows (a multiple of 16); k_tile must be divisible
by both group and ct_k. The latter two are positive multiples of
eight; s and t must be eight. Groups need not divide k slices.
See q4nx_dequant_ref for the packed input and output layouts.
Both arguments are uint8 byte buffers, including the bfp16ebs8 output, so the generic harness compares the complete encoded result byte for byte rather than decoding or quantizing it again. The kernel saves, selects and restores floor rounding itself; no setup kernel is needed.
Source code in python/iron/kernels/quant.py
Core state¶
set_rounding: the core's rounding-mode register.
The AIE core narrows accumulators (an srs shift, a bf16 store) in the
rounding mode its mode register holds, and a fresh core boots in floor.
A kernel that needs another mode names a setter as its contract's setup,
and a design calls it once before the kernel's first call; the mode persists
on that core until something changes it.
RoundingMode ¶
Bases: str, Enum
An aie::rounding_mode, named as the C++ enumerator is.
set_rounding ¶
set_rounding(
mode: RoundingMode = CONV_EVEN,
) -> ExternalFunction
Set the core's rounding mode using always-inline, merge-linked LLVM IR.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
mode
|
RoundingMode
|
The |
CONV_EVEN
|
Returns:
| Type | Description |
|---|---|
ExternalFunction
|
ExternalFunction |
Source code in python/iron/kernels/core.py
Reduction¶
Reduction kernel factories: reduce_add, reduce_min, reduce_max, compute_max.
reduce_add_ref ¶
Numpy reference for reduce_add: per-tile sum in int64.
reduce_min_ref ¶
Numpy reference for reduce_min: per-tile minimum.
reduce_max_ref ¶
Numpy reference for reduce_max: per-tile maximum.
reduce_add ¶
reduce_add(
tile_size: int = 1024,
dtype: type = int32,
vectorized: bool = True,
) -> ExternalFunction
Reduction kernel: sums all elements of a tile to a scalar.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tile_size
|
int
|
Number of elements in the input tile. |
1024
|
dtype
|
type
|
Element data type (only |
int32
|
vectorized
|
bool
|
If |
True
|
Returns:
| Type | Description |
|---|---|
ExternalFunction
|
ExternalFunction configured for the reduce_add kernel. |
Raises:
| Type | Description |
|---|---|
ValueError
|
When |
Source code in python/iron/kernels/reduce.py
reduce_min ¶
reduce_min(
tile_size: int = 1024,
dtype: type = int32,
vectorized: bool = True,
) -> ExternalFunction
Reduction kernel: finds the minimum element of a tile.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tile_size
|
int
|
Number of elements in the input tile. |
1024
|
dtype
|
type
|
Element data type (only |
int32
|
vectorized
|
bool
|
If |
True
|
Returns:
| Type | Description |
|---|---|
ExternalFunction
|
ExternalFunction configured for the reduce_min kernel. |
Raises:
| Type | Description |
|---|---|
ValueError
|
When |
Source code in python/iron/kernels/reduce.py
reduce_max ¶
reduce_max(
tile_size: int = 1024,
dtype: type = int32,
vectorized: bool = True,
) -> ExternalFunction
Reduction kernel: finds the maximum element of a tile (int32 or bfloat16).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tile_size
|
int
|
Number of elements in the input tile. |
1024
|
dtype
|
type
|
Element data type ( |
int32
|
vectorized
|
bool
|
If |
True
|
Returns:
| Type | Description |
|---|---|
ExternalFunction
|
ExternalFunction configured for the reduce_max kernel. |
Raises:
| Type | Description |
|---|---|
ValueError
|
When |
Source code in python/iron/kernels/reduce.py
compute_max ¶
Pairwise scalar max — companion to reduce_max.
Used for multi-core reductions where each core produces a partial max and a final tree reduces them pairwise.
Lives in the same reduce_max.cc as reduce_max,
but uses an unspecialized object independent of reduction tile sizes.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
dtype
|
type
|
Element data type ( |
int32
|
Returns:
| Type | Description |
|---|---|
ExternalFunction
|
ExternalFunction configured for the |
ExternalFunction
|
is |
ExternalFunction
|
(DMA-aligned) tile of |
Raises:
| Type | Description |
|---|---|
ValueError
|
When |
Source code in python/iron/kernels/reduce.py
compute_max_ref ¶
Numpy reference for compute_max.
The kernel compares only element 0 of each (DMA-padded) input tile and
writes element 0 of the output; the reference does the same, returning
max(a[..., 0], b[..., 0]) with a trailing axis of length 1.
Source code in python/iron/kernels/reduce.py
Linear algebra¶
Linear algebra kernel factories: mm, mv, cascade_mm.
StreamDimsABC ¶
Bases: NamedTuple
The three dims_to_stream a matmul design needs, one per operand.
None for an operand a build streams untransformed.
MatrixKernel ¶
Bases: _ZeroInitializedKernel
A kernel whose first three operands are the A, B and C of a product.
The blocking and the DMA transforms are declared once, on the contract's
operand layouts (TensorLayout.block and .stream); these
properties read them back in the form the matrix designs consume.
mac_dims
property
¶
(r, s, t): the MMUL micro-tile of A (r, s), B (s, t) and C (r, t).
stream_dims
property
¶
stream_dims: StreamDimsABC
The dims_to_stream a design applies to A, B and C; None streams as stored.
mm_acc_dtype ¶
Return what mm.cc accumulates in for input_dtype.
accauto is acc32 for 8-bit and acc64 for 16-bit integer inputs, and
float32 for bf16.
Source code in python/iron/kernels/linalg.py
mv_ref ¶
mm_bfp_ref ¶
Numpy reference for mm_bfp: a @ b on bfp16ebs8-quantized operands.
a is (M, K) and b (K, N) float; each is quantized the way
the host encodes it for the kernel (blocks of 8 along K, see
aie.utils.bfp) and the product is accumulated in float64. The
kernel's own output is bfp16ebs8 too, which the tolerance covers.
Source code in python/iron/kernels/linalg.py
mm_bfp_mixed_ref ¶
Numpy reference for mm_bfp with mixed=True.
a arrives bf16 and the core converts it, under the rounding mode
mm_bfp_mixed.cc pins, so it is quantized conv_even here. b is
encoded by the host, which truncates, so it keeps the default as in
mm_bfp_ref.
Pairing A with the wrong mode is the difference between 10 mismatching outputs and 2791, on a 64x64x64 tile of large inputs; leaving A unquantized altogether gives 1369.
Source code in python/iron/kernels/linalg.py
mm_tile_ref ¶
One mm call: a (dim_m, dim_k) tile times a (dim_k, dim_n) one.
Tiles arrive flattened as (calls, ...), one row per call, and one
(dim_m * dim_n,) row comes back per call.
Source code in python/iron/kernels/linalg.py
mv_tile_ref ¶
One mv call: a (dim_m, dim_k) tile times a (dim_k,) vector.
Source code in python/iron/kernels/linalg.py
mm_bfp_tile_ref ¶
One mm_bfp call, on operands quantized as the host encodes them.
Blocks of 8 run along K for both operands, so B is quantized transposed.
With mixed the A tile stays bf16 and the core converts it itself, with
a rounding this does not model -- which is what the wider mixed tolerance
covers. See mm_bfp_mixed_ref for
why modelling it as a host-side quantize is worse, not better.
Source code in python/iron/kernels/linalg.py
mm_stream_dims ¶
mm_stream_dims(
dim_m: int,
dim_k: int,
dim_n: int,
mac_dims,
*,
b_col_maj: bool = False,
c_col_maj: bool = False
) -> StreamDimsABC
DMA dims_to_stream that feed mm.cc its (r, s, t) micro-tiles.
mm.cc consumes A, B and produces C in the micro-tile blocking given by
mac_dims; a plain row-major stream yields wrong numbers, not an error.
Every matmul design (single_core, whole_array, cascade, ...) derives these
same three transforms from mac_dims; kernels.mm(...).stream_dims
carries them so designs do not re-derive them. Keys "A", "B", "C".
b_col_maj describes a B tile stored as (n, k) (the transpose) and
c_col_maj a C tile emitted as (n, m), matching the kernel's
-DB_COL_MAJ / -DC_COL_MAJ builds.
Source code in python/iron/kernels/linalg.py
mm ¶
mm(
dim_m: int = 64,
dim_k: int = 64,
dim_n: int = 64,
input_dtype: type = int16,
output_dtype: type = int16,
vectorized: bool = True,
b_col_maj: bool = False,
c_col_maj: bool = False,
use_chess: bool = False,
emulate_bf16_mmul_with_bfp16: bool = False,
round_conv_even: bool = False,
) -> MatrixKernel
Matrix-multiply kernel: C += A * B.
.zero initializes the accumulator using the independent, reusable
kernels.zero(dim_m * dim_n, output_dtype) kernel. The contract declares
the same initializer for the generic harness.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
dim_m
|
int
|
Number of rows of A / C. |
64
|
dim_k
|
int
|
Number of columns of A / rows of B. |
64
|
dim_n
|
int
|
Number of columns of B / C. |
64
|
input_dtype
|
type
|
Input element type ( |
int16
|
output_dtype
|
type
|
Output element type. |
int16
|
vectorized
|
bool
|
If |
True
|
b_col_maj
|
bool
|
If |
False
|
c_col_maj
|
bool
|
If |
False
|
use_chess
|
bool
|
If |
False
|
emulate_bf16_mmul_with_bfp16
|
bool
|
AIE2P only, bf16 inputs only. When
|
False
|
round_conv_even
|
bool
|
AIE2 only, bf16 inputs only. When |
False
|
Returns:
| Type | Description |
|---|---|
MatrixKernel
|
ExternalFunction configured for the matmul kernel. |
Raises:
| Type | Description |
|---|---|
ValueError
|
When |
Source code in python/iron/kernels/linalg.py
486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 | |
mv ¶
mv(
dim_m: int = 32,
dim_k: int = 32,
input_dtype: type = int16,
output_dtype: type = int32,
vectorized: bool = True,
use_chess: bool = False,
vec_size: int = 64,
output_rows: int | None = None,
) -> ExternalFunction
Matrix-vector multiply kernel: c += A * b.
(np.int16, np.int32) builds aie_kernels/linalg/mv_i16.cc; its
vectorized path reads A word-transposed, which A's layout carries
(contract.layouts[0].stream). Its .zero companion initializes C
with the independent kernels.zero(dim_m, output_dtype).
(bfloat16, bfloat16) builds
aie_kernels/linalg/mv_bf16.cc, IRON's GEMV kernel, whose signature
is (m, row_offset, A, b, c): row_offset shifts the write into
c so one core can fill several output blocks; A is row-major.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
dim_m
|
int
|
Number of rows of A (output vector length). |
32
|
dim_k
|
int
|
Number of columns of A (input vector length). |
32
|
input_dtype
|
type
|
Input element type: |
int16
|
output_dtype
|
type
|
Output element type: |
int32
|
vectorized
|
bool
|
If |
True
|
use_chess
|
bool
|
If |
False
|
vec_size
|
int
|
bf16 only: the kernel's |
64
|
output_rows
|
int | None
|
bf16 only: the rows of a C tile that successive calls
fill |
None
|
Returns:
| Type | Description |
|---|---|
ExternalFunction
|
ExternalFunction configured for the matvec kernel. |
Raises:
| Type | Description |
|---|---|
ValueError
|
When the dtype combination is not supported. |
Source code in python/iron/kernels/linalg.py
639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 | |
mv_bf16_ref ¶
Numpy reference for the bf16 mv: a @ b over m rows, accumulated in float32.
a is row-major (rows, K); only the first m rows are computed,
and the kernel writes them at row_offset into c.
Source code in python/iron/kernels/linalg.py
mm_bfp ¶
mm_bfp(
dim_m: int = 64,
dim_k: int = 64,
dim_n: int = 64,
mixed: bool = False,
) -> MatrixKernel
Block-floating-point matmul C += A @ B on bfp16ebs8 blocks (aie2p only).
mixed=False (aie_kernels/linalg/mm_bfp.cc): A, B and C are
v8bfp16ebs8 blocks, all pre-shuffled into the mmul layout, so no
DMA transform applies (stream_dims is None for every operand).
mixed=True (mm_bfp_mixed.cc): A is bf16 in the (r, s, t)
micro-tile layout, B is bfp16ebs8, C is bf16; stream_dims.A and
.C carry the transforms and .B is None.
Initialize C with the independent kernels.zero factory.
The host holds B transposed (b_col_maj), and
every bfp16ebs8 operand is encoded and shuffled into the mmul tile
layout on the host with aie.utils.bfp, which is what the generic
harness does; the contract's reference multiplies the quantized
operands. These are the kernels
programming_examples/ml/block_datatypes/matrix_multiplication build.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
dim_m
|
int
|
Tile rows of A and C (multiple of 8). |
64
|
dim_k
|
int
|
Tile columns of A / rows of B (multiple of 8). |
64
|
dim_n
|
int
|
Tile columns of B and C (multiple of 8). |
64
|
mixed
|
bool
|
bf16 A and C with bfp16 B. |
False
|
Source code in python/iron/kernels/linalg.py
809 810 811 812 813 814 815 816 817 818 819 820 821 822 823 824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 851 852 853 854 855 856 857 858 859 860 861 862 863 864 865 866 867 868 869 870 871 872 873 874 875 876 877 878 879 880 881 882 883 884 885 886 887 888 889 890 891 892 893 894 895 896 897 | |
mm_bfp_shuffle ¶
mm_bfp_shuffle(
dim_m: int = 64,
dim_k: int = 64,
dim_n: int = 64,
*,
in_shape: tuple | None = None,
out_shape: tuple | None = None,
unshuffle: bool = False
) -> ExternalFunction
Scalar shuffle of a bfp16ebs8 tile into (or out of) the mmul block layout (aie2p).
scalar_shuffle(in, out, tile_width, tile_height, unshuffle) from
mm_bfp.cc; the in-core-shuffle block-datatype examples run it before
mm_bfp. By default the input tile is sized like mm_bfp's A and
the output like its C; in_shape / out_shape (in v8bfp16ebs8
blocks) override that, e.g. (dim_m, dim_k // 8) twice to shuffle an
A tile in place, matching the ObjectFifo types a design already uses.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
dim_m
|
int
|
Tile rows (multiple of 8). |
64
|
dim_k
|
int
|
A's tile columns (multiple of 8). |
64
|
dim_n
|
int
|
C's tile columns (multiple of 8). |
64
|
in_shape
|
tuple | None
|
Input tile shape in blocks; default |
None
|
out_shape
|
tuple | None
|
Output tile shape in blocks; default |
None
|
unshuffle
|
bool
|
Validate the other direction, out of the block layout back to row-major; the harness passes it as the call's last argument. |
False
|
Source code in python/iron/kernels/linalg.py
mha ¶
mha(
dim_m: int = 64,
dim_k: int = 64,
dim_n: int = 64,
pv: bool = False,
b_col_maj: bool = False,
emulate_bf16_mmul_with_bfp16: bool = False,
) -> MatrixKernel
Flash-attention toolkit from aie_kernels/linalg/mha.cc.
One translation unit that includes softmax.cc and mm.cc and
exports the symbols an attention dataflow composes over one micro-tile.
The returned kernel is one of the toolkit's two matmuls, both
accumulating into C, so it is a
MatrixKernel judged like
mm. By default that is the QK^T product
matmul_bf16_bf16_wrapper, mm.cc's bf16 product on its 4x8x8
micro-tile behind an idx_buffer gate (the call runs when
idx[0] <= idx[1]) bound here to [0, 0]. With pv it is the
P*V product matmul_bf16_bf16_rowmaj, mha.cc's own expansion of
the native 8x8x8 micro-tile and ungated. matmul_PV is that same
product preceded by a row rescale, which needs the online softmax's
running state and so is exercised by test_mha_e2e.py instead.
Bind the others from the same object with
fn.object_file.bind(symbol, arg_types):
matmul_bf16_bf16_wrapper_scalar, partial_softmax,
matmul_PV, rescale_O, init_scale_buffer. It declares but
does not define passThroughLine: take that from
passthrough(dtype=np.int32), as IRON's MHA operator does.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
dim_m
|
int
|
Rows of the micro-tile (multiple of 16). |
64
|
dim_k
|
int
|
Depth of the micro-tile (multiple of 8). |
64
|
dim_n
|
int
|
Columns of the micro-tile (multiple of 16). |
64
|
pv
|
bool
|
If |
False
|
b_col_maj
|
bool
|
If |
False
|
emulate_bf16_mmul_with_bfp16
|
bool
|
As for |
False
|
Source code in python/iron/kernels/linalg.py
977 978 979 980 981 982 983 984 985 986 987 988 989 990 991 992 993 994 995 996 997 998 999 1000 1001 1002 1003 1004 1005 1006 1007 1008 1009 1010 1011 1012 1013 1014 1015 1016 1017 1018 1019 1020 1021 1022 1023 1024 1025 1026 1027 1028 1029 1030 1031 1032 1033 1034 1035 1036 1037 1038 1039 1040 1041 1042 1043 1044 1045 1046 1047 1048 1049 1050 1051 1052 1053 1054 1055 1056 1057 1058 1059 1060 1061 1062 1063 1064 1065 1066 1067 1068 1069 1070 1071 1072 1073 1074 1075 1076 1077 1078 1079 1080 1081 1082 1083 1084 1085 1086 | |
mha_softmax ¶
One 64x64 block of mha.cc's online softmax, partial_softmax.
Writes the block's unnormalized weights P = exp2(A * s - m), with
s = log2(e) / 8 and m each query row's running maximum, and
updates the running state scale_buffer: [m, m, l, exp2(m_prev -
m)], 64 rows each. The state is zeroed before every call, so each call
is a first key block against a running maximum of 0; the carry across
blocks is test_mha_e2e.py's. The kernel masks by overwriting
A's masked entries in place.
Which block it is stays a runtime operand, as in the kernel: idx is
(key block, query block) (equal on the causal diagonal, key block
past query block skipped), and S_q_eff/S_kv_eff are the
sequence lengths whose tails pad the block.
Source code in python/iron/kernels/linalg.py
mha_softmax_ref ¶
Numpy reference for mha_softmax: (P, scale_buffer).
True exp2 rather than the device's interpolant, from a zeroed
running state. m is rounded to bf16 before P uses it, as the
kernel stores it. A padded row keeps m = l = 0; a skipped or wholly
padded block leaves the state untouched.
Source code in python/iron/kernels/linalg.py
prefill_fv ¶
Flash-attention prefill toolkit from aie_kernels/linalg/flash_attn_prefill.cc.
One translation unit per geometry, exporting the five steps an attention
prefill dataflow composes over one query chunk. The returned kernel is the
y += S*V step, mm.cc-style bf16 products on the native 8x8x8
micro-tile accumulating into a float32 y, so it is sampled and judged like
mm. Its key-chunk index j is bound to 0.
Bind the others from the same object with
fn.object_file.bind(symbol, arg_types): prefill_round_begin,
prefill_qk_step, prefill_block_mid, prefill_epilogue.
One reference serves both geometries even though they decompose
differently, because PrefillGeom<256>::reorder_s deinterleaves its 2x2
attn_qk output into the same block order the 1x1 geometry produces
directly, and block_mid runs it before this step sees S. The
geometries differ only in operand storage order, which the layouts carry.
head_dim picks the geometry, and two instantiations coexist in one
design: differing -DPREFILL_HEAD_DIM gives each its own object and its
own symbol prefix.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
head_dim
|
int
|
512 for global attention, 256 for sliding-window. |
512
|
Source code in python/iron/kernels/linalg.py
1189 1190 1191 1192 1193 1194 1195 1196 1197 1198 1199 1200 1201 1202 1203 1204 1205 1206 1207 1208 1209 1210 1211 1212 1213 1214 1215 1216 1217 1218 1219 1220 1221 1222 1223 1224 1225 1226 1227 1228 1229 1230 1231 1232 1233 1234 1235 1236 1237 1238 1239 1240 1241 1242 1243 1244 1245 1246 1247 1248 1249 1250 1251 1252 1253 1254 1255 1256 1257 1258 1259 1260 1261 1262 1263 1264 1265 1266 1267 1268 1269 1270 1271 1272 1273 1274 1275 1276 1277 1278 1279 1280 1281 1282 | |
prefill_fv_ref ¶
Numpy reference for prefill_fv: one y += S @ V key chunk.
y is not an argument: it is InOut, so the contract's initializer
zeroes it before each independent call and the reference computes the
whole product.
Accumulated in float32, not the float64 of
mm_tile_ref. The kernel's y is
float32 and its MMUL accumulates there, so a float64 reference is more
precise than the kernel is defined to be and the comparison measures that
gap rather than correctness. The difference is invisible when the product
narrows back to bf16, as every other matmul here does, and decisive when
it does not: on inputs near 1e4 the eight products reach 1e8 and cancel,
leaving a result near zero whose float64 and float32 sums differ by a
whole float32 ULP of the intermediate. Products themselves are exact
either way -- bf16 carries 8 mantissa bits and 8 + 8 < 24.
Source code in python/iron/kernels/linalg.py
mm_bfp_shuffle_ref ¶
Numpy reference for mm_bfp_shuffle: aie.utils.bfp.shuffle of one tile's bytes.
tile is the encoded (tile_height, tile_width) tile as bytes
(tile_width in values); returns the reordered bytes.
Source code in python/iron/kernels/linalg.py
cascade_mm ¶
cascade_mm(
dim_m: int = 64,
dim_k: int = 64,
dim_n: int = 64,
input_dtype: type = int16,
output_dtype: type = int16,
use_chess: bool = False,
) -> _CascadeMatrixKernel
Build the GET half of a cascade matrix multiply: C += A * B + cascade.
cascade_mm.cc emits all three cascade variants (get_only,
put_only, put_get) in one object. This binds get_only (also
available as .get_only), with .put_only and .put_get siblings;
cascade_mm_put is the PUT half
that feeds it, and put_get serves longer chains:
fn.object_file.bind("matmul_scalar_cascade_put_get_<dtype>", fn.arg_types()).
.zero initializes accumulators using independent kernels.zero.
The pair is a two-tile
design, which the generic builder does not run; the device test builds
and judges it (test/python/npu/test_kernels_e2e.py). The partial sum
crosses the cascade as a 32-bit integer lane: with a floating-point
output type the PUT half's product is truncated toward zero.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
dim_m
|
int
|
Number of rows of A / C. |
64
|
dim_k
|
int
|
Number of columns of A / rows of B. |
64
|
dim_n
|
int
|
Number of columns of B / C. |
64
|
input_dtype
|
type
|
Input element type. |
int16
|
output_dtype
|
type
|
Output element type. |
int16
|
use_chess
|
bool
|
If |
False
|
Raises:
| Type | Description |
|---|---|
ValueError
|
When the dtype combination is not supported. |
Source code in python/iron/kernels/linalg.py
1332 1333 1334 1335 1336 1337 1338 1339 1340 1341 1342 1343 1344 1345 1346 1347 1348 1349 1350 1351 1352 1353 1354 1355 1356 1357 1358 1359 1360 1361 1362 1363 1364 1365 1366 1367 1368 1369 1370 1371 1372 1373 1374 1375 1376 1377 1378 1379 1380 1381 1382 1383 1384 1385 1386 1387 1388 1389 1390 1391 1392 1393 1394 1395 1396 1397 1398 1399 1400 1401 1402 1403 1404 1405 1406 1407 1408 1409 1410 1411 1412 1413 1414 1415 1416 1417 1418 | |
cascade_mm_put ¶
cascade_mm_put(
dim_m: int = 64,
dim_k: int = 64,
dim_n: int = 64,
input_dtype: type = int16,
output_dtype: type = int16,
use_chess: bool = False,
) -> MatrixKernel
Build the PUT half of cascade_mm: A * B onto the cascade stream.
Same object and arguments as the GET half. put_only never touches
its third argument (the ABI just mirrors get_only), so the contract
binds it to zeros. Its result leaves on the cascade stream, so it is
judged with its GET half by the device test, not by the generic builder.
Source code in python/iron/kernels/linalg.py
Convolution¶
Convolution kernel factories: conv2dk1/3/14, bottleneck (bn_*) variants.
The conv2dk1, conv2dk1_i8, conv2dk1_skip, conv2dk3 and conv2dk14 factories specialize their dimensions at compile time. Their runtime dimension arguments remain in the ABI and must match the factory dimensions; scales, region checks and channel offsets remain runtime values.
conv2dk1_ref ¶
Numpy reference for conv2dk1: 1x1 conv, requantized.
Layouts are the kernel's: activations [C/8][W][8] (x is one
line of input_width * input_channels values, or (calls, ...) of
them), weights [OC/8][IC/8][ic8][oc8], output [OC/8][W][8] as
uint8. out = sat_u8((sum_ic x * w + 2**(scale-1)) >> scale),
i.e. a fused ReLU. Exact for the scalar path of conv2dk1.cc.
Source code in python/iron/kernels/conv.py
conv2dk1_i8_ref ¶
Numpy reference for conv2dk1_i8: 1x1 conv to int8.
The layouts of conv2dk1_ref with an
int8 output and no ReLU:
out = sat_i8((sum_ic x * w + 2**(scale-1)) >> scale). Exact for the
scalar path of conv2dk1_i8.cc.
Source code in python/iron/kernels/conv.py
conv2dk1_skip_ref ¶
conv2dk1_skip_ref(
x0,
x1,
weights,
skip,
input_width,
input_channels,
output_channels,
scale,
skip_scale,
)
Numpy reference for conv2dk1_skip: 1x1 conv plus residual.
x0 and x1 each hold half the input channels ([IC/16][W][8]
lines, x1 the upper half); weights are [OC/8][IC/8][ic8][oc8]
over all of them and skip is an [OC/8][W][8] line. The conv sum
is requantized and saturated to int8 first, then the residual is
added and the total requantized to uint8:
conv = sat_i8((sum_ic x * w + 2**(scale-1)) >> scale)
out = sat_u8((conv + skip + 2**(skip_scale-1)) >> skip_scale)
Exact for the scalar path of conv2dk1_skip.cc. A shift of 0 is no
shift (the scalar path's 1 << -1 is not defined for it; the vector
path handles 0).
Source code in python/iron/kernels/conv.py
conv2dk3_ref ¶
conv2dk3_ref(
line0,
line1,
line2,
weights,
input_width,
input_channels,
output_channels,
kernel_width,
kernel_height,
check,
scale,
channel_offset,
)
Numpy reference for conv2dk3: 3x3 conv over three lines.
Produces the output line for line1. Activations are [C/8][W][8]
lines; weights [WOC/8][IC/8][3 rows][kernel_width][ic8][oc8] where
WOC may exceed output_channels (channel_offset selects this
call's slice, in units of 8 channels). The spatial border is zero
padded; check is 0 (top: line0 ignored), 1 (middle) or 2
(bottom: line2 ignored), matching the kernel's region enum.
out = sat_u8((sum + 2**(scale-1)) >> scale). Exact for the scalar
path of conv2dk3.cc; kernel_height is accepted for the
signature and is 3.
Source code in python/iron/kernels/conv.py
conv2dk1_skip_init_ref ¶
conv2dk1_skip_init_ref(
x0,
x1,
weights,
skip,
input_width,
input_channels,
output_channels,
skip_input_channels,
scale,
skip_scale,
scale_skip_conv,
)
Numpy reference for conv2dk1_skip_init: 1x1 conv plus a projected residual.
Like conv2dk1_skip_ref, but the
residual is itself a 1x1 conv of skip ([ICs/8][W][8]) with the
weights stored after the main ones ([OC/8][ICs/8][ic8][oc8] at
offset OC * IC):
conv = sat_i8((sum_ic x * w + 2**(scale-1)) >> scale)
proj = sat_i8((sum_ics skip * ws + 2**(scale_skip_conv-1)) >> scale_skip_conv)
out = sat_u8((conv + proj + 2**(skip_scale-1)) >> skip_scale)
Exact for the scalar path of conv2dk1_skip_init.cc; a shift of 0
is no shift.
Source code in python/iron/kernels/conv.py
conv2dk14_ref ¶
Numpy reference for conv2dk14: a KxK patch conv (stride K) to int8.
One call covers T = input_width / kernel_width patches of K*K
RGBA pixels. Layouts are the kernel's: input [T/8][P/2][t8][p2][4]
uint8 (P = K*K pixels of 4 channels), weights
[OC/8][P/2][p2][4][oc8] int8, output [OC/8][T][oc8]
int8: out = sat_i8((sum_{p,c} x * w + 2**(scale-1)) >> scale).
input_channels is accepted for the signature (the pixel is RGBA).
Exact for the scalar path of conv2dk14.cc.
Source code in python/iron/kernels/conv.py
dwconv1d_channels_first ¶
dwconv1d_channels_first(
seq_len: int = 1024,
kernel_size: int = 9,
bias: bool = True,
) -> ExternalFunction
Depthwise 1-D cross-correlation on one bf16 channel.
out[t] = bias + sum_p w[p] * x_pad[t + p] for t < seq_len: a
'same' convolution when x_pad is the channel zero-padded by
(kernel_size - 1) // 2 on each side plus DWCONV1D_TAIL don't-care
elements (programming_examples/ml/dwconv1d builds it that way). The
weight row holds kernel_size taps followed by the bias, whether or
not bias is enabled.
One channel per call with time contiguous, vectorized along time with
scalar taps. See
dwconv1d_channels_last for
the transposed layout, and the "Choosing a depthwise conv1d" section of
aie_kernels/README.md for which to reach for.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
seq_len
|
int
|
Outputs per call (multiple of 16). |
1024
|
kernel_size
|
int
|
Taps, 1 to 17. |
9
|
bias
|
bool
|
Add the trailing weight as a bias. |
True
|
Source code in python/iron/kernels/conv.py
dwconv1d_channels_first_ref ¶
Numpy reference for dwconv1d_channels_first on the padded row(s).
x_pad is (..., seq_len + DWCONV1D_TAIL); w is (..., kernel_size + 1).
Source code in python/iron/kernels/conv.py
dwconv1d ¶
Compatibility alias for dwconv1d_channels_first.
Source code in python/iron/kernels/conv.py
dwconv1d_ref ¶
Compatibility alias for dwconv1d_channels_first_ref.
Source code in python/iron/kernels/conv.py
dwconv1d_channels_last ¶
Depthwise 1-D conv over a channels-last layout, 5 taps.
y[c] = clamp(sum_{t=0..4} w_t[c] * x_t[c], lo, hi) for c < channels:
one output timestep across every channel, with per-channel taps. The five
taps arrive as five separate base pointers, oldest first, so a depth-5
ObjectFifo is itself the sliding window.
Counterpart to
dwconv1d_channels_first;
layout picks the vectorization axis, so neither subsumes the other. See
the "Choosing a depthwise conv1d" section of aie_kernels/README.md.
The five weight planes are five independent arguments, like the taps, so where they live is the design's business: they need not be one buffer, or evenly spaced.
lo/hi are runtime buffers the design writes, bound here to
+/-_CLAMP_LIMIT so the kernel is judged against a reference clamping to
the same pair.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
channels
|
int
|
Channels per call (multiple of 32). |
256
|
clamp
|
bool
|
Clamp the result to the runtime |
True
|
Source code in python/iron/kernels/conv.py
dwconv1d_channels_last_ref ¶
dwconv1d_channels_last_ref(
w_0,
w_1,
w_2,
w_3,
w_4,
x_0,
x_1,
x_2,
x_3,
x_4,
*,
lo,
hi,
clamp: bool
)
Numpy reference for dwconv1d_channels_last: one timestep over all channels.
Each tap is an independent plane, so this is a plain sum_t w_t * x_t.
Accumulated in float32, which is what the kernel's accfloat is, then
narrowed once on store; clamp applies after the narrowing, as the
kernel's aie::clamp does on the already-bf16 vector.
Source code in python/iron/kernels/conv.py
conv2dk1 ¶
conv2dk1(
input_width: int = 32,
input_channels: int = 64,
output_channels: int = 64,
act_dtype: type = int8,
) -> ExternalFunction
1x1 convolution kernel.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
input_width
|
int
|
Spatial width of the input. |
32
|
input_channels
|
int
|
Number of input channels. |
64
|
output_channels
|
int
|
Number of output channels. |
64
|
act_dtype
|
type
|
Activation data type ( |
int8
|
Returns:
| Type | Description |
|---|---|
ExternalFunction
|
ExternalFunction configured for the conv2dk1 kernel. |
Raises:
| Type | Description |
|---|---|
ValueError
|
When |
Source code in python/iron/kernels/conv.py
conv2dk3 ¶
conv2dk3(
input_width: int = 32,
input_channels: int = 64,
output_channels: int = 64,
act_dtype: type = int8,
weight_output_channels: int | None = None,
) -> ExternalFunction
3x3 convolution kernel.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
input_width
|
int
|
Spatial width of the input. |
32
|
input_channels
|
int
|
Number of input channels. |
64
|
output_channels
|
int
|
Number of output channels produced by this call. |
64
|
act_dtype
|
type
|
Activation data type ( |
int8
|
weight_output_channels
|
int | None
|
Total number of output channels stored in the
weights buffer. Defaults to |
None
|
Returns:
| Type | Description |
|---|---|
ExternalFunction
|
ExternalFunction configured for the conv2dk3 kernel. |
Raises:
| Type | Description |
|---|---|
ValueError
|
When |
Source code in python/iron/kernels/conv.py
conv2dk1_skip ¶
conv2dk1_skip(
input_width: int = 32,
input_channels: int = 64,
output_channels: int = 64,
act_dtype: type = int8,
) -> ExternalFunction
1x1 convolution kernel with skip (residual) connection.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
input_width
|
int
|
Spatial width of the input. |
32
|
input_channels
|
int
|
Number of input channels. |
64
|
output_channels
|
int
|
Number of output channels. |
64
|
act_dtype
|
type
|
Activation data type ( |
int8
|
Returns:
| Type | Description |
|---|---|
ExternalFunction
|
ExternalFunction configured for the conv2dk1_skip kernel. |
Raises:
| Type | Description |
|---|---|
ValueError
|
When |
Note
The activations are uint8 in two half-channel tensors whatever
act_dtype, which types the residual (skip) only. The generic
harness packs tensors of one type into one fifo, so an int8
residual streams beside the uint8 activations in a second fifo.
Source code in python/iron/kernels/conv.py
conv2dk1_i8 ¶
conv2dk1_i8(
input_width: int = 32,
input_channels: int = 64,
output_channels: int = 64,
) -> ExternalFunction
1x1 convolution kernel with int8 activations/weights/output.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
input_width
|
int
|
Spatial width of the input. |
32
|
input_channels
|
int
|
Number of input channels. |
64
|
output_channels
|
int
|
Number of output channels. |
64
|
Returns:
| Type | Description |
|---|---|
ExternalFunction
|
ExternalFunction configured for the conv2dk1_i8 kernel. |
Source code in python/iron/kernels/conv.py
conv2dk14 ¶
conv2dk14(
input_width: int = 224,
input_channels: int = 16,
output_channels: int = 16,
kernel_width: int = 14,
) -> ExternalFunction
14x14 convolution kernel.
The source lives under aie_kernels/conv/ and builds for aie2 as
well, where the vector path has its own AIE2 variant.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
input_width
|
int
|
Spatial width of the input. |
224
|
input_channels
|
int
|
Number of input channels. |
16
|
output_channels
|
int
|
Number of output channels. |
16
|
kernel_width
|
int
|
Width (and height) of the convolution kernel. |
14
|
Returns:
| Type | Description |
|---|---|
ExternalFunction
|
ExternalFunction configured for the conv2dk14 kernel. |
Source code in python/iron/kernels/conv.py
conv2dk1_skip_init ¶
conv2dk1_skip_init(
input_width: int = 32,
input_channels: int = 64,
output_channels: int = 64,
act_dtype: type = int8,
skip_input_channels: int | None = None,
) -> ExternalFunction
1x1 convolution kernel with skip-init connection.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
input_width
|
int
|
Spatial width of the input. |
32
|
input_channels
|
int
|
Number of input channels. |
64
|
output_channels
|
int
|
Number of output channels. |
64
|
act_dtype
|
type
|
Activation data type ( |
int8
|
skip_input_channels
|
int | None
|
Number of input channels for the skip-projection
1x1 conv whose weights are concatenated after the main conv
weights in the same buffer. Defaults to |
None
|
Returns:
| Type | Description |
|---|---|
ExternalFunction
|
ExternalFunction configured for the conv2dk1_skip_init kernel. |
Raises:
| Type | Description |
|---|---|
ValueError
|
When |
Source code in python/iron/kernels/conv.py
755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 811 812 813 814 815 816 817 818 819 820 821 822 823 824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 | |
bn_conv2dk1_relu_ref ¶
Numpy reference for bn_conv2dk1_relu.
The layouts of conv2dk1_ref, rounding
half to even: out = sat_u8(round_even(sum_ic x * w, scale)).
Source code in python/iron/kernels/conv.py
bn_conv2dk1_i8_ref ¶
Numpy reference for bn_conv2dk1_i8.
uint8 activations, int8 output:
out = sat_i8(round_even(sum_ic x * w, scale)).
Source code in python/iron/kernels/conv.py
bn_conv2dk1_skip_ref ¶
bn_conv2dk1_skip_ref(
x,
weights,
skip,
input_width,
input_channels,
output_channels,
scale,
skip_scale,
)
Numpy reference for bn_conv2dk1_skip.
skip is an [OC/8][W][8] line of either signedness:
Both shifts must be at least 1.
Source code in python/iron/kernels/conv.py
bn_conv2dk3_ref ¶
bn_conv2dk3_ref(
line0,
line1,
line2,
weights,
input_width,
input_channels,
output_channels,
kernel_width,
kernel_height,
check,
scale,
channel_offset,
)
Numpy reference for bn_conv2dk3: 3x3 stride-2 conv.
The layouts and check of conv2dk3_ref;
output x reads input pixels 2x-1 .. 2x+1 and the output line is
input_width / 2 wide: out = sat_u8(round_even(sum, scale)).
Source code in python/iron/kernels/conv.py
bn_conv2dk3_dw_ref ¶
bn_conv2dk3_dw_ref(
line0,
line1,
line2,
weights,
input_width,
input_channels,
output_channels,
kernel_width,
kernel_height,
check,
scale,
channel_offset,
*,
stride: int = 1
)
Numpy reference for bn_conv2dk3_dw: depthwise 3x3 + ReLU.
Lines are [C/8][W][8] uint8, weights [C/8][3 rows][3][c8],
check as in conv2dk3_ref:
out = sat_u8(round_even(sum, scale)) over output_channels
channels and input_width / stride pixels. kernel_width,
kernel_height and channel_offset are accepted for the signature.
Source code in python/iron/kernels/conv.py
bn_conv2dk3_dw_out_split_ref ¶
bn_conv2dk3_dw_out_split_ref(
line0,
line1,
line2,
weights,
input_width,
input_channels,
output_channels,
kernel_width,
kernel_height,
check,
scale,
channel_offset,
)
Numpy reference for bn_conv2dk3_dw_out_split.
The stride-1 bn_conv2dk3_dw_ref,
its output channels split in two halves, one per output.
Source code in python/iron/kernels/conv.py
bn_conv2dk1_relu_xy_pool_padded_ref ¶
bn_conv2dk1_relu_xy_pool_padded_ref(
x,
weights,
input_width,
input_channels,
output_channels,
output_channels_padd,
scale,
y_index,
output_split,
weight_index,
)
Numpy reference for bn_conv2dk1_relu_xy_pool_padded into a zeroed output.
Each output channel of this call's slice (output_channels /
output_split of them, at slice weight_index) sums
sat_u8(round_even(sum_ic x * w, scale)) over the input_width
pixels. On the last row (y_index == input_width - 1) the sum is
averaged over 49 pixels in float32 as the kernel does: an average
whose first decimal is 5 rounds to even, any other rounds half up.
Channels outside the slice stay 0; output_channels_padd must not
exceed output_channels.
Source code in python/iron/kernels/conv.py
bn_fc_relu_ui16_pad_ref ¶
bn_fc_relu_ui16_pad_ref(
x,
weights,
input_width,
input_channels,
input_channels_pad,
output_channels,
scale,
)
Numpy reference for bn_fc_relu_ui16_pad.
A 1x1 conv of uint16 activations whose weights are laid out for
input_channels_pad input channels ([OC/8][ICP/8][ic8][oc8], the
first input_channels used); the output is uint16 holding
sat_u8(round_even(sum_ic x * w, scale)).
Source code in python/iron/kernels/conv.py
bn_conv2dk1_relu ¶
bn_conv2dk1_relu(
input_width: int = 32,
input_channels: int = 64,
output_channels: int = 64,
) -> ExternalFunction
Bottleneck 1x1 conv + ReLU kernel (int8 in, uint8 out).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
input_width
|
int
|
Spatial width of the input. |
32
|
input_channels
|
int
|
Number of input channels. |
64
|
output_channels
|
int
|
Number of output channels. |
64
|
Returns:
| Type | Description |
|---|---|
ExternalFunction
|
ExternalFunction configured for the bn_conv2dk1_relu kernel. |
Source code in python/iron/kernels/conv.py
bn_conv2dk3 ¶
bn_conv2dk3(
input_width: int = 32,
input_channels: int = 64,
output_channels: int = 64,
weight_output_channels: int | None = None,
) -> ExternalFunction
Bottleneck 3x3 conv with stride-2 kernel (int8 in, uint8 out).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
input_width
|
int
|
Spatial width of the input. |
32
|
input_channels
|
int
|
Number of input channels. |
64
|
output_channels
|
int
|
Number of output channels produced by this call. |
64
|
weight_output_channels
|
int | None
|
Total number of output channels stored in the
weights buffer, as for |
None
|
Returns:
| Type | Description |
|---|---|
ExternalFunction
|
ExternalFunction configured for the bn_conv2dk3 kernel. |
Source code in python/iron/kernels/conv.py
bn_conv2dk1_i8 ¶
bn_conv2dk1_i8(
input_width: int = 32,
input_channels: int = 64,
output_channels: int = 64,
) -> ExternalFunction
Bottleneck 1x1 conv kernel (uint8 in, int8 out).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
input_width
|
int
|
Spatial width of the input. |
32
|
input_channels
|
int
|
Number of input channels. |
64
|
output_channels
|
int
|
Number of output channels. |
64
|
Returns:
| Type | Description |
|---|---|
ExternalFunction
|
ExternalFunction configured for the bn_conv2dk1_i8 kernel. |
Source code in python/iron/kernels/conv.py
bn_conv2dk1_skip ¶
bn_conv2dk1_skip(
input_width: int = 32,
input_channels: int = 64,
output_channels: int = 64,
skip_dtype: type = uint8,
) -> ExternalFunction
Bottleneck 1x1 conv with skip connection (uint8 in).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
input_width
|
int
|
Spatial width of the input. |
32
|
input_channels
|
int
|
Number of input channels. |
64
|
output_channels
|
int
|
Number of output channels. |
64
|
skip_dtype
|
type
|
Skip connection data type ( |
uint8
|
Returns:
| Type | Description |
|---|---|
ExternalFunction
|
ExternalFunction configured for the bn_conv2dk1_skip kernel. |
Raises:
| Type | Description |
|---|---|
ValueError
|
When |
Source code in python/iron/kernels/conv.py
bn_conv2dk3_dw ¶
bn_conv2dk3_dw(
input_width: int = 32,
input_channels: int = 64,
output_channels: int = 64,
stride: int = 1,
) -> ExternalFunction
Bottleneck depthwise 3x3 conv + ReLU kernel (uint8 in/out).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
input_width
|
int
|
Spatial width of the input. |
32
|
input_channels
|
int
|
Number of input channels. |
64
|
output_channels
|
int
|
Number of output channels. |
64
|
stride
|
int
|
Convolution stride (1 or 2). |
1
|
Returns:
| Type | Description |
|---|---|
ExternalFunction
|
ExternalFunction configured for the bn_conv2dk3_dw kernel. |
Raises:
| Type | Description |
|---|---|
ValueError
|
When |
Source code in python/iron/kernels/conv.py
bn_conv2dk1_relu_xy_pool_padded ¶
bn_conv2dk1_relu_xy_pool_padded(
input_width: int = 7,
input_channels: int = 80,
output_channels: int = 1280,
weight_chunk_count: int | None = None,
) -> ExternalFunction
Fused 1x1 conv + ReLU + xy-pool with channel padding (int8 in, uint16 out).
A post-stage kernel that fuses a pointwise (1x1) convolution, ReLU activation, and global xy avg-pool into a single pass, with output channels padded to a DMA-friendly multiple. Sized for MobileNet V3's post-bottleneck stage where the final 1x1 expand-conv collapses the 7x7 feature map into a 1x1 vector.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
input_width
|
int
|
Spatial width of the input. |
7
|
input_channels
|
int
|
Number of input channels. |
80
|
output_channels
|
int
|
Logical output channels (e.g. 1280). Sets both
the output buffer length AND, when |
1280
|
weight_chunk_count
|
int | None
|
Override the weight buffer's element count when
the design streams weights in chunks (cascade/output-split).
|
None
|
Returns:
| Type | Description |
|---|---|
ExternalFunction
|
ExternalFunction configured for the fused conv+relu+xy_pool kernel. |
Source code in python/iron/kernels/conv.py
bn_conv2dk1_partial_put_i8 ¶
bn_conv2dk1_partial_put_i8(
input_width: int = 7,
input_channels: int = 80,
weight_count: int = 4800,
*,
block_index: int = 13
) -> ExternalFunction
Cascade-PUT half of a width-split 1x1 conv on int8 activations.
The PUT tile of a two-tile cascade-split pointwise conv: consumes a
width slice of the activation, multiplies against its weight half,
and emits the partial sum onto the cascade stream (no separate
output buffer — cascade-only). Sister of
bn_conv2dk1_partial_get_relu_i8.
Currently defined in the .cc only for MobileNet V3's bn13 / bn14
(one wrapper symbol per block); block_index selects which.
Generalising this to arbitrary block names would require adding a
non-prefixed wrapper to bn_conv2dk1_i8.cc.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
input_width
|
int
|
Spatial width of the input slice. |
7
|
input_channels
|
int
|
Number of input channels. |
80
|
weight_count
|
int
|
Per-call weight chunk size in elements (the design streams weights in chunks; full weight tensor is shared across multiple kernel invocations). |
4800
|
block_index
|
int
|
|
13
|
Returns:
| Type | Description |
|---|---|
ExternalFunction
|
ExternalFunction configured for the PUT tile. |
Raises:
| Type | Description |
|---|---|
ValueError
|
When |
Source code in python/iron/kernels/conv.py
bn_conv2dk1_partial_get_relu_i8 ¶
bn_conv2dk1_partial_get_relu_i8(
input_width: int = 7,
input_channels: int = 80,
output_channels: int = 480,
weight_count: int = 4800,
*,
block_index: int = 13
) -> ExternalFunction
Cascade-GET half of a width-split 1x1 conv + ReLU on int8 activations.
The GET tile of a two-tile cascade-split pointwise conv: consumes
the cascade partial sum from its sister PUT tile, finishes the dot
product against its weight half, applies ReLU, and writes the full
output buffer. Sister of bn_conv2dk1_partial_put_i8.
Currently defined in the .cc only for MobileNet V3's bn13 / bn14
(one wrapper symbol per block); block_index selects which.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
input_width
|
int
|
Spatial width of the input slice. |
7
|
input_channels
|
int
|
Number of input channels. |
80
|
output_channels
|
int
|
Number of output channels (full L1 output width). |
480
|
weight_count
|
int
|
Per-call weight chunk size in elements. |
4800
|
block_index
|
int
|
|
13
|
Returns:
| Type | Description |
|---|---|
ExternalFunction
|
ExternalFunction configured for the GET tile. |
Raises:
| Type | Description |
|---|---|
ValueError
|
When |
Source code in python/iron/kernels/conv.py
bn_conv2dk3_dw_out_split ¶
bn_conv2dk3_dw_out_split(
input_width: int = 7,
input_channels: int = 480,
output_split_channels: int = 240,
*,
block_index: int = 13
) -> ExternalFunction
Depthwise 3x3 stride-1 conv with split output stream (uint8 in/out).
A variant of bn_conv2dk3_dw (stride=1) that writes its output
to TWO separate buffers — the channel dimension is split in half so
downstream cascade-PUT tiles can each consume one slice. Used by
MobileNet V3's bn13 / bn14 depthwise stage to feed the L3 cascade.
Currently defined in the .cc only via per-block extern wrappers
(BN13 or BN14 macro picks the symbol prefix); block_index selects
which.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
input_width
|
int
|
Spatial width of the input. |
7
|
input_channels
|
int
|
Number of input channels (== output channels — depthwise). |
480
|
output_split_channels
|
int
|
Channels per output slice (half of
|
240
|
block_index
|
int
|
|
13
|
Returns:
| Type | Description |
|---|---|
ExternalFunction
|
ExternalFunction configured for the split-output DW kernel. |
Raises:
| Type | Description |
|---|---|
ValueError
|
When |
Source code in python/iron/kernels/conv.py
1472 1473 1474 1475 1476 1477 1478 1479 1480 1481 1482 1483 1484 1485 1486 1487 1488 1489 1490 1491 1492 1493 1494 1495 1496 1497 1498 1499 1500 1501 1502 1503 1504 1505 1506 1507 1508 1509 1510 1511 1512 1513 1514 1515 1516 1517 1518 1519 1520 1521 1522 1523 1524 1525 1526 1527 1528 1529 1530 1531 1532 1533 1534 | |
bn_conv2dk1_input_split_partial_put_ui8 ¶
bn_conv2dk1_input_split_partial_put_ui8(
input_width: int = 7,
input_channels: int = 240,
weight_count: int = 9600,
*,
block_index: int = 13
) -> ExternalFunction
Input-split cascade-PUT half of a 1x1 conv on uint8 activations.
Like bn_conv2dk1_partial_put_i8 but consumes a CHANNEL slice
(input-split) of a uint8 activation instead of a width slice of int8.
Used by MobileNet V3's bn13 / bn14 L3 stage.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
input_width
|
int
|
Spatial width of the input slice. |
7
|
input_channels
|
int
|
Number of input channels (one half of the full input after split). |
240
|
weight_count
|
int
|
Per-call weight chunk size in elements. |
9600
|
block_index
|
int
|
|
13
|
Returns:
| Type | Description |
|---|---|
ExternalFunction
|
ExternalFunction configured for the input-split PUT tile. |
Raises:
| Type | Description |
|---|---|
ValueError
|
When |
Source code in python/iron/kernels/conv.py
bn_conv2dk1_input_split_partial_skip_get ¶
bn_conv2dk1_input_split_partial_skip_get(
input_width: int = 7,
input_channels: int = 240,
output_channels: int = 80,
weight_count: int = 9600,
*,
block_index: int = 13
) -> ExternalFunction
Input-split cascade-GET half of a 1x1 conv + skip-add (uint8 in, int8 out).
The GET tile completes the cascade-split 1x1 + ReLU + residual add
pattern: consumes the partial sum from its sister PUT tile, finishes
the dot product, adds a skip row of int8 activations, and writes int8
output. Sister of bn_conv2dk1_input_split_partial_put_ui8.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
input_width
|
int
|
Spatial width of the input slice. |
7
|
input_channels
|
int
|
Number of input channels (one half after split). |
240
|
output_channels
|
int
|
Final output channels. |
80
|
weight_count
|
int
|
Per-call weight chunk size in elements. |
9600
|
block_index
|
int
|
|
13
|
Returns:
| Type | Description |
|---|---|
ExternalFunction
|
ExternalFunction configured for the input-split skip-GET tile. |
Raises:
| Type | Description |
|---|---|
ValueError
|
When |
Source code in python/iron/kernels/conv.py
bn_fc_relu_ui16_pad ¶
bn_fc_relu_ui16_pad(
input_channels: int = 1280,
output_channels: int = 16,
weight_chunk_count: int | None = None,
) -> ExternalFunction
Fully-connected layer (1x1 conv on (1,1,C)) + ReLU, uint16 in/out, with padding.
A post-stage FC kernel used by MobileNet V3's classifier head. Input is
a (1,1,input_channels) feature vector held as uint16; output is
output_channels uint16 logits. Weights stored in a padded layout
(the input_channels_pad runtime arg selects the actual stride).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
input_channels
|
int
|
Number of input channels (e.g. 1280). |
1280
|
output_channels
|
int
|
Number of output channels per call (slice width, since the full FC is split across multiple tiles). |
16
|
weight_chunk_count
|
int | None
|
Override the weight buffer's element count when
the design streams weights in chunks (cascade/ping-pong).
|
None
|
Returns:
| Type | Description |
|---|---|
ExternalFunction
|
ExternalFunction configured for the post-L2 FC kernel. |
Source code in python/iron/kernels/conv.py
Activation functions¶
Activation kernel factories and NumPy reference implementations.
softmax ¶
Softmax activation kernel for bf16 tiles.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tile_size
|
int
|
Number of elements per tile, passed at run time; a positive multiple of 32, the kernel's vector step. |
1024
|
Returns:
| Type | Description |
|---|---|
ExternalFunction
|
ExternalFunction configured for the softmax kernel. |
Source code in python/iron/kernels/activation.py
gelu ¶
GELU activation kernel (tanh approximation) for bf16 tiles (must be 1024).
Source code in python/iron/kernels/activation.py
silu ¶
SiLU (Swish) activation kernel for bf16 tiles (must be 1024).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tile_size
|
int
|
Elements per call (must be 1024). |
1024
|
use_lut
|
bool
|
Compute tanh from the interpolated LUT rather than AIE2P's vtanh instruction, which makes this 7.5x closer to the true function and judged against an exact model of it. Moot on aie2. |
False
|
Source code in python/iron/kernels/activation.py
silu_sized ¶
SiLU (Swish) for bf16 tiles, with a compiled-in element count.
Runtime-size sibling of silu; design
keeps the (in, out, size) ABI. Positive whole vectors are required
(16 on aie2, 32 on aie2p).
Source code in python/iron/kernels/activation.py
gelu_sized ¶
GELU (tanh approx) for bf16 tiles, with a compiled-in element count.
Runtime-size sibling of gelu; design
keeps the (in, out, size) ABI. Positive whole vectors only: multiples
of 16 on aie2 or 32 on aie2p.
Source code in python/iron/kernels/activation.py
swiglu ¶
SwiGLU gated activation kernel for bf16 tiles (must be 1024).
out = (x * w1) * silu(x * w2); see swiglu_ref.
Source code in python/iron/kernels/activation.py
bf16_exp ¶
Element-wise exponential kernel for bf16 tiles (must be 1024).
Computes exp(clip(x, -88, 88)): the kernel saturates rather than
overflowing for real inputs, including infinities. On AIE2P a
range-reduced polynomial and integer exponent reconstruction replace
the hardware exp2 approximation, preserving subnormal outputs. See
bf16_exp_ref for why that
clamp matches the AIE2 table's domain.
Source code in python/iron/kernels/activation.py
exp2f_vec ¶
Software f32 2**x kernel: a degree-5 minimax poly, not a LUT.
A float32-output alternative to bf16_exp, sharing its AIE2P range-reduced
polynomial but with a separately configurable input domain. See
aie_kernels/activation/exp2f_vec.cc for the accuracy rationale and the
noinline codegen hazard this kernel carries.
The same source builds for aie2.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tile_size
|
int
|
Number of elements per tile; must be a multiple of 16 (the kernel's vector width). |
1024
|
min_x
|
float
|
Input is clamped to this before evaluation. The default
-111 is the lowest exponent that still holds the kernel's
8.9e-5 relative error; -126 is the hard floor (one f32
exponent field), reachable at up to 6.5e-3. See
|
-111.0
|
Returns:
| Type | Description |
|---|---|
ExternalFunction
|
ExternalFunction configured for the exp2f_vec kernel. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If tile_size is not a multiple of 16, or min_x is below -126. |
Source code in python/iron/kernels/activation.py
tanh ¶
Tanh for bf16 tiles of a positive multiple of 32 elements.
The count is compiled in; retain
tile_size as a trailing int argument (e.g. via
transform_parallel(pass_size_to_kernel=True)).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tile_size
|
int
|
Elements per call (a positive multiple of 32). |
1024
|
use_lut
|
bool
|
Compute tanh from the interpolated LUT rather than AIE2P's
vtanh instruction. Moot on aie2, which only has the LUT. See
|
False
|
Source code in python/iron/kernels/activation.py
sigmoid ¶
Sigmoid for bf16 tiles of a positive multiple of 32 elements.
The count is compiled in; retain tile_size as a trailing ABI argument.
Source code in python/iron/kernels/activation.py
leaky_relu ¶
Leaky ReLU for bf16 tiles of at least 64 elements, in multiples of 32.
The count is compiled in, but the ABI retains (tile_size, alpha) as
trailing int/bfloat16 arguments. The slope remains runtime-valued.
Source code in python/iron/kernels/activation.py
relu_ref ¶
Numpy reference for a ReLU kernel — element-wise max(x, 0).
Exact; tolerance comparison is not needed. See aie.utils.verify
for the relaxed bf16/LUT-style comparators most kernels here want.
Source code in python/iron/kernels/activation.py
silu_ref ¶
Numpy reference for silu (Swish) — x * sigmoid(x).
LUT-approximation territory; pair with rtol=0.128 (the default
in count_mismatches) when verifying.
Source code in python/iron/kernels/activation.py
gelu_ref ¶
Numpy reference for gelu.
Tanh approximation 0.5 * x * (1 + tanh(sqrt(2/pi) * (x + 0.044715 * x^3))).
Matches the C++ kernel's tanh-GELU formula. It is evaluated in float64:
in float32, 1 + tanh cancels for x below about -4.5 and leaves
values over ten times too large.
Source code in python/iron/kernels/activation.py
bf16_exp_lut_ref ¶
Model of getExpBf16, the LUT exponential the bf16_exp kernel uses.
The kernel clamps to +/-_EXP_BF16_CLAMP, converts to Q8 with a floor
(bfloat16_to_int(x, 8)), then reads the byte halves of that fixed-point
key as two table indices and multiplies: exp(x) = exp(int) * exp(frac).
Both tables hold exactly bfloat16(exp(.)), so they are written here as
that rule rather than as 512 opaque floats -- the unreachable middle of the
integer table (keys the clamp cannot produce) is the only part that is not
an exponential, and it is never read.
The product is exact in f32 (two bf16 operands), so the model differs from
the device only where the hardware flushes the one subnormal table entry,
bfloat16(exp(-88)).
Source code in python/iron/kernels/activation.py
sigmoid_lut_ref ¶
Model of sigmoid built with use_lut=True.
Follows activation/sigmoid.cc step for step: x/2 is exact (0.5 is a power
of two), the accumulator overload of tanh_bf16_v16 narrows to bf16
before the table, and the +1 and *0.5 stay in the accumulator so
there is a single store rounding at the end.
Source code in python/iron/kernels/activation.py
silu_lut_ref ¶
Model of silu built with use_lut=True.
activation/silu.cc narrows the sigmoid factor to bf16 before the final multiply, so that rounding is modelled too, not folded away. The sigmoid is exactly 0 from x = -8 down, and x is clamped there before the multiply, so -inf gives 0 rather than NaN.
Source code in python/iron/kernels/activation.py
swiglu_lut_ref ¶
Model of swiglu built with use_lut=True.
activation/swiglu.cc narrows after every multiply -- x*w1, x*w2, the
sigmoid factor and the silu product each land in a bf16 register before
the next step -- which is what this reproduces. x*w2 is clamped at -8
before its multiply, as in silu_lut_ref, and where the silu product is 0
the output is 0, so an overflowed x*w1 does not make inf * 0.
Source code in python/iron/kernels/activation.py
tanh_lut_ref ¶
Numpy model of getTanhBf16, the interpolated-LUT tanh.
The kernel clamps x to the table's range [-4, 4 - 1/64], then evaluates
slope[e] * x + offset[e] for e = floor(4x) + 16: 32 segments of
width 0.25 over [-4, 4). The end segments are the constants -1 and +1,
so the clamp changes no finite result, and +-inf gives +-1 rather than
0 * inf. The product is exact in f32 (bf16 carries 8 mantissa bits and
8 + 8 < 24), so the only rounding is the accumulator's store back to bf16,
which is why the build using this is judged at one ulp rather than a
percentage.
This is what tanh computes with
use_lut=True, and what it always computes on aie2. The default aie2p
build uses the vtanh instruction instead, which is a coarser
approximation with no published spec, so it is judged against
tanh_ref and a measured bound.
Source code in python/iron/kernels/activation.py
tanh_ref ¶
Numpy reference for tanh — element-wise tanh(x).
LUT/native-approximation territory; pair with rtol=0.128 when verifying.
Source code in python/iron/kernels/activation.py
sigmoid_ref ¶
Numpy reference for sigmoid — 1 / (1 + exp(-x)).
LUT-approximation territory; pair with rtol=0.128 when verifying.
Source code in python/iron/kernels/activation.py
leaky_relu_ref ¶
Numpy reference for leaky_relu.
x if x > 0 else alpha * x. alpha must match the slope the design
passes to the kernel at runtime. Exact up to bf16 rounding; pair with a
small rtol when verifying.
Source code in python/iron/kernels/activation.py
swiglu_ref ¶
Numpy reference for swiglu: (x * w1) * silu(x * w2).
swiglu.cc forms the two products in bf16, then silu of the second
through the tanh LUT (0.5 * (1 + tanh(z / 2))). The reference rounds
the two products to bf16 as the kernel does and computes the rest in
float32; LUT-approximation territory, pair with rtol=0.128.
Source code in python/iron/kernels/activation.py
bf16_exp_ref ¶
Numpy reference for bf16_exp — element-wise exp(x).
LUT approximation territory; pair with the canonical 12.8% relative
tolerance and stop_at_nonfinite=True (the default in
count_mismatches) when verifying.
exp(clip(x, -88, 88)), not plain exp(x): the kernel clamps to
EXP_BF16_CLAMP before its Q8 fixed-point table lookup (see
aie_runtime_lib/AIE2/lut_based_ops.h), so it saturates rather than
overflowing. +88 is the largest value the tables carry.
exp(-88) is a nonzero bf16 subnormal (about 6.06e-39), not
zero: AIE2P preserves it through integer exponent reconstruction,
whereas the AIE2 LUT may flush the tail under the absolute tolerance.
The clamp also keeps the reference itself in range:
exp(88) = 1.65e+38 fits float32 where exp(89) would not.
Source code in python/iron/kernels/activation.py
exp2f_vec_ref ¶
Numpy reference for exp2f_vec: exact 2**x.
Unlike the LUT-based refs above, this is float64 2**x (not a
reimplementation of the on-device poly): the kernel targets ~8.9e-5
relative error by design, several orders tighter than the LUT-based
kernels' 12.8% default, so pair with a correspondingly tight
tolerance (e.g. rtol=1e-3) rather than the LUT default.
The kernel clamps its input to min_x before evaluating (see the
factory's min_x), so the reference does the same: 2**-5000 is
2**min_x on the device, not zero. Pass the factory's min_x.
Source code in python/iron/kernels/activation.py
softmax_ref ¶
Numpy reference for softmax.
The AIE kernel computes softmax independently per tile_size-element
tile (no cross-tile reduction), so the reference splits x the same
way before applying the float32 softmax. x.size must be a
multiple of tile_size.
Source code in python/iron/kernels/activation.py
Normalization¶
rms_norm, rms_norm_eps, and layer_norm support both aie2 and aie2p.
They accept tile_size (default 1024) or its compatibility alias cols.
rope in Data movement has the same size API and supports both interleaved
and two_halves=True layouts. The transformer module re-exports the canonical
bf16 norm and RoPE factories and references; importing through either module
does not select a different implementation. Norm references accept eps.
Normalization kernel factories + numpy references: rms_norm, layer_norm.
rms_norm ¶
RMS-norm a bf16 row on aie2/aie2p; (in, out, cols), eps=1e-5.
cols is a compatibility alias for tile_size. Positive row lengths
need not be vector-aligned: the kernel handles scalar tails.
Source code in python/iron/kernels/norm.py
rms_norm_eps ¶
RMS-norm a bf16 row (gamma=1); design passes (in, out, cols, epsilon).
Source code in python/iron/kernels/norm.py
layer_norm ¶
Layer-norm a bf16 row; (in, out, cols), gamma=1, beta=0, eps=1e-5.
cols aliases tile_size, a positive multiple of 16 on aie2 or 32 on
aie2p (the source processes whole vectors, without a scalar tail).
Source code in python/iron/kernels/norm.py
rms_norm_ref ¶
Numpy reference for rms_norm: x / sqrt(mean(x**2) + eps).
Source code in python/iron/kernels/norm.py
layer_norm_ref ¶
Numpy reference for layer_norm: (x - mean) / sqrt(var + eps).
Source code in python/iron/kernels/norm.py
Transformer blocks¶
Transformer building blocks: rms_norm, layer_norm (bf16, f32, affine+cast), rope, mm_activation_epilogue.
The bf16 norms and RoPE are re-exported from norm and datamovement;
they support aie2 and aie2p, with cols as an alias for tile_size.
The activation epilogue and the f32/affine norms take their aie2p source on
both generations. Each processes one row
(cols elements) per call; the row length is a scalar Param the
factory binds to cols. These are the kernels
programming_examples/ml/{norm,rope,mm_activation_epilogue} build.
layer_norm_f32 ¶
Row-wise LayerNorm on float32 in and out (gamma = 1, beta = 0, eps 1e-5).
A separate factory rather than a dtype of
layer_norm: this one is held to atol 1e-3
(2e-6 on aie2) instead of the bf16 tolerance, which its reference meets
only by computing the variance two-pass in float64. Merging them would
put that numerical difference behind a dtype switch.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
cols
|
int
|
Elements per row (multiple of 16). |
4096
|
Source code in python/iron/kernels/transformer.py
layer_norm_affine_cast ¶
Row-wise LayerNorm, f32 in, per-column gamma/beta, bf16 out.
The second argument holds gamma (cols values) followed by beta
(cols values) as float32: a tensor Param, which the generic
builder bakes into a core buffer.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
cols
|
int
|
Elements per row (multiple of 16). |
4096
|
Source code in python/iron/kernels/transformer.py
mm_activation_epilogue ¶
GEMM epilogue on float32 rows: identity (0), SiLU (1), tanh-GELU (2) or ReLU (3) by mode.
One resident kernel whose mode is a runtime argument, so a design can
switch activations without recompiling
(programming_examples/ml/mm_activation_epilogue).
AIE2 has no tanh instruction, so there SiLU and GELU read getTanhBf16's table, and its tuned build is judged against a model of that arithmetic instead of the true functions.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tile_size
|
int
|
Elements per call (multiple of 16). |
1024
|
Source code in python/iron/kernels/transformer.py
layer_norm_f32_ref ¶
Numpy reference for layer_norm_f32.
Centered two-pass variance in float64 so it stays exact on the non-zero-mean input the f32 kernel is exercised with.
Source code in python/iron/kernels/transformer.py
layer_norm_affine_cast_ref ¶
Numpy reference for layer_norm_affine_cast.
gamma_beta is gamma then beta, each cols float32 values.
Source code in python/iron/kernels/transformer.py
mm_activation_epilogue_ref ¶
Numpy reference for mm_activation_epilogue.
mode 0 identity, 1 x * sigmoid(x), 2 tanh-approximation GELU,
3 max(x, 0).
Source code in python/iron/kernels/transformer.py
mm_activation_epilogue_lut_ref ¶
Model of mm_activation_epilogue on aie2.
Follows mm_activation_epilogue.cc's roundings around getTanhBf16
(tanh_lut_ref). SiLU splits
x into hi, its top 16 bits, and lo, bf16(x - hi), and
multiplies each by the bf16 sigmoid (bf16(t + 1)) / 2, where t is
the table's tanh of bf16(x) / 2 narrowed to bf16. hi is finite for
any finite x, so huge inputs give about x or 0 rather than NaN;
+-inf still gives NaN. GELU runs in bf16: x, x * x and the inner
polynomial are each rounded before the next step, and the output is
bf16(x / 2) * bf16(t + 1), all reading x clamped at -8 so -inf
gives 0. Both return +0 where IEEE arithmetic gives -0, as the accumulator
does. Identity and ReLU are exact.
Source code in python/iron/kernels/transformer.py
Vision¶
Vision kernel factories: color conversion, threshold, filter2d, add_weighted.
rgba2hue ¶
Convert a line of RGBA pixels to hue values (full-range, 0..255).
Source code in python/iron/kernels/vision.py
threshold ¶
threshold(
line_width: int = 1920,
dtype: type = uint8,
use_chess: bool = False,
) -> ExternalFunction
Apply a threshold operation to a line of pixels.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
line_width
|
int
|
Number of elements per line. |
1920
|
dtype
|
type
|
Element data type, |
uint8
|
use_chess
|
bool
|
When |
False
|
Raises:
| Type | Description |
|---|---|
ValueError
|
When |
Source code in python/iron/kernels/vision.py
bitwise_or ¶
bitwise_or(
line_width: int = 1920,
dtype: type = uint8,
use_chess: bool = False,
) -> ExternalFunction
Element-wise bitwise OR of two lines.
Source code in python/iron/kernels/vision.py
bitwise_and ¶
bitwise_and(
line_width: int = 1920,
dtype: type = uint8,
use_chess: bool = False,
) -> ExternalFunction
Element-wise bitwise AND of two lines.
Source code in python/iron/kernels/vision.py
gray2rgba ¶
Convert a grayscale line to RGBA.
Source code in python/iron/kernels/vision.py
rgba2gray ¶
Convert an RGBA line to grayscale.
Source code in python/iron/kernels/vision.py
filter2d ¶
Apply a 3x3 2D convolution filter across three input lines.
Source code in python/iron/kernels/vision.py
add_weighted ¶
add_weighted(
line_width: int = 1920,
dtype: type = uint8,
use_chess: bool = False,
) -> ExternalFunction
Weighted addition of two lines with a gamma offset.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
line_width
|
int
|
Number of elements per line. |
1920
|
dtype
|
type
|
Element data type, |
uint8
|
use_chess
|
bool
|
When |
False
|
Raises:
| Type | Description |
|---|---|
ValueError
|
When |
Source code in python/iron/kernels/vision.py
gray2rgba_ref ¶
Numpy reference for gray2rgba: (y, y, y, 255) per pixel.
Source code in python/iron/kernels/vision.py
rgba2gray_ref ¶
Numpy reference for rgba2gray: fixed-point BT.470 luma.
Y = (9798 R + 19235 G + 3736 B + 2**14) >> 15, saturated to uint8;
alpha is ignored. Within one LSB of the kernel (rounding of the final
shift).
Source code in python/iron/kernels/vision.py
rgba2hue_ref ¶
Numpy reference for rgba2hue: full-range hue.
rgba2hue.cc multiplies by a Q7.9 reciprocal rather than dividing, so
with d = max - min of R, G, B and inv = 85 * 512 / d the hue is
(offset * 512 + c * inv) >> 10 for whichever channel holds the max:
c = G - B at offset 1, B - R at 171, R - G at 341.
Each offset carries the + 1 that rounds the final halving, so there is
one rounding step rather than two. The cast to uint8 wraps, so a
negative hue (R max, G < B) comes out as 256 + h -- the right circular
value. Gray pixels (d == 0) are hue 0, and a max held by both G and R
goes to G, as the kernel's select order does. inv truncates, which
leaves hue up to one LSB below the exact value.
Source code in python/iron/kernels/vision.py
threshold_ref ¶
Numpy reference for threshold (OpenCV semantics).
ttype: 0 binary (x > thresh ? maxval : 0), 1 binary inverted,
2 truncate (min(x, thresh)), 3 to-zero (x > thresh ? x : 0),
4 to-zero inverted. Exact.
Source code in python/iron/kernels/vision.py
bitwise_or_ref ¶
Numpy reference for bitwise_or. Exact.
bitwise_and_ref ¶
Numpy reference for bitwise_and. Exact.
add_weighted_ref ¶
Numpy reference for add_weighted: Q2.14 blend.
out = sat(((alpha * a + beta * b) >> 14) + gamma) with alpha and
beta as Q2.14 fixed point (8192 is 0.5) and gamma in output
units, as OpenCV's addWeighted has it; the kernel reads gamma as
the data type, so for uint8 data -56 is 200. The vector path in
addWeighted.cc rounds the shift down; its scalar path rounds to
nearest, so the two differ by at most one LSB.
Source code in python/iron/kernels/vision.py
filter2d_ref ¶
Numpy reference for filter2d: 3x3 correlation of three lines.
Produces the middle line. The vector path keeps only the top byte of
each int16 coefficient (k >> 8 as int8) and shifts the sum by
4, so a Q4.12 kernel such as 4096 * [[0, 1, 0], [1, -4, 1], [0, 1, 0]]
lands on integer taps. Borders replicate the edge pixel; the result is
saturated to uint8. Within one LSB of the kernel.
Source code in python/iron/kernels/vision.py
Contracts and the generic design builder¶
Every factory the generic builder can build carries a KernelContract on the
returned ExternalFunction (fn.contract). The field-by-field account is in
Kernel Library;
the reference below is generated from the dataclass, so it cannot drift from
it. Multi-dtype factories publish the combinations they support as
factory.dtypes, and the factories above export their references as *_ref
functions (add_ref, reduce_max_ref, mm_ref, ...), so host code never
reimplements the math.
Shared helpers for the kernels submodules.
KernelContract
dataclass
¶
KernelContract(
roles: tuple[type, ...],
reference: (
Callable[..., ndarray | tuple[ndarray, ...]] | None
) = None,
tolerance: Tolerance | None = None,
ops_per_call: int | None = None,
out_valid: int | None = None,
sample: Callable[..., list] | None = None,
acc_dtype: type | None = None,
reduction: int | None = None,
setup: Callable[[], object] | None = None,
stack_bytes: int | None = None,
unsupported: str | None = None,
layouts: tuple[TensorLayout | None, ...] = (),
parameter_bindings: tuple[tuple[int, object], ...] = (),
initializers: tuple[tuple[int, Callable], ...] = (),
out_offset: tuple[int, int] | None = None,
trace: Trace | None = None,
uses_lut: bool = False,
)
What a kernel computes, declared next to the factory that builds it.
arg_types fixes each argument's shape and dtype; the contract adds
what types cannot say, so aie.iron.algorithms.kernel_design can
build, run and judge any factory from this one declaration.
Attributes:
| Name | Type | Description |
|---|---|---|
roles |
tuple[type, ...]
|
|
reference |
Callable[..., ndarray | tuple[ndarray, ...]] | None
|
The host implementation, and the arithmetic model (a
saturating kernel's reference clips). Called with every unbound
non-output argument in order: |
tolerance |
Tolerance | None
|
How close the device must come; |
ops_per_call |
int | None
|
Arithmetic operations per call; |
out_valid |
int | None
|
Meaningful leading elements of a DMA-padded output tile;
|
sample |
Callable[..., list] | None
|
|
acc_dtype |
type | None
|
The accumulator type, or |
reduction |
int | None
|
Terms summed into one output element per call; |
setup |
Callable[[], object] | None
|
A kernel to run once on the core first ( |
stack_bytes |
int | None
|
Core stack a Worker calling this kernel needs, when more than the target's default. Say where the number came from. |
unsupported |
str | None
|
Why the builder cannot run this kernel, or |
layouts |
tuple[TensorLayout | None, ...]
|
A |
parameter_bindings |
tuple[tuple[int, object], ...]
|
|
initializers |
tuple[tuple[int, Callable], ...]
|
|
out_offset |
tuple[int, int] | None
|
|
trace |
Trace | None
|
The |
uses_lut |
bool
|
Whether the kernel gathers through an |
Overflow, rounding and NaN handling are not declared twice: the reference is the arithmetic model and the tolerance the slack against it.
out_indices
property
¶
Output argument positions, in declaration order.
out_index
property
¶
Position of the one output, written (Out) or accumulated into (InOut).
Raises for a kernel with several outputs: code that must handle any
kernel reads out_indices.
accumulates
property
¶
Whether the kernel reads its output back (InOut), as C += A * B does.
reference_indices ¶
Argument positions handed to reference, in order.
An InOut output is excluded like an Out one: the reference
computes the result from the declared initializer's state. The
builder initializes the buffer before each independent tile call.
Source code in python/iron/kernels/_common.py
validate_types ¶
Validate contracts against NumPy tensor aliases and scalar dtypes.
Raw MLIR types remain usable by ExternalFunction, but the host contract requires NumPy declarations for sampling, layouts and references.
Source code in python/iron/kernels/_common.py
TensorLayout
dataclass
¶
TensorLayout(
shape: tuple[int, ...],
pack: Callable | None = None,
unpack: Callable | None = None,
stream: list | None = None,
block: tuple[int, ...] | None = None,
)
How a kernel wants one tensor operand laid out.
shape is the logical tile. pack and unpack are the reversible
host codec between (calls, *shape) and (calls, storage_elements);
identity is the default. stream is the DMA transform
(dims_to_stream) a design applies on the hop that feeds this operand
to the kernel or drains it, None when the operand streams as stored;
block is the micro-tile the kernel consumes or produces, (r, s)
for an MMUL operand. The codec is built from the same two facts, so the
host and the design agree by construction. None of this is an algorithm
or a whole-problem iteration schedule.
Param ¶
Read-only test-fixture parameter; its ABI determines scalar or tensor.
The generic harness holds its value fixed across calls. This is not a C++ operand lifetime: direct designs may pass a new value on every kernel call.
Trace
dataclass
¶
How a kernel's event0()/event1() markers bracket one call.
Trace.whole_call(): one pair brackets every call of the entry symbol
and nothing it calls emits another, so each trace interval is one call.
Trace.none(reason): a call emits no marker. Trace.partial(reason):
markers exist but do not bracket each call exactly once (around an inner
loop, or skipped on an early return), so intervals cannot be attributed
to calls. test_kernel_trace_markers.py checks the declaration against
the compiled IR of every library build.
aie.iron.algorithms.kernel_design turns any contract-bearing factory into a
design of one Worker, built on the same single-core pipeline as
transform, for_each and reduce. What a kernel can answer about itself
-- its
reference result, its safe input range, which arguments are parameters,
how to judge a device output -- lives on
ExternalFunction instead, so bringing up a
kernel needs no test harness. See
Kernel Library for
the add-a-kernel procedure and the test tiers built on it.
Import the builder with from aie.iron.algorithms import kernel_design as kd
and call kd.design(...), or import design directly from aie.iron.algorithms.
The former aie.utils.kernel_harness module has been removed.
Build, sample and check independent kernel calls from their declarations.
Each call consumes one tile per In, reads constant Param values, and
writes one tile per output. Layout codecs convert logical tiles to kernel
storage on the host. The Worker and the runtime sequence are the shared
single-core pipeline's (_pipeline). Whole-problem tiling, reductions
across kernel calls and multi-core schedules belong to algorithms, not this
kernel-validation harness.
CallCycles
dataclass
¶
CallCycles(
kernel: tuple[int, ...] = (),
initializers: dict[int, tuple[int, ...]] = dict(),
setup: tuple[int, ...] = (),
truncated: bool = False,
untimed: str | None = None,
)
One traced run's intervals, split by the kernel that emitted them.
kernel holds one interval per call of the measured kernel, in call
order; initializers the same for each traced initializer, keyed by
the InOut argument it initializes; setup the setup kernel's one
interval when it is traced. A trace that fills its buffer keeps a prefix
of the stream, which the split still labels correctly, and truncated
says the lists are short. untimed is the contract's reason when the
kernel's markers do not bracket its calls; then nothing ran.
design ¶
design(
factory,
*,
calls=1,
scalars=(),
shape=None,
params=None,
aiecc_flags=None,
guard=False,
**factory_kwargs
)
Wrap tile calls; params/scalars supply unbound tensor/scalar Params.
This harness embeds these values for every call; changing them recompiles the design. Direct designs can supply different operands on each call.
With guard=True the core writes GUARD_BYTES of 0x55 after each
output tile in its memory before every call and drains them with the
tile, so a kernel that writes past its output shows up on the host:
size the outputs with output_size(..., guard=True) and split them
with strip_guard. bfp outputs carry no guard.
Source code in python/iron/algorithms/kernel_design.py
host_args ¶
Describe physical host buffers, not kernel arguments or constant Params.
Same-type streamed inputs share a buffer; each output has its own.
Descriptors expose direction, shape, dtype and n_elements.
Source code in python/iron/algorithms/kernel_design.py
sample_inputs ¶
One array per unbound In/tensor Param; only In has a call dimension.
Source code in python/iron/algorithms/kernel_design.py
host_layout ¶
Pack declared layouts and interleave same-type In tiles; omit Param buffers.
Source code in python/iron/algorithms/kernel_design.py
output_size ¶
Storage elements per output; a tuple for multiple outputs.
Source code in python/iron/algorithms/kernel_design.py
upload ¶
Return design inputs and output tensor(s); splat multiple outputs when calling.
Source code in python/iron/algorithms/kernel_design.py
cycles_per_call ¶
cycles_per_call(
design_,
inputs,
out_size,
out_dtype,
*,
fn,
trace_size,
workdir,
calls=1
) -> CallCycles
Trace one run and split its intervals by kernel (CallCycles).
Only a kernel whose contract declares Trace.whole_call() is timed;
one declaring none or partial returns its reason without running,
and one declaring nothing raises. Traced initializers and a traced setup
kernel are split off by position; a partial one raises, because its
intervals cannot be labeled. Size trace_size from
traced_intervals.
Source code in python/iron/algorithms/kernel_design.py
traced_intervals ¶
How many event0/event1 intervals a traced run of fn emits.
A trace_size that holds fewer truncates the run; 0 when the
kernel itself is not timed.
Source code in python/iron/algorithms/kernel_design.py
split_intervals ¶
Label an interval stream: setup intervals, then per_call per call.
The harness runs the setup kernel once, then each call runs its traced
initializers in contract order and the kernel last, so interval j of
call n sits at setup + n * per_call + j. Returns (setup
intervals, one tuple per per-call kernel, truncated). A stream longer
than that is some kernel emitting markers it does not declare.
Source code in python/iron/algorithms/kernel_design.py
test/python/npu/test_kernels_perf.py times the library: correctness
first, then cycles, wall time and build size, gated by a device preflight
and a measurement-sanity test. It is an ordinary pytest module, so -k
selects cases and the session's exit status decides whether any numbers are
written. --baseline-sources DIR measures every selected case a second
time with its kernels from DIR and compares the two runs' raw output words
and cycles in --perf-meta. The timing helpers live here:
Benchmarking helpers for NPU kernel callables.
Stats keeps the raw sample list and robust statistics (median, MAD,
p95, coefficient of variation) beside avg/min/max, so callers can report
jitter, not just central tendency. The numbers are numpy's.
Stats
dataclass
¶
Stats(
avg_us: float,
min_us: float,
max_us: float,
median_us: float = 0.0,
mad_us: float = 0.0,
p95_us: float = 0.0,
cov: float = 0.0,
n: int = 0,
samples_us: list[float] = list(),
)
as_dict ¶
Flat dict for JSON emission (samples omitted).
Source code in python/utils/benchmark.py
Preflight
dataclass
¶
What the runtime says about the device a run is about to use.
run_iters ¶
run_iters(
fn: Callable,
*args,
warmup: int = 0,
iters: int = 1,
arg_sets: list[tuple] | None = None,
**kwargs
) -> BenchmarkResult
Invoke fn warmup + iters times, reporting timings.
End-to-end latency is measured around the Python call. If the return
value carries an npu_time (nanoseconds, captured by the runtime
around kernel.wait()), it is reported separately so callers can see
the host-side overhead delta.
arg_sets: an optional list of positional-argument tuples to
rotate through, one per iteration, so consecutive runs do not hit the
same host buffers (buffer rotation; see the CUTLASS measurement
guidelines). When given, *args must be empty.
Source code in python/utils/benchmark.py
preflight ¶
preflight() -> Preflight
Describe the device through the host runtime (any backend).
Source code in python/utils/benchmark.py
provenance ¶
Return a one-line description of what produced a measurement.
The git commit (GITHUB_SHA or git rev-parse HEAD), the Peano that
compiles the kernels (peano_version()), the installed mlir_aie
version, the kernel tree
(MLIR_AIE_KERNEL_SOURCES when set, and kernel_tree_digest()),
and any extra fields (device="NPU Strix", pmode="performance")
as key value pairs. A benchmark row records it so a number can be
traced to a toolchain and to the kernel sources.
A package that is not installed is left out rather than recorded as
unknown. CI builds mlir_aie from source and puts it on PYTHONPATH,
so there is no distribution to read a version from, and every published
row would otherwise carry a word that reads like a lookup failure. The
commit already identifies that build.
Source code in python/utils/benchmark.py
kernel_tree_digest ¶
Return a 12-hex digest of the kernel sources the library factories compile.
Every file under aie_kernels_dir() and aie_runtime_lib_dir(), by
relative path and content. The commit alone cannot say which kernels
ran: MLIR_AIE_KERNEL_SOURCES can name another tree, and a checkout
can carry uncommitted edits. None when neither directory exists.
Source code in python/utils/benchmark.py
Static checks¶
These compiler-remark checks are available on demand through
python -m aie.utils.compile.remarks; there is no static-check CI workflow.
The trace-marker check (test/python/test_kernel_trace_markers.py) runs in
lit on every PR, through trace_markers below.
Static kernel checks: compile every library kernel with Peano and read its remarks.
python -m aie.utils.compile.remarks --target aie2p --out static.json --out-pm static-pm.json --meta static-meta.json
python -m aie.utils.compile.remarks --target aie2p --only '^gelu' --out static.json --baseline-sources ../mlir-aie-base
CPU-only. Every factory in aie.iron.kernels (at its defaults and for each
entry of its .dtypes table) is compiled exactly as the JIT compiles it
(aie.utils.compile.utils.cxx_core_compile_command), plus the
optimization-record flags below, and the records become per-kernel series
for benchmark-action: a Peano bump that changes a loop's schedule shows up
here before anyone looks at device numbers. With MLIR_AIE_KERNEL_SOURCES
set to a checkout, that checkout's aie_kernels/ and aie_runtime_lib/
are compiled instead of the installed copies.
Record shapes, as llvm-aie 22.0.0.2026090201 emits them (they are Peano's, not LLVM's documented ones):
| Pass | Kind / Name | Args | Tracked as |
|---|---|---|---|
pipeliner |
Passed / schedule |
II, NS, Loop, Pipeliner, prologue/epilogue bundles |
loop/<fn>/<bb>/II (rest as hover text) |
pipeliner |
Missed / canPipelineLoop |
"Failed to pipeline loop"; located by DebugLoc only |
unpipelined_loops (keyed L<line>) |
pipeliner |
Analysis / schedule |
MII, SwpMaxMii, "Unable to find schedule" |
schedule_notes in the meta file |
aie-hardware-loops |
Analysis / analysis |
BasicBlock, Zero-Overhead-Loop |
non_zol_loops, and loop/<fn>/<bb>/not_zol per loop |
aie-asm-printer |
Analysis / analysis |
BasicBlock, BundleCount, ByteCount |
pm_bytes (summed over the shipped functions) |
aie-multi-slot-pseudo |
Missed / missing-memory-bank |
Instruction |
missing_bank_loads |
| stderr | -Wpass-failed |
a #pragma clang loop / AIE_* macro the compiler dropped |
pass_failed_warnings, text kept |
| the object | llvm-readobj sections, symbols, relocations |
what the entry symbol reaches | the shipped functions; libcalls (e.g. __divsf3) |
| the object | llvm-readobj --stack-sizes (-fstack-size-section) |
frame bytes per function | stack_bytes on the deepest path from the entry |
The loop counts and pm_bytes cover only the functions the entry symbol
reaches in the object, which are the ones the core link keeps.
stack_bytes over the contract's stack_bytes (else the device default)
prints a warning: the design reserves that much and an overflow corrupts the
neighbouring memory without a fault.
--baseline-sources DIR compiles everything a second time from DIR
and prints each row that differs, for a before/after of a kernel change;
--keep DIR keeps the objects, which --meta names per build.
The loop-scheduling pass reports as pipeliner (a postpipeliner
filter records nothing); it names loops by machine basic block
(bb.1.for.body.i) while the other passes use the IR block
(for.body.i), so the prefix is stripped to join them.
unpipelined_loops counts every loop the pipeliner declined, outer loops
included, so its change is the signal, not its value.
The integer series alert on any increase; pm_bytes goes to its own file
(--out-pm) so it can carry a percentage threshold. Nothing gates: under
GitHub Actions a dropped pragma is a ::warning at its file and line, a
kernel that fails to compile an ::error, and the run then exits 3 and
writes nothing.
LoopInfo
dataclass
¶
LoopInfo(
function: str,
block: str,
ii: int | None = None,
ns: int | None = None,
prologue_bundles: int | None = None,
epilogue_bundles: int | None = None,
pipelined: bool | None = None,
pipeliner: str | None = None,
missed_reason: str | None = None,
zol: bool | None = None,
bundle_count: int | None = None,
byte_count: int | None = None,
file: str | None = None,
line: int | None = None,
)
StaticReport
dataclass
¶
StaticReport(
loops: dict[tuple[str, str], LoopInfo] = dict(),
pm_bytes_by_function: dict[str, int] = dict(),
missing_bank_loads: int = 0,
pass_failed_warnings: int = 0,
pass_failed: list[str] = list(),
schedule_notes: list[str] = list(),
shipped: set[str] | None = None,
libcalls: list[str] = list(),
stack_bytes: int | None = None,
)
parse_yaml ¶
parse_yaml(
path: str | Path, report: StaticReport | None = None
) -> StaticReport
Source code in python/utils/compile/remarks.py
parse_stderr ¶
parse_stderr(
text: str, report: StaticReport
) -> StaticReport
Source code in python/utils/compile/remarks.py
report_rows ¶
report_rows(
report: StaticReport, prefix: str, extra: str
) -> list[dict]
benchmark-action rows for one kernel build. Smaller is better throughout.
Source code in python/utils/compile/remarks.py
workflow_annotations ¶
workflow_annotations(
name: str,
report: StaticReport | None,
detail: str,
root: str | None = None,
seen: set[tuple] | None = None,
) -> list[str]
Annotations for one kernel build: its dropped pragmas, or its compile failure.
One source is compiled once per factory build and per target, so the same
dropped pragma comes back several times a run; seen (shared across
builds) keeps each file, line and message to its first annotation.
Source code in python/utils/compile/remarks.py
compile_command ¶
Return the exact Peano command the JIT would run for ext_fn, plus remark flags.
Inline-source kernels (the aie2 LUT activations) are written out under the
kernel's symbol name first, as the JIT does. Kernels built with
use_chess are rejected: the remarks are Peano's.
Source code in python/utils/compile/remarks.py
analyze ¶
analyze(
ext_fn, target: str, workdir: Path
) -> tuple[StaticReport | None, str]
Compile one kernel and parse its records; (None, reason) when it fails to compile.
Source code in python/utils/compile/remarks.py
kernel_builds ¶
Yield (name, ExternalFunction) for every factory build the library offers.
Each exported factory at its defaults, plus one build per entry of its
.dtypes table; a factory that refuses the current device
(NotImplementedError) is skipped. Remarks depend on source and flags,
not on the shape a test runs, so this is the whole surface.
Source code in python/utils/compile/remarks.py
linked ¶
linked(obj: Path, entry: str) -> Linked
Functions and undefined symbols entry reaches in the object obj.
Source code in python/utils/compile/remarks.py
parse_readobj ¶
parse_readobj(doc: dict, entry: str) -> Linked
linked on the parsed llvm-readobj --elf-output-style=JSON document.
Source code in python/utils/compile/remarks.py
trace_markers ¶
Classify how event0()/event1() bracket a call of entry.
"whole_call" when one pair brackets every call and nothing else
emits a marker, "none" when a call emits no marker at all, and
otherwise a sentence saying what the markers do instead.
Source code in python/utils/compile/remarks.py
trace_shape ¶
Compile ext_fn to -O2 IR and classify its entry's markers (trace_markers).
Source code in python/utils/compile/remarks.py
entry_symbol ¶
Return the symbol the kernel source defines, before the JIT's per-build prefix.
Host-side helpers¶
aie.utils.bfp is the host side of the block-floating-point kernels: the
bfp16ebs8 codec and the tile shuffle the mm_bfp DMA layout needs. It is
the Python counterpart of programming_examples/ml/block_datatypes/helper.h,
which the examples' C++ hosts use; a Python host encodes and checks with
this module.
bfp16ebs8 on the host: encode, decode and the mmul-block shuffle.
v8bfp16ebs8 is the AIE2P block-floating-point type: eight values share
one 8-bit exponent and each carries an 8-bit two's-complement mantissa, so a
block of 8 values is 9 bytes ([exponent, m0, ..., m7]). These are numpy
ports of floatToBfp16, bfp16ebs8ToFloat and
shuffleMatrixForBfp16ebs8 in
programming_examples/ml/block_datatypes/helper.h, bit for bit (the host
test compiles that header and compares), including the header's rounding:
mantissas truncate toward negative infinity, and a value more than 31
binades below its block's maximum becomes 0 (positive) or -1 LSB (negative).
NumPy structured arrays describe the packed storage. Neither NumPy nor ml_dtypes provides shared-exponent arithmetic for this format; the codec below supplies that conversion, not a replacement scalar dtype.
The block-floating-point matmul kernels (aie.iron.kernels.mm_bfp) load
8x8 sub-tiles as one 72-byte block vector, which a DMA cannot gather at
9-byte granularity, so tiles are rearranged by shuffle on the host: within each
(tile_height, tile_width) tile the 8-row by 8-block sub-tiles are made
contiguous in raster order. quantize is what a kernel sees of a
float input, and what a reference should multiply.
encode ¶
float32 (..., n) with n % 8 == 0 -> uint8 (..., n * 9 // 8).
Blocks are taken along the last axis. Inputs must be finite (the C++ silently drops inf and NaN, which shifts every later value).
floor is the header's truncation. conv_even rounds each mantissa
to nearest, ties to even, as amd/IRON's f32_to_bfp16ebs8 packs
weights, byte for byte. A mantissa that rounds to +128 saturates to 127
there, so decode(encode(x, rounding="conv_even")) differs from
quantize, which models the core raising the exponent instead.
Source code in python/utils/bfp.py
decode ¶
uint8 (..., n * 9 // 8) -> float32 (..., n); the exact inverse map of a block.
Source code in python/utils/bfp.py
quantize ¶
Return what a kernel reads of x, for a given conversion rounding mode.
Which mode applies is a property of who converts, not of where. floor
is decode(encode(x)): this module's encoder truncates toward negative
infinity, and a core converting in floor mode agrees with it -- the
q4nx_dequant kernel pins floor and its reference matches the device
byte for byte. conv_even models a core converting with
round-to-nearest-ties-to-even, which is what mm_bfp's mixed kernel
pins so its K reduction does not accumulate a one-sided bias.
The mode is load-bearing, not a detail: on 64x64x64 mixed tiles of random and large inputs, the right one reproduces the kernel's bf16 output bit for bit, and floor mismatches about 3550 of each tile's 4096 outputs.
A mantissa that rounds up to +128 does not fit the 8-bit field, so the block's exponent goes up by one and the block is requantized. Clamping instead would cost a whole step to the one element the shared exponent was chosen for. The carry is one-sided, as on the core: -128 fits, so a block whose most negative value rounds to -128 keeps its exponent.
Source code in python/utils/bfp.py
shuffle ¶
shuffle(
b,
width: int,
height: int,
tile_width: int,
tile_height: int,
*,
unshuffle=False
) -> ndarray
Reorder an encoded (height, width) matrix into (or out of) the mmul tile layout.
b is the encoded matrix (height rows of width * 9 // 8 bytes,
flat or 2-D; width and tile_width count values, so both are
multiples of 8 and the tile is tile_height rows, a multiple of 8).
Within each tile, every 8-row by 8-value sub-tile (72 bytes) becomes
contiguous, sub-tiles in raster order; the tiles themselves stay where
they are, so a DMA that copies a tile row by row delivers the sub-tiles
the kernel's block-vector loads expect. unshuffle is the inverse.
Source code in python/utils/bfp.py
Tolerance-based output verification helpers for examples and tests.
Mirrors the canonical test_utils::nearly_equal semantics used across the
C++ testbenches so Python migrations of those examples behave identically:
|a - b| < max(atol, rtol * (|a| + |b|))
Defaults match the C++ default of rtol=0.128, which is the widely-used
relative tolerance for bfloat16 / LUT-approximated kernels (exp, softmax,
gelu, silu, swiglu, ...).
Tolerance
dataclass
¶
Tolerance(
rtol: float | None = None,
atol: float | None = None,
ulps: int | None = None,
range_frac: float | None = None,
max_mismatch_frac: float = 0.0,
note: str = "",
bound: Callable | None = None,
)
How close a kernel's output must be to its reference.
Exactly one of four kinds, chosen by which fields are set:
- exact -- no field set: bit-equal after casting the reference to the output dtype. Integers, selections (relu, max), lossless copies.
- ulps --
ulpsset: bf16 outputs withinulpsunits in the last place of the correctly rounded reference.atolmay be set alongside as a floor, admitting an element that meets either -- what a kernel needs when the device flushes subnormals to zero, since a flushed value is a full 100% relative and dozens of ulps from the reference but absolutely negligible.rtolstays unset, or the kind is relative. - relative --
rtol/atolset: the canonical|a - b| < max(atol, rtol * (|a| + |b|))ofnearly_equal. Integer outputs are compared with the same formula in exact integer arithmetic, solsb(atol = n + 0.5) admits ann-LSB slack for fixed-point pixel kernels whose rounding shift is not modeled; under exact and ulps integers stay bit-equal. - bound --
boundset:|a - b| <= bound(*args)element by element, whereargsare the reference's own arguments. For a floating-point kernel whose error is set by where its input falls rather than by what it produced -- an approximation exact in one range and a few ulps off in another -- so any onertol/atolis either loose everywhere or unsound somewhere.comparetakes the evaluated bound;ExternalFunction.judgeevaluates it from the inputs it is given.
Non-finite values are never skipped: NaN must meet NaN, and an infinity must meet an infinity of the same sign, under every kind.
range_frac adds a floor scaled to the reference's own range: an element
also passes at |a - b| <= range_frac * max|b|. It is for a kernel whose
error is set by the magnitudes it worked from rather than by the magnitude
it produced -- a dot product whose terms cancel to near zero is no less
accurate than its neighbours, but an elementwise relative bound reads it as
100% wrong. The scale comes from expected, never from actual, so a
kernel cannot widen its own tolerance by returning something large. Unlike
atol it follows the data: the same fraction holds whether the outputs
run to 34 or to 3.4e9, where a fixed floor would be either dead or
permissive. Set it from a measured worst case, and say so in note.
max_mismatch_frac allows that fraction of elements to miss (LUT tails,
saturation edges). note records where the number came from -- a
docstring, a device run, a testbench default -- so a reviewer can tell an
evidenced tolerance from a guessed one.
bf16_ulps
classmethod
¶
bf16_ulps(
n: int = 1,
*,
atol: float | None = None,
max_mismatch_frac: float = 0.0,
note: str = ""
) -> "Tolerance"
bf16 outputs within n ulps of the correctly rounded reference.
atol is an optional strict bound (error < atol) an element may
meet instead of the ulp bound. Use it for the device's subnormal flush
to zero, set to the smallest normal bf16 so it admits the flushed
values and nothing above them.
Source code in python/utils/verify.py
bounded
classmethod
¶
Each element within bound(*args) of the reference, absolutely.
bound takes the reference's arguments and returns one
non-negative bound per output element, shaped like the reference's
result (a tuple of them for several outputs). Derive it from the
kernel's arithmetic and say in note which parts were measured.
Source code in python/utils/verify.py
lsb
classmethod
¶
Integer outputs within n least-significant bits of the reference.
For fixed-point kernels whose final saturating shift may round or
truncate (the AIE srs rounding mode is a core setting the kernel
does not fix). rtol is zero: the slack is absolute.
Source code in python/utils/verify.py
default_for
classmethod
¶
Return the contract a kernel gets when it declares none.
Integer and boolean outputs are bit-exact. bfloat16 outputs get the
repository's canonical rtol=0.128, the C++ testbench default that
the LUT-approximated kernels document; float32 outputs are held to
rtol=1e-4 (a few float32 ULPs of accumulation-order slack, far
inside what a bf16 tolerance would hide) and float16 to 1e-2.
Source code in python/utils/verify.py
Verdict
dataclass
¶
Verdict(
ok: bool,
n_checked: int,
n_mismatch: int,
max_abs_err: float,
max_ulp_err: int | None,
first_bad_index: int | None,
detail: str,
)
Outcome of compare. Truthy when the comparison passed.
compare ¶
compare(
actual,
expected,
tol: Tolerance | None = None,
*,
range_axis: int | None = None,
bound=None
) -> Verdict
Compare a kernel's actual output with a reference under tol.
expected may be higher precision than actual (a float64 sum, an
int64 product); it is cast to actual.dtype, so the kernel is held to
what a correctly rounded implementation would produce. With tol=None
the output dtype's Tolerance.default_for applies.
range_axis selects the axis reduced to compute range_frac's
reference scale. For (calls, tile) arrays, use 1 to scale each call
independently. The default uses the whole reference array.
bound is a bound tolerance's per-element limit, already
evaluated on the inputs and broadcastable to expected.
This measures; it does not model. What a kernel does on overflow, on a
narrowing store, or with subnormal inputs belongs in the reference that
produced expected -- a saturating kernel's reference clips, a
denormal-flushing kernel's reference flushes. A reference that leaves the
output range is reported as such when the comparison fails, since it
means the reference is under-specified rather than the kernel wrong.
Source code in python/utils/verify.py
451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 | |
bf16_ulp_distance ¶
Element-wise distance between two bf16 arrays in units in the last place.
Bit patterns are mapped to a monotonic integer scale (sign-magnitude to two's-complement style) so the distance is a plain subtraction. -0 and +0 map to the same point, so a kernel that produces the other zero is not penalized.