Tensor Access Patterns (taplib)¶
taplib describes how DMAs walk tensors. A DMA buffer descriptor executes a
strided walk: an element offset plus parallel sizes and strides,
outermost dimension first. A TensorAccessPattern is exactly that walk over
a tensor of known shape. The tilings designs use can be built from
TensorAccessPattern.full(dims) (the row-major walk over the whole tensor)
with a few operations on those integers.
Building patterns¶
| To | Use |
|---|---|
| Walk a whole tensor row by row | TensorAccessPattern.full(dims) |
| Restrict the walk with NumPy indexing | tap[2:6, ::2], tap[i] |
| Reorder or reverse the dimensions | tap.permute((1, 0)), tap.T |
| Split or merge a dimension | tap.split(dim, inner), tap.merge(dim) |
| Use the fewest dimensions for the same walk | tap.coalesce() |
| Walk the same data again | tap.repeat(n) (a stride-0 outermost dimension) |
| Have a memtile pad the stream | tap.pad([(before, after), ...]) |
| Cut a tensor into equal tiles | tap.tile(tile_dims) |
Cut one dimension into k equal chunks |
tap.partition(k) |
| Read a tile-blocked buffer back in row-major order | tap.inverse() |
tile() turns a rank-r pattern into a rank-2r one: the grid
dimensions first, then the tile dimensions, so the whole result walks every
tile in row-major order. Everything else is ordinary indexing and
reordering of those dimensions:
| To | Use |
|---|---|
| One tile | tiles[i, j] |
| A row of tiles | tiles[i] |
A block of tiles, or every S-th tile |
tiles[a:b, j::S] |
| Visit tiles column by column | tiles.permute((1, 0, 2, 3)) |
| Walk each tile column-major | tiles.permute((0, 1, 3, 2)) |
| Walk a group of tiles again | tiles[i].repeat(n) |
Patterns are immutable: every operation returns a new one, and a list of
patterns is just a Python list.
from aie.helpers.taplib import TensorAccessPattern
tiles = TensorAccessPattern.full((16, 16)).tile((4, 4))
print(tiles) # TensorAccessPattern([16, 16], offset=0, sizes=[4, 4, 4, 4], strides=[64, 4, 16, 1])
print(tiles[1, 2]) # TensorAccessPattern([16, 16], offset=72, sizes=[4, 4], strides=[16, 1])
print(tiles[:2, :2])
# TensorAccessPattern([16, 16], offset=0, sizes=[2, 2, 4, 4], strides=[64, 4, 16, 1])
print(TensorAccessPattern.full((1024,)).partition(4)[2])
# TensorAccessPattern([1024], offset=512, sizes=[256], strides=[1])
Tiling a tiled pattern tiles hierarchically. tiles.tile((1, 1, 4, 4))
keeps the grid and cuts each tile into 4 x 4 sub-tiles, so the walk goes
tile by tile and, inside each tile, sub-tile by sub-tile.
tiles[i, j].tile((4, 4)) does the same for one tile. Such a walk can need
more dimensions than one buffer descriptor holds, even after coalesce(). A
shim fill() or drain() with a static walk is then split into several
transfers. An
ObjectFifo pattern has to fit in one descriptor (four dimensions on a
memtile, three on a core tile).
What a DMA can walk¶
A DMA moves whole 4-byte words. taplib itself accepts any walk, and the compiler rejects one the hardware cannot do when it lowers the design:
- Every size, and every stride other than an innermost 1, must span whole words. A bf16 run is an even number of elements, and an int8 run a multiple of four.
- The innermost dimension steps a word at a time. For elements that are not 32 bits wide, its stride must be 1.
So for bf16 or int8 data, .T and permute() of a tile are rejected, and so
is a ::2 slice of the innermost dimension. For tiles[0, 0].T on a bf16
tensor, a shim fill() fails with:
The same walk as a memtile's to_stream fails with:
To transpose sub-word data, let the DMA move s x s blocks whose rows are
whole words, and transpose each block in the kernel. The transposes example
below does this.
Using patterns¶
A pattern goes wherever IRON takes a DMA walk:
fill()anddrain()in a runtime sequence take it astap=.- An
ObjectFifotakes it asto_stream(how the producer's DMA reads its object onto the stream) and asfrom_stream/from_stream_per_cons(how a consumer's DMA writes the stream into its object). The pattern walks a tensor the size of what each transfer moves, from offset 0: one object, or one segment of it on a join's output or a distribute's input. A padded pattern asto_streamalso pads the stream on a MemTile. tap.transformation_dimsgives the((size, stride), ...)pairs, for code that still wants them.
Checking a data path on the host¶
On its way from the host to a core, a tensor passes through up to four DMAs, each walking it with its own pattern: the shim reads the host tensor onto the stream, the memtile writes it into an object, the memtile reads that object back out, and the core writes it into its own object. On hardware a mistake in any of them just looks like scrambled data in the kernel.
gather(tensor) returns the stream a DMA emits when it walks tensor with
a pattern, and scatter(stream) the object a DMA stores when it writes
stream with a pattern. Chaining them with the patterns a design uses shows exactly
what each core receives, with no hardware. This is the transposes design's
--strategy=combined path: the memtile shuffles each tile into s x s
blocks so that the kernel only has to transpose each block in place.
import numpy as np
from aie.helpers.taplib import TensorAccessPattern
M, K, m, n, s = 64, 64, 16, 16, 8
host = np.arange(M * K).reshape(M, K)
shim = TensorAccessPattern.full((M, K)).tile((m, n))
memtile_in = TensorAccessPattern.full((n, m)).tile((s, s)).permute((1, 2, 0, 3))
for t, tile in enumerate(shim.gather(host).reshape(-1, m * n)):
i, j = divmod(t, K // n)
obj = memtile_in.scatter(tile) # the memtile's (n, m) object
blocks = obj.reshape(n // s, s, m // s, s)
kernel_out = blocks.transpose(0, 3, 2, 1).reshape(n, m)
assert (kernel_out == host[i * m : (i + 1) * m, j * n : (j + 1) * n].T).all()
For a padded pattern, gather(tensor, pad_value=) fills the padded
positions, and padded_sizes is the shape the receiving object must have.
To inspect a single walk, use accesses(), access_order(),
access_count(), compare_access_orders() and visualize().
Staged patterns in a dispatch-time sequence¶
Every operation also accepts staged values: an aie.ir.Value (a
DispatchTime[T] scalar a runtime sequence receives, or any arithmetic on
one) can stand in for a size, stride, offset, index, slice bound or repeat
count. The arithmetic is emitted as arith ops where it is used, and
each check that would have raised ValueError becomes a cf.assert
guard. A fully static specialization folds the guard away, or fails at
generation time if it is false. The dispatch-time builder instead returns
no stream, and the host refuses the call with the guard's message.
def seq(a_h, b_h, start, n, in_prod, out_cons):
# The buffer as max_tiles equal chunks; the chunk index is staged
# arithmetic (start + loop iv), so the tap's offset is too.
chunks = TensorAccessPattern.full((max_tiles * tile_size,)).partition(max_tiles)
for tile in range_(n):
tap = chunks[start + tile]
tg = TaskGroup()
out_cons.drain(b_h, tap=tap, wait=True, group=tg)
in_prod.fill(a_h, tap=tap, group=tg)
tg.finish()
require(cond, message) from aie.iron adds a guard of your own, such as a
shape constraint the design depends on. cond can depend on dispatch-time
values. whole_array.py (below) has
require(M % (m * n_aie_rows) == 0, "M must be a multiple of m * n_aie_rows")
with a dispatch-time M, and a call whose M breaks it is refused with that
message. The message is a plain string fixed when the design is generated, so
it cannot include the value the call passed.
Some rules of thumb:
- Sizes, strides, offsets and loop bounds can all be dispatch-time values,
but the number of transfers a loop step issues cannot. A
range_over a dispatch-time bound is emitted as a loop, so its body is generated once. Every step issues the same transfers, waited on the same way, in the same order, and theTaskGroups carried from one step to the next are a fixed set of values (see Runtime tasks). This comes from how the sequence is generated, not from the DMA. So when the data does not divide evenly into blocks, the loop covers the whole blocks, and the ragged last block is one more transfer after it, underwith if_(rem > 0):, with its size computed fromrem. - The structure of a pattern stays a Python value: its rank, slice steps,
padding amounts, and the sizes and strides of the dimensions
merge()combines. - Staged values of any integer type are accepted. The compiler hoists their arithmetic out of the buffer-descriptor block, and a value too wide for its descriptor field is refused at dispatch instead of being truncated.
programming_examples/basic/matrix_multiplication/whole_array/whole_array.py
is the whole-array GEMM written this way, with dispatch-time M, K and N.
Checking a dispatch-time design¶
instructions(**scalars) on an @iron.jit design returns the instruction
words a call with those values would run, built exactly as a call builds
them, so no NPU is needed. A dispatch-time stream is not byte-identical to
a static specialization's: it draws buffer descriptors from a pool, polls for
room in a channel's task queue and assembles descriptor words at build time.
aie.utils.txn_trace reduces either stream to the events the hardware acts
on (queue pushes resolved to the chain of transfers they start, token waits,
and other register writes such as runtime parameters) and compares those:
from aie.utils.txn_trace import compare, explain
words = tiled_copy.specialize().instructions(n_tiles=3, start_tile=1)
static = tiled_copy.specialize(n_tiles=3, start_tile=1).instructions()
assert compare(words, static) == []
print(explain(words)) # one line per event
python -m aie.utils.txn_trace insts.bin [other.bin] does the same from
the command line. test/python/dispatch_taplib_copy.py and the whole-array
GEMM's tests/dispatch_txn.py are complete examples.
API reference¶
TensorAccessPattern ¶
TensorAccessPattern(
tensor_dims: Sequence[IntLike],
offset: IntLike,
sizes: Sequence[IntLike],
strides: Sequence[IntLike],
)
A strided walk over a tensor: offset plus outermost-first sizes and strides.
Build one with full() and refine it, or give the numbers directly.
Instances are immutable; every method returns a new pattern.
Two patterns are equal when they walk the same tensor the same way, i.e.
when their coalesce() forms match. Staged patterns are equal only to
themselves.
Create an access pattern.
Values may be Python ints or staged runtime values; a check on a staged value becomes a dispatch-time guard.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tensor_dims
|
Sequence[IntLike]
|
Shape of the tensor the pattern walks. |
required |
offset
|
IntLike
|
Element offset of the first element visited. |
required |
sizes
|
Sequence[IntLike]
|
Extent of each dimension, outermost first. |
required |
strides
|
Sequence[IntLike]
|
Element step of each dimension, outermost first. |
required |
Raises:
| Type | Description |
|---|---|
TypeError
|
If a value is not an integer. |
ValueError
|
If |
Source code in python/helpers/taplib/tap.py
transformation_dims
property
¶
The pattern as ((size, stride), ...), outermost first.
padding
property
¶
(before, after) padding per dimension, or None if the pattern is unpadded.
padded_sizes
property
¶
Extent of each dimension on the (possibly padded) stream.
numel
property
¶
Number of tensor elements the walk visits (repeats counted, padding not).
T
property
¶
The walk with its dimensions reversed, like numpy.ndarray.T.
TensorAccessPattern.full((M, N)).T walks an (M, N) tensor
column by column.
full
classmethod
¶
full(tensor_dims: Sequence[IntLike]) -> TensorAccessPattern
Return the row-major walk over a whole tensor of shape tensor_dims.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tensor_dims
|
Sequence[IntLike]
|
Shape of the tensor. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
TensorAccessPattern |
TensorAccessPattern
|
A pattern that visits every element once, in order. |
Source code in python/helpers/taplib/tap.py
permute ¶
permute(axes: Sequence[int]) -> TensorAccessPattern
Reorder dimensions: result dimension i is this pattern's dimension axes[i].
Whether a DMA can execute the result depends on its 4-byte word
granule: every non-unit stride must span whole words, and for elements
other than 32 bits the innermost stride must be 1. The DMA verifier and
the dynamic lowering enforce this, so .T of a bf16 or int8 tile is
rejected when the design is lowered.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
axes
|
Sequence[int]
|
A permutation of |
required |
Returns:
| Name | Type | Description |
|---|---|---|
TensorAccessPattern |
TensorAccessPattern
|
The reordered walk. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Source code in python/helpers/taplib/tap.py
split ¶
split(dim: int, inner: IntLike) -> TensorAccessPattern
Split dimension dim of size n into (n // inner, inner).
The outer part strides by inner * stride; the inner part keeps the
original stride.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
dim
|
int
|
The dimension to split. |
required |
inner
|
IntLike
|
Size of the new inner dimension; must divide |
required |
Returns:
| Name | Type | Description |
|---|---|---|
TensorAccessPattern |
TensorAccessPattern
|
The walk with one more dimension. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Source code in python/helpers/taplib/tap.py
merge ¶
merge(dim: int) -> TensorAccessPattern
Merge dimensions dim and dim + 1 into one.
Only legal when they are contiguous, i.e. strides[dim] ==
sizes[dim + 1] * strides[dim + 1], or when either has size 1.
Whether to merge is a structural
decision, so the values must be concrete.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
dim
|
int
|
The outer of the two dimensions. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
TensorAccessPattern |
TensorAccessPattern
|
The walk with one fewer dimension. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the two dimensions are not contiguous. |
TypeError
|
If either dimension is staged. |
Source code in python/helpers/taplib/tap.py
coalesce ¶
coalesce() -> TensorAccessPattern
Return the same walk in the fewest dimensions.
Size-1 dimensions are dropped and every contiguous adjacent pair is merged; the result visits the same elements in the same order. A pair with a staged size or stride is kept. The DMA builders apply this themselves to a walk deeper than the hardware takes.
Returns:
| Name | Type | Description |
|---|---|---|
TensorAccessPattern |
TensorAccessPattern
|
The same walk with no size-1 or mergeable dimensions. |
Source code in python/helpers/taplib/tap.py
repeat ¶
repeat(count: IntLike) -> TensorAccessPattern
Walk the whole pattern count times: a new outermost dimension with stride 0.
On a shim DMA the outermost dimension becomes the queue repeat; on a memtile it is a plain zero-stride dimension.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
count
|
IntLike
|
Number of walks; must be >= 1. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
TensorAccessPattern |
TensorAccessPattern
|
The walk with a new outermost dimension of |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Source code in python/helpers/taplib/tap.py
tile ¶
tile(tile_dims: Sequence[IntLike]) -> TensorAccessPattern
Divide every dimension into tiles of tile_dims.
Dimension i of size n_i becomes a grid dimension of n_i // t_i
tiles and a tile dimension of t_i elements. The grid dimensions
come first, so a rank-r pattern becomes a rank-2r one that walks
the tiles in row-major order: tiles[i, j] is the tile at grid
position (i, j) and tiles[i] the i-th row of tiles. Reorder
the grid with permute() ((1, 0, 2, 3) walks a column of tiles at
a time) and take strided or partial groups of tiles by slicing.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tile_dims
|
Sequence[IntLike]
|
One tile extent per dimension; each must divide the dimension. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
TensorAccessPattern |
TensorAccessPattern
|
The grid dimensions followed by the tile dimensions. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Source code in python/helpers/taplib/tap.py
partition ¶
partition(
parts: IntLike, dim: int = 0
) -> TensorAccessPattern
Split dimension dim into parts equal contiguous pieces.
Like np.array_split on an evenly divisible axis, it splits the
outermost dimension unless told otherwise. The result has a new
outermost dimension of parts, so
TensorAccessPattern.full((N,)).partition(k)[i] is the i-th of
k equal chunks of a flat range. Other dimensions are kept whole.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
parts
|
IntLike
|
Number of pieces; must divide the dimension. |
required |
dim
|
int
|
The dimension to split. Defaults to 0. |
0
|
Returns:
| Name | Type | Description |
|---|---|---|
TensorAccessPattern |
TensorAccessPattern
|
The pieces, one per index of the new outermost dimension. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Source code in python/helpers/taplib/tap.py
inverse ¶
inverse() -> TensorAccessPattern
Return the walk that puts this walk's stream back in row-major order.
A DMA that streams a tensor with a pattern visiting every element
exactly once stores a buffer B with B[k] = tensor.flat[walk[k]].
The inverse walks B and emits the tensor in row-major order, so
p.inverse().gather(p.gather(x)) is x.reshape(-1). For the
tiling full((m, n)).tile((r, t)) it is the "un-blocking" walk a
memtile applies to a core's blocked output: sizes [m//r, r, n//t, t]
with strides [r*n, t, r*t, 1].
Returns:
| Name | Type | Description |
|---|---|---|
TensorAccessPattern |
TensorAccessPattern
|
The inverse walk, over a buffer of the same shape. |
Raises:
| Type | Description |
|---|---|
TypeError
|
If the pattern is staged. |
ValueError
|
If the pattern is padded or does not visit every element of its tensor exactly once. |
Source code in python/helpers/taplib/tap.py
pad ¶
pad(
padding: Sequence[Sequence[int]],
) -> TensorAccessPattern
Surround every walk of each dimension with constant elements.
A memtile MM2S channel can pad the stream it emits: for dimension
i it inserts before constant elements ahead of each pass over
the dimension and after behind it, so the stream is
prod(padded_sizes) elements long. Give a padded pattern to an
ObjectFifo as to_stream and it sets the padding too (the pad value
is set on the fifo). Padding is applied last: a padded pattern cannot
be reshaped further.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
padding
|
Sequence[Sequence[int]]
|
One |
required |
Returns:
| Name | Type | Description |
|---|---|---|
TensorAccessPattern |
TensorAccessPattern
|
This walk with its padding. |
Raises:
| Type | Description |
|---|---|
TypeError
|
If a padding count is staged. |
ValueError
|
If the pattern is already padded, an entry is not a pair of counts >= 0, or there is not one entry per dimension. |
Source code in python/helpers/taplib/tap.py
gather ¶
Return the stream a DMA walking tensor with this pattern emits.
This is what an ObjectFifo given this pattern as to_stream sends,
padding included.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tensor
|
array_like
|
A tensor of shape |
required |
pad_value
|
optional
|
The value padded positions hold. Defaults to 0. |
0
|
Returns:
| Type | Description |
|---|---|
ndarray
|
np.ndarray: A 1-D array of |
Raises:
| Type | Description |
|---|---|
TypeError
|
If the pattern is staged. |
ValueError
|
If |
Source code in python/helpers/taplib/tap.py
scatter ¶
Write stream into a tensor the way a DMA walking this pattern does.
This is what an ObjectFifo given this pattern as from_stream stores:
stream element k lands at the k-th position of the walk.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
stream
|
array_like
|
The stream, one element per step of the walk. |
required |
out
|
ndarray | None
|
The tensor to write into; elements the
walk does not visit keep their values. Defaults to a zero tensor
of shape |
None
|
Returns:
| Type | Description |
|---|---|
ndarray
|
np.ndarray: The written tensor. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the pattern is padded (only an emitting DMA pads) or the stream length does not match the walk. |
Source code in python/helpers/taplib/tap.py
accesses ¶
Return the access_order and access_count arrays.
The access_order array numbers the accesses to each element of the tensor in walk order, -1 where the walk never goes; an element accessed more than once holds its last number. The access_count array holds the number of times the walk accesses each element.
Returns:
| Type | Description |
|---|---|
tuple[ndarray, ndarray]
|
tuple[np.ndarray, np.ndarray]: access_order and access_count, each
of shape |
Raises:
| Type | Description |
|---|---|
TypeError
|
If the pattern is staged. |
Source code in python/helpers/taplib/tap.py
access_order ¶
Return the access_order array of accesses().
Returns:
| Type | Description |
|---|---|
ndarray
|
np.ndarray: The walk number of each element's last access, -1 if unvisited. |
access_count ¶
Return the access_count array of accesses().
Returns:
| Type | Description |
|---|---|
ndarray
|
np.ndarray: The number of times the walk accesses each element. |
compare_access_orders ¶
compare_access_orders(other: TensorAccessPattern) -> bool
Return whether two patterns walk the same elements in the same order.
Patterns with different sizes and strides can still be functionally equivalent; this compares the walks themselves, padding included.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
other
|
TensorAccessPattern
|
The pattern to compare to. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
bool |
bool
|
Whether the two walks visit the same elements in the same order. |
Raises:
| Type | Description |
|---|---|
TypeError
|
If |
Source code in python/helpers/taplib/tap.py
visualize ¶
visualize(
show_arrows: bool | None = None,
title: str | None = None,
file_path: str | None = None,
show_plot: bool = True,
plot_access_count: bool = False,
) -> None
Plot the access order (and optionally count) of the walk over its tensor.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
show_arrows
|
bool | None
|
Draw arrows between consecutively accessed elements. Defaults to None (only for small tensors). |
None
|
title
|
str | None
|
Title of the plot. Defaults to the pattern's repr. |
None
|
file_path
|
str | None
|
Path to save the plot to. Defaults to None. |
None
|
show_plot
|
bool
|
Show the plot, e.g. in a Jupyter notebook. Defaults to True. |
True
|
plot_access_count
|
bool
|
Plot the access count as well as the access order. Defaults to False. |
False
|
Raises:
| Type | Description |
|---|---|
NotImplementedError
|
If the tensor has more than 2 dimensions. |
Source code in python/helpers/taplib/tap.py
animate ¶
animate(
frame_dims: int = 1,
title: str | None = None,
animate_access_count: bool = False,
) -> FuncAnimation
Animate the walk, one frame per index of its frame_dims outermost dimensions.
For a tiling full((M, N)).tile((m, n)), frame_dims=2 shows one
tile per frame and frame_dims=1 one row of tiles.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
frame_dims
|
int
|
Number of outermost dimensions the frames step through. Defaults to 1. |
1
|
title
|
str | None
|
The title of the animation. Defaults to None. |
None
|
animate_access_count
|
bool
|
Animate the access count as well as the access order. Defaults to False. |
False
|
Returns:
| Name | Type | Description |
|---|---|---|
FuncAnimation |
FuncAnimation
|
A handle to the animation. |
Raises:
| Type | Description |
|---|---|
NotImplementedError
|
If the tensor has more than 2 dimensions. |
ValueError
|
If |
Source code in python/helpers/taplib/tap.py
Utilities¶
Validation and stride helpers shared by the access-pattern algebra.
validate_and_clean_sizes_strides ¶
validate_and_clean_sizes_strides(
sizes: Sequence[IntLike], strides: Sequence[IntLike]
) -> tuple[list[IntLike], list[IntLike]]
Validate sizes and strides, and zero the strides of leading unit dimensions.
A check on a staged value becomes a dispatch-time guard.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
sizes
|
Sequence[IntLike]
|
Extent of each dimension, outermost first. |
required |
strides
|
Sequence[IntLike]
|
Element step of each dimension, outermost first. |
required |
Returns:
| Type | Description |
|---|---|
tuple[list[IntLike], list[IntLike]]
|
tuple[list[IntLike], list[IntLike]]: The sizes and the cleaned strides. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the lists are empty or differ in length, a size is below 1 or a stride below 0. |
Source code in python/helpers/taplib/utils.py
zero_leading_unit_strides ¶
Zero the stride of each unit dimension from start up to the first that steps.
A unit dimension never steps, so this makes equal walks compare equal. The innermost stride is left as is. Rank and unit-ness are structural, so a staged size ends the scan.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
sizes
|
Sequence
|
Extent of each dimension, outermost first. |
required |
strides
|
Sequence
|
Element step of each dimension, outermost first. |
required |
start
|
int
|
The first dimension to consider. Defaults to 0. |
0
|
Returns:
| Name | Type | Description |
|---|---|---|
list |
list
|
The strides, with those of the leading unit dimensions zeroed. |
Source code in python/helpers/taplib/utils.py
validate_tensor_dims ¶
Check that a tensor has at least one dimension and every dimension is >= 1.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tensor_dims
|
Sequence[IntLike]
|
Tensor dimensions to check. |
required |
Returns:
| Type | Description |
|---|---|
list[IntLike]
|
list[IntLike]: The tensor dimensions. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If there are no dimensions or a dimension is below 1. |
Source code in python/helpers/taplib/utils.py
validate_offset ¶
Check that offset is an element index into a tensor of shape tensor_dims.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
offset
|
IntLike
|
The offset to check. |
required |
tensor_dims
|
Sequence[IntLike]
|
Shape of the tensor. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
IntLike |
IntLike
|
The offset. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the offset is negative or past the last element. |
Source code in python/helpers/taplib/utils.py
row_major_strides ¶
Row-major (C-order) element strides for a tensor of shape dims.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
dims
|
Sequence
|
Tensor dimensions; entries may be staged values. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
list |
list
|
One stride per dimension, outermost first. |
Source code in python/helpers/taplib/utils.py
validate_permutation ¶
Check that axes is a permutation of range(rank).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
axes
|
Sequence[int]
|
The permutation to check. |
required |
rank
|
int
|
Number of dimensions permuted. |
required |
what
|
str
|
Name of the argument, for the error message. |
required |
Returns:
| Type | Description |
|---|---|
tuple[int, ...]
|
tuple[int, ...]: The permutation as a tuple of ints. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Source code in python/helpers/taplib/utils.py
Instruction-stream tracing¶
Decode and semantically compare NPU TXN instruction streams.
A runtime sequence compiled once with DispatchTime scalars and a fully
static specialization of the same sequence program the same DMA transfers, but
their word streams are not byte-identical: the dynamic path draws buffer
descriptors from a free-list pool (different BD ids), polls for room in a
channel's task queue, assembles BD words at build time, and leaves the
buffer-address word to the address patch. This module replays a stream into
the register state it programs and reduces it to a list of events, the
things the hardware actually acts on:
push: a queue write on a channel, resolved to the chain of BDs it runs (for each, transfer length, address as (host argument, byte offset) or an absolute address, the addressing dimensions in a normalized form, and iteration), the repeat count, whether a task-completion token is issued, and the channel control bits.wait: a task-completion-token wait.write: a register write outside the DMA BD and channel registers (runtime parameters, locks, anything else the sequence programs).loadpdi/preempt/scratchpad/update_reg: kept verbatim.
Polls only wait for state another op produces, so they are not events. A push of a BD the stream never wrote is an error rather than a transfer of zeros: its contents come from somewhere the stream does not show.
Two streams are equivalent when their headers name the same device and their
event lists are equal. compare reports the first divergence; explain
prints the events for a human.
The decoding covers the AIE2 / AIE2p shim, memtile and core DMA register
layouts used by npu1 and npu2 (register fields from aie-rt's
xaiemlgbl_params.h).
Opcode ¶
Bases: IntEnum
TXN opcodes, from include/aie/Runtime/TxnEncoding.h.
Header
dataclass
¶
Header(
major: int,
minor: int,
dev_gen: int,
num_rows: int,
num_cols: int,
num_mem_tile_rows: int,
num_ops: int,
size: int,
)
The 4-word TXN header.
Op
dataclass
¶
Op(
kind: str,
pos: int,
words: tuple[int, ...],
addr: int = 0,
value: int = 0,
mask: int = 0,
data: tuple[int, ...] = (),
)
One decoded TXN instruction.
Transfer
dataclass
¶
Transfer(
length: int,
address: tuple,
dims: tuple[tuple[int, int], ...],
outer_stride: int,
iteration: tuple[int, int],
packet: tuple[int, ...],
flags: tuple[int, ...],
burst_axcache: tuple[int, int],
padding: tuple[tuple[int, int], ...] = (),
)
A buffer descriptor as the DMA sees it.
The form is independent of how the stream encoded it.
decode ¶
decode(words: Sequence[int] | ndarray) -> list[Op]
Split a TXN stream (header included) into instructions.
The header's size and op count must match the stream.
Source code in python/utils/txn_trace.py
351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 | |
trace ¶
Replay a stream and return the events it triggers, in order.
Source code in python/utils/txn_trace.py
590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 | |
compare ¶
Return the differences between two streams' devices and events.
Empty when the streams are equivalent.
Source code in python/utils/txn_trace.py
explain ¶
Human-readable listing of a stream's events (or raw ops).