Skip to content

IRON API (aie.iron)

IRON is the high-level Python interface for programming AMD Ryzen™ AI NPUs. You describe a design as Python objects — tiles, workers, data movement, and a host runtime — and IRON compiles it to an optimized xclbin + instruction stream via the MLIR-AIE toolchain.

Structural design objects are resolvable: they lower to MLIR operations when the design is compiled. This page also includes host-side utilities, type markers, and runtime task handles. For the direct MLIR op wrappers that IRON lowers to, see the Dialect op wrappers page.

Import the public design abstractions from aie.iron (e.g. from aie.iron import Worker, ObjectFifo, Runtime), or use import aie.iron as iron for decorators and tensor factories. Supporting types are also documented below; not all are re-exported at package level.

Prefer ObjectFifo for synchronized streaming, Worker for compute, and Runtime(sequence, fn_args) for host-side transfers. Pass runtime buffer types and the ObjectFifo handles used by the sequence through fn_args, and pass workers to Program(workers=...). Use Flow, TileDma, and Lock when explicit routing or DMA control is the teaching goal; use raw dialect operations only where those abstractions do not expose the required operation.


Core design abstractions

The objects most designs are built from.

Program

Program

Program(
    device: Device | None,
    rt: Runtime,
    workers: list | None = None,
)

Construct a Program with all design information needed to run the design on a device.

Note

MLIR verification (ctx.module.operation.verify()) is performed inside resolve_program, not during construction.

Parameters:

Name Type Description Default
device Device

The device used to generate the final MLIR for the design. Accepts the Device | None returned by iron.get_current_device directly and raises if no device has been selected, so callers need not narrow it first.

required
rt Runtime

The runtime object for the design.

required
workers list[Worker] | None

The Workers to run on the device. Defaults to None (no workers). Workers are passed here explicitly rather than started from within the runtime sequence.

None

Raises:

Type Description
ValueError

If device is None (no NPU device was selected/detected).

Source code in python/iron/program.py
def __init__(
    self,
    device: Device | None,
    rt: Runtime,
    workers: "list | None" = None,
):
    """Construct a Program with all design information needed to run the design on a device.

    !!! note
        MLIR verification (`ctx.module.operation.verify()`) is performed inside
        [`resolve_program`][iron.program.Program.resolve_program], not during construction.

    Args:
        device (Device): The device used to generate the final MLIR for the design.
            Accepts the ``Device | None`` returned by ``iron.get_current_device``
            directly and raises if no device has been selected, so callers need
            not narrow it first.
        rt (Runtime): The runtime object for the design.
        workers (list[Worker] | None, optional): The Workers to run on the
            device. Defaults to None (no workers). Workers are passed here
            explicitly rather than started from within the runtime sequence.

    Raises:
        ValueError: If ``device`` is None (no NPU device was selected/detected).
    """
    if device is None:
        raise ValueError(
            "Program requires a device, but none was selected. Pass an explicit "
            "Device, or ensure an NPU runtime is available for "
            "iron.get_current_device()."
        )
    self._device = device
    self._rt = rt
    self._workers = list(workers) if workers is not None else []
    self._trace_size = None
    self._trace_workers = None
    self._reuse_output_buffer = False
    self._egress_shim_col = 0
    self._coretile_events = None
    self._coremem_events = None
    self._memtile_events = None
    self._shimtile_events = None
    self._core_trace_mode = TraceMode.EventTime

enable_trace

enable_trace(
    trace_size: int | None = None,
    workers: list | None = None,
    reuse_output_buffer: bool = False,
    coretile_events: list | None = None,
    coremem_events: list | None = None,
    memtile_events: list | None = None,
    shimtile_events: list | None = None,
    egress_shim_col: int = 0,
    core_trace_mode=EventTime,
)

Enable hardware tracing for this program.

Configures the AIE trace units and routes trace packets to DDR via the shim DMA. Lives on Program (not Runtime) because it configures both the traced Workers' tiles and the Runtime's trace-buffer sequencing.

Parameters:

Name Type Description Default
trace_size int

Size of the trace buffer in bytes.

None
workers list[Worker] | None

Specific workers to trace. If None, all workers with trace set will be traced. Defaults to None.

None
reuse_output_buffer bool

When False (default), trace lowering appends a dedicated trace-buffer argument to the runtime_sequence; it lands at the tail so enabling trace never perturbs the data arguments' indices. When True, trace data is written into the tail of the last output buffer, saving a host buffer. Defaults to False.

False
coretile_events list | None

List of up to 8 core tile trace events. See the AIEX dialect reference for available events under (type)EventAIE such as CoreEventAIE. Defaults to None (uses hardware defaults).

None
coremem_events list | None

List of up to 8 core memory trace events. Defaults to None (uses hardware defaults).

None
memtile_events list | None

List of up to 8 mem tile trace events. Defaults to None (uses hardware defaults).

None
shimtile_events list | None

List of up to 8 shim tile trace events. Defaults to None (uses hardware defaults).

None
egress_shim_col int

Column of the shim tile used to egress trace packets to DDR. Defaults to 0.

0
core_trace_mode TraceMode

Trace mode for core tiles. Defaults to Event-Time.

EventTime
Source code in python/iron/program.py
def enable_trace(
    self,
    trace_size: int | None = None,
    workers: list | None = None,
    reuse_output_buffer: bool = False,
    coretile_events: list | None = None,
    coremem_events: list | None = None,
    memtile_events: list | None = None,
    shimtile_events: list | None = None,
    egress_shim_col: int = 0,
    core_trace_mode=TraceMode.EventTime,
):
    """Enable hardware tracing for this program.

    Configures the AIE trace units and routes trace packets to DDR via the shim DMA.
    Lives on Program (not Runtime) because it configures both the traced
    Workers' tiles and the Runtime's trace-buffer sequencing.

    Args:
        trace_size (int): Size of the trace buffer in bytes.
        workers (list[Worker] | None, optional): Specific workers to trace. If None,
            all workers with ``trace`` set will be traced. Defaults to None.
        reuse_output_buffer (bool, optional): When False (default), trace
            lowering appends a dedicated trace-buffer argument to the
            runtime_sequence; it lands at the tail so enabling trace never
            perturbs the data arguments' indices. When True, trace data is
            written into the tail of the last output buffer, saving a host
            buffer. Defaults to False.
        coretile_events (list | None, optional): List of up to 8 core tile trace events.
            See [the AIEX dialect reference](../AIEXDialect.md) for available
            events under (type)EventAIE such as CoreEventAIE.
            Defaults to None (uses hardware defaults).
        coremem_events (list | None, optional): List of up to 8 core memory trace events.
            Defaults to None (uses hardware defaults).
        memtile_events (list | None, optional): List of up to 8 mem tile trace events.
            Defaults to None (uses hardware defaults).
        shimtile_events (list | None, optional): List of up to 8 shim tile trace events.
            Defaults to None (uses hardware defaults).
        egress_shim_col (int, optional): Column of the shim tile used to
            egress trace packets to DDR. Defaults to 0.
        core_trace_mode (TraceMode, optional): Trace mode for core tiles.
            Defaults to Event-Time.
    """
    self._trace_size = trace_size
    self._trace_workers = workers
    self._reuse_output_buffer = reuse_output_buffer
    self._coretile_events = coretile_events
    self._coremem_events = coremem_events
    self._memtile_events = memtile_events
    self._shimtile_events = shimtile_events
    self._core_trace_mode = core_trace_mode
    self._egress_shim_col = egress_shim_col

resolve_program

resolve_program(device_name='main')

Resolve the program components in order to generate MLIR.

Tiles are emitted as aie.logical_tile ops. The --aie-place-tiles pass in the compilation pipeline converts them to aie.tile ops.

Returns:

Name Type Description
module Module

The module containing the MLIR context information.

Source code in python/iron/program.py
def resolve_program(self, device_name="main"):
    """Resolve the program components in order to generate MLIR.

    Tiles are emitted as aie.logical_tile ops. The --aie-place-tiles pass
    in the compilation pipeline converts them to aie.tile ops.

    Returns:
        module (Module): The module containing the MLIR context information.
    """
    with mlir_mod_ctx() as ctx:
        # Create a fresh device instance of the same type to avoid stale MLIR operations
        # This preserves the device configuration while ensuring clean state
        device_type = type(self._device)
        # For dynamically created device classes, the constructor takes no arguments
        self._device = device_type()  # pyright: ignore[reportCallIssue]

        # Resolve parameters known up front (Worker fn_args) at module
        # scope now. aiex.scratchpad_parameter ops are global across all
        # devices because the scratchpad is a single hardware resource
        # shared by all PDIs.
        for w in self._workers:
            for arg in w.flat_fn_args:
                if isinstance(arg, ScratchpadParameter):
                    arg.resolve()

        @device(self._device.resolve(), sym_name=device_name)
        def device_body():
            # Collect all fifos. Runtime-driven fifos already have their shim
            # endpoints bound (Runtime registered its fn_args at construction),
            # so they resolve here with both ends known -- the sequence body
            # is emitted after workers (self._rt.resolve() below),
            # so body verbs that read worker-side state (barrier locks, worker
            # Buffer placement) see it resolved.
            all_fifos = set()
            all_fifos.update(self._rt.fifos)
            for w in self._workers:
                all_fifos.update(w.fifos)

            # Sort fifos for deterministic resolve
            all_fifos = sorted(all_fifos, key=lambda obj: obj.name)

            # Collect all tiles. Two workers landing on the same compute
            # tile (pinned or after placement) is caught by the aie.device
            # verifier's one-core-per-tile check, so no Python-side guard.
            all_tiles = []
            for w in self._workers:
                all_tiles.append(w.tile)
                # Generic: any user-side Resolvable in fn_args may declare
                # additional tile dependencies via tiles(). Default is [].
                for arg in w.flat_fn_args:
                    if isinstance(arg, Resolvable):
                        all_tiles.extend(arg.tiles())
            for f in all_fifos:
                all_tiles.extend([e.tile for e in f.all_of_endpoints()])
                # Shared-memory delegate tile (ObjectFifo.delegate_tile kwarg)
                # may not appear in any prod/cons endpoint, so pick it up
                # explicitly so resolve_tile() runs on it before fifo resolution.
                if f._object_fifo._delegate_tile is not None:
                    all_tiles.append(f._object_fifo._delegate_tile)
            # Lower-level: explicit Flow / TileDma / Lock primitives
            # contribute tiles too.
            for fl in self._rt.flows:
                all_tiles.extend(fl.all_tiles())
            for td in self._rt.tile_dmas:
                all_tiles.extend(td.all_tiles())
            for lk in self._rt.locks:
                all_tiles.append(lk.tile)

            # Resolve tiles
            for t in all_tiles:
                self._device.resolve_tile(t)

            # Generate fifos
            for f in all_fifos:
                f.resolve()

            # Generate explicit Locks (must come before TileDma + Worker
            # bodies that reference them; Buffers attached to worker
            # fn_args are still resolved in the worker loop below).
            for lk in self._rt.locks:
                lk.resolve()

            # Resolve any Buffers referenced by explicit TileDma programs
            # (those aren't reached via worker.fn_args).
            for td in self._rt.tile_dmas:
                bufs, _ = td.all_buffers_and_locks()
                for b in bufs:
                    if b.tile is None:
                        b._tile = td.tile
                    b.resolve()

            # generate functions - this may call resolve() more than once on the same fifo, but that's ok
            for w in self._workers:
                for arg in w.flat_fn_args:
                    if isinstance(arg, FuncBase):
                        arg.emit()
                    elif isinstance(arg, Resolvable):
                        if (
                            arg not in self._rt.flows
                            and arg not in self._rt.tile_dmas
                        ):
                            arg.resolve()

            # Generate core programs
            for w in self._workers:
                w.resolve()

            # Emit aie.cascade_flow ops for each Worker's outgoing edges.
            # Must run after worker.resolve() so both tiles are placed.
            for w in self._workers:
                for cf in w._outgoing_cascades:
                    cf.resolve()

            # Generate explicit per-tile DMA programs (lower-level peers
            # of ObjectFifo, paired with Flow + Lock).
            self._rt.resolve_tile_dmas()

            # Generate trace routes
            # TODO Need to iterate over all tiles or workers & fifos to make list of tiles to trace
            #      Alternatively, we merge the mechanism for packet routed objfifos so we use unique
            #      route IDs for trace as well

            # Scan workers and build list of tiles to trace
            tiles_to_trace = []
            if self._trace_workers is not None:
                for w in self._trace_workers:
                    tiles_to_trace.append(w.tile.op)
            else:
                for w in self._workers:
                    if w.trace is not None:
                        tiles_to_trace.append(w.tile.op)
            if self._trace_size is not None and self._trace_size > 0:
                trace_utils.configure_trace(
                    tiles_to_trace,
                    coretile_events=self._coretile_events,
                    coremem_events=self._coremem_events,
                    memtile_events=self._memtile_events,
                    shimtile_events=self._shimtile_events,
                    core_trace_mode=self._core_trace_mode,
                )

            # Emit the runtime sequence body after workers, their locks, and
            # worker Buffers are resolved, so body verbs that read that
            # state (barrier.set, inline_ops over a worker Buffer) are valid.
            # Its shim DMAs reference fifos by symbol name (forward ref), so
            # emitting after the fifo ops is fine.
            #
            # On the full-ELF path the runtime sequence must load its own
            # PDI (no xclbin configures the device), so pass the device
            # symbol as the load_pdi reference. The flag is injected into
            # the compile context by CompilableDesign.
            load_pdi_device_ref = (
                device_name if get_compile_arg("_iron_full_elf") else None
            )
            self._rt.resolve(
                trace_size=self._trace_size,
                reuse_output_buffer=self._reuse_output_buffer,
                egress_shim_col=self._egress_shim_col,
                load_pdi_device_ref=load_pdi_device_ref,
            )

            # Flow transfers name their allocations while the sequence runs.
            # Resolve both at device scope using those symbol references.
            for fl in self._rt.flows:
                fl.resolve()

        # Resolve parameters only discoverable once the sequence body has
        # traced (offset_parameter= passed directly to fill()/drain(),
        # rather than declared up front via Worker fn_args). device_body's
        # own insertion point is scoped to its @device region, so by now
        # the ambient insertion point is back to module scope -- the same
        # place the fn_args-declared parameters above were resolved.
        for p in self._rt._scratchpad_parameters:
            p.resolve()

        self._print_verify(ctx)
        return ctx.module

Worker

Worker and WorkerRuntimeBarrier: compute-core tasks and runtime synchronization primitives.

Worker

Worker(
    core_fn: Callable | None,
    fn_args: list | None = None,
    tile: Tile | None = AnyComputeTile,
    while_true: bool = True,
    stack_size: int | None = None,
    data_size: int | None = None,
    trace: int | None = None,
    trace_events: list | None = None,
    dynamic_objfifo_lowering: bool | None = None,
)

Bases: ObjectFifoEndpoint

A task to be run on an AIE compute core.

A Worker takes a core_fn callable and the arguments it needs (ObjectFIFO handles, Buffers, Kernels, etc.). Each Worker is placed on a single compute tile, either explicitly via tile or automatically by the --aie-place-tiles compiler pass.

Construct a Worker.

Parameters:

Name Type Description Default
core_fn Callable | None

The task to run on a core. If None, a busy-loop (while(true): pass) core will be generated.

required
fn_args list | None

Pointers to arguments, which should include all context the core_fn needs to run. Defaults to None (empty list).

None
tile Tile

The compute tile for the Worker. Also accepts None (treated as AnyComputeTile). Defaults to AnyComputeTile.

AnyComputeTile
while_true bool

If true, will wrap the core_fn in a while(true) loop to ensure it runs until reconfiguration. Defaults to True.

True
stack_size int

The stack_size in bytes for the worker. Defaults to AIETargetModel::getDefaultCoreStackSize() (currently 1024 bytes).

None
data_size int

Bytes of data memory to reserve for this core's compiled sections (.data/.rodata/.bss), beyond the stack. The buffer allocator packs the tile's buffers around the reservation. None leaves the core whatever contiguous run the buffers leave. Defaults to None.

None
trace int

If >0, enable tracing for this worker.

None
trace_events list | None

Custom list of trace events for this worker. Defaults to None.

None
dynamic_objfifo_lowering bool | None

Per-core override for the aie-objectFifo-stateful-transform pass's lowering choice. True forces dynamic (loop-preserving) lowering for this core; False forces static LCM-based unrolling. None (default) leaves the choice to the compiler's global --dynamic-objFifos flag. Note: the per-core attribute is only honored when the global flag is false; when global is true the attribute is ignored. Defaults to None.

None

Raises:

Type Description
ValueError

Parameters are validated.

Source code in python/iron/worker.py
def __init__(
    self,
    core_fn: Callable | None,
    fn_args: list | None = None,
    tile: Tile | None = AnyComputeTile,
    while_true: bool = True,
    stack_size: int | None = None,
    data_size: int | None = None,
    trace: int | None = None,
    trace_events: list | None = None,
    dynamic_objfifo_lowering: bool | None = None,
):
    """Construct a Worker.

    Args:
        core_fn (Callable | None): The task to run on a core. If None, a busy-loop (`while(true): pass`) core will be generated.
        fn_args (list | None, optional): Pointers to arguments, which should include all context the core_fn needs to run. Defaults to None (empty list).
        tile (Tile, optional): The compute tile for the Worker. Also accepts None (treated as AnyComputeTile). Defaults to AnyComputeTile.
        while_true (bool, optional): If true, will wrap the core_fn in a while(true) loop to ensure it runs until reconfiguration. Defaults to True.
        stack_size (int, optional): The stack_size in bytes for the worker. Defaults to AIETargetModel::getDefaultCoreStackSize() (currently 1024 bytes).
        data_size (int, optional): Bytes of data memory to reserve for this
            core's compiled sections (.data/.rodata/.bss), beyond the stack. The
            buffer allocator packs the tile's buffers around the reservation. None
            leaves the core whatever contiguous run the buffers leave.
            Defaults to None.
        trace (int, optional): If >0, enable tracing for this worker.
        trace_events (list | None, optional): Custom list of trace events for this worker. Defaults to None.
        dynamic_objfifo_lowering (bool | None, optional): Per-core override for the
            ``aie-objectFifo-stateful-transform`` pass's lowering choice. ``True`` forces
            dynamic (loop-preserving) lowering for this core; ``False`` forces static
            LCM-based unrolling. ``None`` (default) leaves the choice to the compiler's
            global ``--dynamic-objFifos`` flag. Note: the per-core attribute is only
            honored when the global flag is ``false``; when global is ``true`` the
            attribute is ignored. Defaults to None.

    Raises:
        ValueError: Parameters are validated.
    """
    for arg in flatten_fn_args(
        [
            fn_args or [],
            tile,
            while_true,
            stack_size,
            data_size,
            trace,
            trace_events,
            dynamic_objfifo_lowering,
        ]
    ):
        if isinstance(arg, _DispatchParameter):
            arg._misuse()
    if tile is None:
        tile = AnyComputeTile
    if tile.tile_type is not None and tile.tile_type != AIETileType.CoreTile:
        raise ValueError(
            f"Worker requires a compute tile, but got tile_type={tile.tile_type}"
        )
    if stack_size is not None:
        if not isinstance(stack_size, int) or isinstance(stack_size, bool):
            raise ValueError(
                f"Worker stack_size must be an int, but got "
                f"{type(stack_size).__name__}"
            )
        if stack_size < 1:
            raise ValueError(
                f"Worker stack_size must be >= 1, but got {stack_size}"
            )
    if data_size is not None:
        if not isinstance(data_size, int) or isinstance(data_size, bool):
            raise ValueError(
                f"Worker data_size must be an int, but got "
                f"{type(data_size).__name__}"
            )
        if data_size < 0:
            raise ValueError(
                f"Worker data_size must be >= 0, but got " f"{data_size}"
            )
    # Store the user's Tile directly when it is already typed as CoreTile.
    # This preserves Python object identity so a Buffer and a Worker that
    # share the same Tile object resolve to a single LogicalTileOp. When a
    # fresh copy is needed (untyped tile or the singleton default) use
    # with_type() — it always returns a new object.
    if tile.tile_type == AIETileType.CoreTile and tile is not AnyComputeTile:
        self._tile = tile
    else:
        self._tile = tile.with_type(AIETileType.CoreTile)
    self._while_true = while_true
    self.stack_size = stack_size
    self.data_size = data_size
    self._dynamic_objfifo_lowering = dynamic_objfifo_lowering
    self.trace = trace
    self.trace_events = trace_events

    # If no core_fn is given, make a simple while(true) loop.
    if core_fn is None:

        def do_nothing_core_fun(*args) -> None:
            for _ in range_(sys.maxsize):
                pass

        self.core_fn = do_nothing_core_fun
    else:
        self.core_fn = core_fn
    self.fn_args = fn_args if fn_args is not None else []
    self._fifos = []
    self._buffers = []
    self._barriers = []
    # CascadeFlow objects whose source is this Worker. Populated by
    # CascadeFlow(src, dst).__init__ and consumed by Program.resolve()
    # to emit aie.cascade_flow ops after worker placement.
    self._outgoing_cascades: list = []

    # Check arguments to the core. Some information is saved for resolution.
    # fn_args may nest lists (e.g. one fifo per column); iterate the flattened
    # leaves for registration while the core_fn still receives the structure.
    for arg in flatten_fn_args(self.fn_args):
        if isinstance(arg, ObjectFifoHandle):
            arg.endpoint = self
            self._fifos.append(arg)
        elif isinstance(arg, Buffer):
            # A Buffer pinned to an EXPLICIT tile may legitimately be shared
            # across Workers: AIE compute tiles can read a neighbor tile's L1
            # directly, so a producer core's output buffer can be an input to a
            # consumer core on an adjacent tile. In that case the FIRST worker that
            # references it "owns"/places it and later workers are non-owning
            # readers. We only forbid sharing for AUTO-PLACED buffers (no explicit
            # tile), where two owners would race to pin it to different tiles.
            # Note: ``_tile`` alone is not a reliable signal — the owning Worker
            # auto-pins ``_tile`` to its own tile below — so we key off
            # ``_explicit_tile``, which records the user's construction-time intent.
            if arg._owner_worker is not None and arg._owner_worker is not self:
                if not arg._explicit_tile:
                    raise ValueError(
                        f"Buffer '{arg._name}' has no explicit tile and is shared "
                        f"across Workers; pin it to a tile (Buffer(tile=...)) so "
                        f"placement is unambiguous."
                    )
                # shared reader: keep original owner, just record the reference.
                self._buffers.append(arg)
            else:
                arg._owner_worker = self
                self._buffers.append(arg)
                # If the Buffer has no tile, pin it to the Worker's tile as a
                # convenience.  If the user pinned it explicitly to a neighbor
                # tile (AIE compute tiles can read N/S/E/W neighbors' L1
                # directly), honor that placement — Program.resolve discovers
                # the neighbor tile via Buffer.tiles().
                if arg._tile is None:
                    arg._tile = self._tile
        elif isinstance(arg, ScratchpadParameter):
            pass  # ScratchpadParameters are device-level symbols; no tile placement needed
        elif isinstance(arg, ObjectFifo):
            # This is an easy error to make, so we catch it early
            raise ValueError(
                "Cannot give an ObjectFifo directly to a worker; "
                "must give an ObjectFifoHandle obtained through "
                "ObjectFifo.prod() or ObjectFifo.cons()"
            )
        elif isinstance(arg, WorkerRuntimeBarrier):
            self._barriers.append(arg)

tile property

tile: Tile

The compute tile this Worker is placed on.

flat_fn_args property

flat_fn_args: list

fn_args with any nested lists/tuples flattened to their leaves.

Use this (not fn_args) when iterating to register/resolve individual arguments; fn_args keeps its structure for the core_fn call.

fifos property

fifos: list[ObjectFifoHandle]

Returns a list of ObjectFifoHandles given to the Worker via fn_args.

Returns:

Type Description
list[ObjectFifoHandle]

list[ObjectFifoHandle]: ObjectFifoHandles used by the Worker.

buffers property

buffers: list[Buffer]

Returns a list of Buffers given to the Worker via fn_args.

Returns:

Type Description
list[Buffer]

list[Buffer]: Buffer used by the Worker.

grid staticmethod

grid(
    rows: int,
    cols: int,
    factory: Callable[[int, int], Worker],
) -> list[list[Worker]]

Build a 2D grid of Workers; factory(r, c) returns one Worker.

Replaces the common pattern:

ws = [Worker(...) for i in range(R) for j in range(C)]
ws[i * C + j]  # 1-D index arithmetic

with:

ws = Worker.grid(R, C, lambda r, c: Worker(...))
ws[i][j]       # natural 2-D access

Parameters:

Name Type Description Default
rows int

Outer-dimension count (e.g. column index).

required
cols int

Inner-dimension count (e.g. channel index).

required
factory Callable[[int, int], Worker]

Called once per cell with (r, c); must return a Worker.

required

Returns:

Type Description
list[list[Worker]]

rows-by-cols nested list of Worker instances.

Source code in python/iron/worker.py
@staticmethod
def grid(
    rows: int,
    cols: int,
    factory: Callable[[int, int], "Worker"],
) -> list[list["Worker"]]:
    """Build a 2D grid of Workers; ``factory(r, c)`` returns one Worker.

    Replaces the common pattern:

    ```python
    ws = [Worker(...) for i in range(R) for j in range(C)]
    ws[i * C + j]  # 1-D index arithmetic
    ```

    with:

    ```python
    ws = Worker.grid(R, C, lambda r, c: Worker(...))
    ws[i][j]       # natural 2-D access
    ```

    Args:
        rows: Outer-dimension count (e.g. column index).
        cols: Inner-dimension count (e.g. channel index).
        factory: Called once per cell with ``(r, c)``; must return a Worker.

    Returns:
        ``rows``-by-``cols`` nested list of Worker instances.
    """
    return [[factory(r, c) for c in range(cols)] for r in range(rows)]

WorkerRuntimeBarrier

WorkerRuntimeBarrier(initial_value: int = 0)

A barrier allowing individual workers to synchronize with the runtime sequence.

Initialize a WorkerRuntimeBarrier.

Parameters:

Name Type Description Default
initial_value int

The initial lock value. Defaults to 0.

0
Source code in python/iron/worker.py
def __init__(self, initial_value: int = 0):
    """Initialize a WorkerRuntimeBarrier.

    Args:
        initial_value (int, optional): The initial lock value. Defaults to 0.
    """
    self.initial_value = initial_value
    self.worker_locks = []

wait_for_value

wait_for_value(value: int)

Wait for the barrier to be set to value.

Should be called from inside a core function.

Parameters:

Name Type Description Default
value int

The value to wait for.

required
Source code in python/iron/worker.py
def wait_for_value(self, value: int):
    """Wait for the barrier to be set to `value`.

    Should be called from inside a core function.

    Args:
        value (int): The value to wait for.
    """
    # Here this is assuming that the we are currently placing the last added lock
    # And therefore that wait_for_value operations are placed just after their corresponding Worker...
    # This is a pretty bad assumption, think about an alternative way to solve this
    if len(self.worker_locks) == 0:
        raise ValueError(
            "No workers have been registered for this barrier. Need to pass the barrier as an argument to the worker."
        )
    use_lock(self.worker_locks[-1], LockAction.Acquire, value=value)

set

set(value: int)

Set the barrier to value from within a runtime sequence body.

Parameters:

Name Type Description Default
value int

The value to set the barrier to.

required
Source code in python/iron/worker.py
def set(self, value: int):
    """Set the barrier to ``value`` from within a runtime sequence body.

    Args:
        value (int): The value to set the barrier to.
    """
    _BarrierSetOp(self, value).resolve()

release_with_value

release_with_value(value: int)

Release and decrement the barrier by value inside the core.

Parameters:

Name Type Description Default
value int

The value to decrement by in Release.

required
Source code in python/iron/worker.py
def release_with_value(self, value: int):
    """Release and decrement the barrier by `value` inside the core.

    Args:
        value (int): The value to decrement by in Release.
    """
    if len(self.worker_locks) == 0:
        raise ValueError(
            "No workers have been registered for this barrier. Need to pass the barrier as an argument to the worker."
        )
    use_lock(self.worker_locks[-1], LockAction.Release, value=value)

ObjectFifo

ObjectFifo

ObjectFifo(
    obj_type: type[ndarray],
    *,
    depth: int | None = 2,
    name: str | None = None,
    dims_to_stream: StreamDims | None = None,
    dims_from_stream_per_cons: StreamDims | None = None,
    plio: bool = False,
    pad_dimensions: PadDims | None = None,
    pad_value: int = 0,
    disable_synchronization: bool = False,
    repeat_count: int | None = None,
    delegate_tile: Tile | None = None,
    via_DMA: bool = False,
    init_values: list[ndarray] | None = None,
    consumer_obj_type: type[ndarray] | None = None,
    aie_stream: tuple[int, int] | None = None,
    packet: bool = False,
    packet_id: int | None = None
)

Bases: Resolvable

A synchronized, explicit dataflow channel between IRON program components such as Workers and the Runtime.

Internally, an ObjectFifo is a circular buffer with a given depth and element type. Its users are explicitly either a producer or a consumer, and each user holds an ObjectFifoHandle carrying its (possibly unplaced) tile.

Example
of = ObjectFifo(np.ndarray[(1024,), np.dtype[np.int32]], name="in")
producer = of.prod()   # one producer handle
consumer = of.cons()   # one or more consumer handles

Construct an ObjectFifo.

Parameters:

Name Type Description Default
obj_type type[ndarray]

The type of each buffer in the ObjectFifo

required
depth int | None

The default depth of the ObjectFifo endpoints. Defaults to 2.

2
name str | None

The name of the ObjectFifo. If None is given, a unique name will be generated. Defaults to None.

None
dims_to_stream StreamDims | None

Data layout transformations applied when data is pushed onto the AXI stream, described as pairs of (size, stride) from highest to lowest dimension. Defaults to None.

None
dims_from_stream_per_cons StreamDims | None

List of data layout transformations applied by each consumer when data is read from the AXI stream, described as pairs of (size, stride) from highest to lowest dimension. Defaults to None.

None
plio bool

Whether the ObjectFifo uses PLIO connections. Defaults to False.

False
pad_dimensions PadDims | None

Per-dimension zero-padding applied to the buffer, described as (pad_before, pad_after) pairs from highest to lowest dimension. Lowers to the padDimensions attribute on the underlying aie.objectfifo op. Defaults to None.

None
pad_value int

Per-element constant value used to fill the region created by pad_dimensions. Packed into the raw 32-bit CONSTANT_PAD_VALUE register using this fifo's element width. Only valid together with pad_dimensions (MemTile padding). Defaults to 0.

0
disable_synchronization bool

When True, disables lock-based synchronization on the ObjectFifo. Defaults to False.

False
repeat_count int | None

If set, the sending end replicates each object this many times before moving on to the next one. The receiving end covers a whole batch in one acquire, so its lock initializers scale to match. Distinct from iter_count (BD-chain iteration count). Defaults to None.

None
delegate_tile Tile | None

Shared-memory delegate tile. When set, the ObjectFifo's underlying buffer pool is allocated on this tile's memory module instead of the default placement. Lowers to aie.objectfifo.allocate. Only valid when both producer and consumer have shared-memory access to the delegate tile (e.g. self-loop fifos where prod == cons, or fifos between adjacent tiles spilling to a neighboring MemTile). The delegate is the storage location, not a producer- or consumer-side concept; the underlying op verifier rejects this if either endpoint cannot share memory with the delegate. Defaults to None.

None
via_DMA bool

When True, force the ObjectFifo to route through DMA even when producer and consumer share memory (where a lock-only path would otherwise be used). Lowers to the via_DMA attribute on the underlying aie.objectfifo op. Defaults to False.

False
init_values list[ndarray] | None

Per-buffer static initial values for the producer endpoint. One ndarray per producer-side buffer; the producer tile must be able to hold static data at design startup (e.g. a MemTile). Lowers to the initValues attribute on the underlying aie.objectfifo op. Defaults to None.

None
consumer_obj_type type[ndarray] | None

Consumer element type for asymmetric transfer granularity. When set, the producer sends obj_type-sized transfers and the consumer receives consumer_obj_type-sized transfers. Producer element count must be an integer multiple of consumer element count. Defaults to None.

None
aie_stream tuple[int, int] | None

Mark the fifo as a direct AIE-stream connection by stamping the aie_stream / aie_stream_port attributes (end, port) on the underlying aie.objectfifo op. Use with kernels that emit on the wire via put_ms() instead of going through an L1 buffer. Defaults to None.

None
packet bool

Route this ObjectFifo as an aie.packet_flow, sharing the stream with other packet flows instead of reserving a circuit for it. Decided per fifo, so a design may mix packet- and circuit-switched fifos. Defaults to False.

False
packet_id int | None

Pin the 5-bit header the source stamps, for designs that route on the id (e.g. a MemTile dispatching to one of several cores). Requires packet; when absent, allocation picks an id no other flow is using. Defaults to None.

None

Raises:

Type Description
ValueError

If depth is provided and is less than 1.

Source code in python/iron/dataflow/objectfifo.py
def __init__(
    self,
    obj_type: type[np.ndarray],
    *,
    depth: int | None = 2,
    name: str | None = None,
    dims_to_stream: StreamDims | None = None,
    dims_from_stream_per_cons: StreamDims | None = None,
    plio: bool = False,
    pad_dimensions: PadDims | None = None,
    pad_value: int = 0,
    disable_synchronization: bool = False,
    repeat_count: int | None = None,
    delegate_tile: Tile | None = None,
    via_DMA: bool = False,
    init_values: list[np.ndarray] | None = None,
    consumer_obj_type: type[np.ndarray] | None = None,
    aie_stream: tuple[int, int] | None = None,
    packet: bool = False,
    packet_id: int | None = None,
):
    """Construct an ObjectFifo.

    Args:
        obj_type (type[np.ndarray]): The type of each buffer in the ObjectFifo
        depth (int | None, optional): The default depth of the ObjectFifo endpoints. Defaults to 2.
        name (str | None, optional): The name of the ObjectFifo. If None is given, a unique name will be generated. Defaults to None.
        dims_to_stream (StreamDims | None, optional): Data layout transformations applied
            when data is pushed onto the AXI stream, described as pairs of (size, stride)
            from highest to lowest dimension. Defaults to None.
        dims_from_stream_per_cons (StreamDims | None, optional): List of data layout
            transformations applied by each consumer when data is read from the AXI
            stream, described as pairs of (size, stride) from highest to lowest
            dimension. Defaults to None.
        plio (bool, optional): Whether the ObjectFifo uses PLIO connections. Defaults to False.
        pad_dimensions (PadDims | None, optional): Per-dimension zero-padding applied to the
            buffer, described as ``(pad_before, pad_after)`` pairs from highest to lowest
            dimension. Lowers to the ``padDimensions`` attribute on the underlying
            ``aie.objectfifo`` op. Defaults to None.
        pad_value (int, optional): Per-element constant value used to fill the region created
            by ``pad_dimensions``. Packed into the raw 32-bit CONSTANT_PAD_VALUE register
            using this fifo's element width. Only valid together with ``pad_dimensions`` (MemTile
            padding). Defaults to 0.
        disable_synchronization (bool, optional): When True, disables lock-based
            synchronization on the ObjectFifo. Defaults to False.
        repeat_count (int | None, optional): If set, the sending end replicates
            each object this many times before moving on to the next one. The
            receiving end covers a whole batch in one acquire, so its lock
            initializers scale to match. Distinct from ``iter_count``
            (BD-chain iteration count). Defaults to None.
        delegate_tile (Tile | None, optional): Shared-memory delegate tile. When set, the
            ObjectFifo's underlying buffer pool is allocated on this tile's memory module
            instead of the default placement. Lowers to ``aie.objectfifo.allocate``. *Only
            valid when both producer and consumer have shared-memory access to the
            delegate tile* (e.g. self-loop fifos where prod == cons, or fifos between
            adjacent tiles spilling to a neighboring MemTile). The delegate is the storage
            location, not a producer- or consumer-side concept; the underlying op verifier
            rejects this if either endpoint cannot share memory with the delegate.
            Defaults to None.
        via_DMA (bool, optional): When True, force the ObjectFifo to route through DMA
            even when producer and consumer share memory (where a lock-only path would
            otherwise be used). Lowers to the ``via_DMA`` attribute on the underlying
            ``aie.objectfifo`` op. Defaults to False.
        init_values (list[np.ndarray] | None, optional): Per-buffer static initial values
            for the producer endpoint. One ndarray per producer-side buffer; the producer
            tile must be able to hold static data at design startup (e.g. a MemTile).
            Lowers to the ``initValues`` attribute on the underlying ``aie.objectfifo``
            op. Defaults to None.
        consumer_obj_type (type[np.ndarray] | None, optional): Consumer element type for
            asymmetric transfer granularity. When set, the producer sends obj_type-sized
            transfers and the consumer receives consumer_obj_type-sized transfers.
            Producer element count must be an integer multiple of consumer element count.
            Defaults to None.
        aie_stream (tuple[int, int] | None, optional): Mark the fifo as a direct
            AIE-stream connection by stamping the ``aie_stream`` / ``aie_stream_port``
            attributes ``(end, port)`` on the underlying ``aie.objectfifo`` op. Use with
            kernels that emit on the wire via ``put_ms()`` instead of going through an L1
            buffer. Defaults to None.
        packet (bool, optional): Route this ObjectFifo as an ``aie.packet_flow``, sharing
            the stream with other packet flows instead of reserving a circuit for it.
            Decided per fifo, so a design may mix packet- and circuit-switched fifos.
            Defaults to False.
        packet_id (int | None, optional): Pin the 5-bit header the source stamps, for
            designs that route on the id (e.g. a MemTile dispatching to one of several
            cores). Requires ``packet``; when absent, allocation picks an id no other
            flow is using. Defaults to None.

    Raises:
        ValueError: If ``depth`` is provided and is less than 1.
    """
    self._depth = depth
    if self._depth is not None and self._depth < 1:
        raise ValueError(
            f"Default ObjectFifo depth must be > 0, but got {self._depth}"
        )
    self._obj_type = obj_type
    self._dims_to_stream = dims_to_stream
    self._dims_from_stream_per_cons = dims_from_stream_per_cons
    self._plio = plio
    self._pad_dimensions = pad_dimensions
    self._pad_value = pad_value
    if name is None:
        self.name = f"of{next(ObjectFifo._of_index)}"
    else:
        self.name = name
    self._op: ObjectFifoCreateOp | None = None
    self._prod: ObjectFifoHandle | None = None
    self._cons: list[ObjectFifoHandle] = []
    self._resolving = False
    self._iter_count: int | None = None
    self._repeat_count: int | None = repeat_count
    self._disable_synchronization: bool = disable_synchronization
    # Delegate tile for shared-memory buffer placement (lowers to aie.objectfifo.allocate).
    # Must be resolved before resolve() runs — Program.resolve() picks this up via
    # ObjectFifo._delegate_tile when collecting tiles to assign MLIR ops to.
    self._delegate_tile: Tile | None = delegate_tile
    self._via_DMA: bool = via_DMA
    self._init_values: list[np.ndarray] | None = init_values
    self._consumer_obj_type: type[np.ndarray] | None = consumer_obj_type
    self._aie_stream: tuple[int, int] | None = aie_stream
    self._packet: bool = packet
    self._packet_id: int | None = packet_id

depth property

depth: int | None

The default depth of the ObjectFifo. This may be overridden by an ObjectFifoHandle upon construction.

dims_from_stream_per_cons property

dims_from_stream_per_cons: StreamDims | None

The default dimensions from stream per consumer value. This may be overridden by an ObjectFifoHandle of type consumer.

dims_to_stream property

dims_to_stream: StreamDims | None

The dimensions to stream value. This will be shared by the ObjectFifoHandle of type producer.

shape property

shape: Sequence[int]

The shape of each buffer belonging to the ObjectFifo.

dtype property

dtype: type[NpuDType]

The per-element data type of each element in each buffer belonging to the ObjectFifo.

obj_type property

obj_type: type[ndarray]

The tensor type of each buffer belonging to the ObjectFifo.

set_iter_count

set_iter_count(iter_count: int)

Set how many times each end of the ObjectFifo cycles through its buffers.

Parameters:

Name Type Description Default
iter_count int

Passes each end's BD chain makes before it stops. Both ends carry depth * repeat_count * iter_count objects, so a sending end replicating each object repeat_count times makes that many fewer passes than the receiving end opposite it. - Must be in range [1, 256]

required

Raises:

Type Description
ValueError

If iter_count is outside the valid range [1, 256]

Source code in python/iron/dataflow/objectfifo.py
def set_iter_count(self, iter_count: int):
    """Set how many times each end of the ObjectFifo cycles through its buffers.

    Args:
        iter_count (int): Passes each end's BD chain makes before it stops.
            Both ends carry ``depth * repeat_count * iter_count`` objects, so
            a sending end replicating each object ``repeat_count`` times
            makes that many fewer passes than the receiving end opposite it.
            - Must be in range [1, 256]

    Raises:
        ValueError: If iter_count is outside the valid range [1, 256]
    """
    if not iter_count or iter_count < 1 or iter_count > 256:
        raise ValueError("Iter count must be in [1, 256] range.")

    self._iter_count = iter_count

prod

prod(
    depth: int | None = None,
    channel: int | None = None,
    tile: Tile | None = None,
) -> ObjectFifoHandle

Return an ObjectFifoHandle of type producer.

Each ObjectFifo may have only one producer handle, so if one already exists, a new reference to this handle will be returned.

Parameters:

Name Type Description Default
depth int | None

The depth of the buffers at the endpoint corresponding to the producer handle. Defaults to None.

None
channel int | None

Pin the producer endpoint's DMA channel instead of first-free assignment. Defaults to None (auto-assign).

None
tile Tile | None

When this handle drives an ObjectFifo from the runtime (passed in Runtime fn_args), the shim tile its host-side DMA binds to. Defaults to None (any available shim tile).

None

Raises:

Type Description
ValueError

Arguments are validated

ValueError

If depth was not specified on ObjectFifo construction, depth must be specified here.

Returns:

Name Type Description
ObjectFifoHandle ObjectFifoHandle

The producer handle to this ObjectFifo.

Source code in python/iron/dataflow/objectfifo.py
def prod(
    self,
    depth: int | None = None,
    channel: int | None = None,
    tile: Tile | None = None,
) -> ObjectFifoHandle:
    """Return an ObjectFifoHandle of type producer.

    Each ObjectFifo may have only one producer handle, so if one already
    exists, a new reference to this handle will be returned.

    Args:
        depth (int | None, optional): The depth of the buffers at the endpoint corresponding to the producer handle. Defaults to None.
        channel (int | None, optional): Pin the producer endpoint's DMA channel instead of first-free assignment. Defaults to None (auto-assign).
        tile (Tile | None, optional): When this handle drives an ObjectFifo
            from the runtime (passed in ``Runtime`` ``fn_args``), the shim tile
            its host-side DMA binds to. Defaults to None (any available shim tile).

    Raises:
        ValueError: Arguments are validated
        ValueError: If depth was not specified on ObjectFifo construction, depth must be specified here.

    Returns:
        ObjectFifoHandle: The producer handle to this ObjectFifo.
    """
    if self._prod:
        if depth is None:
            if self._depth is None:
                raise ValueError("If depth is None, then depth must be specified.")
            else:
                depth = self._depth
        elif depth < 1:
            raise ValueError(f"Depth must be > 1, but got {depth}")
        if channel is not None and self._prod.channel != channel:
            raise ValueError(
                f"Producer handle for {self.name} already pinned to channel "
                f"{self._prod.channel}, cannot re-pin to {channel}."
            )
        if tile is not None and not _same_shim_pin(self._prod._shim_tile, tile):
            raise ValueError(
                f"Producer handle for {self.name} already pinned to shim tile "
                f"{self._prod._shim_tile}, cannot re-pin to {tile}."
            )
    else:
        self._prod = ObjectFifoHandle(self, True, depth, channel=channel, tile=tile)
    return self._prod

cons

cons(
    depth: int | None = None,
    dims_from_stream: StreamDims | None = None,
    channel: int | None = None,
    tile: Tile | None = None,
) -> ObjectFifoHandle

Return an ObjectFifoHandle of type consumer.

Each ObjectFifo may have multiple consumers, so this will return a new consumer handle every time it is called.

Parameters:

Name Type Description Default
depth int | None

The depth of the buffers at the endpoint corresponding to this consumer handle. Defaults to None.

None
dims_from_stream StreamDims | None

Dimensions from stream for this consumer. Defaults to None.

None
channel int | None

Pin this consumer endpoint's DMA channel instead of first-free assignment. Defaults to None (auto-assign).

None
tile Tile | None

When this handle drains an ObjectFifo to the runtime (passed in Runtime fn_args), the shim tile its host-side DMA binds to. Defaults to None (any available shim tile).

None

Raises:

Type Description
ValueError

Arguments are validated

Returns:

Name Type Description
ObjectFifoHandle ObjectFifoHandle

A consumer handle to this ObjectFifo.

Source code in python/iron/dataflow/objectfifo.py
def cons(
    self,
    depth: int | None = None,
    dims_from_stream: StreamDims | None = None,
    channel: int | None = None,
    tile: Tile | None = None,
) -> ObjectFifoHandle:
    """Return an ObjectFifoHandle of type consumer.

    Each ObjectFifo may have multiple consumers, so this will return a new
    consumer handle every time it is called.

    Args:
        depth (int | None, optional): The depth of the buffers at the endpoint corresponding to this consumer handle. Defaults to None.
        dims_from_stream (StreamDims | None, optional): Dimensions from stream for this consumer. Defaults to None.
        channel (int | None, optional): Pin this consumer endpoint's DMA channel instead of first-free assignment. Defaults to None (auto-assign).
        tile (Tile | None, optional): When this handle drains an ObjectFifo to
            the runtime (passed in ``Runtime`` ``fn_args``), the shim tile its
            host-side DMA binds to. Defaults to None (any available shim tile).

    Raises:
        ValueError: Arguments are validated

    Returns:
        ObjectFifoHandle: A consumer handle to this ObjectFifo.
    """
    if depth is None:
        if self._depth is None:
            raise ValueError("If depth is None, then depth must be specified.")
        else:
            depth = self._depth

    if dims_from_stream is None:
        dims_from_stream = self._dims_from_stream_per_cons
    self._cons.append(
        ObjectFifoHandle(
            self,
            is_prod=False,
            depth=depth,
            dims_from_stream=dims_from_stream,
            channel=channel,
            tile=tile,
        )
    )
    return self._cons[-1]

tiles

tiles(cons_only: bool = False) -> list[Tile]

Return the placement tiles corresponding to the endpoints of all handles of this ObjectFifo.

Raises:

Type Description
ValueError

A producer handle must be constructed.

ValueError

At least one consumer handle must be constructed.

Returns:

Type Description
list[Tile]

list[Tile]: A list of tiles of the endpoints of this ObjectFifo.

Source code in python/iron/dataflow/objectfifo.py
def tiles(self, cons_only: bool = False) -> list[Tile]:
    """Return the placement tiles corresponding to the endpoints of all handles of this ObjectFifo.

    Raises:
        ValueError: A producer handle must be constructed.
        ValueError: At least one consumer handle must be constructed.

    Returns:
        list[Tile]: A list of tiles of the endpoints of this ObjectFifo.
    """
    tiles = []
    if not cons_only:
        if self._prod is None:
            raise ValueError(
                "Cannot return prod.tile.op because prod was not created."
            )
        if self._prod.endpoint is None:
            raise ValueError(f"Prod endpoint not set for {self}")
        assert self._prod.endpoint.tile is not None
        tiles += [self._prod.endpoint.tile]
    if self._cons == []:
        raise ValueError("Cannot return cons tiles because cons were not created.")
    for cons in self._cons:
        if cons.endpoint is None:
            raise ValueError(f"Cons endpoint not set for {self}")
        assert cons.endpoint.tile is not None
        tiles.append(cons.endpoint.tile)
    return tiles

ObjectFifoHandle

ObjectFifoHandle(
    of: ObjectFifo,
    is_prod: bool,
    depth: int | None = None,
    dims_from_stream: StreamDims | None = None,
    channel: int | None = None,
    tile: Tile | None = None,
)

Bases: Resolvable

A handle to an ObjectFifo, of type producer or consumer.

Producer and consumer handles are what Worker core functions call acquire and release on to move data through the fifo. Obtain them via ObjectFifo.prod() and ObjectFifo.cons().

Construct an ObjectFifoHandle.

Parameters:

Name Type Description Default
of ObjectFifo

The ObjectFifo to construct the handle for.

required
is_prod bool

Whether the handle should be producer or consumer handle.

required
depth int | None

The depth of the ObjectFifo at this endpoint. Defaults to None.

None
dims_from_stream StreamDims | None

A unique dimensions from stream. This is only valid for consumer handles. Defaults to None.

None
channel int | None

Pin this endpoint's DMA channel instead of first-free assignment. Defaults to None (auto-assign).

None
tile Tile | None

Shim tile for a runtime-driven endpoint (see prod()/cons()). Defaults to None.

None

Raises:

Type Description
ValueError

Arguments are validated.

Source code in python/iron/dataflow/objectfifo.py
def __init__(
    self,
    of: ObjectFifo,
    is_prod: bool,
    depth: int | None = None,
    dims_from_stream: StreamDims | None = None,
    channel: int | None = None,
    tile: Tile | None = None,
):
    """Construct an ObjectFifoHandle.

    Args:
        of (ObjectFifo): The ObjectFifo to construct the handle for.
        is_prod (bool): Whether the handle should be producer or consumer handle.
        depth (int | None, optional): The depth of the ObjectFifo at this endpoint. Defaults to None.
        dims_from_stream (StreamDims | None, optional): A unique dimensions from stream. This is only valid for consumer handles. Defaults to None.
        channel (int | None, optional): Pin this endpoint's DMA channel instead of first-free assignment. Defaults to None (auto-assign).
        tile (Tile | None, optional): Shim tile for a runtime-driven endpoint (see prod()/cons()). Defaults to None.

    Raises:
        ValueError: Arguments are validated.
    """
    if depth is None:
        if of.depth:
            depth = of.depth
        else:
            raise ValueError(
                "Must specify either ObjectFifoHandle depth or ObjectFifo default depth; both are None."
            )
    if depth < 1:
        raise ValueError(f"Depth must be > 0 but is {depth}")
    self._port: ObjectFifoPort = (
        ObjectFifoPort.Produce if is_prod else ObjectFifoPort.Consume
    )
    if is_prod and dims_from_stream:
        raise ValueError("Can only specify dims_from_stream for cons handles")
    elif not is_prod and not dims_from_stream:
        dims_from_stream = of.dims_from_stream_per_cons

    self._is_prod = is_prod
    self._object_fifo = of
    self._depth = depth
    self._channel = channel
    self._shim_tile = tile
    self._endpoint = None
    self._dims_from_stream = dims_from_stream

name property

name: str

The name of the ObjectFifo.

channel property

channel: int | None

The pinned DMA channel for this handle's endpoint, or None to auto-assign.

obj_type property

obj_type: type[ndarray]

The per-buffer type of the ObjectFifo.

shape property

shape: Sequence[int]

The per-buffer shape of the ObjectFifo.

dtype property

dtype: type[NpuDType]

The per-element datatype of the ObjectFifo.

handle_type property

handle_type: str

A string referencing the type of this ObjectFifoHandle.

depth property

depth: int

The depth of this ObjectFifoHandle.

dims_from_stream property

dims_from_stream: StreamDims | None

The dimensions from stream of a consumer ObjectFifoHandle.

endpoint property writable

endpoint: ObjectFifoEndpoint | None

The endpoint of this ObjectFifoHandle.

acquire

acquire(num_elem: int) -> list

Acquire access to some elements of the ObjectFifo, using ObjectFifo synchronization to moderate access.

Parameters:

Name Type Description Default
num_elem int

Number of elements to acquire. If some elements are already acquired, only the additional elements needed to reach a total of num_elem are acquired.

required

Raises:

Type Description
ValueError

Number of elements cannot exceed ObjectFifo depth.

Returns:

Type Description
list

An indexable handle to the acquired elements: a single element when num_elem == 1, or an indexable view when num_elem > 1.

Source code in python/iron/dataflow/objectfifo.py
def acquire(
    self,
    num_elem: int,
) -> list:
    """Acquire access to some elements of the ObjectFifo, using ObjectFifo synchronization to moderate access.

    Args:
        num_elem (int): Number of elements to acquire. If some elements are already
            acquired, only the additional elements needed to reach a total of
            ``num_elem`` are acquired.

    Raises:
        ValueError: Number of elements cannot exceed ObjectFifo depth.

    Returns:
        An indexable handle to the acquired elements: a single element when ``num_elem == 1``, or an indexable view when ``num_elem > 1``.
    """
    if self._depth < num_elem:
        raise ValueError(
            f"Number of elements to acquire {num_elem} must be smaller than depth {self._depth}"
        )
    return self._object_fifo._acquire(self._port, num_elem)

release

release(num_elem: int) -> None

Release access to some elements of the ObjectFifo, allowing the other endpoint of the ObjectFifo to acquire them.

Parameters:

Name Type Description Default
num_elem int

Number of elements to release.

required

Raises:

Type Description
ValueError

Number of elements cannot exceed ObjectFifo depth.

Source code in python/iron/dataflow/objectfifo.py
def release(
    self,
    num_elem: int,
) -> None:
    """Release access to some elements of the ObjectFifo, allowing the other endpoint of the ObjectFifo to acquire them.

    Args:
        num_elem (int): Number of elements to release.

    Raises:
        ValueError: Number of elements cannot exceed ObjectFifo depth.
    """
    if self._depth < num_elem:
        raise ValueError(
            f"Number of elements to release {num_elem} must be smaller than depth {self._depth}"
        )
    self._object_fifo._release(self._port, num_elem)

fill

fill(
    source,
    tap=None,
    wait: bool = False,
    packet: tuple[int, int] | None = None,
    offset_parameter=None,
    group=None,
    sizes=None,
    strides=None,
    offset=None,
    transfer_len=None,
    managed: bool = True,
)

Fill this producer ObjectFifo with data from the source runtime buffer.

Call from within a Runtime sequence body on a producer handle. See _emit_transfer for the shared arguments; returns a Task handle to the transfer.

Source code in python/iron/dataflow/objectfifo.py
def fill(
    self,
    source,
    tap=None,
    wait: bool = False,
    packet: tuple[int, int] | None = None,
    offset_parameter=None,
    group=None,
    sizes=None,
    strides=None,
    offset=None,
    transfer_len=None,
    managed: bool = True,
):
    """Fill this producer ObjectFifo with data from the ``source`` runtime buffer.

    Call from within a [`Runtime`][iron.Runtime] sequence body on a producer
    handle. See ``_emit_transfer`` for the shared arguments; returns a
    [`Task`][iron.runtime.dmataskhandle.Task] handle to the transfer.
    """
    if not self._is_prod:
        raise ValueError("fill() is only valid on a producer ObjectFifoHandle")

    return self._emit_transfer(
        source,
        tap,
        wait,
        packet,
        offset_parameter,
        group,
        sizes,
        strides,
        offset,
        transfer_len,
        managed,
    )

drain

drain(
    dest,
    tap=None,
    wait: bool = False,
    packet: tuple[int, int] | None = None,
    offset_parameter=None,
    group=None,
    sizes=None,
    strides=None,
    offset=None,
    transfer_len=None,
    managed: bool = True,
)

Drain this consumer ObjectFifo, writing data to the dest runtime buffer.

Call from within a Runtime sequence body on a consumer handle. See _emit_transfer for the shared arguments; returns a Task handle to the transfer.

Source code in python/iron/dataflow/objectfifo.py
def drain(
    self,
    dest,
    tap=None,
    wait: bool = False,
    packet: tuple[int, int] | None = None,
    offset_parameter=None,
    group=None,
    sizes=None,
    strides=None,
    offset=None,
    transfer_len=None,
    managed: bool = True,
):
    """Drain this consumer ObjectFifo, writing data to the ``dest`` runtime buffer.

    Call from within a [`Runtime`][iron.Runtime] sequence body on a consumer
    handle. See ``_emit_transfer`` for the shared arguments; returns a
    [`Task`][iron.runtime.dmataskhandle.Task] handle to the transfer.
    """
    if self._is_prod:
        raise ValueError("drain() is only valid on a consumer ObjectFifoHandle")

    return self._emit_transfer(
        dest,
        tap,
        wait,
        packet,
        offset_parameter,
        group,
        sizes,
        strides,
        offset,
        transfer_len,
        managed,
    )

all_of_endpoints

all_of_endpoints() -> list[ObjectFifoEndpoint]

All endpoints belonging to an ObjectFifo.

Source code in python/iron/dataflow/objectfifo.py
def all_of_endpoints(self) -> list[ObjectFifoEndpoint]:
    """All endpoints belonging to an ObjectFifo."""
    return self._object_fifo._get_endpoint(
        is_prod=True
    ) + self._object_fifo._get_endpoint(is_prod=False)

join

join(
    offsets: list[int],
    tile: Tile | None = AnyMemTile,
    depths: list[int] | None = None,
    obj_types: list[type[ndarray]] | None = None,
    names: list[str] | None = None,
    dims_to_stream: list[StreamDims] | None = None,
    dims_from_stream: list[StreamDims] | None = None,
    plio: bool = False,
    repeat_counts: list[int | None] | None = None,
) -> list[ObjectFifo]

Construct multiple ObjectFifos which feed data into a ObjectFifoHandle.

Note that this function is only valid for producer ObjectFifoHandles.

Parameters:

Name Type Description Default
offsets list[int]

Offsets into the current producer, each corresponding to a new consumer.

required
tile Tile

The tile where the Join operation occurs. Also accepts None (treated as AnyMemTile). Defaults to AnyMemTile.

AnyMemTile
depths list[int] | None

The depth of each new ObjectFifo. Defaults to None.

None
obj_types list[type[ndarray]]

The type of the buffers corresponding to each new ObjectFifo. Defaults to None.

None
names list[str] | None

The name of each new ObjectFifo. If not given, unique names will be generated. Defaults to None.

None
dims_to_stream list[list[Sequence[int] | None]] | None

The dimensionsToStream to assign to each new ObjectFifo. Defaults to None.

None
dims_from_stream list[list[Sequence[int] | None]] | None

The dimensionsFromStream to assign to each new ObjectFifo consumer. Defaults to None.

None
plio bool

Set plio on each new ObjectFifo. Defaults to False.

False
repeat_counts list[int | None] | None

Per-sub-fifo MemTile DMA repeat count (see ObjectFifo.repeat_count). Defaults to None.

None

Raises:

Type Description
ValueError

Arguments are validated

Returns:

Type Description
list[ObjectFifo]

list[ObjectFifo]: A list of newly constructed ObjectFifos whose consumers are used in this join() operation.

Source code in python/iron/dataflow/objectfifo.py
def join(
    self,
    offsets: list[int],
    tile: Tile | None = AnyMemTile,
    depths: list[int] | None = None,
    obj_types: list[type[np.ndarray]] | None = None,
    names: list[str] | None = None,
    dims_to_stream: list[StreamDims] | None = None,
    dims_from_stream: list[StreamDims] | None = None,
    plio: bool = False,
    repeat_counts: list[int | None] | None = None,
) -> list[ObjectFifo]:
    """Construct multiple ObjectFifos which feed data into a ObjectFifoHandle.

    Note that this function is only valid for producer ObjectFifoHandles.

    Args:
        offsets (list[int]): Offsets into the current producer, each corresponding to a new consumer.
        tile (Tile, optional): The tile where the Join operation occurs. Also accepts None (treated as AnyMemTile). Defaults to AnyMemTile.
        depths (list[int] | None, optional): The depth of each new ObjectFifo. Defaults to None.
        obj_types (list[type[np.ndarray]], optional): The type of the buffers corresponding to each new ObjectFifo. Defaults to None.
        names (list[str] | None, optional): The name of each new ObjectFifo. If not given, unique names will be generated. Defaults to None.
        dims_to_stream (list[list[Sequence[int]  |  None]] | None, optional): The dimensionsToStream to assign to each new ObjectFifo. Defaults to None.
        dims_from_stream (list[list[Sequence[int]  |  None]] | None, optional): The
            dimensionsFromStream to assign to each new ObjectFifo consumer. Defaults
            to None.
        plio (bool, optional): Set plio on each new ObjectFifo. Defaults to False.
        repeat_counts (list[int | None] | None, optional): Per-sub-fifo MemTile DMA repeat count (see ObjectFifo.repeat_count). Defaults to None.

    Raises:
        ValueError: Arguments are validated

    Returns:
        list[ObjectFifo]: A list of newly constructed ObjectFifos whose consumers are used in this join() operation.
    """
    if not self._is_prod:
        raise ValueError(f"Cannot join() a {self.handle_type} ObjectFifoHandle")
    num_subfifos = len(offsets)
    if depths is None:
        depths = [self.depth] * num_subfifos
    elif len(depths) != num_subfifos:
        raise ValueError("Number of depths does not match number of offsets")

    if obj_types is None:
        obj_types = [self._object_fifo.obj_type] * num_subfifos
    elif len(obj_types) != num_subfifos:
        raise ValueError("Number of obj_types does not match number of offsets")

    if names is None:
        names = [self._object_fifo.name + f"_join{i}" for i in range(num_subfifos)]
    elif len(names) != num_subfifos:
        raise ValueError("Number of names does not match number of offsets")

    if dims_to_stream is None:
        dims_to_stream = [[]] * num_subfifos
    elif len(dims_to_stream) != num_subfifos:
        raise ValueError(
            "Number of dims to stream does not match number of offsets"
        )

    if dims_from_stream is None:
        dims_from_stream = [[]] * num_subfifos
    elif dims_from_stream and len(dims_from_stream) != num_subfifos:
        raise ValueError(
            "Number of dims_from_stream does not match number of offsets"
        )

    if repeat_counts is None:
        repeat_counts = [None for _ in range(num_subfifos)]
    elif len(repeat_counts) != num_subfifos:
        raise ValueError("Number of repeat_counts does not match number of offsets")

    # Create subfifos
    subfifos = []
    for i in range(num_subfifos):
        subfifos.append(
            ObjectFifo(
                obj_types[i],
                name=names[i],
                depth=depths[i],
                dims_to_stream=dims_to_stream[i],
                plio=plio,
                repeat_count=repeat_counts[i],
            )
        )

    subfifo_cons = [
        s.cons(depth=depths[i], dims_from_stream=dims_from_stream[i])
        for i, s in enumerate(subfifos)
    ]
    _ = ObjectFifoLink(subfifo_cons, self, tile, offsets, [])
    return subfifos

split

split(
    offsets: list[int],
    tile: Tile | None = AnyMemTile,
    depths: list[int] | None = None,
    obj_types: list[type[ndarray]] | None = None,
    names: list[str] | None = None,
    dims_to_stream: list[StreamDims] | None = None,
    dims_from_stream: list[StreamDims] | None = None,
    plio: bool = False,
    repeat_counts: list[int | None] | None = None,
    pad_dimensions: list[PadDims | None] | None = None,
    pad_value: list[int] | None = None,
    channels: list[int | None] | None = None,
) -> list[ObjectFifo]

Split the data from an ObjectFifoConsumer handle by sending it to producers in N newly constructed ObjectFifos.

Note this operation is only valid for ObjectFifoHandles of type consumer.

Parameters:

Name Type Description Default
offsets list[int]

The offset into the current consumer for each new ObjectFifo producer.

required
tile Tile

The tile where the Split operation takes place. Also accepts None (treated as AnyMemTile). Defaults to AnyMemTile.

AnyMemTile
depths list[int] | None

The depth of each new ObjectFifo. Defaults to None.

None
obj_types list[type[ndarray]]

The buffer type of each new ObjectFifo. Defaults to None.

None
names list[str] | None

The name of each new ObjectFifo. If not given, a unique name will be generated. Defaults to None.

None
dims_to_stream list[StreamDims] | None

The dimensions to stream for each new ObjectFifo. Defaults to None.

None
dims_from_stream list[StreamDims] | None

The dimensions from stream for each new ObjectFifo. Defaults to None.

None
plio bool

Set plio on each new ObjectFifo. Defaults to False.

False
repeat_counts list[int | None] | None

Per-sub-fifo MemTile DMA repeat count (see ObjectFifo.repeat_count). Defaults to None.

None
pad_dimensions list[PadDims | None] | None

Per-sub-fifo (before, after) pad counts (see ObjectFifo.pad_dimensions). Defaults to None.

None
pad_value list[int] | None

Per-sub-fifo per-element pad fill value (see ObjectFifo.pad_value). Defaults to None.

None
channels list[int | None] | None

Pin the hardware DMA channel each output ObjectFifo produces on, one per output. split() builds those producer handles itself, so this is the only place to say it. Defaults to None (all compiler-assigned).

None

Raises:

Type Description
ValueError

Arguments are validated.

Returns:

Type Description
list[ObjectFifo]

list[ObjectFifo]: A list of newly constructed ObjectFifos whose producers are used in this split() operation.

Source code in python/iron/dataflow/objectfifo.py
def split(
    self,
    offsets: list[int],
    tile: Tile | None = AnyMemTile,
    depths: list[int] | None = None,
    obj_types: list[type[np.ndarray]] | None = None,
    names: list[str] | None = None,
    dims_to_stream: list[StreamDims] | None = None,
    dims_from_stream: list[StreamDims] | None = None,
    plio: bool = False,
    repeat_counts: list[int | None] | None = None,
    pad_dimensions: list[PadDims | None] | None = None,
    pad_value: list[int] | None = None,
    channels: list[int | None] | None = None,
) -> list[ObjectFifo]:
    """Split the data from an ObjectFifoConsumer handle by sending it to producers in N newly constructed ObjectFifos.

    Note this operation is only valid for ObjectFifoHandles of type consumer.

    Args:
        offsets (list[int]): The offset into the current consumer for each new ObjectFifo producer.
        tile (Tile, optional): The tile where the Split operation takes place. Also accepts None (treated as AnyMemTile). Defaults to AnyMemTile.
        depths (list[int] | None, optional): The depth of each new ObjectFifo. Defaults to None.
        obj_types (list[type[np.ndarray]], optional): The buffer type of each new ObjectFifo. Defaults to None.
        names (list[str] | None, optional): The name of each new ObjectFifo. If not given, a unique name will be generated. Defaults to None.
        dims_to_stream (list[StreamDims] | None, optional): The dimensions to stream for each new ObjectFifo. Defaults to None.
        dims_from_stream (list[StreamDims] | None, optional): The dimensions from stream for each new ObjectFifo. Defaults to None.
        plio (bool, optional): Set plio on each new ObjectFifo. Defaults to False.
        repeat_counts (list[int | None] | None, optional): Per-sub-fifo MemTile DMA repeat count (see ObjectFifo.repeat_count). Defaults to None.
        pad_dimensions (list[PadDims | None] | None, optional): Per-sub-fifo (before, after) pad counts (see ObjectFifo.pad_dimensions). Defaults to None.
        pad_value (list[int] | None, optional): Per-sub-fifo per-element pad fill value (see ObjectFifo.pad_value). Defaults to None.

        channels (list[int | None] | None, optional): Pin the hardware DMA
            channel each output ObjectFifo produces on, one per output.
            split() builds those producer handles itself, so this is the
            only place to say it. Defaults to None (all compiler-assigned).

    Raises:
        ValueError: Arguments are validated.

    Returns:
        list[ObjectFifo]: A list of newly constructed ObjectFifos whose producers are used in this split() operation.
    """
    if self._is_prod:
        raise ValueError(f"Cannot split() a {self.handle_type} ObjectFifoHandle")
    num_subfifos = len(offsets)
    if depths is None:
        depths = [self.depth] * num_subfifos
    elif len(depths) != num_subfifos:
        raise ValueError("Number of depths does not match number of offsets")

    if obj_types is None:
        obj_types = [self._object_fifo.obj_type] * num_subfifos
    elif len(obj_types) != num_subfifos:
        raise ValueError("Number of obj_types does not match number of offsets")

    if names is None:
        names = [self._object_fifo.name + f"_split{i}" for i in range(num_subfifos)]
    elif len(names) != num_subfifos:
        raise ValueError("Number of names does not match number of offsets")

    if dims_to_stream is None:
        dims_to_stream = [[]] * num_subfifos
    elif len(dims_to_stream) != num_subfifos:
        raise ValueError(
            "Number of dims_to_stream arrays does not match number of offsets"
        )

    if dims_from_stream is None:
        dims_from_stream = [[]] * num_subfifos
    elif len(dims_from_stream) != num_subfifos:
        raise ValueError(
            "Number of dims_from_stream arrays does not match number of offsets"
        )

    if repeat_counts is None:
        repeat_counts = [None for _ in range(num_subfifos)]
    elif len(repeat_counts) != num_subfifos:
        raise ValueError("Number of repeat_counts does not match number of offsets")

    if pad_dimensions is None:
        pad_dimensions = [None for _ in range(num_subfifos)]
    elif len(pad_dimensions) != num_subfifos:
        raise ValueError(
            "Number of pad_dimensions does not match number of offsets"
        )

    if pad_value is None:
        pad_value = [0 for _ in range(num_subfifos)]
    elif len(pad_value) != num_subfifos:
        raise ValueError("Number of pad_value does not match number of offsets")

    # Create subfifos
    subfifos = []
    for i in range(num_subfifos):
        subfifos.append(
            ObjectFifo(
                obj_types[i],
                name=names[i],
                depth=depths[i],
                dims_to_stream=dims_to_stream[i],
                dims_from_stream_per_cons=dims_from_stream[i],
                plio=plio,
                repeat_count=repeat_counts[i],
                pad_dimensions=pad_dimensions[i],
                pad_value=pad_value[i],
            )
        )

    # Create link and set it as endpoints
    pinned: list[int | None] = (
        [None] * len(subfifos) if channels is None else list(channels)
    )
    if len(pinned) != len(subfifos):
        raise ValueError(
            f"split() got {len(pinned)} channels for {len(subfifos)} "
            "outputs; give one per output or none at all."
        )
    # A subfifo's producer handle is built here, so a caller wanting its
    # channel pinned has nowhere else to say it -- prod() refuses to re-pin.
    subfifo_prods = [s.prod(channel=c) for s, c in zip(subfifos, pinned)]
    _ = ObjectFifoLink(self, subfifo_prods, tile, [], offsets)
    return subfifos

forward

forward(
    tile: Tile | None = AnyMemTile,
    obj_type: type[ndarray] | None = None,
    depth: int | None = None,
    name: str | None = None,
    dims_to_stream: StreamDims | None = None,
    dims_from_stream: StreamDims | None = None,
    plio: bool = False,
    repeat_count: int | None = None,
    pad_dimensions: PadDims | None = None,
    pad_value: int = 0,
    channel: int | None = None,
) -> ObjectFifo

Forward an ObjectFifoHandle of type consumer to a newly-constructed ObjectFifo.

This is a special case of the split() operation where the consumer handle is forwarded to the producer of a newly-constructed ObjectFifo.

Parameters:

Name Type Description Default
tile Tile

The tile for the Forward operation. Also accepts None (treated as AnyMemTile). Defaults to AnyMemTile.

AnyMemTile
obj_type type[ndarray] | None

The object type of the new ObjectFifo. Defaults to None.

None
depth int | None

The depth of the new ObjectFifo. Defaults to None.

None
name str | None

The name of the new ObjectFifo. If None is given, a unique name will be generated. Defaults to None.

None
dims_to_stream StreamDims | None

The dimensions to stream for the new ObjectFifo. Defaults to None.

None
dims_from_stream StreamDims | None

The dimensions from stream for the new ObjectFifo. Defaults to None.

None
plio bool

Set plio on each new ObjectFifo. Defaults to False.

False
repeat_count int | None

MemTile DMA repeat count for the new ObjectFifo (see ObjectFifo.repeat_count). Defaults to None.

None
pad_dimensions PadDims | None

Per-dimension (before, after) constant-pad counts for the forwarded (memtile) ObjectFifo. Defaults to None.

None
pad_value int

Per-element constant fill value for pad_dimensions (see ObjectFifo.pad_value). Defaults to 0.

0
channel int | None

Pin the hardware DMA channel the forwarded ObjectFifo produces on. forward() builds that producer handle itself, so this is the only place to say it. Defaults to None (assigned by the compiler).

None

Raises:

Type Description
ValueError

Arguments are Validated

Returns:

Name Type Description
ObjectFifo ObjectFifo

A newly constructed ObjectFifo whose producer used in this forward() operation.

Source code in python/iron/dataflow/objectfifo.py
def forward(
    self,
    tile: Tile | None = AnyMemTile,
    obj_type: type[np.ndarray] | None = None,
    depth: int | None = None,
    name: str | None = None,
    dims_to_stream: StreamDims | None = None,
    dims_from_stream: StreamDims | None = None,
    plio: bool = False,
    repeat_count: int | None = None,
    pad_dimensions: PadDims | None = None,
    pad_value: int = 0,
    channel: int | None = None,
) -> ObjectFifo:
    """Forward an ObjectFifoHandle of type consumer to a newly-constructed ObjectFifo.

    This is a special case of the split() operation where the consumer handle
    is forwarded to the producer of a newly-constructed ObjectFifo.

    Args:
        tile (Tile, optional): The tile for the Forward operation. Also accepts None (treated as AnyMemTile). Defaults to AnyMemTile.
        obj_type (type[np.ndarray] | None, optional): The object type of the new ObjectFifo. Defaults to None.
        depth (int | None, optional): The depth of the new ObjectFifo. Defaults to None.
        name (str | None, optional): The name of the new ObjectFifo. If None is given, a unique name will be generated. Defaults to None.
        dims_to_stream (StreamDims | None, optional): The dimensions to stream for the new ObjectFifo. Defaults to None.
        dims_from_stream (StreamDims | None, optional): The dimensions from stream for the new ObjectFifo. Defaults to None.
        plio (bool, optional): Set plio on each new ObjectFifo. Defaults to False.
        repeat_count (int | None, optional): MemTile DMA repeat count for the new ObjectFifo (see ObjectFifo.repeat_count). Defaults to None.
        pad_dimensions (PadDims | None, optional): Per-dimension (before, after) constant-pad
            counts for the forwarded (memtile) ObjectFifo. Defaults to None.
        pad_value (int, optional): Per-element constant fill value for pad_dimensions (see
            ObjectFifo.pad_value). Defaults to 0.
        channel (int | None, optional): Pin the hardware DMA channel the
            forwarded ObjectFifo produces on. forward() builds that
            producer handle itself, so this is the only place to say it.
            Defaults to None (assigned by the compiler).

    Raises:
        ValueError: Arguments are Validated

    Returns:
        ObjectFifo: A newly constructed ObjectFifo whose producer used in this forward() operation.
    """
    if self._is_prod:
        raise ValueError(f"Cannot forward a {self.handle_type} ObjectFifoHandle")
    obj_types = [obj_type] if obj_type else None
    depths = [depth] if depth else None
    names = [name] if name else [self._object_fifo.name + "_fwd"]
    dims_to_stream_arg = [dims_to_stream] if dims_to_stream else None
    dims_from_stream_arg = [dims_from_stream] if dims_from_stream else None

    forward_fifo = self.split(
        [0],
        tile=tile,
        obj_types=obj_types,
        depths=depths,
        names=names,
        dims_to_stream=dims_to_stream_arg,
        dims_from_stream=dims_from_stream_arg,
        plio=plio,
        repeat_counts=[repeat_count] if repeat_count is not None else None,
        pad_dimensions=[pad_dimensions] if pad_dimensions is not None else None,
        pad_value=[pad_value] if pad_value else None,
        channels=[channel] if channel is not None else None,
    )
    return forward_fifo[0]
ObjectFifoLink(
    srcs: list[ObjectFifoHandle] | ObjectFifoHandle,
    dsts: list[ObjectFifoHandle] | ObjectFifoHandle,
    tile: Tile | None = AnyMemTile,
    src_offsets: list[int] | None = None,
    dst_offsets: list[int] | None = None,
)

Bases: ObjectFifoEndpoint, Resolvable

This is an object used internally by split(), join() and forward() operations.

Construct an ObjectFifoLink. This is either a many-to-one, one-to-many, or one-to-one operation.

Parameters:

Name Type Description Default
srcs list[ObjectFifoHandle] | ObjectFifoHandle

A list of consumer ObjectFifoHandles to link.

required
dsts list[ObjectFifoHandle] | ObjectFifoHandle

A list of producer ObjectFifoHandles to link.

required
tile Tile

The tile where the link occurs. Also accepts None (treated as AnyMemTile). Defaults to AnyMemTile.

AnyMemTile
src_offsets list[int] | None

If many sources, one offset per source is required to split the destination. Defaults to None (empty list).

None
dst_offsets list[int] | None

If many destinations, one offset per destination is required to split the source. Defaults to None (empty list).

None

Raises:

Type Description
ValueError

Arguments are validated.

Source code in python/iron/dataflow/objectfifo.py
def __init__(
    self,
    srcs: list[ObjectFifoHandle] | ObjectFifoHandle,
    dsts: list[ObjectFifoHandle] | ObjectFifoHandle,
    tile: Tile | None = AnyMemTile,
    src_offsets: list[int] | None = None,
    dst_offsets: list[int] | None = None,
):
    """Construct an ObjectFifoLink. This is either a many-to-one, one-to-many, or one-to-one operation.

    Args:
        srcs (list[ObjectFifoHandle] | ObjectFifoHandle): A list of consumer ObjectFifoHandles to link.
        dsts (list[ObjectFifoHandle] | ObjectFifoHandle): A list of producer ObjectFifoHandles to link.
        tile (Tile, optional): The tile where the link occurs. Also accepts None (treated as AnyMemTile). Defaults to AnyMemTile.
        src_offsets (list[int] | None, optional): If many sources, one offset per source
            is required to split the destination. Defaults to None (empty list).
        dst_offsets (list[int] | None, optional): If many destinations, one offset per
            destination is required to split the source. Defaults to None (empty list).

    Raises:
        ValueError: Arguments are validated.
    """
    self._srcs = single_elem_or_list_to_list(srcs)
    self._dsts = single_elem_or_list_to_list(dsts)
    self._src_offsets = src_offsets if src_offsets is not None else []
    self._dst_offsets = dst_offsets if dst_offsets is not None else []
    self._resolving = False

    if len(self._srcs) < 1:
        raise ValueError("An ObjectFifoLink must have at least one source")
    if len(self._dsts) < 1:
        raise ValueError("An ObjectFifoLink must have at least one destination")
    if len(self._srcs) != 1 and len(self._dsts) != 1:
        raise ValueError(
            "An ObjectFifoLink may only have > 1 of either sources or destinations, but not both"
        )
    if len(self._src_offsets) > 0 and len(self._src_offsets) != len(self._srcs):
        raise ValueError(
            "The number of source offsets does not match the number of sources"
        )
    if len(self._dst_offsets) > 0 and len(self._dst_offsets) != len(self._dsts):
        raise ValueError(
            "The number of destination offsets does not match the number of destinations"
        )
    self._op = None
    for s in self._srcs:
        s.endpoint = self
    for d in self._dsts:
        d.endpoint = self
    if tile is None:
        tile = AnyMemTile
    # Isolate singleton defaults, but retain user tiles shared with Workers
    # or other links so they resolve to the same logical tile.
    if any(
        tile is default for default in (AnyMemTile, AnyComputeTile, AnyShimTile)
    ):
        tile = tile.copy()
    # Respect explicit types and let the device infer fully placed tiles.
    placed = tile.col is not None and tile.row is not None
    if tile.tile_type is None and not placed:
        tile.tile_type = AIETileType.MemTile
    ObjectFifoEndpoint.__init__(self, tile)

Runtime

The host-side orchestration entry point. Calls to producer-handle fill and consumer-handle drain are declared in the sequence body passed to Runtime(seq, fn_args); Workers are passed to Program(workers=...).

Runtime: orchestrates host-side data movement and worker execution for an IRON program.

The runtime sequence is written as a callback body -- a plain Python function whose parameters are the runtime I/O buffers (and, optionally, runtime scalars). The body runs eagerly inside @runtime_sequence at resolve time, mirroring how Worker.core_fn runs inside @core. Because the body executes with live MLIR values in scope, it can use native range_/if_ control flow with fill/drain verbs nested inside -- the dynamic path lowers these to scf.for/scf.if (EmitC C++ TXN), and the static path (Python range/int bounds) elaborates to a flat binary sequence.

IronRuntimeError

Bases: Exception

Raised by the IRON Runtime when resolution encounters an unrecoverable state.

ActiveSequence

ActiveSequence(runtime: 'Runtime')

The state of a runtime sequence body while it is being emitted.

The body's data-movement verbs (fifo.fill/fifo.drain) and TaskGroup reach this object through the active-sequence ContextVar (see _context) rather than a threaded rt reference, so the body signature carries only the runtime buffers.

The body runs exactly once, inside the runtime_sequence op: each verb both binds its ObjectFifo's shim endpoint and emits the shim DMA. The DMA references the fifo by symbol name (a legal MLIR forward reference), so it does not require the fifo to be resolved yet -- the Program resolves fifos and cores afterward, with every runtime endpoint already bound.

Source code in python/iron/runtime/runtime.py
def __init__(self, runtime: "Runtime"):
    self._runtime = runtime
    # The implicit group for fill/drain calls that pass no explicit group.
    self._default_task_group = TaskGroup(next(runtime._task_group_index))
    self._open_task_groups: list[TaskGroup] = []
    self._used_default = False
    self._used_explicit = False

note_fifo

note_fifo(handle: ObjectFifoHandle) -> None

Record that handle is driven from the runtime (its shim endpoint).

Source code in python/iron/runtime/runtime.py
def note_fifo(self, handle: ObjectFifoHandle) -> None:
    """Record that ``handle`` is driven from the runtime (its shim endpoint)."""
    self._runtime._fifos.add(handle)

finish_task_group

finish_task_group(tg: TaskGroup) -> None

Close a task group: await its waited tasks, then free the rest.

Waits are ordered before frees within the group, matching the hardware-safe order the old flat-list runtime used.

Source code in python/iron/runtime/runtime.py
def finish_task_group(self, tg: TaskGroup) -> None:
    """Close a task group: await its waited tasks, then free the rest.

    Waits are ordered before frees within the group, matching the
    hardware-safe order the old flat-list runtime used.
    """
    if tg in self._open_task_groups:
        self._open_task_groups.remove(tg)
    actions = tg._actions
    if not actions:
        return
    wait_tasks = [(fn, a) for (fn, a) in actions if fn == dma_await_task]
    free_tasks = [(fn, a) for (fn, a) in actions if fn == dma_free_task]
    if len(wait_tasks) + len(free_tasks) != len(actions):
        unknown = [
            (fn, a)
            for (fn, a) in actions
            if fn != dma_await_task and fn != dma_free_task
        ]
        raise IronRuntimeError(
            f"Unknown action type detected: {','.join(str(a) for a in unknown)}"
        )
    for fn, a in wait_tasks + free_tasks:
        fn(*a)
    tg._actions = []

emit_transfer

emit_transfer(
    task: DMATask, task_group: TaskGroup | None
) -> None

Emit a DMA transfer and record its await/free action(s) for group close.

A waited transfer is both awaited and freed: the dynamic BD-pool pass only returns an id to the runtime free-list on an explicit free (an await is a pure TCT sync), so a waited-but-never-freed task would leak a pool slot on every rolled-loop iteration. Matches the static path's long-standing "await implies release" convention -- it just does so with an explicit free instead of folding it into the await.

Source code in python/iron/runtime/runtime.py
def emit_transfer(self, task: DMATask, task_group: TaskGroup | None) -> None:
    """Emit a DMA transfer and record its await/free action(s) for group close.

    A waited transfer is both awaited and freed: the dynamic BD-pool pass
    only returns an id to the runtime free-list on an explicit free (an
    await is a pure TCT sync), so a waited-but-never-freed task would leak
    a pool slot on every rolled-loop iteration. Matches the static path's
    long-standing "await implies release" convention -- it just does so
    with an explicit free instead of folding it into the await.
    """
    task.resolve()
    if task_group is not None:
        self._used_explicit = True
        group = task_group
    else:
        self._used_default = True
        group = self._default_task_group
    if task.will_wait():
        group._actions.append((dma_await_task, [task.task]))
    group._actions.append((dma_free_task, [task.task]))

finalize

finalize() -> None

Close bookkeeping after the body runs.

Source code in python/iron/runtime/runtime.py
def finalize(self) -> None:
    """Close bookkeeping after the body runs."""
    explicit_open = [
        tg for tg in self._open_task_groups if tg is not self._default_task_group
    ]
    if explicit_open:
        tgs = ", ".join(str(t) for t in explicit_open)
        raise IronRuntimeError(f"Failed to close task groups: {tgs}")
    if (
        self._runtime._strict_task_groups
        and self._used_default
        and self._used_explicit
    ):
        raise IronRuntimeError(
            "Mixing explicit task groups and the default task group is "
            "prohibited. Please assign all tasks to a task group."
        )
    # Flush any transfers left in the default group (no explicit finish).
    if self._default_task_group._actions:
        self.finish_task_group(self._default_task_group)

Runtime

Runtime(
    seq_fn: Callable,
    fn_args: "Sequence | None" = None,
    *,
    strict_task_groups: bool = True
)

Bases: Resolvable

The host-side sequence of data-movement operations that execute an IRON design.

A Runtime describes what the host does at runtime: filling input ObjectFifos with data and draining results back to host buffers. The sequence is the seq_fn callback passed to the constructor; its body reads the runtime buffers as parameters and moves data with fifo.fill(...) / fifo.drain(...).

Create a runtime from its sequence body and fn_args.

Mirrors Worker(core_fn, fn_args): seq_fn runs inside @runtime_sequence at resolve time, called as seq_fn(*fn_args), and can use native range_/if_ and fifo.fill/fifo.drain. Objects in fn_args (ObjectFifoHandles, Buffers, ...) are registered eagerly at construction -- fifo shim endpoints bind now (from prod(tile=)/cons(tile=)) so the Program resolves fifos and cores before the body emits, letting body verbs read resolved worker state (e.g. barrier.set).

Each fn_args entry is one of:

  • an unbound DispatchTime parameter: replaced with its live SSA scalar. Forward each parameter once, in any order; parameter identity determines the host binding, not its position in fn_args.
  • a type (a tensor type or a scalar type like np.int32): declares a runtime input and is replaced with a live SSA value bound to a new runtime_sequence block arg -- a tensor type becomes a RuntimeData (fill/drain target), a scalar type becomes the bare SSA value (scf survives to the dynamic EmitC path).
  • a concrete int or NumPy integer value: also declares a runtime input, but is folded into a constant instead of a block arg (constant-bound range_/if_ unrolls to the static binary path). One body thus serves both lowerings depending on whether the caller passes a type or an integer here. NumPy integers retain their scalar dtype; plain Python integers use i32 for compatibility with the common np.int32 path.
  • any other object (ObjectFifoHandle, Buffer, Kernel, ScratchpadParameter, WorkerRuntimeBarrier, ...): passed through to the body unchanged, as with Worker.fn_args.

Parameters:

Name Type Description Default
seq_fn Callable

The sequence body; params bound to fn_args in order.

required
fn_args Sequence | None

Types/ints (runtime inputs) and shared objects, in the order seq_fn expects them. Defaults to None (empty list).

None
strict_task_groups bool

Disallow mixing the default and explicit task groups. Defaults to True.

True
Source code in python/iron/runtime/runtime.py
def __init__(
    self,
    seq_fn: Callable,
    fn_args: "Sequence | None" = None,
    *,
    strict_task_groups: bool = True,
) -> None:
    """Create a runtime from its sequence body and fn_args.

    Mirrors [`Worker`][iron.Worker]``(core_fn, fn_args)``: ``seq_fn`` runs
    inside ``@runtime_sequence`` at resolve time, called as
    ``seq_fn(*fn_args)``, and can use native ``range_``/``if_`` and
    ``fifo.fill``/``fifo.drain``. Objects in ``fn_args`` (ObjectFifoHandles,
    Buffers, ...) are registered eagerly at construction -- fifo shim
    endpoints bind now (from ``prod(tile=)``/``cons(tile=)``) so the Program
    resolves fifos and cores before the body emits, letting body verbs read
    resolved worker state (e.g. ``barrier.set``).

    Each ``fn_args`` entry is one of:

    * an unbound **DispatchTime parameter**: replaced with its live SSA
      scalar. Forward each parameter once, in any order; parameter identity
      determines the host binding, not its position in ``fn_args``.
    * a **type** (a tensor type or a scalar type like ``np.int32``): declares
      a runtime input and is replaced with a live SSA value bound to a new
      ``runtime_sequence`` block arg -- a tensor type becomes a
      ``RuntimeData`` (``fill``/``drain`` target), a scalar type becomes the
      bare SSA value (``scf`` survives to the dynamic EmitC path).
    * a concrete **int or NumPy integer value**: also declares a runtime input,
      but is folded into a constant instead of a block arg (constant-bound
      ``range_``/``if_`` unrolls to the static binary path). One body thus
      serves both lowerings depending on whether the caller passes a type or
      an integer here. NumPy integers retain their scalar dtype; plain Python
      integers use i32 for compatibility with the common ``np.int32`` path.
    * any other object (ObjectFifoHandle, Buffer, Kernel, ScratchpadParameter,
      WorkerRuntimeBarrier, ...): passed through to the body unchanged, as
      with ``Worker.fn_args``.

    Args:
        seq_fn (Callable): The sequence body; params bound to ``fn_args`` in order.
        fn_args (Sequence | None): Types/ints (runtime inputs) and shared objects,
            in the order ``seq_fn`` expects them. Defaults to None (empty list).
        strict_task_groups (bool): Disallow mixing the default and explicit task groups. Defaults to True.
    """
    self._seq_fn: Callable = seq_fn
    self._fn_args = list(fn_args) if fn_args is not None else []
    self._dispatch_binding = object()
    dispatch_args = [
        arg for arg in self._fn_args if isinstance(arg, _DispatchParameter)
    ]
    if len({id(arg) for arg in dispatch_args}) != len(dispatch_args):
        raise TypeError(
            "Forward each DispatchTime parameter exactly once to Runtime."
        )
    if len({id(arg.owner) for arg in dispatch_args}) > 1:
        raise TypeError(
            "Runtime cannot mix DispatchTime parameters from different designs."
        )
    for container in self._fn_args:
        if isinstance(container, (list, tuple)):
            for arg in flatten_fn_args(container):
                if isinstance(arg, _DispatchParameter):
                    raise TypeError(
                        f"DispatchTime parameter {arg.name!r} must be a direct Runtime fn_args entry."
                    )
    # A concrete int entry is a folded constant; a type/generic-alias entry
    # is a runtime input; anything else passes through as an object fn_arg.
    self._const_inputs: list[int | np.integer | None] = [
        v if isinstance(v, (int, np.integer)) and not isinstance(v, bool) else None
        for v in self._fn_args
    ]
    self._rt_data: list["RuntimeData | None"] = [
        (
            RuntimeData(
                arg.scalar_type if isinstance(arg, _DispatchParameter) else arg
            )
            if c is None
            and (
                isinstance(arg, (type, _DispatchParameter))
                or get_origin(arg) is np.ndarray
            )
            else None
        )
        for c, arg in zip(self._const_inputs, self._fn_args)
    ]
    # Keep the host scalar ABI in signature order, while the callback and
    # tensor arguments retain fn_args order. Only dispatch slots move.
    dispatch_data = iter(
        data
        for _, data in sorted(
            (
                (arg.position, data)
                for arg, data in zip(self._fn_args, self._rt_data)
                if isinstance(arg, _DispatchParameter)
            ),
            key=lambda item: item[0],
        )
    )
    self._block_data = []
    for arg, data in zip(self._fn_args, self._rt_data):
        if isinstance(arg, _DispatchParameter):
            self._block_data.append(next(dispatch_data))
        elif data is not None:
            if dispatch_args and data.is_scalar:
                raise TypeError(
                    "Runtime cannot mix bare scalar types with DispatchTime parameters."
                )
            self._block_data.append(data)
    self._fifos: set[ObjectFifoHandle] = set()
    self._register_fn_args()
    # Lower-level explicit-routing primitives (peers of ObjectFifo for
    # designs that hand-wire flows + DMA programs instead of letting
    # ObjectFifo manage them).
    self._flows = []
    self._locks = []
    self._tile_dmas = []
    self._resolved_tile_dmas = None
    self._scratchpad_parameters: list[ScratchpadParameter] = []
    self._strict_task_groups = strict_task_groups
    self._task_group_index = itertools.count()

fifos property

fifos: list[ObjectFifoHandle]

The ObjectFifoHandles driven from the runtime by fill()/drain().

add_flow

add_flow(flow) -> None

Register an explicit flow so the Program resolves it alongside the ObjectFifos.

Accepts a Flow or PacketFlow.

Source code in python/iron/runtime/runtime.py
def add_flow(self, flow) -> None:
    """Register an explicit flow so the Program resolves it alongside the ObjectFifos.

    Accepts a [`Flow`][iron.Flow] or [`PacketFlow`][iron.PacketFlow].
    """
    self._flows.append(flow)

add_lock

add_lock(lock) -> None

Register an explicit Lock shared between a Worker and a TileDma.

See TileDma.

Source code in python/iron/runtime/runtime.py
def add_lock(self, lock) -> None:
    """Register an explicit [`Lock`][iron.Lock] shared between a Worker and a TileDma.

    See [`TileDma`][iron.TileDma].
    """
    self._locks.append(lock)

add_tile_dma

add_tile_dma(tile_dma) -> None

Register a TileDma; channels sharing a Tile are combined at resolution.

Source code in python/iron/runtime/runtime.py
def add_tile_dma(self, tile_dma) -> None:
    """Register a TileDma; channels sharing a Tile are combined at resolution."""
    if self._resolved_tile_dmas is not None:
        raise IronRuntimeError("Cannot register TileDma after DMA resolution.")
    self._tile_dmas.append(tile_dma)

resolve_tile_dmas

resolve_tile_dmas() -> None

Validate and emit one DMA region per Tile without changing registrations.

Source code in python/iron/runtime/runtime.py
def resolve_tile_dmas(self) -> None:
    """Validate and emit one DMA region per Tile without changing registrations."""
    from ..dataflow.tile_dma import TileDma

    if self._resolved_tile_dmas is None:
        programs = {}
        coordinates = {}
        for tile_dma in self._tile_dmas:
            tile = tile_dma.tile
            if tile.col is not None and tile.row is not None:
                key = (tile.col, tile.row)
                if key in coordinates and coordinates[key] is not tile:
                    raise IronRuntimeError(
                        f"Two TileDma programs name {tile}, via different "
                        "Tile objects. Share one Tile object for their channels."
                    )
                coordinates[key] = tile
            if tile in programs:
                programs[tile] = TileDma(
                    tile, [*programs[tile].channels, *tile_dma.channels]
                )
            else:
                programs[tile] = tile_dma
        self._resolved_tile_dmas = list(programs.values())
    # Placement coalesces tile ops, not their DMA regions or channel chains.
    for program in self._resolved_tile_dmas:
        program.resolve()

resolve

resolve(
    loc: Location | None = None,
    ip: InsertionPoint | None = None,
    *,
    trace_size: int | None = None,
    reuse_output_buffer: bool = False,
    egress_shim_col: int = 0,
    load_pdi_device_ref: str | None = None
) -> None

Build the runtime_sequence op and run the sequence body inside it.

The body runs exactly once. Each fill/drain verb binds its ObjectFifo's shim endpoint and emits the shim DMA (referencing the fifo by symbol name, a forward reference). The Program calls this before it resolves ObjectFifos and cores, so by the time a fifo is resolved every runtime endpoint -- including those on link siblings -- is already bound.

Parameters:

Name Type Description Default
loc Location | None

Optional MLIR source location for the generated ops.

None
ip InsertionPoint | None

Optional MLIR insertion point for the generated ops.

None
trace_size int | None

Forwarded from Program.enable_trace; see there. None/0 disables tracing.

None
reuse_output_buffer bool

Forwarded from Program.enable_trace; see there.

False
egress_shim_col int

Forwarded from Program.enable_trace; see there.

0
load_pdi_device_ref str | None

On the full-ELF path (no xclbin configures the device), the device symbol to load via npu_load_pdi as the first op in the sequence. None on the xclbin path.

None
Source code in python/iron/runtime/runtime.py
def resolve(
    self,
    loc: ir.Location | None = None,
    ip: ir.InsertionPoint | None = None,
    *,
    trace_size: int | None = None,
    reuse_output_buffer: bool = False,
    egress_shim_col: int = 0,
    load_pdi_device_ref: str | None = None,
) -> None:
    """Build the ``runtime_sequence`` op and run the sequence body inside it.

    The body runs exactly once. Each ``fill``/``drain`` verb binds its
    ObjectFifo's shim endpoint and emits the shim DMA (referencing the fifo
    by symbol name, a forward reference). The Program calls this before it
    resolves ObjectFifos and cores, so by the time a fifo is resolved every
    runtime endpoint -- including those on link siblings -- is already bound.

    Args:
        loc: Optional MLIR source location for the generated ops.
        ip: Optional MLIR insertion point for the generated ops.
        trace_size: Forwarded from
            [`Program.enable_trace`][iron.program.Program.enable_trace]; see
            there. ``None``/``0`` disables tracing.
        reuse_output_buffer: Forwarded from
            [`Program.enable_trace`][iron.program.Program.enable_trace]; see
            there.
        egress_shim_col: Forwarded from
            [`Program.enable_trace`][iron.program.Program.enable_trace]; see
            there.
        load_pdi_device_ref: On the full-ELF path (no xclbin configures the
            device), the device symbol to load via ``npu_load_pdi`` as the
            first op in the sequence. ``None`` on the xclbin path.
    """
    # A runtime_sequence block arg per runtime (type) input; folded-constant
    # inputs contribute no block arg.
    rt_dtypes = [
        try_convert_np_type_to_mlir_type(rt_data.arr_type)
        for rt_data in self._block_data
        if rt_data is not None
    ]
    active = ActiveSequence(self)

    seq_op = RuntimeSequenceOp(sym_name="sequence")
    entry_block = seq_op.body.blocks.append(*rt_dtypes)
    with ir.InsertionPoint(entry_block):
        # Full-ELF designs configure the device themselves: no xclbin
        # pre-loads the PDI, so the sequence must start by loading it.
        if load_pdi_device_ref is not None:
            npu_load_pdi(device_ref=load_pdi_device_ref)

        block_args = iter(entry_block.arguments)
        for rt_data in self._block_data:
            if rt_data is not None:
                rt_data.op = next(block_args)
        for arg in self._fn_args:
            if isinstance(arg, _DispatchParameter):
                arg._bind(self._dispatch_binding)

        if trace_size is not None and trace_size > 0:
            trace_utils.start_trace(
                trace_size=trace_size,
                reuse_output_buffer=reuse_output_buffer,
                routing="single",
                egress_shim_col=egress_shim_col,
            )

        # Build the body's positional args, one per fn_args entry, in order:
        #   * folded-constant entry -> an arith.constant of that value;
        #   * scalar type entry     -> its live SSA value (used in arithmetic
        #                              and range_/if_ bounds);
        #   * tensor type entry     -> its RuntimeData handle (fill/drain);
        #   * anything else         -> passed through unchanged (fifos, etc).
        body_args = []
        for arg, const_val, rt_data in zip(
            self._fn_args, self._const_inputs, self._rt_data
        ):
            if const_val is not None:
                dtype = (
                    type(const_val)
                    if isinstance(const_val, np.integer)
                    else np.int32
                )
                scalar_type = np_dtype_to_mlir_type(dtype)
                value = int(const_val)
                # IntegerAttr's C API accepts int64_t, including the bit
                # pattern of a uint64/index value above INT64_MAX.
                if value > np.iinfo(np.int64).max:
                    value -= 1 << 64
                if (
                    isinstance(scalar_type, ir.IntegerType)
                    and ir.IntegerType(scalar_type).is_unsigned
                ):
                    # arith.constant requires signless integers. EmitC's
                    # ConstantLike op preserves unsigned types and folds.
                    from ...dialects.emitc import (  # pyright: ignore[reportMissingImports]
                        ConstantOp,
                    )

                    body_args.append(
                        ConstantOp(
                            scalar_type, ir.IntegerAttr.get(scalar_type, value)
                        ).result
                    )
                else:
                    body_args.append(constant(value, scalar_type))
            elif rt_data is not None:
                body_args.append(rt_data.op if rt_data.is_scalar else rt_data)
            else:
                body_args.append(arg)

        with active_sequence_scope(active):
            self._seq_fn(*body_args)
            active.finalize()

    self._dedup_runtime_consumers()

sync_parameters

sync_parameters() -> None

Emit aiex.sync_scratchpad_parameters_from_host in the sequence body.

Call after all scratchpad parameters have been written on the host side and before starting workers that read them.

Source code in python/iron/runtime/runtime.py
def sync_parameters() -> None:
    """Emit ``aiex.sync_scratchpad_parameters_from_host`` in the sequence body.

    Call after all scratchpad parameters have been written on the host side and
    before starting workers that read them.
    """
    active_sequence()  # ensure we're inside a sequence body
    _SyncParametersTask().resolve()

Buffer

Use a Buffer for local scratch storage shared by sequential kernel calls in one Worker, as in the edge-detection example. Passing it in the Worker's fn_args associates it with that Worker's tile. Unlike an ObjectFifo, a Buffer does not provide producer/consumer synchronization.

Named memory region accessible by both Workers and the Runtime.

Buffer

Buffer(
    type: type[ndarray] | None = None,
    initial_value: ndarray | None = None,
    name: str | None = None,
    tile: Tile | None = None,
    use_write_rtp: bool = False,
    address: int | None = None,
    mem_bank: int | None = None,
)

Bases: Resolvable

A buffer that is available both to Workers and to the Runtime for operations.

This is often used for Runtime Parameters.

Declare a memory region at the top-level of the design.

The buffer is accessible by both Workers and the Runtime.

Parameters:

Name Type Description Default
type type[ndarray] | None

The type of the buffer. Defaults to None.

None
initial_value ndarray | None

An initial value to set the buffer to. Should be of same datatype and shape as the buffer. Defaults to None.

None
name str | None

The name of the buffer. If none is given, a unique name will be generated. Defaults to None.

None
tile Tile | None

The tile for the buffer. Automatically set to the Worker's tile when the buffer is passed in the Worker's fn_args list. Defaults to None.

None
use_write_rtp bool

If use_write_rtp, write_rtp/read_rtp operations will be generated. Otherwise, traditional write/read operations will be used. Defaults to False.

False
address int | None

Pin the buffer to a fixed L1 address. Needed for host-written RTP buffers the runtime pokes at a hardcoded address. Defaults to None (compiler-assigned).

None
mem_bank int | None

Pin the buffer to a specific L1 memory bank. The pin is a hard constraint: if the bank cannot hold the buffer, the compiler reports an error rather than placing it elsewhere. Defaults to None (compiler-assigned).

None

Raises:

Type Description
ValueError

If neither type nor initial_value is provided, or if address/mem_bank are provided but are not a non-negative int.

Source code in python/iron/buffer.py
def __init__(
    self,
    type: type[np.ndarray] | None = None,
    initial_value: np.ndarray | None = None,
    name: str | None = None,
    tile: Tile | None = None,
    use_write_rtp: bool = False,
    address: int | None = None,
    mem_bank: int | None = None,
):
    """Declare a memory region at the top-level of the design.

    The buffer is accessible by both Workers and the Runtime.

    Args:
        type (type[np.ndarray] | None, optional): The type of the buffer. Defaults to None.
        initial_value (np.ndarray | None, optional): An initial value to set the buffer
            to. Should be of same datatype and shape as the buffer. Defaults to None.
        name (str | None, optional): The name of the buffer. If none is given, a unique
            name will be generated. Defaults to None.
        tile (Tile | None, optional): The tile for the buffer. Automatically set to the
            Worker's tile when the buffer is passed in the Worker's fn_args list.
            Defaults to None.
        use_write_rtp (bool, optional): If use_write_rtp, write_rtp/read_rtp operations
            will be generated. Otherwise, traditional write/read operations will be
            used. Defaults to False.
        address (int | None, optional): Pin the buffer to a fixed L1 address. Needed
            for host-written RTP buffers the runtime pokes at a hardcoded address.
            Defaults to None (compiler-assigned).
        mem_bank (int | None, optional): Pin the buffer to a specific L1 memory bank.
            The pin is a hard constraint: if the bank cannot hold the buffer, the
            compiler reports an error rather than placing it elsewhere. Defaults to
            None (compiler-assigned).

    Raises:
        ValueError: If neither ``type`` nor ``initial_value`` is provided, or if
            ``address``/``mem_bank`` are provided but are not a non-negative int.
    """
    if type is None and initial_value is None:
        raise ValueError("Must provide either type, initial value, or both.")
    if address is not None:
        if not isinstance(address, int) or isinstance(address, bool):
            raise ValueError(
                f"Buffer address must be an int, but got "
                f"{address.__class__.__name__}"
            )
        if address < 0:
            raise ValueError(f"Buffer address must be >= 0, but got {address}")
    if mem_bank is not None:
        if not isinstance(mem_bank, int) or isinstance(mem_bank, bool):
            raise ValueError(
                f"Buffer mem_bank must be an int, but got "
                f"{mem_bank.__class__.__name__}"
            )
        if mem_bank < 0:
            raise ValueError(f"Buffer mem_bank must be >= 0, but got {mem_bank}")
    if type is None:
        assert initial_value is not None
        type = np.ndarray[initial_value.shape, np.dtype[initial_value.dtype.type]]
    self._initial_value = initial_value
    self._name = name
    self._op = None
    self._arr_type = type
    if not self._name:
        self._name = f"buf_{next(Buffer._gbuf_index)}"
    self._use_write_rtp = use_write_rtp
    self._address = address
    self._mem_bank = mem_bank
    self._tile = tile
    # Whether the user pinned this Buffer to an explicit tile at
    # construction.  A Worker may auto-pin _tile later as a
    # convenience, so `_tile is not None` is not a reliable signal of
    # user intent; this flag is.  Only explicitly-placed Buffers may be
    # shared (read) across Workers — see Worker.
    self._explicit_tile = tile is not None
    self._owner_worker: "Worker | None" = None

tile property

tile: Tile | None

The tile this buffer is on.

shape property

shape: Sequence[int]

The shape of the buffer.

dtype property

dtype: type[NpuDType]

The per-element datatype of the buffer.

tiles

tiles() -> list

Tile dependency for Program.resolve tile discovery.

Pinned Buffers (e.g. a compute Worker reading a neighbor tile's L1 directly) need their tile registered with the Device before resolve runs. Worker-attached Buffers without an explicit placement get pinned to the Worker's tile in Worker's constructor, which is already discoverable via Worker.tile; this method just exposes any extra (cross-tile) placements.

Source code in python/iron/buffer.py
def tiles(self) -> list:
    """Tile dependency for Program.resolve tile discovery.

    Pinned Buffers (e.g. a compute [`Worker`][iron.Worker] reading a
    neighbor tile's L1 directly) need their tile registered with the
    Device before `resolve` runs. Worker-attached Buffers without an
    explicit placement get pinned to the Worker's tile in
    [`Worker`][iron.Worker]'s constructor, which is already discoverable
    via `Worker.tile`; this method just exposes any extra (cross-tile)
    placements.
    """
    return [self._tile] if self._tile is not None else []

Kernels

Kernel binds a function symbol; KernelObject owns the shared link artifact. Both are exported from aie.iron. Pass a KernelObject("shared.o") to several Kernel constructors to bind symbols from one precompiled object; ObjectFile("shared.o", symbol_prefix=...) is the same for a prebuilt object whose symbols were renamed under a prefix. For C++ source, ExternalFunction creates the owner, exposed as fn.object_file; fn.object_file.bind(symbol, arg_types) binds another entry point to that owner, applying its symbol prefix. External functions with the same explicit output filename and identical source recipes also share ownership. Conflicting recipes for one output filename are rejected. Resolving a source-backed binding registers its artifact for compilation without requiring the original ExternalFunction to remain alive. Independent operations such as kernels.zero(...) own their own objects.

Kernel and ExternalFunction: wrappers for pre-compiled and C++ AIE compute kernels.

KernelObject dataclass

KernelObject(
    name: str,
    link_with_mode: str | None = None,
    _source: _KernelSource | None = None,
    _compiled_dirs: set[str] = set(),
    *,
    _symbol_prefix: str | None = None
)

One link artifact, shared by every kernel that binds one of its symbols.

A prebuilt object only needs a filename and link policy. ExternalFunction supplies the immutable source recipe; compilation state belongs here, not to any particular exported function.

object_file_name property

object_file_name: str

Filename of the linked artifact.

symbol_prefix property

symbol_prefix: str | None

Symbol namespace shared by all bindings of this artifact.

resolve_symbol

resolve_symbol(name: str) -> str

Return name qualified into this object's symbol namespace.

Source code in python/iron/kernel.py
def resolve_symbol(self, name: str) -> str:
    """Return ``name`` qualified into this object's symbol namespace."""
    if not name:
        raise ValueError("Kernel name cannot be empty.")
    return f"{self.symbol_prefix}_{name}" if self.symbol_prefix else name

bind

bind(
    name: str,
    arg_types: list[ArgType] | None = None,
    *,
    link_with_mode: str | None = None,
    stack_size_override: int | None = None
) -> Kernel

Bind a source-level symbol while retaining this artifact's ownership.

Source code in python/iron/kernel.py
def bind(
    self,
    name: str,
    arg_types: list[ArgType] | None = None,
    *,
    link_with_mode: str | None = None,
    stack_size_override: int | None = None,
) -> "Kernel":
    """Bind a source-level symbol while retaining this artifact's ownership."""
    return Kernel(
        self.resolve_symbol(name),
        self,
        arg_types,
        link_with_mode=link_with_mode,
        stack_size_override=stack_size_override,
    )

ObjectFile

ObjectFile(
    object_file_name: str,
    *,
    symbol_prefix: str | None = None,
    link_with_mode: str | None = None
)

Bases: KernelObject

A prebuilt KernelObject with an optional symbol namespace.

Source code in python/iron/kernel.py
def __init__(
    self,
    object_file_name: str,
    *,
    symbol_prefix: str | None = None,
    link_with_mode: str | None = None,
) -> None:
    super().__init__(object_file_name, link_with_mode, _symbol_prefix=symbol_prefix)

Kernel

Kernel(
    name: str,
    object_file_name: str | KernelObject,
    arg_types: list[ArgType] | None = None,
    *,
    link_with_mode: str | None = None,
    stack_size_override: int | None = None
)

Bases: Resolvable

An AIE core function backed by a pre-compiled object file.

Use ExternalFunction instead when you want to compile from C/C++ source at JIT time.

resolve() emits a func.func private declaration with a link_with attribute naming object_file_name. The aie-assign-core-link-files pass propagates this into the CoreOp's link_files attribute so the linker knows which file to include.

link_with_mode selects how that artifact is consumed: the default (None) object-links it, while "merge" asks aiecc to llvm-link it into the core's LLVM module before codegen. The mode is explicit metadata -- it is never inferred from the file suffix.

Construct a Kernel backed by a pre-compiled object file.

Parameters:

Name Type Description Default
name str

Symbol name of the function as it appears in the object file.

required
object_file_name str | KernelObject

Filename (e.g. "add_one.o") or shared KernelObject of the pre-compiled object file. Must be on the linker search path at compile time.

required
arg_types list[ArgType] | None

Type signature of the function arguments. Defaults to None (empty list).

None
link_with_mode str | None

Optional link policy emitted alongside link_with. "merge" routes the artifact through aiecc's llvm-link merge path; None (the default) object-links it.

None
stack_size_override int | None

Declared upper bound, in bytes, on the stack that this kernel's call subtree uses. Set it for recursion, for an indirect call, or for a link_with_mode="merge" kernel, which aiecc's stack analysis reads as part of the core rather than as a separate object. See Kernel.stack_size_override.

None
Source code in python/iron/kernel.py
def __init__(
    self,
    name: str,
    object_file_name: str | KernelObject,
    arg_types: list[ArgType] | None = None,
    *,
    link_with_mode: str | None = None,
    stack_size_override: int | None = None,
) -> None:
    """Construct a Kernel backed by a pre-compiled object file.

    Args:
        name: Symbol name of the function as it appears in the object file.
        object_file_name: Filename (e.g. ``"add_one.o"``) or shared
            ``KernelObject`` of the pre-compiled object file. Must be on
            the linker search path at compile time.
        arg_types: Type signature of the function arguments.  Defaults to None (empty list).
        link_with_mode: Optional link policy emitted alongside
            ``link_with``.  ``"merge"`` routes the artifact through aiecc's
            ``llvm-link`` merge path; None (the default) object-links it.
        stack_size_override: Declared upper bound, in bytes, on the stack
            that this kernel's call subtree uses. Set it for recursion, for
            an indirect call, or for a ``link_with_mode="merge"`` kernel,
            which aiecc's stack analysis reads as part of the core rather
            than as a separate object. See
            [`Kernel.stack_size_override`][iron.kernel.Kernel.stack_size_override].
    """
    self._init_identity(name, arg_types)
    if isinstance(object_file_name, KernelObject):
        if (
            link_with_mode is not None
            and link_with_mode != object_file_name.link_with_mode
        ):
            raise ValueError(
                "link_with_mode conflicts with the shared KernelObject"
            )
        self._object_file = object_file_name
    else:
        self._object_file = KernelObject(object_file_name, link_with_mode)
    self._stack_size_override = stack_size_override

object_file property

object_file: KernelObject

The artifact owner, also shared by sibling symbol bindings.

object_file_name property

object_file_name: str

Filename of the compiled object file.

link_with_mode: str | None

Link policy emitted with link_with, or None for object linking.

stack_size_override property

stack_size_override: int | None

Declared upper bound on the stack that this kernel's call subtree uses.

With None, aiecc's analysis computes the bound. An explicit value replaces that computed bound, even when it is smaller: it is a declaration, and 0 is legal. See Core Data Memory.

name property

name: str

Symbol name of the function as it appears in the object file.

resolve

resolve(
    loc: Location | None = None,
    ip: InsertionPoint | None = None,
) -> None

Declare this kernel in the enclosing symbol table, once.

The declaration lives in the IR, not on the kernel, so one kernel can be resolved into any number of designs and contexts. A kernel equal to one already declared reuses that declaration; one that shares only the symbol name is rejected.

Source code in python/iron/kernel.py
def resolve(
    self,
    loc: ir.Location | None = None,
    ip: ir.InsertionPoint | None = None,
) -> None:
    """Declare this kernel in the enclosing symbol table, once.

    The declaration lives in the IR, not on the kernel, so one kernel can
    be resolved into any number of designs and contexts. A kernel equal to
    one already declared reuses that declaration; one that shares only the
    symbol name is rejected.
    """
    if self.object_file._source is not None:
        # JIT clears its discovery registry before generating a design.
        # Re-register the artifact even when only a sibling binding survives.
        ExternalFunction._register_object(self)
    point = ip if ip is not None else ir.InsertionPoint.current
    table = _enclosing_symbol_table(point)
    if self._name in table:
        self._check_declaration(table[self._name])
        return
    with point:
        external_func(
            self._name,
            inputs=self._arg_types,
            link_with=self._object_file_name,
            link_with_mode=self._link_with_mode,
            stack_size_override=self._stack_size_override,
        )

arg_shape

arg_shape(arg_index: int = 0) -> tuple[int, ...]

Return the shape tuple of the array argument at arg_index.

Works for both np.ndarray[(...,), np.dtype[T]] parameterized types (the canonical IRON kernel signature) and MLIR MemRefType operands.

Parameters:

Name Type Description Default
arg_index int

Index into arg_types. Defaults to 0.

0

Raises:

Type Description
ValueError

When arg_index is out of range or the argument at that index is not an array type.

Source code in python/iron/kernel.py
def arg_shape(self, arg_index: int = 0) -> tuple[int, ...]:
    """Return the shape tuple of the array argument at `arg_index`.

    Works for both `np.ndarray[(...,), np.dtype[T]]` parameterized
    types (the canonical IRON kernel signature) and MLIR MemRefType
    operands.

    Args:
        arg_index: Index into `arg_types`. Defaults to 0.

    Raises:
        ValueError: When `arg_index` is out of range or the
            argument at that index is not an array type.
    """
    arg = self._resolve_arg(arg_index)
    type_args = getattr(arg, "__args__", None)
    if type_args is not None and len(type_args) > 0:
        shape_arg = type_args[0]
        if isinstance(shape_arg, tuple):
            return shape_arg
    shape = getattr(arg, "shape", None)
    if shape is not None:
        return tuple(shape)
    raise ValueError(
        f"Argument {arg_index} does not have a shape or is not an array type."
    )

arg_dtype

arg_dtype(arg_index: int = 0)

Return the numpy dtype of the array argument at arg_index.

Parameters:

Name Type Description Default
arg_index int

Index into arg_types. Defaults to 0.

0

Raises:

Type Description
ValueError

When arg_index is out of range or the argument at that index is not an array type.

Source code in python/iron/kernel.py
def arg_dtype(self, arg_index: int = 0):
    """Return the numpy dtype of the array argument at `arg_index`.

    Args:
        arg_index: Index into `arg_types`. Defaults to 0.

    Raises:
        ValueError: When `arg_index` is out of range or the
            argument at that index is not an array type.
    """
    arg = self._resolve_arg(arg_index)
    type_args = getattr(arg, "__args__", None)
    if type_args is not None and len(type_args) >= 2:
        dt = type_args[1]
        dt_args = getattr(dt, "__args__", None)
        return _as_dtype(dt_args[0] if dt_args is not None else dt)
    dtype = getattr(arg, "dtype", None)
    if dtype is not None:
        return _as_dtype(dtype)
    raise ValueError(
        f"Argument {arg_index} does not have a dtype or is not an array type."
    )

tile_size

tile_size(arg_index: int = 0) -> int

Return the first dimension of the array argument at arg_index.

Convenience wrapper over arg_shape for the common case of a 1-D buffer argument. tile_size(i) is equivalent to arg_shape(i)[0].

Parameters:

Name Type Description Default
arg_index int

Index into arg_types. Defaults to 0.

0
Source code in python/iron/kernel.py
def tile_size(self, arg_index: int = 0) -> int:
    """Return the first dimension of the array argument at `arg_index`.

    Convenience wrapper over
    [`arg_shape`][iron.kernel.Kernel.arg_shape] for the common case of
    a 1-D buffer argument. `tile_size(i)` is equivalent to
    `arg_shape(i)[0]`.

    Args:
        arg_index: Index into `arg_types`. Defaults to 0.
    """
    shape = self.arg_shape(arg_index)
    if len(shape) == 0:
        raise ValueError(
            f"Argument {arg_index} does not have a shape or is not an array type."
        )
    return shape[0]

arg_types

arg_types() -> list

Return the argument types as declared: np.ndarray[shape, dtype] / scalars.

A copy, and stable: resolving the kernel builds MLIR types from these without replacing them, so a memoized kernel describes itself the same way before and after a build.

Source code in python/iron/kernel.py
def arg_types(self) -> list:
    """Return the argument types as declared: ``np.ndarray[shape, dtype]`` / scalars.

    A copy, and stable: resolving the kernel builds MLIR types from these
    without replacing them, so a memoized kernel describes itself the same
    way before and after a build.
    """
    return self._arg_types.copy()

ExternalFunction

ExternalFunction(
    name: str,
    object_file_name: str | None = None,
    source_file: str | None = None,
    source_string: str | None = None,
    arg_types: list[ArgType] | None = None,
    include_dirs: list[str] | None = None,
    compile_flags: list[str] | None = None,
    *,
    symbol_prefix: str | None = None,
    use_chess: bool = False,
    inline: bool = False,
    stack_size_override: int | None = None,
    contract: Any = None
)

Bases: Kernel

An AIE core function compiled from C/C++ source at JIT time.

Each instance is registered in _instances at construction time so that the @jit decorator can discover and compile all source files before invoking the MLIR compilation pipeline. _instances is cleared at the start of each @jit call to prevent stale registrations from a previous (possibly failed) run.

Use the base Kernel class instead when you have a pre-built object file.

Construct an ExternalFunction compiled from C/C++ source at JIT time.

Parameters:

Name Type Description Default
name str

Symbol name of the function as it will appear in the object file.

required
object_file_name str | None

Output artifact name. Defaults to <effective_name>.o, or <effective_name>.ll with inline=True. With inline=True an explicit name must end in .ll (textual LLVM IR) or .bc (bitcode) -- that suffix selects the emitted format -- and is otherwise rejected.

None
source_file str | None

Path to a C/C++ source file on disk. Mutually exclusive with source_string.

None
source_string str | None

Inline C/C++ source code. Mutually exclusive with source_file.

None
arg_types list[ArgType] | None

Type signature of the function arguments. Defaults to None (empty list).

None
include_dirs list[str] | None

Additional -I directories passed to the chosen compiler (Peano by default; xchesscc when use_chess=True). Relative paths resolve against the current directory at construction. Defaults to None (empty list).

None
compile_flags list[str] | None

Additional flags passed verbatim to the chosen compiler. Defaults to None (empty list).

None
symbol_prefix str | None

Optional prefix for the exported symbol name. When set, the effective symbol name becomes <symbol_prefix>_<name> and the object file is named accordingly. The original name is preserved in _original_name for source file naming.

None
use_chess bool

When True, this ExternalFunction's source is compiled with xchesscc_wrapper instead of Peano's clang++. The JIT compile orchestration auto-detects the design-level toolchain from the registered EFs and switches aiecc's front-end accordingly; mixing chess + peano EFs in one design is rejected loudly because aiecc only invokes one front-end per compile.

False
inline bool

When True, compile the kernel to alwaysinline LLVM IR (.ll) and declare it with link_with_mode = "merge" so aiecc llvm-links it into the core and inlines it, instead of object-linking a separate .o. Removes the func.call boundary and the separate object. Peano path only (the Chess/xchesscc toolchain cannot llvm-link).

False
stack_size_override int | None

Declared upper bound, in bytes, on the stack that this kernel's call subtree uses. See Kernel.stack_size_override. With inline=True, the merged kernel has no separate object, so this bound is the one input aiecc's stack analysis reads for this kernel.

None
contract Any

What the kernel computes, as an aie.iron.kernels.KernelContract: argument roles, a host reference, a tolerance and the operand layouts. The library factories always give one; a hand-built kernel may leave it None and then cannot be built or judged generically.

None
Source code in python/iron/kernel.py
def __init__(
    self,
    name: str,
    object_file_name: str | None = None,
    source_file: str | None = None,
    source_string: str | None = None,
    arg_types: list[ArgType] | None = None,
    include_dirs: list[str] | None = None,
    compile_flags: list[str] | None = None,
    *,
    symbol_prefix: str | None = None,
    use_chess: bool = False,
    inline: bool = False,
    stack_size_override: int | None = None,
    contract: Any = None,
) -> None:
    """Construct an ExternalFunction compiled from C/C++ source at JIT time.

    Args:
        name: Symbol name of the function as it will appear in the object
            file.
        object_file_name: Output artifact name. Defaults to
            ``<effective_name>.o``, or ``<effective_name>.ll`` with
            ``inline=True``. With ``inline=True`` an explicit name must end
            in ``.ll`` (textual LLVM IR) or ``.bc`` (bitcode) -- that suffix
            selects the emitted format -- and is otherwise rejected.
        source_file: Path to a C/C++ source file on disk.  Mutually
            exclusive with ``source_string``.
        source_string: Inline C/C++ source code.  Mutually exclusive with
            ``source_file``.
        arg_types: Type signature of the function arguments.  Defaults to
            None (empty list).
        include_dirs: Additional ``-I`` directories passed to the chosen
            compiler (Peano by default; xchesscc when ``use_chess=True``).
            Relative paths resolve against the current directory at construction.
            Defaults to None (empty list).
        compile_flags: Additional flags passed verbatim to the chosen
            compiler.  Defaults to None (empty list).
        symbol_prefix: Optional prefix for the exported symbol name.  When
            set, the effective symbol name becomes ``<symbol_prefix>_<name>``
            and the object file is named accordingly.  The original name is
            preserved in ``_original_name`` for source file naming.
        use_chess: When ``True``, this ExternalFunction's source is
            compiled with ``xchesscc_wrapper`` instead of Peano's
            ``clang++``.  The JIT compile orchestration auto-detects the
            design-level toolchain from the registered EFs and switches
            aiecc's front-end accordingly; mixing chess + peano EFs in
            one design is rejected loudly because aiecc only invokes one
            front-end per compile.
        inline: When True, compile the kernel to ``alwaysinline`` LLVM IR
            (``.ll``) and declare it with ``link_with_mode = "merge"`` so
            aiecc llvm-links it into the core and inlines it, instead of
            object-linking a separate ``.o``. Removes the ``func.call``
            boundary and the separate object. Peano path only (the
            Chess/xchesscc toolchain cannot llvm-link).
        stack_size_override: Declared upper bound, in bytes, on the stack
            that this kernel's call subtree uses. See
            [`Kernel.stack_size_override`][iron.kernel.Kernel.stack_size_override].
            With ``inline=True``, the merged kernel has no separate object,
            so this bound is the one input aiecc's stack analysis reads for
            this kernel.
        contract: What the kernel computes, as an
            ``aie.iron.kernels.KernelContract``: argument roles, a host
            reference, a tolerance and the operand layouts. The library
            factories always give one; a hand-built kernel may leave it
            ``None`` and then cannot be built or judged generically.
    """
    if inline and use_chess:
        raise ValueError(
            f"ExternalFunction '{name}': inline=True requires the Peano "
            "toolchain and cannot be combined with use_chess=True."
        )
    if inline and symbol_prefix:
        raise NotImplementedError(
            f"ExternalFunction '{name}': inline=True combined with symbol_prefix is "
            "not supported (an inline kernel is emitted as LLVM IR and cannot be "
            "symbol-renamed). Use inline without a symbol_prefix, or drop inline for "
            "this kernel."
        )

    self._original_name = name
    self.contract = contract
    effective_name = f"{symbol_prefix}_{name}" if symbol_prefix else name
    object_file_name_explicit = object_file_name is not None
    if not object_file_name:
        object_file_name = (
            f"{effective_name}.ll" if inline else f"{effective_name}.o"
        )
    elif inline and Path(object_file_name).suffix.lower() not in (".ll", ".bc"):
        # An inline kernel is emitted as LLVM IR, and the suffix picks the
        # format (textual vs bitcode), so a wrong one has no valid reading.
        # Reject it instead of silently renaming the caller's artifact --
        # aiecc routes on the `link_with_mode` attribute, not the suffix, so
        # a rename would buy nothing.  Compared case-insensitively, matching
        # compile_cxx_core_function's own suffix check; the caller's exact
        # spelling is preserved either way.
        raise ValueError(
            f"ExternalFunction '{name}': inline=True emits LLVM IR, so "
            f"object_file_name must end in '.ll' (textual LLVM IR) or "
            f"'.bc' (bitcode); got '{object_file_name}'."
        )
    super().__init__(
        effective_name,
        object_file_name,
        arg_types,
        link_with_mode="merge" if inline else None,
        stack_size_override=stack_size_override,
    )

    if source_file is None and source_string is None:
        raise ValueError("source_file or source_string must be provided.")
    if source_file is not None:
        try:
            source_bytes = Path(source_file).read_bytes()
        except OSError:
            source_bytes = f"<unreadable:{source_file}>".encode()
    else:
        assert source_string is not None
        source_bytes = source_string.encode()
    self._object_file = KernelObject(
        object_file_name,
        "merge" if inline else None,
        _KernelSource(
            str(Path(source_file).resolve()) if source_file is not None else None,
            source_string if source_file is None else None,
            hashlib.sha256(source_bytes).hexdigest(),
            tuple(str(Path(d).absolute()) for d in include_dirs or ()),
            tuple(compile_flags or ()),
            use_chess,
            symbol_prefix,
            name if inline else None,
        ),
    )
    self._cached_digest: str | None = None

    # A translation unit may export several symbols, but an output path
    # must have only one recipe. Auto-suffix default names on conflict;
    # explicit names are never silently changed.
    for existing in ExternalFunction._instances:
        if (
            existing.object_file_name == object_file_name
            and existing.object_file._source != self.object_file._source
        ):
            if object_file_name_explicit:
                raise ValueError(
                    f"ExternalFunction '{effective_name}' would collide with "
                    f"an already-registered instance: same "
                    f"explicit object_file_name='{object_file_name}' but "
                    f"different compile_flags / source.  Distinguish them "
                    f"by passing a distinct `object_file_name=...`."
                )
            suffix = self._content_digest()[:8]
            output_path = Path(object_file_name)
            object_file_name = str(
                output_path.with_name(
                    f"{output_path.stem}_{suffix}{output_path.suffix}"
                )
            )
            self._object_file = replace(self.object_file, name=object_file_name)
            self._cached_digest = None
            break
    for existing in ExternalFunction._instances:
        if existing.object_file_name == self.object_file_name:
            if existing.object_file._source != self.object_file._source:
                raise ValueError(
                    f"ExternalFunction '{effective_name}' would collide on '{self.object_file_name}'"
                )
            self._object_file = existing.object_file
            break
    ExternalFunction._instances.add(self)

source_file property

source_file: str | None

Path to the C/C++ source on disk, or None for inline source.

source_string property

source_string: str | None

Inline C/C++ source text, or None when compiled from a file.

include_dirs property

include_dirs: list[str]

Copy of the extra -I directories passed to the compiler.

compile_flags property

compile_flags: list[str]

Copy of the extra flags passed verbatim to the compiler.

use_chess property

use_chess: bool

True when this kernel's object is built with xchesscc, not Peano.

param_values

param_values(inputs: list) -> list

Pick the Param arrays out of one logical input list.

inputs is one array per unbound In/tensor Param in argument order. A design bakes tensor Param arguments into core buffers rather than streaming them, so it needs them separately.

Source code in python/iron/kernel.py
def param_values(self, inputs: list) -> list:
    """Pick the ``Param`` arrays out of one logical input list.

    ``inputs`` is one array per unbound ``In``/tensor ``Param`` in argument
    order. A design bakes tensor ``Param`` arguments into core buffers
    rather than streaming them, so it needs them separately.
    """
    from .kernels._common import Param, _is_tensor_type

    c = self._require_contract()
    types = self.arg_types()
    c.validate_types(types)
    positions = [i for i in c.reference_indices() if _is_tensor_type(types[i])]
    if len(inputs) != len(positions):
        raise ValueError(f"{self.name}: expected {len(positions)} input arrays")
    return [np.asarray(a) for a, i in zip(inputs, positions) if c.roles[i] is Param]

input_limit

input_limit(
    dtype, *, reduction: int | None = None
) -> int | None

Largest integer magnitude an input may take without overflowing.

From the contract's acc_dtype and reduction (reduction overrides the per-call value, e.g. with the full K of a tiled matmul): with two or more multiplied inputs every product of two limits summed reduction times must fit the accumulator with a factor-4 margin; with one input the sum of reduction limits must. None for float inputs, or when the contract declares no accumulator.

The output dtype does not bound this. What a kernel does when a result leaves the output range is its reference's to model, and clipping inputs to the output range would leave a requantizing kernel's data near zero.

Source code in python/iron/kernel.py
def input_limit(self, dtype, *, reduction: int | None = None) -> int | None:
    """Largest integer magnitude an input may take without overflowing.

    From the contract's ``acc_dtype`` and ``reduction`` (``reduction``
    overrides the per-call value, e.g. with the full ``K`` of a tiled
    matmul): with two or more multiplied inputs every product of two
    limits summed ``reduction`` times must fit the accumulator with a
    factor-4 margin; with one input the sum of ``reduction`` limits must.
    ``None`` for float inputs, or when the contract declares no
    accumulator.

    The output dtype does not bound this. What a kernel does when a
    result leaves the output range is its reference's to model, and
    clipping inputs to the output range would leave a requantizing
    kernel's data near zero.
    """
    from aie.utils.compile.jit.markers import In

    from .kernels._common import Param, _is_tensor_type

    c = self._require_contract()
    types = self.arg_types()
    c.validate_types(types)
    dt = np.dtype(dtype)
    if not np.issubdtype(dt, np.integer):
        return None
    if c.acc_dtype is None or not np.issubdtype(np.dtype(c.acc_dtype), np.integer):
        return None
    n = reduction or c.reduction or 1
    budget = np.iinfo(c.acc_dtype).max // 4
    n_tensors = sum(
        r in (In, Param) and _is_tensor_type(t) for r, t in zip(c.roles, types)
    )
    limit = int(np.sqrt(budget // n)) if n_tensors >= 2 else budget // n
    return max(1, min(limit, int(np.iinfo(dt).max)))

expected

expected(inputs: list, *, scalars: tuple = ())

Return reference output(s), cast to each output argument's dtype.

Source code in python/iron/kernel.py
def expected(self, inputs: list, *, scalars: tuple = ()):
    """Return reference output(s), cast to each output argument's dtype."""
    from aie.helpers.npdtypes import v8bfp16ebs8

    c = self._require_contract()
    if c.reference is None:
        raise ValueError(f"{self.name}: contract has no reference")
    result = c.reference(*self._reference_args(inputs, scalars))
    multiple = len(c.out_indices) > 1
    results = result if multiple else (result,)
    if multiple and (
        not isinstance(results, tuple) or len(results) != len(c.out_indices)
    ):
        raise ValueError("reference must return one tuple entry per output")
    outputs = []
    for i, value in zip(c.out_indices, results):
        dt = self.arg_dtype(i)
        outputs.append(
            np.asarray(value).astype(np.float32 if dt is v8bfp16ebs8 else dt)
        )
    return tuple(outputs) if multiple else outputs[0]

output_dtype

output_dtype(ref_dtype=None)

Host dtype(s) of the device output buffer(s).

Defaults to the declared argument dtypes, with bfp16ebs8 represented as packed bytes. The optional reference dtype override is retained for compatibility. Multiple outputs return a tuple in argument order.

Source code in python/iron/kernel.py
def output_dtype(self, ref_dtype=None):
    """Host dtype(s) of the device output buffer(s).

    Defaults to the declared argument dtypes, with bfp16ebs8 represented
    as packed bytes. The optional reference dtype override is retained
    for compatibility. Multiple outputs return a tuple in argument order.
    """
    import numpy as np
    from aie.helpers.npdtypes import v8bfp16ebs8

    outputs = self._require_contract().out_indices
    multiple = len(outputs) > 1
    dtypes = (
        tuple(self.arg_dtype(i) for i in outputs)
        if ref_dtype is None
        else ref_dtype if multiple else (ref_dtype,)
    )
    if multiple and (
        not isinstance(dtypes, (tuple, list)) or len(dtypes) != len(outputs)
    ):
        raise ValueError("provide one reference dtype per output")
    result = tuple(
        np.uint8 if self.arg_dtype(i) is v8bfp16ebs8 else dt
        for i, dt in zip(outputs, dtypes)
    )
    return result if multiple else result[0]

judge

judge(
    got,
    ref,
    *,
    calls: int = 1,
    tolerance=None,
    inputs: list | None = None,
    scalars: tuple = ()
)

Compare a flat device output against a reference under the contract.

Declared layouts decode each output into logical tiles. DMA padding is trimmed per call. One Verdict summarizes all outputs and is false if any output fails; its detail identifies the failing output. Without streamed inputs, a complete one-call reference may be repeated. With no streamed inputs, a one-tile reference is repeated across calls. A tolerance that is a function of the inputs needs the inputs and scalars the reference was given.

Source code in python/iron/kernel.py
def judge(
    self,
    got,
    ref,
    *,
    calls: int = 1,
    tolerance=None,
    inputs: list | None = None,
    scalars: tuple = (),
):
    """Compare a flat device output against a reference under the contract.

    Declared layouts decode each output into logical tiles. DMA padding
    is trimmed per call. One Verdict summarizes all outputs and is false
    if any output fails; its detail identifies the failing output.
    Without streamed inputs, a complete one-call reference may be repeated.
    With no streamed inputs, a one-tile reference is repeated across calls.
    A tolerance that is a function of the inputs needs the ``inputs`` and
    ``scalars`` the reference was given.
    """
    from aie.utils.compile.jit.markers import In
    from aie.utils.verify import Tolerance, Verdict, compare

    c = self._require_contract()
    multiple = len(c.out_indices) > 1
    actuals, references = (got, ref) if multiple else ((got,), (ref,))
    if (
        not isinstance(actuals, (tuple, list))
        or not isinstance(references, (tuple, list))
        or len(actuals) != len(c.out_indices)
        or len(references) != len(c.out_indices)
    ):
        raise ValueError("provide one actual and reference array per output")
    tol = tolerance or c.tolerance
    bounds = (None,) * len(c.out_indices)
    if tol is not None and tol.kind == "bound":
        if inputs is None:
            raise ValueError(
                f"{self.name}: its tolerance is a function of the inputs; "
                "pass the reference's inputs= (and scalars=) to judge"
            )
        bounds = tol.bound(*self._reference_args(inputs, scalars))
        bounds = bounds if multiple else (bounds,)
    verdicts = []
    for i, actual, reference, bound in zip(
        c.out_indices, actuals, references, bounds
    ):
        layout = c.layouts[i] if c.layouts else None
        got = layout.decode(actual, calls=calls) if layout else np.asarray(actual)
        got, ref = got.reshape(calls, -1), np.asarray(reference)
        if c.out_valid is not None:
            got = got[:, : c.out_valid]
        if In not in c.roles and ref.size == got.shape[1]:
            ref = np.broadcast_to(ref.reshape(1, -1), got.shape)
        else:
            ref = ref.reshape(calls, -1)
        verdicts.append(
            compare(
                got,
                ref,
                tol or Tolerance.default_for(ref.dtype),
                range_axis=1,
                bound=None if bound is None else np.reshape(bound, ref.shape),
            )
        )
    if not multiple:
        return verdicts[0]
    first_bad, offset = None, 0
    for result in verdicts:
        if first_bad is None and result.first_bad_index is not None:
            first_bad = offset + result.first_bad_index
        offset += result.n_checked
    ulps = [v.max_ulp_err for v in verdicts if v.max_ulp_err is not None]
    return Verdict(
        ok=all(verdicts),
        n_checked=sum(v.n_checked for v in verdicts),
        n_mismatch=sum(v.n_mismatch for v in verdicts),
        max_abs_err=max(v.max_abs_err for v in verdicts),
        max_ulp_err=max(ulps) if ulps else None,
        first_bad_index=first_bad,
        detail="; ".join(
            f"output {i} (argument {arg}): {v.detail}"
            for i, (arg, v) in enumerate(zip(c.out_indices, verdicts))
        ),
    )

ScratchpadParameter

ScratchpadParameter: a named runtime value set from the host and read by Workers.

ScratchpadParameter

ScratchpadParameter(name: str, dtype: NpuDType)

Bases: Resolvable

A named runtime parameter communicated from host to AIE cores via the scratchpad mechanism.

Declare a ScratchpadParameter at design time. Pass it to a Worker via fn_args and call read inside the core_fn to obtain its current value. The --aie-lower-scratchpad-parameters pass automatically inserts the necessary lock and scratchpad-sync preamble ops.

Example
import numpy as np
from aie.iron import ScratchpadParameter, Worker, Runtime, Program

seq_len = ScratchpadParameter("seq_len", np.int32)

def core_body(p):
    v = p.read()
    ...

worker = Worker(core_body, [seq_len])

def sequence(out):
    # The compiler automatically inserts the parameter-sync preamble.
    ...

rt = Runtime(sequence, [output_type])

Create a ScratchpadParameter.

Parameters:

Name Type Description Default
name str

Symbol name for the parameter (must be unique within the device).

required
dtype NpuDType

The numpy scalar type (e.g. np.int32, np.int16, bfloat16). np.float32 is not supported -- the scratchpad encoding zeroes the top 2 bits of the value, which clobbers the sign and top exponent bits of an f32.

required
Source code in python/iron/scratchpad_parameter.py
def __init__(self, name: str, dtype: NpuDType):
    """Create a ScratchpadParameter.

    Args:
        name: Symbol name for the parameter (must be unique within the
              device).
        dtype: The numpy scalar type (e.g. ``np.int32``, ``np.int16``,
               ``bfloat16``).  ``np.float32`` is not supported -- the
               scratchpad encoding zeroes the top 2 bits of the value,
               which clobbers the sign and top exponent bits of an f32.
    """
    self._name = name
    self._dtype = dtype
    self._resolved = False

name property

name: str

The symbol name of this parameter.

dtype property

dtype: NpuDType

The numpy scalar type of this parameter.

read

read() -> Value

Emit aiex.read_scratchpad_parameter inside a core body.

Must be called within an active MLIR insertion point (i.e. inside a Worker's core_fn).

Returns:

Type Description
Value

An MLIR SSA value of the parameter's type.

Source code in python/iron/scratchpad_parameter.py
def read(self) -> "ir.Value":
    """Emit `aiex.read_scratchpad_parameter` inside a core body.

    Must be called within an active MLIR insertion point (i.e. inside a
    Worker's `core_fn`).

    Returns:
        An MLIR SSA value of the parameter's type.
    """
    mlir_type = np_dtype_to_mlir_type(self._dtype)
    return aiex.read_scratchpad_parameter(self._name, mlir_type)

resolve

resolve(
    loc: Location | None = None,
    ip: InsertionPoint | None = None,
) -> None

Emit aiex.scratchpad_parameter @name : type at module scope.

Source code in python/iron/scratchpad_parameter.py
def resolve(
    self,
    loc: ir.Location | None = None,
    ip: ir.InsertionPoint | None = None,
) -> None:
    """Emit ``aiex.scratchpad_parameter @name : type`` at module scope."""
    if not self._resolved:
        mlir_type = np_dtype_to_mlir_type(self._dtype)
        aiex.scratchpad_parameter(  # pyright: ignore[reportAttributeAccessIssue]
            self._name, mlir_type, loc=loc, ip=ip
        )
        self._resolved = True

Compile-time & JIT

Decorators and markers for JIT-compiling a design and injecting compile-time constants. These are re-exported into iron from aie.utils.

Note

The JIT entry point and tensor factories below are thin re-exports from the compiled aie.utils package. Their full signatures and source are available in the running package; the summaries here describe the public contract.

Symbol Kind Summary
iron.jit decorator Compile a design on a cache miss, then run it on the attached NPU.
iron.CompilableDesign class Bundle a design generator with its compile-time configuration.
iron.CallableDesign class A compiled, callable design produced from a CompilableDesign.
iron.compileconfig decorator Attach compile-time configuration to a design generator.
iron.get_compile_arg function Dynamically inject a compile-time argument (advanced).
iron.In / iron.Out / iron.InOut markers Type-annotation markers for design inputs/outputs.
iron.CompileTime marker Type-annotation marker for a compile-time constant argument.
iron.DispatchTime marker Integer scalar that can vary per call without recompiling the device program.

For dispatch scalar defaults, specialization, and runtime binding, see Dispatch-time scalars in the runtime data-movement guide.

See the Programming Guide for worked examples of @iron.jit.


Tensor factories

NumPy-like helpers that allocate NPU-accessible host tensors. Re-exported into iron from aie.utils.

Symbol Summary
iron.tensor Wrap existing data as an NPU-accessible tensor.
iron.arange NPU-accessible analogue of numpy.arange.
iron.zeros / iron.ones / iron.full Allocate a tensor filled with 0, 1, or a constant.
iron.zeros_like Allocate a zero tensor matching another's shape/dtype.
iron.rand / iron.randint Allocate a tensor of random floats / integers.

Device management

Symbol Summary
iron.get_current_device Return the currently selected NPU device.
iron.set_current_device Select the NPU device for subsequent allocations.
iron.ensure_current_device Raise if no device is currently selected.

Data type helpers

Utilities for converting between short string names and numpy dtype objects.

str_to_dtype

str_to_dtype(dtype_str: str) -> type

Convert a string representation of a data type to its corresponding dtype object.

Parameters:

Name Type Description Default
dtype_str str

The string representation of the data type.

required

Returns:

Type Description
type

The corresponding numpy/ml_dtypes type object.

Source code in python/iron/dtype.py
def str_to_dtype(dtype_str: str) -> type:
    """Convert a string representation of a data type to its corresponding dtype object.

    Args:
        dtype_str: The string representation of the data type.

    Returns:
        The corresponding numpy/ml_dtypes type object.
    """
    value = None
    try:
        value = dtype_map[dtype_str]
    except KeyError:
        raise ValueError(f"Unrecognized dtype: {dtype_str}")
    return value

dtype_to_str

dtype_to_str(dtype: type) -> str

Convert a dtype object to its string representation.

Parameters:

Name Type Description Default
dtype type

The dtype object to convert.

required

Returns:

Type Description
str

The string representation of the dtype.

Source code in python/iron/dtype.py
def dtype_to_str(dtype: type) -> str:
    """Convert a dtype object to its string representation.

    Args:
        dtype: The dtype object to convert.

    Returns:
        The string representation of the dtype.
    """
    for key, value in dtype_map.items():
        if value == dtype:
            return key
    raise ValueError(f"Unrecognized dtype: {dtype}")

Advanced primitives

Still part of the aie.iron API — but reach for these only when the managed ObjectFifo abstraction is not enough and you need explicit control over routing, DMA descriptors, and locks.

Flow / PacketFlow

Circuit-switched (Flow) and packet-switched (PacketFlow) stream connections, plus the PacketDest endpoint descriptor.

IRON-level circuit- and packet-switched route primitives.

Two classes live here: Flow (circuit-switched) and PacketFlow (packet-switched, with explicit packet IDs), plus the small PacketDest dataclass PacketFlow uses for its destination list. They share a private _emit_shim_dma_alloc helper and are treated as a sibling pair by dataflow/__init__.py; splitting them across two modules would either duplicate the helper or require a third file to hold it.

Both are peers of ObjectFifo in the dataflow namespace. ObjectFifo wraps route + buffers + locks + DMA into one circular-buffer abstraction; Flow / PacketFlow are the lower-level "just declare the route" primitives, paired with explicit TileDma programs (and Buffer / Lock shared state) for designs that need direct control.

Flow

Flow(
    src: Tile,
    dst: Tile,
    *,
    src_port: WireBundle = DMA,
    src_channel: int = 0,
    dst_port: WireBundle = DMA,
    dst_channel: int = 0,
    shim_symbol: str | None = None
)

Bases: Resolvable

An explicit AXI-stream route between a source and destination endpoint.

Connects (src_tile, src_port, src_channel) to (dst_tile, dst_port, dst_channel). Lowers to a single aie.flow op. The user is responsible for arranging matching TileDma channels on the producer and consumer ends.

Construct a Flow.

Parameters:

Name Type Description Default
src Tile

The source tile.

required
dst Tile

The destination tile.

required
src_port WireBundle

The source port bundle. Defaults to DMA.

DMA
src_channel int

The source channel. Defaults to 0.

0
dst_port WireBundle

The destination port bundle. Defaults to DMA.

DMA
dst_channel int

The destination channel. Defaults to 0.

0
shim_symbol str | None

Name for the aie.shim_dma_allocation this Flow emits for its shim endpoint. Only needed to refer to the channel by name from elsewhere (e.g. a raw shim_dma_single_bd_task("symbol", ...)); fill/drain name it themselves. Direction is inferred: shim-as-source → MM2S, shim-as-dest → S2MM.

None
Source code in python/iron/dataflow/flow.py
def __init__(
    self,
    src: Tile,
    dst: Tile,
    *,
    src_port: WireBundle = WireBundle.DMA,
    src_channel: int = 0,
    dst_port: WireBundle = WireBundle.DMA,
    dst_channel: int = 0,
    shim_symbol: str | None = None,
):
    """Construct a Flow.

    Args:
        src (Tile): The source tile.
        dst (Tile): The destination tile.
        src_port (WireBundle): The source port bundle.  Defaults to DMA.
        src_channel (int): The source channel.  Defaults to 0.
        dst_port (WireBundle): The destination port bundle.  Defaults to DMA.
        dst_channel (int): The destination channel.  Defaults to 0.
        shim_symbol (str | None): Name for the ``aie.shim_dma_allocation``
            this Flow emits for its shim endpoint. Only needed to refer to
            the channel by name from elsewhere (e.g. a raw
            ``shim_dma_single_bd_task("symbol", ...)``); ``fill``/``drain``
            name it themselves. Direction is inferred: shim-as-source →
            MM2S, shim-as-dest → S2MM.
    """
    self._src = src
    self._dst = dst
    self._src_port = src_port
    self._src_channel = src_channel
    self._dst_port = dst_port
    self._dst_channel = dst_channel
    self._shim_symbol = shim_symbol
    self._op = None

all_tiles

all_tiles()

Return the tiles this Flow touches — Program uses this to resolve them.

Source code in python/iron/dataflow/flow.py
def all_tiles(self):
    """Return the tiles this Flow touches — Program uses this to resolve them."""
    return [self._src, self._dst]

fill

fill(source, **kwargs)

Send data from the source runtime buffer into this route.

Call from within a Runtime sequence body, on a Flow with exactly one shim endpoint, at the source. See emit_shim_transfer for the keyword arguments; returns a Task handle to the transfer.

Source code in python/iron/dataflow/flow.py
def fill(self, source, **kwargs):
    """Send data from the ``source`` runtime buffer into this route.

    Call from within a [`Runtime`][iron.Runtime] sequence body, on a Flow
    with exactly one shim endpoint, at the source. See ``emit_shim_transfer``
    for the keyword arguments; returns a
    [`Task`][iron.runtime.dmataskhandle.Task] handle to the transfer.
    """
    return self._transfer(source, DMAChannelDir.MM2S, **kwargs)

drain

drain(dest, **kwargs)

Receive data from this route into the dest runtime buffer.

Call from within a Runtime sequence body, on a Flow with exactly one shim endpoint, at the destination. See emit_shim_transfer for the keyword arguments; returns a Task handle to the transfer.

Source code in python/iron/dataflow/flow.py
def drain(self, dest, **kwargs):
    """Receive data from this route into the ``dest`` runtime buffer.

    Call from within a [`Runtime`][iron.Runtime] sequence body, on a Flow
    with exactly one shim endpoint, at the destination. See
    ``emit_shim_transfer`` for the keyword arguments; returns a
    [`Task`][iron.runtime.dmataskhandle.Task] handle to the transfer.
    """
    return self._transfer(dest, DMAChannelDir.S2MM, **kwargs)

PacketDest dataclass

PacketDest(
    tile: Tile, port: WireBundle = DMA, channel: int = 0
)

One destination endpoint of a PacketFlow.

Held as a small dataclass so the PacketFlow constructor's destination list reads cleanly when there are multiple sinks (uncommon, but the underlying op supports it).

PacketFlow

PacketFlow(
    pkt_id: int,
    src: Tile,
    dst: Tile,
    *,
    src_port: WireBundle = DMA,
    src_channel: int = 0,
    dst_port: WireBundle = DMA,
    dst_channel: int = 0,
    extra_dsts: Sequence[PacketDest] = (),
    keep_pkt_header: bool = False,
    shim_symbol: str | None = None
)

Bases: Resolvable

An explicit packet-switched route from a source to one or more destinations.

Connects (src_tile, src_port, src_channel) to each destination endpoint, tagging the stream with pkt_id. Lowers to a single aie.packetflow op holding one aie.packet_source and one aie.packet_dest per destination. The user is responsible for arranging matching TileDma channels on the producer and consumer ends.

Construct a PacketFlow.

Parameters:

Name Type Description Default
pkt_id int

The packet ID — the same byte the routing fabric uses to dispatch. Caller controls the value (often reused across stages so a memtile can re-emit packets keeping the original ID for downstream routing).

required
src Tile

Source tile.

required
dst Tile

Primary destination tile.

required
src_port WireBundle

Source port bundle (as for Flow).

DMA
src_channel int

Source channel (as for Flow).

0
dst_port WireBundle

Destination port bundle (as for Flow).

DMA
dst_channel int

Destination channel (as for Flow).

0
extra_dsts Sequence[PacketDest]

Additional destination endpoints if this packet needs to fan out. Each is a PacketDest.

()
keep_pkt_header bool

If True, downstream tile receives the 4-byte packet header alongside the payload (useful when the receiver needs to re-emit with the same pkt_id). Defaults to False.

False
shim_symbol str | None

Same meaning as on Flow — auto-emit a matching aie.shim_dma_allocation when one endpoint is a shim tile.

None
Source code in python/iron/dataflow/flow.py
def __init__(
    self,
    pkt_id: int,
    src: Tile,
    dst: Tile,
    *,
    src_port: WireBundle = WireBundle.DMA,
    src_channel: int = 0,
    dst_port: WireBundle = WireBundle.DMA,
    dst_channel: int = 0,
    extra_dsts: Sequence[PacketDest] = (),
    keep_pkt_header: bool = False,
    shim_symbol: str | None = None,
):
    """Construct a PacketFlow.

    Args:
        pkt_id: The packet ID — the same byte the routing fabric uses to
            dispatch.  Caller controls the value (often reused across
            stages so a memtile can re-emit packets keeping the original
            ID for downstream routing).
        src: Source tile.
        dst: Primary destination tile.
        src_port: Source port bundle (as for [`Flow`][iron.Flow]).
        src_channel: Source channel (as for [`Flow`][iron.Flow]).
        dst_port: Destination port bundle (as for [`Flow`][iron.Flow]).
        dst_channel: Destination channel (as for [`Flow`][iron.Flow]).
        extra_dsts: Additional destination endpoints if this packet needs
            to fan out. Each is a [`PacketDest`][iron.PacketDest].
        keep_pkt_header: If `True`, downstream tile receives the 4-byte
            packet header alongside the payload (useful when the receiver
            needs to re-emit with the same pkt_id). Defaults to `False`.
        shim_symbol: Same meaning as on [`Flow`][iron.Flow] — auto-emit a
            matching `aie.shim_dma_allocation` when one endpoint is a
            shim tile.
    """
    self._pkt_id = pkt_id
    self._src = src
    self._dst = dst
    self._src_port = src_port
    self._src_channel = src_channel
    self._dst_port = dst_port
    self._dst_channel = dst_channel
    self._extra_dsts: list[PacketDest] = list(extra_dsts)
    self._keep_pkt_header = keep_pkt_header
    self._shim_symbol = shim_symbol
    self._op = None

CascadeFlow

Directed cascade-stream connection between two adjacent Workers.

CascadeFlow: a directed cascade stream connection between two Workers.

CascadeFlow

CascadeFlow(src: 'Worker', dst: 'Worker')

Bases: Resolvable

A directed cascade stream connection from one Worker to another.

Construct one of these for each cascade edge in your design:

CascadeFlow(producer_worker, consumer_worker)

Lowers to aie.cascade_flow(producer.tile, consumer.tile) after both Workers are placed. The kernel functions are responsible for using the put_mcd / get_scd intrinsics to actually drive/read the cascade stream — this object only declares the directed topology edge.

Hardware constraints (enforced by the underlying op verifier):

  • Source and destination tiles must be cardinal-adjacent.
  • Each compute tile has at most one cascade input (from N or W) and one cascade output (to S or E). Multiple cascade outputs from the same tile will fail at lowering, not at construction.
  • ShimTiles and MemTiles do not have cascade interfaces.

Discovery: each newly-constructed CascadeFlow registers itself on its source Worker's _outgoing_cascades list. Program.resolve() walks the runtime's workers and resolves each worker's outgoing cascades after placement — no global registry, no drain step.

Construct a CascadeFlow.

Parameters:

Name Type Description Default
src 'Worker'

Source Worker whose tile drives the cascade stream.

required
dst 'Worker'

Destination Worker whose tile reads the cascade stream.

required
Source code in python/iron/dataflow/cascadeflow.py
def __init__(self, src: "Worker", dst: "Worker"):
    """Construct a CascadeFlow.

    Args:
        src: Source `Worker` whose tile drives the cascade stream.
        dst: Destination `Worker` whose tile reads the cascade stream.
    """
    self._src = src
    self._dst = dst
    # Self-register on the source Worker so Program.resolve() can find
    # us by walking its workers (the same walk it already does).
    src._outgoing_cascades.append(self)

resolve

resolve(loc=None, ip=None) -> None

Emit aie.cascade_flow(src.tile, dst.tile).

Source code in python/iron/dataflow/cascadeflow.py
def resolve(self, loc=None, ip=None) -> None:
    """Emit ``aie.cascade_flow(src.tile, dst.tile)``."""
    _cascade_flow_op(self._src.tile.op, self._dst.tile.op)

TileDma / DmaChannel / Bd

Explicit tile DMA programs: TileDma, DmaChannel, buffer descriptors (Bd), and the Acquire / Release lock actions.

IRON-level explicit per-tile DMA program.

Peer of Worker (which describes the compute body of a tile). A TileDma describes the DMA engine program for the same (or a different) tile — what each hardware DMA channel does, which buffers it reads/writes, and how it synchronizes with the compute side via locks.

Used together with Flow / PacketFlow (which describe the AXI-stream routes) and explicit Buffer + Lock declarations, for designs where ObjectFifo would hide too much to be useful.

Acquire dataclass

Acquire(
    lock: Lock, value: int = 1, greater_equal: bool = True
)

An aie.use_lock(..., AcquireGreaterEqual|Acquire) op at the start of a BD.

Release dataclass

Release(lock: Lock, value: int = 1)

An aie.use_lock(..., Release) op at the end of a BD.

BdIteration dataclass

BdIteration(size: int, stride: int, current: int = 0)

Iteration state of a buffer descriptor for aie.dma_bd.

Lets one BD cover size sub-buffers over size executions instead of an N-deep chain. Values are true/element: the base advances by stride elements each execution and wraps after size executions; current is the starting step (default 0). The lowering applies the hardware -1 bias and element->word scaling. NOTE: the identically-named iteration_* family on the runtime-sequence path uses RAW register values instead -- do not copy numbers between them.

Bd dataclass

Bd(
    buffer: Buffer,
    offset: int = 0,
    length: int | None = None,
    acquires: list[Acquire] = list(),
    releases: list[Release] = list(),
    next: int | str | None = None,
    packet: tuple[int, int] | None = None,
    bd_id: int | None = None,
    sizes: list = list(),
    strides: list = list(),
    pad_dimensions: list[Sequence[int]] | None = None,
    iteration: BdIteration | None = None,
    out_of_order_id: int | None = None,
)

A single buffer-descriptor entry in a DmaChannel's chain.

Lowers to one basic block containing acquires + aie.dma_bd + releases + an aie.next_bd. The next field selects what the next_bd points at:

  • None (default) — follow the channel: the next entry in bds, and from the last entry either back to the head or out of the chain, per DmaChannel.loop.
  • "self" — the BD loops to itself, whatever the rest of the chain does.
  • an int i — point at the i-th BD in this channel's bds list (zero-based). Useful for explicit cycles in a multi-BD chain.

next is ignored on an out-of-order channel because those BDs are chained only for configuration and the hardware selects by header id (see DmaChannel.out_of_order).

DmaChannel dataclass

DmaChannel(
    direction: DMAChannelDir,
    channel: int,
    bds: list[Bd],
    pad_value: int = 0,
    repeat_count: int = 0,
    out_of_order: bool = False,
    loop: bool = True,
)

One hardware DMA channel on a tile, with its BD chain.

Parameters:

Name Type Description Default
direction DMAChannelDir

DMAChannelDir.S2MM (host→tile) or DMAChannelDir.MM2S (tile→host).

required
channel int

hardware channel index.

required
bds list[Bd]

ordered list of Bd entries that form the chain (in-order) or n-way merge (out-of-order).

required
repeat_count int

extra repeats of the task (0 = run once), where the task is the BD chain (in-order) or a merge round (out-of-order). Only meaningful on a chain that ends -- see loop.

0
loop bool

whether the last BD chains back to the first (the default), making the chain endless. An endless chain is one task that never completes: it runs for as long as its locks let it, which is how ObjectFifo expresses the same thing, and repeat_count has nothing to count and is ignored. loop=False ends the chain after its last BD, making it a task that completes and can be re-run -- which is what gives repeat_count meaning, and what a design reproducing a specific descriptor layout wants. Note a chain that ends runs exactly repeat_count + 1 times, so a loop=False channel expected to move more than one buffer needs a matching count; left at 0 it moves one and stops.

True
out_of_order bool

put the channel into out-of-order mode (S2MM only). Each BD receives the packet with bd.bd_id == pkt.out_of_order_id, and the BD chain (next bd) is ignored. Each BD receives its own BdIteration.size packets per merge round. Every BD must be packet-enabled and the ingress flow must set keep_pkt_header=True. Multiple out-of-order channels must have disjoint BD ids. At the hardware level, repeat_count is the total number of packets to accept (0-based). This class converts repeated merge rounds to that total.

False

TileDma

TileDma(tile: Tile, channels: Iterable[DmaChannel])

Bases: Resolvable

Per-tile DMA program.

Lowers to an aie.mem (compute tile), aie.memtile_dma (memtile), or aie.shim_dma (shim tile) region based on the tile's type.

Parameters:

Name Type Description Default
tile Tile

the tile whose DMA hardware this program targets.

required
channels Iterable[DmaChannel]

ordered list of DmaChannel entries.

required
Source code in python/iron/dataflow/tile_dma.py
def __init__(self, tile: Tile, channels: Iterable[DmaChannel]):
    self._tile = tile
    self._channels: list[DmaChannel] = []
    self.add_channels(channels)
    self._resolved = False

add_channel

add_channel(channel: DmaChannel) -> None

Add a channel to this tile's DMA program.

A tile has one DMA program, so a helper that wires transfers one at a time needs somewhere to put the second channel it wants on a tile it has already reached.

Source code in python/iron/dataflow/tile_dma.py
def add_channel(self, channel: DmaChannel) -> None:
    """Add a channel to this tile's DMA program.

    A tile has one DMA program, so a helper that wires transfers one at a
    time needs somewhere to put the second channel it wants on a tile it has
    already reached.
    """
    self.add_channels([channel])

add_channels

add_channels(channels: Iterable[DmaChannel]) -> None

Add channels after checking all hardware channel keys.

Source code in python/iron/dataflow/tile_dma.py
def add_channels(self, channels: Iterable[DmaChannel]) -> None:
    """Add channels after checking all hardware channel keys."""
    channels = list(channels)
    keys = {(channel.direction, channel.channel) for channel in self._channels}
    for channel in channels:
        key = (channel.direction, channel.channel)
        if key in keys:
            raise ValueError(
                f"TileDma for {self._tile} already has "
                f"{channel.direction} channel {channel.channel}."
            )
        keys.add(key)
    self._channels.extend(channels)

all_buffers_and_locks

all_buffers_and_locks()

Iterate every Buffer + Lock this program touches.

Program uses this to make sure they're all resolved before us.

Source code in python/iron/dataflow/tile_dma.py
def all_buffers_and_locks(self):
    """Iterate every Buffer + Lock this program touches.

    Program uses this to make sure they're all resolved before us.
    """
    seen_buffers: list[Buffer] = []
    seen_locks: list[Lock] = []
    for ch in self._channels:
        for bd in ch.bds:
            if bd.buffer not in seen_buffers:
                seen_buffers.append(bd.buffer)
            for use in (*bd.acquires, *bd.releases):
                if use.lock not in seen_locks:
                    seen_locks.append(use.lock)
    return seen_buffers, seen_locks

Lock

IRON-level Lock primitive — a named aie.lock on a specific tile.

Pairs with Buffer for designs that wire DMA / compute synchronization explicitly (via TileDma and Flow) instead of letting ObjectFifo manage it.

Lock

Lock(
    tile: Tile,
    lock_id: int | None = None,
    init: int = 0,
    name: str | None = None,
)

Bases: Resolvable

A named hardware lock on a specific tile.

Construct a Lock.

Parameters:

Name Type Description Default
tile Tile

The tile that owns this lock.

required
lock_id int | None

Hardware lock ID; passed straight through to the underlying aie.lock op. If None (the default), the lowering pass picks one.

None
init int

Initial lock value at design startup. Defaults to 0.

0
name str | None

Symbol name for the lock. A unique name is generated if not provided.

None
Source code in python/iron/lock.py
def __init__(
    self,
    tile: Tile,
    lock_id: int | None = None,
    init: int = 0,
    name: str | None = None,
):
    """Construct a Lock.

    Args:
        tile (Tile): The tile that owns this lock.
        lock_id (int | None): Hardware lock ID; passed straight through to
            the underlying `aie.lock` op. If `None` (the default),
            the lowering pass picks one.
        init (int): Initial lock value at design startup. Defaults to 0.
        name (str | None): Symbol name for the lock. A unique name is
            generated if not provided.
    """
    self._tile = tile
    self._lock_id = lock_id
    self._init = init
    self._name = name or f"lock_{next(Lock._glock_index)}"
    self._op = None

acquire

acquire(value: int = 1) -> None

Emit aie.use_lock(self, AcquireGreaterEqual, value=value).

The default AcquireGreaterEqual mode matches what almost every ObjectFifo / DMA-driven design wants; use acquire_exact for the rarer Acquire (exact-equality) mode.

Source code in python/iron/lock.py
def acquire(self, value: int = 1) -> None:
    """Emit `aie.use_lock(self, AcquireGreaterEqual, value=value)`.

    The default `AcquireGreaterEqual` mode matches what almost every
    ObjectFifo / DMA-driven design wants; use
    [`acquire_exact`][iron.lock.Lock.acquire_exact] for the rarer
    `Acquire` (exact-equality) mode.
    """
    _use_lock(self.op, LockAction.AcquireGreaterEqual, value=value)

acquire_exact

acquire_exact(value: int = 1) -> None

Emit aie.use_lock(self, Acquire, value=value) (exact match).

Source code in python/iron/lock.py
def acquire_exact(self, value: int = 1) -> None:
    """Emit `aie.use_lock(self, Acquire, value=value)` (exact match)."""
    _use_lock(self.op, LockAction.Acquire, value=value)

release

release(value: int = 1) -> None

Emit aie.use_lock(self, Release, value=value).

Source code in python/iron/lock.py
def release(self, value: int = 1) -> None:
    """Emit `aie.use_lock(self, Release, value=value)`."""
    _use_lock(self.op, LockAction.Release, value=value)

Runtime tasks

Lower-level runtime task types scheduled by the Runtime.

RuntimeTask

RuntimeTask(task_group: TaskGroup | None = None)

Bases: Resolvable

A RuntimeTask is a task to be performed during runtime. A task may be synchronous or asynchronous.

Construct a RuntimeTask. It may be associated with a TaskGroup.

Parameters:

Name Type Description Default
task_group TaskGroup | None

The TaskGroup associated with this task. Defaults to None.

None
Source code in python/iron/runtime/task.py
def __init__(self, task_group: TaskGroup | None = None):
    """Construct a RuntimeTask. It may be associated with a TaskGroup.

    Args:
        task_group (TaskGroup | None, optional): The TaskGroup associated with this task. Defaults to None.
    """
    self._task_group = task_group

task_group property

task_group: TaskGroup | None

The TaskGroup associated with this task.

TaskGroup: groups related runtime transfers so they are awaited/freed together.

TaskGroup

TaskGroup(id: int | None = None)

A grouping of runtime data transfers awaited and freed together.

Construct one inside a runtime sequence body and pass it as the group= argument to fifo.fill(...) / fifo.drain(...). Call finish to await the group's waited transfers and free the rest (waits are ordered before frees).

def seq(A, C):
    tg = TaskGroup()
    inA.prod().fill(A, group=tg)
    outC.cons().drain(C, wait=True, group=tg)
    tg.finish()

Construct a TaskGroup, registering it with the active runtime sequence.

Parameters:

Name Type Description Default
id int | None

Group id, unique within a Runtime. Defaults to the active sequence's next id. Passing an explicit id is only needed for the runtime's internal default group.

None
Source code in python/iron/runtime/taskgroup.py
def __init__(self, id: int | None = None):
    """Construct a TaskGroup, registering it with the active runtime sequence.

    Args:
        id (int | None): Group id, unique within a Runtime. Defaults to the
            active sequence's next id. Passing an explicit id is only needed
            for the runtime's internal default group.
    """
    # Actions accumulated for this group: (dma_await_task | dma_free_task, [task]).
    self._actions: list = []
    # Lazy import to avoid a cycle (runtime -> taskgroup -> _context).
    from ._context import _active_sequence

    active = _active_sequence.get()
    if id is None:
        if active is None:
            raise RuntimeError(
                "TaskGroup() must be constructed within the function passed "
                "to Runtime(seq_fn, fn_args)."
            )
        id = next(active._runtime._task_group_index)
    self._group_id = id
    if active is not None and id is not None:
        active.register_task_group(self)

group_id property

group_id: int

The id of the task group.

finish

finish() -> None

Await this group's waited transfers, then free the rest.

Source code in python/iron/runtime/taskgroup.py
def finish(self) -> None:
    """Await this group's waited transfers, then free the rest."""
    from ._context import active_sequence

    active_sequence().finish_task_group(self)

DMATask: a RuntimeTask that generates a shim DMA transfer operation.

DMATask

DMATask(
    alloc: str,
    rt_data: RuntimeData,
    tap: TensorAccessPattern | None = None,
    task_group: TaskGroup | None = None,
    wait: bool = False,
    offset_parameter: str | None = None,
    packet: tuple[int, int] | None = None,
    sizes=None,
    strides=None,
    offset=None,
    transfer_len=None,
)

Bases: RuntimeTask

Construct a RuntimeTask that will resolve to a DMA Operation.

Provide the access pattern either as a static tap (TensorAccessPattern) or as explicit sizes/strides/offset/ transfer_len lists whose entries may be runtime SSA values (for the dynamic lowering). The two forms are mutually exclusive.

Parameters:

Name Type Description Default
alloc str

Name of the shim DMA allocation this transfer drives, i.e. the symbol naming the shim channel. An ObjectFifo with a shim endpoint generates an allocation of its own name; a Flow generates one for its shim_symbol.

required
rt_data RuntimeData

The Runtime buffer associated with the operation.

required
tap TensorAccessPattern | None

The static access pattern. Mutually exclusive with sizes/strides/offset/transfer_len.

None
task_group TaskGroup | None

The task group associated with the operation. Defaults to None.

None
wait bool

Whether this task should conclude with a call to await or a call to free. Defaults to False.

False
offset_parameter str | None

Name of a ScratchpadParameter whose value is used as the element offset for this DMA transfer. Defaults to None.

None
packet tuple[int, int] | None

Stamp the shim DMA's BD with a packet header (pkt_type, pkt_id). Pairs with downstream packet-switched routing (e.g. an ObjectFifo built with packet=True or an explicit PacketFlow). Defaults to None.

None
sizes optional

Explicit access-pattern sizes whose entries may be runtime SSA values. Used instead of tap for the dynamic path.

None
strides optional

Explicit access-pattern strides, paired with sizes. Used instead of tap for the dynamic path.

None
offset optional

Explicit access-pattern offset. Used instead of tap for the dynamic path.

None
transfer_len optional

Explicit access-pattern transfer length. Used instead of tap for the dynamic path.

None
Source code in python/iron/runtime/dmatask.py
def __init__(
    self,
    alloc: str,
    rt_data: RuntimeData,
    tap: TensorAccessPattern | None = None,
    task_group: TaskGroup | None = None,
    wait: bool = False,
    offset_parameter: str | None = None,
    packet: tuple[int, int] | None = None,
    sizes=None,
    strides=None,
    offset=None,
    transfer_len=None,
):
    """Construct a RuntimeTask that will resolve to a DMA Operation.

    Provide the access pattern either as a static ``tap``
    (TensorAccessPattern) or as explicit ``sizes``/``strides``/``offset``/
    ``transfer_len`` lists whose entries may be runtime SSA values (for the
    dynamic lowering). The two forms are mutually exclusive.

    Args:
        alloc (str): Name of the shim DMA allocation this transfer drives,
            i.e. the symbol naming the shim channel. An
            [`ObjectFifo`][iron.ObjectFifo] with a shim endpoint generates an
            allocation of its own name; a [`Flow`][iron.Flow] generates one
            for its ``shim_symbol``.
        rt_data (RuntimeData): The Runtime buffer associated with the operation.
        tap (TensorAccessPattern | None, optional): The static access pattern. Mutually exclusive with sizes/strides/offset/transfer_len.
        task_group (TaskGroup | None, optional): The task group associated with the operation. Defaults to None.
        wait (bool, optional): Whether this task should conclude with a call to await or a call to free. Defaults to False.
        offset_parameter (str | None, optional): Name of a ScratchpadParameter whose
            value is used as the element offset for this DMA transfer. Defaults to None.
        packet (tuple[int, int] | None, optional): Stamp the shim DMA's
            BD with a packet header `(pkt_type, pkt_id)`. Pairs with
            downstream packet-switched routing (e.g. an
            [`ObjectFifo`][iron.ObjectFifo] built with ``packet=True`` or
            an explicit [`PacketFlow`][iron.PacketFlow]). Defaults to None.
        sizes (optional): Explicit access-pattern sizes whose entries may be
            runtime SSA values. Used instead of ``tap`` for the dynamic path.
        strides (optional): Explicit access-pattern strides, paired with
            ``sizes``. Used instead of ``tap`` for the dynamic path.
        offset (optional): Explicit access-pattern offset. Used instead of
            ``tap`` for the dynamic path.
        transfer_len (optional): Explicit access-pattern transfer length.
            Used instead of ``tap`` for the dynamic path.
    """
    if tap is not None and any(
        v is not None for v in (sizes, strides, offset, transfer_len)
    ):
        raise ValueError(
            "DMATask: tap and sizes/strides/offset/transfer_len are mutually "
            "exclusive access-pattern specifications; pass only one."
        )
    self._alloc = alloc
    self._rt_data = rt_data
    self._tap = tap
    self._wait = wait
    self._offset_parameter = offset_parameter
    self._packet = packet
    self._sizes = sizes
    self._strides = strides
    self._offset = offset
    self._transfer_len = transfer_len
    self._task = None
    RuntimeTask.__init__(self, task_group)

task property

task

Return the MLIR op generated by resolve().

This handle is used after resolution for await or free operations as dictated by a TaskGroup Finish operation.

will_wait

will_wait() -> bool

Whether this task should conclude with an await operation.

Source code in python/iron/runtime/dmatask.py
def will_wait(self) -> bool:
    """Whether this task should conclude with an await operation."""
    return self._wait

emit_shim_transfer

emit_shim_transfer(
    alloc: str,
    rt_data,
    tap=None,
    wait: bool = False,
    packet: tuple[int, int] | None = None,
    offset_parameter=None,
    group=None,
    sizes=None,
    strides=None,
    offset=None,
    transfer_len=None,
    managed: bool = True,
) -> Task

Emit one shim DMA transfer on the alloc channel, inside the active sequence.

The body shared by the runtime's data-movement verbs -- ObjectFifoHandle.fill/drain and Flow.fill/drain -- which differ only in how they name the shim channel and in the endpoint bookkeeping they do first.

The access pattern is given either as a static tap or as explicit sizes/strides/offset/transfer_len whose entries may be runtime SSA values (the dynamic path). The two forms are mutually exclusive; when neither is given, a linear transfer of the whole buffer is used.

When managed is True (default), the transfer is enrolled in a TaskGroup (explicit group or the sequence's implicit one), which awaits/frees it at group close. When False, the caller owns the transfer's lifetime via the returned Task's .free()/.await_() -- used for hand-rolled software pipelines that carry the task across scf.for iterations.

Returns:

Name Type Description
Task Task

A handle to the transfer.

Lazy imports break the runtime<->scratchpad import cycle.

Source code in python/iron/runtime/dmatask.py
def emit_shim_transfer(
    alloc: str,
    rt_data,
    tap=None,
    wait: bool = False,
    packet: tuple[int, int] | None = None,
    offset_parameter=None,
    group=None,
    sizes=None,
    strides=None,
    offset=None,
    transfer_len=None,
    managed: bool = True,
) -> Task:
    """Emit one shim DMA transfer on the ``alloc`` channel, inside the active sequence.

    The body shared by the runtime's data-movement verbs --
    ``ObjectFifoHandle.fill``/``drain`` and ``Flow.fill``/``drain`` -- which
    differ only in how they name the shim channel and in the endpoint
    bookkeeping they do first.

    The access pattern is given either as a static ``tap`` or as explicit
    ``sizes``/``strides``/``offset``/``transfer_len`` whose entries may be
    runtime SSA values (the dynamic path). The two forms are mutually exclusive;
    when neither is given, a linear transfer of the whole buffer is used.

    When ``managed`` is True (default), the transfer is enrolled in a TaskGroup
    (explicit ``group`` or the sequence's implicit one), which awaits/frees it at
    group close. When False, the caller owns the transfer's lifetime via the
    returned Task's ``.free()``/``.await_()`` -- used for hand-rolled software
    pipelines that carry the task across ``scf.for`` iterations.

    Returns:
        Task: A handle to the transfer.

    Lazy imports break the runtime<->scratchpad import cycle.
    """
    from ..scratchpad_parameter import ScratchpadParameter
    from ._context import active_sequence

    active = active_sequence()
    rt = active._runtime

    if not isinstance(rt_data, RuntimeData):
        raise ValueError(f"Expected a RuntimeData source/dest, got {rt_data}")
    if rt_data not in rt._rt_data:
        raise ValueError(
            f"{rt_data} is not a RuntimeData object declared by sequence()"
        )

    explicit = any(v is not None for v in (sizes, strides, offset, transfer_len))
    if tap is not None and explicit:
        raise ValueError(
            "Pass either tap or sizes/strides/offset/transfer_len, not both."
        )
    if tap is None and not explicit:
        tap = rt_data.default_tap()

    if not managed and group is not None:
        raise ValueError(
            "An unmanaged transfer (managed=False) is not part of a TaskGroup; "
            "do not also pass group=."
        )

    offset_param_name = None
    if offset_parameter is not None:
        if isinstance(offset_parameter, ScratchpadParameter):
            offset_param_name = offset_parameter.name
            if offset_parameter not in rt._scratchpad_parameters:
                rt._scratchpad_parameters.append(offset_parameter)
        else:
            offset_param_name = offset_parameter

    task = DMATask(
        alloc,
        rt_data,
        tap=tap,
        task_group=group,
        wait=wait,
        offset_parameter=offset_param_name,
        packet=packet,
        sizes=sizes,
        strides=strides,
        offset=offset,
        transfer_len=transfer_len,
    )
    if managed:
        active.emit_transfer(task, group)
    else:
        # Emit the BD only; the caller owns await/free via the Task.
        task.resolve()
    # Wrap the transfer's !index result: it is both the scf iter_arg payload
    # and the operand dma_await_task/dma_free_task accept.
    return Task(task.task.result)

Task: a handle to an in-flight shim DMA transfer.

Returned by fifo.fill/fifo.drain. Its handle is the transfer's !index SSA value, so a Task can be carried across scf.for iterations as a range_ iter_args entry for software-pipelined transfers -- and it carries .free() / .await_() verbs so the loop body does not need to reach for the raw aiex.dma_free_task / aiex.dma_await_task dialect ops.

A Task returned by an unmanaged transfer (managed=False) is not enrolled in a TaskGroup's automatic await/free, so the caller owns its lifetime with these verbs -- exactly what a hand-rolled ping-pong needs.

Task

Task(handle)

A handle to a submitted shim DMA transfer.

Wraps the transfer's !index SSA value (handle). Pass a Task as a range_ iter_args entry to carry an in-flight transfer across loop iterations; range_ unwraps it to its handle for the scf.for iter_arg and re-wraps the block argument as a Task for the body.

Source code in python/iron/runtime/dmataskhandle.py
def __init__(self, handle):
    self._handle = handle

handle property

handle

The transfer's !index SSA value (the scf iter_arg payload).

free

free() -> None

Return this transfer's buffer descriptor to the pool (dma_free_task).

Source code in python/iron/runtime/dmataskhandle.py
def free(self) -> None:
    """Return this transfer's buffer descriptor to the pool (``dma_free_task``)."""
    dma_free_task(self._handle)

await_

await_() -> None

Block until this transfer completes (dma_await_task).

The transfer must have been issued with wait=True so it carries a completion token.

Source code in python/iron/runtime/dmataskhandle.py
def await_(self) -> None:
    """Block until this transfer completes (``dma_await_task``).

    The transfer must have been issued with ``wait=True`` so it carries a
    completion token.
    """
    dma_await_task(self._handle)