.. _xrt_native_apis.rst: .. comment:: SPDX-License-Identifier: Apache-2.0 comment:: Copyright (C) 2019-2021 Xilinx, Inc. All rights reserved. comment:: Copyright (C) 2022-2026 Advanced Micro Devices, Inc. All rights reserved. XRT Native APIs =============== XRT exposes host-side APIs in C++ and Python. Native XRT host code must link against the **xrt_coreutil** library. C++ examples in this guide assume a compiler with **ISO C++17** or newer (for example ``-std=c++17``). Example ``g++`` invocation: .. code-block:: shell g++ -g -std=c++17 -I$XILINX_XRT/include -L$XILINX_XRT/lib -o host.exe host.cpp -lxrt_coreutil -pthread For general host code development, C++-based APIs are recommended, hence this document only describes the C++-based API interfaces. The full Doxygen generated C and C++ API documentation can be found in :doc:`xrt_native.main`. The C++ Class objects used for the APIs are the following: +----------------------+---------------------+------------------------------------------------+ | Core Object | C++ Class | Header Files | +======================+=====================+================================================+ | Device | ``xrt::device`` | ``#include `` | +----------------------+---------------------+------------------------------------------------+ | Buffer | ``xrt::bo`` | ``#include `` | +----------------------+---------------------+------------------------------------------------+ | Kernel | ``xrt::kernel`` | ``#include `` | +----------------------+---------------------+------------------------------------------------+ | Run | ``xrt::run`` | ``#include `` | +----------------------+---------------------+------------------------------------------------+ | Run-list | ``xrt::runlist`` | ``#include `` | +----------------------+---------------------+------------------------------------------------+ | Context | ``xrt::hw_context`` | ``#include `` | +----------------------+---------------------+------------------------------------------------+ | Xclbin | ``xrt::xclbin`` | ``#include `` | +----------------------+---------------------+------------------------------------------------+ | Control code (ELF) | ``xrt::elf`` | ``#include `` | +----------------------+---------------------+------------------------------------------------+ | User-managed Kernel | ``xrt::ip`` | ``#include `` | +----------------------+---------------------+------------------------------------------------+ | AIE Graph | ``xrt::graph`` | ``#include `` | | | | | | | | ``#include `` | +----------------------+---------------------+------------------------------------------------+ The majority of core data structures are defined in the header files under ``$XILINX_XRT/include/xrt/``. Newer features such as ``xrt::ip``, ``xrt::runlist``, ``xrt::elf``, and related types live under ``$XILINX_XRT/include/xrt/experimental/``. APIs in that experimental area are subject to breaking changes. The common host code flow using the above data structures is as follows: - Open AMD **Device** and load a kernel defined either in **ELF**, **XCLBIN** or combination of both. - Create **Buffer** objects to hold data for kernel inputs and outputs - If required use the Buffer class member functions for the data transfer between host and device (before and after the kernel execution). - Use **Kernel** and **Run** objects to offload and manage the compute-intensive tasks running on the device. - Release the **Buffer** object and close the **Device**. Below we will walk through the common API usage to accomplish the above tasks. Device and Context (NPU Flow) ----------------------------- Device and Context classes provide fundamental infrastructure-related interfaces. The primary objectives of the device- and context-related APIs are: - Open a device and create a context on the device - Load a compiled kernel binary (or an elf) onto the device The simplest code to load an elf is as below: .. code:: c++ :number-lines: 10 unsigned int dev_index = 0; auto device = xrt::device(dev_index); xrt::elf elf{"config.elf"}; auto hwctx = xrt::hw_context(device, elf); The above code block shows: - The ``xrt::device`` class's constructor is used to open the device (enumerated as 0) - The ``xrt::elf`` class's constructor is used to load a compiled binary into host memory from the filesystem ("config.elf") - The ``xrt::hw_context`` class's constructor is used to load the compiled binary on the device The class constructor ``xrt::device::device(const std::string& bdf)`` also supports opening a device object from a PCIe BDF passed as a string. .. code:: c++ :number-lines: 10 auto device = xrt::device("0000:03:00.1"); The ``xrt::device::get_info()`` is a useful member function to obtain necessary information about a device. Some of the information such as Name, BDF can be used to select a specific device to load an XCLBIN .. code:: c++ :number-lines: 10 std::cout << "device name: " << device.get_info() << "\n"; std::cout << "device bdf: " << device.get_info() << "\n"; The class constructor ``xrt::elf(const void *data, size_t size)`` also supports creating an elf object from compiled data already in memory. .. code:: c++ :number-lines: 10 void *myctrlcode = mycompiler_out(); xrt::elf elf(myctrlcode, 0x10000); Device and XCLBIN (Classic FPGA Flow) ------------------------------------- Device and XCLBIN classes provide fundamental infrastructure-related interfaces. The primary objectives of the device- and XCLBIN-related APIs are - Open a device - Load a compiled kernel binary (or XCLBIN) onto the device The simplest code to load an XCLBIN is as below: .. code:: c++ :number-lines: 10 unsigned int dev_index = 0; auto device = xrt::device(dev_index); auto xclbin_uuid = device.load_xclbin("kernel.xclbin"); The above code block shows: - The ``xrt::device`` class's constructor is used to open the device (enumerated as 0) - The member function ``xrt::device::load_xclbin`` is used to load the XCLBIN from the filename. - The member function ``xrt::device::load_xclbin`` returns the XCLBIN UUID, which is required to open the kernel (see the Kernel section). The class constructor ``xrt::device::device(const std::string& bdf)`` also supports opening a device object from a PCIe BDF passed as a string. .. code:: c++ :number-lines: 10 auto device = xrt::device("0000:03:00.1"); The ``xrt::device::get_info()`` is a useful member function to obtain necessary information about a device. Some of the information such as Name, BDF can be used to select a specific device to load an XCLBIN .. code:: c++ :number-lines: 10 std::cout << "device name: " << device.get_info() << "\n"; std::cout << "device bdf: " << device.get_info() << "\n"; Buffers ------- Buffers are primarily used to store the input/output data for use by the device. The buffer-related APIs are discussed in the following three subsections: 1. Buffer allocation and deallocation 2. Data transfer using Buffers 3. Miscellaneous other Buffer APIs 1. Buffer allocation and deallocation ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The C++ interface for buffers is illustrated below. The class constructor ``xrt::bo`` is mainly used to allocate a buffer object 4K aligned. By default, a regular buffer is created (optionally the user can create other types of buffers by providing a flag). .. code:: c++ :number-lines: 15 auto bank_grp_arg0 = kernel.group_id(0); // Memory bank index for kernel argument 0 auto bank_grp_arg1 = kernel.group_id(1); // Memory bank index for kernel argument 1 auto input_buffer = xrt::bo(device, buffer_size_in_bytes,bank_grp_arg0); auto output_buffer = xrt::bo(device, buffer_size_in_bytes, bank_grp_arg1); In the above code ``xrt::bo`` buffer objects are created using the class constructor. Please note the following: - As no special flags are used a regular buffer will be created. Regular buffer is most common type of buffer that has a host backing pointer allocated by user space in heap memory and a device buffer allocated in the specified memory bank. - The second argument specifies the buffer size. - The third argument is used to specify the enumerated memory bank index (to specify the buffer location) where the buffer should be allocated. There are two ways to specify the memory bank index - Through kernel arguments: In the above example, the ``xrt::kernel::group_id()`` member function is used to pass the memory bank index. This member function accepts a kernel argument index and detects the corresponding memory bank index by inspecting the XCLBIN. - Passing a memory bank index: The ``xrt::kernel::group_id()`` overload also accepts the memory bank index directly (as reported by ``xrt-smi examine --report memory``). Creating special Buffers ************************ The ``xrt::bo()`` constructors accept additional buffer flags via an ``enum class`` argument. The main enumerator values are: - ``xrt::bo::flags::normal``: Regular buffer (default) - ``xrt::bo::flags::device_only``: Device only buffer (meant to be used only by the kernel, there is no host backing pointer). - ``xrt::bo::flags::host_only``: Host only buffer (buffer resides in the host memory directly transferred to/from the kernel) - ``xrt::bo::flags::p2p``: P2P buffer, A special type of device-only buffer capable of peer-to-peer transfer - ``xrt::bo::flags::cacheable``: Use a cacheable buffer when the host CPU accesses the buffer frequently (typical on edge platforms). .. note:: Buffer flags are specific to the host and device. Not all the flags are honored on all systems. The below example shows creating a P2P buffer on a device memory bank connected to argument 3 of the kernel. .. code:: c++ :number-lines: 15 auto p2p_buffer = xrt::bo(device, buffer_size_in_bytes, xrt::bo::flags::p2p, kernel.group_id(3)); Creating Buffers from the user pointer ************************************** The ``xrt::bo()`` constructor can also be called using a pointer provided by the user. The user pointer must be aligned to 4K boundary. .. code:: c++ :number-lines: 15 // Host Memory pointer aligned to 4K boundary int *host_ptr; posix_memalign(&host_ptr,4096,MAX_LENGTH*sizeof(int)); // Simple example: fill the allocated host memory for(int i=0; i(); for (auto i=0;i(); auto out_bo_map = out_bo.map(); // Prepare input data std::copy(my_float_array,my_float_array+SIZE,inp_bo_map); in_bo.sync("in_sink", XCL_BO_SYNC_BO_GMIO_TO_AIE, SIZE * sizeof(float),0); out_bo.sync("out_sink", XCL_BO_SYNC_BO_AIE_TO_GMIO, SIZE * sizeof(float), 0); The above code shows - Input and output buffer (``in_bo`` and ``out_bo``) to the graph are created and mapped to the user space - The member function ``xrt::aie::bo::sync`` is used for data transfer using the following arguments - The name of the GMIO ports associated with the DMA transfer - The direction of the buffer transfer - GMIO to Graph: ``XCL_BO_SYNC_BO_GMIO_TO_AIE`` - Graph to GMIO: ``XCL_BO_SYNC_BO_AIE_TO_GMIO`` - The size and the offset of the buffer GMIOs and external buffers --------------------------- XRT provides ``xrt::aie::buffer`` for GMIO and external-buffer endpoints. GMIOs and external buffers move data between global memory (for example DDR) and the AI Engine. They help manage data flow so large workloads can be staged without exhausting local tile memory. Construction of ``xrt::aie::buffer`` succeeds only if a GMIO or external buffer with the given name exists in the loaded design. The class overloads ``xrt::aie::buffer::sync(...)`` to move data between global memory and the AIE. - ``xrt::aie::buffer::sync(xrt::bo bo, ...)`` synchronizes between an ``xrt::aie::buffer`` (GMIO or external buffer) and an ``xrt::bo`` in global memory. - ``xrt::aie::buffer::sync(xrt::bo ping, xrt::bo pong, ...)`` attaches ping/pong ``xrt::bo`` buffers to an external buffer for parallel transfers. The example below uses one input and one output GMIO or external buffer: data moves from the global buffer ``in_bo`` into ``gr.in1``. .. code:: c++ :number-lines: 1 auto device = xrt::aie::device(0); auto uuid = device.load_xclbin("kernel.xclbin"); // Create buffer in DDR / global memory and prepare input auto in_bo = xrt::aie::bo (device, SIZE * sizeof (float), 0, 0); auto inp_bo_map = in_bo.map(); std::copy(my_float_array,my_float_array+SIZE,inp_bo_map); // Create buffer in DDR / global memory for output auto out_bo = xrt::aie::bo (device, SIZE * sizeof (float), 0, 0); auto out_bo_map = out_bo.map(); // GMIO / external buffer for input — sync from in_bo auto in_buffer = xrt::aie::buffer(device, uuid, "gr.in1"); in_buffer.sync(in_bo, XCL_BO_SYNC_BO_GMIO_TO_AIE, SIZE * sizeof(float),0); // Run graphs that use the output GMIO / external buffer // GMIO / external buffer for output — sync to out_bo auto out_buffer = xrt::aie::buffer(device, uuid, "gr.out1"); out_buffer.sync(out_bo, XCL_BO_SYNC_BO_AIE_TO_GMIO, SIZE * sizeof(float),0); The class also overloads ``xrt::aie::buffer::async(...)`` to start an asynchronous transfer involving an ``xrt::bo``. - ``xrt::aie::buffer::async(xrt::bo bo, ...)`` starts an asynchronous sync between an ``xrt::aie::buffer`` and global memory. - ``xrt::aie::buffer::async(xrt::bo ping, xrt::bo pong, ...)`` starts an asynchronous sync using ping/pong ``xrt::bo`` objects. Use ``xrt::aie::buffer::wait()`` to wait for the asynchronous operation to finish. The example below is the same scenario as above, using ``async`` and ``wait`` instead of ``sync`` alone. .. code:: c++ :number-lines: 1 auto device = xrt::aie::device(0); auto uuid = device.load_xclbin("kernel.xclbin"); // Create buffer in DDR / global memory and prepare input auto in_bo = xrt::aie::bo (device, SIZE * sizeof (float), 0, 0); auto inp_bo_map = in_bo.map(); std::copy(my_float_array,my_float_array+SIZE,inp_bo_map); // Create buffer in DDR / global memory for output auto out_bo = xrt::aie::bo (device, SIZE * sizeof (float), 0, 0); auto out_bo_map = out_bo.map(); // GMIO / external buffer for input auto in_buffer = xrt::aie::buffer(device, uuid, "gr.in1"); in_buffer.async(in_bo, XCL_BO_SYNC_BO_GMIO_TO_AIE, SIZE * sizeof(float),0); // Run graphs that use the output GMIO / external buffer // GMIO / external buffer for output auto out_buffer = xrt::aie::buffer(device, uuid, "gr.out1"); out_buffer.async(out_bo, XCL_BO_SYNC_BO_AIE_TO_GMIO, SIZE * sizeof(float),0); out_buffer.wait(); Ping-pong buffers ~~~~~~~~~~~~~~~~~ The example below attaches ping-pong ``xrt::bo`` buffers in global memory to an external buffer (``gr.ext1``) for double-buffered input, then syncs the result to ``out_bo`` via ``gr.out1``. .. code:: c++ :number-lines: 1 auto device = xrt::aie::device(0); auto uuid = device.load_xclbin("kernel.xclbin"); // Host buffer and GMIO / external buffer for primary input auto in_bo = xrt::aie::bo (device, SIZE * sizeof (float), 0, 0); auto in_bo_map = in_bo.map(); std::copy(my_float_array,my_float_array+SIZE,in_bo_map); auto in_buffer = xrt::aie::buffer(device, uuid, "gr.in1"); in_buffer.sync(in_bo, XCL_BO_SYNC_BO_GMIO_TO_AIE, SIZE * sizeof(float),0); // Output buffer in global memory auto out_bo = xrt::aie::bo (device, SIZE * sizeof (float), 0, 0); auto out_bo_map = out_bo.map(); // Ping-pong buffers for an external buffer port auto ext1_bo = xrt::aie::bo (device, SIZE * sizeof (float), 0, 0); auto ext2_bo = xrt::aie::bo (device, SIZE * sizeof (float), 0, 0); auto ping_pong_bo = xrt::aie::buffer(device, uuid, "gr.ext1"); ping_pong_bo.sync(ext1_bo, ext2_bo, XCL_BO_SYNC_BO_GMIO_TO_AIE, SIZE * sizeof(float),0); auto out_buffer = xrt::aie::buffer(device, uuid, "gr.out1"); out_buffer.sync(out_bo, XCL_BO_SYNC_BO_AIE_TO_GMIO, SIZE * sizeof(float),0); XRT Error API ------------- In general, XRT APIs can encounter two types of errors: - **Synchronous errors:** The API may throw an exception that host code can catch and handle. - **Asynchronous errors:** Failures reported later from the driver, system, or hardware. XRT provides ``xrt::error`` and related member functions to surface asynchronous errors to user-space host code, which aids debugging. - ``xrt::error::get_error_code()`` — underlying ``xrtErrorCode`` for the error object (constructed from the device and error class, or from an explicit code and timestamp) - ``xrt::error::get_timestamp()`` — timestamp associated with that error - ``xrt::error::to_string()`` — formatted description string for the error object **Note:** Asynchronous error retrieval is still evolving and currently focuses on AIE-related asynchronous errors. Broader coverage is planned for a future release. Example code .. code:: c++ :number-lines: 41 graph.run(runIteration); try { graph.wait(timeout); } catch (const std::system_error& ex) { if (ex.code().value() == ETIME) { xrt::error error(device, XRT_ERROR_CLASS_AIE); auto errCode = error.get_error_code(); auto timestamp = error.get_timestamp(); auto err_str = error.to_string(); /* code to deal with this specific error */ std::cout << err_str << std::endl; } else { /* Something else */ } } The above code shows - After timeout occurs from ``xrt::graph::wait()`` the member functions ``xrt::error`` class are called to retrieve asynchronous error code and timestamp - Member function ``xrt::error::to_string()`` is called to obtain the error string. Profiling --------- In Versal ACAPs with AI Engines, the XRT Profiling class (``xrt::aie::profiling``) and its member functions can be used to configure AI Engine hardware resources for performance profiling and event tracing. Create Profiling Event ~~~~~~~~~~~~~~~~~~~~~~ The ``xrt::aie::profiling`` constructor creates a profiling object, as shown below. .. code:: c :number-lines: 35 auto event = xrt::aie::profiling(device); Use the profiling object to start and stop counters and to read profiling statistics through the profiling APIs. Start Profiling ~~~~~~~~~~~~~~~ The member function ``xrt::aie::profiling::start()`` is used to start performance counters in AI Engine as per the profiling option passed as an argument. This function configures the performance counters in the AI Engine and starts profiling. .. code:: c :number-lines: 45 auto graph = xrt::graph(device, xclbin_uuid, "graph_name"); std::string port1_name = "..."; // PLIO/GMIO port per UG1079 std::string port2_name = "..."; // PLIO/GMIO port per UG1079 uint32_t value = 0; // meaning depends on profiling_option event.start( xrt::aie::profiling::profiling_option::io_total_stream_running_to_idle_cycles, port1_name, port2_name, value); // run graph ... s2mm_run.wait(); Use the same ``xrt::aie::profiling`` object for ``read()`` and ``stop()`` after ``start()``; see ``xrt/xrt_aie.h`` and UG1079 for option and port semantics. Read Profiling ~~~~~~~~~~~~~~ ``xrt::aie::profiling::read()`` returns the current performance counter value for the profiling session on that object. .. code:: c++ :number-lines: 35 uint64_t cycle_count = event.read(); Stop Profiling ~~~~~~~~~~~~~~ The ``xrt::aie::profiling::stop`` function stops the performance profiling associated with the profiling handle and releases the corresponding hardware resources. .. code:: c :number-lines: 35 event.stop(); double throughput = output_size_in_bytes / (cycle_count *0.8 * 1e-3); // Every AIE cycle is 0.8ns in production board std::cout << "Throughput of the graph: " << throughput << " MB/s" << std::endl;