Skip to content

LLMs on NPU

End-to-end decoder-only LLM inference (prefill + autoregressive decode) mapped to the AMD NPU2 in bf16 via MLIR-AIR. Coverage and performance below are refreshed nightly. Per-model details and source are in programming_examples/llms/.

Nightly LLM Benchmark (NPU2)

End-to-end LLM inference performance on the AMD Ryzen AI 5 PRO 340 (Krackan Point, NPU2) benchmark runner โ€” 2ร—32 GB DDR5-5600 SODIMM โ€” refreshed nightly. TTFT is time to first token (prefill latency); Decode is steady-state generation throughput.

Verify: ๐ŸŸข pass   ๐Ÿ”ด fail   โšช skipped.   โ€” not measured.

Single-point rows โ€” metrics that do not yet have a curve below. โ†“ means that model's number for this metric is in a curve instead.

Model Context TTFT (ms) Decode (tok/s) Measured Verify
qwen25_3b_q4 (HF) 5 4057.0 โ†“ 2026-09-18 ๐ŸŸข
qwen3_4b_q4nx (HF) 24 5097.0 โ†“ 2026-09-18 ๐ŸŸข

Decode throughput vs context (tok/s)

Steady-state decode throughput at increasing KV-cache depth.

The models below reimplement AMD NPU LLM designs originally developed by the FastFlowLM team, using the higher-level abstractions of the MLIR-AIR dialect.

Model 1k 2k 4k 8k 16k 32k 64k 128k Verify
gemma3_4b_q4nx (HF) 20.61 19.20 16.90 13.60 9.79 6.28 3.66 1.99 ๐ŸŸข
gemma4_e2b_q4nx (HF) โ€” โ€” โ€” โ€” โ€” โ€” โ€” โ€” ๐ŸŸข
lfm2_1_2b_q4nx (HF) 67.65 62.39 54.18 42.76 30.17 18.99 10.89 5.88 ๐ŸŸข
llama31_8b_q4nx (HF) 13.20 12.54 11.42 9.69 7.44 5.08 3.11 1.75 ๐ŸŸข
llama32_1b (HF) 10.45 8.16 6.22 4.46 2.60 1.18 0.49 0.23 ๐ŸŸข
llama32_1b_int4 (HF) 13.93 11.50 8.00 5.05 2.85 1.09 0.47 0.18 ๐ŸŸข
llama32_1b_q4nx (HF) 68.61 63.33 54.84 43.14 30.41 19.05 10.92 5.89 ๐ŸŸข
llama32_3b (HF) 4.10 3.14 2.38 1.66 0.82 0.37 โ€” โ€” ๐ŸŸข
llama32_3b_q4nx (HF) 28.12 25.78 22.06 17.15 11.89 7.36 4.18 2.24 ๐ŸŸข
phi4_mini_q4nx (HF) 23.64 21.65 18.65 14.63 10.21 6.37 3.63 1.95 ๐ŸŸข
qwen25_0_5b (HF) 10.79 9.68 8.24 6.21 4.04 1.75 0.76 0.37 ๐ŸŸข
qwen25_1_5b (HF) 6.57 5.84 4.60 3.21 1.83 0.78 0.41 0.21 ๐ŸŸข
qwen25_3b (HF) 3.84 3.49 2.67 2.14 1.11 0.48 0.25 0.10 ๐ŸŸข
qwen25_3b_q4 (HF) 21.79 19.93 16.96 13.12 8.99 5.54 3.14 1.68 ๐ŸŸข
qwen25_7b_q4nx (HF) 13.37 12.80 11.81 10.23 8.07 5.68 3.56 2.04 ๐ŸŸข
qwen3_0_6b (HF) 8.65 5.84 4.04 2.64 1.24 0.57 โ€” โ€” ๐ŸŸข
qwen3_1_7b (HF) 5.60 4.91 3.73 2.58 1.23 0.54 โ€” โ€” ๐ŸŸข
qwen3_4b (HF) 3.15 2.38 1.73 1.15 0.57 0.24 โ€” โ€” ๐ŸŸข
qwen3_4b_q4nx (HF) 20.90 19.16 16.40 12.75 8.83 5.47 3.11 1.67 ๐ŸŸข
qwen3_8b_q4nx (HF) 12.70 12.03 10.86 9.12 6.91 4.65 2.81 1.57 ๐ŸŸข
smollm2_1_7b (HF) 5.97 4.99 3.65 2.19 1.19 โ€” โ€” โ€” ๐ŸŸข

โ€” not swept at this context, or an expected failure.

Prefill latency vs padded prefill length (TTFT, ms)

Time to first token at increasing padded prefill length. Each column is a separate build of the prefill ELFs: the prompt is padded to the built length, so TTFT tracks that length rather than the number of real prompt tokens.

Model 512 1k 2k 4k Verify
gemma3_4b_q4nx (HF) 1035 1942 3727 7553 ๐ŸŸข
gemma4_e2b_q4nx (HF) 1041 1917 4042 โ€” ๐ŸŸข
lfm2_1_2b_q4nx (HF) 393 635 1112 81871 ๐ŸŸข
llama31_8b_q4nx (HF) 1369 2490 5980 13307 ๐ŸŸข
llama32_1b (HF) 303 551 1021 2199 ๐ŸŸข
llama32_1b_int4 (HF) 304 531 1027 2201 ๐ŸŸข
llama32_1b_q4nx (HF) 292 483 907 2162 ๐ŸŸข
llama32_3b (HF) 651 1164 2574 6234 ๐ŸŸข
llama32_3b_q4nx (HF) 646 1144 2552 6124 ๐ŸŸข
phi4_mini_q4nx (HF) 788 1359 2906 7013 ๐ŸŸข
qwen25_0_5b (HF) 337 557 859 โ€” ๐ŸŸข
qwen25_1_5b (HF) 683 1192 2201 4710 ๐ŸŸข
qwen25_3b (HF) โ€” โ€” 4108 8804 ๐ŸŸข
qwen25_7b_q4nx (HF) 2056 3944 8034 16816 ๐ŸŸข
qwen3_0_6b (HF) 400 698 1439 3326 ๐ŸŸข
qwen3_1_7b (HF) โ€” โ€” 2039 4482 ๐ŸŸข
qwen3_4b (HF) 1366 2558 5291 11680 ๐ŸŸข
qwen3_8b_q4nx (HF) 1816 3573 7102 15281 ๐ŸŸข
smollm2_1_7b (HF) 498 880 1713 3687 ๐ŸŸข

โ€” not swept at this length, or an expected failure.

Vision-Language-Action policies

Not autoregressive: these emit one action chunk per observation rather than a token stream, so there is no decode throughput and no context to sweep. Action chunk is the latency to the model's first usable output. Prefix is the token count the backbone attends over.

Model Prefix Action chunk (ms) Measured Verify
smolvla 241 682.0 2026-09-18 ๐ŸŸข

๐Ÿ“ˆ Performance history over time โ€” per-nightly TTFT and decode throughput plotted per model.

Last updated 2026-09-18 ยท mlir-air f4ffd43 ยท mlir-aie 10767b5 ยท llvm-aie 22.0.0.2026090201+a36c62b9