LLMs on NPU¶
End-to-end decoder-only LLM inference (prefill + autoregressive decode) mapped to the AMD NPU2 in bf16 via MLIR-AIR. Coverage and performance below are refreshed nightly. Per-model details and source are in programming_examples/llms/.
Nightly LLM Benchmark (NPU2)¶
End-to-end LLM inference performance on the AMD Ryzen AI 5 PRO 340 (Krackan Point, NPU2) benchmark runner โ 2ร32 GB DDR5-5600 SODIMM โ refreshed nightly. TTFT is time to first token (prefill latency); Decode is steady-state generation throughput.
Verify: ๐ข pass ๐ด fail โช skipped. โ not measured.
Single-point rows โ metrics that do not yet have a curve below. โ means that model's number for this metric is in a curve instead.
| Model | Context | TTFT (ms) | Decode (tok/s) | Measured | Verify |
|---|---|---|---|---|---|
| qwen25_3b_q4 (HF) | 5 | 4057.0 | โ | 2026-09-18 | ๐ข |
| qwen3_4b_q4nx (HF) | 24 | 5097.0 | โ | 2026-09-18 | ๐ข |
Decode throughput vs context (tok/s)¶
Steady-state decode throughput at increasing KV-cache depth.
The models below reimplement AMD NPU LLM designs originally developed by the FastFlowLM team, using the higher-level abstractions of the MLIR-AIR dialect.
| Model | 1k | 2k | 4k | 8k | 16k | 32k | 64k | 128k | Verify |
|---|---|---|---|---|---|---|---|---|---|
| gemma3_4b_q4nx (HF) | 20.61 | 19.20 | 16.90 | 13.60 | 9.79 | 6.28 | 3.66 | 1.99 | ๐ข |
| gemma4_e2b_q4nx (HF) | โ | โ | โ | โ | โ | โ | โ | โ | ๐ข |
| lfm2_1_2b_q4nx (HF) | 67.65 | 62.39 | 54.18 | 42.76 | 30.17 | 18.99 | 10.89 | 5.88 | ๐ข |
| llama31_8b_q4nx (HF) | 13.20 | 12.54 | 11.42 | 9.69 | 7.44 | 5.08 | 3.11 | 1.75 | ๐ข |
| llama32_1b (HF) | 10.45 | 8.16 | 6.22 | 4.46 | 2.60 | 1.18 | 0.49 | 0.23 | ๐ข |
| llama32_1b_int4 (HF) | 13.93 | 11.50 | 8.00 | 5.05 | 2.85 | 1.09 | 0.47 | 0.18 | ๐ข |
| llama32_1b_q4nx (HF) | 68.61 | 63.33 | 54.84 | 43.14 | 30.41 | 19.05 | 10.92 | 5.89 | ๐ข |
| llama32_3b (HF) | 4.10 | 3.14 | 2.38 | 1.66 | 0.82 | 0.37 | โ | โ | ๐ข |
| llama32_3b_q4nx (HF) | 28.12 | 25.78 | 22.06 | 17.15 | 11.89 | 7.36 | 4.18 | 2.24 | ๐ข |
| phi4_mini_q4nx (HF) | 23.64 | 21.65 | 18.65 | 14.63 | 10.21 | 6.37 | 3.63 | 1.95 | ๐ข |
| qwen25_0_5b (HF) | 10.79 | 9.68 | 8.24 | 6.21 | 4.04 | 1.75 | 0.76 | 0.37 | ๐ข |
| qwen25_1_5b (HF) | 6.57 | 5.84 | 4.60 | 3.21 | 1.83 | 0.78 | 0.41 | 0.21 | ๐ข |
| qwen25_3b (HF) | 3.84 | 3.49 | 2.67 | 2.14 | 1.11 | 0.48 | 0.25 | 0.10 | ๐ข |
| qwen25_3b_q4 (HF) | 21.79 | 19.93 | 16.96 | 13.12 | 8.99 | 5.54 | 3.14 | 1.68 | ๐ข |
| qwen25_7b_q4nx (HF) | 13.37 | 12.80 | 11.81 | 10.23 | 8.07 | 5.68 | 3.56 | 2.04 | ๐ข |
| qwen3_0_6b (HF) | 8.65 | 5.84 | 4.04 | 2.64 | 1.24 | 0.57 | โ | โ | ๐ข |
| qwen3_1_7b (HF) | 5.60 | 4.91 | 3.73 | 2.58 | 1.23 | 0.54 | โ | โ | ๐ข |
| qwen3_4b (HF) | 3.15 | 2.38 | 1.73 | 1.15 | 0.57 | 0.24 | โ | โ | ๐ข |
| qwen3_4b_q4nx (HF) | 20.90 | 19.16 | 16.40 | 12.75 | 8.83 | 5.47 | 3.11 | 1.67 | ๐ข |
| qwen3_8b_q4nx (HF) | 12.70 | 12.03 | 10.86 | 9.12 | 6.91 | 4.65 | 2.81 | 1.57 | ๐ข |
| smollm2_1_7b (HF) | 5.97 | 4.99 | 3.65 | 2.19 | 1.19 | โ | โ | โ | ๐ข |
โ not swept at this context, or an expected failure.
Prefill latency vs padded prefill length (TTFT, ms)¶
Time to first token at increasing padded prefill length. Each column is a separate build of the prefill ELFs: the prompt is padded to the built length, so TTFT tracks that length rather than the number of real prompt tokens.
| Model | 512 | 1k | 2k | 4k | Verify |
|---|---|---|---|---|---|
| gemma3_4b_q4nx (HF) | 1035 | 1942 | 3727 | 7553 | ๐ข |
| gemma4_e2b_q4nx (HF) | 1041 | 1917 | 4042 | โ | ๐ข |
| lfm2_1_2b_q4nx (HF) | 393 | 635 | 1112 | 81871 | ๐ข |
| llama31_8b_q4nx (HF) | 1369 | 2490 | 5980 | 13307 | ๐ข |
| llama32_1b (HF) | 303 | 551 | 1021 | 2199 | ๐ข |
| llama32_1b_int4 (HF) | 304 | 531 | 1027 | 2201 | ๐ข |
| llama32_1b_q4nx (HF) | 292 | 483 | 907 | 2162 | ๐ข |
| llama32_3b (HF) | 651 | 1164 | 2574 | 6234 | ๐ข |
| llama32_3b_q4nx (HF) | 646 | 1144 | 2552 | 6124 | ๐ข |
| phi4_mini_q4nx (HF) | 788 | 1359 | 2906 | 7013 | ๐ข |
| qwen25_0_5b (HF) | 337 | 557 | 859 | โ | ๐ข |
| qwen25_1_5b (HF) | 683 | 1192 | 2201 | 4710 | ๐ข |
| qwen25_3b (HF) | โ | โ | 4108 | 8804 | ๐ข |
| qwen25_7b_q4nx (HF) | 2056 | 3944 | 8034 | 16816 | ๐ข |
| qwen3_0_6b (HF) | 400 | 698 | 1439 | 3326 | ๐ข |
| qwen3_1_7b (HF) | โ | โ | 2039 | 4482 | ๐ข |
| qwen3_4b (HF) | 1366 | 2558 | 5291 | 11680 | ๐ข |
| qwen3_8b_q4nx (HF) | 1816 | 3573 | 7102 | 15281 | ๐ข |
| smollm2_1_7b (HF) | 498 | 880 | 1713 | 3687 | ๐ข |
โ not swept at this length, or an expected failure.
Vision-Language-Action policies¶
Not autoregressive: these emit one action chunk per observation rather than a token stream, so there is no decode throughput and no context to sweep. Action chunk is the latency to the model's first usable output. Prefix is the token count the backbone attends over.
| Model | Prefix | Action chunk (ms) | Measured | Verify |
|---|---|---|---|---|
| smolvla | 241 | 682.0 | 2026-09-18 | ๐ข |
๐ Performance history over time โ per-nightly TTFT and decode throughput plotted per model.
Last updated 2026-09-18 ยท mlir-air f4ffd43 ยท mlir-aie 10767b5 ยท llvm-aie 22.0.0.2026090201+a36c62b9