Ultralytics YOLO27:

AMD Xilinx Deployment for Ultralytics YOLO with Vitis AI#

Native Ultralytics export coming soon

Native Ultralytics export support for AMD Xilinx devices is coming soon. Until then, this guide explains the AMD Xilinx hardware and software landscape and shows how to deploy Ultralytics YOLO26 today with AMD's Vitis AI tools, starting from an ONNX export or a PyTorch checkpoint.

AMD Xilinx devices power many of the world's industrial cameras, automotive vision systems, robots, drones and medical imaging products. They combine Arm processors with programmable logic and, on newer devices, dedicated AI Engines, so a single chip can capture video, preprocess it, run object detection and act on the result with low, predictable inference latency.

This guide covers what each AMD Xilinx device family is, how AI runs on them, which Ultralytics YOLO operators each accelerator supports, and the step-by-step workflow for deploying YOLO models on Zynq UltraScale+, Kria and Versal hardware.

What is AMD Xilinx?#

Xilinx invented the field-programmable gate array (FPGA) in the 1980s and became a leading supplier of adaptive SoCs and FPGAs. AMD completed its acquisition of Xilinx in February 2022, and the product lines are now sold under the AMD brand as AMD Zynq, AMD Kria, AMD Versal and AMD Vitis.

Xilinx or AMD?

Both names refer to the same products. AMD markets them as "adaptive SoCs and FPGAs", but engineers still widely say "Xilinx". Part numbers keep the XC prefix (for example xczu7ev), and the older Vitis AI repository and Docker images still live under the Xilinx name on GitHub and Docker Hub. This guide uses "AMD Xilinx" so you can find it with either name.

Key Terms and Concepts#

AMD Xilinx deployment uses its own vocabulary. The table below explains every term used in this guide.

TermWhat it means
FPGAField-programmable gate array: a chip whose digital logic is configured after manufacturing by loading a design, called a bitstream. It can implement custom hardware such as video pipelines or neural network accelerators.
Programmable logic (PL)The FPGA fabric inside an AMD Xilinx SoC. On Zynq and Kria devices, the AI accelerator is built in the PL.
Processing system (PS)The hard Arm CPU cores, memory controllers and peripherals of the SoC. It runs Linux, your application, and any model layers the accelerator cannot execute.
Adaptive SoC / MPSoCA system-on-chip that combines a processing system with programmable logic, plus AI Engines on many Versal devices. MPSoC stands for multiprocessor system-on-chip.
AI Engine (AIE, AIE-ML, AIE-MLv2)Arrays of hardened vector processors on many Versal devices, including the Versal AI Edge series covered here, designed for machine learning and signal processing.
DPUDeep Learning Processing Unit: AMD's INT8 neural network accelerator, delivered as IP that is built into the PL (for example DPUCZDX8G on Zynq UltraScale+ and Kria). Sizes such as B512 to B4096 give the peak operations per clock cycle.
NPU / NPU IPNeural Processing Unit: AMD's current-generation inference accelerator, which replaces the DPU in recent Vitis AI releases. AMD describes its NPU IP as a soft accelerator that combines AI Engines with programmable logic, so it also needs a matching hardware design. See the NPU glossary entry.
Vitis AIAMD's toolchain for deploying neural networks on AMD Xilinx devices. It covers quantization, compilation, runtimes, examples and Docker environments.
AMD QuarkAMD's current model quantization library, used by the Versal AI Edge Gen 2 flow to turn an FP32 ONNX model into an INT8 model.
Quantization, PTQ and QATConverting FP32 weights and activations to INT8. Post-training quantization (PTQ) uses calibration images. Quantization-aware training (QAT) fine-tunes the model to recover accuracy.
Calibration imagesA small, representative set of images run through the model during PTQ to choose the INT8 scale for each tensor.
BF16 and mixed precisionBFloat16 is a 16-bit floating-point format that keeps FP32's range. Mixed precision runs most of the network in INT8 and sensitive layers in BF16.
XIRXilinx Intermediate Representation: the graph format the DPU compiler produces and the runtime reads.
.xmodelA serialized XIR graph. The quantizer writes a quantized .xmodel, and the DPU compiler turns it into a compiled .xmodel with DPU instructions, quantized weights and any CPU subgraphs. The compiled model requires the matching DPU configuration.
arch.json / DPU fingerprintThe file that describes a specific DPU configuration. The DPU compiler needs it, and an .xmodel compiled for one fingerprint will not run on another.
SnapshotThe compiled model directory produced by the Versal AI Edge (VEK280) NPU flow. It is tied to one NPU IP variant.
.raiThe compiled model file produced by the Versal AI Edge Gen 2 NPU flow.
VART / VART-MLThe Vitis AI Runtime libraries that load compiled models and run them on the board, with C++ and Python APIs.
ONNX Runtime Vitis AI EPThe VitisAIExecutionProvider for ONNX Runtime, which compiles and runs ONNX models on AMD NPUs.
CPU fallback / graph partitioningWhen the accelerator cannot run an operator, the compiler usually splits the model into accelerator and CPU subgraphs, and every split adds a data transfer that can dominate latency. Some operators instead force the entire model onto the CPU or fail compilation.

AMD Xilinx Device Families for Edge AI#

AMD Xilinx devices for edge AI fall into three families. Zynq and Kria use the DPU in programmable logic, while the Versal AI Edge devices covered here use the NPU on their AI Engines.

FamilyWhat it isApplication CPUAI acceleratorExample boards
Zynq UltraScale+ MPSoCArm CPUs and FPGA logic on one chip, in sizes from ZU1 to ZU19Dual- or quad-core Arm Cortex-A53DPU built in the programmable logicZCU104, ZCU102, custom boards
Kria K26 system-on-moduleProduction-ready module built around a Zynq UltraScale+ MPSoCQuad-core Arm Cortex-A53DPU built in the programmable logicKV260 Vision AI Starter Kit, KR260 Robotics Starter Kit
Versal AI Edge SeriesAdaptive SoCs; AIE-ML parts such as VE2302 and VE2802 run the NPUDual-core Arm Cortex-A72NPU on AIE-ML AI Engines and PLVEK280
Versal AI Edge Series Gen 2Next-generation adaptive SoC with AIE-MLv2 AI EnginesUp to eight Arm Cortex-A78AENPU on AIE-MLv2 AI Engines and PLVEK385

Zynq UltraScale+ MPSoC#

Each Zynq UltraScale+ chip pairs an Arm processing system, with dual-core (CG) or quad-core (EG and EV) Cortex-A53 cores and real-time Cortex-R5F cores, with FPGA logic. EV devices add a hardened H.264/H.265 video codec. To run neural networks, designers build a DPU into the logic next to their camera and video pipelines. On small devices the DPU competes for space with the rest of the design.

Kria System-on-Modules#

The Kria K26 module packages a Zynq UltraScale+ MPSoC, memory and power on a production-ready module, so you avoid designing the processor, memory and power subsystem yourself. The module plugs into a carrier card, either a starter-kit board or your own design. It powers the KV260 Vision AI Starter Kit for smart cameras and the KR260 Robotics Starter Kit for robotics. Because the K26 is built on Zynq UltraScale+, it uses the same DPU flow. The Kria portfolio also includes other modules, so check which processor your module uses before choosing a flow.

Versal Adaptive SoCs#

Versal is AMD's adaptive SoC family. Its AI Edge and AI Core series add hardened AI Engines next to the Arm cores and programmable logic, while some other Versal series have no AI Engines. AMD's NPU IP runs on AI Engines and programmable logic together, and Vitis AI targets the AIE-ML parts of the AI Edge series, such as the VE2302 and VE2802. The Versal AI Edge Series (VEK280 evaluation kit) and the Versal AI Edge Series Gen 2 (VEK385 evaluation kit) are AMD's current targets for edge AI and the focus of current Vitis AI releases.

How AI Runs on AMD Xilinx Devices: DPU vs NPU#

Most AMD Xilinx AI deployments follow the same pattern. The accelerator runs the layers it supports, the Arm CPU runs preprocessing, post-processing and any layers the accelerator cannot execute, and a runtime on the board coordinates the two.

graph LR
    A[Camera / video input]:::start --> B[Arm CPU<br>Linux, preprocessing,<br>post-processing]:::proc
    B <--> C[AI accelerator<br>DPU in programmable logic<br>or NPU on AI Engines + PL]:::out
    B --> D[Application<br>alerts, control, display]:::start

    classDef start fill:#4CAF50,color:#fff
    classDef proc fill:#2196F3,color:#fff
    classDef out fill:#9C27B0,color:#fff

AMD has shipped two generations of accelerator, each with its own toolchain and compiled model file. This guide follows Vitis AI 3.5 for the DPU and Vitis AI 6.3 for the NPU; check AMD's current documentation for later releases.

FlowHardwareToolchainQuantizerCompiled artifactBoard runtimeStatus
DPUZynq UltraScale+, KriaVitis AI 3.5 (Docker)vai_q_pytorch.xmodelVARTFrozen compiler, model zoo and DPU IP
NPU (Versal AI Edge)VEK280 and other Versal AI Edge partsVitis AI 6.3 (Docker)Built into the snapshot flowSnapshotVART-MLActive
NPU (Versal AI Edge Gen 2)VEK385 and other Gen 2 partsVitis AI 6.3 (Docker)AMD Quark.raiONNX Runtime Vitis AI EP or VART-MLActive
The DPU flow is frozen

Vitis AI 3.5 is the last release with DPU compiler and model zoo updates. Later releases in the Xilinx/Vitis-AI repository keep the compiler, model zoo and Zynq UltraScale+ DPU IP unchanged while updating the runtime and compatibility with newer AMD tool versions (see the Vitis AI 5.0 release notes), and AMD's current Vitis AI documentation describes the NPU as the replacement for the deprecated DPU architecture. Existing Zynq UltraScale+ and Kria products can keep shipping on the DPU, but its operator support will not grow, so newer model architectures need the adaptations described in YOLO Model Compatibility.

Ryzen AI laptops use a different stack

AMD Ryzen AI processors in PCs also contain an NPU, but they use the separate Ryzen AI Software stack rather than the embedded Vitis AI flows in this guide. For AMD Instinct and Radeon GPUs, see the AMD GPU integration.

Which Vitis AI Flow Do I Need?#

Pick your flow from the device on your board:

graph TD
    A[Start: which AMD device<br>is on your board?]:::start --> B{Device family?}:::decide
    B -->|Zynq UltraScale+ MPSoC<br>or Kria K26| C[DPU flow<br>Vitis AI 3.5]:::proc
    B -->|Versal AI Edge<br>VEK280| D[NPU snapshot flow<br>Vitis AI 6.3]:::proc
    B -->|Versal AI Edge Gen 2<br>VEK385| E[NPU Quark flow<br>Vitis AI 6.3]:::proc
    C --> F[Train YOLO with Hard-Swish<br>then compile to .xmodel]:::out
    D --> G[Run your model on calibration<br>images to capture a snapshot]:::out
    E --> H[Quantize ONNX with Quark<br>then compile to .rai]:::out

    classDef start fill:#4CAF50,color:#fff
    classDef proc fill:#2196F3,color:#fff
    classDef decide fill:#FF9800,color:#fff
    classDef out fill:#9C27B0,color:#fff

YOLO Model Compatibility and Supported Operators#

An accelerator only speeds up the operators it implements in hardware. When a model contains an unsupported operator, the compiler usually sends that part of the network to the Arm CPU, and each round trip between the accelerator and the CPU adds latency. Some operators cannot be partitioned: on the Versal AI Edge Gen 2 NPU, AMD lists operators such as NonZero and NonMaxSuppression that can force the entire model onto the CPU. Operator support is the most important factor in how well a YOLO model performs on AMD Xilinx hardware.

graph LR
    subgraph S1 [Stock YOLO26 on the DPU]
        A1[Conv]:::out --> A2[SiLU<br>CPU]:::error --> A3[Conv]:::out --> A4[SiLU<br>CPU]:::error --> A5[...]:::proc
    end
    subgraph S2 [Hard-Swish YOLO26 on the DPU]
        B1[Backbone<br>Conv + Hard-Swish<br>DPU]:::out --> B2[C2PSA attention<br>CPU]:::error --> B3[Neck<br>DPU]:::out --> B4[C3k2 attention<br>CPU]:::error --> B5[Detect head<br>DPU]:::out --> B6[Sigmoid and<br>post-processing<br>CPU]:::error
    end

    classDef proc fill:#2196F3,color:#fff
    classDef out fill:#9C27B0,color:#fff
    classDef error fill:#F44336,color:#fff

The table shows where each operator in a YOLO26 model runs. A stock YOLO26n ONNX export contains 87 SiLU activations, each exported as a Sigmoid and a Mul, plus 4 MatMul and 2 Softmax operators from its two attention blocks: the C2PSA block at the end of the backbone (layer 10) and the attention-enabled C3k2 block that produces the P5 output (layer 22).

OperatorWhere it appears in YOLODPU (Zynq UltraScale+, Kria)NPU (Versal AI Edge Gen 2)
Convolution + batch normalizationEvery Conv block✅✅
SiLU activationEvery Conv block (default activation)❌ Runs on CPU; replace with Hard-Swish✅
Hard-Swish, ReLU, ReLU6, LeakyReLUOptional activations set in the model YAML✅ Fused into the convolution✅
SigmoidClass scores in the detection head❌ Runs on CPU (usually part of post-processing)✅
MatMul between two activationsAttention blocks (C2PSA; YOLO26 C3k2)❌ Runs on CPU✅
SoftmaxAttention blocks; DFL in YOLOv8 and YOLO11❌ Runs on CPU✅
Reshape, TransposeAttention blocks⚠️ Fused when possible, otherwise CPU✅
Split, SliceC3k2 and C2f blocks⚠️ Converted to slices; check compiler report✅
Resize (nearest upsample)Neck upsampling✅✅
MaxPool, Concat, AddSPPF block and feature fusion✅✅
TopK, GatherElementsYOLO26 NMS-free head (nms=False)❌ Runs on CPU⚠️ CPU partition on the Arm host
NonMaxSuppressionOnly when exported with nms=True❌ Runs on CPU❌ Can force the model onto CPU

Sources: AMD UG1414 supported operators, PyTorch operator support, and the Versal AI Edge Gen 2 supported, CPU partition and unsupported operator lists. Support also depends on your DPU configuration and graph patterns, so always check the compiler's partition report.

YOLO26 head: no DFL and an optional NMS-free output

YOLO26 removes Distribution Focal Loss (DFL), so unlike YOLO11 and YOLOv8 its box outputs need no softmax decoding. It also adds a second attention block compared with YOLO11, so compare target-specific compiler reports and on-device benchmarks before choosing a model. Exports with nms unset keep the one-to-many head and need NMS on the CPU like other YOLO models. Export with nms=False to use YOLO26's NMS-free one-to-one head instead, which replaces NMS with a lightweight top-k selection that runs on the CPU.

Make YOLO26 DPU-Ready with Hard-Swish#

The DPU fuses only ReLU, ReLU6, LeakyReLU, Hard-Swish and Hard-Sigmoid into its convolutions. Hard-Swish is a hardware-friendly approximation of SiLU, which makes it the natural replacement. Ultralytics model YAML files accept an activation key that changes the default activation of Conv blocks (Model YAML Configuration Guide).

Copy yolo26.yaml to yolo26-hswish.yaml and add one line under the parameters:

# Parameters
nc: 80 # number of classes
activation: nn.Hardswish() # default Conv activation, DPU-native
end2end: True # whether to use end-to-end mode

Then build the model, transfer the pretrained YOLO26 weights and fine-tune on your dataset:

Fine-tune a Hard-Swish YOLO26 model
from ultralytics import YOLO

# Build YOLO26n with Hard-Swish activations; the 'n' in the name selects the nano scale
model = YOLO("yolo26n-hswish.yaml").load("yolo26n.pt")  # transfer pretrained weights

# Fine-tune so the network adapts to Hard-Swish
model.train(data="coco8.yaml", epochs=100, imgsz=640)

Activations have no weights, so all of the pretrained weights transfer. The exported ONNX graph then contains 87 HardSwish operators and no SiLU. Replace coco8.yaml with your own dataset, and compare accuracy against the SiLU model with Val mode before deployment.

Alternatives to retraining
  • Swap at quantization time: set "convert_silu_to_hswish": true in the Vitis AI 3.5 PyTorch quantizer's JSON configuration to replace SiLU during quantization. It saves a training run but usually costs more accuracy, which AMD's fast fine-tuning or QAT can partly recover. See the vai_q_pytorch configuration guide.
  • LeakyReLU: the DPU implements LeakyReLU with a fixed negative slope of 26/256 (about 0.1). If you use LeakyReLU, train with activation: nn.LeakyReLU(0.1015625) so the trained and deployed slopes match.

Handling Attention Blocks on the DPU#

YOLO26 applies attention at the lowest resolution (a 20×20 grid at 640 input) in two places: the C2PSA block at layer 10 and the attention-enabled C3k2 block at layer 22. YOLO11 has one C2PSA block. On a DPU their MatMul and Softmax operators run on the CPU, which splits the model into alternating DPU and CPU subgraphs. You have three options:

  1. Accept the CPU blocks. The compiled .xmodel then contains CPU subgraphs, so run it with AMD's Graph Runner, which executes DPU and CPU subgraphs together when a CPU implementation exists for every operator; otherwise you must implement and register the missing operators. At 20×20 the attention computation is small, but each extra transfer between the DPU and the CPU adds latency, so measure it on your board.

  2. Use an attention-free YAML. In your Hard-Swish YAML, replace the C2PSA layer with nn.Identity so the layer indices used by Concat and Detect stay valid, and disable attention in layer 22:

    backbone:
        # ... layers 0-9 unchanged
        - [-1, 1, nn.Identity, []] # 10 C2PSA removed; keeps later layer indices valid
    
    head:
        # ... layers 11-21 unchanged
        - [-1, 1, C3k2, [1024, True, 0.5, False]] # 22 (P5/32-large), attention disabled
        - [[16, 19, 22], 1, Detect, [nc]] # Detect(P3, P4, P5)

    The exported ONNX graph then contains no MatMul or Softmax operators. The attention weights no longer apply (624 of 666 YOLO26n weights transfer), so fine-tune longer and compare accuracy with Val mode.

  3. Use an attention-free model such as YOLOv8, which AMD has used in its own DPU examples.

On Versal AI Edge Gen 2, attention operators are listed as NPU-supported, so these changes are usually unnecessary. Confirm placement in the compiler report, because AMD notes that supported operators can still fall back to the CPU because of configuration or memory constraints.

Model Compatibility at a Glance#

ModelDPU (Zynq UltraScale+, Kria)NPU (Versal AI Edge Gen 2)
YOLO26Train with Hard-Swish; two attention blocks become CPU subgraphs; no DFLExpected to run without changes; validate on your board
YOLO11Train with Hard-Swish; C2PSA and DFL softmax run on CPUExpected to run without changes; validate on your board
YOLOv8Train with Hard-Swish; DFL softmax runs on CPUAMD's YOLOv8m tutorial (Vitis AI 6.3, VEK385, INT8 with BF16 tail): compiler report shows 1,181 operators (99.915%) and 99.994% of GOPs on the NPU, no model changes

Deploy YOLO26 on AMD Xilinx Today#

Until native export is available, deployment follows four steps:

graph LR
    A[1. Train or fine-tune<br>Ultralytics YOLO]:::start --> B{Target?}:::decide
    B -->|Versal NPU| C[2. Export to ONNX<br>model.export]:::proc
    B -->|Zynq or Kria DPU| D[2. Keep the trained<br>PyTorch checkpoint]:::proc
    C --> E[3. Quantize and compile<br>Vitis AI 6.3 Docker]:::proc
    D --> F[3. Quantize and compile<br>Vitis AI 3.5 Docker]:::proc
    E --> G[4. Run on the board<br>VART-ML or ONNX Runtime]:::out
    F --> H[4. Run on the board<br>VART]:::out
    G -.->|accuracy check| A
    H -.->|accuracy check| A

    classDef start fill:#4CAF50,color:#fff
    classDef proc fill:#2196F3,color:#fff
    classDef decide fill:#FF9800,color:#fff
    classDef out fill:#9C27B0,color:#fff

Step 1: Train or Fine-Tune Your Model#

Train on your own data with Train mode or on the Ultralytics Platform. For DPU targets, start from the Hard-Swish YAML. Record a baseline with Val mode so you can measure the accuracy impact of quantization later.

Step 2: Export to ONNX for NPU Targets#

ONNX is the common input to AMD's NPU flows. The DPU flow quantizes the trained PyTorch checkpoint directly in the Vitis AI 3.5 Docker image, so DPU users can skip this step. Export with a fixed batch size of 1 and an opset that AMD supports; AMD's YOLOv8m tutorial for Versal AI Edge Gen 2 uses opset 17.

Export
from ultralytics import YOLO

# Load the model you trained in Step 1
model = YOLO("runs/detect/train/weights/best.pt")

# Export to ONNX with a static shape for the AMD compiler
model.export(format="onnx", opset=17, imgsz=640)  # creates 'best.onnx' next to 'best.pt'

See the ONNX integration and export arguments for all options. With nms unset, run NMS on the CPU after inference; for YOLO26, nms=False selects the NMS-free head instead. Do not embed NMS with nms=True, because AMD lists NonMaxSuppression among the operators that can force the entire model onto the CPU.

Keep AMD's ONNX Runtime build

AMD's Docker images ship their own ONNX Runtime build with the Vitis AI Execution Provider. Ultralytics checks for ONNX Runtime during export and may install the stock package over it. Export on any machine and copy the .onnx file into the container, or set YOLO_AUTOINSTALL=false when you run Ultralytics inside AMD's Docker image.

Step 3: Quantize and Compile with Vitis AI#

The workflows below use AMD's Docker images on an x86-64 Linux host. You do not need the board for this step. Choose the tab for your device:

  1. Start AMD's Vitis AI 6.3 Docker image for Versal AI Edge Gen 2. See system requirements.
  2. Quantize the ONNX model to INT8 with AMD Quark using the VINT8 configuration. AMD's minimum configuration also requires Int32Bias=False, enable_npu_cnn=True, DedicatedQDQPair=True and QuantizeAllOpTypes=True. Quark reads calibration data through a data reader that you write, so apply the same preprocessing as inference: letterbox resizing to the export size, RGB channel order, 0–1 scaling and NCHW layout, on representative images from your dataset.
  3. Exclude the post-processing subgraph from quantization. AMD's YOLOv8m tutorial warns that quantizing it causes missed detections. In that YOLOv8m example, the compiler then runs the tail in BF16 on the NPU; unsupported tail operators, such as YOLO26's top-k selection, still run on the CPU.
  4. Choose the board runtime before you compile. Standard compilation works with ONNX Runtime, which runs NPU-incompatible operators, such as YOLO26's top-k selection, on the CPU itself, and with VART-ML only when every operator runs on the NPU. To run a model that keeps CPU operators under VART-ML, add AMD's CPU partition passes to vitisai_config.json. Those artifacts cannot run through ONNX Runtime.
  5. Compile by creating an ONNX Runtime session with the VitisAIExecutionProvider and a vitisai_config.json that names your target device. Compilation writes a .rai file to the cache directory. See compiling a model.

To skip quantization, compile the FP32 ONNX model directly and the compiler converts it to BF16. Compilation requires an AMD AI Engine compiler license; see AMD's licensing page.

Step 4: Run and Validate on the Board#

Prepare the board first. It must run a hardware design and Linux image that contain the accelerator configuration you compiled for, plus the matching Vitis AI runtime. See AMD's setup guides for Zynq UltraScale+ and Kria DPU targets, Versal AI Edge (VEK280) and Versal AI Edge Gen 2 (VEK385).

Then copy the artifacts your runtime needs:

FlowArtifacts to copy to the boardBoard runtime
DPU (Zynq UltraScale+, Kria)Compiled .xmodelVART; Graph Runner for CPU subgraphs
NPU (Versal AI Edge, VEK280)Snapshot directoryVART-ML
NPU (Versal AI Edge Gen 2), ORTFP32 or quantized ONNX model used for compilation, vitisai_config.json and the compiled cache directoryONNX Runtime with the Vitis AI EP
NPU (Versal AI Edge Gen 2), VART-ML.rai file (with CPU partition passes if any operator runs on the CPU), plus a VART-ML runner configurationVART-ML

For the NPU flows, the exported ONNX graph already decodes boxes and applies the class-score sigmoid, so the host only interprets the output:

  • nms unset: detection models output a (1, 4 + nc, anchors) tensor of xywh boxes and per-class scores. Select the best class per anchor, convert boxes to corners, filter by confidence and run NMS; the Ultralytics non_max_suppression function performs all of these steps.
  • nms=False (YOLO26): the model outputs a (1, max_det, 6) tensor of [x1, y1, x2, y2, score, class] rows, which only needs a confidence threshold.

In both cases, rescale boxes from the letterboxed input back to the original image. Only graphs cut before the decode step need box decoding on the host. On the DPU, VART buffers hold fixed-point INT8 values: query each tensor's shape and fix_point scale, quantize inputs and dequantize outputs before applying the steps above. The Graph Runner returns the full graph's outputs, while a DPU-only runner returns intermediate DPU subgraph outputs that your code must finish computing. The layouts above are the ONNX (CPU view) layouts: VART-ML defaults to hardware tensor views whose shape, data type and memory layout can differ, so configure the runner's input and output tensor types as CPU views or convert the hardware format yourself (see AMD's VART-ML architecture overview).

Compare on-device accuracy with the FP32 baseline from Step 1 using your own validation set and the same performance metrics, such as mAP. AMD's published YOLOv8m results on the VEK385 show the accuracy cost of INT8 deployment:

YOLOv8m configurationHardwaremAP50-95 (COCO)
FP32 ONNXHost CPU49.95
BF16VEK385 NPU50.29
VINT8 quantized, FP32 tailHost CPU48.75
VINT8 with BF16 tailVEK385 NPU48.38

Source: AMD YOLOv8m tutorial for Versal AI Edge Gen 2, which also reports a 10.69 ms average inference time over 100 VART runs at dp_size=1.

Licensing for commercial products

Shipping Ultralytics YOLO inside a commercial AMD Xilinx product requires either compliance with the AGPL-3.0 license or an Ultralytics Enterprise License.

Real-World Applications#

AMD Xilinx devices are common wherever vision AI must run in real time, at low power, close to the sensor:

Summary#

AMD Xilinx devices run YOLO models through two accelerator generations. The DPU on Zynq UltraScale+ and Kria uses the frozen Vitis AI 3.5 flow and produces .xmodel files. It needs DPU-native activations such as Hard-Swish, and it runs attention on the CPU. The NPU on Versal AI Edge and Versal AI Edge Gen 2 uses current Vitis AI releases. The Gen 2 NPU supports SiLU and attention operators and runs AMD's YOLOv8m example on the VEK385 almost entirely on the NPU, while operator coverage on the earlier VEK280 NPU depends on the Vitis AI version and precision.

Native Ultralytics export for AMD Xilinx devices is coming soon. Until then, train with Ultralytics, export to ONNX for Versal NPU targets or keep the PyTorch checkpoint for DPU targets, and compile with Vitis AI as described above. For other deployment targets, see the model deployment options guide, deployment best practices, and accelerator integrations such as Hailo, Rockchip RKNN and Axelera.

FAQ#

  • Yes. AMD completed its acquisition of Xilinx in February 2022, and Xilinx products are now sold as AMD adaptive SoCs and FPGAs: AMD Zynq, AMD Kria, AMD Versal and AMD Vitis. Engineers still widely use the Xilinx name, and part numbers keep the XC prefix.

  • Not yet. Native Ultralytics export for AMD Xilinx devices is coming soon. Today, export to ONNX with model.export(format="onnx") for Versal NPU targets, or quantize the trained PyTorch checkpoint with vai_q_pytorch for Zynq UltraScale+ and Kria DPU targets, then compile with AMD's Vitis AI tools as shown in Deploy YOLO26 on AMD Xilinx Today.

  • The DPU (Deep Learning Processing Unit) is AMD's earlier INT8 accelerator. It is built in the programmable logic of Zynq UltraScale+ and Kria devices, compiled with Vitis AI 3.5, and produces .xmodel files. The NPU is its replacement in current Vitis AI releases. On Versal AI Edge devices it combines hardened AI Engines with programmable logic, supports INT8, BF16 and mixed precision, and supports more operators, including SiLU and, on Versal AI Edge Gen 2, attention.

  • An .xmodel is a serialized XIR graph used by the AMD DPU toolchain. The quantizer writes a quantized .xmodel, and the vai_c_xir compiler turns it into a compiled .xmodel that contains the DPU instruction stream, quantized INT8 weights and any subgraphs that must run on the CPU. The compiled file targets one specific DPU configuration, described by an arch.json fingerprint, and runs on the board through the Vitis AI Runtime (VART), or through the Graph Runner when it contains CPU subgraphs.

  • No. The DPU accelerates only ReLU, ReLU6, LeakyReLU, Hard-Swish and Hard-Sigmoid activations, and it runs Sigmoid, Softmax and MatMul between two activations on the CPU. Train YOLO with activation: nn.Hardswish() in the model YAML to keep convolutions on the DPU, and see Handling Attention Blocks on the DPU for attention options. Versal AI Edge Gen 2 NPUs support SiLU, Softmax and MatMul natively.

  • The KV260 uses a Zynq UltraScale+ MPSoC with a DPU, so follow the DPU flow. Train a Hard-Swish YOLO26 model, quantize it with vai_q_pytorch in the Vitis AI 3.5 Docker image, compile it with vai_c_xir using the KV260's arch.json, and run the resulting .xmodel on the board with VART, or with the Graph Runner if it contains CPU subgraphs.

  • No. Quantization and compilation run in AMD's Vitis AI Docker images on an x86-64 Linux host. You need the board only to run the compiled model and measure on-device latency and accuracy.

  • Any task can run if its operators compile for your accelerator. Unsupported operators usually run on the CPU, but some can force the entire model onto the CPU or fail compilation. Object detection is the most common workload and the one AMD's own examples use. For segmentation, pose and other tasks, check the compiler's partition report to confirm that the heavy layers run on the accelerator.

Contributors

Comments