Ultralytics YOLO27:

使用 NVIDIA DALI 加速 GPU 预处理#

在生产环境中部署 Ultralytics YOLO 模型时,预处理 往往会成为瓶颈。虽然 TensorRT 可以在几毫秒内完成模型推理,但基于 CPU 的预处理(调整大小、填充、归一化)每张图像可能需要 2–10 毫秒,尤其是在高分辨率下。NVIDIA DALI(数据加载库)通过将整个预处理流程移至 GPU 来解决这个问题。

本指南将带你构建能够精确复现 Ultralytics YOLO 预处理的 DALI 流水线,将其与 model.predict() 集成,处理视频流,并通过 Triton Inference Server 实现端到端部署。

本指南适合哪些人?

本指南适用于在生产环境中部署 YOLO 模型,且 CPU 预处理经测量确实成为瓶颈的工程师——通常是使用 TensorRT 在 NVIDIA GPU 上进行部署、高吞吐量视频流水线,或 Triton Inference Server 部署的场景。如果你使用 model.predict() 进行常规推理,并且没有预处理瓶颈,那么默认的 CPU 流水线就能很好地满足需求。

快速摘要
  • 正在构建 DALI 流水线? 使用 fn.resize(mode="not_larger") + fn.crop(out_of_bounds_policy="pad") + fn.crop_mirror_normalize,在 GPU 上复现 YOLO 的 letterbox 预处理。
  • 要与 Ultralytics 集成? 将 DALI 输出作为 torch.Tensor 传给 model.predict()——Ultralytics 会自动跳过图像预处理。
  • 要使用 Triton 部署? 使用 DALI 后端搭配 TensorRT 集成模型,实现零 CPU 预处理。

为什么使用 DALI 进行 YOLO 预处理#

在典型的 YOLO 推理流水线中,预处理步骤由 CPU 执行:

  1. 解码图像(JPEG/PNG)
  2. 调整大小,同时保持宽高比
  3. 填充至目标尺寸(letterbox)
  4. 归一化像素值,从 [0, 255] 到 [0, 1]
  5. 转换数据布局,从 HWC 转为 CHW

使用 DALI 后,所有这些操作都在 GPU 上运行,从而消除 CPU 瓶颈。以下场景尤其适合使用 DALI:

场景DALI 的优势
GPU 推理速度快TensorRT 引擎的推理时间不到 1 毫秒,因此 CPU 预处理会成为主要开销
高分辨率输入1080p 和 4K 视频流需要耗费大量计算资源进行尺寸调整
批量大小大服务器端推理会并行处理大量图像
CPU 核心数有限像 NVIDIA Jetson 这样的边缘设备,或每个 GPU 配备的 CPU 核心数较少的高密度 GPU 服务器

前置条件#

仅支持 Linux

NVIDIA DALI 仅支持 Linux,无法在 Windows 或 macOS 上使用。

安装所需的软件包:

pip install ultralytics
pip install --extra-index-url https://pypi.nvidia.com nvidia-dali-cuda130

要求:

  • NVIDIA GPU(计算能力 5.0+ / Maxwell 或更新架构)
  • CUDA 11.0+、12.0+ 或 13.0+
  • Python 3.10–3.14
  • Linux 操作系统

了解 YOLO 预处理#

在构建 DALI 流水线之前,值得先准确了解 Ultralytics 在预处理期间执行的操作。关键类是 ultralytics/data/augment.py 中的 LetterBox:

from ultralytics.data.augment import LetterBox

letterbox = LetterBox(
    new_shape=(640, 640),  # 目标尺寸
    center=True,  # 将图像居中(两侧填充相同的宽度)
    stride=32,  # 步幅对齐
    padding_value=114,  # 灰色填充(114, 114, 114)
)

ultralytics/engine/predictor.py 中的完整预处理流程会执行以下步骤:

步骤操作CPU 函数DALI 对应项
1Letterbox 调整大小cv2.resizefn.resize(mode="not_larger")
2居中填充cv2.copyMakeBorderfn.crop(out_of_bounds_policy="pad")
3BGR → RGBim[..., ::-1]fn.decoders.image(output_type=types.RGB)
4HWC → CHW + 归一化 /255np.transpose + tensor / 255fn.crop_mirror_normalize(std=[255,255,255])

letterbox 操作通过以下方式保持宽高比:

  1. 计算缩放比例:r = min(target_h / h, target_w / w)
  2. 调整至 (round(w * r), round(h * r))
  3. 用灰色填充剩余空间(114),使图像达到目标尺寸
  4. 将图像居中,让两侧的填充宽度相同

用于 YOLO 的 DALI 流水线#

推荐的 DALI 流水线会复现 Ultralytics 默认的 LetterBox(center=True) 行为,也就是标准 YOLO 推理所使用的行为。

居中流水线(推荐,与 Ultralytics LetterBox 一致)#

此版本会精确复现 Ultralytics 默认的居中填充预处理,与 LetterBox(center=True) 一致:

带居中填充的 DALI 流水线(推荐)
from nvidia import dali
from nvidia.dali import fn, types

@dali.pipeline_def(batch_size=8, num_threads=4, device_id=0)
def yolo_dali_pipeline_centered(image_dir, target_size=640):
    """DALI pipeline replicating YOLO preprocessing with centered padding.

    Matches Ultralytics LetterBox(center=True) behavior exactly.
    """
    # 在 GPU 上读取并解码图像
    jpegs, _ = fn.readers.file(file_root=image_dir, random_shuffle=False, name="Reader")
    images = fn.decoders.image(jpegs, device="mixed", output_type=types.RGB)

    # 保持宽高比进行尺寸调整
    resized = fn.resize(
        images,
        resize_x=target_size,
        resize_y=target_size,
        mode="not_larger",
        interp_type=types.INTERP_LINEAR,
        antialias=False,  # 匹配 cv2.INTER_LINEAR(不进行抗锯齿)
    )

    # 使用 fn.crop 和 out_of_bounds_policy 进行居中填充
    # 当裁剪尺寸大于图像尺寸时,fn.crop 会将图像居中并对称填充
    padded = fn.crop(
        resized,
        crop=(target_size, target_size),
        out_of_bounds_policy="pad",
        fill_values=114,  # YOLO 填充值
    )

    # 归一化并转换数据布局
    output = fn.crop_mirror_normalize(
        padded,
        dtype=types.FLOAT,
        output_layout="CHW",
        mean=[0.0, 0.0, 0.0],
        std=[255.0, 255.0, 255.0],
    )
    return output
什么时候只用 `fn.pad` 就够了?

如果不需要与 LetterBox(center=True) 完全一致,可以使用 fn.pad(...) 而不是 fn.crop(..., out_of_bounds_policy="pad") 来简化填充步骤。此变体只会填充右侧和底部边缘,对于自定义部署流水线可能已经够用,但无法精确匹配 Ultralytics 默认的居中 letterbox 行为。

为什么使用 `fn.crop` 进行居中填充?

DALI 的 fn.pad 运算符只会在右侧和底部边缘添加填充。要实现居中填充(与 Ultralytics LetterBox(center=True) 一致),请将 fn.crop 与 out_of_bounds_policy="pad" 配合使用。在默认的 crop_pos_x=0.5 和 crop_pos_y=0.5 设置下,图像会自动居中并进行对称填充。

抗锯齿不匹配

DALI 的 fn.resize 默认启用抗锯齿(antialias=True),而 OpenCV 的 cv2.resize 搭配 INTER_LINEAR 时不会应用抗锯齿。务必在 DALI 中设置 antialias=False,以匹配 CPU 流水线。若省略此设置,会产生细微的数值差异,进而影响模型精度。

运行流水线#

构建并运行 DALI 流水线
# Build and run the pipeline
pipe = yolo_dali_pipeline_centered(image_dir="/path/to/images", target_size=640)
pipe.build()

# Get a batch of preprocessed images
(output,) = pipe.run()

# Convert to numpy or PyTorch tensors
batch_np = output.as_cpu().as_array()  # Shape: (batch_size, 3, 640, 640)
print(f"Output shape: {batch_np.shape}, dtype: {batch_np.dtype}")
print(f"Value range: [{batch_np.min():.4f}, {batch_np.max():.4f}]")

在 Ultralytics Predict 中使用 DALI#

你可以将预处理后的 PyTorch 张量直接传给 model.predict()。传入 torch.Tensor 时,Ultralytics 会跳过图像预处理(letterbox、BGR→RGB、HWC→CHW 和 /255 归一化),只会在将张量传给模型之前执行设备传输和数据类型转换。

在这种情况下,Ultralytics 无法获取原始图像尺寸,因此检测框坐标会以 640×640 letterbox 空间为基准返回。要将坐标映射回原始图像空间,请使用 scale_boxes,它会处理 LetterBox 所使用的精确舍入逻辑:

from ultralytics.utils.ops import scale_boxes

# 框:形状为 (N, 4) 的张量,采用 xyxy 格式,坐标位于 640x640 letterbox 空间
# 将框从 letterbox 尺寸 (640, 640) 缩放回原始尺寸 (orig_h, orig_w)
boxes = scale_boxes((640, 640), boxes, (orig_h, orig_w))

这适用于所有外部预处理路径——直接输入张量、视频流和 Triton 部署。

DALI + Ultralytics 预测
from nvidia.dali.plugin.pytorch import DALIGenericIterator

from ultralytics import YOLO

# Load model
model = YOLO("yolo26n.pt")

# Create DALI iterator
pipe = yolo_dali_pipeline_centered(image_dir="/path/to/images", target_size=640)
pipe.build()
dali_iter = DALIGenericIterator(pipe, ["images"], reader_name="Reader")

# Run inference with DALI-preprocessed tensors
for batch in dali_iter:
    images = batch[0]["images"]  # Already on GPU, shape (B, 3, 640, 640)
    results = model.predict(images, verbose=False)
    for result in results:
        print(f"Detected {len(result.boxes)} objects")
预处理开销为零

将 torch.Tensor 传给 model.predict() 时,图像预处理耗时约为 ~0.004ms(基本为零);相比之下,CPU 预处理约需 ~1-10ms。张量必须采用 BCHW 格式,数据类型为 float32(或 float16),并归一化至 [0, 1]。Ultralytics 仍会自动处理设备传输和数据类型转换。

使用 DALI 处理视频流#

要实时处理视频,请使用 fn.external_source 从任意来源馈送帧——例如 OpenCV、GStreamer 或自定义采集库:

用于视频流预处理的 DALI 流水线
from nvidia import dali
from nvidia.dali import fn, types

@dali.pipeline_def(batch_size=1, num_threads=4, device_id=0)
def yolo_video_pipeline(target_size=640):
    """DALI pipeline for processing video frames from external source."""
    # 用于馈送来自 OpenCV、GStreamer 等的帧的外部数据源。
    frames = fn.external_source(device="cpu", name="input")
    frames = fn.reshape(frames, layout="HWC")

    # 移至 GPU 并进行预处理
    frames_gpu = frames.gpu()
    resized = fn.resize(
        frames_gpu,
        resize_x=target_size,
        resize_y=target_size,
        mode="not_larger",
        interp_type=types.INTERP_LINEAR,
        antialias=False,
    )
    padded = fn.crop(
        resized,
        crop=(target_size, target_size),
        out_of_bounds_policy="pad",
        fill_values=114,
    )
    output = fn.crop_mirror_normalize(
        padded,
        dtype=types.FLOAT,
        output_layout="CHW",
        mean=[0.0, 0.0, 0.0],
        std=[255.0, 255.0, 255.0],
    )
    return output

使用 DALI 的 Triton Inference Server#

在生产部署中,将 DALI 预处理与 TensorRT 推理结合,并在 Triton Inference Server 中使用集成模型。这样可以完全消除 CPU 预处理——输入原始 JPEG 字节,输出检测结果,所有处理都在 GPU 上完成。

模型仓库结构#

model_repository/
├── dali_preprocessing/
│   ├── 1/
│   │   └── model.dali
│   └── config.pbtxt
├── yolo_trt/
│   ├── 1/
│   │   └── model.plan
│   └── config.pbtxt
└── ensemble_dali_yolo/
    ├── 1/                  # Empty directory (required by Triton)
    └── config.pbtxt

步骤 1:创建 DALI Pipeline#

为 Triton DALI 后端序列化 DALI pipeline:

为 Triton 序列化 DALI pipeline
from nvidia import dali
from nvidia.dali import fn, types

@dali.pipeline_def(batch_size=8, num_threads=4, device_id=0)
def triton_dali_pipeline():
    """DALI preprocessing pipeline for Triton deployment."""
    # 输入:来自 Triton 的原始编码图像字节
    images = fn.external_source(device="cpu", name="DALI_INPUT_0")
    images = fn.decoders.image(images, device="mixed", output_type=types.RGB)

    resized = fn.resize(
        images,
        resize_x=640,
        resize_y=640,
        mode="not_larger",
        interp_type=types.INTERP_LINEAR,
        antialias=False,
    )
    padded = fn.crop(
        resized,
        crop=(640, 640),
        out_of_bounds_policy="pad",
        fill_values=114,
    )
    output = fn.crop_mirror_normalize(
        padded,
        dtype=types.FLOAT,
        output_layout="CHW",
        mean=[0.0, 0.0, 0.0],
        std=[255.0, 255.0, 255.0],
    )
    return output

# 将 pipeline 序列化到模型仓库
pipe = triton_dali_pipeline()
pipe.serialize(filename="model_repository/dali_preprocessing/1/model.dali")

步骤 2:将 YOLO 导出为 TensorRT#

将 YOLO 模型导出为 TensorRT engine
from pathlib import Path

from ultralytics import YOLO

model = YOLO("yolo26n.pt")
engine_path = model.export(
    format="engine", imgsz=640, quantize=16, batch=8, dynamic=True, nms=False
)  # 无 NMS(N, 300, 6);TensorRT >= 8.5

# Ultralytics 会在 .engine 文件前添加元数据头;请移除该头,以便 Triton 加载原始 TensorRT plan
with open(engine_path, "rb") as f:
    meta_len = int.from_bytes(f.read(4), byteorder="little")  # JSON 元数据头的长度
    f.seek(4 + meta_len)
    plan = f.read()
Path("model_repository/yolo_trt/1").mkdir(parents=True, exist_ok=True)
Path("model_repository/yolo_trt/1/model.plan").write_bytes(plan)

步骤 3:配置 Triton#

dali_preprocessing/config.pbtxt:

name: "dali_preprocessing"
backend: "dali"
max_batch_size: 8
input [
  {
    name: "DALI_INPUT_0"
    data_type: TYPE_UINT8
    dims: [ -1 ]
  }
]
output [
  {
    name: "DALI_OUTPUT_0"
    data_type: TYPE_FP32
    dims: [ 3, 640, 640 ]
  }
]

yolo_trt/config.pbtxt:

name: "yolo_trt"
platform: "tensorrt_plan"
max_batch_size: 8
input [
  {
    name: "images"
    data_type: TYPE_FP32
    dims: [ 3, 640, 640 ]
  }
]
output [
  {
    name: "output0"
    data_type: TYPE_FP32
    dims: [ 300, 6 ]
  }
]

ensemble_dali_yolo/config.pbtxt:

name: "ensemble_dali_yolo"
platform: "ensemble"
max_batch_size: 8
input [
  {
    name: "INPUT"
    data_type: TYPE_UINT8
    dims: [ -1 ]
  }
]
output [
  {
    name: "OUTPUT"
    data_type: TYPE_FP32
    dims: [ 300, 6 ]
  }
]
ensemble_scheduling {
  step [
    {
      model_name: "dali_preprocessing"
      model_version: -1
      input_map {
        key: "DALI_INPUT_0"
        value: "INPUT"
      }
      output_map {
        key: "DALI_OUTPUT_0"
        value: "preprocessed_image"
      }
    },
    {
      model_name: "yolo_trt"
      model_version: -1
      input_map {
        key: "images"
        value: "preprocessed_image"
      }
      output_map {
        key: "output0"
        value: "OUTPUT"
      }
    }
  ]
}
集成映射的工作原理

集成模型通过虚拟张量名称连接各个模型。DALI 步骤中的 output_map 值 "preprocessed_image" 与 TensorRT 步骤中的 input_map 值 "preprocessed_image" 相匹配。这些名称是任意的,用于将一个步骤的输出连接到下一步骤的输入——它们无需与任何模型的内部张量名称匹配。

步骤 4:发送推理请求#

为什么使用 `tritonclient` 而不是 `YOLO('http://...')`?

Ultralytics 提供内置 Triton 支持,可自动处理预处理和后处理。不过,它无法与 DALI 集成模型配合使用,因为 YOLO() 会发送预处理后的 float32 张量,而集成模型需要原始 JPEG 字节。DALI 集成模型请直接使用 tritonclient;不使用 DALI 的标准部署则使用内置集成。

向 Triton 集成模型发送图像
import numpy as np
import tritonclient.http as httpclient

client = httpclient.InferenceServerClient(url="localhost:8000")

# Load image as raw bytes (JPEG/PNG encoded)
image_data = np.fromfile("image.jpg", dtype="uint8")
image_data = np.expand_dims(image_data, axis=0)  # Add batch dimension

# Create input
input_tensor = httpclient.InferInput("INPUT", image_data.shape, "UINT8")
input_tensor.set_data_from_numpy(image_data)

# Run inference through the ensemble
result = client.infer(model_name="ensemble_dali_yolo", inputs=[input_tensor])
detections = result.as_numpy("OUTPUT")  # Shape: (1, 300, 6) -> [x1, y1, x2, y2, conf, class_id]

# Filter by confidence (no NMS needed for the nms=False export)
detections = detections[0]  # First image
detections = detections[detections[:, 4] > 0.25]  # Confidence threshold
print(f"Detected {len(detections)} objects")
批量处理 JPEG 图像

向 Triton 发送一批 JPEG 图像时,请将所有编码后的字节数组填充到相同长度(即该批次中的最大字节数)。Triton 要求输入张量的批次形状一致。

支持的任务#

DALI 预处理适用于使用标准 LetterBox pipeline 的所有 YOLO 任务:

任务支持说明
检测✅标准 letterbox 预处理
实例分割✅与目标检测相同的预处理
语义分割✅与目标检测相同的图像预处理
深度估计✅与目标检测相同的图像预处理
分类❌使用 torchvision transforms(中心裁剪),而不是 letterbox
姿态估计✅与目标检测相同的预处理
定向目标检测 (OBB)✅与目标检测相同的预处理

局限性#

  • 仅限 Linux:DALI 不支持 Windows 或 macOS
  • 需要 NVIDIA GPU:不支持仅使用 CPU 的备用方案
  • 静态 pipeline:pipeline 结构在构建时定义,无法动态更改
  • fn.pad 只在右侧/底部填充:如需居中填充,请将 fn.crop 与 out_of_bounds_policy="pad" 配合使用
  • 不支持 rect 模式:DALI pipeline 会生成固定尺寸的输出(例如 640×640)。不支持会生成可变尺寸输出(例如 384×640)的 auto=True rect 模式。请注意,虽然 TensorRT 支持动态输入形状,但固定尺寸的 DALI pipeline 与固定尺寸的 engine 搭配更自然,可实现最大吞吐量
  • 多个实例的内存占用:在 Triton 中将 instance_group 与大于 1 的 count 配合使用,可能导致内存占用过高。DALI 模型请使用默认实例组

常见问题#

  • 具体收益取决于你的 pipeline。当使用 TensorRT 进行 GPU 推理的速度已经很快时,耗时 2–10 毫秒的 CPU 预处理可能会成为主要开销。DALI 通过在 GPU 上运行预处理来消除这一瓶颈。高分辨率输入(1080p、4K)、较大的批次大小,以及每个 GPU 配备的 CPU 核心数有限的系统,通常能获得最大收益。

  • 可以。使用 DALIGenericIterator 获取预处理后的 torch.Tensor 输出,然后将其传递给 model.predict()。不过,使用 TensorRT 模型时,性能收益最大,因为这类模型的推理速度已经很快,CPU 预处理会成为瓶颈。

  • fn.pad 只在右侧和底部边缘添加填充。fn.crop 与 out_of_bounds_policy="pad" 配合使用时,会将图像居中,并在四周对称填充,效果与 Ultralytics 的 LetterBox(center=True) 行为一致。

  • 几乎完全一致。在 fn.resize 中设置 antialias=False,即可匹配 OpenCV 的 cv2.INTER_LINEAR。由于 GPU 与 CPU 的运算方式不同,可能会出现轻微的浮点差异(< 0.001),但对检测精度没有可测量的影响。

  • CV-CUDA 是另一个用于 GPU 加速视觉处理的 NVIDIA 库。它提供逐算子控制(类似在 GPU 上运行的 OpenCV),而不是采用 DALI 的 pipeline 方式。CV-CUDA 的 cvcuda.copymakeborder() 支持显式设置各边的填充,因此可以轻松实现居中 letterbox。基于 pipeline 的工作流(尤其是搭配 Triton)建议选择 DALI;自定义推理代码需要精细的算子级控制时,则建议选择 CV-CUDA。

评论