使用 NVIDIA DALI 加速 GPU 预处理#
在生产环境中部署 Ultralytics YOLO 模型时,预处理 往往会成为瓶颈。虽然 TensorRT 可以在几毫秒内完成模型推理,但基于 CPU 的预处理(调整大小、填充、归一化)每张图像可能需要 2–10 毫秒,尤其是在高分辨率下。NVIDIA DALI(数据加载库)通过将整个预处理流程移至 GPU 来解决这个问题。
本指南将带你构建能够精确复现 Ultralytics YOLO 预处理的 DALI 流水线,将其与 model.predict() 集成,处理视频流,并通过 Triton Inference Server 实现端到端部署。
本指南适用于在生产环境中部署 YOLO 模型,且 CPU 预处理经测量确实成为瓶颈的工程师——通常是使用 TensorRT 在 NVIDIA GPU 上进行部署、高吞吐量视频流水线,或 Triton Inference Server 部署的场景。如果你使用 model.predict() 进行常规推理,并且没有预处理瓶颈,那么默认的 CPU 流水线就能很好地满足需求。
- 正在构建 DALI 流水线? 使用
fn.resize(mode="not_larger")+fn.crop(out_of_bounds_policy="pad")+fn.crop_mirror_normalize,在 GPU 上复现 YOLO 的 letterbox 预处理。 - 要与 Ultralytics 集成? 将 DALI 输出作为
torch.Tensor传给model.predict()——Ultralytics 会自动跳过图像预处理。 - 要使用 Triton 部署? 使用 DALI 后端搭配 TensorRT 集成模型,实现零 CPU 预处理。
为什么使用 DALI 进行 YOLO 预处理#
在典型的 YOLO 推理流水线中,预处理步骤由 CPU 执行:
- 解码图像(JPEG/PNG)
- 调整大小,同时保持宽高比
- 填充至目标尺寸(letterbox)
- 归一化像素值,从
[0, 255]到[0, 1] - 转换数据布局,从 HWC 转为 CHW
使用 DALI 后,所有这些操作都在 GPU 上运行,从而消除 CPU 瓶颈。以下场景尤其适合使用 DALI:
| 场景 | DALI 的优势 |
|---|---|
| GPU 推理速度快 | TensorRT 引擎的推理时间不到 1 毫秒,因此 CPU 预处理会成为主要开销 |
| 高分辨率输入 | 1080p 和 4K 视频流需要耗费大量计算资源进行尺寸调整 |
| 批量大小大 | 服务器端推理会并行处理大量图像 |
| CPU 核心数有限 | 像 NVIDIA Jetson 这样的边缘设备,或每个 GPU 配备的 CPU 核心数较少的高密度 GPU 服务器 |
前置条件#
NVIDIA DALI 仅支持 Linux,无法在 Windows 或 macOS 上使用。
安装所需的软件包:
pip install ultralytics
pip install --extra-index-url https://pypi.nvidia.com nvidia-dali-cuda130要求:
- NVIDIA GPU(计算能力 5.0+ / Maxwell 或更新架构)
- CUDA 11.0+、12.0+ 或 13.0+
- Python 3.10–3.14
- Linux 操作系统
了解 YOLO 预处理#
在构建 DALI 流水线之前,值得先准确了解 Ultralytics 在预处理期间执行的操作。关键类是 ultralytics/data/augment.py 中的 LetterBox:
from ultralytics.data.augment import LetterBox
letterbox = LetterBox(
new_shape=(640, 640), # 目标尺寸
center=True, # 将图像居中(两侧填充相同的宽度)
stride=32, # 步幅对齐
padding_value=114, # 灰色填充(114, 114, 114)
)ultralytics/engine/predictor.py 中的完整预处理流程会执行以下步骤:
| 步骤 | 操作 | CPU 函数 | DALI 对应项 |
|---|---|---|---|
| 1 | Letterbox 调整大小 | cv2.resize | fn.resize(mode="not_larger") |
| 2 | 居中填充 | cv2.copyMakeBorder | fn.crop(out_of_bounds_policy="pad") |
| 3 | BGR → RGB | im[..., ::-1] | fn.decoders.image(output_type=types.RGB) |
| 4 | HWC → CHW + 归一化 /255 | np.transpose + tensor / 255 | fn.crop_mirror_normalize(std=[255,255,255]) |
letterbox 操作通过以下方式保持宽高比:
- 计算缩放比例:
r = min(target_h / h, target_w / w) - 调整至
(round(w * r), round(h * r)) - 用灰色填充剩余空间(
114),使图像达到目标尺寸 - 将图像居中,让两侧的填充宽度相同
用于 YOLO 的 DALI 流水线#
推荐的 DALI 流水线会复现 Ultralytics 默认的 LetterBox(center=True) 行为,也就是标准 YOLO 推理所使用的行为。
居中流水线(推荐,与 Ultralytics LetterBox 一致)#
此版本会精确复现 Ultralytics 默认的居中填充预处理,与 LetterBox(center=True) 一致:
from nvidia import dali
from nvidia.dali import fn, types
@dali.pipeline_def(batch_size=8, num_threads=4, device_id=0)
def yolo_dali_pipeline_centered(image_dir, target_size=640):
"""DALI pipeline replicating YOLO preprocessing with centered padding.
Matches Ultralytics LetterBox(center=True) behavior exactly.
"""
# 在 GPU 上读取并解码图像
jpegs, _ = fn.readers.file(file_root=image_dir, random_shuffle=False, name="Reader")
images = fn.decoders.image(jpegs, device="mixed", output_type=types.RGB)
# 保持宽高比进行尺寸调整
resized = fn.resize(
images,
resize_x=target_size,
resize_y=target_size,
mode="not_larger",
interp_type=types.INTERP_LINEAR,
antialias=False, # 匹配 cv2.INTER_LINEAR(不进行抗锯齿)
)
# 使用 fn.crop 和 out_of_bounds_policy 进行居中填充
# 当裁剪尺寸大于图像尺寸时,fn.crop 会将图像居中并对称填充
padded = fn.crop(
resized,
crop=(target_size, target_size),
out_of_bounds_policy="pad",
fill_values=114, # YOLO 填充值
)
# 归一化并转换数据布局
output = fn.crop_mirror_normalize(
padded,
dtype=types.FLOAT,
output_layout="CHW",
mean=[0.0, 0.0, 0.0],
std=[255.0, 255.0, 255.0],
)
return output如果不需要与 LetterBox(center=True) 完全一致,可以使用 fn.pad(...) 而不是 fn.crop(..., out_of_bounds_policy="pad") 来简化填充步骤。此变体只会填充右侧和底部边缘,对于自定义部署流水线可能已经够用,但无法精确匹配 Ultralytics 默认的居中 letterbox 行为。
DALI 的 fn.pad 运算符只会在右侧和底部边缘添加填充。要实现居中填充(与 Ultralytics LetterBox(center=True) 一致),请将 fn.crop 与 out_of_bounds_policy="pad" 配合使用。在默认的 crop_pos_x=0.5 和 crop_pos_y=0.5 设置下,图像会自动居中并进行对称填充。
DALI 的 fn.resize 默认启用抗锯齿(antialias=True),而 OpenCV 的 cv2.resize 搭配 INTER_LINEAR 时不会应用抗锯齿。务必在 DALI 中设置 antialias=False,以匹配 CPU 流水线。若省略此设置,会产生细微的数值差异,进而影响模型精度。
运行流水线#
# Build and run the pipeline
pipe = yolo_dali_pipeline_centered(image_dir="/path/to/images", target_size=640)
pipe.build()
# Get a batch of preprocessed images
(output,) = pipe.run()
# Convert to numpy or PyTorch tensors
batch_np = output.as_cpu().as_array() # Shape: (batch_size, 3, 640, 640)
print(f"Output shape: {batch_np.shape}, dtype: {batch_np.dtype}")
print(f"Value range: [{batch_np.min():.4f}, {batch_np.max():.4f}]")在 Ultralytics Predict 中使用 DALI#
你可以将预处理后的 PyTorch 张量直接传给 model.predict()。传入 torch.Tensor 时,Ultralytics 会跳过图像预处理(letterbox、BGR→RGB、HWC→CHW 和 /255 归一化),只会在将张量传给模型之前执行设备传输和数据类型转换。
在这种情况下,Ultralytics 无法获取原始图像尺寸,因此检测框坐标会以 640×640 letterbox 空间为基准返回。要将坐标映射回原始图像空间,请使用 scale_boxes,它会处理 LetterBox 所使用的精确舍入逻辑:
from ultralytics.utils.ops import scale_boxes
# 框:形状为 (N, 4) 的张量,采用 xyxy 格式,坐标位于 640x640 letterbox 空间
# 将框从 letterbox 尺寸 (640, 640) 缩放回原始尺寸 (orig_h, orig_w)
boxes = scale_boxes((640, 640), boxes, (orig_h, orig_w))这适用于所有外部预处理路径——直接输入张量、视频流和 Triton 部署。
from nvidia.dali.plugin.pytorch import DALIGenericIterator
from ultralytics import YOLO
# Load model
model = YOLO("yolo26n.pt")
# Create DALI iterator
pipe = yolo_dali_pipeline_centered(image_dir="/path/to/images", target_size=640)
pipe.build()
dali_iter = DALIGenericIterator(pipe, ["images"], reader_name="Reader")
# Run inference with DALI-preprocessed tensors
for batch in dali_iter:
images = batch[0]["images"] # Already on GPU, shape (B, 3, 640, 640)
results = model.predict(images, verbose=False)
for result in results:
print(f"Detected {len(result.boxes)} objects")将 torch.Tensor 传给 model.predict() 时,图像预处理耗时约为 ~0.004ms(基本为零);相比之下,CPU 预处理约需 ~1-10ms。张量必须采用 BCHW 格式,数据类型为 float32(或 float16),并归一化至 [0, 1]。Ultralytics 仍会自动处理设备传输和数据类型转换。
使用 DALI 处理视频流#
要实时处理视频,请使用 fn.external_source 从任意来源馈送帧——例如 OpenCV、GStreamer 或自定义采集库:
from nvidia import dali
from nvidia.dali import fn, types
@dali.pipeline_def(batch_size=1, num_threads=4, device_id=0)
def yolo_video_pipeline(target_size=640):
"""DALI pipeline for processing video frames from external source."""
# 用于馈送来自 OpenCV、GStreamer 等的帧的外部数据源。
frames = fn.external_source(device="cpu", name="input")
frames = fn.reshape(frames, layout="HWC")
# 移至 GPU 并进行预处理
frames_gpu = frames.gpu()
resized = fn.resize(
frames_gpu,
resize_x=target_size,
resize_y=target_size,
mode="not_larger",
interp_type=types.INTERP_LINEAR,
antialias=False,
)
padded = fn.crop(
resized,
crop=(target_size, target_size),
out_of_bounds_policy="pad",
fill_values=114,
)
output = fn.crop_mirror_normalize(
padded,
dtype=types.FLOAT,
output_layout="CHW",
mean=[0.0, 0.0, 0.0],
std=[255.0, 255.0, 255.0],
)
return output使用 DALI 的 Triton Inference Server#
在生产部署中,将 DALI 预处理与 TensorRT 推理结合,并在 Triton Inference Server 中使用集成模型。这样可以完全消除 CPU 预处理——输入原始 JPEG 字节,输出检测结果,所有处理都在 GPU 上完成。
模型仓库结构#
model_repository/
├── dali_preprocessing/
│ ├── 1/
│ │ └── model.dali
│ └── config.pbtxt
├── yolo_trt/
│ ├── 1/
│ │ └── model.plan
│ └── config.pbtxt
└── ensemble_dali_yolo/
├── 1/ # Empty directory (required by Triton)
└── config.pbtxt步骤 1:创建 DALI Pipeline#
为 Triton DALI 后端序列化 DALI pipeline:
from nvidia import dali
from nvidia.dali import fn, types
@dali.pipeline_def(batch_size=8, num_threads=4, device_id=0)
def triton_dali_pipeline():
"""DALI preprocessing pipeline for Triton deployment."""
# 输入:来自 Triton 的原始编码图像字节
images = fn.external_source(device="cpu", name="DALI_INPUT_0")
images = fn.decoders.image(images, device="mixed", output_type=types.RGB)
resized = fn.resize(
images,
resize_x=640,
resize_y=640,
mode="not_larger",
interp_type=types.INTERP_LINEAR,
antialias=False,
)
padded = fn.crop(
resized,
crop=(640, 640),
out_of_bounds_policy="pad",
fill_values=114,
)
output = fn.crop_mirror_normalize(
padded,
dtype=types.FLOAT,
output_layout="CHW",
mean=[0.0, 0.0, 0.0],
std=[255.0, 255.0, 255.0],
)
return output
# 将 pipeline 序列化到模型仓库
pipe = triton_dali_pipeline()
pipe.serialize(filename="model_repository/dali_preprocessing/1/model.dali")步骤 2:将 YOLO 导出为 TensorRT#
from pathlib import Path
from ultralytics import YOLO
model = YOLO("yolo26n.pt")
engine_path = model.export(
format="engine", imgsz=640, quantize=16, batch=8, dynamic=True, nms=False
) # 无 NMS(N, 300, 6);TensorRT >= 8.5
# Ultralytics 会在 .engine 文件前添加元数据头;请移除该头,以便 Triton 加载原始 TensorRT plan
with open(engine_path, "rb") as f:
meta_len = int.from_bytes(f.read(4), byteorder="little") # JSON 元数据头的长度
f.seek(4 + meta_len)
plan = f.read()
Path("model_repository/yolo_trt/1").mkdir(parents=True, exist_ok=True)
Path("model_repository/yolo_trt/1/model.plan").write_bytes(plan)步骤 3:配置 Triton#
dali_preprocessing/config.pbtxt:
name: "dali_preprocessing"
backend: "dali"
max_batch_size: 8
input [
{
name: "DALI_INPUT_0"
data_type: TYPE_UINT8
dims: [ -1 ]
}
]
output [
{
name: "DALI_OUTPUT_0"
data_type: TYPE_FP32
dims: [ 3, 640, 640 ]
}
]yolo_trt/config.pbtxt:
name: "yolo_trt"
platform: "tensorrt_plan"
max_batch_size: 8
input [
{
name: "images"
data_type: TYPE_FP32
dims: [ 3, 640, 640 ]
}
]
output [
{
name: "output0"
data_type: TYPE_FP32
dims: [ 300, 6 ]
}
]ensemble_dali_yolo/config.pbtxt:
name: "ensemble_dali_yolo"
platform: "ensemble"
max_batch_size: 8
input [
{
name: "INPUT"
data_type: TYPE_UINT8
dims: [ -1 ]
}
]
output [
{
name: "OUTPUT"
data_type: TYPE_FP32
dims: [ 300, 6 ]
}
]
ensemble_scheduling {
step [
{
model_name: "dali_preprocessing"
model_version: -1
input_map {
key: "DALI_INPUT_0"
value: "INPUT"
}
output_map {
key: "DALI_OUTPUT_0"
value: "preprocessed_image"
}
},
{
model_name: "yolo_trt"
model_version: -1
input_map {
key: "images"
value: "preprocessed_image"
}
output_map {
key: "output0"
value: "OUTPUT"
}
}
]
}集成模型通过虚拟张量名称连接各个模型。DALI 步骤中的 output_map 值 "preprocessed_image" 与 TensorRT 步骤中的 input_map 值 "preprocessed_image" 相匹配。这些名称是任意的,用于将一个步骤的输出连接到下一步骤的输入——它们无需与任何模型的内部张量名称匹配。
步骤 4:发送推理请求#
Ultralytics 提供内置 Triton 支持,可自动处理预处理和后处理。不过,它无法与 DALI 集成模型配合使用,因为 YOLO() 会发送预处理后的 float32 张量,而集成模型需要原始 JPEG 字节。DALI 集成模型请直接使用 tritonclient;不使用 DALI 的标准部署则使用内置集成。
import numpy as np
import tritonclient.http as httpclient
client = httpclient.InferenceServerClient(url="localhost:8000")
# Load image as raw bytes (JPEG/PNG encoded)
image_data = np.fromfile("image.jpg", dtype="uint8")
image_data = np.expand_dims(image_data, axis=0) # Add batch dimension
# Create input
input_tensor = httpclient.InferInput("INPUT", image_data.shape, "UINT8")
input_tensor.set_data_from_numpy(image_data)
# Run inference through the ensemble
result = client.infer(model_name="ensemble_dali_yolo", inputs=[input_tensor])
detections = result.as_numpy("OUTPUT") # Shape: (1, 300, 6) -> [x1, y1, x2, y2, conf, class_id]
# Filter by confidence (no NMS needed for the nms=False export)
detections = detections[0] # First image
detections = detections[detections[:, 4] > 0.25] # Confidence threshold
print(f"Detected {len(detections)} objects")向 Triton 发送一批 JPEG 图像时,请将所有编码后的字节数组填充到相同长度(即该批次中的最大字节数)。Triton 要求输入张量的批次形状一致。
支持的任务#
DALI 预处理适用于使用标准 LetterBox pipeline 的所有 YOLO 任务:
| 任务 | 支持 | 说明 |
|---|---|---|
| 检测 | ✅ | 标准 letterbox 预处理 |
| 实例分割 | ✅ | 与目标检测相同的预处理 |
| 语义分割 | ✅ | 与目标检测相同的图像预处理 |
| 深度估计 | ✅ | 与目标检测相同的图像预处理 |
| 分类 | ❌ | 使用 torchvision transforms(中心裁剪),而不是 letterbox |
| 姿态估计 | ✅ | 与目标检测相同的预处理 |
| 定向目标检测 (OBB) | ✅ | 与目标检测相同的预处理 |
局限性#
- 仅限 Linux:DALI 不支持 Windows 或 macOS
- 需要 NVIDIA GPU:不支持仅使用 CPU 的备用方案
- 静态 pipeline:pipeline 结构在构建时定义,无法动态更改
fn.pad只在右侧/底部填充:如需居中填充,请将fn.crop与out_of_bounds_policy="pad"配合使用- 不支持 rect 模式:DALI pipeline 会生成固定尺寸的输出(例如 640×640)。不支持会生成可变尺寸输出(例如 384×640)的
auto=Truerect 模式。请注意,虽然 TensorRT 支持动态输入形状,但固定尺寸的 DALI pipeline 与固定尺寸的 engine 搭配更自然,可实现最大吞吐量 - 多个实例的内存占用:在 Triton 中将
instance_group与大于 1 的count配合使用,可能导致内存占用过高。DALI 模型请使用默认实例组
常见问题#
可以。使用
DALIGenericIterator获取预处理后的torch.Tensor输出,然后将其传递给model.predict()。不过,使用 TensorRT 模型时,性能收益最大,因为这类模型的推理速度已经很快,CPU 预处理会成为瓶颈。fn.pad只在右侧和底部边缘添加填充。fn.crop与out_of_bounds_policy="pad"配合使用时,会将图像居中,并在四周对称填充,效果与 Ultralytics 的LetterBox(center=True)行为一致。几乎完全一致。在
fn.resize中设置antialias=False,即可匹配 OpenCV 的cv2.INTER_LINEAR。由于 GPU 与 CPU 的运算方式不同,可能会出现轻微的浮点差异(< 0.001),但对检测精度没有可测量的影响。