Ultralytics YOLO27:
Get Started

实用工具简介#

YOLO model code with 3D perspective visualization

ultralytics 软件包提供了各种实用工具,帮助支持、增强和加速你的工作流。虽然还有许多其他可用工具,但本指南重点介绍了一些对开发者最有用的工具,可作为使用 Ultralytics 工具进行编程的实用参考。



观看: Ultralytics 实用工具 | 自动标注、Explorer API 和数据集转换

数据#

自动标注 / 注释#

数据集注释 是一个资源密集且耗时的过程。如果你有一个使用合理数量的数据训练的 Ultralytics YOLO 目标检测模型,可以将该模型与 SAM 配合使用,以分割格式自动标注更多数据。

from ultralytics.data.annotator import auto_annotate

auto_annotate(
    data="path/to/new/data",
    det_model="yolo26n.pt",
    sam_model="mobile_sam.pt",
    device="cuda",
    output_dir="path/to/save_labels",
)

此函数不返回任何值。更多详情:

可视化数据集注释#

此函数会在训练前可视化图像上的 YOLO 注释,帮助你识别并修正可能导致检测结果错误的注释。它会绘制边界框、使用类别名称标记对象,并根据背景亮度调整文本颜色,以提高可读性。

from ultralytics.data.utils import visualize_image_annotations

label_map = {  # 定义包含所有已注释类别标签的标签映射。
    0: "person",
    1: "car",
}

# Visualize
visualize_image_annotations(
    "path/to/image.jpg",  # 输入图像路径。
    "path/to/annotations.txt",  # 图像的注释文件路径。
    label_map,
)

将分割掩码转换为 YOLO 格式#

分割掩码转为 YOLO 格式

使用此工具将分割掩码图像数据集转换为 Ultralytics YOLO 分割格式。此函数会读取包含二值格式掩码图像的目录,并将这些图像转换为 YOLO 分割格式。

转换后的掩码将保存在指定的输出目录中。

from ultralytics.data.converter import convert_segment_masks_to_yolo_seg

# `classes` 是数据集中的类别总数。
# COCO 数据集包含 80 个类别。
convert_segment_masks_to_yolo_seg(masks_dir="path/to/masks_dir", output_dir="path/to/output_dir", classes=80)

将 COCO 转换为 YOLO 格式#

使用此工具将 COCO JSON 注释转换为 YOLO 格式。对于目标检测(边界框)数据集,请将 use_segments 和 use_keypoints 都设为 False。

from ultralytics.data.converter import convert_coco

convert_coco(
    "coco/annotations/",
    use_segments=False,
    use_keypoints=False,
    cls91to80=True,
)

如需了解有关 convert_coco 函数的更多信息,请访问参考页面。

获取边界框尺寸#

import cv2

from ultralytics import YOLO
from ultralytics.utils.plotting import Annotator

model = YOLO("yolo26n.pt")  # Load pretrain or fine-tune model

# Process the image
source = cv2.imread("path/to/image.jpg")
results = model(source)

# Extract results
annotator = Annotator(source, example=model.names)

for box in results[0].boxes.xyxy.cpu():
    width, height, area = annotator.get_bbox_dimension(box)
    print(f"Bounding Box Width {width.item()}, Height {height.item()}, Area {area.item()}")

将边界框转换为分割段#

对于现有的 x y w h 边界框数据,可以使用 yolo_bbox2segment 函数将其转换为分割段。请按以下方式组织图像和注释文件:

data
|__ images
    ├─ 001.jpg
    ├─ 002.jpg
    ├─ ..
    └─ NNN.jpg
|__ labels
    ├─ 001.txt
    ├─ 002.txt
    ├─ ..
    └─ NNN.txt
from ultralytics.data.converter import yolo_bbox2segment

yolo_bbox2segment(
    im_dir="path/to/images",
    save_dir=None,  # 保存到图像目录旁边的 "labels-segment" 目录中
    sam_model="sam_b.pt",
)

如需了解该函数的更多信息,请访问 yolo_bbox2segment 参考页面。

将分割段转换为边界框#

如果你的数据集采用分割数据集格式,可以使用此函数轻松将其转换为直立(或水平)边界框(x y w h 格式)。

import numpy as np

from ultralytics.utils.ops import segments2boxes

segments = np.array(
    [
        [805, 392, 797, 400, 812, 402, 808, 714, 808, 392],
        [115, 398, 113, 400, 150, 410, 150, 400, 149, 298],
        [267, 412, 265, 413, 300, 420, 300, 413, 299, 412],
    ],
    dtype=np.float32,
)

segments2boxes([s.reshape(-1, 2) for s in segments])
# >>> array([[804.5, 553. ,  15. , 322. ],
#           [131.5, 354. ,  37. , 112. ],
#           [282.5, 416. ,  35. ,   8. ]],
# dtype=float32) # xywh 边界框

如需了解此函数的工作原理,请访问参考页面。

实用工具#

图像压缩#

压缩单个图像文件以减小其大小,同时保持宽高比和画质。如果输入图像的尺寸小于最大尺寸,则不会调整其大小。

from pathlib import Path

from ultralytics.data.utils import compress_one_image

for f in Path("path/to/dataset").rglob("*.jpg"):
    compress_one_image(f)

自动拆分数据集#

自动将数据集拆分为 train/val/test 子集,并将拆分结果保存到 autosplit_*.txt 文件中。此工具会创建持久化的拆分文件;而 fraction 训练参数会为一次运行选择可复现的子集。

from ultralytics.data.split import autosplit

autosplit(
    path="path/to/images",
    weights=(0.9, 0.1, 0.0),  # (train, validation, test) 比例拆分
    annotated_only=False,  # 为 True 时,仅拆分具有注释文件的图像
)

如需了解此函数的更多详情,请参阅参考页面。

分割多边形转二值掩码#

将单个多边形(以列表形式表示)转换为指定图像尺寸的二值掩码。多边形应为一个扁平的一维数组,其中包含 N 个坐标,并以交替排列的 x, y 值表示多边形轮廓。

警告

N 必须始终为偶数。

import numpy as np

from ultralytics.data.utils import polygon2mask

imgsz = (1080, 810)
polygon = np.array([805, 392, 797, 400, ..., 808, 714, 808, 392])  # (N,) 扁平 x、y 坐标

mask = polygon2mask(
    imgsz,  # tuple
    [polygon],  # 以列表形式输入
    color=255,  # 8 位二值
    downsample_ratio=1,
)

边界框#

边界框(水平)实例#

要管理边界框数据,Bboxes 类可以帮助你在不同框坐标格式之间转换、缩放框尺寸、计算面积、添加偏移量等。

import numpy as np

from ultralytics.utils.instance import Bboxes

boxes = Bboxes(
    bboxes=np.array(
        [
            [22.878, 231.27, 804.98, 756.83],
            [48.552, 398.56, 245.35, 902.71],
            [669.47, 392.19, 809.72, 877.04],
            [221.52, 405.8, 344.98, 857.54],
            [0, 550.53, 63.01, 873.44],
            [0.0584, 254.46, 32.561, 324.87],
        ]
    ),
    format="xyxy",
)

boxes.areas()
# >>> array([ 4.1104e+05,       99216,       68000,       55772,       20347,      2288.5])

boxes.convert("xywh")
print(boxes.bboxes)
# >>> array(
#     [[ 413.93, 494.05,  782.1, 525.56],
#      [ 146.95, 650.63,  196.8, 504.15],
#      [  739.6, 634.62, 140.25, 484.85],
#      [ 283.25, 631.67, 123.46, 451.74],
#      [ 31.505, 711.99,  63.01, 322.91],
#      [  16.31, 289.67, 32.503,  70.41]]
# )

如需了解更多属性和方法,请参阅 Bboxes 参考部分。

提示

以下许多函数以及其他函数都可以通过 Bboxes 类访问;如果你更喜欢直接使用函数,请参阅接下来的小节,了解如何单独导入这些函数。

缩放边界框#

放大或缩小图像时,可以使用 ultralytics.utils.ops.scale_boxes 相应地缩放边界框坐标,使其与图像匹配。

import cv2 as cv
import numpy as np

from ultralytics.utils.ops import scale_boxes

image = cv.imread("ultralytics/assets/bus.jpg")
h, w, c = image.shape
resized = cv.resize(image, None, fx=1.2, fy=1.2)
new_h, new_w, _ = resized.shape

xyxy_boxes = np.array(
    [
        [22.878, 231.27, 804.98, 756.83],
        [48.552, 398.56, 245.35, 902.71],
        [669.47, 392.19, 809.72, 877.04],
        [221.52, 405.8, 344.98, 857.54],
        [0, 550.53, 63.01, 873.44],
        [0.0584, 254.46, 32.561, 324.87],
    ]
)

new_boxes = scale_boxes(
    img1_shape=(h, w),  # 原始图像尺寸
    boxes=xyxy_boxes,  # 原始图像中的边界框
    img0_shape=(new_h, new_w),  # 调整后的图像尺寸(缩放到)
    ratio_pad=None,
    padding=False,
    xywh=False,
)

print(new_boxes)
# >>> array(
#     [[  27.454,  277.52,  965.98,   908.2],
#     [   58.262,  478.27,  294.42,  1083.3],
#     [   803.36,  470.63,  971.66,  1052.4],
#     [   265.82,  486.96,  413.98,    1029],
#     [        0,  660.64,  75.612,  1048.1],
#     [   0.0701,  305.35,  39.073,  389.84]]
# )

边界框格式转换#

XYXY → XYWH#

将边界框坐标从 (x1, y1, x2, y2) 格式转换为 (x, y, width, height) 格式,其中 (x1, y1) 是左上角,(x2, y2) 是右下角。

import numpy as np

from ultralytics.utils.ops import xyxy2xywh

xyxy_boxes = np.array(
    [
        [22.878, 231.27, 804.98, 756.83],
        [48.552, 398.56, 245.35, 902.71],
        [669.47, 392.19, 809.72, 877.04],
        [221.52, 405.8, 344.98, 857.54],
        [0, 550.53, 63.01, 873.44],
        [0.0584, 254.46, 32.561, 324.87],
    ]
)
xywh = xyxy2xywh(xyxy_boxes)

print(xywh)
# >>> array(
#     [[ 413.93,  494.05,   782.1, 525.56],
#     [  146.95,  650.63,   196.8, 504.15],
#     [   739.6,  634.62,  140.25, 484.85],
#     [  283.25,  631.67,  123.46, 451.74],
#     [  31.505,  711.99,   63.01, 322.91],
#     [   16.31,  289.67,  32.503,  70.41]]
# )

所有边界框转换#

from ultralytics.utils.ops import (
    ltwh2xywh,
    ltwh2xyxy,
    xywh2ltwh,  # xywh → 左上角、w、h
    xywh2xyxy,
    xywhn2xyxy,  # 归一化坐标 → 像素坐标
    xyxy2ltwh,  # xyxy → 左上角、w、h
    xyxy2xywhn,  # 像素坐标 → 归一化坐标
)

for func in (ltwh2xywh, ltwh2xyxy, xywh2ltwh, xywh2xyxy, xywhn2xyxy, xyxy2ltwh, xyxy2xywhn):
    print(help(func))  # 打印函数文档字符串

请参阅每个函数的文档字符串,或访问 ultralytics.utils.ops 参考页面了解更多信息。

绘图#

注释工具#

Ultralytics 提供了一个 Annotator 类,用于为各种数据类型添加注释。此类最适合用于目标检测边界框、姿态关键点和有向边界框。

边界框注释#

使用 Ultralytics YOLO 的 Python 示例 🚀
import cv2 as cv
import numpy as np

from ultralytics.utils.plotting import Annotator, colors

names = {
    0: "person",
    5: "bus",
    11: "stop sign",
}

image = cv.imread("ultralytics/assets/bus.jpg")
ann = Annotator(
    image,
    line_width=None,  # default auto-size
    font_size=None,  # default auto-size
    font="Arial.ttf",  # must be ImageFont compatible
    pil=False,  # use PIL, otherwise uses OpenCV
)

xyxy_boxes = np.array(
    [
        [5, 22.878, 231.27, 804.98, 756.83],  # class-idx x1 y1 x2 y2
        [0, 48.552, 398.56, 245.35, 902.71],
        [0, 669.47, 392.19, 809.72, 877.04],
        [0, 221.52, 405.8, 344.98, 857.54],
        [0, 0, 550.53, 63.01, 873.44],
        [11, 0.0584, 254.46, 32.561, 324.87],
    ]
)

for nb, box in enumerate(xyxy_boxes):
    c_idx, *box = box
    label = f"{str(nb).zfill(2)}:{names.get(int(c_idx))}"
    ann.box_label(box, label, color=colors(c_idx, bgr=True))

image_with_bboxes = ann.result()

使用检测结果时,可以使用 model.names 中的名称。 如需进一步了解,请参阅 Annotator 参考页面。

Ultralytics 扫掠注释#

使用 Ultralytics 实用工具进行扫掠注释
import sys

import cv2
import numpy as np

from ultralytics import YOLO
from ultralytics.solutions.solutions import SolutionAnnotator
from ultralytics.utils.plotting import colors

# User defined video path and model file
cap = cv2.VideoCapture("path/to/video.mp4")
model = YOLO(model="yolo26s-seg.pt")  # Model file, e.g., yolo26s.pt or yolo26m-seg.pt

if not cap.isOpened():
    print("Error: Could not open video.")
    sys.exit()

# Initialize the video writer object.
w, h, fps = (int(cap.get(x)) for x in (cv2.CAP_PROP_FRAME_WIDTH, cv2.CAP_PROP_FRAME_HEIGHT, cv2.CAP_PROP_FPS))
video_writer = cv2.VideoWriter("ultralytics.avi", cv2.VideoWriter_fourcc(*"mp4v"), fps, (w, h))

masks = None  # Initialize variable to store masks data
f = 0  # Initialize frame count variable for enabling mouse event.
line_x = w  # Store width of line.
dragging = False  # Initialize bool variable for line dragging.
classes = model.names  # Store model classes names for plotting.
window_name = "Ultralytics Sweep Annotator"

def drag_line(event, x, _, flags, param):
    """Mouse callback function to enable dragging a vertical sweep line across the video frame."""
    global line_x, dragging
    if event == cv2.EVENT_LBUTTONDOWN or (flags & cv2.EVENT_FLAG_LBUTTON):
        line_x = max(0, min(x, w))
        dragging = True

while cap.isOpened():  # Loop over the video capture object.
    ret, im0 = cap.read()
    if not ret:
        break
    f = f + 1  # Increment frame count.
    count = 0  # Re-initialize count variable on every frame for precise counts.
    results = model.track(im0, persist=True)[0]

    if f == 1:
        cv2.namedWindow(window_name)
        cv2.setMouseCallback(window_name, drag_line)

    annotator = SolutionAnnotator(im0)

    if results.boxes.is_track:
        if results.masks is not None:
            masks = [np.array(m, dtype=np.int32) for m in results.masks.xy]

        boxes = results.boxes.xyxy.tolist()
        track_ids = results.boxes.id.int().cpu().tolist()
        clss = results.boxes.cls.cpu().tolist()

        for mask, box, cls, t_id in zip(masks or [None] * len(boxes), boxes, clss, track_ids):
            color = colors(t_id, True)  # Assign different color to each tracked object.
            label = f"{classes[cls]}:{t_id}"
            if mask is not None and mask.size > 0:
                if box[0] > line_x:
                    count += 1
                    cv2.polylines(im0, [mask], True, color, 2)
                    x, y = mask.min(axis=0)
                    (w_m, _), _ = cv2.getTextSize(label, cv2.FONT_HERSHEY_SIMPLEX, 0.5, 1)
                    cv2.rectangle(im0, (x, y - 20), (x + w_m, y), color, -1)
                    cv2.putText(im0, label, (x, y - 5), cv2.FONT_HERSHEY_SIMPLEX, 0.5, (255, 255, 255), 1)
            else:
                if box[0] > line_x:
                    count += 1
                    annotator.box_label(box=box, color=color, label=label)

    # Generate draggable sweep line
    annotator.sweep_annotator(line_x=line_x, line_y=h, label=f"COUNT:{count}")

    cv2.imshow(window_name, im0)
    video_writer.write(im0)
    if cv2.waitKey(1) & 0xFF == ord("q"):
        break

# Release the resources
cap.release()
video_writer.release()
cv2.destroyAllWindows()

如需了解有关 sweep_annotator 方法的更多详情,请访问参考部分此处。

自适应标签注释#

SolutionAnnotator.adaptive_label 会绘制标签,其形状由 shape 参数决定(该参数取代了较旧的 circle_label 和 text_label 方法):

  • 矩形:annotator.adaptive_label(box, label=names[int(cls)], color=colors(cls, True), shape="rect")
  • 圆形:annotator.adaptive_label(box, label=names[int(cls)], color=colors(cls, True), shape="circle")


观看: Python 文本和圆形标注深度指南 | 实时演示 | Ultralytics 标注 🚀
使用 Ultralytics 实用工具进行自适应标签注释
import cv2

from ultralytics import YOLO
from ultralytics.solutions.solutions import SolutionAnnotator
from ultralytics.utils.plotting import colors

model = YOLO("yolo26s.pt")
names = model.names
cap = cv2.VideoCapture("path/to/video.mp4")

w, h, fps = (int(cap.get(x)) for x in (cv2.CAP_PROP_FRAME_WIDTH, cv2.CAP_PROP_FRAME_HEIGHT, cv2.CAP_PROP_FPS))
writer = cv2.VideoWriter("Ultralytics circle annotation.avi", cv2.VideoWriter_fourcc(*"MJPG"), fps, (w, h))

while True:
    ret, im0 = cap.read()
    if not ret:
        break

    annotator = SolutionAnnotator(im0)
    results = model.predict(im0)[0]
    boxes = results.boxes.xyxy.cpu()
    clss = results.boxes.cls.cpu().tolist()

    for box, cls in zip(boxes, clss):
        annotator.adaptive_label(box, label=names[int(cls)], color=colors(cls, True), shape="circle")
    writer.write(im0)
    cv2.imshow("Ultralytics circle annotation", im0)

    if cv2.waitKey(1) & 0xFF == ord("q"):
        break

writer.release()
cap.release()
cv2.destroyAllWindows()

如需了解更多信息,请参阅 SolutionAnnotator 参考页面。

其他#

代码性能分析#

使用 with 或将其用作装饰器,检查代码运行/处理所需的时间。

from ultralytics.utils.ops import Profile

with Profile(device="cuda:0") as dt:
    pass  # 用于测量的操作

print(dt)
# >>> "Elapsed time is 9.5367431640625e-07 s"

Ultralytics 支持的格式#

需要在 Ultralytics 中以编程方式使用支持的图像或视频格式?如有需要,请使用这些常量:

from ultralytics.data.utils import IMG_FORMATS, VID_FORMATS

print(IMG_FORMATS)
# {'avif', 'bmp', 'dng', 'heic', 'heif', 'jp2', 'jpeg', 'jpg', 'mpo', 'png', 'tif', 'tiff', 'webp'}

print(VID_FORMATS)
# {'asf', 'avi', 'gif', 'm4v', 'mkv', 'mov', 'mp4', 'mpeg', 'mpg', 'ts', 'wmv', 'webm'}

向上取整到可整除数#

计算大于或等于 x 且能被 y 整除的最小整数。

from ultralytics.utils.ops import make_divisible

make_divisible(7, 3)
# >>> 9
make_divisible(7, 2)
# >>> 8

常见问题#

  • Ultralytics 包含一系列实用工具,旨在简化并优化机器学习工作流。主要工具包括用于标注数据集的自动标注、使用 convert_coco 将 COCO 转换为 YOLO 格式、压缩图像,以及自动拆分数据集。这些工具可以减少手动操作、确保一致性并提高数据处理效率。

  • 如果你有一个预训练的 Ultralytics YOLO 目标检测模型,可以将它与 SAM 模型配合使用,以分割格式自动标注数据集。示例如下:

    from ultralytics.data.annotator import auto_annotate
    
    auto_annotate(
        data="path/to/new/data",
        det_model="yolo26n.pt",
        sam_model="mobile_sam.pt",
        device="cuda",
        output_dir="path/to/save_labels",
    )

    如需了解更多详情,请参阅 auto_annotate 参考部分,或者使用 Ultralytics Platform 这一托管式免代码替代方案,通过 SAM 2.1、SAM 3 或 SAM 3.1 进行点击式掩码标注,也可以使用预训练和微调 YOLO 模型对 detect、segment 和 OBB 任务进行预测。

  • 要将 COCO JSON 标注转换为用于目标检测的 YOLO 格式,你可以使用 convert_coco 实用工具。示例代码如下:

    from ultralytics.data.converter import convert_coco
    
    convert_coco(
        "coco/annotations/",
        use_segments=False,
        use_keypoints=False,
        cls91to80=True,
    )

    如需了解更多信息,请访问 convert_coco 参考页面。

  • Ultralytics Platform 提供自动数据集分析:Charts 选项卡会显示拆分分布、数量最多的类别、图像尺寸直方图以及标注位置的 2D 热力图,帮助你在训练前发现数据不平衡和异常值。

  • 要将现有边界框数据(x y w h 格式)转换为分割段,你可以使用 yolo_bbox2segment 函数。请确保文件分别存放在图像目录和标签目录中。

    from ultralytics.data.converter import yolo_bbox2segment
    
    yolo_bbox2segment(
        im_dir="path/to/images",
        save_dir=None,  # 保存到图像目录旁边的 "labels-segment" 目录中
        sam_model="sam_b.pt",
    )

    如需了解更多信息,请访问 yolo_bbox2segment 参考页面。

评论