Ultralytics YOLO27:
Get Started

KITTI 深度数据集#

KITTI 数据集是一个真实户外自动驾驶基准,采集于行驶中的车辆,覆盖卡尔斯鲁厄市区及周边。对于单目深度估计,真值深度由 Velodyne HDL-64 LiDAR 扫描仪获取,并使用 Uhrig 等人 2017 年提出的方法进行稠密化。生成的深度图仍然是稀疏的,约 16–20% 的像素带有有效深度值。KITTI 是 YOLO26-Depth 预训练混合数据中的真实户外驾驶数据源,也用作 KITTI Eigen 评测基准。

主要功能#

  • 真实户外驾驶场景,深度范围最高约 80 m,远超典型室内范围。
  • 深度真值由 Velodyne HDL-64 LiDAR 获取,并使用 Uhrig 等人 2017 年提出的 Sparsity Invariant CNNs 方法进行稠密化。
  • 稀疏监督:每张图像中只有约 16–20% 的像素带有有效深度值;无效像素会从损失计算和指标计算中屏蔽。
  • 立体图像对(左 image_02 和右 image_03)提供额外视角用于训练。
  • 深度值以 uint16 PNG 格式存储,每米 256 个单位(depth_scale: 256),遵循 Ultralytics 深度数据集格式。

数据集结构#

Ultralytics 使用的 KITTI 深度数据分为两个子集:

  1. 训练集:55,198 张图像(左 image_02 和右 image_03)。为确保评测公平,所有 28 段 KITTI Eigen 测试行程均不用于训练。
  2. 评测集:KITTI Eigen 测试集——其中包含 652 帧具有改进真值的左摄像头图像。评测时将深度上限设为 80 m,并在预测结果与真值之间进行中值(仅尺度)对齐。

深度范围最高约为 80 m,数据集 YAML(depth-kitti.yaml)会相应设置 max_depth: 80。首次使用时,YAML 会自动下载改进真值(~14 GB)和匹配的原始行程数据(~175 GB),并将它们转换为 Ultralytics 深度数据集格式。

在 YOLO26-Depth 中的作用#

KITTI 为 YOLO26-Depth 预训练混合数据提供真实户外驾驶监督,补充了以室内数据为主的数据源。它也用于报告驾驶场景深度精度,是标准的 KITTI Eigen 基准。

KITTI 是深度头不受限(log 模式)的一个重要例证:固定的 10 m 输出上限无法表示 80 m 的驾驶场景。有关深度头输出范围和 max_depth 处理方式的详情,请参阅深度任务页面。

结果#

按模型规模统计的 KITTI Eigen delta1 精度(越高越好):

模型KITTI Eigen δ1
YOLO26n-Depth0.878
YOLO26s-Depth0.879
YOLO26m-Depth0.913
YOLO26l-Depth0.926
YOLO26x-Depth0.932

在包含 652 帧的标准划分上使用 imgsz=768 和 rect=False 测得。这些结果无法与已发表的 KITTI 数值直接比较:val 会将每张图像拉伸为正方形 imgsz,并将稀疏真值通过最近邻重采样到该尺寸;而参考评测器会将预测结果缩放回原生真值分辨率;DepthMetrics 会针对 max_depth 屏蔽真值,而改进真值协议会屏蔽 gt > 0,并且只对预测结果设上限;此外,已发布权重使用先前的划分进行训练,其中有 72 帧测试图像被纳入训练集。

数据集 YAML#

使用 YAML 文件定义数据集配置。文件包含数据集路径、类别以及最大深度等其他相关信息。

ultralytics/cfg/datasets/depth-kitti.yaml
# Ultralytics 🚀 AGPL-3.0 License - https://ultralytics.com/license

# KITTI dataset for monocular depth estimation — real outdoor driving, Velodyne HDL-64 LiDAR (densified), sparse, up to ~80 m
# Documentation: https://docs.ultralytics.com/datasets/depth/kitti
# Example usage: yolo depth train data=depth-kitti.yaml model=yolo26n-depth.pt
# parent
# ├── ultralytics
# └── datasets
#     └── depth-kitti  ← downloads here (~190 GB archives, ~140 GB converted)
#         ├── images/{train,val}  # RGB images
#         └── depth/{train,val}   # paired 16-bit *.png depth maps (images/ -> depth/)
# If an interrupted download leaves a partial dataset, delete the depth-kitti dir and re-run to rebuild it.
# The download only runs when the val images are missing (train is never checked), so a depth-kitti dir
# built before this split fix keeps the old split until it is deleted and rebuilt.

path: depth-kitti # dataset root dir (relative to Ultralytics settings 'datasets_dir')
train: images/train # train images (relative to 'path') 55198 images
val: images/val # val images (relative to 'path') 652 images (KITTI Eigen test split, left camera)
max_depth: 80 # (m) maximum valid depth; GT beyond this is excluded from val metrics

nc: 1
names:
  0: depth

channels: 3
depth_scale: 256 # KITTI PNG value 256 = 1 meter

# Download script/URL (optional)
download: |
  import shutil
  from pathlib import Path

  from ultralytics.utils import TQDM
  from ultralytics.utils.downloads import download

  # The 28 KITTI Eigen test scenes as "<drive> <frame indices>" -> val; every other annotated drive -> train. val is
  # the 697-frame Eigen et al. (NeurIPS 2014) test split restricted to frames with improved GT here: 652 left-camera.
  EIGEN_TEST = """
  2011_09_26_drive_0002 6 9 12 15 18 21 24 27 30 33 36 39 42 45 48 51 54 57 60 63 69
  2011_09_26_drive_0009 16 32 48 64 80 96 112 128 144 160 176 196 212 228 244 260 276 292 308 324 340 356 372 388
  2011_09_26_drive_0013 5 10 20 30 35 40 45 50 60 65 70 80 85 90 95 100 105 110 115 120 125 130 135
  2011_09_26_drive_0020 6 9 12 15 18 21 27 33 36 42 45 48 51 54 57 60 63 66 69 72 75 78
  2011_09_26_drive_0023 18 36 54 72 90 108 126 144 162 180 198 216 252 270 288 306 324 342 360 378 396 414 432 450 468
  2011_09_26_drive_0027 7 14 21 28 35 42 49 56 63 70 77 84 91 98 112 119 126 133 140 147 161 168 175 182
  2011_09_26_drive_0029 14 28 42 56 70 84 98 112 126 140 154 168 182 196 268 296 310 324 338 352 366 380 394 408
  2011_09_26_drive_0036 32 64 96 128 160 192 224 256 288 320 352 384 416 448 480 512 544 576 608 640 672 704 768
  2011_09_26_drive_0046 5 10 15 20 25 30 35 40 45 50 55 60 65 70 75 80 85 90 95 100 105 110 115
  2011_09_26_drive_0048 5 6 7 8 9 10 11 12 13 14 15 16
  2011_09_26_drive_0052 6 8 10 12 14 16 18 20 22 26 28 30 32 36 38 40 42 44 46 48 50 52
  2011_09_26_drive_0056 11 22 33 44 55 66 77 88 99 110 121 143 154 165 176 187 198 209 220 231 242 253 264 275 286
  2011_09_26_drive_0059 14 28 42 56 70 84 98 112 126 140 154 182 196 210 224 238 274 288 302 316 330 344 358
  2011_09_26_drive_0064 22 44 66 88 110 154 176 198 220 242 264 286 308 330 352 374 396 418 440 462 484 506 528 550
  2011_09_26_drive_0084 49 62 75 88 101 114 127 140 153 179 192 205 218 231 244 257 270 283 296 309 322 335 348 361 374
  2011_09_26_drive_0086 7 34 61 88 115 142 169 196 223 250 277 304 331 358 385 412 439 466 493 520 574 601 628 655 682
  2011_09_26_drive_0093 16 32 48 64 80 96 112 128 144 160 176 192 208 224 240 256 305 321 337 353 369 385 401 417
  2011_09_26_drive_0096 19 38 57 76 95 114 133 152 171 190 209 228 247 266 285 304 323 342 361 380 399 418 437 456
  2011_09_26_drive_0101 80 114 148 182 216 284 318 352 386 420 454 488 522 556 590 624 658 692 726 760 794 828 862 896 930
  2011_09_26_drive_0106 15 35 43 51 59 67 75 83 91 99 107 115 123 131 139 147 155 163 171 179 187 195 203 211 219
  2011_09_26_drive_0117 26 52 78 104 130 156 182 208 234 260 286 312 338 364 390 416 442 468 494 546 572 598 624 650
  2011_09_28_drive_0002 6 9 12 18 21 24 30 33 36 39 45 48 51 54 57 60 63 72 78 81 84 87 90
  2011_09_29_drive_0071 36 72 108 144 180 216 252 288 324 360 396 432 468 504 540 576 612 735 771 807 879 915 951
  2011_09_30_drive_0016 11 22 33 44 55 66 77 88 110 121 132 143 154 165 176 187 198 209 220 231 242 253 264
  2011_09_30_drive_0018 107 214 428 535 642 749 856 963 1070 1177 1284 1391 1498 1605 1712 1819 1926 2033 2140 2247 2419 2526 2633 2740
  2011_09_30_drive_0027 41 82 123 164 205 246 287 328 369 410 451 492 533 574 615 656 753 794 835 876 917 958 1040 1081
  2011_10_03_drive_0027 181 362 543 734 915 1096 1277 1458 1639 1820 2001 2363 2544 2725 2906 3087 3268 3449 3630 3811 3992 4173 4354 4535
  2011_10_03_drive_0047 32 64 96 128 160 192 224 288 320 352 384 416 448 480 512 544 576 608 640 672 704 736 768 800"""
  EIGEN_TEST = {ln.split()[0]: {int(i) for i in ln.split()[1:]} for ln in EIGEN_TEST.strip().splitlines()}
  # Annotated drives excluded from the release training set (raw data was unavailable at build time)
  SKIP_DRIVES = {"2011_09_28_drive_0104", "2011_09_28_drive_0135", "2011_09_28_drive_0177", "2011_09_28_drive_0214"}

  # Download the improved sparse GT (~14 GB) and the raw recordings it covers (~175 GB)
  dir = Path(yaml["path"])  # dataset root dir
  base = "https://s3.eu-central-1.amazonaws.com/avg-kitti"
  download([f"{base}/data_depth_annotated.zip"], dir=dir / "source", delete=True, exist_ok=True)
  gt_drives = sorted(p for p in (dir / "source" / "data_depth_annotated").glob("*/*_sync") if p.name[:-5] not in SKIP_DRIVES)
  urls = [f"{base}/raw_data/{p.name[:-5]}/{p.name}.zip" for p in gt_drives]
  download(urls, dir=dir / "source" / "raw", delete=True, exist_ok=True, threads=4)

  # Convert: GT PNGs are uint16 meters*256; RGB frames come from the matching raw drive
  for split in ("train", "val"):
      (dir / "images" / split).mkdir(parents=True, exist_ok=True)
      (dir / "depth" / split).mkdir(parents=True, exist_ok=True)
  for gt in TQDM(gt_drives, desc="Converting"):
      drive = gt.name[:-5]  # 2011_09_26_drive_0002
      split = "val" if drive in EIGEN_TEST else "train"
      # val is the canonical Eigen test split: left camera only, and only its 652 listed frames
      for cam in ("image_02",) if split == "val" else ("image_02", "image_03"):
          for png in sorted((gt / "proj_depth" / "groundtruth" / cam).glob("*.png")):
              if split == "val" and int(png.stem) not in EIGEN_TEST[drive]:
                  continue
              name = f"{drive}_{cam[-2:]}_{png.stem}"
              png.replace(dir / "depth" / split / f"{name}.png")
              raw = dir / "source" / "raw" / drive[:10] / gt.name / cam / "data" / png.name
              raw.replace(dir / "images" / split / f"{name}.png")
  shutil.rmtree(dir / "source")

用法#

如需在 KITTI 数据集上训练 YOLO26n-Depth 模型 100 个轮次,可以使用以下代码片段。可用参数的完整列表请参阅模型训练页面。

训练示例
from ultralytics import YOLO

# 加载预训练的深度模型
model = YOLO("yolo26n-depth.pt")

# 在 KITTI 上训练模型
results = model.train(data="depth-kitti.yaml", epochs=100, imgsz=640)

预训练模型#

首次通过名称引用预训练 YOLO26-Depth 模型时,会从 Ultralytics v8.4.0 资源发布页自动下载:

引用与致谢#

如果你在研究或开发工作中使用 KITTI 数据集,请引用以下论文:

引用格式
@article{geiger2013vision,
      title={Vision meets Robotics: The KITTI Dataset},
      author={Geiger, Andreas and Lenz, Philip and Stiller, Christoph and Urtasun, Raquel},
      journal={The International Journal of Robotics Research},
      year={2013},
      publisher={SAGE Publications}
}

@inproceedings{uhrig2017sparsity,
      title={Sparsity Invariant CNNs},
      author={Uhrig, Jonas and Schneider, Nick and Schneider, Lukas and Franke, Uwe and Brox, Thomas and Geiger, Andreas},
      booktitle={International Conference on 3D Vision (3DV)},
      year={2017}
}

我们感谢卡尔斯鲁厄理工学院和芝加哥丰田技术学院创建并维护 KITTI 数据集,也感谢 Uhrig 等人提出深度稠密化方法,使更稠密的监督成为可能。

常见问题#

  • KITTI 是 YOLO26-Depth 预训练混合数据中的真实户外驾驶数据源,也用作 KITTI Eigen 评测基准。它的 Velodyne LiDAR 深度范围约为 80 m,因此深度头不设上限,而不是限制在固定的室内范围内。

  • Ultralytics 配置使用来自左右摄像头的 55,198 张训练图像,排除全部 28 段 KITTI Eigen 测试行程,并使用 652 帧具有改进真值的左摄像头 Eigen 测试图像进行评测。深度 PNG 每米使用 256 个单位(depth_scale: 256),YAML 会设置 max_depth: 80。

  • 运行 yolo depth train data=depth-kitti.yaml model=yolo26n-depth.pt epochs=100 imgsz=640,或使用用法部分中的 Python 示例。训练页面列出了所有可用参数。

  • 结果中的数值来自 Ultralytics 验证器,使用 imgsz=768 和 rect=False,将图像缩放到正方形输入,针对 max_depth 屏蔽真值,并在每张图像上进行指标计算;计算前会按图像进行尺度对齐,然后对整个验证集取平均。KITTI 参考评测器采用不同的协议,因此这些数值无法直接比较。

评论