Ultralytics YOLO27:
Get Started

KITTI 깊이 데이터셋#

KITTI 데이터셋은 카를스루에 시내와 주변에서 주행 차량으로 촬영한 실제 야외 자율주행 벤치마크입니다. 단안 깊이 추정의 ground truth 깊이는 Velodyne HDL-64 LiDAR 스캐너로 얻었으며, Uhrig 외(2017)의 방법으로 조밀화했습니다. 결과 깊이 맵은 여전히 희소하며, 유효한 깊이값이 있는 픽셀은 약 16~20%입니다. KITTI는 YOLO26-Depth 사전 학습 혼합 데이터의 실제 야외 주행 데이터 소스이며, KITTI Eigen 평가 벤치마크로도 사용됩니다.

주요 기능#

  • 일반적인 실내 범위를 훨씬 뛰어넘어 최대 약 80m 깊이까지 포함하는 실제 야외 주행 장면입니다.
  • Velodyne HDL-64 LiDAR로 얻은 깊이 ground truth를 Uhrig 외(2017)의 Sparsity Invariant CNNs 접근법으로 조밀화했습니다.
  • 희소한 지도 학습: 이미지당 픽셀의 약 16~20%에만 유효한 깊이값이 있으며, 유효하지 않은 픽셀은 손실 및 지표 계산에서 마스킹됩니다.
  • 스테레오 이미지 쌍(왼쪽 image_02 및 오른쪽 image_03)은 학습에 추가 시점을 제공합니다.
  • 깊이값은 미터당 256단위(depth_scale: 256)를 사용하는 uint16 PNG로 저장하며, Ultralytics 깊이 데이터셋 형식을 따릅니다.

데이터셋 구조#

Ultralytics에서 사용하는 KITTI 깊이 데이터는 두 개의 하위 집합으로 나뉩니다:

  1. 학습 분할: 이미지 55,198개(왼쪽 image_02 및 오른쪽 image_03). 공정한 평가를 위해 KITTI Eigen 테스트 주행 28개는 모두 학습에서 제외됩니다.
  2. 평가 분할: KITTI Eigen 테스트 분할의 왼쪽 카메라 프레임 중 개선된 ground truth가 있는 652개입니다. 평가에는 80m 깊이 상한과 예측값 및 ground truth 간의 중앙값(스케일만) 정렬을 사용합니다.

깊이 범위는 약 80m까지이며, 데이터셋 YAML(depth-kitti.yaml)은 이에 맞게 max_depth: 80을 설정합니다. 처음 사용할 때 YAML은 개선된 ground truth(~14 GB)와 해당 원시 주행 데이터(~175 GB)를 자동으로 다운로드하고 Ultralytics 깊이 데이터셋 형식으로 변환합니다.

YOLO26-Depth에서의 역할#

KITTI는 YOLO26-Depth 사전 학습 혼합 데이터에서 실제 야외 주행 지도 학습 데이터를 제공하여 실내 위주 데이터 소스를 보완합니다. 또한 주행 장면 깊이 정확도를 보고하는 표준 KITTI Eigen 벤치마크로 사용됩니다.

KITTI는 깊이 헤드가 제한되지 않는(log 모드) 이유를 잘 보여주는 대표 사례입니다. 고정된 10m 출력 상한으로는 80m 주행 장면을 표현할 수 없습니다. 헤드 출력 범위와 max_depth 처리에 대한 자세한 내용은 깊이 작업 페이지를 참조하세요.

결과#

모델 크기별 KITTI Eigen delta1 정확도(높을수록 좋음):

모델KITTI Eigen δ1
YOLO26n-Depth0.878
YOLO26s-Depth0.879
YOLO26m-Depth0.913
YOLO26l-Depth0.926
YOLO26x-Depth0.932

imgsz=768 및 rect=False을 사용해 표준 652프레임 분할에서 측정했습니다. 다음과 같은 차이 때문에 공개된 KITTI 수치와 직접 비교할 수 없습니다. val에서는 각 이미지를 정사각형 imgsz로 늘리고 희소 ground truth를 최근접 보간으로 이에 맞춰 리샘플링하지만, 기준 평가기는 예측값을 원래 ground truth 해상도로 다시 조정합니다. DepthMetrics은 ground truth를 max_depth에 따라 마스킹하지만, 개선된 ground truth 프로토콜은 gt > 0를 마스킹하고 예측값에만 상한을 적용합니다. 또한 공개 가중치는 이전 분할로 학습되었으며, 이 분할에는 테스트 프레임 72개가 학습 세트에 포함되어 있었습니다.

데이터셋 YAML#

데이터셋 구성을 정의하는 데 YAML 파일을 사용합니다. YAML 파일에는 데이터셋 경로, 클래스 및 최대 깊이와 같은 기타 관련 정보가 포함됩니다.

ultralytics/cfg/datasets/depth-kitti.yaml
# Ultralytics 🚀 AGPL-3.0 License - https://ultralytics.com/license

# KITTI dataset for monocular depth estimation — real outdoor driving, Velodyne HDL-64 LiDAR (densified), sparse, up to ~80 m
# Documentation: https://docs.ultralytics.com/datasets/depth/kitti
# Example usage: yolo depth train data=depth-kitti.yaml model=yolo26n-depth.pt
# parent
# ├── ultralytics
# └── datasets
#     └── depth-kitti  ← downloads here (~190 GB archives, ~140 GB converted)
#         ├── images/{train,val}  # RGB images
#         └── depth/{train,val}   # paired 16-bit *.png depth maps (images/ -> depth/)
# If an interrupted download leaves a partial dataset, delete the depth-kitti dir and re-run to rebuild it.
# The download only runs when the val images are missing (train is never checked), so a depth-kitti dir
# built before this split fix keeps the old split until it is deleted and rebuilt.

path: depth-kitti # dataset root dir (relative to Ultralytics settings 'datasets_dir')
train: images/train # train images (relative to 'path') 55198 images
val: images/val # val images (relative to 'path') 652 images (KITTI Eigen test split, left camera)
max_depth: 80 # (m) maximum valid depth; GT beyond this is excluded from val metrics

nc: 1
names:
  0: depth

channels: 3
depth_scale: 256 # KITTI PNG value 256 = 1 meter

# Download script/URL (optional)
download: |
  import shutil
  from pathlib import Path

  from ultralytics.utils import TQDM
  from ultralytics.utils.downloads import download

  # The 28 KITTI Eigen test scenes as "<drive> <frame indices>" -> val; every other annotated drive -> train. val is
  # the 697-frame Eigen et al. (NeurIPS 2014) test split restricted to frames with improved GT here: 652 left-camera.
  EIGEN_TEST = """
  2011_09_26_drive_0002 6 9 12 15 18 21 24 27 30 33 36 39 42 45 48 51 54 57 60 63 69
  2011_09_26_drive_0009 16 32 48 64 80 96 112 128 144 160 176 196 212 228 244 260 276 292 308 324 340 356 372 388
  2011_09_26_drive_0013 5 10 20 30 35 40 45 50 60 65 70 80 85 90 95 100 105 110 115 120 125 130 135
  2011_09_26_drive_0020 6 9 12 15 18 21 27 33 36 42 45 48 51 54 57 60 63 66 69 72 75 78
  2011_09_26_drive_0023 18 36 54 72 90 108 126 144 162 180 198 216 252 270 288 306 324 342 360 378 396 414 432 450 468
  2011_09_26_drive_0027 7 14 21 28 35 42 49 56 63 70 77 84 91 98 112 119 126 133 140 147 161 168 175 182
  2011_09_26_drive_0029 14 28 42 56 70 84 98 112 126 140 154 168 182 196 268 296 310 324 338 352 366 380 394 408
  2011_09_26_drive_0036 32 64 96 128 160 192 224 256 288 320 352 384 416 448 480 512 544 576 608 640 672 704 768
  2011_09_26_drive_0046 5 10 15 20 25 30 35 40 45 50 55 60 65 70 75 80 85 90 95 100 105 110 115
  2011_09_26_drive_0048 5 6 7 8 9 10 11 12 13 14 15 16
  2011_09_26_drive_0052 6 8 10 12 14 16 18 20 22 26 28 30 32 36 38 40 42 44 46 48 50 52
  2011_09_26_drive_0056 11 22 33 44 55 66 77 88 99 110 121 143 154 165 176 187 198 209 220 231 242 253 264 275 286
  2011_09_26_drive_0059 14 28 42 56 70 84 98 112 126 140 154 182 196 210 224 238 274 288 302 316 330 344 358
  2011_09_26_drive_0064 22 44 66 88 110 154 176 198 220 242 264 286 308 330 352 374 396 418 440 462 484 506 528 550
  2011_09_26_drive_0084 49 62 75 88 101 114 127 140 153 179 192 205 218 231 244 257 270 283 296 309 322 335 348 361 374
  2011_09_26_drive_0086 7 34 61 88 115 142 169 196 223 250 277 304 331 358 385 412 439 466 493 520 574 601 628 655 682
  2011_09_26_drive_0093 16 32 48 64 80 96 112 128 144 160 176 192 208 224 240 256 305 321 337 353 369 385 401 417
  2011_09_26_drive_0096 19 38 57 76 95 114 133 152 171 190 209 228 247 266 285 304 323 342 361 380 399 418 437 456
  2011_09_26_drive_0101 80 114 148 182 216 284 318 352 386 420 454 488 522 556 590 624 658 692 726 760 794 828 862 896 930
  2011_09_26_drive_0106 15 35 43 51 59 67 75 83 91 99 107 115 123 131 139 147 155 163 171 179 187 195 203 211 219
  2011_09_26_drive_0117 26 52 78 104 130 156 182 208 234 260 286 312 338 364 390 416 442 468 494 546 572 598 624 650
  2011_09_28_drive_0002 6 9 12 18 21 24 30 33 36 39 45 48 51 54 57 60 63 72 78 81 84 87 90
  2011_09_29_drive_0071 36 72 108 144 180 216 252 288 324 360 396 432 468 504 540 576 612 735 771 807 879 915 951
  2011_09_30_drive_0016 11 22 33 44 55 66 77 88 110 121 132 143 154 165 176 187 198 209 220 231 242 253 264
  2011_09_30_drive_0018 107 214 428 535 642 749 856 963 1070 1177 1284 1391 1498 1605 1712 1819 1926 2033 2140 2247 2419 2526 2633 2740
  2011_09_30_drive_0027 41 82 123 164 205 246 287 328 369 410 451 492 533 574 615 656 753 794 835 876 917 958 1040 1081
  2011_10_03_drive_0027 181 362 543 734 915 1096 1277 1458 1639 1820 2001 2363 2544 2725 2906 3087 3268 3449 3630 3811 3992 4173 4354 4535
  2011_10_03_drive_0047 32 64 96 128 160 192 224 288 320 352 384 416 448 480 512 544 576 608 640 672 704 736 768 800"""
  EIGEN_TEST = {ln.split()[0]: {int(i) for i in ln.split()[1:]} for ln in EIGEN_TEST.strip().splitlines()}
  # Annotated drives excluded from the release training set (raw data was unavailable at build time)
  SKIP_DRIVES = {"2011_09_28_drive_0104", "2011_09_28_drive_0135", "2011_09_28_drive_0177", "2011_09_28_drive_0214"}

  # Download the improved sparse GT (~14 GB) and the raw recordings it covers (~175 GB)
  dir = Path(yaml["path"])  # dataset root dir
  base = "https://s3.eu-central-1.amazonaws.com/avg-kitti"
  download([f"{base}/data_depth_annotated.zip"], dir=dir / "source", delete=True, exist_ok=True)
  gt_drives = sorted(p for p in (dir / "source" / "data_depth_annotated").glob("*/*_sync") if p.name[:-5] not in SKIP_DRIVES)
  urls = [f"{base}/raw_data/{p.name[:-5]}/{p.name}.zip" for p in gt_drives]
  download(urls, dir=dir / "source" / "raw", delete=True, exist_ok=True, threads=4)

  # Convert: GT PNGs are uint16 meters*256; RGB frames come from the matching raw drive
  for split in ("train", "val"):
      (dir / "images" / split).mkdir(parents=True, exist_ok=True)
      (dir / "depth" / split).mkdir(parents=True, exist_ok=True)
  for gt in TQDM(gt_drives, desc="Converting"):
      drive = gt.name[:-5]  # 2011_09_26_drive_0002
      split = "val" if drive in EIGEN_TEST else "train"
      # val is the canonical Eigen test split: left camera only, and only its 652 listed frames
      for cam in ("image_02",) if split == "val" else ("image_02", "image_03"):
          for png in sorted((gt / "proj_depth" / "groundtruth" / cam).glob("*.png")):
              if split == "val" and int(png.stem) not in EIGEN_TEST[drive]:
                  continue
              name = f"{drive}_{cam[-2:]}_{png.stem}"
              png.replace(dir / "depth" / split / f"{name}.png")
              raw = dir / "source" / "raw" / drive[:10] / gt.name / cam / "data" / png.name
              raw.replace(dir / "images" / split / f"{name}.png")
  shutil.rmtree(dir / "source")

사용법#

KITTI 데이터셋으로 YOLO26n-Depth 모델을 100 에폭 동안 학습하려면 다음 코드 조각을 사용할 수 있습니다. 사용 가능한 인수의 전체 목록은 모델 학습 페이지를 참조하세요.

학습 예제
from ultralytics import YOLO

# 사전 학습된 깊이 모델을 불러옵니다.
model = YOLO("yolo26n-depth.pt")

# KITTI에서 모델 학습
results = model.train(data="depth-kitti.yaml", epochs=100, imgsz=640)

사전 학습 모델#

사전 학습된 YOLO26-Depth 모델은 이름으로 처음 참조할 때 Ultralytics v8.4.0 assets release에서 자동으로 다운로드됩니다:

인용 및 감사의 글#

연구 또는 개발 작업에서 KITTI 데이터셋을 사용하는 경우 다음 논문을 인용해 주세요:

인용
@article{geiger2013vision,
      title={Vision meets Robotics: The KITTI Dataset},
      author={Geiger, Andreas and Lenz, Philip and Stiller, Christoph and Urtasun, Raquel},
      journal={The International Journal of Robotics Research},
      year={2013},
      publisher={SAGE Publications}
}

@inproceedings{uhrig2017sparsity,
      title={Sparsity Invariant CNNs},
      author={Uhrig, Jonas and Schneider, Nick and Schneider, Lukas and Franke, Uwe and Brox, Thomas and Geiger, Andreas},
      booktitle={International Conference on 3D Vision (3DV)},
      year={2017}
}

KITTI 데이터셋을 만들고 유지 관리한 Karlsruhe Institute of Technology와 Toyota Technological Institute at Chicago, 그리고 더 조밀한 지도 학습을 가능하게 한 깊이 조밀화 방법을 개발한 Uhrig 외 연구진에게 감사드립니다.

자주 묻는 질문#

  • KITTI는 YOLO26-Depth 사전 학습 혼합 데이터의 실제 야외 주행 데이터 소스이며, KITTI Eigen 평가 벤치마크로도 사용됩니다. Velodyne LiDAR 깊이는 약 80m까지 측정되므로, 깊이 헤드는 일반적인 실내 범위로 제한되지 않고 상한이 없습니다.

  • Ultralytics 구성에서는 왼쪽 및 오른쪽 카메라의 학습 이미지 55,198개를 사용하고 KITTI Eigen 테스트 주행 28개를 모두 제외하며, 개선된 ground truth가 있는 Eigen 테스트용 왼쪽 카메라 프레임 652개를 평가에 사용합니다. 깊이 PNG는 미터당 256단위(depth_scale: 256)를 사용하며 YAML은 max_depth: 80을 설정합니다.

  • yolo depth train data=depth-kitti.yaml model=yolo26n-depth.pt epochs=100 imgsz=640을 실행하거나 Usage 섹션의 Python 예제를 사용하세요. 사용 가능한 인수는 모두 Training 페이지에 나와 있습니다.

  • 결과의 값은 imgsz=768 및 rect=False를 사용하는 Ultralytics 검증기에서 산출됩니다. 검증기는 이미지를 정사각형 입력 크기로 조정하고, max_depth에 따라 ground truth를 마스킹하며, 이미지별 스케일 정렬을 적용해 이미지별 지표를 계산한 뒤 검증 세트 전체의 평균을 구합니다. KITTI 기준 평가기는 다른 프로토콜을 사용하므로 수치를 직접 비교할 수 없습니다.

댓글