KITTI Depth 데이터셋#
KITTI 데이터셋은 칼스루에(Karlsruhe) 도시와 그 주변을 주행하는 차량에서 수집한 실제 야외 자율주행 벤치마크입니다. 단안 깊이 추정의 경우, 정답(ground-truth) 깊이는 Velodyne HDL-64 LiDAR 스캐너에서 추출되었으며 Uhrig et al. 2017의 방법을 사용하여 밀집화되었습니다. 결과 생성된 깊이 맵은 여전히 희소하며, 약 16~20%의 픽셀만 유효한 깊이 값을 가집니다. KITTI는 YOLO26-Depth 사전 훈련 믹스에서 유일한 실제 야외 장거리 소스이며, KITTI Eigen 평가 벤치마크 역할도 합니다.
주요 특징#
- 일반적인 실내 범위를 훨씬 뛰어넘는, 최대 약 80m에 이르는 깊이를 포함하는 실제 야외 주행 장면입니다.
- Velodyne HDL-64 LiDAR에서 획득하고 Uhrig 외 2017의 Sparsity Invariant CNNs 접근 방식을 사용하여 밀집화된 깊이 정답 데이터입니다.
- 희소 지도(Sparse supervision): 이미지당 약 16~20%의 픽셀만이 유효한 깊이 값을 가집니다. 유효하지 않은 픽셀은 손실 및 지표에서 마스킹 처리됩니다.
- 스테레오 이미지 쌍(왼쪽
image_02및 오른쪽image_03)은 훈련을 위한 추가 시점을 제공합니다. - 깊이 값은 미터 단위의
.npyfloat32 배열로 저장되며, Ultralytics depth dataset format을 따릅니다.
데이터셋 구조#
Ultralytics에서 사용하는 KITTI 깊이 데이터는 두 개의 하위 세트로 나뉩니다:
- Training split: 55,198개 이미지 (왼쪽
image_02및 오른쪽image_03). 공정한 평가를 위해 모든 28개의 KITTI Eigen 테스트 주행 데이터는 학습에서 제외됩니다. - Evaluation split: KITTI Eigen 테스트 분할 - 개선된 ground truth가 있는 652개의 왼쪽 카메라 프레임입니다. 평가는 80m 깊이 상한(depth cap)과 예측값과 ground truth 간의 중앙값(스케일 전용) 정렬을 사용합니다.
깊이 범위는 약 80m에 달하며, 데이터셋 YAML(depth-kitti.yaml)은 이에 맞춰 max_depth: 80을 설정합니다.
YOLO26-Depth에서의 역할#
KITTI는 주로 실내 소스로 구성된 YOLO26-Depth 사전 학습 조합에서 유일한 실제 야외 장거리 지도 데이터를 제공합니다. 또한 주행 장면 깊이 정확도를 보고하기 위한 표준 KITTI Eigen 벤치마크이기도 합니다.
KITTI는 왜 깊이 헤드가 바운드되지 않는지(log 모드) 보여주는 핵심적인 예시입니다. 고정된 10m 출력 상한선은 80m 주행 장면을 표현할 수 없습니다. 헤드 출력 범위 및 max_depth 처리에 대한 자세한 내용은 depth task page를 참조하세요.
결과#
모델 크기별 KITTI Eigen delta1 정확도(높을수록 좋음):
| 모델 | KITTI Eigen δ1 |
|---|---|
| YOLO26n-Depth | 0.878 |
| YOLO26s-Depth | 0.879 |
| YOLO26m-Depth | 0.913 |
| YOLO26l-Depth | 0.926 |
| YOLO26x-Depth | 0.932 |
imgsz=768 및 rect=False을 사용하여 652프레임 표준 분할에서 측정되었습니다. 이는 발표된 KITTI 수치와 직접 비교할 수 없습니다. 검증(val)은 각 이미지를 정사각형 imgsz(으)로 늘리고 희소한 ground truth를 그 위에 가장 가까운 이웃 방식으로 리샘플링합니다. 반면 참조 평가기는 예측을 원래 ground truth 해상도로 다시 크기 조절하고; DepthMetrics은 max_depth에 대해 ground truth를 마스킹하는 반면, 개선된 ground truth 프로토콜은 gt > 0를 마스킹하고 예측값에만 상한을 적용합니다. 또한 이미지별 지표를 평균화하는 대신 분할의 모든 유효한 픽셀을 풀링하며, 출시된 가중치는 이러한 테스트 프레임 중 72개를 학습 세트에 포함했던 이전 분할로 학습되었습니다.
데이터셋 YAML#
데이터셋 구성을 정의하기 위해 YAML(Yet Another Markup Language) 파일이 사용됩니다. 이 파일에는 데이터셋의 경로, 클래스, 그리고 최대 깊이와 같은 기타 관련 정보가 포함되어 있습니다.
# Ultralytics 🚀 AGPL-3.0 License - https://ultralytics.com/license
# KITTI dataset for monocular depth estimation — real outdoor driving, Velodyne HDL-64 LiDAR (densified), sparse, up to ~80 m
# Documentation: https://docs.ultralytics.com/datasets/depth/kitti
# Example usage: yolo depth train data=depth-kitti.yaml model=yolo26n-depth.pt
# parent
# ├── ultralytics
# └── datasets
# └── depth-kitti ← downloads here (~190 GB archives, ~140 GB converted)
# ├── images/{train,val} # RGB images
# └── depth/{train,val} # paired *.npy, float32 meters (images/ -> depth/)
# If an interrupted download leaves a partial dataset, delete the depth-kitti dir and re-run to rebuild it.
# The download only runs when the val images are missing (train is never checked), so a depth-kitti dir
# built before this split fix keeps the old split until it is deleted and rebuilt.
path: depth-kitti # dataset root dir (relative to Ultralytics settings 'datasets_dir')
train: images/train # train images (relative to 'path') 55198 images
val: images/val # val images (relative to 'path') 652 images (KITTI Eigen test split, left camera)
max_depth: 80 # (m) maximum valid depth; GT beyond this is excluded from val metrics
nc: 1
names:
0: depth
channels: 3
# Download script/URL (optional)
download: |
import shutil
from pathlib import Path
import cv2
import numpy as np
from ultralytics.utils import TQDM
from ultralytics.utils.downloads import download
# The 28 KITTI Eigen test scenes as "<drive> <frame indices>" -> val; every other annotated drive -> train. val is
# the 697-frame Eigen et al. (NeurIPS 2014) test split restricted to frames with improved GT here: 652 left-camera.
EIGEN_TEST = """
2011_09_26_drive_0002 6 9 12 15 18 21 24 27 30 33 36 39 42 45 48 51 54 57 60 63 69
2011_09_26_drive_0009 16 32 48 64 80 96 112 128 144 160 176 196 212 228 244 260 276 292 308 324 340 356 372 388
2011_09_26_drive_0013 5 10 20 30 35 40 45 50 60 65 70 80 85 90 95 100 105 110 115 120 125 130 135
2011_09_26_drive_0020 6 9 12 15 18 21 27 33 36 42 45 48 51 54 57 60 63 66 69 72 75 78
2011_09_26_drive_0023 18 36 54 72 90 108 126 144 162 180 198 216 252 270 288 306 324 342 360 378 396 414 432 450 468
2011_09_26_drive_0027 7 14 21 28 35 42 49 56 63 70 77 84 91 98 112 119 126 133 140 147 161 168 175 182
2011_09_26_drive_0029 14 28 42 56 70 84 98 112 126 140 154 168 182 196 268 296 310 324 338 352 366 380 394 408
2011_09_26_drive_0036 32 64 96 128 160 192 224 256 288 320 352 384 416 448 480 512 544 576 608 640 672 704 768
2011_09_26_drive_0046 5 10 15 20 25 30 35 40 45 50 55 60 65 70 75 80 85 90 95 100 105 110 115
2011_09_26_drive_0048 5 6 7 8 9 10 11 12 13 14 15 16
2011_09_26_drive_0052 6 8 10 12 14 16 18 20 22 26 28 30 32 36 38 40 42 44 46 48 50 52
2011_09_26_drive_0056 11 22 33 44 55 66 77 88 99 110 121 143 154 165 176 187 198 209 220 231 242 253 264 275 286
2011_09_26_drive_0059 14 28 42 56 70 84 98 112 126 140 154 182 196 210 224 238 274 288 302 316 330 344 358
2011_09_26_drive_0064 22 44 66 88 110 154 176 198 220 242 264 286 308 330 352 374 396 418 440 462 484 506 528 550
2011_09_26_drive_0084 49 62 75 88 101 114 127 140 153 179 192 205 218 231 244 257 270 283 296 309 322 335 348 361 374
2011_09_26_drive_0086 7 34 61 88 115 142 169 196 223 250 277 304 331 358 385 412 439 466 493 520 574 601 628 655 682
2011_09_26_drive_0093 16 32 48 64 80 96 112 128 144 160 176 192 208 224 240 256 305 321 337 353 369 385 401 417
2011_09_26_drive_0096 19 38 57 76 95 114 133 152 171 190 209 228 247 266 285 304 323 342 361 380 399 418 437 456
2011_09_26_drive_0101 80 114 148 182 216 284 318 352 386 420 454 488 522 556 590 624 658 692 726 760 794 828 862 896 930
2011_09_26_drive_0106 15 35 43 51 59 67 75 83 91 99 107 115 123 131 139 147 155 163 171 179 187 195 203 211 219
2011_09_26_drive_0117 26 52 78 104 130 156 182 208 234 260 286 312 338 364 390 416 442 468 494 546 572 598 624 650
2011_09_28_drive_0002 6 9 12 18 21 24 30 33 36 39 45 48 51 54 57 60 63 72 78 81 84 87 90
2011_09_29_drive_0071 36 72 108 144 180 216 252 288 324 360 396 432 468 504 540 576 612 735 771 807 879 915 951
2011_09_30_drive_0016 11 22 33 44 55 66 77 88 110 121 132 143 154 165 176 187 198 209 220 231 242 253 264
2011_09_30_drive_0018 107 214 428 535 642 749 856 963 1070 1177 1284 1391 1498 1605 1712 1819 1926 2033 2140 2247 2419 2526 2633 2740
2011_09_30_drive_0027 41 82 123 164 205 246 287 328 369 410 451 492 533 574 615 656 753 794 835 876 917 958 1040 1081
2011_10_03_drive_0027 181 362 543 734 915 1096 1277 1458 1639 1820 2001 2363 2544 2725 2906 3087 3268 3449 3630 3811 3992 4173 4354 4535
2011_10_03_drive_0047 32 64 96 128 160 192 224 288 320 352 384 416 448 480 512 544 576 608 640 672 704 736 768 800"""
EIGEN_TEST = {ln.split()[0]: {int(i) for i in ln.split()[1:]} for ln in EIGEN_TEST.strip().splitlines()}
# Annotated drives excluded from the release training set (raw data was unavailable at build time)
SKIP_DRIVES = {"2011_09_28_drive_0104", "2011_09_28_drive_0135", "2011_09_28_drive_0177", "2011_09_28_drive_0214"}
# Download the improved sparse GT (~14 GB) and the raw recordings it covers (~175 GB)
dir = Path(yaml["path"]) # dataset root dir
base = "https://s3.eu-central-1.amazonaws.com/avg-kitti"
download([f"{base}/data_depth_annotated.zip"], dir=dir / "source", delete=True, exist_ok=True)
gt_drives = sorted(p for p in (dir / "source" / "data_depth_annotated").glob("*/*_sync") if p.name[:-5] not in SKIP_DRIVES)
urls = [f"{base}/raw_data/{p.name[:-5]}/{p.name}.zip" for p in gt_drives]
download(urls, dir=dir / "source" / "raw", delete=True, exist_ok=True, threads=4)
# Convert: GT PNGs are uint16 meters*256; RGB frames come from the matching raw drive
for split in ("train", "val"):
(dir / "images" / split).mkdir(parents=True, exist_ok=True)
(dir / "depth" / split).mkdir(parents=True, exist_ok=True)
for gt in TQDM(gt_drives, desc="Converting"):
drive = gt.name[:-5] # 2011_09_26_drive_0002
split = "val" if drive in EIGEN_TEST else "train"
# val is the canonical Eigen test split: left camera only, and only its 652 listed frames
for cam in ("image_02",) if split == "val" else ("image_02", "image_03"):
for png in sorted((gt / "proj_depth" / "groundtruth" / cam).glob("*.png")):
if split == "val" and int(png.stem) not in EIGEN_TEST[drive]:
continue
name = f"{drive}_{cam[-2:]}_{png.stem}"
depth = cv2.imread(str(png), cv2.IMREAD_ANYDEPTH).astype(np.float32) / 256.0
np.save(dir / "depth" / split / f"{name}.npy", depth)
raw = dir / "source" / "raw" / drive[:10] / gt.name / cam / "data" / png.name
raw.replace(dir / "images" / split / f"{name}.png")
shutil.rmtree(dir / "source")사용법#
KITTI 데이터셋에서 YOLO26n-Depth 모델을 100 epochs 동안 훈련하려면 다음 코드 조각을 사용할 수 있습니다. 사용 가능한 인수의 전체 목록은 모델 Training 페이지를 참조하세요.
from ultralytics import YOLO
# Load a pretrained depth model
model = YOLO("yolo26n-depth.pt")
# Train the model on KITTI
results = model.train(data="depth-kitti.yaml", epochs=100, imgsz=640)사전 학습된 모델#
사전 학습된 YOLO26-Depth 모델은 이름으로 처음 참조될 때 Ultralytics v8.4.0 assets release에서 자동으로 다운로드됩니다:
인용 및 감사의 글#
연구 또는 개발 작업에서 KITTI 데이터셋을 사용하는 경우, 다음 논문을 인용해 주십시오:
@article{geiger2013vision,
title={Vision meets Robotics: The KITTI Dataset},
author={Geiger, Andreas and Lenz, Philip and Stiller, Christoph and Urtasun, Raquel},
journal={The International Journal of Robotics Research},
year={2013},
publisher={SAGE Publications}
}
@inproceedings{uhrig2017sparsity,
title={Sparsity Invariant CNNs},
author={Uhrig, Jonas and Schneider, Nick and Schneider, Lukas and Franke, Uwe and Brox, Thomas and Geiger, Andreas},
booktitle={International Conference on 3D Vision (3DV)},
year={2017}
}KITTI 데이터셋을 생성하고 유지 관리해 준 Karlsruhe Institute of Technology와 Toyota Technological Institute at Chicago, 그리고 밀도 높은 지도가 가능하도록 깊이 밀도화 방법을 제공한 Uhrig et al.에게 감사를 표합니다.