Link to this sectionTập dữ liệu độ sâu KITTI#
Tập dữ liệu KITTI là một bộ tiêu chuẩn lái xe tự động ngoài trời trong thế giới thực, được ghi lại từ một phương tiện đang di chuyển trong và xung quanh thành phố Karlsruhe. Đối với ước tính độ sâu đơn ảnh, độ sâu thực tế (ground-truth) được thu thập từ máy quét LiDAR Velodyne HDL-64 và được làm dày bằng phương pháp của Uhrig et al. 2017. Các bản đồ độ sâu kết quả vẫn ở dạng thưa thớt, với khoảng 16–20% pixel mang giá trị độ sâu hợp lệ. KITTI là nguồn dữ liệu ngoài trời tầm xa thực tế duy nhất trong quá trình tiền huấn luyện YOLO26-Depth và cũng đóng vai trò là tiêu chuẩn đánh giá KITTI Eigen.
Link to this sectionTính năng chính#
- Các cảnh lái xe ngoài trời thực tế với độ sâu lên tới khoảng 80 m, xa hơn nhiều so với phạm vi trong nhà thông thường.
- Độ sâu thực tế thu được từ LiDAR Velodyne HDL-64 và được làm dày bằng phương pháp Sparsity Invariant CNNs của Uhrig et al. 2017.
- Giám sát thưa thớt: chỉ khoảng 16–20% pixel trên mỗi hình ảnh mang giá trị độ sâu hợp lệ; các pixel không hợp lệ sẽ bị loại bỏ khỏi hàm loss và các số liệu đánh giá.
- Các cặp hình ảnh stereo (trái
image_02và phảiimage_03) cung cấp các góc nhìn bổ sung để huấn luyện. - Các giá trị độ sâu được lưu trữ dưới dạng mảng float32
.npytheo đơn vị mét, tuân theo định dạng tập dữ liệu độ sâu Ultralytics.
Link to this sectionCấu trúc tập dữ liệu#
Dữ liệu độ sâu KITTI được Ultralytics sử dụng được chia thành hai tập con:
- Tập huấn luyện (Training split): 60.040 hình ảnh (trái
image_02và phảiimage_03). 28 lần lái thử KITTI Eigen được loại trừ khỏi quá trình huấn luyện để đảm bảo sự công bằng khi đánh giá. - Tập đánh giá (Evaluation split): tập kiểm thử KITTI Eigen, gồm 32.378 hình ảnh (cả hai camera). Quá trình đánh giá sử dụng phương pháp cắt Garg, giới hạn độ sâu 80 m và căn chỉnh log-least-squares giữa các dự đoán và giá trị thực tế.
Phạm vi độ sâu đạt khoảng 80 m và tệp YAML của tập dữ liệu (depth-kitti.yaml) thiết lập max_depth: 80 cho phù hợp.
Link to this sectionVai trò trong YOLO26-Depth#
KITTI cung cấp nguồn giám sát tầm xa, ngoài trời thực tế duy nhất trong quá trình tiền huấn luyện YOLO26-Depth, bổ sung cho các nguồn dữ liệu chủ yếu ở trong nhà. Đây cũng là tiêu chuẩn đánh giá KITTI Eigen dùng để báo cáo độ chính xác về độ sâu trong cảnh lái xe.
KITTI là ví dụ chính cho thấy lý do tại sao đầu ra độ sâu lại không bị giới hạn (chế độ log): ngưỡng đầu ra cố định 10 m không thể biểu diễn các cảnh lái xe 80 m. Xem trang tác vụ độ sâu để biết chi tiết về phạm vi đầu ra và cách xử lý max_depth.
Link to this sectionKết quả#
Độ chính xác KITTI Eigen delta1 theo kích thước mô hình (càng cao càng tốt):
| Mô hình | KITTI Eigen δ1 |
|---|---|
| YOLO26n-Depth | 0.888 |
| YOLO26s-Depth | 0.882 |
| YOLO26m-Depth | 0.924 |
| YOLO26l-Depth | 0.927 |
| YOLO26x-Depth | 0.939 |
Link to this sectionYAML tập dữ liệu#
Tệp YAML (Yet Another Markup Language) được sử dụng để xác định cấu hình tập dữ liệu. Tệp này chứa thông tin về đường dẫn tập dữ liệu, các lớp (classes) và các thông tin liên quan khác như độ sâu tối đa.
# Ultralytics 🚀 AGPL-3.0 License - https://ultralytics.com/license
# KITTI dataset for monocular depth estimation — real outdoor driving, Velodyne HDL-64 LiDAR (densified), sparse, up to ~80 m
# Documentation: https://docs.ultralytics.com/datasets/depth/kitti
# Example usage: yolo depth train data=depth-kitti.yaml model=yolo26n-depth.pt
# parent
# ├── ultralytics
# └── datasets
# └── depth-kitti ← downloads here (~190 GB archives, ~230 GB converted)
# ├── images/{train,val} # RGB images
# └── depth/{train,val} # paired *.npy, float32 meters (images/ -> depth/)
# If an interrupted download leaves a partial dataset, delete the depth-kitti dir and re-run to rebuild it.
path: depth-kitti # dataset root dir (relative to Ultralytics settings 'datasets_dir')
train: images/train # train images (relative to 'path') 60040 images
val: images/val # val images (relative to 'path') 32378 images (KITTI Eigen test split)
max_depth: 80 # (m) maximum valid depth; GT beyond this is excluded from val metrics
nc: 1
names:
0: depth
channels: 3
# Download script/URL (optional)
download: |
import shutil
from pathlib import Path
import cv2
import numpy as np
from ultralytics.utils import TQDM
from ultralytics.utils.downloads import download
# Drives of the 28 KITTI Eigen test scenes -> val; every other annotated drive -> train
VAL_DRIVES = """2011_09_26_drive_0002 2011_09_26_drive_0009 2011_09_26_drive_0013 2011_09_26_drive_0020
2011_09_26_drive_0023 2011_09_26_drive_0027 2011_09_26_drive_0029 2011_09_26_drive_0036 2011_09_26_drive_0046
2011_09_26_drive_0048 2011_09_26_drive_0052 2011_09_26_drive_0056 2011_09_26_drive_0059 2011_09_26_drive_0064
2011_09_26_drive_0084 2011_09_26_drive_0086 2011_09_26_drive_0093 2011_09_26_drive_0096 2011_09_26_drive_0101
2011_09_26_drive_0106 2011_09_26_drive_0117 2011_09_28_drive_0002 2011_09_29_drive_0026 2011_09_29_drive_0071
2011_09_30_drive_0016 2011_10_03_drive_0034 2011_10_03_drive_0042 2011_10_03_drive_0047""".split()
# Annotated drives excluded from the release training set (raw data was unavailable at build time)
SKIP_DRIVES = {"2011_09_28_drive_0104", "2011_09_28_drive_0135", "2011_09_28_drive_0177", "2011_09_28_drive_0214"}
# Download the improved sparse GT (~14 GB) and the raw recordings it covers (~175 GB)
dir = Path(yaml["path"]) # dataset root dir
base = "https://s3.eu-central-1.amazonaws.com/avg-kitti"
download([f"{base}/data_depth_annotated.zip"], dir=dir / "source", delete=True, exist_ok=True)
gt_drives = sorted(p for p in (dir / "source" / "data_depth_annotated").glob("*/*_sync") if p.name[:-5] not in SKIP_DRIVES)
urls = [f"{base}/raw_data/{p.name[:-5]}/{p.name}.zip" for p in gt_drives]
download(urls, dir=dir / "source" / "raw", delete=True, exist_ok=True, threads=4)
# Convert: GT PNGs are uint16 meters*256; RGB frames come from the matching raw drive
for split in ("train", "val"):
(dir / "images" / split).mkdir(parents=True, exist_ok=True)
(dir / "depth" / split).mkdir(parents=True, exist_ok=True)
for gt in TQDM(gt_drives, desc="Converting"):
drive = gt.name[:-5] # 2011_09_26_drive_0002
split = "val" if drive in VAL_DRIVES else "train"
for cam in ("image_02", "image_03"):
for png in sorted((gt / "proj_depth" / "groundtruth" / cam).glob("*.png")):
name = f"{drive}_{cam[-2:]}_{png.stem}"
depth = cv2.imread(str(png), cv2.IMREAD_ANYDEPTH).astype(np.float32) / 256.0
np.save(dir / "depth" / split / f"{name}.npy", depth)
raw = dir / "source" / "raw" / drive[:10] / gt.name / cam / "data" / png.name
raw.replace(dir / "images" / split / f"{name}.png")
shutil.rmtree(dir / "source")Link to this sectionCách sử dụng#
Để huấn luyện mô hình YOLO26n-Depth trên tập dữ liệu KITTI trong 100 epoch, bạn có thể sử dụng các đoạn mã sau. Để xem danh sách đầy đủ các đối số khả dụng, hãy tham khảo trang Huấn luyện mô hình.
from ultralytics import YOLO
# Load a pretrained depth model
model = YOLO("yolo26n-depth.pt")
# Train the model on KITTI
results = model.train(data="depth-kitti.yaml", epochs=100, imgsz=640)Link to this sectionCác model đã được tiền huấn luyện#
Các mô hình YOLO26-Depth đã tiền huấn luyện tự động tải xuống từ bản phát hành tài nguyên v8.4.0 của Ultralytics khi được tham chiếu theo tên lần đầu tiên:
Link to this sectionTrích dẫn và Ghi nhận#
Nếu bạn sử dụng tập dữ liệu KITTI trong công việc nghiên cứu hoặc phát triển của mình, vui lòng trích dẫn các bài báo sau:
@article{geiger2013vision,
title={Vision meets Robotics: The KITTI Dataset},
author={Geiger, Andreas and Lenz, Philip and Stiller, Christoph and Urtasun, Raquel},
journal={The International Journal of Robotics Research},
year={2013},
publisher={SAGE Publications}
}
@inproceedings{uhrig2017sparsity,
title={Sparsity Invariant CNNs},
author={Uhrig, Jonas and Schneider, Nick and Schneider, Lukas and Franke, Uwe and Brox, Thomas and Geiger, Andreas},
booktitle={International Conference on 3D Vision (3DV)},
year={2017}
}Chúng tôi xin ghi nhận Viện Công nghệ Karlsruhe và Viện Công nghệ Toyota tại Chicago vì đã tạo ra và duy trì tập dữ liệu KITTI, cũng như Uhrig et al. vì phương pháp làm dày độ sâu giúp việc giám sát mật độ trở nên khả thi.