Hypersim Depth 데이터셋#
Hypersim은 종합적인 실내 장면 이해를 위해 설계된 사실적인 실내 장면 합성 데이터셋입니다. 전문적으로 제작된 Evermotion 3D 인테리어를 레이 트레이싱해 이미지를 생성하므로, 매우 사실적인 렌더링과 완벽하고 조밀한 픽셀 단위 depth 정답 데이터를 함께 제공합니다.
Hypersim은 완전한 합성 데이터셋이므로 모든 픽셀에 정확한 depth 값이 있으며, 센서 노이즈나 측정값 누락, 측정 아티팩트가 없습니다. 따라서 단안 depth estimation 모델을 학습할 때 깨끗하고 조밀한 실내 기하 정보를 제공하는 탁월한 데이터 소스입니다.
주요 기능#
- 전문가가 제작한 Evermotion 3D 인테리어를 레이 트레이싱한 사실적인 합성 실내 장면입니다.
- 실내 환경만 포함합니다.
- 센서 노이즈나 누락된 값이 없는 조밀하고 정확한 픽셀 단위 depth 정답 데이터입니다.
- depth 범위는 대부분 ≤10 m이며, 더 먼 거리도 일부 포함합니다.
- Ultralytics depth 학습 데이터 구성에 이미지 74,619장(학습 68,242장 / 검증 6,377장)을 제공합니다.
데이터셋 구조#
Hypersim depth 데이터셋은 두 개의 하위 집합으로 나뉩니다:
- Train: 학습용 depth 맵이 함께 제공되는 이미지 68,242장입니다.
- Val: 학습 중 검증에 사용하는 depth 맵이 함께 제공되는 이미지 6,377장입니다.
각 RGB 이미지에는 크기가 조정된 uint16 depth PNG가 연결되어 있으며, Ultralytics depth dataset format을 따릅니다.
데이터 가져오기#
Hypersim은 자동 다운로드를 지원하지 않습니다. 장면별 ZIP 파일 약 ~1.9 TB로 구성된 데이터셋은 Apple이 CC BY-SA 3.0 license에 따라 배포하며, ml-hypersim repository의 공식 스크립트로 다운로드합니다(선택적 다운로드 도구는 contrib/99991에서 사용할 수 있습니다):
git clone https://github.com/apple/ml-hypersim && cd ml-hypersim
python code/python/tools/dataset_download_images.py --downloads_dir ./downloads --decompress_dir ./scenesRGB 프레임은 frame.*.tonemap.jpg 미리보기이며 depth 데이터는 frame.*.depth_meters.hdf5에 있습니다. 저장된 값은 평면 depth가 아니라 카메라 중심까지의 레이 거리입니다. 저장하기 전에 변환하고 NaN 픽셀(창문, 하늘)은 0(유효하지 않은 값)으로 매핑하세요. 학습/검증 할당에는 공식 metadata_images_split_scene_v1.csv 장면 분할을 사용하세요. Ultralytics depth dataset format으로 변환하는 예시는 다음과 같습니다:
import shutil
from pathlib import Path
import h5py
import numpy as np
from ultralytics.data.utils import save_depth_png
W, H, FOCAL = 1024, 768, 886.81 # Hypersim camera intrinsics
x, y = np.meshgrid(np.linspace(-W / 2, W / 2, W), np.linspace(-H / 2, H / 2, H))
ray2plane = (FOCAL / np.sqrt(x**2 + y**2 + FOCAL**2)).astype(np.float32) # distance → planar depth
src, dst = Path("scenes"), Path("datasets/depth-hypersim")
out = "train" # assign per scene from metadata_images_split_scene_v1.csv
(dst / f"images/{out}").mkdir(parents=True, exist_ok=True)
(dst / f"depth/{out}").mkdir(parents=True, exist_ok=True)
for h5 in sorted(src.rglob("*.depth_meters.hdf5")):
scene, cam, frame = h5.parts[-4], h5.parent.name.split("_geometry")[0], h5.name.split(".")[1]
rgb = h5.parents[1] / f"{cam}_final_preview" / f"frame.{frame}.tonemap.jpg"
dist = np.asarray(h5py.File(h5)["dataset"], np.float32)
depth = np.nan_to_num(dist * ray2plane, nan=0.0) # NaN (sky, glass) → 0 = invalid
name = f"{scene}_{cam}_{frame}"
save_depth_png(dst / f"depth/{out}/{name}.png", depth)
shutil.copy(rgb, dst / f"images/{out}/{name}.jpg")YOLO26-Depth에서의 역할#
Hypersim은 Ultralytics YOLO26-Depth 모델 사전 학습에 사용되는 광범위한 다중 데이터셋 혼합 데이터(약 ~2.19M 이미지)의 학습 소스 중 하나입니다. 이 데이터 구성에서 Hypersim은 노이즈가 더 많은 실제 센서 데이터와 상호 보완되는 깨끗하고 조밀한 실내 기하 정보를 제공합니다.
이 구성에는 별도의 홀드아웃 Hypersim 벤치마크가 없습니다. 대신 학습된 모델은 NYU Depth V2, KITTI, Make3D, ETH3D, iBims-1 등 표준 단안 depth 벤치마크에서 평가됩니다.
데이터셋 YAML#
데이터셋 구성을 정의할 때 YAML 파일을 사용합니다. 이 파일에는 데이터셋 경로, 클래스 및 기타 관련 정보가 포함됩니다. Hypersim의 경우 depth-hypersim.yaml 파일에서 경로와 단일 depth 클래스를 정의합니다.
# Ultralytics 🚀 AGPL-3.0 License - https://ultralytics.com/license
# Hypersim dataset for monocular depth estimation — photorealistic synthetic indoor (ray-traced), dense depth up to ~10 m
# Documentation: https://docs.ultralytics.com/datasets/depth/hypersim
# Example usage: yolo depth train data=depth-hypersim.yaml model=yolo26n-depth.pt
# No autodownload — obtain the source data (see docs) and arrange it as below.
# parent
# ├── ultralytics
# └── datasets
# └── depth-hypersim
# ├── images/{train,val} # RGB images
# └── depth/{train,val} # paired 16-bit *.png depth maps (images/ -> depth/)
path: depth-hypersim # dataset root dir (relative to Ultralytics settings 'datasets_dir')
train: images/train # train images (relative to 'path') 68242 images
val: images/val # val images (relative to 'path') 6377 images
nc: 1
names:
0: depth
channels: 3사용법#
이미지 크기 640으로 Hypersim 데이터셋에서 YOLO26n-Depth 모델을 학습하려면 다음 코드 예제를 사용할 수 있습니다. 사용 가능한 인수의 전체 목록은 모델 Training 페이지를 참조하세요.
from ultralytics import YOLO
# 모델 로드
model = YOLO("yolo26n-depth.pt") # 사전 학습된 모델을 불러옵니다(학습에 권장).
# 모델 학습
results = model.train(data="depth-hypersim.yaml", epochs=100, imgsz=640)사전 학습 모델#
YOLO26 depth 제품군(yolo26n-depth.pt, yolo26s-depth.pt, yolo26m-depth.pt, yolo26l-depth.pt, yolo26x-depth.pt)은 Ultralytics 릴리스에서 자동으로 다운로드되며, Hypersim이 포함된 광범위한 다중 데이터셋 혼합 데이터로 학습됩니다. Ultralytics Platform에서 YOLO26x-depth를 살펴보세요.
인용 및 감사의 글#
연구 또는 개발 작업에서 Hypersim 데이터셋을 사용하는 경우 다음 논문을 인용해 주세요:
@inproceedings{roberts2021hypersim,
title={Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene Understanding},
author={Mike Roberts and Jason Ramapuram and Anurag Ranjan and Atulit Kumar and Miguel Angel Bautista and Nathan Paczan and Russ Webb and Joshua M. Susskind},
booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)},
year={2021}
}컴퓨터 비전 커뮤니티에서 사실적인 합성 실내 데이터셋 Hypersim을 사용할 수 있도록 공개해 주신 제작자 여러분께 감사드립니다.
자주 묻는 질문#
Hypersim은 Apple이 제공하는 사실적인 합성 실내 데이터셋으로, 전문적으로 제작된 Evermotion 3D 인테리어를 레이 트레이싱해 생성했습니다. 모든 픽셀에 정확한 depth 값이 있고 센서 노이즈나 측정값 누락이 없으므로, 이미지 74,619장(학습 68,242장, 검증 6,377장)이 YOLO26-Depth 학습 데이터 구성의 노이즈가 더 많은 실제 센서 데이터와 상호 보완되는 깨끗하고 조밀한 실내 기하 정보를 제공합니다.
Hypersim은 자동 다운로드를 지원하지 않습니다. ml-hypersim repository의 공식 스크립트를 사용해 장면 데이터(약 ~1.9 TB)를 가져온 다음, Obtain the Data에 있는 스크립트로
depth_meters.hdf5레이 거리를 평면 depth로 변환하고 밀리미터 단위 PNG 파일로 저장하세요. 창문과 하늘처럼 NaN인 픽셀은0으로 매핑해 유효하지 않은 값으로 처리하세요.