Enterprise-ready security: ISO 27001 + SOC 2 Type I compliant.

Link to this sectionMonocular Depth Estimation#

Monocular depth estimation examples

Monocular depth estimation predicts a per-pixel depth map from a single RGB image. Each output pixel holds a depth value in meters representing the estimated distance from the camera to that surface point.

The output of a depth model is a dense float map of shape (H, W) aligned to the input image. This per-pixel representation makes monocular depth estimation well-suited for 3D scene reconstruction, robot navigation, AR/VR content creation, and any application that requires spatial layout from a single camera.

Tip

Use task=depth or the yolo depth CLI task for monocular depth estimation. YOLO26 depth model files use the -depth suffix, such as yolo26n-depth.pt.

Link to this sectionModels#

YOLO26 depth models pretrained on a broad multi-dataset mix (indoor + outdoor, ~2.19M images) are shown below. The metrics columns are reported on the NYU Depth V2 Eigen test split.

Models download automatically from the latest Ultralytics release on first use.

Modelsize
(pixels)
delta1NYUabs_relNYUrmseNYUparams
(M)
FLOPs
(B)
YOLO26n-depth7680.8820.1090.4146.446.9
YOLO26s-depth7680.8960.1040.39913.267.9
YOLO26m-depth7680.9210.0890.36423.3130.7
YOLO26l-depth7680.9300.0830.35127.7157.2
YOLO26x-depth7680.9330.0800.34457.0302.0
  • delta1NYU is the percentage of pixels where the predicted depth is within a factor of 1.25 of the ground truth, on the NYU Depth V2 Eigen test split (654 images) with multi-scale + horizontal-flip TTA and log-least-squares alignment.
  • Single-scale accuracy without TTA is reproducible with yolo depth val model=yolo26n-depth.pt data=nyu-depth.yaml imgsz=768 device=0 (substitute model= for each size), which uses median (scale-only) alignment and scores lower: delta1 0.785 (n), 0.786 (s), 0.827 (m), 0.839 (l), 0.843 (x).
  • abs_rel is the mean absolute relative error between predicted and ground-truth depth values.
  • rmse is the root mean squared error in meters.
  • params and FLOPs are measured at 768×768, the training resolution of the released weights.

Link to this sectionDepth range and the log-depth head#

The depth head predicts exp(logit)unbounded (~0.02–150 m) — and decouples scene shape from absolute scale: the network predicts a relative log-depth field, and absolute meters are set by a separate two-parameter transform (exp(a·log d + b)) recovered at evaluation, by lightweight calibration, or by fine-tuning. The common alternative, a bounded sigmoid × max_depth head, instead bakes a fixed ceiling into the architecture, so any depth beyond max_depth is clipped — which prevents training on, and predicting, longer-range scenes.

Link to this sectionWhy the head is unbounded: evidence across depth ranges#

Controlled A/B — same data, same schedule, only the head differs. Training both heads from scratch on an identical mix of indoor (≤10 m) and outdoor (≤80 m) data:

HeadOutput rangeMixed-range val δ1Training behavior
sigmoid × max_depth (10 m)bounded 0–10 m0.221plateaus at epoch 1
log (unbounded)0–~150 m0.367 (+66%)improves throughout

The bounded head cannot represent the >10 m majority of outdoor pixels, so it stops improving almost immediately; the log head learns from the full range.

The failure mode it fixes — single-dataset fine-tuning, stratified by range. Fine-tuning a bounded (10 m) head on individual datasets works when the data fits the cap and collapses when it exceeds it:

DatasetDepth rangeBounded-head δ1
SUN RGB-D≤10 m0.78 ✅
Hypersimmostly ≤10 m0.74 ✅
KITTI~80 m0.08 ❌ (abs_rel ≈ 7.0 — the 80/10 scale mismatch)
vKITTI280 m0.04 ❌

The collapse lands exactly at the cap boundary. The log head has no such boundary.

Released models — cross-range performance. The shipped log-head family, evaluated across benchmarks spanning 10 m (NYU, iBims-1) to 80 m (KITTI), with scale-aligned δ1:

BenchmarkRangeYOLO26x-depth (log)prior bounded release
NYU Eigen10 m0.9330.934
iBims-110 m0.9610.945
ETH3D~60 m0.9590.931
Make3D~70 m0.3020.296
KITTI Eigen80 m0.9420.891
Mean0.8190.799

The largest gains are on the longer-range outdoor benchmarks (KITTI, ETH3D) — exactly where a fixed 10 m ceiling hurts most — while indoor performance is retained.

Link to this sectionTrain#

Train YOLO26n-depth on the Depth8 dataset for 100 epochs at image size 640. For a full list of available arguments see the Configuration page.

Example
from ultralytics import YOLO

# Load a model
model = YOLO("yolo26n-depth.yaml")  # build a new model from YAML
model = YOLO("yolo26n-depth.pt")  # load a pretrained model (recommended for training)
model = YOLO("yolo26n-depth.yaml").load("yolo26n-depth.pt")  # build from YAML and transfer weights

# Train the model
results = model.train(data="depth8.yaml", epochs=100, imgsz=640)

See full train mode details in the Train page.

Link to this sectionFine-tuning on your own data#

When adapting a pretrained depth model to a custom dataset, lower the learning rate and use the AdamW optimizer. The default optimizer settings are tuned for training from scratch (SGD with lr0=0.01); applied to an already-converged depth model they can overwrite the pretrained knowledge and degrade results — especially when fine-tuning on a single domain.

Recommended fine-tuning recipe
from ultralytics import YOLO

model = YOLO("yolo26s-depth.pt")  # start from pretrained weights

model.train(
    data="path/to/your_dataset.yaml",
    epochs=20,
    imgsz=640,
    optimizer="AdamW",
    lr0=1e-4,  # ~100x below the from-scratch default
    warmup_bias_lr=1e-4,  # keep warmup gentle too
)

Additional tips:

  • Augmentation is controlled by the standard args (degrees, translate, scale, shear, perspective, flipud, fliplr, hsv_h, hsv_s, hsv_v); the geometric warp and flips are applied identically to the paired depth map. To see the exact recipe used for the released YOLO26-Depth weights, inspect the train_args stored in the checkpoint — see Inspecting YOLO26 Checkpoint Training Args.
  • mosaic, mixup, cutmix, and copy_paste are not implemented for depth. The depth dataset loader automatically sets these probabilities to 0, so passing them has no effect. These augmentations are not supported because they combine multiple images, which would produce invalid paired depth maps.
  • Any depth range works out of the box. The log-head models predict unbounded depth, so they adapt to short-range (macro) or long-range (outdoor/driving) data without changes. Setting max_depth: in your dataset YAML (in meters) bounds which GT pixels count toward validation metrics.
  • Retain general performance. If you need the model to stay accurate on scenes beyond your training set, mix a small fraction (~5–10%) of diverse general-purpose images into your training data; this substantially reduces forgetting during fine-tuning.
  • Train from scratch (model=yolo26s-depth.yaml) only if your domain is very different and you have a large dataset — there the default SGD lr0=0.01 is appropriate, since there are no pretrained weights to preserve.

Link to this sectionCalibrating the depth scale#

The depth head separates shape (relative scene structure) from scale (absolute meters). If a model already produces good relative depth on your scenes but the absolute values are off for your camera, you can correct the scale in seconds with model.calibrate() — a closed-form fit of a two-parameter log-affine against a small labeled set, with no gradient training and no change to the network weights, so it cannot degrade the relative structure.

Scale calibration
from ultralytics import YOLO

model = YOLO("yolo26s-depth.pt")

# Fit scale on a labeled split (~100+ images is enough), then persist it
model.calibrate(data="path/to/your_dataset.yaml")
model.save("yolo26s-depth-calibrated.pt")

Calibration needs ground-truth depth to fit against, so it runs on a labeled split — it is not something that can happen at blind inference. Use it when relative depth is already good and only the scale/range is wrong; if the relative structure itself needs to change for your domain, fine-tune instead.

Training does this for you automatically: after model.train(...) completes, the best and last checkpoints are calibrated on the validation set so they output metric-scaled depth out of the box.

The released yolo26*-depth.pt checkpoints ship with this calibration already baked in, fit on the pretraining validation mix. It is a single global scale across all domains, so for the most accurate absolute depth on a specific camera or scene type, run model.calibrate() on a small labeled split from your own data — it replaces the baked-in fit.

Link to this sectionDataset format#

Depth estimation datasets pair each RGB image with a corresponding depth file. Depth targets are stored as .npy arrays containing float32 values in meters. The dataset YAML points to an images/ directory; the loader derives the depth file path by replacing the images component with depth and replacing the image extension with .npy.

dataset/
├── images/
│   ├── train/
│   └── val/
└── depth/
    ├── train/
    └── val/

For example, an image at images/train/scene_001.jpg is paired with a depth map at depth/train/scene_001.npy. See the Depth Estimation Dataset Guide for the full format specification.

Link to this sectionVal#

Validate a trained YOLO26n-depth model accuracy on a depth estimation dataset. Pass data explicitly so validation uses the intended dataset YAML. The released weights are trained at imgsz=768, so validate and predict at that size for best accuracy.

Example
from ultralytics import YOLO

# Load a model
model = YOLO("yolo26n-depth.pt")  # load an official model
model = YOLO("path/to/best.pt")  # load a custom model

# Validate the model
metrics = model.val(data="nyu-depth.yaml")
metrics.delta1  # percentage of pixels within threshold δ=1.25
metrics.abs_rel  # mean absolute relative error
metrics.rmse  # root mean squared error (meters)
metrics.silog  # scale-invariant logarithmic error

Link to this sectionPredict#

Use a trained YOLO26n-depth model to run predictions on images.

Example
from ultralytics import YOLO

# Load a model
model = YOLO("yolo26n-depth.pt")  # load an official model
model = YOLO("path/to/best.pt")  # load a custom model

# Predict with the model
results = model("https://ultralytics.com/images/bus.jpg")  # predict on an image

# Access the results
for result in results:
    depth_map = result.depth.data.cpu().numpy()  # torch.Tensor -> NumPy float32, shape (H, W), meters

See full predict mode details in the Predict page.

Link to this sectionResults Output#

YOLO depth estimation returns one Results object per image. Each result stores one dense float depth map for the full image.

AttributeTypeShapeDescription
result.depthDepthMap(H,W)Dense per-pixel depth map.
result.depth.datatorch.Tensor(H,W)Depth values in meters; call .cpu().numpy() for NumPy.
result.boxes--No instance boxes.
result.masks--No instance masks.

For task-specific Results fields across every task, see the Predict Results by Task section.

Link to this sectionExport#

Export a YOLO26n-depth model to a different format like ONNX, CoreML, etc.

Example
from ultralytics import YOLO

# Load a model
model = YOLO("yolo26n-depth.pt")  # load an official model
model = YOLO("path/to/best.pt")  # load a custom model

# Export the model
model.export(format="onnx")

Available YOLO26 depth estimation export formats are in the table below. You can export to any format using the format argument, i.e., format='onnx' or format='engine'. You can predict or validate directly on exported models, i.e., yolo predict model=yolo26n-depth.onnx. Usage examples are shown for your model after export completes.

Formatformat ArgumentModelMetadataArguments
PyTorch-yolo26n-depth.pt-
TorchScripttorchscriptyolo26n-depth.torchscriptimgsz, quantize, dynamic, nms, batch, device
ONNXonnxyolo26n-depth.onnximgsz, quantize, dynamic, simplify, opset, nms, batch, data, fraction, device
OpenVINOopenvinoyolo26n-depth_openvino_model/imgsz, quantize, dynamic, nms, batch, data, fraction, device
TensorRTengineyolo26n-depth.engineimgsz, quantize, dynamic, simplify, opset, workspace, nms, batch, data, fraction, device
CoreMLcoremlyolo26n-depth.mlpackageimgsz, dynamic, quantize, nms, batch, device
TF SavedModelsaved_modelyolo26n-depth_saved_model/imgsz, keras, quantize, nms, batch, data, fraction, device
TF GraphDefpbyolo26n-depth.pbimgsz, batch, device
TF Edge TPUedgetpuyolo26n-depth_edgetpu.tfliteimgsz, quantize, data, fraction, device
PaddlePaddlepaddleyolo26n-depth_paddle_model/imgsz, batch, device
MNNmnnyolo26n-depth.mnnimgsz, batch, dynamic, quantize, nms, device
NCNNncnnyolo26n-depth_ncnn_model/imgsz, quantize, batch, device
IMX500imxyolo26n-depth_imx_model/imgsz, quantize, data, fraction, nms, device
RKNNrknnyolo26n-depth_rknn_model/imgsz, batch, name, quantize, data, fraction, device
ExecuTorchexecutorchyolo26n-depth_executorch_model/imgsz, batch, device
Axeleraaxelerayolo26n-depth_axelera_model/imgsz, batch, quantize, data, fraction, device
DEEPXdeepxyolo26n-depth_deepx_model/imgsz, quantize, data, optimize, device
Qualcomm QNNqnnyolo26n-depth_qnn.onnximgsz, batch, name, quantize, data, fraction, device
LiteRTlitertyolo26n-depth.tfliteimgsz, quantize, batch, data, fraction, device
Hailohailoyolo26n-depth_hailo_model/imgsz, name, quantize, data, fraction, simplify, conf, iou

See full export details in the Export page.

Link to this sectionFAQ#

Link to this sectionHow do I train a YOLO26 depth estimation model on a custom dataset?#

Prepare paired RGB images and .npy depth files, then create a dataset YAML pointing to your images/ directory. The loader finds depth files automatically by replacing images with depth in the path and swapping the image extension for .npy.

Start from pretrained weights and use a low learning rate with AdamW so the fine-tune retains what the model already knows (see Fine-tuning on your own data for why):

Example
from ultralytics import YOLO

# Load a pretrained YOLO26 depth model
model = YOLO("yolo26s-depth.pt")

# Fine-tune on a custom depth dataset
results = model.train(
    data="path/to/your_dataset.yaml",
    epochs=20,
    imgsz=640,
    optimizer="AdamW",
    lr0=1e-4,
    warmup_bias_lr=1e-4,
)

Check the Configuration page for more available arguments.

Link to this sectionWhat metrics does YOLO26 depth estimation report?#

Depth estimation validation reports the metric set used by Depth Anything and related monocular-depth work:

  • delta1 / delta2 / delta3 — percentage of pixels where the ratio of predicted to ground-truth depth (or its inverse) is below 1.25, 1.25², and 1.25³ respectively. Higher is better.
  • abs_rel — mean absolute relative error. Lower is better.
  • rmse — root mean squared error in meters. Lower is better.
  • silog — scale-invariant logarithmic error. Lower is better.

Each prediction is median-aligned to its ground truth per image, and the statistics are then pooled over every valid pixel of the validation set (images with more valid depth pixels weigh proportionally more). Papers that instead average per-image metrics can report slightly different values on the same predictions.

Link to this sectionWhat is the depth map output format?#

The depth map is stored as a torch.Tensor of shape (H, W) where each value is the predicted depth in meters (use result.depth.data.cpu().numpy() for a NumPy float32 array). Access it through result.depth.data after running prediction. The map is aligned to the input image resolution.

Link to this sectionWhat datasets are supported for depth estimation?#

Ultralytics YOLO26 provides built-in dataset YAML configurations for several depth estimation datasets including NYU Depth V2, KITTI, Hypersim, SUN RGB-D, and ARKitScenes. See the Depth Estimation Dataset Guide for the full list and format details.

Link to this sectionHow do I validate a pretrained YOLO26 depth estimation model?#

Validate a pretrained YOLO26 depth model by supplying the dataset YAML used for evaluation:

Example
from ultralytics import YOLO

# Load a pretrained model
model = YOLO("yolo26n-depth.pt")

# Validate the model
metrics = model.val(data="nyu-depth.yaml")
print("delta1:", metrics.delta1)
print("abs_rel:", metrics.abs_rel)
print("rmse:", metrics.rmse)

Link to this sectionHow can I export a YOLO26 depth estimation model to ONNX format?#

Export a YOLO26 depth model to ONNX with Python or CLI:

Example
from ultralytics import YOLO

# Load a pretrained model
model = YOLO("yolo26n-depth.pt")

# Export the model to ONNX format
model.export(format="onnx")

For more details on exporting to various formats, refer to the Export page.

Contributors

Comments