Ultralytics YOLO26とTriton Inference Server#
Triton Inference Server(旧称TensorRT Inference Server)は、NVIDIAが開発したオープンソースソフトウェアソリューションです。NVIDIA GPU向けに最適化されたクラウド推論ソリューションを提供します。Tritonを使用すると、本番環境でAIモデルを大規模にデプロイしやすくなります。Ultralytics YOLO26をTriton Inference Serverと統合することで、スケーラブルで高性能なディープラーニング推論ワークロードをデプロイできます。このガイドでは、統合のセットアップとテストの手順を説明します。
視聴: NVIDIA Triton Inference Serverを使い始める
Triton Inference Serverとは何ですか?#
Triton Inference Serverは、さまざまなAIモデルを本番環境にデプロイするために設計されています。PyTorch、TensorFlow、ONNX、OpenVINO、TensorRTなど、幅広いディープラーニングおよび機械学習フレームワークをサポートしています。主な用途は次のとおりです。
- 1つのサーバーインスタンスから複数のモデルをサービングする
- サーバーを再起動せずにモデルを動的に読み込み、アンロードする
- 複数のモデルを組み合わせて結果を得るアンサンブル推論
- A/Bテストとローリングアップデートのためのモデルのバージョン管理
Triton Inference Serverの主なメリット#
Triton Inference ServerをUltralytics YOLO26と併用すると、次のような利点があります。
- 自動バッチ処理: 複数のAIリクエストをまとめてから処理することで、レイテンシを短縮し、推論速度を向上させます
- Kubernetesとの統合: クラウドネイティブ設計により、Kubernetesとシームレスに連携し、AIアプリケーションの管理とスケーリングを行えます
- ハードウェア固有の最適化: NVIDIA GPUの性能を最大限に活用します
- フレームワークの柔軟性: PyTorch、TensorFlow、ONNX、OpenVINO、TensorRTなど、複数のAIフレームワークをサポートします
- オープンソースでカスタマイズ可能: 特定のニーズに合わせて変更できるため、さまざまなAIアプリケーションに柔軟に対応できます
前提条件#
作業を進める前に、次の前提条件を満たしていることを確認してください。
- Docker(>= 28.2.0、CDIによるGPUアクセスにはNVIDIA Container Toolkit >= 1.18が必要)またはPodmanがマシンにインストールされていること
ultralyticsをインストールします。pip install ultralyticstritonclientをインストールします。pip install tritonclient[all]
Triton Inference Serverのセットアップ#
次のセットアップブロックを実行すると、Ultralytics YOLO26をONNXにエクスポートし、Tritonのモデルリポジトリを構築して、Triton Inference Serverを起動できます。
スクリプト内のruntimeスイッチを使用して、コンテナエンジンを選択します。
- Dockerの場合は
runtime = "docker"を設定します - Podmanの場合は
runtime = "podman"を設定します
import contextlib
import subprocess
import time
from pathlib import Path
from tritonclient.http import InferenceServerClient
from ultralytics import YOLO
runtime = "docker" # set to "podman" to use Podman
# 1) Exporting YOLO26 to ONNX Format
# Load a model
model = YOLO("yolo26n.pt") # load an official model
# Retrieve metadata during export. Metadata needs to be added to config.pbtxt. See next section.
metadata = []
def export_cb(exporter):
metadata.append(exporter.metadata)
model.add_callback("on_export_end", export_cb)
# Export the model
onnx_file = model.export(format="onnx", dynamic=True)
# 2) Setting Up Triton Model Repository
# Define paths
model_name = "yolo"
triton_repo_path = Path("tmp") / "triton_repo"
triton_model_path = triton_repo_path / model_name
# Create directories
(triton_model_path / "1").mkdir(parents=True, exist_ok=True)
# Move ONNX model to Triton Model path
Path(onnx_file).rename(triton_model_path / "1" / "model.onnx")
# Create config file
(triton_model_path / "config.pbtxt").touch()
data = """
# Add metadata
parameters {
key: "metadata"
value {
string_value: "%s"
}
}
# Enable TensorRT acceleration (requires a GPU and TensorRT-enabled Triton; remove this block for CPU-only serving)
# The first run will be slow due to TensorRT engine conversion
optimization {
execution_accelerators {
gpu_execution_accelerator {
name: "tensorrt"
parameters {
key: "precision_mode"
value: "FP16"
}
parameters {
key: "max_workspace_size_bytes"
value: "3221225472"
}
parameters {
key: "trt_engine_cache_enable"
value: "1"
}
parameters {
key: "trt_engine_cache_path"
value: "/models/yolo/1"
}
}
}
}
""" % metadata[0] # noqa
with open(triton_model_path / "config.pbtxt", "w") as f:
f.write(data)
# 3) Running Triton Inference Server
# Define image https://catalog.ngc.nvidia.com/orgs/nvidia/containers/tritonserver
tag = "nvcr.io/nvidia/tritonserver:26.02-py3" # 16.17 GB (Compressed Size)
subprocess.call(f"{runtime} pull {tag}", shell=True)
# CDI GPU request works identically on Docker and Podman
gpu_flags = "--device nvidia.com/gpu=all"
container_name = "triton_server"
# Note: The :z flag on the volume mount is necessary for systems with SELinux (like Fedora/RHEL)
subprocess.call(
f"{runtime} run -d --rm --name {container_name} {gpu_flags} -v {triton_repo_path.absolute()}:/models:z -p 8000:8000 {tag} tritonserver --model-repository=/models",
shell=True,
)
# Wait for the Triton server to start
triton_client = InferenceServerClient(url="127.0.0.1:8000", verbose=False, ssl=False)
# Wait until model is ready
for _ in range(10):
with contextlib.suppress(Exception):
assert triton_client.is_model_ready(model_name)
break
time.sleep(1)推論の実行#
Triton Serverのモデルを使用して推論を実行します。
from ultralytics import YOLO
# Triton Serverのモデルを読み込む
model = YOLO("http://127.0.0.1:8000/yolo", task="detect")
# サーバーで推論を実行する
results = model("path/to/image.jpg")コンテナをクリーンアップします。
import subprocess
runtime = "docker" # set to "podman" to use Podman
container_name = "triton_server" # Kill the named container
subprocess.call(f"{runtime} kill {container_name}", shell=True)TensorRTによる最適化(任意)#
さらに高いパフォーマンスを得るには、Triton Inference ServerでTensorRTを使用できます。TensorRTは、NVIDIA GPU専用に構築された高性能ディープラーニングオプティマイザーで、推論速度を大幅に向上させることができます。
TritonでTensorRTを使用する主なメリットは次のとおりです。
- 最適化されていないモデルと比較して、推論速度が最大36倍向上します
- GPUを最大限に活用するためのハードウェア固有の最適化
- 精度を維持しながら、低精度形式(INT8、FP16)をサポートします
- レイヤーを融合して計算オーバーヘッドを削減します
上記のconfig.pbtxtでは、ONNXモデルに対するTritonのTensorRTアクセラレーターがすでに有効になっています。代わりにUltralyticsでTensorRTを直接使用するには、Ultralytics YOLO26モデルをTensorRT形式にエクスポートしてください。
from ultralytics import YOLO
# YOLO26モデルを読み込む
model = YOLO("yolo26n.pt")
# モデルをTensorRT形式にエクスポートします
model.export(format="engine") # 'yolo26n.engine'を作成しますTensorRTによる最適化の詳細は、TensorRTインテグレーションガイドを参照してください。
これで、Ultralytics YOLO26モデルをTriton Inference Serverにデプロイして実行し、スケーラブルで高性能な推論を実現できます。詳しくは公式Tritonドキュメントを参照するか、Ultralyticsコミュニティに問い合わせてください。
よくある質問#
Ultralytics YOLO26をNVIDIA Triton Inference Serverでセットアップするには、いくつかの重要な手順があります。
-
YOLO26をONNX形式にエクスポートする:
from ultralytics import YOLO # モデルを読み込む model = YOLO("yolo26n.pt") # 公式モデルを読み込みます # モデルをONNXフォーマットにエクスポート onnx_file = model.export(format="onnx", dynamic=True) -
Tritonモデルリポジトリをセットアップする:
from pathlib import Path # パスを定義する model_name = "yolo" triton_repo_path = Path("tmp") / "triton_repo" triton_model_path = triton_repo_path / model_name # ディレクトリを作成する (triton_model_path / "1").mkdir(parents=True, exist_ok=True) Path(onnx_file).rename(triton_model_path / "1" / "model.onnx") (triton_model_path / "config.pbtxt").touch() -
Triton Serverを起動する:
import contextlib import subprocess import time from tritonclient.http import InferenceServerClient # Define image https://catalog.ngc.nvidia.com/orgs/nvidia/containers/tritonserver tag = "nvcr.io/nvidia/tritonserver:26.02-py3" runtime = "docker" # set to "podman" to use Podman subprocess.call(f"{runtime} pull {tag}", shell=True) # CDI GPU request works identically on Docker and Podman gpu_flags = "--device nvidia.com/gpu=all" container_name = "triton_server" subprocess.call( f"{runtime} run -d --rm --name {container_name} {gpu_flags} -v {triton_repo_path.absolute()}:/models:z -p 8000:8000 {tag} tritonserver --model-repository=/models", shell=True, ) triton_client = InferenceServerClient(url="127.0.0.1:8000", verbose=False, ssl=False) for _ in range(10): with contextlib.suppress(Exception): assert triton_client.is_model_ready(model_name) break time.sleep(1)
このセットアップを使用すると、高性能なAIモデル推論のために、Triton Inference Server上でUltralytics YOLO26モデルを効率的に大規模デプロイできます。
-
Ultralytics YOLO26をNVIDIA Triton Inference Serverと統合すると、次のような利点があります。
- スケーラブルなAI推論: Tritonでは、1つのサーバーインスタンスから複数のモデルをサービングできます。モデルの動的な読み込みとアンロードにも対応しており、多様なAIワークロードに対して高いスケーラビリティを発揮します。
- 高性能: NVIDIA GPU向けに最適化されたTriton Inference Serverは、高速な推論処理を実現し、物体検出などのリアルタイムアプリケーションに最適です。
- アンサンブルとモデルのバージョン管理: Tritonのアンサンブルモードでは複数のモデルを組み合わせて結果を向上できます。また、モデルのバージョン管理によりA/Bテストとローリングアップデートが可能です。
- 自動バッチ処理: Tritonは複数の推論リクエストを自動的にまとめることで、スループットを大幅に向上させ、レイテンシを短縮します。
- デプロイの簡素化: システム全体を刷新することなく、AIワークフローを段階的に最適化できるため、効率的なスケーリングが容易になります。
Ultralytics YOLO26をTritonでセットアップして実行する詳しい手順は、Triton Inference Serverのセットアップと推論の実行を参照してください。
Ultralytics YOLO26モデルを NVIDIATriton Inference Serverにデプロイする前にONNX(Open Neural Network Exchange)形式を使用すると、次のような重要なメリットがあります。
- 相互運用性: ONNX形式は、異なるディープラーニングフレームワーク(PyTorch、TensorFlowなど)間での移行に対応しており、幅広い互換性を確保できます。
- 最適化: Tritonを含む多くのデプロイ環境はONNX向けに最適化されているため、推論を高速化し、パフォーマンスを向上できます。
- デプロイの容易さ: ONNXは多くのフレームワークやプラットフォームで広くサポートされており、さまざまなオペレーティングシステムやハードウェア構成へのデプロイを簡素化します。
- フレームワークからの独立性: ONNXに変換すると、モデルは元のフレームワークに依存しなくなり、移植性が高まります。
- 標準化: ONNXは標準化された表現を提供し、異なるAIフレームワーク間の互換性の問題の解決に役立ちます。
モデルをエクスポートするには、次のコマンドを使用します。
from ultralytics import YOLO model = YOLO("yolo26n.pt") onnx_file = model.export(format="onnx", dynamic=True)ONNX統合ガイドの手順に従って、処理を完了できます。
はい、NVIDIA Triton Inference ServerでUltralytics YOLO26モデルを使用して推論を実行できます。Triton Model Repositoryにモデルをセットアップし、サーバーを起動したら、次のようにモデルを読み込んで推論を実行できます。
from ultralytics import YOLO # Triton Serverのモデルを読み込む model = YOLO("http://127.0.0.1:8000/yolo", task="detect") # サーバーで推論を実行する results = model("path/to/image.jpg")この方法を使うと、使い慣れたUltralytics YOLOインターフェースを利用しながら、Tritonの最適化を活用できます。
Ultralytics YOLO26には、デプロイにおいてTensorFlowやPyTorchのモデルと比べて、次のような独自の利点があります。
- リアルタイム性能: リアルタイム物体検出タスク向けに最適化されたUltralytics YOLO26は、最先端の精度と速度を実現しており、ライブ動画分析を必要とするアプリケーションに最適です。
- 使いやすさ: Ultralytics YOLO26はTriton Inference Serverとシームレスに統合でき、多様なエクスポート形式(ONNX、TensorRT)をサポートしているため、さまざまなデプロイシナリオに柔軟に対応できます。
- 高度な機能: Triton経由でサービングすると、動的なモデル読み込み、モデルのバージョン管理、アンサンブル推論を利用できます。これらは、スケーラブルで信頼性の高いAIデプロイに不可欠です。
- シンプルなAPI: Ultralytics APIは、さまざまなデプロイ先で一貫したインターフェースを提供し、学習コストと開発時間を削減します。
- エッジ最適化: Ultralytics YOLO26モデルはエッジデプロイを考慮して設計されており、リソースが限られたデバイスでも優れた性能を発揮します。
詳細については、モデルのエクスポートガイドでデプロイオプションを比較してください。