Ultralytics YOLO26을 사용하는 Triton Inference Server#
Triton Inference Server(이전 명칭: TensorRT Inference Server)는 NVIDIA가 개발한 오픈 소스 소프트웨어 솔루션입니다. NVIDIA GPU에 최적화된 클라우드 추론 솔루션을 제공합니다. Triton은 프로덕션 환경에서 AI 모델을 대규모로 배포하는 작업을 간소화합니다. Ultralytics YOLO26을 Triton Inference Server와 통합하면 확장 가능하고 고성능인 딥러닝 추론 워크로드를 배포할 수 있습니다. 이 가이드에서는 통합을 설정하고 테스트하는 단계를 제공합니다.
Watch: Getting Started with NVIDIA Triton Inference Server.
Triton Inference Server란 무엇입니까?#
Triton Inference Server는 다양한 AI 모델을 프로덕션 환경에 배포하도록 설계되었습니다. PyTorch, TensorFlow, ONNX, OpenVINO, TensorRT를 비롯한 다양한 딥러닝 및 머신러닝 프레임워크를 지원합니다. 주요 사용 사례는 다음과 같습니다.
- 단일 서버 인스턴스에서 여러 모델 제공
- 서버를 다시 시작하지 않고 모델을 동적으로 로드 및 언로드
- 여러 모델을 함께 사용하여 결과를 얻을 수 있는 앙상블 추론
- A/B 테스트 및 롤링 업데이트를 위한 모델 버전 관리
Triton Inference Server의 주요 이점#
Ultralytics YOLO26을 Triton Inference Server와 함께 사용하면 다음과 같은 여러 이점이 있습니다.
- 자동 배칭: 처리 전에 여러 AI 요청을 하나로 그룹화하여 지연 시간을 줄이고 추론 속도를 향상합니다.
- Kubernetes 통합: 클라우드 네이티브 설계로 Kubernetes와 원활하게 연동하여 AI 애플리케이션을 관리하고 확장합니다.
- 하드웨어별 최적화: NVIDIA GPU를 최대한 활용하여 최고의 성능을 제공합니다.
- 프레임워크 유연성: PyTorch, TensorFlow, ONNX, OpenVINO, TensorRT를 비롯한 여러 AI 프레임워크를 지원합니다.
- 오픈 소스 및 사용자 지정 가능: 특정 요구 사항에 맞게 수정할 수 있어 다양한 AI 애플리케이션에 유연하게 적용할 수 있습니다.
필수 조건#
계속하기 전에 다음 사전 요구 사항이 충족되었는지 확인합니다.
- Docker(>= 28.2.0, CDI GPU 액세스를 위한 NVIDIA Container Toolkit >= 1.18 포함) 또는 Podman이 시스템에 설치되어 있어야 합니다.
ultralytics을 설치합니다.pip install ultralyticstritonclient을 설치합니다.pip install tritonclient[all]
Triton Inference Server 설정#
다음 전체 설정 블록을 실행하여 Ultralytics YOLO26을 ONNX로 내보내고, Triton 모델 리포지토리를 빌드한 다음 Triton Inference Server를 시작합니다.
스크립트에서 runtime 스위치를 사용하여 컨테이너 엔진을 선택합니다.
- Docker에는
runtime = "docker"을 설정합니다. - Podman에는
runtime = "podman"을 설정합니다.
import contextlib
import subprocess
import time
from pathlib import Path
from tritonclient.http import InferenceServerClient
from ultralytics import YOLO
runtime = "docker" # set to "podman" to use Podman
# 1) Exporting YOLO26 to ONNX Format
# Load a model
model = YOLO("yolo26n.pt") # load an official model
# Retrieve metadata during export. Metadata needs to be added to config.pbtxt. See next section.
metadata = []
def export_cb(exporter):
metadata.append(exporter.metadata)
model.add_callback("on_export_end", export_cb)
# Export the model
onnx_file = model.export(format="onnx", dynamic=True)
# 2) Setting Up Triton Model Repository
# Define paths
model_name = "yolo"
triton_repo_path = Path("tmp") / "triton_repo"
triton_model_path = triton_repo_path / model_name
# Create directories
(triton_model_path / "1").mkdir(parents=True, exist_ok=True)
# Move ONNX model to Triton Model path
Path(onnx_file).rename(triton_model_path / "1" / "model.onnx")
# Create config file
(triton_model_path / "config.pbtxt").touch()
data = """
# Add metadata
parameters {
key: "metadata"
value {
string_value: "%s"
}
}
# Enable TensorRT acceleration (requires a GPU and TensorRT-enabled Triton; remove this block for CPU-only serving)
# The first run will be slow due to TensorRT engine conversion
optimization {
execution_accelerators {
gpu_execution_accelerator {
name: "tensorrt"
parameters {
key: "precision_mode"
value: "FP16"
}
parameters {
key: "max_workspace_size_bytes"
value: "3221225472"
}
parameters {
key: "trt_engine_cache_enable"
value: "1"
}
parameters {
key: "trt_engine_cache_path"
value: "/models/yolo/1"
}
}
}
}
""" % metadata[0] # noqa
with open(triton_model_path / "config.pbtxt", "w") as f:
f.write(data)
# 3) Running Triton Inference Server
# Define image https://catalog.ngc.nvidia.com/orgs/nvidia/containers/tritonserver
tag = "nvcr.io/nvidia/tritonserver:26.02-py3" # 16.17 GB (Compressed Size)
subprocess.call(f"{runtime} pull {tag}", shell=True)
# CDI GPU request works identically on Docker and Podman
gpu_flags = "--device nvidia.com/gpu=all"
container_name = "triton_server"
# Note: The :z flag on the volume mount is necessary for systems with SELinux (like Fedora/RHEL)
subprocess.call(
f"{runtime} run -d --rm --name {container_name} {gpu_flags} -v {triton_repo_path.absolute()}:/models:z -p 8000:8000 {tag} tritonserver --model-repository=/models",
shell=True,
)
# Wait for the Triton server to start
triton_client = InferenceServerClient(url="127.0.0.1:8000", verbose=False, ssl=False)
# Wait until model is ready
for _ in range(10):
with contextlib.suppress(Exception):
assert triton_client.is_model_ready(model_name)
break
time.sleep(1)추론 실행#
Triton Server 모델을 사용하여 추론을 실행합니다.
from ultralytics import YOLO
# Load the Triton Server model
model = YOLO("http://127.0.0.1:8000/yolo", task="detect")
# Run inference on the server
results = model("path/to/image.jpg")컨테이너를 정리합니다.
import subprocess
runtime = "docker" # set to "podman" to use Podman
container_name = "triton_server" # Kill the named container
subprocess.call(f"{runtime} kill {container_name}", shell=True)TensorRT 최적화(선택 사항)#
더 높은 성능을 위해 Triton Inference Server에서 TensorRT를 사용할 수 있습니다. TensorRT는 NVIDIA GPU용으로 특별히 구축된 고성능 딥러닝 옵티마이저로, 추론 속도를 크게 향상할 수 있습니다.
TensorRT를 Triton과 함께 사용할 때의 주요 이점은 다음과 같습니다.
- 최적화되지 않은 모델보다 최대 36배 빠른 추론
- GPU 활용도를 극대화하는 하드웨어별 최적화
- 정확도를 유지하면서 낮은 정밀도 형식(INT8, FP16) 지원
- 계산 오버헤드를 줄이는 레이어 융합
TensorRT를 직접 사용하려면 Ultralytics YOLO26 모델을 TensorRT 형식으로 내보낼 수 있습니다.
from ultralytics import YOLO
# Load the YOLO26 model
model = YOLO("yolo26n.pt")
# Export the model to TensorRT format
model.export(format="engine") # creates 'yolo26n.engine'TensorRT 최적화에 대한 자세한 내용은 TensorRT 통합 가이드를 참조하십시오.
이제 Triton Inference Server에서 Ultralytics YOLO26 모델을 배포하고 실행하여 확장 가능하고 고성능인 추론을 수행할 수 있습니다. 자세한 내용은 Triton 공식 문서를 참조하거나 Ultralytics 커뮤니티에 도움을 요청하십시오.
FAQ#
Ultralytics YOLO26을 NVIDIA Triton Inference Server와 함께 설정하려면 몇 가지 주요 단계를 수행해야 합니다.
-
YOLO26을 ONNX 형식으로 내보내기:
from ultralytics import YOLO # Load a model model = YOLO("yolo26n.pt") # load an official model # Export the model to ONNX format onnx_file = model.export(format="onnx", dynamic=True) -
Triton 모델 리포지토리 설정:
from pathlib import Path # Define paths model_name = "yolo" triton_repo_path = Path("tmp") / "triton_repo" triton_model_path = triton_repo_path / model_name # Create directories (triton_model_path / "1").mkdir(parents=True, exist_ok=True) Path(onnx_file).rename(triton_model_path / "1" / "model.onnx") (triton_model_path / "config.pbtxt").touch() -
Triton Server 실행:
import contextlib import subprocess import time from tritonclient.http import InferenceServerClient # Define image https://catalog.ngc.nvidia.com/orgs/nvidia/containers/tritonserver tag = "nvcr.io/nvidia/tritonserver:26.02-py3" runtime = "docker" # set to "podman" to use Podman subprocess.call(f"{runtime} pull {tag}", shell=True) # CDI GPU request works identically on Docker and Podman gpu_flags = "--device nvidia.com/gpu=all" container_name = "triton_server" subprocess.call( f"{runtime} run -d --rm --name {container_name} {gpu_flags} -v {triton_repo_path.absolute()}:/models:z -p 8000:8000 {tag} tritonserver --model-repository=/models", shell=True, ) triton_client = InferenceServerClient(url="127.0.0.1:8000", verbose=False, ssl=False) for _ in range(10): with contextlib.suppress(Exception): assert triton_client.is_model_ready(model_name) break time.sleep(1)
이 설정을 사용하면 Triton Inference Server에서 Ultralytics YOLO26 모델을 대규모로 효율적으로 배포하여 고성능 AI 모델 추론을 수행할 수 있습니다.
-
Ultralytics YOLO26을 NVIDIA Triton Inference Server와 통합하면 다음과 같은 여러 이점이 있습니다.
- 확장 가능한 AI 추론: Triton은 단일 서버 인스턴스에서 여러 모델을 제공하고 모델을 동적으로 로드 및 언로드할 수 있어 다양한 AI 워크로드에 대해 높은 확장성을 제공합니다.
- 고성능: NVIDIA GPU에 최적화된 Triton Inference Server는 고속 추론 작업을 보장하므로 객체 감지와 같은 실시간 애플리케이션에 적합합니다.
- 앙상블 및 모델 버전 관리: Triton의 앙상블 모드를 사용하면 여러 모델을 결합하여 결과를 향상할 수 있으며, 모델 버전 관리를 통해 A/B 테스트와 롤링 업데이트를 지원합니다.
- 자동 배칭: Triton은 여러 추론 요청을 자동으로 하나로 그룹화하여 처리량을 크게 향상하고 지연 시간을 줄입니다.
- 간소화된 배포: 전체 시스템을 개편하지 않고도 AI 워크플로를 점진적으로 최적화할 수 있어 효율적으로 확장하기가 더 쉽습니다.
Ultralytics YOLO26을 Triton에서 설정하고 실행하는 자세한 지침은 Triton Inference Server 설정 및 추론 실행을 참조하십시오.
Ultralytics YOLO26 모델을 NVIDIA Triton Inference Server에 배포하기 전에 ONNX 형식을 사용하면 다음과 같은 주요 이점이 있습니다.
- 상호 운용성: ONNX 형식은 서로 다른 딥러닝 프레임워크(PyTorch, TensorFlow 등) 간 전송을 지원하여 더 폭넓은 호환성을 보장합니다.
- 최적화: Triton을 비롯한 많은 배포 환경이 ONNX에 맞게 최적화되어 더 빠른 추론과 향상된 성능을 제공합니다.
- 간편한 배포: ONNX는 다양한 프레임워크와 플랫폼에서 폭넓게 지원되므로 여러 운영 체제와 하드웨어 구성에서 배포 프로세스를 간소화합니다.
- 프레임워크 독립성: ONNX로 변환하면 모델이 더 이상 원래 프레임워크에 종속되지 않아 이식성이 향상됩니다.
- 표준화: ONNX는 서로 다른 AI 프레임워크 간의 호환성 문제를 해결하는 데 도움이 되는 표준화된 표현을 제공합니다.
모델을 내보내려면 다음을 사용합니다.
from ultralytics import YOLO model = YOLO("yolo26n.pt") onnx_file = model.export(format="onnx", dynamic=True)ONNX 통합 가이드의 단계에 따라 프로세스를 완료할 수 있습니다.
예, NVIDIA Triton Inference Server에서 Ultralytics YOLO26 모델을 사용하여 추론을 실행할 수 있습니다. Triton 모델 리포지토리에 모델을 설정하고 서버를 실행한 후에는 다음과 같이 모델을 로드하고 추론을 실행할 수 있습니다.
from ultralytics import YOLO # Load the Triton Server model model = YOLO("http://127.0.0.1:8000/yolo", task="detect") # Run inference on the server results = model("path/to/image.jpg")이 방법을 사용하면 익숙한 Ultralytics YOLO 인터페이스를 활용하면서 Triton의 최적화 기능을 이용할 수 있습니다.
배포 측면에서 Ultralytics YOLO26은 TensorFlow 및 PyTorch 모델과 비교하여 다음과 같은 고유한 이점을 제공합니다.
- 실시간 성능: 실시간 객체 감지 작업에 최적화된 Ultralytics YOLO26은 최첨단 정확도와 속도를 제공하므로 실시간 비디오 분석이 필요한 애플리케이션에 적합합니다.
- 간편한 사용: Ultralytics YOLO26은 Triton Inference Server와 원활하게 통합되며 다양한 내보내기 형식(ONNX, TensorRT)을 지원하므로 다양한 배포 시나리오에 유연하게 대응할 수 있습니다.
- 고급 기능: Triton을 통한 서빙은 동적 모델 로딩, 모델 버전 관리, 앙상블 추론을 제공하며, 이는 확장 가능하고 신뢰할 수 있는 AI 배포에 매우 중요합니다.
- 간소화된 API: Ultralytics API는 다양한 배포 대상에서 일관된 인터페이스를 제공하여 학습 곡선과 개발 시간을 줄입니다.
- 엣지 최적화: Ultralytics YOLO26 모델은 엣지 배포를 고려하여 설계되었으므로 리소스가 제한된 디바이스에서도 뛰어난 성능을 제공합니다.
자세한 내용은 모델 내보내기 가이드에서 배포 옵션을 비교해 보십시오.