Ultralytics YOLO26을 사용하는 Triton Inference Server#
Triton Inference Server(이전 명칭: TensorRT Inference Server)는 NVIDIA에서 개발한 오픈 소스 소프트웨어 솔루션입니다. NVIDIA GPU에 최적화된 클라우드 추론 솔루션을 제공합니다. Triton은 프로덕션 환경에서 AI 모델을 대규모로 배포하는 과정을 간소화합니다. Ultralytics YOLO26을 Triton Inference Server와 통합하면 확장 가능하고 성능이 뛰어난 딥러닝 추론 워크로드를 배포할 수 있습니다. 이 가이드에서는 통합을 설정하고 테스트하는 단계를 안내합니다.
시청: NVIDIA Triton Inference Server 시작하기.
Triton Inference Server란 무엇인가요?#
Triton Inference Server는 다양한 AI 모델을 프로덕션 환경에 배포하도록 설계되었습니다. PyTorch, TensorFlow, ONNX, OpenVINO, TensorRT를 비롯한 다양한 딥러닝 및 머신러닝 프레임워크를 지원합니다. 주요 사용 사례는 다음과 같습니다.
- 단일 서버 인스턴스에서 여러 모델 제공
- 서버를 다시 시작하지 않고 모델을 동적으로 로드 및 언로드
- 여러 모델을 함께 사용해 결과를 얻는 앙상블 추론
- A/B 테스트 및 롤링 업데이트를 위한 모델 버전 관리
Triton Inference Server의 주요 이점#
Ultralytics YOLO26을 Triton Inference Server와 함께 사용하면 다음과 같은 여러 이점이 있습니다.
- 자동 배칭: 여러 AI 요청을 처리 전에 하나로 묶어 지연 시간을 줄이고 추론 속도를 높입니다.
- Kubernetes 통합: 클라우드 네이티브 설계로 Kubernetes와 원활하게 작동하여 AI 애플리케이션을 관리하고 확장할 수 있습니다.
- 하드웨어별 최적화: NVIDIA GPU를 최대한 활용해 최고의 성능을 제공합니다.
- 프레임워크 유연성: PyTorch, TensorFlow, ONNX, OpenVINO 및 TensorRT를 비롯한 여러 AI 프레임워크를 지원합니다.
- 오픈 소스 및 사용자 지정 가능: 다양한 AI 애플리케이션의 유연성을 보장하며 특정 요구 사항에 맞게 수정할 수 있습니다.
사전 요구 사항#
진행하기 전에 다음 사전 요구 사항을 충족하는지 확인하세요.
- 컴퓨터에 Docker(>= 28.2.0, CDI GPU 액세스용 NVIDIA Container Toolkit >= 1.18 포함) 또는 Podman이 설치되어 있어야 합니다.
ultralytics을 설치합니다.pip install ultralyticstritonclient을 설치합니다.pip install tritonclient[all]
Triton Inference Server 설정#
이 전체 설정 블록을 실행하여 Ultralytics YOLO26을 ONNX로 내보내고, Triton 모델 리포지토리를 빌드한 다음 Triton Inference Server를 시작하세요.
스크립트에서 runtime 스위치를 사용해 컨테이너 엔진을 선택하세요.
- Docker를 사용하려면
runtime = "docker"을 설정합니다. - Podman을 사용하려면
runtime = "podman"을 설정합니다.
import contextlib
import subprocess
import time
from pathlib import Path
from tritonclient.http import InferenceServerClient
from ultralytics import YOLO
runtime = "docker" # set to "podman" to use Podman
# 1) Exporting YOLO26 to ONNX Format
# Load a model
model = YOLO("yolo26n.pt") # load an official model
# Retrieve metadata during export. Metadata needs to be added to config.pbtxt. See next section.
metadata = []
def export_cb(exporter):
metadata.append(exporter.metadata)
model.add_callback("on_export_end", export_cb)
# Export the model
onnx_file = model.export(format="onnx", dynamic=True)
# 2) Setting Up Triton Model Repository
# Define paths
model_name = "yolo"
triton_repo_path = Path("tmp") / "triton_repo"
triton_model_path = triton_repo_path / model_name
# Create directories
(triton_model_path / "1").mkdir(parents=True, exist_ok=True)
# Move ONNX model to Triton Model path
Path(onnx_file).rename(triton_model_path / "1" / "model.onnx")
# Create config file
(triton_model_path / "config.pbtxt").touch()
data = """
# Add metadata
parameters {
key: "metadata"
value {
string_value: "%s"
}
}
# Enable TensorRT acceleration (requires a GPU and TensorRT-enabled Triton; remove this block for CPU-only serving)
# The first run will be slow due to TensorRT engine conversion
optimization {
execution_accelerators {
gpu_execution_accelerator {
name: "tensorrt"
parameters {
key: "precision_mode"
value: "FP16"
}
parameters {
key: "max_workspace_size_bytes"
value: "3221225472"
}
parameters {
key: "trt_engine_cache_enable"
value: "1"
}
parameters {
key: "trt_engine_cache_path"
value: "/models/yolo/1"
}
}
}
}
""" % metadata[0] # noqa
with open(triton_model_path / "config.pbtxt", "w") as f:
f.write(data)
# 3) Running Triton Inference Server
# Define image https://catalog.ngc.nvidia.com/orgs/nvidia/containers/tritonserver
tag = "nvcr.io/nvidia/tritonserver:26.02-py3" # 16.17 GB (Compressed Size)
subprocess.call(f"{runtime} pull {tag}", shell=True)
# CDI GPU request works identically on Docker and Podman
gpu_flags = "--device nvidia.com/gpu=all"
container_name = "triton_server"
# Note: The :z flag on the volume mount is necessary for systems with SELinux (like Fedora/RHEL)
subprocess.call(
f"{runtime} run -d --rm --name {container_name} {gpu_flags} -v {triton_repo_path.absolute()}:/models:z -p 8000:8000 {tag} tritonserver --model-repository=/models",
shell=True,
)
# Wait for the Triton server to start
triton_client = InferenceServerClient(url="127.0.0.1:8000", verbose=False, ssl=False)
# Wait until model is ready
for _ in range(10):
with contextlib.suppress(Exception):
assert triton_client.is_model_ready(model_name)
break
time.sleep(1)추론 실행#
Triton Server 모델을 사용해 추론을 실행합니다.
from ultralytics import YOLO
# Triton Server 모델 로드
model = YOLO("http://127.0.0.1:8000/yolo", task="detect")
# 서버에서 추론 실행
results = model("path/to/image.jpg")컨테이너를 정리합니다.
import subprocess
runtime = "docker" # set to "podman" to use Podman
container_name = "triton_server" # Kill the named container
subprocess.call(f"{runtime} kill {container_name}", shell=True)TensorRT 최적화(선택 사항)#
성능을 더욱 높이려면 Triton Inference Server에서 TensorRT를 사용할 수 있습니다. TensorRT는 NVIDIA GPU 전용으로 개발된 고성능 딥러닝 최적화 도구로, 추론 속도를 크게 높일 수 있습니다.
Triton에서 TensorRT를 사용할 때의 주요 이점은 다음과 같습니다.
- 최적화되지 않은 모델보다 추론 속도가 최대 36배 빠릅니다.
- GPU를 최대한 활용하기 위한 하드웨어별 최적화를 제공합니다.
- 정확도를 유지하면서 저정밀도 형식(INT8, FP16)을 지원합니다.
- 계산 오버헤드를 줄이기 위해 레이어를 융합합니다.
위의 config.pbtxt은 이미 ONNX 모델에 Triton의 TensorRT 가속기를 활성화합니다. 대신 Ultralytics에서 TensorRT를 직접 실행하려면 Ultralytics YOLO26 모델을 TensorRT 형식으로 내보내세요.
from ultralytics import YOLO
# YOLO26 모델을 불러옵니다
model = YOLO("yolo26n.pt")
# 모델을 TensorRT 형식으로 내보내기
model.export(format="engine") # 'yolo26n.engine' 생성TensorRT 최적화에 관한 자세한 내용은 TensorRT 통합 가이드를 참조하세요.
이제 확장 가능하고 성능이 뛰어난 추론을 위해 Ultralytics YOLO26 모델을 Triton Inference Server에 배포하고 실행할 수 있습니다. 자세한 내용은 공식 Triton 문서를 참조하거나 Ultralytics 커뮤니티에 도움을 요청하세요.
자주 묻는 질문#
Ultralytics YOLO26을 NVIDIA Triton Inference Server와 함께 설정하려면 몇 가지 주요 단계를 따라야 합니다.
-
YOLO26을 ONNX 형식으로 내보내기:
from ultralytics import YOLO # 모델 로드 model = YOLO("yolo26n.pt") # 공식 모델을 불러옵니다 # 모델을 ONNX 형식으로 내보내기 onnx_file = model.export(format="onnx", dynamic=True) -
Triton 모델 리포지토리 설정:
from pathlib import Path # 경로 정의 model_name = "yolo" triton_repo_path = Path("tmp") / "triton_repo" triton_model_path = triton_repo_path / model_name # 디렉터리 생성 (triton_model_path / "1").mkdir(parents=True, exist_ok=True) Path(onnx_file).rename(triton_model_path / "1" / "model.onnx") (triton_model_path / "config.pbtxt").touch() -
Triton Server 실행:
import contextlib import subprocess import time from tritonclient.http import InferenceServerClient # Define image https://catalog.ngc.nvidia.com/orgs/nvidia/containers/tritonserver tag = "nvcr.io/nvidia/tritonserver:26.02-py3" runtime = "docker" # set to "podman" to use Podman subprocess.call(f"{runtime} pull {tag}", shell=True) # CDI GPU request works identically on Docker and Podman gpu_flags = "--device nvidia.com/gpu=all" container_name = "triton_server" subprocess.call( f"{runtime} run -d --rm --name {container_name} {gpu_flags} -v {triton_repo_path.absolute()}:/models:z -p 8000:8000 {tag} tritonserver --model-repository=/models", shell=True, ) triton_client = InferenceServerClient(url="127.0.0.1:8000", verbose=False, ssl=False) for _ in range(10): with contextlib.suppress(Exception): assert triton_client.is_model_ready(model_name) break time.sleep(1)
이 설정을 사용하면 고성능 AI 모델 추론을 위해 Ultralytics YOLO26 모델을 Triton Inference Server에 효율적으로 대규모 배포할 수 있습니다.
-
Ultralytics YOLO26을 NVIDIA Triton Inference Server와 통합하면 다음과 같은 여러 이점이 있습니다.
- 확장 가능한 AI 추론: Triton은 단일 서버 인스턴스에서 여러 모델을 제공하고 모델을 동적으로 로드 및 언로드할 수 있어 다양한 AI 워크로드에 맞게 확장할 수 있습니다.
- 고성능: NVIDIA GPU에 최적화된 Triton Inference Server는 고속 추론 작업을 보장하므로 객체 감지와 같은 실시간 애플리케이션에 적합합니다.
- 앙상블 및 모델 버전 관리: Triton의 앙상블 모드에서는 여러 모델을 결합해 결과를 개선할 수 있으며, 모델 버전 관리 기능은 A/B 테스트 및 롤링 업데이트를 지원합니다.
- 자동 배칭: Triton은 여러 추론 요청을 자동으로 하나로 묶어 처리량을 크게 높이고 지연 시간을 줄입니다.
- 간소화된 배포: 시스템을 전면 개편하지 않고 AI 워크플로를 단계적으로 최적화할 수 있어 효율적인 확장이 쉬워집니다.
Triton에서 Ultralytics YOLO26을 설정하고 실행하는 자세한 방법은 Triton Inference Server 설정 및 추론 실행을 참조하세요.
Ultralytics YOLO26 모델을 NVIDIA Triton Inference Server에 배포하기 전에 ONNX(개방형 신경망 교환) 형식을 사용하면 다음과 같은 주요 이점이 있습니다.
- 상호 운용성: ONNX 형식은 서로 다른 딥러닝 프레임워크(예: PyTorch, TensorFlow) 간 전송을 지원하여 호환성을 넓힙니다.
- 최적화: Triton을 비롯한 여러 배포 환경은 ONNX에 최적화되어 있어 추론 속도와 성능을 높일 수 있습니다.
- 간편한 배포: ONNX는 다양한 프레임워크와 플랫폼에서 폭넓게 지원되므로 여러 운영 체제 및 하드웨어 구성에서 배포 과정을 간소화합니다.
- 프레임워크 독립성: ONNX로 변환하면 모델이 더 이상 원래 프레임워크에 종속되지 않아 이식성이 높아집니다.
- 표준화: ONNX는 표준화된 표현을 제공하여 서로 다른 AI 프레임워크 간의 호환성 문제를 해결하는 데 도움이 됩니다.
모델을 내보내려면 다음을 사용합니다.
from ultralytics import YOLO model = YOLO("yolo26n.pt") onnx_file = model.export(format="onnx", dynamic=True)ONNX 통합 가이드의 단계를 따르면 프로세스를 완료할 수 있습니다.
예, Ultralytics YOLO26 모델을 NVIDIA Triton Inference Server에서 사용하여 추론을 실행할 수 있습니다. Triton Model Repository에서 모델을 설정하고 서버를 실행한 다음, 다음과 같이 모델을 로드하고 추론을 실행할 수 있습니다.
from ultralytics import YOLO # Triton Server 모델 로드 model = YOLO("http://127.0.0.1:8000/yolo", task="detect") # 서버에서 추론 실행 results = model("path/to/image.jpg")이 접근 방식을 사용하면 익숙한 Ultralytics YOLO 인터페이스를 활용하면서 Triton의 최적화 기능도 이용할 수 있습니다.
Ultralytics YOLO26은 배포 시 TensorFlow 및 PyTorch 모델과 비교해 몇 가지 고유한 장점을 제공합니다.
- 실시간 성능: 실시간 객체 감지 작업에 최적화된 Ultralytics YOLO26은 최첨단 정확도와 속도를 제공하므로 실시간 비디오 분석이 필요한 애플리케이션에 적합합니다.
- 사용 편의성: Ultralytics YOLO26은 Triton Inference Server와 원활하게 통합되며 다양한 내보내기 형식(ONNX, TensorRT)을 지원하므로 여러 배포 시나리오에 유연하게 대응합니다.
- 고급 기능: Triton을 통한 서빙은 동적 모델 로딩, 모델 버전 관리, 앙상블 추론을 지원합니다. 이러한 기능은 확장 가능하고 안정적인 AI 배포에 필수적입니다.
- 간소화된 API: Ultralytics API는 다양한 배포 대상에서 일관된 인터페이스를 제공하여 학습 곡선과 개발 시간을 줄여줍니다.
- 엣지 최적화: Ultralytics YOLO26 모델은 엣지 배포를 고려해 설계되었으며, 리소스가 제한된 장치에서도 뛰어난 성능을 제공합니다.
자세한 내용은 모델 내보내기 가이드에서 배포 옵션을 비교해 보세요.