使用 Ultralytics 在 Vertex AI 上部署预训练 YOLO 模型进行推理#
本指南将向你介绍如何使用 Ultralytics 将预训练 YOLO26 模型容器化、为其构建 FastAPI 推理服务器,并通过该推理服务器将模型部署到 Google Cloud Vertex AI。示例实现将介绍 YOLO26 的目标检测用例,但相同的原则也适用于其他 YOLO 任务。
开始之前,你需要创建一个 Google Cloud Platform (GCP) 项目。作为新用户,你可以免费获得价值 300 美元的 GCP 额度;这笔额度足以测试一个正在运行的配置,之后你可以将其扩展到其他 YOLO26 用例,包括训练、批量推理和流式推理。
你将学到什么#
- 使用 FastAPI 为 Ultralytics YOLO26 模型创建推理后端。
- 创建 GCP Artifact Registry 仓库来存储 Docker 镜像。
- 构建包含模型的 Docker 镜像,并将其推送到 Artifact Registry。
- 在 Vertex AI 中导入模型。
- 创建 Vertex AI 端点并部署模型。
- 使用 Ultralytics 完全掌控模型:你可以使用自定义推理逻辑,全面控制预处理、后处理和响应格式。
- 其余工作交由 Vertex AI 处理:它会自动扩缩容,同时让你灵活配置计算资源、内存和 GPU。
- 原生 GCP 集成与安全性:可无缝配置 Cloud Storage、BigQuery、Cloud Functions、VPC 控制、IAM 政策和审计日志。
前置条件#
- 在你的机器上安装 Docker。
- 安装 Google Cloud SDK,并完成身份验证以使用 gcloud CLI。
- 强烈建议你先阅读 Ultralytics Docker 快速入门指南,因为你需要按照本指南扩展其中一个 Ultralytics 官方 Docker 镜像。
1. 使用 FastAPI 创建推理后端#
首先,你需要创建一个 FastAPI 应用,用来处理 YOLO26 模型的推理请求。该应用将负责模型加载、图像预处理和推理(预测)逻辑。
Vertex AI 合规性基础知识#
Vertex AI 要求你的容器实现两个特定端点:
-
健康检查端点(
/health):服务就绪时必须返回 HTTP 状态200 OK。 -
预测端点(
/predict):接受包含 base64 编码图像和可选参数的结构化预测请求。具体限制取决于端点类型,请参阅负载大小限制。/predict端点的请求负载应遵循以下 JSON 结构:{ "instances": [{ "image": "base64_encoded_image" }], "parameters": { "confidence": 0.5 } }
项目文件夹结构#
大部分构建工作将在 Docker 容器内完成,Ultralytics 也会加载预训练 YOLO26 模型,因此本地文件夹结构可以保持简单:
YOUR_PROJECT/
├── src/
│ ├── __init__.py
│ ├── app.py # Core YOLO26 inference logic
│ └── main.py # FastAPI inference server
├── tests/
├── .env # Environment variables for local development
├── Dockerfile # Container configuration
├── LICENSE # AGPL-3.0 License
└── pyproject.toml # Python dependencies and project configUltralytics YOLO26 模型和框架采用 AGPL-3.0 许可,其中包含重要的合规要求。请务必阅读 Ultralytics 文档,了解如何遵守许可条款。
创建包含依赖项的 pyproject.toml#
为了方便管理项目,请创建一个 pyproject.toml 文件,并添加以下依赖项:
[project]
name = "YOUR_PROJECT_NAME"
version = "0.0.1"
description = "YOUR_PROJECT_DESCRIPTION"
requires-python = ">=3.10,<3.13"
dependencies = [
"ultralytics>=8.4.0",
"fastapi[all]>=0.89.1",
"uvicorn[standard]>=0.20.0",
"pillow>=9.0.0",
"loguru",
]
[build-system]
requires = ["setuptools>=61.0"]
build-backend = "setuptools.build_meta"uvicorn将用于运行 FastAPI 服务器。loguru将用于在 FastAPI 服务器中进行日志记录。pillow将用于图像处理,但你不局限于 PIL 图像——Ultralytics 还支持许多其他格式。
使用 Ultralytics YOLO26 编写推理逻辑#
现在你已经设置好项目结构和依赖项,可以实现 YOLO26 的核心推理逻辑了。创建一个 src/app.py 文件,使用 Ultralytics Python API 处理模型加载、图像处理和预测。
# src/app.py
from ultralytics import YOLO
# Model initialization and readiness state
model_yolo = None
_model_ready = False
def _initialize_model():
"""Initialize the YOLO model."""
global model_yolo, _model_ready
try:
# Use pretrained YOLO26n model from Ultralytics base image
model_yolo = YOLO("yolo26n.pt")
_model_ready = True
except Exception as e:
print(f"Error initializing YOLO model: {e}")
_model_ready = False
model_yolo = None
# Initialize model on module import
_initialize_model()
def is_model_ready() -> bool:
"""Check if the model is ready for inference."""
return _model_ready and model_yolo is not None容器启动时会加载一次模型,并由所有请求共享。如果模型需要处理大量推理任务,建议在后续步骤向 Vertex AI 导入模型时选择内存更大的机器类型。
接下来,使用 pillow 创建两个用于输入和输出图像处理的实用函数。YOLO26 原生支持 PIL 图像。
def get_image_from_bytes(binary_image: bytes) -> Image.Image:
"""Convert image from bytes to PIL RGB format."""
input_image = Image.open(io.BytesIO(binary_image)).convert("RGB")
return input_imagedef get_bytes_from_image(image: Image.Image) -> bytes:
"""Convert PIL image to bytes."""
return_image = io.BytesIO()
image.save(return_image, format="JPEG", quality=85)
return_image.seek(0)
return return_image.getvalue()最后,实现用于处理目标检测的 run_inference 函数。在此示例中,我们将从模型预测结果中提取边界框、类别名称和置信度分数。该函数将返回一个字典,其中包含检测结果和原始结果,以便进一步处理或添加标注。
def run_inference(input_image: Image.Image, confidence_threshold: float = 0.5) -> Dict[str, Any]:
"""Run inference on an image using YOLO26n model."""
# Check if model is ready
if not is_model_ready():
print("Model not ready for inference")
return {"detections": [], "results": None}
try:
# Make predictions and get raw results
results = model_yolo.predict(
imgsz=640, source=input_image, conf=confidence_threshold, save=False, augment=False, verbose=False
)
# Extract detections (bounding boxes, class names, and confidences)
detections = []
if results and len(results) > 0:
result = results[0]
if result.boxes is not None and len(result.boxes.xyxy) > 0:
boxes = result.boxes
# Convert tensors to numpy for processing
xyxy = boxes.xyxy.cpu().numpy()
conf = boxes.conf.cpu().numpy()
cls = boxes.cls.cpu().numpy().astype(int)
# Create detection dictionaries
for i in range(len(xyxy)):
detection = {
"xmin": float(xyxy[i][0]),
"ymin": float(xyxy[i][1]),
"xmax": float(xyxy[i][2]),
"ymax": float(xyxy[i][3]),
"confidence": float(conf[i]),
"class": int(cls[i]),
"name": model_yolo.names.get(int(cls[i]), f"class_{int(cls[i])}"),
}
detections.append(detection)
return {
"detections": detections,
"results": results, # Keep raw results for annotation
}
except Exception as e:
# If there's an error, return empty structure
print(f"Error in YOLO detection: {e}")
return {"detections": [], "results": None}你也可以选择添加一个函数,使用 Ultralytics 内置的绘图方法为图像添加边界框和标签。如果你想在预测响应中返回带标注的图像,这会很有用。
def get_annotated_image(results: list) -> Image.Image:
"""Get annotated image using Ultralytics built-in plot method."""
if not results or len(results) == 0:
raise ValueError("No results provided for annotation")
result = results[0]
# 使用 Ultralytics 内置的绘图方法并输出 PIL 图像
return result.plot(pil=True)使用 FastAPI 创建 HTTP 推理服务器#
现在你已经完成 YOLO26 的核心推理逻辑,可以创建 FastAPI 应用来提供推理服务了。应用将包含 Vertex AI 所需的健康检查和预测端点。
首先,创建 src/main.py,添加导入项,创建 FastAPI 应用,并为 Vertex AI 配置日志记录。由于 Vertex AI 会将 stderr 视为错误输出,将日志输出到 stdout 是合理的做法。
# src/main.py
import sys
from fastapi import FastAPI
from loguru import logger
app = FastAPI()
# 配置日志记录器
logger.remove()
logger.add(
sys.stdout,
colorize=True,
format="<green>{time:HH:mm:ss}</green> | <level>{message}</level>",
level=10,
)
logger.add("log.log", rotation="1 MB", level="DEBUG", compression="zip")为了完全满足 Vertex AI 的合规要求,请在环境变量中定义必需的端点,并设置请求大小限制。建议在生产部署中使用 Vertex AI 私有端点。这样,除了具备完善的安全性和访问控制外,请求负载上限也会更高(10 MB,而公共端点为 1.5 MB)。
# Vertex AI 环境变量
AIP_HTTP_PORT = int(os.getenv("AIP_HTTP_PORT", "8080"))
AIP_HEALTH_ROUTE = os.getenv("AIP_HEALTH_ROUTE", "/health")
AIP_PREDICT_ROUTE = os.getenv("AIP_PREDICT_ROUTE", "/predict")
# 请求大小限制(私有端点为 10 MB,公共端点为 1.5 MB)
MAX_REQUEST_SIZE = 10 * 1024 * 1024 # 10 MB(字节数)添加两个 Pydantic 模型,用于验证请求和响应:
# 用于请求/响应的 Pydantic 模型
class PredictionRequest(BaseModel):
instances: list
parameters: Optional[Dict[str, Any]] = None
class PredictionResponse(BaseModel):
predictions: list添加健康检查端点以验证模型是否就绪。这对 Vertex AI 很重要:如果没有专用的健康检查,其编排器就会向随机套接字发送 ping,无法判断模型是否已准备好进行推理。检查成功时必须返回 200 OK,失败时则返回 503 Service Unavailable:
# 健康检查端点
@app.get(AIP_HEALTH_ROUTE, status_code=status.HTTP_200_OK)
def health_check():
"""Health check endpoint for Vertex AI."""
if not is_model_ready():
raise HTTPException(status_code=503, detail="Model not ready")
return {"status": "healthy"}现在,你已经具备实现预测端点所需的一切,可以处理推理请求了。该端点将接收图像文件、运行推理并返回结果。请注意,图像必须进行 base64 编码,这会使负载大小额外增加最多 33%。
@app.post(AIP_PREDICT_ROUTE, response_model=PredictionResponse)
async def predict(request: PredictionRequest):
"""Prediction endpoint for Vertex AI."""
try:
predictions = []
for instance in request.instances:
if isinstance(instance, dict):
if "image" in instance:
image_data = base64.b64decode(instance["image"])
input_image = get_image_from_bytes(image_data)
else:
raise HTTPException(status_code=400, detail="Instance must contain 'image' field")
else:
raise HTTPException(status_code=400, detail="Invalid instance format")
# Extract YOLO26 parameters if provided
parameters = request.parameters or {}
confidence_threshold = parameters.get("confidence", 0.5)
return_annotated_image = parameters.get("return_annotated_image", False)
# Run inference with YOLO26n model
result = run_inference(input_image, confidence_threshold=confidence_threshold)
detections_list = result["detections"]
# Format predictions for Vertex AI
detections = []
for detection in detections_list:
formatted_detection = {
"class": detection["name"],
"confidence": detection["confidence"],
"bbox": {
"xmin": detection["xmin"],
"ymin": detection["ymin"],
"xmax": detection["xmax"],
"ymax": detection["ymax"],
},
}
detections.append(formatted_detection)
# Build prediction response
prediction = {"detections": detections, "detection_count": len(detections)}
# Add annotated image if requested and detections exist
if (
return_annotated_image
and result["results"]
and result["results"][0].boxes is not None
and len(result["results"][0].boxes) > 0
):
annotated_image = get_annotated_image(result["results"])
img_bytes = get_bytes_from_image(annotated_image)
prediction["annotated_image"] = base64.b64encode(img_bytes).decode("utf-8")
predictions.append(prediction)
logger.info(
f"Processed {len(request.instances)} instances, found {sum(len(p['detections']) for p in predictions)} total detections"
)
return PredictionResponse(predictions=predictions)
except HTTPException:
# Re-raise HTTPException as-is (don't catch and convert to 500)
raise
except Exception as e:
logger.error(f"Prediction error: {e}")
raise HTTPException(status_code=500, detail=f"Prediction failed: {e}")最后,添加应用入口点以运行 FastAPI 服务器。
if __name__ == "__main__":
import uvicorn
logger.info(f"Starting server on port {AIP_HTTP_PORT}")
logger.info(f"Health check route: {AIP_HEALTH_ROUTE}")
logger.info(f"Predict route: {AIP_PREDICT_ROUTE}")
uvicorn.run(app, host="0.0.0.0", port=AIP_HTTP_PORT)现在,你已经有了一个完整的 FastAPI 应用,可以处理 YOLO26 推理请求。你可以先安装依赖项,然后使用 uv 等工具运行服务器,在本地进行测试。
# Install dependencies
uv pip install -e .
# Run the FastAPI server directly
uv run src/main.py要测试服务器,可以使用 cURL 查询 /health 和 /predict 端点。在 tests 文件夹中放入一张测试图像。然后在终端中运行以下命令:
# Test health endpoint
curl http://localhost:8080/health
# Test predict endpoint with base64 encoded image
curl -X POST -H "Content-Type: application/json" -d "{\"instances\": [{\"image\": \"$(base64 -i tests/test_image.jpg)\"}]}" http://localhost:8080/predict你应该会收到包含检测对象的 JSON 响应。首次请求可能会有短暂延迟,因为 Ultralytics 需要拉取并加载 YOLO26 模型。
2. 使用你的应用扩展 Ultralytics Docker 镜像#
Ultralytics 提供了多个 Docker 镜像,可作为应用镜像的基础。这些镜像包含 Ultralytics 和所需的 CUDA 库,因此你只需在其上添加应用即可。
要充分发挥 Ultralytics YOLO 模型的能力,进行 GPU 推理时应选择针对 CUDA 优化的镜像。不过,如果 CPU 推理足以满足任务需求,也可以选择仅包含 CPU 的镜像,以节省计算资源:
- Dockerfile:针对 YOLO26 单 GPU/多 GPU 训练和推理优化的镜像。
- Dockerfile-cpu:仅包含 CPU、用于 YOLO26 推理的镜像。
为你的应用创建 Docker 镜像#
在项目根目录中创建一个 Dockerfile,并添加以下内容:
# Extends official Ultralytics Docker image for YOLO26
FROM ultralytics/ultralytics:latest
ENV PYTHONUNBUFFERED=1
# Install FastAPI and dependencies
RUN uv pip install --system fastapi[all] uvicorn[standard] loguru
WORKDIR /app
COPY src/ ./src/
COPY pyproject.toml ./
# Install the application package
RUN uv pip install --system -e .
RUN mkdir -p /app/logs
ENV PYTHONPATH=/app/src
# Port for Vertex AI
EXPOSE 8080
# Start the inference server
ENTRYPOINT ["python", "src/main.py"]此示例使用 Ultralytics 官方 Docker 镜像 ultralytics:latest 作为基础镜像。该镜像已包含 YOLO26 模型和所有必需的依赖项。服务器入口点与我们在本地测试 FastAPI 应用时使用的相同。
构建并测试 Docker 镜像#
现在,你可以使用以下命令构建 Docker 镜像:
docker build --platform linux/amd64 -t IMAGE_NAME:IMAGE_VERSION .将 IMAGE_NAME 和 IMAGE_VERSION 替换为你需要的值,例如 yolo26-fastapi:0.1。请注意,如果要部署到 Vertex AI,必须为 linux/amd64 架构构建镜像。如果你在 Apple Silicon Mac 或其他非 x86 架构的设备上构建镜像,则需要显式设置 --platform 参数。
镜像构建完成后,你可以在本地测试 Docker 镜像:
docker run --platform linux/amd64 -p 8080:8080 IMAGE_NAME:IMAGE_VERSION现在,你的 Docker 容器正在端口 8080 上运行 FastAPI 服务器,可以接收推理请求。你可以使用之前相同的 cURL 命令测试 /health 和 /predict 端点:
# Test health endpoint
curl http://localhost:8080/health
# Test predict endpoint with base64 encoded image
curl -X POST -H "Content-Type: application/json" -d "{\"instances\": [{\"image\": \"$(base64 -i tests/test_image.jpg)\"}]}" http://localhost:8080/predict3. 将 Docker 镜像上传到 GCP Artifact Registry#
要在 Vertex AI 中导入容器化模型,你需要将 Docker 镜像上传到 Google Cloud Artifact Registry。如果你还没有 Artifact Registry 仓库,需要先创建一个。
在 Google Cloud Artifact Registry 中创建仓库#
在 Google Cloud Console 中打开 Artifact Registry 页面。如果你是首次使用 Artifact Registry,系统可能会提示你先启用 Artifact Registry API。
- 选择“创建仓库”。
- 输入仓库名称。选择所需区域,其他选项使用默认设置,除非你需要进行特定更改。
区域选择可能会影响机器的可用性,以及非 Enterprise 用户可使用的某些计算资源限制。更多信息请参阅 Vertex AI 官方文档:Vertex AI 配额和限制
- 创建仓库后,请将 PROJECT_ID、Location(Region)和 Repository Name 保存到你的密钥库或
.env文件中。之后为 Docker 镜像添加标签并将其推送到 Artifact Registry 时会用到这些信息。
对 Artifact Registry 验证 Docker 身份#
对刚创建的 Artifact Registry 仓库验证 Docker 客户端身份。在终端中运行以下命令:
gcloud auth configure-docker YOUR_REGION-docker.pkg.dev为镜像添加标签并推送到 Artifact Registry#
为 Docker 镜像添加标签,并将其推送到 Google Artifact Registry。
建议每次更新镜像时都使用唯一标签。包括 Vertex AI 在内的大多数 GCP 服务都依赖镜像标签进行自动版本控制和扩缩容,因此最好使用语义化版本号或基于日期的标签。
使用 Artifact Registry 仓库 URL 为镜像添加标签。将占位符替换为之前保存的值。
docker tag IMAGE_NAME:IMAGE_VERSION YOUR_REGION-docker.pkg.dev/YOUR_PROJECT_ID/YOUR_REPOSITORY_NAME/IMAGE_NAME:IMAGE_VERSION将带标签的镜像推送到 Artifact Registry 仓库。
docker push YOUR_REGION-docker.pkg.dev/YOUR_PROJECT_ID/YOUR_REPOSITORY_NAME/IMAGE_NAME:IMAGE_VERSION等待进程完成。现在应该可以在 Artifact Registry 仓库中看到该图像。
如需了解如何在 Artifact Registry 中处理图像的详细说明,请参阅 Artifact Registry 文档:推送和拉取图像。
4. 在 Vertex AI 中导入模型#
使用刚刚推送的 Docker 镜像,现在可以在 Vertex AI 中导入模型。
-
在 Google Cloud 导航菜单中,依次前往 Vertex AI > Model Registry。或者,在 Google Cloud Console 顶部的搜索栏中搜索“Vertex AI”。
-
点击“导入”。
-
选择“作为新模型导入”。
-
选择区域。你可以选择与 Artifact Registry 仓库相同的区域,但应根据所在区域可用的机器类型和配额来决定。
-
选择“导入现有模型容器”。
-
在“容器映像”字段中,浏览你之前创建的 Artifact Registry 仓库,然后选择刚刚推送的映像。
-
向下滚动到“环境变量”部分,输入你在 FastAPI 应用中定义的 predict 和 health 端点以及端口。
-
点击“导入”。Vertex AI 需要几分钟来注册模型并为部署做好准备。导入完成后,你会收到电子邮件通知。
5. 创建 Vertex AI Endpoint 并部署模型#
在 Vertex AI 术语中,Endpoint 指已部署的模型,因为它代表你发送推理请求的 HTTP 端点;而 Model 则是存储在 Model Registry 中经过训练的机器学习制品。
要部署模型,你需要在 Vertex AI 中创建一个 Endpoint。
-
在 Vertex AI 导航菜单中,前往 Endpoints。选择导入模型时使用的区域。点击 Create。
-
输入端点名称。
-
对于访问权限,Vertex AI 建议使用私有 Vertex AI 端点。除了安全方面的优势外,选择私有端点还可以获得更高的负载上限;不过,你需要配置 VPC 网络和防火墙规则,才能允许访问该端点。请参阅 Vertex AI 文档,了解有关私有端点的更多说明。
-
点击“继续”。
-
在“模型设置”对话框中,选择你之前导入的模型。现在,你可以为模型配置机器类型、内存和 GPU 设置。如果预计推理负载较高,请预留充足的内存,以确保没有 I/O 瓶颈,从而保障 YOLO26 的正常性能。
-
在“加速器类型”中,选择要用于推理的 GPU 类型。如果不确定要选择哪种 GPU,可以先使用 NVIDIA T4,它支持 CUDA。
请注意,某些区域的计算配额非常有限,因此你可能无法在所在区域选择某些机器类型或 GPU。如果这会造成严重影响,请将部署区域更改为配额更高的区域。详情请参阅 Vertex AI 官方文档:Vertex AI 配额和限制。
- 选择机器类型后,你可以点击 Continue。此时,你可以选择启用 Vertex AI 中的模型监控——这项附加服务会跟踪模型性能并提供模型行为洞察。此功能为可选项,且会产生额外费用,请根据需求选择。点击 Create。
Vertex AI 部署模型需要几分钟时间(某些区域最长需要 30 分钟)。部署完成后,你会收到电子邮件通知。
6. 测试已部署的模型#
部署完成后,Vertex AI 会提供一个示例 API 界面供你测试模型。
要测试远程推理,你可以使用提供的 cURL 命令,也可以创建另一个 Python 客户端库,向已部署的模型发送请求。请记住,在将图像发送到 /predict 端点之前,需要先将图像编码为 base64。
与本地测试类似,首次请求可能会有短暂延迟,因为 Ultralytics 需要在运行中的容器内拉取并加载 YOLO26 模型。
你已成功在 Google Cloud Vertex AI 上使用 Ultralytics 部署预训练 YOLO26 模型。
常见问题#
可以;不过,你需要先将模型导出为与 Vertex AI 兼容的格式,例如 TensorFlow、Scikit-learn 或 XGBoost。Google Cloud 提供了在 Vertex 上运行
.pt模型的指南,其中全面介绍了转换流程:在 Vertex AI 上运行 PyTorch 模型。请注意,最终配置将仅依赖 Vertex AI 标准服务层,不支持 Ultralytics 框架的高级功能。由于 Vertex AI 完全支持容器化模型,并可根据部署配置自动扩缩,因此无需将模型转换为其他格式,即可充分利用 Ultralytics YOLO 模型的全部能力。
FastAPI 可为推理工作负载提供高吞吐量。异步支持让你能够同时处理多个请求,而不会阻塞主线程;在提供计算机视觉模型服务时,这一点至关重要。
FastAPI 自动验证请求和响应,可减少生产环境推理服务中的运行时错误。对于输入格式一致性至关重要的目标检测 API,这尤其有价值。
FastAPI 为推理流程带来的计算开销极小,因此可以将更多资源留给模型执行和图像处理任务。
FastAPI 还支持 SSE(服务器发送事件),适用于流式推理场景。
这实际上是 Google Cloud Platform 的一项灵活性功能:你使用的每项服务都需要选择一个区域。对于在 Vertex AI 上部署容器化模型这一任务,最重要的是为 Model Registry 选择区域。该区域将决定模型部署可用的机器类型和配额。
此外,如果你要扩展此配置,并在 Cloud Storage 或 BigQuery 中存储预测数据或结果,则需要使用与 Model Registry 相同的区域,以最大限度地降低延迟并确保数据访问具有高吞吐量。