Apple Core AI Integration#
coreai-core publishes macosx_26_0_arm64 and manylinux_2_34_x86_64 wheels, so export runs on Apple silicon Macs and x86_64 Linux with glibc 2.34 or newer. The exported .aimodel runs on iOS 27 and macOS 27. The Ultralytics iOS SDK (8.9.15 and later) and Flutter plugin (0.6.15 and later) load .aimodel assets as an opt-in on iOS 27 devices; Core ML remains their default.
Core AI is Apple's new framework for running neural networks directly on Apple silicon. It introduces the .aimodel model format, a modern Swift inference API, PyTorch-based conversion tools, ahead-of-time compilation, model specialization, and dedicated debugging and profiling tools.
Apple describes Core AI as the next evolution of on-device AI execution and the inference framework behind on-device Apple Intelligence. It is designed for current neural network architectures, from compact vision models to large generative models, and can schedule work across the CPU, GPU, and Apple Neural Engine (ANE).
Core AI is a new deployment path rather than a new name for Core ML. The frameworks use different model formats, conversion tools, runtime APIs, and application-integration patterns.
Core AI and Core ML Compared#
| Capability | Core AI | Core ML |
|---|---|---|
| Model artifact | .aimodel | .mlpackage or .mlmodel |
| Ultralytics export | Available with format=coreai | Available with format=coreml |
| Apple runtime API | AIModel, InferenceFunction, and NDArray | MLModel, often through VNCoreMLModel and VNCoreMLRequest |
| Conversion workflow | PyTorch torch.export through coreai-torch | TorchScript conversion through coremltools |
| Primary focus | Modern neural networks and generative AI | Broad machine learning deployment, including neural and non-neural models |
| Image integration | Applications prepare tensors or use Core AI image descriptors and buffers | Direct integration with the Vision framework for image scaling, orientation, and requests |
| Hardware | CPU, GPU, and Apple Neural Engine | CPU, GPU, and Apple Neural Engine |
| Model preparation | Specialization at installation or first use, with optional ahead-of-time compilation | Xcode or on-device model compilation |
| Custom operations | Custom Core AI lowerings and Metal kernels | Core ML custom layers and supported MIL operations |
| Deployment availability | New Apple operating-system generation; currently beta | Broad support across existing Apple operating systems |
| Ultralytics iOS and Flutter SDKs | Opt-in on iOS 27 and later devices | Fully supported, and the default |
Core ML remains the appropriate choice when an application needs broad device coverage, Vision framework integration, or model types such as decision trees and tabular pipelines. Apple continues to support Core ML and directs developers with non-neural model types to it.
How the Core AI Format Works#
The Core AI authoring workflow starts from a PyTorch model:
PyTorch model
↓ torch.export
ExportedProgram
↓ coreai-torch
Core AI program
↓ optimize and save
.aimodel
↓ specialize or compile ahead of time
Apple silicon executableApple's coreai-torch package converts a torch.export.ExportedProgram by lowering PyTorch ATen operations into Core AI operations. Unsupported operations can be implemented with a custom lowering or custom Metal kernel.
The resulting .aimodel is an unspecialized model asset. When an application prepares the model, Core AI specializes it for the target device. Applications can let this happen on first use, request specialization earlier, or ship an ahead-of-time compiled model to reduce initial loading time.
In Swift, applications load the asset with the Core AI framework, select an inference function, provide typed NDArray inputs, and receive named outputs. This is different from wrapping a Core ML model in a Vision request, so adopting Core AI requires an application runtime designed for .aimodel assets.
For implementation details, see Apple's documentation for AIModel, model specialization and caching, and ahead-of-time compilation.
Exporting YOLO26 Models to Core AI#
from ultralytics import YOLO
model = YOLO("yolo26n.pt")
model.export(format="coreai") # creates 'yolo26n.aimodel'
model.export(format="coreai", quantize=16) # FP16 asset
# Run the exported model
coreai_model = YOLO("yolo26n.aimodel")
results = coreai_model("https://ultralytics.com/images/bus.jpg")For the full argument list see Export mode. The graph is static: it is traced at
the imgsz given to export, so predict on that same size. Ultralytics metadata travels inside the
asset's own metadata.json, so class names, stride and task survive the round trip.
Choosing the head#
With nms=False, YOLO26 exports its end-to-end head, which selects detections inside the graph. Core AI has
no top-k primitive, so that selection lowers to a full sort and is charged a fixed cost at the Apple
Neural Engine partition boundary — about 1.7 ms, regardless of max_det. Exporting with
nms=None emits the raw (1, 84, 8400) predictions instead and leaves non-maximum suppression
to the predictor:
yolo export model=yolo26n.pt format=coreai nms=None quantize=16On an iPhone 17 Pro running iOS 27.0, YOLO26n at 640 measures 3.01 ms with the head in the graph and
1.28 ms without it (FP16, ahead-of-time compiled, three interleaved blocks of 50 iterations). Both
round-trip through YOLO(...) for inference. Use nms=False when a single graph call must return
finished detections, or keep the default nms=None for external NMS.
On iOS 27 or macOS 27, an application would then load and run the exported asset through Apple's Core AI Swift API. Exported assets use the entrypoint main, take a single images input of shape [batch, 3, imgsz, imgsz], and return output0 (instance segmentation models also return output1):
import CoreAI
let modelURL = Bundle.main.url(forResource: "yolo26n", withExtension: "aimodel")!
let model = try await AIModel(contentsOf: modelURL)
guard let function = try model.loadFunction(named: "main") else {
throw AppError.missingInferenceFunction
}
let outputs = try await function.run(inputs: ["images": imageTensor])Unlike the current Core ML and Vision workflow, the Core AI path in the Ultralytics iOS SDK does its own letterbox preprocessing and NDArray construction, reads the same Ultralytics metadata from the asset's metadata.json, and reuses the Core ML output decoders. It loads when an app passes an .aimodel path or an .aimodel.zip URL. Apple provides current API details in the Core AI framework documentation and working model examples in the Core AI models repository.
Measured Performance#
End-to-end single-image inference for YOLO26n FP16 (quantize=16) Core ML and Core AI exports with the default raw
head (nms=None) on a Mac mini with an Apple M4 (4 Performance and 6
Efficiency CPU cores, 10-core GPU, 16-core Neural Engine), 16 GB memory and macOS 27.0, using ultralytics 8.4.168,
coremltools 9.0 for Core ML inference, and coreai-torch 0.4.3 with coreai-core 1.0.0b3 for Core AI inference
on Python 3.13. Each cell shows the total time (preprocessing + inference + postprocessing) with the per-stage
split beneath it.
| Model | Task | size (pixels) | Core ML CPUCPU_ONLY(ms) | Core ML CPU + ANE preferredCPU_AND_NE(ms) | Core AI CPUcpu_only()(ms) | Core AI CPU + ANE preferredneural_engine()(ms) |
|---|---|---|---|---|---|---|
| YOLO26n | Detect | 640 | 14.4 0.6 / 13.4 / 0.4 | 7.6 0.6 / 6.7 / 0.4 | 16.8 0.6 / 15.9 / 0.3 | 2.8 0.6 / 2.0 / 0.2 |
| YOLO26n-seg | Segment | 640 | 18.7 0.6 / 16.5 / 1.5 | 9.2 0.6 / 7.1 / 1.5 | 25.3 0.6 / 23.2 / 1.5 | 5.1 0.6 / 3.1 / 1.4 |
| YOLO26n-sem | Semantic | 640 | 33.9 1.3 / 32.2 / 0.4 | 73.7 1.4 / 71.9 / 0.4 | 47.1 1.2 / 38.5 / 7.4 | 18.7 1.2 / 11.1 / 6.4 |
| YOLO26n-depth | Depth | 640 | 36.8 0.8 / 35.5 / 0.5 | 12.1 0.9 / 10.7 / 0.5 | 40.2 0.8 / 39.0 / 0.5 | 7.5 0.7 / 6.3 / 0.5 |
| YOLO26n-cls | Classify | 224 | 3.8 1.9 / 1.8 / 0.0 | 3.3 1.9 / 1.4 / 0.0 | 3.0 1.9 / 1.0 / 0.0 | 2.4 1.9 / 0.6 / 0.0 |
| YOLO26n-pose | Pose | 640 | 15.5 0.6 / 14.6 / 0.3 | 7.0 0.6 / 6.2 / 0.3 | 17.8 0.5 / 17.0 / 0.3 | 2.7 0.5 / 2.0 / 0.2 |
| YOLO26n-obb | OBB | 640 | 32.7 1.3 / 31.1 / 0.2 | 16.9 1.5 / 15.2 / 0.2 | 37.2 1.1 / 35.9 / 0.2 | 5.7 1.3 / 4.3 / 0.1 |
- Speed values are single-image burst latencies: the mean of 15
predictcalls after 3 warmup calls onbus.jpgthrough the Ultralytics Python API, with each model and compute unit in a fresh process. CPU/accelerator order alternated between tasks in one sequential sweep. Core ML rows load withcoremltools.ComputeUnit.CPU_ONLYorCPU_AND_NE; Core AI rows specialize withSpecializationOptions.cpu_only()orSpecializationOptions.from_preferred_compute_unit_kind(ComputeUnitKind.neural_engine()), with final operation placement controlled by each framework. - Detect, segment, classify, pose and OBB returned the same predictions in both formats on every compute unit. The Core ML FP16 semantic model runs slower with the Neural Engine preferred than on CPU only on this Mac, and Core AI semantic postprocessing takes 6.4 to 7.4 ms against 0.4 ms for Core ML.
- Compare the on-device iPhone 17 Pro results in the CoreML integration.
Advantages of Core AI#
Core AI offers several promising advantages for future Ultralytics deployment:
- Modern PyTorch export path: Conversion starts from
torch.export, preserving a more expressive PyTorch graph than the tracing workflow used by many existing exporters. - Fine-grained runtime control: Applications can manage specialization, compiled-model caches, inference functions, memory, and compute placement.
- Advanced model support: Stateful execution, dynamic shapes, multiple functions in one artifact, and custom Metal kernels are designed for modern vision and generative architectures.
- Dedicated developer tools: The Core AI Debugger can inspect graphs and tensor values and trace them back to the originating Python code. Xcode and Instruments provide runtime profiling.
- Zero-copy opportunities: Core AI exposes storage and buffer controls intended to reduce copies between camera, graphics, and inference workloads.
- Apple-silicon optimization: Device specialization lets Apple optimize a model for the CPU, GPU, and Neural Engine available on the specific device.
- Flexible compression: Apple's Core AI Optimization tools support quantization, palettization, and pruning, including low-bit weight formats.
These capabilities could be particularly useful for future YOLO models with dynamic execution, larger multimodal components, or custom operations that do not map cleanly to existing Core ML operations.
Current Disadvantages and Limitations#
Core AI is not currently a replacement for the production Core ML path:
- New operating systems required: The public framework targets the iOS 27 and macOS 27 generation, while Core ML supports a much larger installed base.
- Beta software: Apple's Core AI framework and parts of its Python toolchain are still preliminary and may change before their stable releases.
- Narrower export environment:
coreai-torchcurrently requires Python 3.11 to 3.14, plus recent PyTorch versions, which is much narrower than Ultralytics' supported Python and PyTorch range. - Limited export platforms:
coreai-corepublishesmacosx_26_0_arm64andmanylinux_2_34_x86_64wheels only, soformat=coreaineeds an Apple silicon Mac on macOS 26 or later, or x86_64 Linux with glibc 2.34 or newer (for example Ubuntu 22.04+). Running an.aimodelstill requires Apple hardware. - Opt-in in the Ultralytics SDKs, not the default: On an iPhone 17 Pro the on-device comparison puts Core AI level with Core ML for the full pipeline rather than ahead, with semantic, depth and CPU-only inference slower and FP16 assets about twice the download size, so the iOS SDK and Flutter plugin keep Core ML as the default.
- Use the raw head for the SDKs: With
nms=False, YOLO26n takes about twice as long on Core AI as on Core ML (3.06 against 1.53 ms in one comparison run), and the FP16 end-to-end pose model returns no detections under Core AI's default placement on iOS 27.0 (apple/coreai-torch#115). The raw head (nms=None, the default) avoids both, and the SDKs run NMS in Swift. See Choosing the head. - No iOS Simulator runtime: The iOS Simulator SDK does not include Core AI.
- Application migration required: A
.aimodelcannot be substituted for an.mlpackage; model loading, preprocessing, inference calls, metadata handling, and output decoding need a Core AI implementation outside the Ultralytics SDKs, which provide one. - Limited production evidence: Performance, power use, first-run specialization time, accuracy, and compression need validation across the supported YOLO task and device matrix.
- No NMS pipeline: Core ML can package an NMS stage for older YOLO detection models. Core AI exports raw one-to-many predictions by default; use
nms=Falsefor YOLO26's NMS-free head. Embedded NMS (nms=True) anddynamic=Trueare not supported.coreai-torchhas no lowering fortorchvision::nms, so NMS stays on the host. - Fixed input size: The exported graph is traced at one
imgszand has no dynamic shapes, so predict at the size it was exported with. - FP16 assets can abort on load: Some FP16
.aimodelassets fail to load their Apple Neural Engine program and MPSGraph raises a failed assertion, which ends the process rather than falling back. This happens inside Apple's runtime, before any Ultralytics code runs, and the same asset loads with a CPU-only specialization. Fall back to FP32 if an asset does this; the FP16 SDK assets checked on an iPhone 17 Pro (every nano model, detect at every size, and the largest model of each task) loaded without an abort.
Which Apple Format Should You Use?#
Use Core ML today when you need:
- Deployment across current and older Apple operating systems
- The default path of the Ultralytics iOS or Flutter SDK
- Vision framework image handling
- Tested FP16 and INT8 YOLO deployment
- Embedded NMS for compatible legacy detection models
Evaluate Core AI when you can require iOS 27 or macOS 27 and need:
- The newest Apple on-device neural network runtime
- Explicit specialization and cache management
- Advanced dynamic or stateful model execution
- Custom Core AI operations or Metal kernels
- Detailed Core AI graph debugging and runtime profiling
Core ML and Core AI are expected to coexist while applications transition. Supporting Core AI does not immediately remove the need for Core ML because their deployment targets and application contracts differ.
Ultralytics Roadmap#
The dedicated coreai export target is implemented: export and numerical validation cover the supported YOLO26 task models and run continuously in Ultralytics CI on macOS 26, and FP16 latency is measured on-device. The remaining roadmap before Core AI reaches parity with the Core ML path:
- Closing the end-to-end head latency gap on the Apple Neural Engine (apple/coreai-torch#66) and the Neural Engine correctness issues (apple/coreai-torch#115, #116).
- INT8 and palettized Core AI assets; the SDK assets are FP16, about twice the download size of the INT8 Core ML assets.
- Stable Apple framework and conversion-tool releases (the iOS 27 and macOS 27 generation is currently in beta).
- Memory, power, and specialization benchmarks across the supported device matrix.
The Ultralytics iOS SDK and Flutter plugin load Core AI as an opt-in on iOS 27 and later; Core ML remains their default and the recommended target for coverage below iOS 27. Follow the Ultralytics roadmap and release notes for the remaining items.
Additional Resources#
- Apple Core AI overview
- Core AI framework documentation
- Core AI PyTorch Extensions
- Core AI Optimization
- Apple Core AI models repository
- Ultralytics Core ML integration
FAQ#
Yes. Export with
model.export(format="coreai")oryolo export format=coreaion an Apple silicon Mac running macOS 26 or later, or on x86_64 Linux with glibc 2.34 or newer; the exported.aimodelruns on iOS 27 and macOS 27. The Ultralytics iOS and Flutter SDKs load it as an opt-in on iOS 27 devices. For their default path, and for operating systems below that generation, export Core ML.mlpackagefiles withformat="coreml".Not immediately. Core AI is Apple's newer path for modern neural networks, while Core ML remains supported and provides broader operating-system coverage, Vision integration, and non-neural model support.
No. They contain different model representations and are loaded by different frameworks. Conversion must start from the source model through the appropriate Apple toolchain.
The initial integration is expected to coexist with Core ML. Any future replacement decision depends on operating-system adoption, stable tooling, and on-device performance.