We ported a complete Stable Diffusion XL (Lightning) pipeline to Core AI which we previously had on Core ML. The pipeline comprises text encoders, a ViT-H image encoder, ControlNet, the UNet and the VAE decoder. It works very well: on iPhone 17 Pro and M1 iPad Pro it generates 1.4–1.6x faster than our Core ML pipeline at the same image quality. However, along the way, we hit seven issues. Each one is filed with a small self-contained repro. We're posting them here together for visibility, with our workarounds, in case others run into the same problems.
Ahead-of-time compilation (coreai-build compile)
Each of these gives wrong results on the Neural Engine with no error. The same .aimodel specialized on the device is correct.
- FB25067476: A 3×3 fp16 conv with a large output returns a wrong bottom half (16.7 dB vs 60.7 dB for the top half). This depends on the chip: 256×512×512 fails on M3 Pro and A19 Pro, while 256×384×384 fails only on M1. Our workaround: split such convs by output channels so that no single output exceeds about 128×384×384.
- FB25067497: A nearest 2× upsample followed by a 3×3 fp16 conv returns output unrelated to the correct result (correlation 0.02) on all three chips. The upsample alone and the conv alone are correct.
- Our workaround:* write the upsample as
repeat_interleave(4)+pixel_shuffle(2).
- Our workaround:* write the upsample as
- FB25067530: A 3×3 stride-2 conv with 8-bit palettized weights is wrong (16 dB) on all three chips. The same conv with fp16 weights, or palettized with stride 1, is correct.
- Our workaround:* a 2×2 pixel unshuffle followed by a 2×2 conv.
Caching and loading
- FB25066222: The cache for ahead-of-time compiled
.aimodelcfiles is keyed by the outer program only (main.hash), not the weights, so a weights-only model update silently runs the old weights. On iOS, the Neural Engine program cache is also keyed by function name per app, and survivesAIModelCache.deleteAll()and deleting the app. Our workaround: append a hash of the model's content to every function name. - FB25066664: On macOS, ahead-of-time builds of transformer models load 3–4× slower than on-device specialization, and parallel loads finish one at a time. Precompiling therefore made our Mac first launch slower (1152 s vs about 650 s), not faster.
Neural Engine runtime (Mac)
- FB25067019: The built-in GroupNorm returns wrong results on the Mac's Neural Engine for large inputs (
GroupNorm(32, 256)on 1×256×256×256: 21.7 dB), while the GPU and CPU are correct (84 dB). Our workaround: compute the means and variance explicitly. - FB25064649: A ViT-H vision tower with an 8-bit palettized 14×14/stride-14 patch conv aborts the process on the Mac's Neural Engine (MPSGraph ANE error -19). Our workaround: keep the patch conv in fp16, or use a pixel unshuffle followed by a 1×1 conv.
Environment
- macOS 27.0.1 (26A434), iOS/iPadOS 27.0.1 (24A446)
- Xcode 27.0, Metal Toolchain 27.1.266.1 (coreai-build 3600.83.1), coreai-torch 0.4.3
- Tested on M3 Pro, iPhone 17 Pro and M1 iPad Pro
We're happy to provide more data. We'd also like to hear whether others see the ahead-of-time issues on other chips.