iPhone 18 Pro: First Core ML load on Neural Engine up to 65× slower than on iPhone 17 Pro

On iPhone 18 Pro (A20 Pro), the first MLModel load of a model on the Neural Engine can take minutes where an iPhone 17 Pro needs seconds. Cached loads are instant, and once loaded the model runs faster than on the iPhone 17 Pro, so only the on-device specialization is affected. This looks like the same problem as A20 pro devices take too long on CoreML model specialization and iPhone 18 pro takes very long time to load models. The "loading" time there is likely this specialization too: it happens inside MLModel(contentsOf:) for an already compiled .mlmodelc.

I can reproduce it reliably with attention in the form Apple recommends for the Neural Engine (split einsum, as in Deploying Transformers on the Apple Neural Engine). In this example most of the extra time comes from the softmax over the key axis, which in that layout is the channel axis. Other ops and model types may well be affected too; this is simply the smallest case we could build that shows the problem clearly. The repro uses four self-attention layers (10 heads × 64, 4096 tokens, fp16) built with coremltools and loaded cold with .cpuAndNeuralEngine. MLComputePlan places every op on the Neural Engine on all three devices. iOS 27.0.1 (24A446) on both iPhones:

                                     iPhone 18 Pro        iPhone 17 Pro        M3 Pro Mac
                                     cold     predict     cold    predict      cold    predict
Split-einsum attention, 4 layers     53.4 s   21.6 ms     0.8 s   111.4 ms     1.0 s   105.8 ms
Q.K einsum only                       5.1 s   24.0 ms     1.2 s    36.8 ms     1.4 s    41.8 ms
Q.K einsum + softmax                 36.1 s   61.5 ms     1.7 s   141.0 ms     2.0 s   128.3 ms
Scores keys-last (softmax last axis) 10.5 s  109.0 ms     0.6 s   198.4 ms     0.7 s   180.2 ms
scaled_dot_product_attention          1.4 s  140.4 ms     0.4 s   184.9 ms     0.5 s   148.1 ms

Cached loads take 0.01–0.03 s everywhere.

What I observed:

  • Compile time grows linearly with the number of attention layers.
  • Every alternative attention layout I tried that compiles quickly runs 5–6× slower on the A20 Pro, so I found no workaround that keeps its speed.
  • During the slow load, ANECompilerService runs almost entirely on the efficiency cores. Requesting a higher QoS, MLOptimizationHints (.fastPrediction, reshapeFrequency), computeUnits = .all and dense instead of palettized weights don't help.
  • Loading several models in parallel additionally gets ANECompilerService killed by jetsam (per-process-limit).
  • The specialization cache doesn't survive reinstalling the app.

In our app (Stable Diffusion XL), each UNet chunk with attention takes 110–150 s instead of 6–16 s, so the first launch takes about 25 minutes in the foreground, and in our tests again after every reinstall. The same networks converted to Core AI specialize normally on this device (see our Core AI porting notes).

Filed as FB25106520 with a self-contained repro: a coremltools script, an iOS app and a Mac runner, plus results from all three devices.

Environment: iPhone 18 Pro (iPhone19,2) and iPhone 17 Pro (iPhone18,1) on iOS 27.0.1 (24A446), M3 Pro on macOS 27.0.1 (26A434), Xcode 27.0, coremltools 8.3.0.

Is this a known A20 Pro issue, and is there a recommended way to avoid it until it's fixed? Falling back to the GPU or CPU isn't a good option for us: it avoids the slow load, but gives up the Neural Engine's efficiency, which matters for on-device generation on a phone, and its speed (a full SDXL UNet step on the iPhone 18 Pro takes 509 ms on the Neural Engine vs 1,639 ms with .cpuAndGPU). We'd also like to hear from others which ops their slow first loads involve, to see whether this is specific to attention and softmax or broader.

Happens to me as well, I opened my thread here: https://developer.apple.com/forums/thread/848590

No response from Apple.

But this is definitely confirmed. Not sure how they missed this and didnt bother testing anything coreml related.

This also happens on M6 devices as well. great job Apple.

Some further info on M6, (but also expeiencing the same issue on A20):

Same issue here on an M6 Mac mini (Mac18,5, Core AI architecture h18g), macOS 27.0 (26A425), Xcode 27.0, so it isn’t limited to A20 Pro. In our case, attention isn’t involved: the slow graphs are long unrolled LSTM recurrences.

Model: a music source-separation network split into 22 FP16 Core ML components, loaded with .cpuAndNeuralEngine. MLComputePlan places every device-assigned operation on the Neural Engine.

Cold first load on M6, per component (cached reloads are 0.01–0.04 s):

6 stage prefixes (57-step bidirectional LSTM): 12.7–12.9 s each
3 recurrent cells (52 steps): 8.0–8.2 s each
3 recurrent cells (34 steps): 4.9–5.0 s each
Conv encoders/decoder: 1.1–3.1 s each
Whole model (22 components): about 123–145 s
The same compiled .mlmodelc files on M4:

An isolated 26-step LSTM: 1.3 s on M4 vs 4.4 s on M6.
The 26/17-step variant of the whole model: about 25 s on M4 vs about 103 s on M6.
Compile time grows with the number of unrolled steps. For one LSTM: 8 steps 1.2 s, 19 steps 3.1 s, 57 steps 11.7 s. Almost all of it is ANECompilerService CPU time. Loading components in parallel only reduced the total from 145 s to 121 s, because the compiler largely processes them one at a time.

Compiler crash: a 338-step LSTM graph that loads in 21 s on M4 crashes ANECompilerService on M6, with “Thread stack size exceeded due to excessive recursion” in SubgraphIdentification::ComputeReachableMap (called from ComputeReachabilityMaps / AddBoundaryTensorsForAcyclicCluster). Core ML retries, caches nothing, and every later load takes about 113 s again.

Core AI didn’t help us. For the 57-step prefix component:

Core ML: 12.8 s
Core AI precompiled with coreai-build compile --architecture h18g --preferred-compute neural-engine: 16.3 s (still runs ANECompilerService on first load)
Core AI source .aimodel specialized on device: 161.6 s

Just to note, model specialization on an iphone 16 pro max finished in about 30 seconds (iOS 27)

iPhone 18 Pro: First Core ML load on Neural Engine up to 65× slower than on iPhone 17 Pro
 
 
Q