Environment
-
iPhone 17 Pro Max, iOS 27.0.1 ([build])
-
Reference: MacBook Pro (M3 Pro), macOS 26.4.1
-
ML Program model, MLComputeUnitsCPUAndNeuralEngine. A single prediction takes about 2.8 ms on the ANE.
Description
On iOS 27, Core ML predictions submitted to the Neural Engine from multiple threads at the same time no longer overlap. The ANE runs them strictly one after another. Adding threads gives no extra throughput, and per-call latency grows linearly with the number of requests in flight. The same code still gets a clear concurrency gain on macOS 26.4.1, and our app's ANE-parallel path on this iPhone became noticeably slower after the iOS 27 update.
On iOS 27, latency is about N × single-call time, which means requests are queued and never overlap. Running two different models concurrently on the ANE shows the same behavior.
What we tried, with no change in result
- A separate MLModel instance per thread.
- The async
predictionFromFeatures:options:completionHandler:API.
Questions
Is this serialization an intended change in iOS 27? Is there a supported way to get concurrent ANE execution back?