Posts under Machine Learning & AI topic

Post

Replies

Boosts

Views

Activity

iPhone 18 Pro: First Core ML load on Neural Engine up to 65× slower than on iPhone 17 Pro
On iPhone 18 Pro (A20 Pro), the first MLModel load of a model on the Neural Engine can take minutes where an iPhone 17 Pro needs seconds. Cached loads are instant, and once loaded the model runs faster than on the iPhone 17 Pro, so only the on-device specialization is affected. This looks like the same problem as A20 pro devices take too long on CoreML model specialization and iPhone 18 pro takes very long time to load models. The "loading" time there is likely this specialization too: it happens inside MLModel(contentsOf:) for an already compiled .mlmodelc. I can reproduce it reliably with attention in the form Apple recommends for the Neural Engine (split einsum, as in Deploying Transformers on the Apple Neural Engine). In this example most of the extra time comes from the softmax over the key axis, which in that layout is the channel axis. Other ops and model types may well be affected too; this is simply the smallest case we could build that shows the problem clearly. The repro uses four self-attention layers (10 heads × 64, 4096 tokens, fp16) built with coremltools and loaded cold with .cpuAndNeuralEngine. MLComputePlan places every op on the Neural Engine on all three devices. iOS 27.0.1 (24A446) on both iPhones: iPhone 18 Pro iPhone 17 Pro M3 Pro Mac cold predict cold predict cold predict Split-einsum attention, 4 layers 53.4 s 21.6 ms 0.8 s 111.4 ms 1.0 s 105.8 ms Q.K einsum only 5.1 s 24.0 ms 1.2 s 36.8 ms 1.4 s 41.8 ms Q.K einsum + softmax 36.1 s 61.5 ms 1.7 s 141.0 ms 2.0 s 128.3 ms Scores keys-last (softmax last axis) 10.5 s 109.0 ms 0.6 s 198.4 ms 0.7 s 180.2 ms scaled_dot_product_attention 1.4 s 140.4 ms 0.4 s 184.9 ms 0.5 s 148.1 ms Cached loads take 0.01–0.03 s everywhere. What I observed: Compile time grows linearly with the number of attention layers. Every alternative attention layout I tried that compiles quickly runs 5–6× slower on the A20 Pro, so I found no workaround that keeps its speed. During the slow load, ANECompilerService runs almost entirely on the efficiency cores. Requesting a higher QoS, MLOptimizationHints (.fastPrediction, reshapeFrequency), computeUnits = .all and dense instead of palettized weights don't help. Loading several models in parallel additionally gets ANECompilerService killed by jetsam (per-process-limit). The specialization cache doesn't survive reinstalling the app. In our app (Stable Diffusion XL), each UNet chunk with attention takes 110–150 s instead of 6–16 s, so the first launch takes about 25 minutes in the foreground, and in our tests again after every reinstall. The same networks converted to Core AI specialize normally on this device (see our Core AI porting notes). Filed as FB25106520 with a self-contained repro: a coremltools script, an iOS app and a Mac runner, plus results from all three devices. Environment: iPhone 18 Pro (iPhone19,2) and iPhone 17 Pro (iPhone18,1) on iOS 27.0.1 (24A446), M3 Pro on macOS 27.0.1 (26A434), Xcode 27.0, coremltools 8.3.0. Is this a known A20 Pro issue, and is there a recommended way to avoid it until it's fixed? Falling back to the GPU or CPU isn't a good option for us: it avoids the slow load, but gives up the Neural Engine's efficiency, which matters for on-device generation on a phone, and its speed (a full SDXL UNet step on the iPhone 18 Pro takes 509 ms on the Neural Engine vs 1,639 ms with .cpuAndGPU). We'd also like to hear from others which ops their slow first loads involve, to see whether this is specific to attention and softmax or broader.
3
4
368
17h
iOS 27: concurrent Core ML predictions on the Neural Engine are serialized
Environment iPhone 17 Pro Max, iOS 27.0.1 ([build]) Reference: MacBook Pro (M3 Pro), macOS 26.4.1 ML Program model, MLComputeUnitsCPUAndNeuralEngine. A single prediction takes about 2.8 ms on the ANE. Description On iOS 27, Core ML predictions submitted to the Neural Engine from multiple threads at the same time no longer overlap. The ANE runs them strictly one after another. Adding threads gives no extra throughput, and per-call latency grows linearly with the number of requests in flight. The same code still gets a clear concurrency gain on macOS 26.4.1, and our app's ANE-parallel path on this iPhone became noticeably slower after the iOS 27 update. On iOS 27, latency is about N × single-call time, which means requests are queued and never overlap. Running two different models concurrently on the ANE shows the same behavior. What we tried, with no change in result A separate MLModel instance per thread. The async predictionFromFeatures:options:completionHandler: API. Questions Is this serialization an intended change in iOS 27? Is there a supported way to get concurrent ANE execution back?
0
0
219
21h
iOS 27: English sentence embedding model unavailable on physical devices
I’m seeing NLEmbedding.sentenceEmbedding(for: .english) return nil on physical devices running iOS 27, both in debug builds under Xcode and in production. The same failure has been observed across thousands of users on iOS 27. One affected device runs iOS 27.0.1 (24A446). Minimal example with diagnostics: import NaturalLanguage let language = NLLanguage.english print("Current revision:", NLEmbedding.currentSentenceEmbeddingRevision(for: language)) print("Supported revisions:", Array(NLEmbedding.supportedSentenceEmbeddingRevisions(for: language))) if let model = NLEmbedding.sentenceEmbedding(for: language) { print("Loaded revision:", model.revision, "dimension:", model.dimension) } else { print("Sentence embedding model unavailable") } On affected devices, the current revision is 1 and supported revisions are [1], but model loading returns nil. The framework logs: Unable to locate Asset for sentence embedding model for local en. The failure occurs during model loading, before any call to vector(for:). This uses Apple’s built-in English sentence model, with no custom embedding file. Is this a known issue on iOS 27? Is there a supported way to request or restore the missing sentence-model asset, or a recommended recovery approach?
0
0
437
1d
PCC rate limits and blocking Siri AI as a result
Hello, I have PCC deployed in my 3 apps and my testing for these apps has caused me to hit the rate limit many times. When this rate limit is hit, Siri AI (as well as HomeKit PCC features) stops working completely for me until the rate limit is reset. I have been forced to scale back PCC and testing in my apps as this situation is quite opaque. My logs say I’m being rate limited when using PCC in these cases. Example log I see: 10:51:47.407 | RemoteManager executing remote streaming request 10:51:47.407 ModelCatalog | Looking up asset bundle com.apple.fm.language.instruct_server_v2.fm_api in Model Catalog 10:51:47.410 DaemonSession | Asset com.apple.fm.language.instruct_server_v2.fm_api has inference providers: 10:51:47.410 DaemonSession | - private-ml-client 10:51:47.411 ExecutionGroup | executing request C9B08798-6D68-4819-83C9-89B2EF5278C4 10:51:47.431 InferenceProviderConnection | InferenceProvider requestInference (C9B08798-6D68-4819-83C9-89B2EF5278C4) finished 10:51:47.431 RemoteManager | Request failed on private-ml-client with error: InferenceError::rateLimited::PrivateCloudComputeError: deniedDueToUserDeviceRateLimit, falling back 10:51:47.431 RemoteIPC Tasks | Failed to execute remote streamingRequest with error: InferenceError::rateLimited::PrivateCloudComputeError: deniedDueToUserDeviceRateLimit 10:51:56 modelmanagerd [com.apple.modelmanager:RemoteIPC Tasks] Failed to execute remote oneShotRequest with error: InferenceError::rateLimited::PrivateCloudComputeError: deniedDueToUserDev 11:23:39 privatecloudcomputed [com.apple.privatecloudcompute:RateLimiter] rate limit applied from cached denials What are the rate limits? How many requests a day? Per second? Also, I subscribed to the highest iCloud storage tier in hopes of higher rate limit but there doesn't seem to be any change. https://support.apple.com/en-us/127901 This issue started when iOS27 became publicly available - I was never rate limited when on the iOS27 beta.
2
0
825
1d
Crashed: AXSpeech EXC_BAD_ACCESS KERN_INVALID_ADDRESS 0x000056f023efbeb0
Application is getting Crashed: AXSpeech EXC_BAD_ACCESS KERN_INVALID_ADDRESS 0x000056f023efbeb0 Crashed: AXSpeech 0 libobjc.A.dylib 0x4820 objc_msgSend + 32 1 libsystem_trace.dylib 0x6c34 _os_log_fmt_flatten_object + 116 2 libsystem_trace.dylib 0x5344 _os_log_impl_flatten_and_send + 1884 3 libsystem_trace.dylib 0x4bd0 _os_log + 152 4 libsystem_trace.dylib 0x9c48 _os_log_error_impl + 24 5 TextToSpeech 0xd0a8c _pcre2_xclass_8 6 TextToSpeech 0x3bc04 TTSSpeechUnitTestingMode 7 TextToSpeech 0x3f128 TTSSpeechUnitTestingMode 8 AXCoreUtilities 0xad38 -[NSArray(AXExtras) ax_flatMappedArrayUsingBlock:] + 204 9 TextToSpeech 0x3eb18 TTSSpeechUnitTestingMode 10 TextToSpeech 0x3c948 TTSSpeechUnitTestingMode 11 TextToSpeech 0x48824 AXAVSpeechSynthesisVoiceFromTTSSpeechVoice 12 TextToSpeech 0x49804 AXAVSpeechSynthesisVoiceFromTTSSpeechVoice 13 Foundation 0xf6064 __NSThreadPerformPerform + 264 14 CoreFoundation 0x37acc CFRUNLOOP_IS_CALLING_OUT_TO_A_SOURCE0_PERFORM_FUNCTION + 28 15 CoreFoundation 0x36d48 __CFRunLoopDoSource0 + 176 16 CoreFoundation 0x354fc __CFRunLoopDoSources0 + 244 17 CoreFoundation 0x34238 __CFRunLoopRun + 828 18 CoreFoundation 0x33e18 CFRunLoopRunSpecific + 608 19 Foundation 0x2d4cc -[NSRunLoop(NSRunLoop) runMode:beforeDate:] + 212 20 TextToSpeech 0x24b88 TTSCFAttributedStringCreateStringByBracketingAttributeWithString 21 Foundation 0xb3154 NSThread__start + 732 com.livingMedia.AajTakiPhone_issue_3ceba855a8ad2d1af83655803dc13f70_crash_session_9081fa41ced440ae9a57c22cb432f312_DNE_0_v2_stacktrace.txt 22 libsystem_pthread.dylib 0x24d4 _pthread_start + 136 23 libsystem_pthread.dylib 0x1a10 thread_start + 8
5
1
2.6k
2d
iPhone 18 pro takes very long time to load models
I think this could be relate to previous post here https://developer.apple.com/forums/thread/848590 But in my testing it seems like most of the time spent was in loading but not compilation. Our app has multiple models but so far it seems like one of them has this issue. The model itself was only 2.2MB, but somehow loading took 50s after compilation? And on 17e it only took 2s to load. But I cannot attach the model here so please advice.
1
1
595
2d
A20 pro devices take too long on CoreML model specialization
Hi, I'm a ML developer specializing in local audio related models. I upgraded from a 16 pro max to an 18 pro max. I'm not sure what the reason is, but model specialization (happens only on the first load, cpu and neural engine compute units) is sometimes 2x - 5x longer on the 18 pro max vs the 16 pro max. Both devices are on iOS 27. For example, I have a stem separation model running, and the 18 pro max took 171 seconds (including inference) and the 16 pro max took 37 seconds (including inference). 18 pro max inference is ~55x RTFx and the 16 pro max is ~40 RTFx for this particular model, but I've noticed it as well in TTS models etc. Please look into this, I know the 18 pro max has double the neural engine cores, so I'm not sure if the complexity makes specialization take longer. Subsequent loads are great, same high rtfx for the 18 PM (only 0.5s load), and this is where it beats the 16 PM handily. I attached 2 images showing the 18 PM on a first load (model specialization likely) vs a subsequent run. Please look into this. (Note this is not just my implementation, I tested a range of different models, all with the same outcome).
1
5
561
2d
Core AI on iOS/macOS 27: seven issues found while porting an SDXL pipeline
We ported a complete Stable Diffusion XL (Lightning) pipeline to Core AI which we previously had on Core ML. The pipeline comprises text encoders, a ViT-H image encoder, ControlNet, the UNet and the VAE decoder. It works very well: on iPhone 17 Pro and M1 iPad Pro it generates 1.4–1.6x faster than our Core ML pipeline at the same image quality. However, along the way, we hit seven issues. Each one is filed with a small self-contained repro. We're posting them here together for visibility, with our workarounds, in case others run into the same problems. Ahead-of-time compilation (coreai-build compile) Each of these gives wrong results on the Neural Engine with no error. The same .aimodel specialized on the device is correct. FB25067476: A 3×3 fp16 conv with a large output returns a wrong bottom half (16.7 dB vs 60.7 dB for the top half). This depends on the chip: 256×512×512 fails on M3 Pro and A19 Pro, while 256×384×384 fails only on M1. Our workaround: split such convs by output channels so that no single output exceeds about 128×384×384. FB25067497: A nearest 2× upsample followed by a 3×3 fp16 conv returns output unrelated to the correct result (correlation 0.02) on all three chips. The upsample alone and the conv alone are correct. Our workaround:* write the upsample as repeat_interleave(4) + pixel_shuffle(2). FB25067530: A 3×3 stride-2 conv with 8-bit palettized weights is wrong (16 dB) on all three chips. The same conv with fp16 weights, or palettized with stride 1, is correct. Our workaround:* a 2×2 pixel unshuffle followed by a 2×2 conv. Caching and loading FB25066222: The cache for ahead-of-time compiled .aimodelc files is keyed by the outer program only (main.hash), not the weights, so a weights-only model update silently runs the old weights. On iOS, the Neural Engine program cache is also keyed by function name per app, and survives AIModelCache.deleteAll() and deleting the app. Our workaround: append a hash of the model's content to every function name. FB25066664: On macOS, ahead-of-time builds of transformer models load 3–4× slower than on-device specialization, and parallel loads finish one at a time. Precompiling therefore made our Mac first launch slower (1152 s vs about 650 s), not faster. Neural Engine runtime (Mac) FB25067019: The built-in GroupNorm returns wrong results on the Mac's Neural Engine for large inputs (GroupNorm(32, 256) on 1×256×256×256: 21.7 dB), while the GPU and CPU are correct (84 dB). Our workaround: compute the means and variance explicitly. FB25064649: A ViT-H vision tower with an 8-bit palettized 14×14/stride-14 patch conv aborts the process on the Mac's Neural Engine (MPSGraph ANE error -19). Our workaround: keep the patch conv in fp16, or use a pixel unshuffle followed by a 1×1 conv. Environment macOS 27.0.1 (26A434), iOS/iPadOS 27.0.1 (24A446) Xcode 27.0, Metal Toolchain 27.1.266.1 (coreai-build 3600.83.1), coreai-torch 0.4.3 Tested on M3 Pro, iPhone 17 Pro and M1 iPad Pro We're happy to provide more data. We'd also like to hear whether others see the ahead-of-time issues on other chips.
0
3
70
2d
iOS 27.2 beta: App Shortcut phrase fails, but a named personal shortcut runs the same intent
I'm investigating a Siri invocation failure with an explicitly declared App Shortcut phrase. The same underlying App Intent works when run directly in Shortcuts and when invoked by the name of a saved personal shortcut. Environment: iPhone 17 Pro Max, iOS 27.2 beta (24B5099f), clean restore without a backup. iPhone and Siri languages: Italian. Standard Siri, no Siri AI. Xcode 27.2 (27B5028f), iphoneos 27.2 SDK. Italian App Shortcut phrase and app metadata. PetroCheck 1.3 (637), already open during the failing invocation. The published phrase is “Trova carburanti vicini con PetroCheck”, declared using \(.applicationName). This is a phrase-based App Shortcut; the intent does not adopt an App Schema. What I observe: Saying the App Shortcut phrase produces a generic Siri error: “mi dispiace, ma si è verificato un errore”. Running the action directly in Shortcuts succeeds and records execution in the app's intent journal. Saving a personal shortcut named “Diagnostica Petro” with just that action, then saying “Esegui Diagnostica Petro”, also succeeds and records execution. Returning only a minimal Text snippet from the original intent does not resolve the phrase invocation failure. For two captured failing phrase invocations, the device logs show BackgroundShortcutRunner failing to resolve the workflow reference: -[WFWorkflowDatabaseRunDescriptor(Conversion) workflowReferenceWithDatabase:error:] Couldn't find shortcut with descriptor: <private> reason: unable to resolve workflow reference from descriptor The first of these attempts has no new entry at the recorded start of perform(). The descriptor itself is redacted, so I cannot identify which reference Siri selected. A subsequent attempt also contains these assistantd messages shortly before the descriptor error: Found no AppShortcutTargets! Could not cast to VoiceCommand task to create AppShortcut invocation AppShortcuts enablement result=false I am including these as observations, without interpreting them as proof of a disabled setting. Calling updateAppShortcutParameters() at launch did not resolve the failure. I checked the compiled device bundle's App Intents metadata: the action is discoverable, its fuel parameter is optional, and the declared shortcut points to the correct intent. The Italian training metadata contains the phrase and the application name PetroCheck. This verifies the compiled metadata, not the device's registration database. I then built a separate app, “Prova Petro”, with a different bundle identifier, a fresh phrase, one intent with no parameters, and a dialog-only result. It has no location, networking, snippet or dependencies from the original app. The Siri phrase fails in this app too. I have not yet captured its intent diary or host logs, so I cannot claim that the minimal app fails at exactly the same stage. These are the core declarations from the compiled minimal project; the sample's journal calls are omitted here: import AppIntents struct RoutingProbeIntent: AppIntent { static let title: LocalizedStringResource = "Verifica collegamento Siri" static let supportedModes: IntentModes = [.background] func perform() async throws -> some IntentResult & ProvidesDialog { return .result(dialog: "Il comando Siri ha raggiunto Prova Petro.") } } struct ProbeShortcuts: AppShortcutsProvider { static var appShortcuts: [AppShortcut] { AppShortcut( intent: RoutingProbeIntent(), phrases: ["Verifica Siri con \(.applicationName)"], shortTitle: "Verifica collegamento Siri", systemImageName: "checkmark.circle" ) } } The minimal app's display name is Prova Petro, and the spoken phrase is “Verifica Siri con Prova Petro”. Its App.init() calls ProbeShortcuts.updateAppShortcutParameters(). The extracted metadata confirms one shortcut, zero parameters, dialog-only output, and the Italian application name and phrase. Has anyone reproduced this difference between an App Shortcut phrase and a named personal shortcut on iOS 27.2 beta, particularly with Italian phrases? Feedback Assistant: FB25077916. The complete minimal Xcode project and scoped diagnostics are attached to the report. Is there anything missing from this setup, or a supported way to diagnose the unresolved shortcut reference? Comparisons with other OS builds, languages or devices would be useful. I have not yet established a regression against a stable OS release. Prova Petro minimal source and Siri routing logs
0
0
37
2d
JEV
(I don't follow the AI stuff here, so sorry if this is a stupid question. Or the wrong category.) There is a new fangled AI mode called JEV. Can the current Apple Intelligence libraries do something like it, or is this a WWDC27 thing?
2
0
846
3d
Encrypted Core ML loading capacity drops after process termination and recovers after reboot (-42905)
We have a standalone reproduction of encrypted Core ML model loading capacity decreasing after the app is terminated during an unfinished load. Feedback ID: FB25001494 The Feedback Assistant report includes a minimal UIKit/Objective-C project, a small model generated entirely from seeded random constants, and 40 standalone experiment logs. Environment iPhone 14 Pro (iPhone15,2) iOS 16.0 (20A357) Both synchronous and asynchronous Core ML loading APIs MLComputeUnitsAll Loading runs on a background worker, with only one outstanding request during each interruption trial No application-level loading timeout Other iOS versions have not yet been verified with this standalone sample. Error NSError domain: com.apple.CoreML Code: 9 The error description contains "Failed to set up decrypt context" and "error:-42905". Reproduction Reboot the device and complete normal model loading/releasing to ensure the encryption key is available. Measure capacity by sequentially loading and retaining encrypted MLModel instances until the first failure, then release all successful instances. Launch a fresh process and terminate it with SIGKILL approximately 50 ms after starting its first model load, before the load completes. Repeat this interruption in five fresh processes. Launch another process and measure capacity again. Observed results Initial capacity: 99, 99 in two measurements. After five interrupted synchronous loads: 94, 94. After five additional interrupted asynchronous loads: 90, 90. After 40 normal asynchronous loads/releases: still 90, 90. After five more interrupted synchronous loads: 85, 85. After at least 60 seconds with no sample process running: still 85, 85. After rebooting the device: 99, 99. Timing matters: five interruptions at approximately 5 ms did not reduce capacity. We do not claim that every interrupted load loses exactly one resource. Controls and interpretation Normal controls and capacity probes release all successfully loaded models before their processes are terminated. Autorelease pools and associated-object deallocation witnesses are used to check model lifetime; these do not directly inspect internal decrypt sessions. The capacity-limit error while deliberately retaining many models is expected. The unexpected behavior is that terminating an unfinished load reduces the repeatable capacity available to subsequent fresh processes. This suggests a cleanup issue across process termination, but the internal cause has not been established. Questions Is this a known issue, and if it has been fixed, which iOS version contains the fix? Is there a supported recovery mechanism that does not require rebooting the device? Is there a recommended loading or lifecycle workaround? Switching between synchronous and asynchronous loading did not eliminate the behavior. Has anyone reproduced this specific interrupted-load behavior on a newer iOS version? Related discussions https://developer.apple.com/forums/thread/740731 https://developer.apple.com/forums/thread/678599 We also reviewed thread 707622, where moving the autorelease pool inside the loop resolved retained-model exhaustion. Our sample includes release controls and specifically tests capacity after termination of an unfinished load.
0
0
1k
1w
SIRI AI AND APPLE INTELLIGENCE
Since I updated to this os 27 in my iPhone 16 plus the siri ai and Apple Intelligence is got freeze in “ Adding support for Siri is in progress. Siri will be unavailable until the update is complete. “ I updated on 21/09/2026 today date is 30/09/2026 I tried all troubleshooting methods and watch a bunch of YouTube videos, but still stuck in the same position. I contacted Apple support. But , they also do nothing. Please anyone help me to how to get siri ai beta @appleindia @applesupport
0
0
1.1k
1w
Adding MCP and connector support to your own Foundation Models apps
Circling back on the LocalLM Lab arc. With v0.7, we've moved from prompt experimentation into real app development on Apple's Foundation Models local AI. The LocalLM Lab SDK lets you build that same on-device model and MCP client this thread has covered directly into your own app, with real tool and data access (Slack, Todoist, GitHub, Notion, Linear, plus Calendar, Reminders, Contacts and Location). And you can ship your app including through the Mac App Store. This is a big improvement over version 0.6, where the localai-cli toolkit needed LocalLM Lab installed and running. On the other hand, the SDK (LocalLMLabSDKCore) doesn't relay through anything; it links FoundationModels and a real MCP client directly into your own binary and is totally self-contained. The example included in the SDK, Plate Today, has actually been built into a sandboxed test app and verified working, with a signed path to a Mac App Store .pkg (Apple Distribution signing + provisioning profile pipeline). That's "verified signable and sandbox-compatible," to be precise. Entitlements (from personal experience: always a complicated topic): com.apple.security.app-sandbox + com.apple.security.network.client for the app itself, plus the standard personal-information entitlements per connector used (com.apple.security.personal-information.calendars, .addressbook, .location) and matching NS*UsageDescription strings in Info.plist. The one worth flagging specifically: the network entitlement is easy to miss and fails silently rather than throwing. Without it, MCP connections and Weather calls just hang with no error surfaced. OAuth handling requires the app delegate callback (application(_:open:)), not SwiftUI's .onOpenURL. Worth knowing before wiring it up if you're SwiftUI-only. Full entitlements list + SDK guide: https://github.com/ancientcomputing/locallm/blob/main/docs/sdk-guide.md Feature page: thisbrain.ai/locallm/sdk.html I hope the availability of the SDK (free, Apache 2.0 license) will give folks further incentive to explore local AI-enabled applications on the Mac. What else would you want to do that the SDK doesn't currently support? File picker? Calendar/Reminders/Contacts edits & writes?
6
1
3.6k
1w
Create ML Object Tracking training stuck at 97.4% - When pressing 'Resume Training' it PAUSES automatically after 3 seconds
Hi, I’m training an Object Tracking reference object in Create ML using Extended training mode + All Angles for use with visionOS high-frame-rate object tracking. The training ran successfully for approximately 48 hours and reached 97.4%. At that point, Create ML automatically paused. The issue is now completely reproducible: whenever I click Resume, training runs for approximately 3 seconds and then automatically returns to Paused at exactly 97.4%. I captured the system logs while reproducing the issue. MLRecipeExecutionService crashes with: LayerVariable.swift:116: Fatal error: The new value must have the same shape as the current value ([40, 1, 3, 3]), but it has [96, 1, 3, 3]. Immediately afterwards, Create ML reports an interrupted XPC connection to MLRecipeExecutionService. Shortly before the crash, the service also logs: Detector training already finished but its loss file was removed from the cache; reporting loss as 0. and: Error cleaning up detector images: NSCocoaErrorDomain Code=4 NSPOSIXErrorDomain Code=2 "No such file or directory" I would really like to avoid discarding ~48 hours of training if the existing tracker/detector checkpoints can still be recovered. Thanks!
0
0
283
1w
Exploring Apple Silicon + MLX for a persistent local AI companion architecture
I’m developing an independent project in Scotland called Isla Watson. The architecture is built around a simple principle: the model is replaceable; the identity is not. Long-term memory, persistent internal state and identity are designed to remain outside the foundation model, allowing local models to act as replaceable reasoning and language components without resetting the companion. I’m now exploring whether Apple Silicon and MLX could provide the long-term local compute platform for the system — including specialist Mac nodes for reasoning, memory, speech and perception, with distributed inference when larger models are required. A particular area of interest is whether multiple Macs can be used in two complementary ways: as independent specialist agents during normal operation; and as a distributed MLX inference group when a larger model exceeds the capacity of one machine. The first technical study I’d like to establish is a reproducible 1-node → 2-node baseline, measuring model capacity, unified-memory use, time to first token, generation throughput, power consumption, agent concurrency and distributed scaling efficiency. The wider research goal is to keep persistent identity and state independent from whichever foundation model is currently providing language and reasoning. I’d particularly value guidance from anyone working with MLX distributed inference, Thunderbolt/RDMA multi-Mac setups, or local agent architectures. I’ve also posted an architecture-level overview in the MLX GitHub community and have a one-page public brief available for anyone interested in the wider design. https://github.com/ml-explore/mlx/discussions/4482
2
0
431
1w
Model Guardrails Too Restrictive?
I'm experimenting with using the Foundation Models framework to do news summarization in an RSS app but I'm finding that a lot of articles are getting kicked back with a vague message about guardrails. This seems really common with political news but we're talking mainstream stuff, i.e. Politico, etc. If the models are this restrictive, this will be tough to use. Is this intended? FB17904424
10
5
2.6k
1w
tensorflow 2.20 broken support
Hi, testing latest tensorflow-metal plugin with tensorflow 2.20 doesn't work.. using python Python 3.12.11 (main, Jun 3 2025, 15:41:47) [Clang 17.0.0 (clang-1700.0.13.3)] on darwin simple testing shows error: import tensorflow as tf Traceback (most recent call last): File "", line 1, in File "/Users/obg/npu/venv-tf/lib/python3.12/site-packages/tensorflow/init.py", line 438, in _ll.load_library(_plugin_dir) File "/Users/obg/npu/venv-tf/lib/python3.12/site-packages/tensorflow/python/framework/load_library.py", line 151, in load_library py_tf.TF_LoadLibrary(lib) tensorflow.python.framework.errors_impl.NotFoundError: dlopen(/Users/obg/npu/venv-tf/lib/python3.12/site-packages/tensorflow-plugins/libmetal_plugin.dylib, 0x0006): Library not loaded: @rpath/_pywrap_tensorflow_internal.so Referenced from: <8B62586B-B082-3113-93AB-FD766A9960AE> /Users/obg/npu/venv-tf/lib/python3.12/site-packages/tensorflow-plugins/libmetal_plugin.dylib Reason: tried: '/Users/obg/npu/venv-tf/lib/python3.12/site-packages/tensorflow-plugins/../_solib_darwin_arm64/_U@local_Uconfig_Utf_S_S_C_Upywrap_Utensorflow_Uinternal___Uexternal_Slocal_Uconfig_Utf/_pywrap_tensorflow_internal.so' (no such file), '/Users/obg/npu/venv-tf/lib/python3.12/site-packages/tensorflow-plugins/../_solib_darwin_arm64/_U@local_Uconfig_Utf_S_S_C_Upywrap_Utensorflow_Uinternal___Uexternal_Slocal_Uconfig_Utf/_pywrap_tensorflow_internal.so' (no such file), '/opt/homebrew/lib/_pywrap_tensorflow_internal.so' (no such file), '/System/Volumes/Preboot/Cryptexes/OS/opt/homebrew/lib/_pywrap_tensorflow_internal.so' (no such file) tf.config.experimental.list_physical_devices('GPU') Traceback (most recent call last): File "", line 1, in NameError: name 'tf' is not defined I fixed this error by copying _pywrap_tensorflow_internal.so where it's searched.. 1)mkdir /Users/obg/npu/venv-tf/lib/python3.12/site-packages/tensorflow-plugins/../_solib_darwin_arm64 2)mkdir /Users/obg/npu/venv-tf/lib/python3.12/site-packages/tensorflow-plugins/../_solib_darwin_arm64/_U@local_Uconfig_Utf_S_S_C_Upywrap_Utensorflow_Uinternal___Uexternal_Slocal_Uconfig_Utf/ 3)cp /Users/obg/npu/venv-tf/lib/python3.12/site-packages/tensorflow/python/_pywrap_tensorflow_internal.so /Users/obg/npu/venv-tf/lib/python3.12/site-packages/tensorflow-plugins/../_solib_darwin_arm64/_U@local_Uconfig_Utf_S_S_C_Upywrap_Utensorflow_Uinternal___Uexternal_Slocal_Uconfig_Utf/ then fails symbol not found: Symbol not found: __ZN10tensorflow28_AttrValue_default_instance_E in libmetal_plugin.dylib full log: with import tensorflow as tf Traceback (most recent call last): File "", line 1, in File "/Users/obg/npu/venv-tf/lib/python3.12/site-packages/tensorflow/init.py", line 438, in _ll.load_library(_plugin_dir) File "/Users/obg/npu/venv-tf/lib/python3.12/site-packages/tensorflow/python/framework/load_library.py", line 151, in load_library py_tf.TF_LoadLibrary(lib) tensorflow.python.framework.errors_impl.NotFoundError: dlopen(/Users/obg/npu/venv-tf/lib/python3.12/site-packages/tensorflow-plugins/libmetal_plugin.dylib, 0x0006): Symbol not found: __ZN10tensorflow28_AttrValue_default_instance_E Referenced from: <8B62586B-B082-3113-93AB-FD766A9960AE> /Users/obg/npu/venv-tf/lib/python3.12/site-packages/tensorflow-plugins/libmetal_plugin.dylib Expected in: <2FF91C8B-0CB6-3E66-96B7-092FDF36772E> /Users/obg/npu/venv-tf/lib/python3.12/site-packages/_solib_darwin_arm64/_U@local_Uconfig_Utf_S_S_C_Upywrap_Utensorflow_Uinternal___Uexternal_Slocal_Uconfig_Utf/_pywrap_tensorflow_internal.so
4
0
2.7k
1w
Private Cloud Compute throws guardrailViolation on benign song analysis (FB24938334)
On iOS 27.0, PrivateCloudComputeLanguageModel refuses a large share of harmless requests, and the same requests often succeed when simply re-run. Our app writes a short, general-audience explanation of what a song is about. When the cloud model refuses, the on-device model with .permissiveContentTransformations answers the same prompt without trouble. PCC has no guardrail configuration, so there's nothing on our side to adjust. • Two refusals came within the first sentence of plainly benign text: "Tiny Dancer" (Elton John) at 167 characters, and "In the Ghetto" (Elvis Presley) at 169 characters, while describing snow on a Chicago morning. • Others include "Already Gone" (Eagles), "Do Ya" (ELO) and "We Didn't Start the Fire" (Billy Joel). • It's non-deterministic: "Question" (The Moody Blues) was refused, then succeeded 52 seconds later with an identical request. • Most failures are "Streamed response may contain sensitive or unsafe content", arriving mid-generation, so the rejection seems to target the model's own output rather than the input. • Rewording the instructions to steer the model toward mainstream, general-audience language didn't change the refusal rate. Filed as FB24938334 with 8 logFeedbackAttachment captures (.triggeredGuardrailUnexpectedly), each including the rolled-back rejected draft. Siri language English (US), iPhone 16 Pro. Is PCC's guardrail policy expected to be tuned for content-transformation tasks like this, or is there a recommended pattern for them?
0
0
294
1w
iPhone 18 Pro: First Core ML load on Neural Engine up to 65× slower than on iPhone 17 Pro
On iPhone 18 Pro (A20 Pro), the first MLModel load of a model on the Neural Engine can take minutes where an iPhone 17 Pro needs seconds. Cached loads are instant, and once loaded the model runs faster than on the iPhone 17 Pro, so only the on-device specialization is affected. This looks like the same problem as A20 pro devices take too long on CoreML model specialization and iPhone 18 pro takes very long time to load models. The "loading" time there is likely this specialization too: it happens inside MLModel(contentsOf:) for an already compiled .mlmodelc. I can reproduce it reliably with attention in the form Apple recommends for the Neural Engine (split einsum, as in Deploying Transformers on the Apple Neural Engine). In this example most of the extra time comes from the softmax over the key axis, which in that layout is the channel axis. Other ops and model types may well be affected too; this is simply the smallest case we could build that shows the problem clearly. The repro uses four self-attention layers (10 heads × 64, 4096 tokens, fp16) built with coremltools and loaded cold with .cpuAndNeuralEngine. MLComputePlan places every op on the Neural Engine on all three devices. iOS 27.0.1 (24A446) on both iPhones: iPhone 18 Pro iPhone 17 Pro M3 Pro Mac cold predict cold predict cold predict Split-einsum attention, 4 layers 53.4 s 21.6 ms 0.8 s 111.4 ms 1.0 s 105.8 ms Q.K einsum only 5.1 s 24.0 ms 1.2 s 36.8 ms 1.4 s 41.8 ms Q.K einsum + softmax 36.1 s 61.5 ms 1.7 s 141.0 ms 2.0 s 128.3 ms Scores keys-last (softmax last axis) 10.5 s 109.0 ms 0.6 s 198.4 ms 0.7 s 180.2 ms scaled_dot_product_attention 1.4 s 140.4 ms 0.4 s 184.9 ms 0.5 s 148.1 ms Cached loads take 0.01–0.03 s everywhere. What I observed: Compile time grows linearly with the number of attention layers. Every alternative attention layout I tried that compiles quickly runs 5–6× slower on the A20 Pro, so I found no workaround that keeps its speed. During the slow load, ANECompilerService runs almost entirely on the efficiency cores. Requesting a higher QoS, MLOptimizationHints (.fastPrediction, reshapeFrequency), computeUnits = .all and dense instead of palettized weights don't help. Loading several models in parallel additionally gets ANECompilerService killed by jetsam (per-process-limit). The specialization cache doesn't survive reinstalling the app. In our app (Stable Diffusion XL), each UNet chunk with attention takes 110–150 s instead of 6–16 s, so the first launch takes about 25 minutes in the foreground, and in our tests again after every reinstall. The same networks converted to Core AI specialize normally on this device (see our Core AI porting notes). Filed as FB25106520 with a self-contained repro: a coremltools script, an iOS app and a Mac runner, plus results from all three devices. Environment: iPhone 18 Pro (iPhone19,2) and iPhone 17 Pro (iPhone18,1) on iOS 27.0.1 (24A446), M3 Pro on macOS 27.0.1 (26A434), Xcode 27.0, coremltools 8.3.0. Is this a known A20 Pro issue, and is there a recommended way to avoid it until it's fixed? Falling back to the GPU or CPU isn't a good option for us: it avoids the slow load, but gives up the Neural Engine's efficiency, which matters for on-device generation on a phone, and its speed (a full SDXL UNet step on the iPhone 18 Pro takes 509 ms on the Neural Engine vs 1,639 ms with .cpuAndGPU). We'd also like to hear from others which ops their slow first loads involve, to see whether this is specific to attention and softmax or broader.
Replies
3
Boosts
4
Views
368
Activity
17h
iOS 27: concurrent Core ML predictions on the Neural Engine are serialized
Environment iPhone 17 Pro Max, iOS 27.0.1 ([build]) Reference: MacBook Pro (M3 Pro), macOS 26.4.1 ML Program model, MLComputeUnitsCPUAndNeuralEngine. A single prediction takes about 2.8 ms on the ANE. Description On iOS 27, Core ML predictions submitted to the Neural Engine from multiple threads at the same time no longer overlap. The ANE runs them strictly one after another. Adding threads gives no extra throughput, and per-call latency grows linearly with the number of requests in flight. The same code still gets a clear concurrency gain on macOS 26.4.1, and our app's ANE-parallel path on this iPhone became noticeably slower after the iOS 27 update. On iOS 27, latency is about N × single-call time, which means requests are queued and never overlap. Running two different models concurrently on the ANE shows the same behavior. What we tried, with no change in result A separate MLModel instance per thread. The async predictionFromFeatures:options:completionHandler: API. Questions Is this serialization an intended change in iOS 27? Is there a supported way to get concurrent ANE execution back?
Replies
0
Boosts
0
Views
219
Activity
21h
iOS 27: English sentence embedding model unavailable on physical devices
I’m seeing NLEmbedding.sentenceEmbedding(for: .english) return nil on physical devices running iOS 27, both in debug builds under Xcode and in production. The same failure has been observed across thousands of users on iOS 27. One affected device runs iOS 27.0.1 (24A446). Minimal example with diagnostics: import NaturalLanguage let language = NLLanguage.english print("Current revision:", NLEmbedding.currentSentenceEmbeddingRevision(for: language)) print("Supported revisions:", Array(NLEmbedding.supportedSentenceEmbeddingRevisions(for: language))) if let model = NLEmbedding.sentenceEmbedding(for: language) { print("Loaded revision:", model.revision, "dimension:", model.dimension) } else { print("Sentence embedding model unavailable") } On affected devices, the current revision is 1 and supported revisions are [1], but model loading returns nil. The framework logs: Unable to locate Asset for sentence embedding model for local en. The failure occurs during model loading, before any call to vector(for:). This uses Apple’s built-in English sentence model, with no custom embedding file. Is this a known issue on iOS 27? Is there a supported way to request or restore the missing sentence-model asset, or a recommended recovery approach?
Replies
0
Boosts
0
Views
437
Activity
1d
PCC rate limits and blocking Siri AI as a result
Hello, I have PCC deployed in my 3 apps and my testing for these apps has caused me to hit the rate limit many times. When this rate limit is hit, Siri AI (as well as HomeKit PCC features) stops working completely for me until the rate limit is reset. I have been forced to scale back PCC and testing in my apps as this situation is quite opaque. My logs say I’m being rate limited when using PCC in these cases. Example log I see: 10:51:47.407 | RemoteManager executing remote streaming request 10:51:47.407 ModelCatalog | Looking up asset bundle com.apple.fm.language.instruct_server_v2.fm_api in Model Catalog 10:51:47.410 DaemonSession | Asset com.apple.fm.language.instruct_server_v2.fm_api has inference providers: 10:51:47.410 DaemonSession | - private-ml-client 10:51:47.411 ExecutionGroup | executing request C9B08798-6D68-4819-83C9-89B2EF5278C4 10:51:47.431 InferenceProviderConnection | InferenceProvider requestInference (C9B08798-6D68-4819-83C9-89B2EF5278C4) finished 10:51:47.431 RemoteManager | Request failed on private-ml-client with error: InferenceError::rateLimited::PrivateCloudComputeError: deniedDueToUserDeviceRateLimit, falling back 10:51:47.431 RemoteIPC Tasks | Failed to execute remote streamingRequest with error: InferenceError::rateLimited::PrivateCloudComputeError: deniedDueToUserDeviceRateLimit 10:51:56 modelmanagerd [com.apple.modelmanager:RemoteIPC Tasks] Failed to execute remote oneShotRequest with error: InferenceError::rateLimited::PrivateCloudComputeError: deniedDueToUserDev 11:23:39 privatecloudcomputed [com.apple.privatecloudcompute:RateLimiter] rate limit applied from cached denials What are the rate limits? How many requests a day? Per second? Also, I subscribed to the highest iCloud storage tier in hopes of higher rate limit but there doesn't seem to be any change. https://support.apple.com/en-us/127901 This issue started when iOS27 became publicly available - I was never rate limited when on the iOS27 beta.
Replies
2
Boosts
0
Views
825
Activity
1d
Crashed: AXSpeech EXC_BAD_ACCESS KERN_INVALID_ADDRESS 0x000056f023efbeb0
Application is getting Crashed: AXSpeech EXC_BAD_ACCESS KERN_INVALID_ADDRESS 0x000056f023efbeb0 Crashed: AXSpeech 0 libobjc.A.dylib 0x4820 objc_msgSend + 32 1 libsystem_trace.dylib 0x6c34 _os_log_fmt_flatten_object + 116 2 libsystem_trace.dylib 0x5344 _os_log_impl_flatten_and_send + 1884 3 libsystem_trace.dylib 0x4bd0 _os_log + 152 4 libsystem_trace.dylib 0x9c48 _os_log_error_impl + 24 5 TextToSpeech 0xd0a8c _pcre2_xclass_8 6 TextToSpeech 0x3bc04 TTSSpeechUnitTestingMode 7 TextToSpeech 0x3f128 TTSSpeechUnitTestingMode 8 AXCoreUtilities 0xad38 -[NSArray(AXExtras) ax_flatMappedArrayUsingBlock:] + 204 9 TextToSpeech 0x3eb18 TTSSpeechUnitTestingMode 10 TextToSpeech 0x3c948 TTSSpeechUnitTestingMode 11 TextToSpeech 0x48824 AXAVSpeechSynthesisVoiceFromTTSSpeechVoice 12 TextToSpeech 0x49804 AXAVSpeechSynthesisVoiceFromTTSSpeechVoice 13 Foundation 0xf6064 __NSThreadPerformPerform + 264 14 CoreFoundation 0x37acc CFRUNLOOP_IS_CALLING_OUT_TO_A_SOURCE0_PERFORM_FUNCTION + 28 15 CoreFoundation 0x36d48 __CFRunLoopDoSource0 + 176 16 CoreFoundation 0x354fc __CFRunLoopDoSources0 + 244 17 CoreFoundation 0x34238 __CFRunLoopRun + 828 18 CoreFoundation 0x33e18 CFRunLoopRunSpecific + 608 19 Foundation 0x2d4cc -[NSRunLoop(NSRunLoop) runMode:beforeDate:] + 212 20 TextToSpeech 0x24b88 TTSCFAttributedStringCreateStringByBracketingAttributeWithString 21 Foundation 0xb3154 NSThread__start + 732 com.livingMedia.AajTakiPhone_issue_3ceba855a8ad2d1af83655803dc13f70_crash_session_9081fa41ced440ae9a57c22cb432f312_DNE_0_v2_stacktrace.txt 22 libsystem_pthread.dylib 0x24d4 _pthread_start + 136 23 libsystem_pthread.dylib 0x1a10 thread_start + 8
Replies
5
Boosts
1
Views
2.6k
Activity
2d
iPhone 18 pro takes very long time to load models
I think this could be relate to previous post here https://developer.apple.com/forums/thread/848590 But in my testing it seems like most of the time spent was in loading but not compilation. Our app has multiple models but so far it seems like one of them has this issue. The model itself was only 2.2MB, but somehow loading took 50s after compilation? And on 17e it only took 2s to load. But I cannot attach the model here so please advice.
Replies
1
Boosts
1
Views
595
Activity
2d
A20 pro devices take too long on CoreML model specialization
Hi, I'm a ML developer specializing in local audio related models. I upgraded from a 16 pro max to an 18 pro max. I'm not sure what the reason is, but model specialization (happens only on the first load, cpu and neural engine compute units) is sometimes 2x - 5x longer on the 18 pro max vs the 16 pro max. Both devices are on iOS 27. For example, I have a stem separation model running, and the 18 pro max took 171 seconds (including inference) and the 16 pro max took 37 seconds (including inference). 18 pro max inference is ~55x RTFx and the 16 pro max is ~40 RTFx for this particular model, but I've noticed it as well in TTS models etc. Please look into this, I know the 18 pro max has double the neural engine cores, so I'm not sure if the complexity makes specialization take longer. Subsequent loads are great, same high rtfx for the 18 PM (only 0.5s load), and this is where it beats the 16 PM handily. I attached 2 images showing the 18 PM on a first load (model specialization likely) vs a subsequent run. Please look into this. (Note this is not just my implementation, I tested a range of different models, all with the same outcome).
Replies
1
Boosts
5
Views
561
Activity
2d
Core AI on iOS/macOS 27: seven issues found while porting an SDXL pipeline
We ported a complete Stable Diffusion XL (Lightning) pipeline to Core AI which we previously had on Core ML. The pipeline comprises text encoders, a ViT-H image encoder, ControlNet, the UNet and the VAE decoder. It works very well: on iPhone 17 Pro and M1 iPad Pro it generates 1.4–1.6x faster than our Core ML pipeline at the same image quality. However, along the way, we hit seven issues. Each one is filed with a small self-contained repro. We're posting them here together for visibility, with our workarounds, in case others run into the same problems. Ahead-of-time compilation (coreai-build compile) Each of these gives wrong results on the Neural Engine with no error. The same .aimodel specialized on the device is correct. FB25067476: A 3×3 fp16 conv with a large output returns a wrong bottom half (16.7 dB vs 60.7 dB for the top half). This depends on the chip: 256×512×512 fails on M3 Pro and A19 Pro, while 256×384×384 fails only on M1. Our workaround: split such convs by output channels so that no single output exceeds about 128×384×384. FB25067497: A nearest 2× upsample followed by a 3×3 fp16 conv returns output unrelated to the correct result (correlation 0.02) on all three chips. The upsample alone and the conv alone are correct. Our workaround:* write the upsample as repeat_interleave(4) + pixel_shuffle(2). FB25067530: A 3×3 stride-2 conv with 8-bit palettized weights is wrong (16 dB) on all three chips. The same conv with fp16 weights, or palettized with stride 1, is correct. Our workaround:* a 2×2 pixel unshuffle followed by a 2×2 conv. Caching and loading FB25066222: The cache for ahead-of-time compiled .aimodelc files is keyed by the outer program only (main.hash), not the weights, so a weights-only model update silently runs the old weights. On iOS, the Neural Engine program cache is also keyed by function name per app, and survives AIModelCache.deleteAll() and deleting the app. Our workaround: append a hash of the model's content to every function name. FB25066664: On macOS, ahead-of-time builds of transformer models load 3–4× slower than on-device specialization, and parallel loads finish one at a time. Precompiling therefore made our Mac first launch slower (1152 s vs about 650 s), not faster. Neural Engine runtime (Mac) FB25067019: The built-in GroupNorm returns wrong results on the Mac's Neural Engine for large inputs (GroupNorm(32, 256) on 1×256×256×256: 21.7 dB), while the GPU and CPU are correct (84 dB). Our workaround: compute the means and variance explicitly. FB25064649: A ViT-H vision tower with an 8-bit palettized 14×14/stride-14 patch conv aborts the process on the Mac's Neural Engine (MPSGraph ANE error -19). Our workaround: keep the patch conv in fp16, or use a pixel unshuffle followed by a 1×1 conv. Environment macOS 27.0.1 (26A434), iOS/iPadOS 27.0.1 (24A446) Xcode 27.0, Metal Toolchain 27.1.266.1 (coreai-build 3600.83.1), coreai-torch 0.4.3 Tested on M3 Pro, iPhone 17 Pro and M1 iPad Pro We're happy to provide more data. We'd also like to hear whether others see the ahead-of-time issues on other chips.
Replies
0
Boosts
3
Views
70
Activity
2d
iOS 27.2 beta: App Shortcut phrase fails, but a named personal shortcut runs the same intent
I'm investigating a Siri invocation failure with an explicitly declared App Shortcut phrase. The same underlying App Intent works when run directly in Shortcuts and when invoked by the name of a saved personal shortcut. Environment: iPhone 17 Pro Max, iOS 27.2 beta (24B5099f), clean restore without a backup. iPhone and Siri languages: Italian. Standard Siri, no Siri AI. Xcode 27.2 (27B5028f), iphoneos 27.2 SDK. Italian App Shortcut phrase and app metadata. PetroCheck 1.3 (637), already open during the failing invocation. The published phrase is “Trova carburanti vicini con PetroCheck”, declared using \(.applicationName). This is a phrase-based App Shortcut; the intent does not adopt an App Schema. What I observe: Saying the App Shortcut phrase produces a generic Siri error: “mi dispiace, ma si è verificato un errore”. Running the action directly in Shortcuts succeeds and records execution in the app's intent journal. Saving a personal shortcut named “Diagnostica Petro” with just that action, then saying “Esegui Diagnostica Petro”, also succeeds and records execution. Returning only a minimal Text snippet from the original intent does not resolve the phrase invocation failure. For two captured failing phrase invocations, the device logs show BackgroundShortcutRunner failing to resolve the workflow reference: -[WFWorkflowDatabaseRunDescriptor(Conversion) workflowReferenceWithDatabase:error:] Couldn't find shortcut with descriptor: <private> reason: unable to resolve workflow reference from descriptor The first of these attempts has no new entry at the recorded start of perform(). The descriptor itself is redacted, so I cannot identify which reference Siri selected. A subsequent attempt also contains these assistantd messages shortly before the descriptor error: Found no AppShortcutTargets! Could not cast to VoiceCommand task to create AppShortcut invocation AppShortcuts enablement result=false I am including these as observations, without interpreting them as proof of a disabled setting. Calling updateAppShortcutParameters() at launch did not resolve the failure. I checked the compiled device bundle's App Intents metadata: the action is discoverable, its fuel parameter is optional, and the declared shortcut points to the correct intent. The Italian training metadata contains the phrase and the application name PetroCheck. This verifies the compiled metadata, not the device's registration database. I then built a separate app, “Prova Petro”, with a different bundle identifier, a fresh phrase, one intent with no parameters, and a dialog-only result. It has no location, networking, snippet or dependencies from the original app. The Siri phrase fails in this app too. I have not yet captured its intent diary or host logs, so I cannot claim that the minimal app fails at exactly the same stage. These are the core declarations from the compiled minimal project; the sample's journal calls are omitted here: import AppIntents struct RoutingProbeIntent: AppIntent { static let title: LocalizedStringResource = "Verifica collegamento Siri" static let supportedModes: IntentModes = [.background] func perform() async throws -> some IntentResult & ProvidesDialog { return .result(dialog: "Il comando Siri ha raggiunto Prova Petro.") } } struct ProbeShortcuts: AppShortcutsProvider { static var appShortcuts: [AppShortcut] { AppShortcut( intent: RoutingProbeIntent(), phrases: ["Verifica Siri con \(.applicationName)"], shortTitle: "Verifica collegamento Siri", systemImageName: "checkmark.circle" ) } } The minimal app's display name is Prova Petro, and the spoken phrase is “Verifica Siri con Prova Petro”. Its App.init() calls ProbeShortcuts.updateAppShortcutParameters(). The extracted metadata confirms one shortcut, zero parameters, dialog-only output, and the Italian application name and phrase. Has anyone reproduced this difference between an App Shortcut phrase and a named personal shortcut on iOS 27.2 beta, particularly with Italian phrases? Feedback Assistant: FB25077916. The complete minimal Xcode project and scoped diagnostics are attached to the report. Is there anything missing from this setup, or a supported way to diagnose the unresolved shortcut reference? Comparisons with other OS builds, languages or devices would be useful. I have not yet established a regression against a stable OS release. Prova Petro minimal source and Siri routing logs
Replies
0
Boosts
0
Views
37
Activity
2d
JEV
(I don't follow the AI stuff here, so sorry if this is a stupid question. Or the wrong category.) There is a new fangled AI mode called JEV. Can the current Apple Intelligence libraries do something like it, or is this a WWDC27 thing?
Replies
2
Boosts
0
Views
846
Activity
3d
Encrypted Core ML loading capacity drops after process termination and recovers after reboot (-42905)
We have a standalone reproduction of encrypted Core ML model loading capacity decreasing after the app is terminated during an unfinished load. Feedback ID: FB25001494 The Feedback Assistant report includes a minimal UIKit/Objective-C project, a small model generated entirely from seeded random constants, and 40 standalone experiment logs. Environment iPhone 14 Pro (iPhone15,2) iOS 16.0 (20A357) Both synchronous and asynchronous Core ML loading APIs MLComputeUnitsAll Loading runs on a background worker, with only one outstanding request during each interruption trial No application-level loading timeout Other iOS versions have not yet been verified with this standalone sample. Error NSError domain: com.apple.CoreML Code: 9 The error description contains "Failed to set up decrypt context" and "error:-42905". Reproduction Reboot the device and complete normal model loading/releasing to ensure the encryption key is available. Measure capacity by sequentially loading and retaining encrypted MLModel instances until the first failure, then release all successful instances. Launch a fresh process and terminate it with SIGKILL approximately 50 ms after starting its first model load, before the load completes. Repeat this interruption in five fresh processes. Launch another process and measure capacity again. Observed results Initial capacity: 99, 99 in two measurements. After five interrupted synchronous loads: 94, 94. After five additional interrupted asynchronous loads: 90, 90. After 40 normal asynchronous loads/releases: still 90, 90. After five more interrupted synchronous loads: 85, 85. After at least 60 seconds with no sample process running: still 85, 85. After rebooting the device: 99, 99. Timing matters: five interruptions at approximately 5 ms did not reduce capacity. We do not claim that every interrupted load loses exactly one resource. Controls and interpretation Normal controls and capacity probes release all successfully loaded models before their processes are terminated. Autorelease pools and associated-object deallocation witnesses are used to check model lifetime; these do not directly inspect internal decrypt sessions. The capacity-limit error while deliberately retaining many models is expected. The unexpected behavior is that terminating an unfinished load reduces the repeatable capacity available to subsequent fresh processes. This suggests a cleanup issue across process termination, but the internal cause has not been established. Questions Is this a known issue, and if it has been fixed, which iOS version contains the fix? Is there a supported recovery mechanism that does not require rebooting the device? Is there a recommended loading or lifecycle workaround? Switching between synchronous and asynchronous loading did not eliminate the behavior. Has anyone reproduced this specific interrupted-load behavior on a newer iOS version? Related discussions https://developer.apple.com/forums/thread/740731 https://developer.apple.com/forums/thread/678599 We also reviewed thread 707622, where moving the autorelease pool inside the loop resolved retained-model exhaustion. Our sample includes release controls and specifically tests capacity after termination of an unfinished load.
Replies
0
Boosts
0
Views
1k
Activity
1w
SIRI AI AND APPLE INTELLIGENCE
Since I updated to this os 27 in my iPhone 16 plus the siri ai and Apple Intelligence is got freeze in “ Adding support for Siri is in progress. Siri will be unavailable until the update is complete. “ I updated on 21/09/2026 today date is 30/09/2026 I tried all troubleshooting methods and watch a bunch of YouTube videos, but still stuck in the same position. I contacted Apple support. But , they also do nothing. Please anyone help me to how to get siri ai beta @appleindia @applesupport
Replies
0
Boosts
0
Views
1.1k
Activity
1w
Adding MCP and connector support to your own Foundation Models apps
Circling back on the LocalLM Lab arc. With v0.7, we've moved from prompt experimentation into real app development on Apple's Foundation Models local AI. The LocalLM Lab SDK lets you build that same on-device model and MCP client this thread has covered directly into your own app, with real tool and data access (Slack, Todoist, GitHub, Notion, Linear, plus Calendar, Reminders, Contacts and Location). And you can ship your app including through the Mac App Store. This is a big improvement over version 0.6, where the localai-cli toolkit needed LocalLM Lab installed and running. On the other hand, the SDK (LocalLMLabSDKCore) doesn't relay through anything; it links FoundationModels and a real MCP client directly into your own binary and is totally self-contained. The example included in the SDK, Plate Today, has actually been built into a sandboxed test app and verified working, with a signed path to a Mac App Store .pkg (Apple Distribution signing + provisioning profile pipeline). That's "verified signable and sandbox-compatible," to be precise. Entitlements (from personal experience: always a complicated topic): com.apple.security.app-sandbox + com.apple.security.network.client for the app itself, plus the standard personal-information entitlements per connector used (com.apple.security.personal-information.calendars, .addressbook, .location) and matching NS*UsageDescription strings in Info.plist. The one worth flagging specifically: the network entitlement is easy to miss and fails silently rather than throwing. Without it, MCP connections and Weather calls just hang with no error surfaced. OAuth handling requires the app delegate callback (application(_:open:)), not SwiftUI's .onOpenURL. Worth knowing before wiring it up if you're SwiftUI-only. Full entitlements list + SDK guide: https://github.com/ancientcomputing/locallm/blob/main/docs/sdk-guide.md Feature page: thisbrain.ai/locallm/sdk.html I hope the availability of the SDK (free, Apache 2.0 license) will give folks further incentive to explore local AI-enabled applications on the Mac. What else would you want to do that the SDK doesn't currently support? File picker? Calendar/Reminders/Contacts edits & writes?
Replies
6
Boosts
1
Views
3.6k
Activity
1w
Create ML Object Tracking training stuck at 97.4% - When pressing 'Resume Training' it PAUSES automatically after 3 seconds
Hi, I’m training an Object Tracking reference object in Create ML using Extended training mode + All Angles for use with visionOS high-frame-rate object tracking. The training ran successfully for approximately 48 hours and reached 97.4%. At that point, Create ML automatically paused. The issue is now completely reproducible: whenever I click Resume, training runs for approximately 3 seconds and then automatically returns to Paused at exactly 97.4%. I captured the system logs while reproducing the issue. MLRecipeExecutionService crashes with: LayerVariable.swift:116: Fatal error: The new value must have the same shape as the current value ([40, 1, 3, 3]), but it has [96, 1, 3, 3]. Immediately afterwards, Create ML reports an interrupted XPC connection to MLRecipeExecutionService. Shortly before the crash, the service also logs: Detector training already finished but its loss file was removed from the cache; reporting loss as 0. and: Error cleaning up detector images: NSCocoaErrorDomain Code=4 NSPOSIXErrorDomain Code=2 "No such file or directory" I would really like to avoid discarding ~48 hours of training if the existing tracker/detector checkpoints can still be recovered. Thanks!
Replies
0
Boosts
0
Views
283
Activity
1w
Exploring Apple Silicon + MLX for a persistent local AI companion architecture
I’m developing an independent project in Scotland called Isla Watson. The architecture is built around a simple principle: the model is replaceable; the identity is not. Long-term memory, persistent internal state and identity are designed to remain outside the foundation model, allowing local models to act as replaceable reasoning and language components without resetting the companion. I’m now exploring whether Apple Silicon and MLX could provide the long-term local compute platform for the system — including specialist Mac nodes for reasoning, memory, speech and perception, with distributed inference when larger models are required. A particular area of interest is whether multiple Macs can be used in two complementary ways: as independent specialist agents during normal operation; and as a distributed MLX inference group when a larger model exceeds the capacity of one machine. The first technical study I’d like to establish is a reproducible 1-node → 2-node baseline, measuring model capacity, unified-memory use, time to first token, generation throughput, power consumption, agent concurrency and distributed scaling efficiency. The wider research goal is to keep persistent identity and state independent from whichever foundation model is currently providing language and reasoning. I’d particularly value guidance from anyone working with MLX distributed inference, Thunderbolt/RDMA multi-Mac setups, or local agent architectures. I’ve also posted an architecture-level overview in the MLX GitHub community and have a one-page public brief available for anyone interested in the wider design. https://github.com/ml-explore/mlx/discussions/4482
Replies
2
Boosts
0
Views
431
Activity
1w
我基于Swift语言和CoreML框架开发了一个ai记忆模型
可以部署在本地,可以自动提取agent的上下文和工作并整理成记忆,用以约束智能体的工作,突破智能体上下文和llm上下文的限制,目前实测可以存储2.8万条记忆,提取速度38ms,我想跟大家讨论一下,他的应用面如何?
Replies
1
Boosts
0
Views
115
Activity
1w
New Siri AI Stuck
I’ve tried every online recommendation but nothing works, someone helps me out , my WiFi is pretty good, 100GB storage space and I have left the phone plugged in for 5 consecutive days
Replies
0
Boosts
0
Views
653
Activity
1w
Model Guardrails Too Restrictive?
I'm experimenting with using the Foundation Models framework to do news summarization in an RSS app but I'm finding that a lot of articles are getting kicked back with a vague message about guardrails. This seems really common with political news but we're talking mainstream stuff, i.e. Politico, etc. If the models are this restrictive, this will be tough to use. Is this intended? FB17904424
Replies
10
Boosts
5
Views
2.6k
Activity
1w
tensorflow 2.20 broken support
Hi, testing latest tensorflow-metal plugin with tensorflow 2.20 doesn't work.. using python Python 3.12.11 (main, Jun 3 2025, 15:41:47) [Clang 17.0.0 (clang-1700.0.13.3)] on darwin simple testing shows error: import tensorflow as tf Traceback (most recent call last): File "", line 1, in File "/Users/obg/npu/venv-tf/lib/python3.12/site-packages/tensorflow/init.py", line 438, in _ll.load_library(_plugin_dir) File "/Users/obg/npu/venv-tf/lib/python3.12/site-packages/tensorflow/python/framework/load_library.py", line 151, in load_library py_tf.TF_LoadLibrary(lib) tensorflow.python.framework.errors_impl.NotFoundError: dlopen(/Users/obg/npu/venv-tf/lib/python3.12/site-packages/tensorflow-plugins/libmetal_plugin.dylib, 0x0006): Library not loaded: @rpath/_pywrap_tensorflow_internal.so Referenced from: <8B62586B-B082-3113-93AB-FD766A9960AE> /Users/obg/npu/venv-tf/lib/python3.12/site-packages/tensorflow-plugins/libmetal_plugin.dylib Reason: tried: '/Users/obg/npu/venv-tf/lib/python3.12/site-packages/tensorflow-plugins/../_solib_darwin_arm64/_U@local_Uconfig_Utf_S_S_C_Upywrap_Utensorflow_Uinternal___Uexternal_Slocal_Uconfig_Utf/_pywrap_tensorflow_internal.so' (no such file), '/Users/obg/npu/venv-tf/lib/python3.12/site-packages/tensorflow-plugins/../_solib_darwin_arm64/_U@local_Uconfig_Utf_S_S_C_Upywrap_Utensorflow_Uinternal___Uexternal_Slocal_Uconfig_Utf/_pywrap_tensorflow_internal.so' (no such file), '/opt/homebrew/lib/_pywrap_tensorflow_internal.so' (no such file), '/System/Volumes/Preboot/Cryptexes/OS/opt/homebrew/lib/_pywrap_tensorflow_internal.so' (no such file) tf.config.experimental.list_physical_devices('GPU') Traceback (most recent call last): File "", line 1, in NameError: name 'tf' is not defined I fixed this error by copying _pywrap_tensorflow_internal.so where it's searched.. 1)mkdir /Users/obg/npu/venv-tf/lib/python3.12/site-packages/tensorflow-plugins/../_solib_darwin_arm64 2)mkdir /Users/obg/npu/venv-tf/lib/python3.12/site-packages/tensorflow-plugins/../_solib_darwin_arm64/_U@local_Uconfig_Utf_S_S_C_Upywrap_Utensorflow_Uinternal___Uexternal_Slocal_Uconfig_Utf/ 3)cp /Users/obg/npu/venv-tf/lib/python3.12/site-packages/tensorflow/python/_pywrap_tensorflow_internal.so /Users/obg/npu/venv-tf/lib/python3.12/site-packages/tensorflow-plugins/../_solib_darwin_arm64/_U@local_Uconfig_Utf_S_S_C_Upywrap_Utensorflow_Uinternal___Uexternal_Slocal_Uconfig_Utf/ then fails symbol not found: Symbol not found: __ZN10tensorflow28_AttrValue_default_instance_E in libmetal_plugin.dylib full log: with import tensorflow as tf Traceback (most recent call last): File "", line 1, in File "/Users/obg/npu/venv-tf/lib/python3.12/site-packages/tensorflow/init.py", line 438, in _ll.load_library(_plugin_dir) File "/Users/obg/npu/venv-tf/lib/python3.12/site-packages/tensorflow/python/framework/load_library.py", line 151, in load_library py_tf.TF_LoadLibrary(lib) tensorflow.python.framework.errors_impl.NotFoundError: dlopen(/Users/obg/npu/venv-tf/lib/python3.12/site-packages/tensorflow-plugins/libmetal_plugin.dylib, 0x0006): Symbol not found: __ZN10tensorflow28_AttrValue_default_instance_E Referenced from: <8B62586B-B082-3113-93AB-FD766A9960AE> /Users/obg/npu/venv-tf/lib/python3.12/site-packages/tensorflow-plugins/libmetal_plugin.dylib Expected in: <2FF91C8B-0CB6-3E66-96B7-092FDF36772E> /Users/obg/npu/venv-tf/lib/python3.12/site-packages/_solib_darwin_arm64/_U@local_Uconfig_Utf_S_S_C_Upywrap_Utensorflow_Uinternal___Uexternal_Slocal_Uconfig_Utf/_pywrap_tensorflow_internal.so
Replies
4
Boosts
0
Views
2.7k
Activity
1w
Private Cloud Compute throws guardrailViolation on benign song analysis (FB24938334)
On iOS 27.0, PrivateCloudComputeLanguageModel refuses a large share of harmless requests, and the same requests often succeed when simply re-run. Our app writes a short, general-audience explanation of what a song is about. When the cloud model refuses, the on-device model with .permissiveContentTransformations answers the same prompt without trouble. PCC has no guardrail configuration, so there's nothing on our side to adjust. • Two refusals came within the first sentence of plainly benign text: "Tiny Dancer" (Elton John) at 167 characters, and "In the Ghetto" (Elvis Presley) at 169 characters, while describing snow on a Chicago morning. • Others include "Already Gone" (Eagles), "Do Ya" (ELO) and "We Didn't Start the Fire" (Billy Joel). • It's non-deterministic: "Question" (The Moody Blues) was refused, then succeeded 52 seconds later with an identical request. • Most failures are "Streamed response may contain sensitive or unsafe content", arriving mid-generation, so the rejection seems to target the model's own output rather than the input. • Rewording the instructions to steer the model toward mainstream, general-audience language didn't change the refusal rate. Filed as FB24938334 with 8 logFeedbackAttachment captures (.triggeredGuardrailUnexpectedly), each including the rolled-back rejected draft. Siri language English (US), iPhone 16 Pro. Is PCC's guardrail policy expected to be tuned for content-transformation tasks like this, or is there a recommended pattern for them?
Replies
0
Boosts
0
Views
294
Activity
1w