What is the supported DriverKit Stop/drain sequence for an IOUserClient operation queue?

Environment:

  • macOS 26.6.2 (25G83), Apple silicon
  • Xcode 26.6 (17F113)
  • DriverKit SDK 25.5

I am implementing a DriverKit IOService with an IOUserClient. This is a lifecycle and object-ownership question independent of the device protocol.

The intended design admits at most one user client during a provider lifetime. Lifecycle methods run on the provider’s default queue, while IOUserClient ExternalMethod requests run on a separate serial IODispatchQueue. At most one device request may be in flight.

The shutdown invariant we need is:

  1. Stop accepting new requests.
  2. Allow every accepted request to complete exactly once, or cancel it.
  3. Observe completion of the operation queue’s cancellation handler.
  4. Call the inherited Stop implementation last.
  5. Perform no provider access afterward.

The relevant public documentation is:

IOService::Stop: https://developer.apple.com/documentation/driverkit/ioservice/stop

IODispatchQueue::Cancel: https://developer.apple.com/documentation/driverkit/iodispatchqueue/cancel

IOService::SetDispatchQueue: https://developer.apple.com/documentation/driverkit/ioservice/setdispatchqueue

For the normal path, the proposed sequence is conceptually:

Stop(provider): close request admission operationQueue->Cancel(cancellationHandler) wait for the cancellation handler from the separate queue super::Stop(provider)

I need clarification of the complete supported public API contract:

  1. If IODispatchQueue::Cancel returns a non-success result, is its cancellation handler still guaranteed to execute? If it is not, what supported action lets Stop keep the provider and user client valid until previously accepted work is no longer capable of accessing them?

  2. Is it supported for the provider and its one user client to share the provider-owned serial operation queue? If the IOUserClient stops independently, must it own and cancel a separate queue, or is there a supported per-client drain mechanism that does not cancel provider-owned work?

  3. Is the driver’s public IOService::Stop override guaranteed to run on every termination path where accepted user-client work must be drained, including when the provider is already inactive or the DriverKit server has slept? If not, which public lifecycle callback supplies that drain point?

  4. Is blocking the provider’s default queue inside Stop while awaiting the cancellation handler from a separate operation queue the supported interpretation of “wait for your cancellation handlers”? If not, what public continuation mechanism should be used before calling inherited Stop?

We also observed one power-management panic after sleep/wake:

HiMDScsiDriver::setPowerState(..., 0 -> 4) timed out after 20342 ms

The DEXT does not currently override SetPowerState. This panic motivates the lifecycle review, but I am not treating it as proof that the Stop/drain design caused the timeout.

I am looking specifically for a supported public DriverKit sequence. I do not want to rely on private framework entry points or infer object-lifetime guarantees from a successful build or experiment.

Answered by DTS Engineer in 904952022

First, I'd strongly recommend you take a look at the "Managing Device Removal" section of "IOKit Fundamentals", particularly "The Phases of Device Removal". That's the foundation of how this works, and I'll be referring back to it below.

So, with that context, the first thing to understand is that (kernel) IOService:stop and (DriverKit) IOService:Stop are NOT direct equivalents. Within the kernel, "stop" is actually the first step in final object destruction and, most critically, it won't be called until normal activity has stopped and I/O cleared. In practical terms, the reaching "stop" means that your driver has ALREADY stopped "working" and is functionally "dead".

That's NOT true of (DriverKit) IOService:Stop. More specifically, "Stop" is actually called during phase 2 as part of termination. Note these comments in the class reference:

"Before terminating the object in the provider, the system calls this method to stop the service associated with that object."

and

"After you call super, it is a programmer error to access the provider object."

Flipping those statements around, until your DEXT calls "super:Stop", its provider is still valid and fully usable. Similarly:

"If your driver has any in-progress asynchronous tasks, cancel those tasks and wait for DriverKit to call the associated cancellation handler before calling the super version of this method."

Meaning, until your DEXT calls "super::Stop"... it hasn't actually "stopped".

Understanding that last point is critical. Just like IOKit, device termination is a process your DEXT is part of, not something that's "done" to your DEXT. Indeed, the most common way device termination fails is that your DEXT doesn't tear down, leaving it "live" indefinitely.

Shifting into specifics:

Is blocking the provider’s default queue inside Stop while awaiting the cancellation handler from a separate operation queue the supported interpretation of “wait for your cancellation handlers”? If not, what public continuation mechanism should be used before calling inherited Stop?

So, I think the documentation was somewhat poorly phrased when it said:

"Use your implementation of this method to stop all activity and put your driver in a quiescent state.... wait for DriverKit… calling the super version of this method"

That natural reading of that is that you should block inside "Stop()", but there's actually no reason to do so. What's more typical, assuming the implementation isn't trivial, is to do what our IOUserClient sample code does, which is to fire off its cleanup work, then call "super::Stop()" when a callback determines that work is done.

If IODispatchQueue::Cancel returns a non-success result, is its cancellation handler still guaranteed to execute?

As far as I can tell, our "::Cancel" methods never actually fail. More specifically, the majority of them are hard-coded to return "kIOReturnSuccess". Most of them don't have a failure path at all, and the few exceptions I've found assert on any failure instead of returning. Somewhat amusingly, it looks like the author of IOInterruptDispatchSource() shared your concern, as its only failure path is an assert checking that a different cancel method didn't fail.

Honestly, I think I'd copy that approach and use an assert to confirm "kIOReturnSuccess". If cancellation fails, that’s a change/bug you need to investigate and resolve, not a normal behavior you can anticipate.

Is it supported for the provider and its one-user client to share the provider-owned serial operation queue?

This depends entirely on the driver and its user client. It's fairly typical for the provider to own all "work", with the user client simply passing commands to it. There are other cases where the user client is doing substantial tracking alongside its provider. It really just depends on what you're trying to do.

If the IOUserClient stops independently,

Keep in mind that user client termination is relatively common, since it can be triggered by things like the connecting process termination.

must it own and cancel a separate queue,

Most user clients are doing some amount of work to track the work they're managing on behalf of their client; however, there's a lot of variation in how complex that actually is.

or is there a supported per-client drain mechanism that does not cancel provider-owned work?

We don't have a specific API for this, as the details vary too much between use cases and implementations.

__
Kevin Elliott
DTS Engineer, CoreOS/Hardware

Accepted Answer

First, I'd strongly recommend you take a look at the "Managing Device Removal" section of "IOKit Fundamentals", particularly "The Phases of Device Removal". That's the foundation of how this works, and I'll be referring back to it below.

So, with that context, the first thing to understand is that (kernel) IOService:stop and (DriverKit) IOService:Stop are NOT direct equivalents. Within the kernel, "stop" is actually the first step in final object destruction and, most critically, it won't be called until normal activity has stopped and I/O cleared. In practical terms, the reaching "stop" means that your driver has ALREADY stopped "working" and is functionally "dead".

That's NOT true of (DriverKit) IOService:Stop. More specifically, "Stop" is actually called during phase 2 as part of termination. Note these comments in the class reference:

"Before terminating the object in the provider, the system calls this method to stop the service associated with that object."

and

"After you call super, it is a programmer error to access the provider object."

Flipping those statements around, until your DEXT calls "super:Stop", its provider is still valid and fully usable. Similarly:

"If your driver has any in-progress asynchronous tasks, cancel those tasks and wait for DriverKit to call the associated cancellation handler before calling the super version of this method."

Meaning, until your DEXT calls "super::Stop"... it hasn't actually "stopped".

Understanding that last point is critical. Just like IOKit, device termination is a process your DEXT is part of, not something that's "done" to your DEXT. Indeed, the most common way device termination fails is that your DEXT doesn't tear down, leaving it "live" indefinitely.

Shifting into specifics:

Is blocking the provider’s default queue inside Stop while awaiting the cancellation handler from a separate operation queue the supported interpretation of “wait for your cancellation handlers”? If not, what public continuation mechanism should be used before calling inherited Stop?

So, I think the documentation was somewhat poorly phrased when it said:

"Use your implementation of this method to stop all activity and put your driver in a quiescent state.... wait for DriverKit… calling the super version of this method"

That natural reading of that is that you should block inside "Stop()", but there's actually no reason to do so. What's more typical, assuming the implementation isn't trivial, is to do what our IOUserClient sample code does, which is to fire off its cleanup work, then call "super::Stop()" when a callback determines that work is done.

If IODispatchQueue::Cancel returns a non-success result, is its cancellation handler still guaranteed to execute?

As far as I can tell, our "::Cancel" methods never actually fail. More specifically, the majority of them are hard-coded to return "kIOReturnSuccess". Most of them don't have a failure path at all, and the few exceptions I've found assert on any failure instead of returning. Somewhat amusingly, it looks like the author of IOInterruptDispatchSource() shared your concern, as its only failure path is an assert checking that a different cancel method didn't fail.

Honestly, I think I'd copy that approach and use an assert to confirm "kIOReturnSuccess". If cancellation fails, that’s a change/bug you need to investigate and resolve, not a normal behavior you can anticipate.

Is it supported for the provider and its one-user client to share the provider-owned serial operation queue?

This depends entirely on the driver and its user client. It's fairly typical for the provider to own all "work", with the user client simply passing commands to it. There are other cases where the user client is doing substantial tracking alongside its provider. It really just depends on what you're trying to do.

If the IOUserClient stops independently,

Keep in mind that user client termination is relatively common, since it can be triggered by things like the connecting process termination.

must it own and cancel a separate queue,

Most user clients are doing some amount of work to track the work they're managing on behalf of their client; however, there's a lot of variation in how complex that actually is.

or is there a supported per-client drain mechanism that does not cancel provider-owned work?

We don't have a specific API for this, as the details vary too much between use cases and implementations.

__
Kevin Elliott
DTS Engineer, CoreOS/Hardware

Thank you. This resolves the normal Stop path, asynchronous cleanup pattern, Cancel-success expectation, and provider-owned work model.

We have a corrected lifecycle, but one blocker remains with one remaining clarification from question 3:

Does the public IOService::Stop phase-2/provider-validity guidance also apply when the provider is already inactive or the DriverKit user server has slept?

In those cases, is the driver’s public Stop override still guaranteed to be invoked, with the provider remaining valid until the driver’s asynchronous cancellation callback calls super::Stop?

I am asking only which public lifecycle guarantee the driver may rely on, not how any private framework entry point is implemented.

Does the public IOService::Stop phase-2/provider-validity guidance also apply when the provider is already inactive or the DriverKit user server has slept?

In those cases, is the driver’s public Stop override still guaranteed to be invoked, with the provider remaining valid until the driver’s asynchronous cancellation callback calls super::Stop?

Yes. There's only one valid[1] driver termination sequence, namely the one that goes through "Stop".

[1] Strictly speaking, your DEXT can also crash or be terminated. If/when that happens, the details of what the kernel driver will do is determined by the individual driver family, but in most cases the hardware will either become nonfunctional (until hot plug/reboot) or destroy itself automatically. However, what happens after your DEXT is "gone" isn't really your DEXT's problem.

__
Kevin Elliott
DTS Engineer, CoreOS/Hardware

Thank you, Kevin. That fully resolves the remaining ambiguity for us.

Your explanation of the single valid termination sequence, the provider’s validity until super::Stop, and the callback-driven cleanup pattern gives us the lifecycle contract we needed to proceed safely.

I also appreciate the clarification about abnormal DEXT termination. We will treat that as a failure condition rather than attempting to incorporate it into the normal shutdown protocol.

Thank you for taking the time to answer the follow-up so precisely.

I have one USB timing follow-up while validating the shutdown sequence discussed here. The earlier Stop/lifetime guidance is resolved.

For USBDriverKit IOUSBHostPipe::CompleteAsyncIO, is completionTimestamp guaranteed to use the same clock epoch and tick units as mach_absolute_time() on supported macOS releases? The DriverKit 25.5 header and documentation describe it as the absolute time the transfer completed, but I could not find an explicit clock definition.

The intended observation is one successfully submitted AsyncIO read. When child Stop closes admission to new requests, we would immediately record mach_absolute_time(), then compare that value with the transfer completionTimestamp. We need to establish whether the transfer was still in progress at that point; a callback delivered later could describe a transfer that had already completed.

If that timestamp comparison is unsupported, which documented API or callback observation can establish this ordering without relying on a marker before submission or on callback delivery time?

API reference: https://developer.apple.com/documentation/usbdriverkit/iousbhostpipe/completeasyncio

For USBDriverKit IOUSBHostPipe::CompleteAsyncIO, is completionTimestamp guaranteed to use the same clock epoch and tick units as mach_absolute_time() on supported macOS releases?

Our documentation should probably be more explicit about these details, but whenever we refer to "absolute time", we generally mean "mach_absolute_time". There's only one clock in the system (all other times are derived from it), so the only difference between clocks is how they've been shifted to different formats or time bases. I'll also note that there's generally enough divergence between them that inferring the source is relatively straightforward, particularly if you take a moment to force divergence. For example, mach_absolute_time and mach_continuous_time start at the same point, but quickly diverge as soon as the system is allowed to sleep.

In any case, its current value from the kernel is literally:

completionTimestamp = mach_absolute_time();

...and I can't really think of any reason we'd change that.

We need to establish whether the transfer was still in progress at that point; a callback delivered later could describe a transfer that had already completed.

What's your underlying goal here?

I think the tricky part here is that there are lots of points in this process that inject small amounts of unpredictable latency, which makes precise time comparisons a lot more complicated. For example, the timestamp you receive in CompleteAsyncIO was added in an async callback, which means it’s technically marking some point "after" the command complete, not the EXACT time the command completed. Similarly, the queuing of events going into your DEXT means that "Stop" only happens "after" existing work is finished, not a particular instant in time.

Now, in real-world terms those timing differences are infinitesimal and basically irrelevant, but if you're trying to "prove" something about the very specific timing of events, then they very well could matter.

__
Kevin Elliott
DTS Engineer, CoreOS/Hardware

Thanks Kevin, that clears it up - particularly where the completion timestamp is actually recorded.

The underlying goal is to validate safe shutdown of our USB DriverKit driver. We want to check that it stops accepting requests, lets outstanding work finish, and only then releases resources and calls inherited Stop, following your earlier guidance.

I think we made the timing question more exacting than it needed to be. We don’t need to identify the precise instant USB activity ends; we need to verify the cleanup ordering and measure how long the sequence takes.

We’ll keep the lifecycle ordering checks separate from the timing measurements, then use repeated tests on real machines to establish observed ranges and look for stalls or races. Those measurements won’t establish universal timing guarantees, but they should give us useful practical evidence and we are dealing with a known-set of devices, 10 unique models, so that is a very manageable set.

Thanks again for the explanation and your patience - I think we have enough to work from here.

What is the supported DriverKit Stop/drain sequence for an IOUserClient operation queue?
 
 
Q