Supported completeness and lifecycle guarantees for es_new_descendants_client

Hello Apple Developer Technical Support,

I am evaluating es_new_descendants_client for a local command runner that must report success only after its workload and every process descended from that workload have exited. If observation is incomplete or ambiguous, the runner must report failure. This is a design inquiry, not a report of a reproduced operating-system defect; no entitled prototype has been tested.

The proposed observer would create its client and subscribe to lifecycle notifications before launching any workload. It would maintain a registry using process-lifetime identities, add processes on creation and remove them on exit. An unmatched event, missing required field, detected loss or observer failure would invalidate the run. It would consider closure only after all registered workload processes had exited. We have not established that these rules are sufficient.

Could you clarify which of the following properties are supported API guarantees, and identify any that applications must not rely on? A documented reference or an explicit statement that a guarantee is unavailable would both help. Please identify applicable macOS/SDK versions and any known version-dependent limitations.

1. Membership and creation-event coverage

Does the observed subtree retain a process and all of its future descendants after its original parent exits, it is reparented, it double-forks, or it changes process group/session with setpgid or setsid? Could a process remain observable for exit while creation events for its children become invisible?

For a workload launched after successful subscription, does every successful process-creation path—including fork, vfork and posix_spawn—produce a lifecycle event sufficient to register the new process before closure can be declared? Which event and identity fields should be used for each path, including a child that exits without a successful exec? Does the calling observer receive the necessary event for its own initial workload launch?

2. Ordering and the meaning of exit

Is there a supported per-client ordering guarantee that every child-creation event from a process is delivered before that process's exit notification, including concurrent creation and exit? Can the child's events arrive before the event that introduces that child? Please distinguish kernel enqueue order, handler delivery order and any processing order the application must impose.

At what lifecycle boundary is ES_EVENT_TYPE_NOTIFY_EXIT generated? Does it establish that the identified process can no longer execute or initiate writes, or can relevant activity continue after the notification? We would not equate process exit with filesystem durability or completion of work already delegated to other processes.

3. Muting and other visibility filters

Does a newly created descendants client have default process, path or target-path mutes that can suppress fork/exit notifications? What supported sequence of configuration and inspection calls establishes complete lifecycle visibility before launch, including mute inversion and executable-path changes?

Apart from subscription and muting, are there policy, security, rate-limit or client-type exclusions that can suppress those events? Which suppressed events, if any, are intentionally absent from the sequence counter rather than reported as drops?

4. Sequence numbers and loss detection

The global_seq_num documentation requires message version greater than 4. Is that field guaranteed for descendants-client lifecycle messages? Do notifications concerning the calling observer and its descendants use the same per-client sequence?

How can a client establish a valid initial baseline and detect loss before its first received message? Is every dropped subscribed, unmuted lifecycle event reflected in the next delivered sequence number? What counter reset, wraparound or client-recreation rules must be handled?

Would the proposed registry rule make terminal loss fail safely—for example, a lost final exit leaves a process registered—under the supported ordering and visibility semantics? Or is there a counterexample in which the registry can become empty while an unobserved descendant survives?

5. Synchronization, observer failure and delegated work

Does es_sync_client provide any loss/completeness information beyond draining preceding queued messages? Its documented callbacks also run for a destroyed or null client, so we would not interpret callback arrival alone as successful completion. Is there a supported mechanism to distinguish a healthy drain from invalidation?

What does “instigates” cover for this client? In particular, can it observe or attribute work executed by existing launchd/XPC services, or by unrelated processes receiving file descriptors? We would treat such work as outside a lineage-only closure claim unless it is explicitly covered or independently excluded.

Does this client provide any supported protection against a same-UID workload stopping, killing or otherwise interfering with its observer, or must that isolation be supplied separately? Observer failure would invalidate the run; we are not assuming ES supplies a write barrier for evidence files.

6. Supported cleanup and deployment

Is there a supported public mechanism to signal a non-child descendant by process-lifetime identity, without a PID-reuse race between observing it and sending a signal? Is there a recommended approach if the observer cannot wait on that process? We do not want to depend on private libproc functions as an application contract.

Finally, is this use case eligible for com.apple.developer.endpoint-security.client in a standalone signed command-line observer, and what supported signing/provisioning or packaging requirements apply? This is a request for guidance, not an entitlement application.

Our central question is whether supported APIs can establish complete descendant-process closure under these constraints. If they cannot, we would appreciate a clear statement of that limitation or a supported alternative.

Thank you.

Documentation consulted:

Answered by DTS Engineer in 907383022

Does the observed subtree retain a process and all of its future descendants after its original parent exits, it is reparented, it double-forks, or it changes process group/session with setpgid or setsid? Could a process remain observable for exit while creation events for its children become invisible?

It is not possible for a child process to become invisible to its managing client.

For a workload launched after successful subscription, does every successful process-creation path—including fork, vfork, and posix_spawn—produce a lifecycle event sufficient to register the new process before closure can be declared?

Endpoint security is built on top of kauth's hooks, which means its monitoring occurs at critical junction points which simply cannot be bypassed. However, that also means that there isn't always a direct mapping for every system call.

Case in point, there is no event for posix_spawn because posix_spawn is actually implemented by simply calling "fork" and then "exec" in the kernel, which is where ES authorizes these events. However, keep in mind that this also means the functions being auth'd in the kernel don't necessarily match your expectations. See this forum thread on setuid and posix_spawn for example.

Which event and identity fields should be used for each path, including a child that exits without a successful exec?

As I mentioned above, there are only fork and exec are the two ES events. All process creation goes through those two functions.

Does the calling observer receive the necessary event for its own initial workload launch?

I'm not sure what you mean here.

Is there a supported per-client ordering guarantee that every child-creation event from a process is delivered before that process's exit notification, including concurrent creation and exit?

You need to differentiate between auth events and notifications here. Syscalls block waiting on auth events, so you can't really receive an auth event "after" a process has exited. Theoretically, it's possible that timing might align such that exit occurred while your app was processing that auth event, but in practice, I think the notification latency is high enough that this won't actually happen.

Can the child's events arrive before the event that introduces that child?

Notifications are designed to be higher latency than auth events, so it's certainly possible for the auth event of an action to arrive before the notify generated by an earlier syscall. However, in the case of fork, I think you're virtually certain to receive fork before any new process activity occurs simply because of the time required for the new process to become functional.

At what lifecycle boundary is ES_EVENT_TYPE_NOTIFY_EXIT generated?

It's sent by the kernel as it's going through the process of destroying the process. The specific point is here.

https://github.com/apple-oss-distributions/xnu/blob/f6217f891ac0bb64f3d375211650a4c1ff8ca1ea/bsd/kern/kern_exit.c#L2488

Does it establish that the identified process can no longer execute or initiate writes, or can relevant activity continue after the notification?

Very much so.

Does a newly created descendant client have default process, path, or target-path mutes that can suppress fork/exit notifications? What supported sequence of configuration and inspection calls establishes complete lifecycle visibility before launch, including mute inversion and executable-path changes?

I don't think any of the default mutes would apply to your child processes; however, the more relevant point is that you can fully configure everything before you spawn your first child.

Apart from subscription and muting, are there policy, security, rate-limit, or client-type exclusions that can suppress those events? Which suppressed events, if any, are intentionally absent from the sequence counter rather than reported as drops?

In the case of es_new_descendants_client, you can directly control your own deadline, including making it unbounded. Assuming "reasonable" client design, I don't think you'd actually miss anything.

  1. Sequence numbers and loss detection

I think all of these questions are well covered by the header file, but please let me know if you're still confused after looking through it.

Its documented callbacks also run for a destroyed or null client, so we would not interpret callback arrival alone as successful completion. Is there a supported mechanism to distinguish a healthy drain from invalidation?

I'm deeply confused by this question. The only reason an ES client would be destroyed is because you destroyed it. es_sync_client runs on client destruction as a way to simplify resource management and cleanup, not because there's any confusion about why the client was destroyed.

Does this client provide any supported protection against a same-UID workload stopping, killing, or otherwise interfering with its observer, or must that isolation be supplied separately?

The client can protect itself from any manipulation coming from the children it's monitoring, but it cannot protect itself from the broader system.

Is there a supported public mechanism to signal a non-child descendant by process-lifetime identity, without a PID-reuse race between observing it and sending a signal?

What are you trying to do here?

Finally, is this use case eligible for com.apple.developer.endpoint-security.client in a standalone signed command-line observer, and what supported signing/provisioning or packaging requirements apply?

Our signing infrastructure is going to require the client be inside a bundle.

__
Kevin Elliott
DTS Engineer, CoreOS/Hardware

Accepted Answer

Does the observed subtree retain a process and all of its future descendants after its original parent exits, it is reparented, it double-forks, or it changes process group/session with setpgid or setsid? Could a process remain observable for exit while creation events for its children become invisible?

It is not possible for a child process to become invisible to its managing client.

For a workload launched after successful subscription, does every successful process-creation path—including fork, vfork, and posix_spawn—produce a lifecycle event sufficient to register the new process before closure can be declared?

Endpoint security is built on top of kauth's hooks, which means its monitoring occurs at critical junction points which simply cannot be bypassed. However, that also means that there isn't always a direct mapping for every system call.

Case in point, there is no event for posix_spawn because posix_spawn is actually implemented by simply calling "fork" and then "exec" in the kernel, which is where ES authorizes these events. However, keep in mind that this also means the functions being auth'd in the kernel don't necessarily match your expectations. See this forum thread on setuid and posix_spawn for example.

Which event and identity fields should be used for each path, including a child that exits without a successful exec?

As I mentioned above, there are only fork and exec are the two ES events. All process creation goes through those two functions.

Does the calling observer receive the necessary event for its own initial workload launch?

I'm not sure what you mean here.

Is there a supported per-client ordering guarantee that every child-creation event from a process is delivered before that process's exit notification, including concurrent creation and exit?

You need to differentiate between auth events and notifications here. Syscalls block waiting on auth events, so you can't really receive an auth event "after" a process has exited. Theoretically, it's possible that timing might align such that exit occurred while your app was processing that auth event, but in practice, I think the notification latency is high enough that this won't actually happen.

Can the child's events arrive before the event that introduces that child?

Notifications are designed to be higher latency than auth events, so it's certainly possible for the auth event of an action to arrive before the notify generated by an earlier syscall. However, in the case of fork, I think you're virtually certain to receive fork before any new process activity occurs simply because of the time required for the new process to become functional.

At what lifecycle boundary is ES_EVENT_TYPE_NOTIFY_EXIT generated?

It's sent by the kernel as it's going through the process of destroying the process. The specific point is here.

https://github.com/apple-oss-distributions/xnu/blob/f6217f891ac0bb64f3d375211650a4c1ff8ca1ea/bsd/kern/kern_exit.c#L2488

Does it establish that the identified process can no longer execute or initiate writes, or can relevant activity continue after the notification?

Very much so.

Does a newly created descendant client have default process, path, or target-path mutes that can suppress fork/exit notifications? What supported sequence of configuration and inspection calls establishes complete lifecycle visibility before launch, including mute inversion and executable-path changes?

I don't think any of the default mutes would apply to your child processes; however, the more relevant point is that you can fully configure everything before you spawn your first child.

Apart from subscription and muting, are there policy, security, rate-limit, or client-type exclusions that can suppress those events? Which suppressed events, if any, are intentionally absent from the sequence counter rather than reported as drops?

In the case of es_new_descendants_client, you can directly control your own deadline, including making it unbounded. Assuming "reasonable" client design, I don't think you'd actually miss anything.

  1. Sequence numbers and loss detection

I think all of these questions are well covered by the header file, but please let me know if you're still confused after looking through it.

Its documented callbacks also run for a destroyed or null client, so we would not interpret callback arrival alone as successful completion. Is there a supported mechanism to distinguish a healthy drain from invalidation?

I'm deeply confused by this question. The only reason an ES client would be destroyed is because you destroyed it. es_sync_client runs on client destruction as a way to simplify resource management and cleanup, not because there's any confusion about why the client was destroyed.

Does this client provide any supported protection against a same-UID workload stopping, killing, or otherwise interfering with its observer, or must that isolation be supplied separately?

The client can protect itself from any manipulation coming from the children it's monitoring, but it cannot protect itself from the broader system.

Is there a supported public mechanism to signal a non-child descendant by process-lifetime identity, without a PID-reuse race between observing it and sending a signal?

What are you trying to do here?

Finally, is this use case eligible for com.apple.developer.endpoint-security.client in a standalone signed command-line observer, and what supported signing/provisioning or packaging requirements apply?

Our signing infrastructure is going to require the client be inside a bundle.

__
Kevin Elliott
DTS Engineer, CoreOS/Hardware

Hi Kevin,

Thank you. I checked ESClient.h, ESMessage.h and ESTypes.h. The caller-notification scope answers my initial-launch question: I meant the observer creating its first workload after subscribing. I also understand your clarification about client destruction; we would retain the same live client through the drain and invalidate our run on shutdown.

The remaining question is whether we can soundly conclude that no workload descendant remains alive, rather than relying on likely delivery timing. This is still a design inquiry; we have not tested an entitled prototype.

1. Notify ordering and an empty registry

The headers mark fork and exit as notify-only. ESClient.h documents synchronous enqueueing before a syscall returns and serial handler delivery. We would configure visibility before launch, process fork/exec/exit notifications serially, account for exec identity transitions, and reject sequence gaps or unmatched events. We would not test for completion until the initial launch had been registered and drained.

For a tracked parent P that creates C and exits, is C's fork notification guaranteed to be enqueued and delivered before P's exit notification and before any notification from C, including when C exits without exec? Or is additional ordering logic required?

Specifically, with the same live client, can an empty registry followed by a completed es_sync_client drain and no detected loss still coexist with an unregistered live descendant? We would not treat an AUTH event as a substitute for a missing fork notification.

2. Sequence boundaries

I understand the counters detect gaps between delivered messages. What initial counter value should a newly subscribed client expect, so loss before its first received message is detectable? Does a sync marker provide any way to detect dropped messages after the last delivered message? If not, is the registry argument sufficient because a missing exit leaves a registered process outstanding, or can a missing creation notification defeat that argument?

3. Exit and timeout cleanup

The linked XNU revision calls mac_proc_notify_exit at line 2488. To disambiguate my earlier two-part question: once NOTIFY_EXIT is delivered, can that process still execute user code or initiate new writes? We are excluding storage durability and previously delegated work from this claim.

On timeout, we want to terminate any surviving descendant without accidentally signalling an unrelated process that reused its PID. A descendant might have been reparented, so the observer cannot necessarily wait on it. Is there a supported public way to signal that process using a stable lifetime identity? If not, we would report cleanup as unconfirmed rather than claim closure.

Please identify the applicable macOS/SDK versions for these guarantees. We understand that signing requires a bundle; this is not an entitlement application.

Thank you.

For a tracked parent P that creates C and exits, is C's fork notification guaranteed to be enqueued and delivered before P's exit notification and before any notification from C, including when C exits without exec? Or is additional ordering logic required?

The problem with both is with the word "guaranteed". Under any kind of "normal" usage pattern, I'd expect the fork to be enqueued and delivered before the exit as well as any notifications from C. Making that "guaranteed" requires thinking through a very large number of edge cases, all of which are very tricky to validate and prove. Putting that another way, I'm fairly confident that it's not possible for a notification to occur after an exit, but PROVING that across all possible configurations, circumstances, and system versions is much harder than you seem to think.

However, the bigger issue is that, from the ES client’s perspective, this entire line of thinking is just a bit silly. That is, the solution to this concern:

is C's fork notification guaranteed to be enqueued and delivered before P's exit notification

...isn't to try and reason or predict about the system’s behavior, but is instead to simply "wait". The event delivery pipeline is sufficiently high performance that there isn't going to be any significant delay in delivery. In terms of API, if you want to be sure that an exited process is "cleared”, then es_sync_client can ensure that, since its firing guarantees that all actions that occurred prior to that will have reached your client.

I understand the counters detect gaps between delivered messages.

Let me be clear, this is not "normal" behavior. That is, the "baseline" ES client behavior is to deliver "all" messages, with the client being responsible for processing events fast enough that the system doesn't lag. For messages that are considered particularly critical, this often means creating multiple ES client instances so that each client can receive and process events in parallel. In the case of es_new_descendants_client, you can manage your own deadline configuration such that dropped messages never occur.

What initial counter value should a newly subscribed client expect, so loss before its first received message is detectable?

You cannot lose messages before the first is received.

Does a sync marker provide any way to detect dropped messages after the last delivered message?

No, but, again, if you're concerned about dropped messages, build your client so that nothing gets dropped.

Once NOTIFY_EXIT is delivered, can that process still execute user code or initiate new writes?

No, the process is no longer capable of being executed at the point NOTIFY_EXIT occurs.

On timeout, we want to terminate any surviving descendant without accidentally signalling an unrelated process that reused its PID. A descendant might have been reparented, so the observer cannot necessarily wait on it. Is there a supported public way to signal that process using a stable lifetime identity? We do not want to depend on private libproc functions as an application contract.

While libproc is generally worth avoiding, I think I'd consider "proc_signal_with_audittoken"/"proc_terminate_with_audittoken" legitimate exceptions. Either one solves the problem with more safety and less effort than any alternative. I'll also point out that the main issue with libproc is that the API has been documented as subject to change without notice:

/*
 * This header file contains private interfaces to obtain process information.
 * These interfaces are subject to change in future releases.
 */

...but that doesn't mean every API is equally "dangerous" or likely to change.

In this case, we need an API that sends a signal to an audit token, so both functions take an audit token... and a signal number. There isn't much that can go wrong there, nor can I see any reason for it to change, short of some kind of VERY large-scale redesign of the entire architecture.

I've pinged Quinn to see if he disagrees, but I think this is a very reasonable use of libproc. I'd probably create a dedicated unit test that confirms it works as expected, but I don't think it's all that likely to ever fail.

__
Kevin Elliott
DTS Engineer, CoreOS/Hardware

I think this is a very reasonable use of libproc

Agreed.

The situation with libproc is much like the situation with Mach. We strive to maintain binary compatibility but that’s tricky because it’s very tightly bound to the kernel’s implementation.

Oh, one last thing: proc_signal_with_audittoken was added in some time during the macOS 14 release cycle, so you can’t use it before then. But I think you’ll be OK in this case because es_new_descendants_client requires macOS 27 anyway.

Share and Enjoy
—
Quinn “The Eskimo!” @ Developer Technical Support @ Apple
let myEmail = "eskimo" + "1" + "@" + "apple.com"

Supported completeness and lifecycle guarantees for es_new_descendants_client
 
 
Q