Supported completeness and lifecycle guarantees for es_new_descendants_client

Hello Apple Developer Technical Support,

I am evaluating es_new_descendants_client for a local command runner that must report success only after its workload and every process descended from that workload have exited. If observation is incomplete or ambiguous, the runner must report failure. This is a design inquiry, not a report of a reproduced operating-system defect; no entitled prototype has been tested.

The proposed observer would create its client and subscribe to lifecycle notifications before launching any workload. It would maintain a registry using process-lifetime identities, add processes on creation and remove them on exit. An unmatched event, missing required field, detected loss or observer failure would invalidate the run. It would consider closure only after all registered workload processes had exited. We have not established that these rules are sufficient.

Could you clarify which of the following properties are supported API guarantees, and identify any that applications must not rely on? A documented reference or an explicit statement that a guarantee is unavailable would both help. Please identify applicable macOS/SDK versions and any known version-dependent limitations.

1. Membership and creation-event coverage

Does the observed subtree retain a process and all of its future descendants after its original parent exits, it is reparented, it double-forks, or it changes process group/session with setpgid or setsid? Could a process remain observable for exit while creation events for its children become invisible?

For a workload launched after successful subscription, does every successful process-creation path—including fork, vfork and posix_spawn—produce a lifecycle event sufficient to register the new process before closure can be declared? Which event and identity fields should be used for each path, including a child that exits without a successful exec? Does the calling observer receive the necessary event for its own initial workload launch?

2. Ordering and the meaning of exit

Is there a supported per-client ordering guarantee that every child-creation event from a process is delivered before that process's exit notification, including concurrent creation and exit? Can the child's events arrive before the event that introduces that child? Please distinguish kernel enqueue order, handler delivery order and any processing order the application must impose.

At what lifecycle boundary is ES_EVENT_TYPE_NOTIFY_EXIT generated? Does it establish that the identified process can no longer execute or initiate writes, or can relevant activity continue after the notification? We would not equate process exit with filesystem durability or completion of work already delegated to other processes.

3. Muting and other visibility filters

Does a newly created descendants client have default process, path or target-path mutes that can suppress fork/exit notifications? What supported sequence of configuration and inspection calls establishes complete lifecycle visibility before launch, including mute inversion and executable-path changes?

Apart from subscription and muting, are there policy, security, rate-limit or client-type exclusions that can suppress those events? Which suppressed events, if any, are intentionally absent from the sequence counter rather than reported as drops?

4. Sequence numbers and loss detection

The global_seq_num documentation requires message version greater than 4. Is that field guaranteed for descendants-client lifecycle messages? Do notifications concerning the calling observer and its descendants use the same per-client sequence?

How can a client establish a valid initial baseline and detect loss before its first received message? Is every dropped subscribed, unmuted lifecycle event reflected in the next delivered sequence number? What counter reset, wraparound or client-recreation rules must be handled?

Would the proposed registry rule make terminal loss fail safely—for example, a lost final exit leaves a process registered—under the supported ordering and visibility semantics? Or is there a counterexample in which the registry can become empty while an unobserved descendant survives?

5. Synchronization, observer failure and delegated work

Does es_sync_client provide any loss/completeness information beyond draining preceding queued messages? Its documented callbacks also run for a destroyed or null client, so we would not interpret callback arrival alone as successful completion. Is there a supported mechanism to distinguish a healthy drain from invalidation?

What does “instigates” cover for this client? In particular, can it observe or attribute work executed by existing launchd/XPC services, or by unrelated processes receiving file descriptors? We would treat such work as outside a lineage-only closure claim unless it is explicitly covered or independently excluded.

Does this client provide any supported protection against a same-UID workload stopping, killing or otherwise interfering with its observer, or must that isolation be supplied separately? Observer failure would invalidate the run; we are not assuming ES supplies a write barrier for evidence files.

6. Supported cleanup and deployment

Is there a supported public mechanism to signal a non-child descendant by process-lifetime identity, without a PID-reuse race between observing it and sending a signal? Is there a recommended approach if the observer cannot wait on that process? We do not want to depend on private libproc functions as an application contract.

Finally, is this use case eligible for com.apple.developer.endpoint-security.client in a standalone signed command-line observer, and what supported signing/provisioning or packaging requirements apply? This is a request for guidance, not an entitlement application.

Our central question is whether supported APIs can establish complete descendant-process closure under these constraints. If they cannot, we would appreciate a clear statement of that limitation or a supported alternative.

Thank you.

Documentation consulted:

Does the observed subtree retain a process and all of its future descendants after its original parent exits, it is reparented, it double-forks, or it changes process group/session with setpgid or setsid? Could a process remain observable for exit while creation events for its children become invisible?

It is not possible for a child process to become invisible to its managing client.

For a workload launched after successful subscription, does every successful process-creation path—including fork, vfork, and posix_spawn—produce a lifecycle event sufficient to register the new process before closure can be declared?

Endpoint security is built on top of kauth's hooks, which means its monitoring occurs at critical junction points which simply cannot be bypassed. However, that also means that there isn't always a direct mapping for every system call.

Case in point, there is no event for posix_spawn because posix_spawn is actually implemented by simply calling "fork" and then "exec" in the kernel, which is where ES authorizes these events. However, keep in mind that this also means the functions being auth'd in the kernel don't necessarily match your expectations. See this forum thread on setuid and posix_spawn for example.

Which event and identity fields should be used for each path, including a child that exits without a successful exec?

As I mentioned above, there are only fork and exec are the two ES events. All process creation goes through those two functions.

Does the calling observer receive the necessary event for its own initial workload launch?

I'm not sure what you mean here.

Is there a supported per-client ordering guarantee that every child-creation event from a process is delivered before that process's exit notification, including concurrent creation and exit?

You need to differentiate between auth events and notifications here. Syscalls block waiting on auth events, so you can't really receive an auth event "after" a process has exited. Theoretically, it's possible that timing might align such that exit occurred while your app was processing that auth event, but in practice, I think the notification latency is high enough that this won't actually happen.

Can the child's events arrive before the event that introduces that child?

Notifications are designed to be higher latency than auth events, so it's certainly possible for the auth event of an action to arrive before the notify generated by an earlier syscall. However, in the case of fork, I think you're virtually certain to receive fork before any new process activity occurs simply because of the time required for the new process to become functional.

At what lifecycle boundary is ES_EVENT_TYPE_NOTIFY_EXIT generated?

It's sent by the kernel as it's going through the process of destroying the process. The specific point is here.

https://github.com/apple-oss-distributions/xnu/blob/f6217f891ac0bb64f3d375211650a4c1ff8ca1ea/bsd/kern/kern_exit.c#L2488

Does it establish that the identified process can no longer execute or initiate writes, or can relevant activity continue after the notification?

Very much so.

Does a newly created descendant client have default process, path, or target-path mutes that can suppress fork/exit notifications? What supported sequence of configuration and inspection calls establishes complete lifecycle visibility before launch, including mute inversion and executable-path changes?

I don't think any of the default mutes would apply to your child processes; however, the more relevant point is that you can fully configure everything before you spawn your first child.

Apart from subscription and muting, are there policy, security, rate-limit, or client-type exclusions that can suppress those events? Which suppressed events, if any, are intentionally absent from the sequence counter rather than reported as drops?

In the case of es_new_descendants_client, you can directly control your own deadline, including making it unbounded. Assuming "reasonable" client design, I don't think you'd actually miss anything.

  1. Sequence numbers and loss detection

I think all of these questions are well covered by the header file, but please let me know if you're still confused after looking through it.

Its documented callbacks also run for a destroyed or null client, so we would not interpret callback arrival alone as successful completion. Is there a supported mechanism to distinguish a healthy drain from invalidation?

I'm deeply confused by this question. The only reason an ES client would be destroyed is because you destroyed it. es_sync_client runs on client destruction as a way to simplify resource management and cleanup, not because there's any confusion about why the client was destroyed.

Does this client provide any supported protection against a same-UID workload stopping, killing, or otherwise interfering with its observer, or must that isolation be supplied separately?

The client can protect itself from any manipulation coming from the children it's monitoring, but it cannot protect itself from the broader system.

Is there a supported public mechanism to signal a non-child descendant by process-lifetime identity, without a PID-reuse race between observing it and sending a signal?

What are you trying to do here?

Finally, is this use case eligible for com.apple.developer.endpoint-security.client in a standalone signed command-line observer, and what supported signing/provisioning or packaging requirements apply?

Our signing infrastructure is going to require the client be inside a bundle.

__
Kevin Elliott
DTS Engineer, CoreOS/Hardware

Hi Kevin,

Thank you. I checked ESClient.h, ESMessage.h and ESTypes.h. The caller-notification scope answers my initial-launch question: I meant the observer creating its first workload after subscribing. I also understand your clarification about client destruction; we would retain the same live client through the drain and invalidate our run on shutdown.

The remaining question is whether we can soundly conclude that no workload descendant remains alive, rather than relying on likely delivery timing. This is still a design inquiry; we have not tested an entitled prototype.

1. Notify ordering and an empty registry

The headers mark fork and exit as notify-only. ESClient.h documents synchronous enqueueing before a syscall returns and serial handler delivery. We would configure visibility before launch, process fork/exec/exit notifications serially, account for exec identity transitions, and reject sequence gaps or unmatched events. We would not test for completion until the initial launch had been registered and drained.

For a tracked parent P that creates C and exits, is C's fork notification guaranteed to be enqueued and delivered before P's exit notification and before any notification from C, including when C exits without exec? Or is additional ordering logic required?

Specifically, with the same live client, can an empty registry followed by a completed es_sync_client drain and no detected loss still coexist with an unregistered live descendant? We would not treat an AUTH event as a substitute for a missing fork notification.

2. Sequence boundaries

I understand the counters detect gaps between delivered messages. What initial counter value should a newly subscribed client expect, so loss before its first received message is detectable? Does a sync marker provide any way to detect dropped messages after the last delivered message? If not, is the registry argument sufficient because a missing exit leaves a registered process outstanding, or can a missing creation notification defeat that argument?

3. Exit and timeout cleanup

The linked XNU revision calls mac_proc_notify_exit at line 2488. To disambiguate my earlier two-part question: once NOTIFY_EXIT is delivered, can that process still execute user code or initiate new writes? We are excluding storage durability and previously delegated work from this claim.

On timeout, we want to terminate any surviving descendant without accidentally signalling an unrelated process that reused its PID. A descendant might have been reparented, so the observer cannot necessarily wait on it. Is there a supported public way to signal that process using a stable lifetime identity? If not, we would report cleanup as unconfirmed rather than claim closure.

Please identify the applicable macOS/SDK versions for these guarantees. We understand that signing requires a bundle; this is not an entitlement application.

Thank you.

Supported completeness and lifecycle guarantees for es_new_descendants_client
 
 
Q