fseventsd cleanup cursor exhaustion and notification stalls on macOS 26.7.1

Feedback ID: FB25090184

I've been tracking recurring fseventsd memory growth and slow file-change notifications on macOS 26.7.1 (25G241), with massive Bazel build trees. There seems to be a repeatable boundary inside event-table cleanup. I may be missing another reset path, but the live observations match the disassembly closely.

In fseventsd 1413.160.2, ARM64E UUID CDD7E21B-7F60-3DD4-AA03-8B36745DC0F4, cleanup at RVA 0x4604 uses a 32-bit scan cursor at table+0x14. The guard at 0x464c/0x4650 skips scanning above 0xfffffdff. Each complete pass advances by 512. Bucket indexing masks the cursor, but the stored cursor keeps increasing. Starting at zero, it stops at 0xfffffe00 after 8388607 complete passes.

Table resize resets the cursor. Two observed October 3 episodes reached exactly that value, accumulated zero-reference records, and recovered without a restart when the table doubled. The resize loop rehashes records rather than freeing them, so resumed cleanup after the reset seems to explain the recovery.

The reviewed resize limit is 262144 buckets. On October 6, after the table reached that size, the next cursor exhaustion persisted. Physical footprint grew from 26.6 MB to 2.31 GB in about six minutes, and fresh FSEvents probes reached an 8-second timeout. A bounded walk found 199944 zero-reference records among 200000 visited. The walk was partial and unsuspended, with linkage errors; that is not a whole-heap count.

Restarting only fseventsd restored advancing cleanup, approximately 5 MB footprint and millisecond probes. The client mix changed after restart, so this is not a controlled fixed-client comparison. Finder stayed running, and Desktop add/delete checks worked on both sides. An earlier stale-Desktop symptom may be a separate issue.

The boundary and table-size prediction were recorded before the live observations. The Feedback report has the timeline, graphs, counter snapshots, record walks, stack excerpts, disassembly and read-only sampler source. No daemon binary or target memory was changed. Live counter access used a temporary SIP debugging exception; a severe memory/notification failure was also captured before that exception. There is not yet a small standalone workload reproducer or a patched-build comparison.

Is the scan cursor supposed to wrap or reset independently of table growth? If this is the expected mechanism, is there another cleanup path that should keep reclaiming zero-reference records once the table can no longer resize?

fseventsd cleanup cursor exhaustion and notification stalls on macOS 26.7.1
 
 
Q