Durability of a directory entry after `rename()` on APFS — what is documented?

What we are building

We maintain an append-only on-disk store. Every update has the same shape:

  1. write a temporary file and make its contents durable;
  2. rename() it to its final name;
  3. only then publish the new state to the rest of the system.

Step 3 is the one we must get right. Once we publish, other components treat the preceding step as committed. If power is lost after we publish but the renamed entry was not yet durable, we would have published a state the filesystem can no longer support after reboot.

Our question is therefore narrow: after rename() returns, what does Apple document about when the resulting directory entry may be treated as durable across power loss? Two cases matter, and we would like them distinguished wherever the answer differs:

  • the target name did not exist before — an entry is added;
  • the target name already existed and is replaced — an entry is replaced and an older link is removed.

We are not asking you to review or endorse our protocol — only what the platform documents.

What we have read, and where we stopped

  • fsync(2) here causes "all modified data and attributes of fildes to be moved to a permanent storage device", then warns that after power loss an application "may find that only some or none of their data was written".
  • fcntl(2) says of F_FULLFSYNC that "data that had been fsync'd on the same device before is guaranteed to be persisted when this call returns", and lists APFS as supported.
  • rename(2) "guarantees that an instance of new will always exist, even if the system should crash in the middle of the operation".

In what we read, we did not find a statement connecting a successful synchronisation call on a descriptor that refers to a directory with the durability of a name added, removed or replaced in that directory. We are not claiming no such statement exists — only that we did not find one, which is why we are asking.

Three questions, independent of one another

Q1 — How are these calls supported on a directory descriptor?

On APFS, how are fsync(2) and fcntl(fd, F_FULLFSYNC) supported when fd refers to a directory (for example one obtained with O_DIRECTORY)? Please treat the two calls separately, since they carry different documented guarantees. "Supported", "not supported" and "not specified for this use" are all useful answers to us.

Q2 — If such a call succeeds, what does it cover?

If the call in Q1 returns successfully, which changes are guaranteed to have been completed — the directory's own attributes; the entries added, removed or replaced; related metadata? We are asking for the documented scope, not the implementation.

Q3 — Is there a documented sequence an application can rely on?

Is there a sequence of calls that Apple documents for an application that needs the name resulting from a rename() to survive power loss before it publishes its next state? If so, we would appreciate the reference, the conditions it depends on — filesystem, device, call ordering — and its stated limits. If there is no public commitment of this kind, we would rather be told so plainly, together with whatever level you are able to confirm, including "this is not currently specified" if that is the accurate answer.

One distinction we do not want to get wrong

rename(2) uses the word crash. We do not want to read a statement about crash behaviour as one about power loss, since the two can differ at the device layer. If a documented guarantee covers one and not the other, that distinction is exactly what we need.

We also note the device-level caveat in the same manual page — certain drives "have also been known to ignore the request to flush their buffered data" — and we are not asking you to speak for any particular drive.

Environment

macOS 26.3, build 25D125, APFS — read on our own development machine and reported here as such; not independently verified by a third party.

Deliberately outside this request

A rename() between two different directories touches a source directory and a target directory. Whether a guarantee would have to cover both is tracked separately and is not asked here, so that an answer about one directory is not read as covering two.

We have not fixed where the temporary file lives, and we are not asking you to choose that layout. Nor do we assume your answer is independent of it: if any of the above depends on whether the source and target directories are the same directory or two different ones, please say so explicitly. We would rather be told the answer is conditional than read an unconditional answer into it.

First, let me start here:

rename(2) "guarantees that an instance of new will always exist, even if the system should crash in the middle of the operation".

What this actually means is that, if you start with two objects ("old" and "new") and do:

rename(old,new)

then, no matter what happens, one of two things will be true:

  1. The rename "happened", so "new" now has the contents of "old" and "old" no longer exists.

  2. The rename did not "happen" and "old" and "new" have the same state they previously had.

File systems are designed around this requirement, with the "modern" approach being that the file system will either

  • Confirm that its current state is valid.

  • Discard partial change to the point that its state is once again valid.

In the context of APFS, that means the file system will always mount into a "valid" state that existed at some point in time. Putting that in concrete terms, if a directory changes with this sequence:

State 1:

file_1

State 2:

file_1
file_2 (create file_2)

State 3:

file_1
file_3 (rename file_2-> file_3)

...then the volume will remount in one of those three states. It will NOT remount as:

file_1
file_2
file_3

...as that state never existed.

Our question is therefore narrow: after rename() returns, what does Apple document about when the resulting directory entry may be treated as durable across power loss?

So, strictly speaking, we can't ACTUALLY make a strong promise about any of this. The core issue is exactly what you referenced here:

We also note the device-level caveat in the same manual page — certain drives "have also been known to ignore the request to flush their buffered data" — and we are not asking you to speak for any particular drive.

The basic issue here is that we don't ACTUALLY have any way to KNOW that the underlying hardware will flush its buffers the way it "should" have; then we wouldn't have the APIs we've ended up with. You can see how the history played out in our documentation. In the beginning, there was fsync.

In the beginning, fsync did exactly what it said it did:

"Note that while fsync() will flush all data from the host to the drive"

Great! The drive now has all the data! However... the problem is:

"...the drive itself may not physically write the data to the platters for quite some time and it may be written in an out-of-order sequence."

Originally, the combination of expense and relative performance meant that drive-level caching was fairly minimal. It did exist, but it was small enough that it didn't really matter— drive capacitors could flush it to disk and, even without that, small cache sizes also meant that write reordering didn't really happen.

...but then buses got faster, memory got a WHOLE lot faster, and drives... didn't change that much, at least not compared to the other factors. Cheap memory meant that you could make a drive feel a LOT faster by sticking a bunch of memory in front of it and completing writes as soon as they hit memory. I/O patterns tend to be bursty, so there was a good chance you'd be able to drain the cache before the next I/O burst, and you could also use write reordering to optimize your actual disk writes.

That works great... until you lose power and all that memory goes away. That problem became big enough that drive-level commands were created that said "Hey, flush your cache!” The problem is that we couldn't just blindly issue the command, as cache sizes had grown to the point where a full flush could take time. So, we introduced F_FULLFSYNC, which:

"...then asks the drive to flush all buffered data to the permanent storage device (arg is ignored). As this drains the entire queue of the device and acts as a barrier, data that had been fsync'd on the same device before is guaranteed to be persisted when this call returns."

Of course, the word "guaranteed" isn't actually true. After all, immediately after that it says:

"Certain FireWire drives have also been known to ignore the request to flush their buffered data."

Calling out "FireWire" is a historical detail. FireWire drives had particular issues with this because most of their bridge boards didn't provide a way to transfer the command to the underlying device. USB wasn't called out because, at the time, its storage adoption was still limited and its bus performance was still low enough that you couldn't send data fast enough for the cache to become an issue.

However, the problem is definitely NOT limited to FireWire and is in fact still with us today.

Summing all that up:

fsync() -> Tell the filesystem to flush its own buffer out to disk.

F_FULLFSYNC-> Tell the underlying device to flush its own cache.

If the call in Q1 returns successfully, which changes are guaranteed to have been completed — the directory's own attributes; the entries added, removed, or replaced; related metadata?

As far as the system is concerned, "everything". The problem here is that the system isn't the problem; the final hardware target is. If you're asking about writing to our own storage, then all data should be fully committed to disk once F_FULLFSYNC returns.

Is there a sequence of calls that Apple documents for an application that needs the name resulting from a rename() to survive power loss before it publishes its next state?

If you want to ensure the changes are fully committed, call F_FULLFSYNC.

__
Kevin Elliott
DTS Engineer, CoreOS/Hardware

Durability of a directory entry after `rename()` on APFS — what is documented?
 
 
Q