Atomic Audit Writes on NVMe Without fsync Every Turn
Engineering around power loss for audit logs on ZFS-backed NVMe
Audit logs are the one component where losing a write is not an option. If the power dies mid-turn, you need to know exactly what happened up to that instant. The naive answer is fsync() after every append, but on a busy agent swarm that turns into a synchronous bottleneck that caps your throughput at whatever your storage can flush per millisecond.
This article is about a different approach: structuring your audit log so that a sudden power loss cannot corrupt or lose entries, even when you don't call fsync on every turn. The techniques here are generic, but I'll ground them in the specific stack I run: three Hetzner servers, a WireGuard mesh, ZFS on NVMe, and a Postgres database that doubles as the audit store.
Why fsync-per-turn is the wrong default
fsync() forces the operating system to flush dirty pages to the storage device and wait for the device to acknowledge that the data is physically on non-volatile media. That sounds like exactly what you want for an audit log, but it has a hidden cost: it converts every append into a round-trip to the disk. On a typical NVMe drive, that's on the order of tens of microseconds, but the real cost is that fsync forces the drive to flush its internal write cache, which often involves a full cache flush command. That can take hundreds of microseconds, and if you're doing it per turn, you're throwing away most of the drive's potential.
Moreover, in a distributed system, fsync on one node doesn't protect you against a network partition or a crash on another node. The audit log is only as durable as the replication strategy you've built around it. So fsync per turn is both too expensive and not sufficient on its own.
The goal is to make the common path (each turn) not require fsync, while still guaranteeing that if power is lost, the log is consistent and every acknowledged entry is recoverable.
What exactly breaks without fsync?
When you write to a file, the data goes into the page cache in RAM. The kernel decides when to write those dirty pages to disk. If power is lost before that happens, the data is gone. Even if the data has been written to the disk, the filesystem metadata (like the file size or the directory entry) might not be updated, leading to a file that is shorter than expected or contains garbage at the end.
For an audit log that is append-only, the typical failure modes are:
- Lost writes: An entry that was acknowledged to the client never makes it to disk.
- Partial writes: A torn write where only part of an entry is on disk, leaving a corrupt record.
- Reordering: Entries appear in a different order than they were written, which is disastrous for an audit trail.
To avoid fsync per turn, you need to design the log so that these failures are either impossible or detectable and recoverable.
The append-only file with a sidecar index
The first step is to separate the audit data from the metadata. Instead of appending to a single file and relying on the filesystem to update the file size atomically, you can write each entry as a self-describing record that includes a length prefix and a checksum. Then you maintain a separate index file that records the offset and length of each entry, but you don't need to fsync that index on every append.
Why does this help? Because the index can be rebuilt from the data file if it's lost or corrupted. The data file itself is append-only, and you can detect torn writes by checking the length prefix against the actual bytes and verifying the checksum.
Here's a concrete example in Python, but the principle applies to any language:
import struct
import hashlib
import os
class AuditLog:
def __init__(self, path):
self.data_path = path + ".data"
self.index_path = path + ".index"
self.data_fh = open(self.data_path, "ab")
self.index_fh = open(self.index_path, "ab")
def append(self, entry: bytes):
# Compute a simple checksum (e.g., SHA-256)
checksum = hashlib.sha256(entry).digest()
# Build a record: [length][checksum][entry]
length = len(entry)
record = struct.pack("!I", length) + checksum + entry
# Append to data file
self.data_fh.write(record)
# Flush to OS (not fsync)
self.data_fh.flush()
# Append to index file: [offset][length]
offset = self.data_fh.tell() - len(record)
index_entry = struct.pack("!QQ", offset, len(record))
self.index_fh.write(index_entry)
self.index_fh.flush()
# No fsync here
def recover(self):
# Rebuild index from data file if needed
self.data_fh.seek(0)
index = []
while True:
header = self.data_fh.read(4)
if len(header) < 4:
break
length = struct.unpack("!I", header)[0]
# Sanity check length (avoid garbage)
if length > 10_000_000: # adjust as needed
break
checksum = self.data_fh.read(32)
entry = self.data_fh.read(length)
if len(entry) < length:
# Torn write, stop
break
if hashlib.sha256(entry).digest() != checksum:
# Corrupt, stop
break
index.append((self.data_fh.tell() - length - 36, length + 36))
# Continue
# Write a fresh index file
with open(self.index_path, "wb") as f:
for offset, length in index:
f.write(struct.pack("!QQ", offset, length))
return indexIn this design, fsync is only called during recover() or at shutdown. On normal operation, we call flush() which pushes data from the application's user-space buffers to the kernel, but does not force it to disk. If power is lost, the data file may have a truncated last entry, but recover() will detect that and rebuild the index, discarding the partial entry.
ZFS: making the filesystem your ally
If you're running ZFS (which I do on the Hetzner boxes), you get some additional guarantees. ZFS is a copy-on-write filesystem that uses checksums on every block. When you write a file, the data is written to new blocks, and only after the write completes does the metadata point to the new blocks. This means that a crash cannot leave a file in a state where it contains a mix of old and new data; it's either the old version or the new version.
But ZFS doesn't automatically make fsync unnecessary. By default, ZFS still buffers writes in the transaction group (txg) and flushes them every few seconds. If power is lost, the last few seconds of writes may be lost. To avoid that, you'd need to force a txg sync, which is essentially what fsync does.
However, you can configure ZFS to be more aggressive about syncing, but that defeats the purpose. Instead, you can use ZFS's built-in snapshotting and replication to provide durability at a different level: even if you lose a few seconds of data, you can recover from a snapshot that was taken just before the crash. But for an audit log, losing a few seconds might be unacceptable.
So the real answer is to combine the append-only design with a periodic fsync at a coarser granularity. For example, you might fsync every 1000 entries or every second, whichever comes first. That gives you a bounded window of potential loss, which you can tune based on your requirements.
The power-loss-safe ordering trick
One subtle issue is ordering. If you write entry A, then entry B, and power is lost, you might end up with B on disk but not A. That would be a reordering, which is fatal for an audit log. How can you prevent that without fsync?
The trick is to use a monotonically increasing sequence number that is part of the entry itself. When reading the log, you can detect gaps and reordering by checking the sequence numbers. If you see entry 5 before entry 4, you know something is wrong and you can trigger a recovery.
But that only detects the problem after the fact. To prevent it, you need to ensure that the write of entry A is visible to the filesystem before entry B is written. That's hard to guarantee without fsync because the kernel can reorder writes to the same file.
One approach is to write each entry to a separate file, and then use a directory rename to make a batch of entries visible atomically. For example, you accumulate entries in a temporary file, and periodically (say every second) you rename it to a new name that includes the sequence range. The rename operation is atomic on POSIX filesystems, and ZFS respects that. If power is lost, you either have the old file or the new file, but never a partial batch.
This is a common pattern in log-structured storage systems. It gives you atomicity at the granularity of a batch, not per entry, but that's often acceptable. You can configure the batch size to match your tolerance for loss.
Alternative: use a database with WAL
If you're already using Postgres (as I do for other data), you can leverage its write-ahead log (WAL) to get durability without fsync on every turn. Postgres's WAL is designed to be crash-safe: every transaction is logged to the WAL before it is committed, and the WAL is flushed to disk at a configurable interval (the synchronous_commit setting).
If you set synchronous_commit = off, Postgres will acknowledge commits without waiting for the WAL to be flushed, but it guarantees that the WAL will be flushed within a few milliseconds. That gives you a bounded window of loss, similar to the periodic fsync approach, but it's handled by the database engine, which is well-tested.
For an audit log, you could insert each entry as a row in a table with a sequence number and a timestamp. Postgres guarantees that committed transactions are not lost if you use synchronous_commit = on, but that's effectively fsync per turn. With off, you get higher throughput at the risk of losing the last few milliseconds of commits.
However, there is a subtlety: if you acknowledge a turn to the client before the WAL is flushed, and then the power fails, the client may believe the entry is durable when it's not. To handle that, you need to design your protocol so that the client can query for the last durable sequence number, which you can get by running SELECT pg_current_wal_lsn() after a flush.
What I actually do in practice
On my three Hetzner boxes, I run a lightweight audit service that writes to a local file using the append-only format described above. The files are stored on ZFS, and I take snapshots every hour and replicate them to another node using zfs send. I also run a nightly job that replays the log into Postgres for querying, using the sequence numbers to deduplicate.
For the agent swarm, each turn produces an audit record that includes the agent ID, the action, the input hash, and the output hash. The record is small (a few hundred bytes), so I can afford to batch them. I flush to the OS on every append, but only fsync every 1000 entries or when the process is shutting down. I also make sure to fsync before sending an acknowledgment to the client if the turn is marked as critical.
I've tested this by pulling the power plug on a test node (not in production, mind you) and verifying that the log is recoverable. In my tests, the worst case is losing the last few entries that were in the page cache, but the log remains consistent and the index can be rebuilt.
Tradeoffs and tuning knobs
The main tradeoff is between throughput and the maximum window of data loss. If you can tolerate losing, say, 1 second of audit data, you can set a timer that calls fsync every second. If you need tighter guarantees, you can reduce that interval, but you'll pay a throughput cost.
Here are some knobs to consider:
- Batch size: How many entries accumulate before you force a sync? Larger batches mean fewer syncs but longer windows.
- Sync interval: A timer-based approach can smooth out bursts.
- Checksum strength: A simple CRC32 might be enough, but if you're paranoid, use SHA-256. The cost is only on recovery, not on the write path.
- Replication: If you have multiple nodes, you can replicate the log in real-time using something like
tail -fandnc, but that introduces network latency.
Recovery procedure
When a node comes back after a power loss, the first thing to do is run the recovery function on the audit log. This scans the data file, validates checksums, and rebuilds the index. Any torn or corrupt entries at the end are discarded. The recovered log is then compared with the replicated copy on another node to reconcile any gaps.
In my recovery script, I also recompute the sequence numbers and check for monotonicity. If I find a gap, I know that some entries were lost, and I can either accept that or try to reconstruct them from other sources (like the agent's own logs).
Conclusion
Avoiding fsync on every turn is not about being sloppy with durability; it's about being smart about where you spend your synchronous I/O. By using an append-only format with checksums, a sidecar index that can be rebuilt, and ZFS's copy-on-write semantics, you can make power loss a recoverable event rather than a catastrophe.
The key insight is to separate the data from the metadata and to make the data self-describing. That way, even if the filesystem metadata is stale, you can reconstruct the log from the raw bytes. And by batching your syncs, you can tune the tradeoff between performance and the window of potential loss.
If you're building an audit log for an autonomous system, I encourage you to think about these patterns before you reach for fsync on every turn. Your throughput will thank you, and your auditors will still be able to trust the log.