Local NVMe Audit Logs: Surviving Power Loss for EU AI Act Inference

A durable, zero-dependency pattern for on-prem AI audit trails

by

Local NVMe Audit Logs: Surviving Power Loss for EU AI Act Inference

The EU AI Act mandates audit trails for high-risk AI systems. If you're running self-hosted LLM inference on bare metal or VPS, you need logs that survive a power loss. Cloud object storage is off the table for sovereignty reasons. This article shows a concrete pattern using local NVMe drives, PostgreSQL, and synchronous writes — no black boxes, no magic.

Why NVMe?

NVMe drives with power-loss protection (PLP) capacitors are commodity hardware now. Consumer NVMe drives often lack PLP, but enterprise-class drives (e.g., Samsung PM9A3, Kioxia CD6) include capacitors that flush the write cache during power loss. Pair that with a filesystem like ext4 or XFS in ordered mode (default), and you get crash-consistent writes. For audit logs, that's sufficient: you won't lose committed records.

The Pattern

We'll log each inference request and response to a local PostgreSQL instance with synchronous_commit = on, writing to a dedicated audit table. The database itself lives on the NVMe array. We also write a raw JSON log to a flat file on the same NVMe, as a secondary, immutable copy. Both paths use fsync.

PostgreSQL Audit Table

CREATE TABLE inference_audit (
    id BIGSERIAL PRIMARY KEY,
    request_id UUID NOT NULL UNIQUE,
    model TEXT NOT NULL,
    prompt_hash BYTEA NOT NULL,  -- SHA-256 of prompt
    response_hash BYTEA NOT NULL, -- SHA-256 of output
    prompt_tokens INTEGER,
    completion_tokens INTEGER,
    created_at TIMESTAMPTZ NOT NULL DEFAULT NOW(),
    metadata JSONB
);

-- Index for time-range queries
CREATE INDEX idx_audit_created_at ON inference_audit (created_at);

Set synchronous_commit = on in postgresql.conf. Also set wal_sync_method = fdatasync (default on Linux). This ensures each transaction is flushed to disk before returning.

Flat File Log

For an extra layer, log each event as a newline-delimited JSON line to a file. Use O_SYNC when opening the file, or call fsync after each write. In Python:

import json, os, fcntl

log_fd = os.open('/var/log/inference/audit.log', os.O_WRONLY | os.O_CREAT | os.O_APPEND | os.O_SYNC, 0o640)

def log_audit(event: dict):
    line = json.dumps(event, separators=(',', ':')) + '\n'
    os.write(log_fd, line.encode())
    os.fsync(log_fd)  # redundant with O_SYNC, but explicit

O_SYNC forces each write to complete only when data is on non-volatile storage. Combined with NVMe PLP, a power loss at any point leaves the last write either fully recorded or not at all — no partial lines.

Verification: Crash Test

We tested this on a Dell R750 with Samsung PM9A3 NVMe (PLP-capable). The workload: vLLM serving Llama 3.1 70B, logging each request to both PostgreSQL and flat file. We pulled the power cord mid-request. After reboot:

  • PostgreSQL recovered cleanly via WAL replay. The last committed transaction was present; partial writes were rolled back.
  • The flat file ended with a complete JSON line (the one that was being written before power loss). No corruption, no partial line.

This works because O_SYNC + PLP ensures the write cache is flushed. Without PLP, consumer NVMe drives may lose the last few milliseconds of data even with O_SYNC (the drive may report completion before data is in NAND). So: use enterprise NVMe.

Rotation and Retention

For flat files, use logrotate with copytruncate and delaycompress:

/var/log/inference/audit.log {
    daily
    rotate 90
    compress
    delaycompress
    copytruncate
    postrotate
        # reopen file descriptor in your app (SIGUSR1 or similar)
    endscript
}

For PostgreSQL, partition the audit table by time (e.g., monthly) and drop old partitions. This avoids bloat and keeps query performance predictable.

EU AI Act Compliance Notes

The EU AI Act requires that audit logs be kept for at least 6 months for high-risk systems (Article 12). This pattern gives you:

  • Integrity: Cryptographic hashes of prompts and responses (store the hash, not plaintext, unless you need to retain the actual data — check your use case).
  • Availability: Local storage with RAID1 (mirroring) on NVMe adds resilience.
  • Confidentiality: Encrypt the flat file at rest (LUKS) and use PostgreSQL TDE or pg_tde for the database.

If you need to prove that logs haven't been tampered with, sign each log line with a key stored in a TPM or HSM. That's a separate article.

Alternatives and Trade-offs

  • WAL-based logging: You could parse PostgreSQL's WAL for audit events, but that's complex and fragile.
  • Kafka with fsync: Overkill for a single node.
  • SQLite with WAL mode: Good for single-process apps, but not for concurrent vLLM workers.

PostgreSQL + flat file is simple, auditable, and survives power loss.

Conclusion

Self-hosted inference doesn't mean sacrificing audit durability. With enterprise NVMe, synchronous writes, and a bit of code, you get logs that survive a power loss without cloud dependencies. No marketing, no fluff — just a pattern that works.

Tested with: vLLM 0.6.3, PostgreSQL 16, Linux 6.8, ext4, Samsung PM9A3.

#audit#eu-ai-act#nvme#self-hosted#storage
Share — X / Twitter · LinkedIn · HN · Email
Damir Radulić
Founder of RiNET. On the Croatian internet since 1996 (Kvarner Net). In Amsterdam now, building autonomous AI infrastructure that runs on Monday morning when nobody's watching — sovereign stacks, agent swarms, LoRA fine-tuning, civic-intelligence platforms.

Related