The Audit Log as a Legal Artifact: Designing for EU AI Act Conformity from Day Zero

Why your self-hosted AI stack needs a tamper-proof, queryable log from the first inference

by
The Audit Log as a Legal Artifact: Designing for EU AI Act Conformity from Day Zero

If you self-host LLMs or run agent swarms in production inside the EU, you are already in scope of the EU AI Act. The regulation does not care whether you are a startup or a multinational. What it cares about is traceability. And the single most important technical control you can implement today is a proper audit log.

Most teams treat audit logs as an afterthought — a syslog stream that gets rotated into a bucket and forgotten. Under the AI Act, that approach is a liability. The Act explicitly requires that high-risk AI systems maintain logs that are "comprehensive, tamper-proof, and available for inspection by national authorities."

This article walks through what that means in practice, and how to build an audit log that satisfies both the letter and the spirit of the regulation — from day zero.

What the EU AI Act Actually Says About Logs

The relevant provisions are scattered across Articles 12, 19, and 61 of the final text, plus Annex IV on technical documentation. The key requirements for audit logs are:

  • Automatic logging: The system must automatically record events during operation.
  • Traceability: Logs must allow reconstruction of the system's behavior, including inputs, outputs, and context.
  • Tamper-proofing: Logs must be protected against modification or deletion by unauthorized parties — and ideally by the operator themselves.
  • Retention: Logs must be kept for a period appropriate to the risk level — typically at least 6 months for high-risk systems, up to 5 years.
  • Accessibility: Logs must be provided to national competent authorities upon request, in a machine-readable format.

Note that these requirements apply to "high-risk" AI systems. If you are deploying an LLM for resume screening, credit scoring, medical diagnosis, or any use case listed in Annex III, you are high-risk. If you are a chatbot for internal IT support, you might be low-risk — but the regulation still expects basic logging for transparency.

Why Postgres Is Your Best Bet for Audit Logs

You already run Postgres for your application state. Use it for audit logs too. The benefits are immediate:

  • ACID compliance ensures that once a log entry is committed, it is durable.
  • Row-level security can restrict who can insert vs. who can read vs. who can delete.
  • Triggers and rules can enforce immutability at the database level.
  • pgvector can store embeddings of prompts and responses for later similarity search (useful for detecting drift or abuse).
  • Foreign Data Wrappers can push logs to cold storage or S3-compatible object stores for long-term retention.

A dedicated audit schema keeps things clean:

CREATE SCHEMA audit;

CREATE TABLE audit.log (
    id              BIGSERIAL PRIMARY KEY,
    event_time      TIMESTAMPTZ NOT NULL DEFAULT now(),
    event_type      TEXT NOT NULL,
    system_id       UUID NOT NULL,
    user_id         TEXT,
    session_id      TEXT,
    request_id      UUID NOT NULL DEFAULT gen_random_uuid(),
    input           JSONB,
    output          JSONB,
    metadata        JSONB,
    checksum        TEXT NOT NULL,
    prev_checksum   TEXT REFERENCES audit.log(checksum),
    UNIQUE (request_id)
);

CREATE INDEX idx_log_event_time ON audit.log (event_time);
CREATE INDEX idx_log_system_id ON audit.log (system_id);
CREATE INDEX idx_log_user_id ON audit.log (user_id);

The checksum column is a SHA-256 hash of the entire row content (excluding the checksum itself and the previous checksum). The prev_checksum column points to the previous row's checksum, forming a hash chain. This is the tamper-proofing mechanism.

Hash Chains: The Minimal Tamper-Proofing You Can't Skip

A hash chain works like a blockchain without the consensus overhead. Each log entry contains the hash of the previous entry. If an attacker modifies an old entry, the hash of that entry changes, breaking the chain for every subsequent entry.

Here's how to compute the checksum in a PostgreSQL trigger:

CREATE OR REPLACE FUNCTION audit.compute_checksum()
RETURNS TRIGGER AS $$
DECLARE
    row_data TEXT;
    prev_hash TEXT;
BEGIN
    -- Get the previous checksum from the last row
    SELECT checksum INTO prev_hash
    FROM audit.log
    ORDER BY id DESC
    LIMIT 1;
    
    NEW.prev_checksum := COALESCE(prev_hash, 'genesis');
    
    -- Build a canonical string of the row
    row_data := format('%s|%s|%s|%s|%s|%s|%s|%s|%s|%s',
        NEW.event_time,
        NEW.event_type,
        NEW.system_id,
        NEW.user_id,
        NEW.session_id,
        NEW.request_id,
        NEW.input::text,
        NEW.output::text,
        NEW.metadata::text,
        NEW.prev_checksum
    );
    
    NEW.checksum := encode(sha256(row_data::bytea), 'hex');
    RETURN NEW;
END;
$$ LANGUAGE plpgsql;

CREATE TRIGGER trg_audit_checksum
    BEFORE INSERT ON audit.log
    FOR EACH ROW
    EXECUTE FUNCTION audit.compute_checksum();

This trigger runs on every insert. It reads the previous row's checksum, concatenates all fields, and computes SHA-256. The result is stored in the new row.

To verify the chain integrity:

WITH RECURSIVE chain AS (
    SELECT id, checksum, prev_checksum, event_time
    FROM audit.log
    WHERE id = (SELECT max(id) FROM audit.log)
    UNION ALL
    SELECT l.id, l.checksum, l.prev_checksum, l.event_time
    FROM audit.log l
    JOIN chain c ON l.checksum = c.prev_checksum
)
SELECT * FROM chain;

If any row is missing or its checksum doesn't match, the query will return fewer rows than expected, or fail to traverse.

What to Log: Every Prompt, Every Response, Every Decision

The Act requires traceability. That means you must log:

  • Every input to the model: the full prompt, including system messages and context.
  • Every output from the model: the raw generated text, plus any post-processing.
  • Confidence scores or logprobs, if available — these are critical for audit of high-risk decisions.
  • Model version and deployment ID: so you can reproduce the exact inference later.
  • Latency and failure modes: timeouts, retries, fallback models.
  • Human-in-the-loop actions: if a human reviewed or overrode the output, log that too.

For agent swarms, log the entire chain of tool calls and responses. Each step should have its own log entry with a parent request ID.

Retention and Rotation: Don't Delete, Archive

The Act does not specify exact retention periods for all cases, but for high-risk systems, expect at least 6 months. Some national regulators may demand up to 5 years. Design for the maximum.

  • Use partitioning in Postgres to drop old data efficiently — but instead of dropping, move to a cheaper storage tier.
  • Use pg_partman to automate time-based partitioning on event_time.
  • After the retention window, move partitions to an S3-compatible object store (MinIO, Ceph) using pg_dump or a foreign data wrapper.
  • Keep a catalog of archived partitions with checksums so you can prove the archive is intact.

Example partition setup:

CREATE TABLE audit.log (
    id              BIGSERIAL,
    event_time      TIMESTAMPTZ NOT NULL,
    -- ... other columns
    PRIMARY KEY (id, event_time)
) PARTITION BY RANGE (event_time);

CREATE TABLE audit.log_2024_q1 PARTITION OF audit.log
    FOR VALUES FROM ('2024-01-01') TO ('2024-04-01');

-- Repeat for each quarter

Querying the Log: From Compliance to Debugging

A well-structured audit log is not just for regulators. It is your best debugging tool. When a model behaves unexpectedly, the log is your time machine.

  • Find all requests by a specific user in the last 24 hours.
  • Find all requests that triggered a specific guardrail (e.g., toxicity filter).
  • Find all requests that used a deprecated model version.
  • Find all requests where the output was overridden by a human.

Because the log is in Postgres, you can use standard SQL. No need for a separate log aggregation tool. You can also join with your application tables to get full context.

Integrating with Inference Servers

Your inference server must emit structured logs. Many inference engines support JSON logging. For example, you can enable JSON output and pipe to a custom logger that inserts into Postgres.

A sidecar process (e.g., a Python script with asyncpg) can consume the log stream and batch-insert into the audit table. A systemd service can tail the log file and insert periodically.

Access Control: Who Can Read the Log?

Not everyone should be able to read the log. The Act requires that logs be accessible to "competent authorities" — not to all employees.

  • Grant INSERT to the service account that writes logs.
  • Grant SELECT only to specific roles: auditor, compliance_officer, regulator.
  • Use row-level security to restrict access to logs older than N days for non-auditors.
  • Never allow DELETE or UPDATE on audit tables. Enforce this with a trigger that rejects any non-INSERT operation.
CREATE OR REPLACE FUNCTION audit.prevent_modification()
RETURNS TRIGGER AS $$
BEGIN
    RAISE EXCEPTION 'audit.log is immutable';
END;
$$ LANGUAGE plpgsql;

CREATE TRIGGER trg_prevent_update
    BEFORE UPDATE ON audit.log
    FOR EACH ROW EXECUTE FUNCTION audit.prevent_modification();

CREATE TRIGGER trg_prevent_delete
    BEFORE DELETE ON audit.log
    FOR EACH ROW EXECUTE FUNCTION audit.prevent_modification();

Real-World Example: A Self-Hosted Resume Screening System

Consider a company that uses a fine-tuned LLM to screen job applications. This is a high-risk use case per Annex III.

Day zero setup:

  1. Deploy the inference server with the custom model.
  2. The application sends prompts via REST API. Each request gets a UUID.
  3. The application inserts into audit.log the prompt, the model version, the user ID (recruiter), and the candidate ID.
  4. The inference server logs the raw output and logprobs, which are also inserted.
  5. A human reviewer later marks a candidate as "pass" or "fail". That decision is logged with the same request ID.
  6. Months later, a regulator asks for all logs related to a specific candidate. The company runs a SQL query and exports the chain of entries as JSON.
  7. The regulator verifies the hash chain to ensure no tampering.

Without this setup, the company would have to scramble to reconstruct what happened — and may fail the audit.

Cost and Performance Considerations

Logging everything is expensive. Each prompt and response can be thousands of tokens. Storing them in JSONB in Postgres will consume disk space.

  • Estimate storage needs based on expected request volume and average log entry size. Postgres can handle large volumes with proper indexing and partitioning.
  • Use compression: set ALTER TABLE audit.log SET (toast_tuple_target = 0); to allow TOAST compression on large JSONB fields.
  • Archive old partitions to cheaper storage (MinIO) after the active retention period.
  • Consider using a separate Postgres instance for audit logs if your main database is I/O constrained. But for most self-hosted setups, a single instance with proper indexing is fine.

The Bottom Line

The EU AI Act is not coming — it is here. The first enforcement dates hit in 2025. If you are building AI systems today, you must design your audit log as a legal artifact from day zero.

Postgres gives you the tools to do it right: ACID transactions, hash chains for tamper-proofing, partitioning for retention, and row-level security for access control. There is no excuse to bolt on a flimsy log later.

Start with a schema, a trigger, and a sidecar ingester. That is 50 lines of code and a systemd unit. Your future self — and your regulator — will thank you.

#audit#compliance#data-sovereignty#eu-ai-act#self-hosted-ai
Share — X / Twitter · LinkedIn · HN · Email
Damir Radulić
Founder of RiNET. On the Croatian internet since 1996 (Kvarner Net). In Amsterdam now, building autonomous AI infrastructure that runs on Monday morning when nobody's watching — sovereign stacks, agent swarms, LoRA fine-tuning, civic-intelligence platforms.

Related