Cross-Border Inference Logs: A Schema That Tracks Every Inference Without Leaking Weights

Auditable logging for sovereign AI under the EU AI Act

by

Cross-Border Inference Logs: A Schema That Tracks Every Inference Without Leaking Weights

When you run a model in one jurisdiction and serve inference to another, the EU AI Act demands an audit trail. But logging every input and output can leak intellectual property—weights, training data, or business logic. This article presents a minimal schema that satisfies Article 12 (record-keeping) while keeping your model sovereign.

The Problem: Audit vs. Secrecy

The EU AI Act requires high-risk systems to log:

  • Timestamp and identity of the operator
  • Input data and output results
  • Any deviations or failures

If you log raw inputs and outputs, you expose the model's behavior. An adversary with enough logs can reverse-engineer weights or extract training data. On-prem deployment helps, but cross-border inference means logs cross borders too. You need a schema that proves compliance without revealing secrets.

Core Principle: Log Hashes, Not Data

Instead of storing raw inputs/outputs, store cryptographic hashes. SHA-256 is fine for non-repudiation; for privacy-sensitive fields, use HMAC with a per-deployment key. The key stays in the jurisdiction where the model runs.

The Schema

We use PostgreSQL with pgcrypto for hashing. The table inference_log lives on the same database as your model registry—never replicated across borders.

CREATE EXTENSION IF NOT EXISTS pgcrypto;

CREATE TYPE inference_status AS ENUM ('success', 'error', 'timeout', 'blocked');

CREATE TABLE inference_log (
    id              BIGSERIAL PRIMARY KEY,
    deployment_id   UUID NOT NULL REFERENCES deployment(id),
    operator_id     UUID NOT NULL REFERENCES operator(id),
    request_hash    BYTEA NOT NULL,  -- SHA-256 of canonical JSON input
    response_hash   BYTEA,           -- SHA-256 of canonical JSON output; NULL on error
    status          inference_status NOT NULL,
    duration_ms     INTEGER NOT NULL,
    input_tokens    INTEGER,
    output_tokens   INTEGER,
    model_version   TEXT NOT NULL,
    jurisdiction    TEXT NOT NULL,   -- ISO 3166-1 alpha-2 of serving location
    created_at      TIMESTAMPTZ NOT NULL DEFAULT NOW(),
    UNIQUE(request_hash, deployment_id, created_at)
);

-- Index for audit queries
CREATE INDEX idx_inference_log_created_at ON inference_log (created_at);
CREATE INDEX idx_inference_log_operator ON inference_log (operator_id, created_at);

Why Hashes?

  • Non-repudiation: The hash proves a specific input was processed at a given time. You can later verify by re-hashing the original input (if you still have it).
  • No weight leakage: The hash reveals nothing about the model's internal state. Even if logs are leaked, an attacker cannot reconstruct weights.
  • GDPR-friendly: Hashed data is pseudonymous; you can delete the original input on request while keeping the log for audit.

Canonical JSON

To ensure consistent hashing, define a canonical form: sort keys alphabetically, use JSON without whitespace, and specify number precision. Example in Go:

import (
    "crypto/sha256"
    "encoding/json"
    "sort"
)

func canonicalHash(v interface{}) []byte {
    b, _ := json.Marshal(v) // Go's json.Marshal produces canonical keys
    h := sha256.Sum256(b)
    return h[:]
}

Cross-Border Considerations

When inference crosses borders (e.g., model in Germany, user in France), the log stays in Germany. The audit authority in France can request a log extract, but they receive only hashes and metadata. To verify, the operator must provide the original input (which they may have stored separately under their own retention policy).

Jurisdiction Field

Always log the ISO code of the server that ran the inference. This satisfies Article 3(1) of the EU AI Act regarding geographic scope. If the model moves, the jurisdiction changes.

Querying for Audit

An auditor asks: "Show me all inferences by operator X between 2024-01-01 and 2024-01-31."

SELECT id, request_hash, response_hash, status, duration_ms, input_tokens, output_tokens, model_version, jurisdiction, created_at
FROM inference_log
WHERE operator_id = 'uuid-of-X'
  AND created_at >= '2024-01-01'::timestamptz
  AND created_at < '2024-02-01'::timestamptz
ORDER BY created_at;

The auditor gets hashes. To verify integrity, they ask the operator to provide the original input for a specific hash. The operator can prove they had that input by re-hashing and matching.

Handling Errors and Blocked Requests

When inference fails (timeout, content filter, model error), log the error status and a hash of the error message (not the raw error). This prevents leaking internal stack traces.

-- Example: blocked request due to content policy
INSERT INTO inference_log (deployment_id, operator_id, request_hash, status, duration_ms, model_version, jurisdiction)
VALUES ('d1', 'op1', digest('{"prompt":"..."}'::text, 'sha256'), 'blocked', 12, 'v1.0', 'DE');

Retention and Deletion

Logs must be retained for the life of the system plus 5 years (per Article 12). Use PostgreSQL table partitioning by month for efficient deletion. Example:

CREATE TABLE inference_log_y2024m01 PARTITION OF inference_log
    FOR VALUES FROM ('2024-01-01') TO ('2024-02-01');

When retention expires, drop the partition—not DELETE. This avoids vacuum bloat.

Integration with vLLM and llama.cpp

Both vLLM and llama.cpp support custom logging callbacks. Wire them to write to the inference_log table via a local HTTP endpoint or direct database connection.

vLLM Example (Python)

from vllm import LLM, SamplingParams
import hashlib, json, psycopg2

def log_inference(input_text, output_text, status, duration_ms, tokens):
    conn = psycopg2.connect("dbname=audit")
    cur = conn.cursor()
    request_hash = hashlib.sha256(json.dumps({"prompt": input_text}, sort_keys=True).encode()).digest()
    response_hash = hashlib.sha256(json.dumps({"output": output_text}, sort_keys=True).encode()).digest()
    cur.execute("""
        INSERT INTO inference_log (deployment_id, operator_id, request_hash, response_hash, status, duration_ms, input_tokens, output_tokens, model_version, jurisdiction)
        VALUES (%s, %s, %s, %s, %s, %s, %s, %s, %s, %s)
    """, (deployment_id, operator_id, request_hash, response_hash, status, duration_ms, tokens[0], tokens[1], model_version, jurisdiction))
    conn.commit()
    cur.close()
    conn.close()

Security Considerations

  • HMAC for sensitive fields: If you must log user IDs or other PII, use HMAC with a server-side secret. The secret never leaves the jurisdiction.
  • Separation of duties: The database that stores logs should be separate from the model serving database. Use a dedicated user with write-only access to the inference_log table.
  • Encryption at rest: Use pgcrypto or disk-level encryption for the log database.

Compliance Checklist

  1. Article 12 (Record-keeping): Logs exist with timestamps, operator identity, and hashes of inputs/outputs.
  2. Article 13 (Transparency): The hash scheme is documented and auditable.
  3. Article 14 (Human oversight): Operators can retrieve logs per user request.
  4. Article 15 (Accuracy/robustness): Error logs capture failures without leaking internals.

Conclusion

Hashing inputs and outputs is the simplest way to satisfy EU AI Act audit requirements without exposing model weights. Use PostgreSQL with pgcrypto, partition by time, and never replicate logs across borders. This schema works with vLLM, llama.cpp, and any inference engine that supports custom logging. Your model stays sovereign; your auditors get what they need.

No fluff. Just logs.

#audit#compliance#data-sovereignty#eu-ai-act#inference-logging
Share — X / Twitter · LinkedIn · HN · Email
Damir Radulić
Founder of RiNET. On the Croatian internet since 1996 (Kvarner Net). In Amsterdam now, building autonomous AI infrastructure that runs on Monday morning when nobody's watching — sovereign stacks, agent swarms, LoRA fine-tuning, civic-intelligence platforms.

Related