Cross-Border Inference Logs: A Schema That Tracks Every Inference Without Leaking Weights
Auditable logging for sovereign AI under the EU AI Act
Cross-Border Inference Logs: A Schema That Tracks Every Inference Without Leaking Weights
When you run a model in one jurisdiction and serve inference to another, the EU AI Act demands an audit trail. But logging every input and output can leak intellectual property—weights, training data, or business logic. This article presents a minimal schema that satisfies Article 12 (record-keeping) while keeping your model sovereign.
The Problem: Audit vs. Secrecy
The EU AI Act requires high-risk systems to log:
- Timestamp and identity of the operator
- Input data and output results
- Any deviations or failures
If you log raw inputs and outputs, you expose the model's behavior. An adversary with enough logs can reverse-engineer weights or extract training data. On-prem deployment helps, but cross-border inference means logs cross borders too. You need a schema that proves compliance without revealing secrets.
Core Principle: Log Hashes, Not Data
Instead of storing raw inputs/outputs, store cryptographic hashes. SHA-256 is fine for non-repudiation; for privacy-sensitive fields, use HMAC with a per-deployment key. The key stays in the jurisdiction where the model runs.
The Schema
We use PostgreSQL with pgcrypto for hashing. The table inference_log lives on the same database as your model registry—never replicated across borders.
CREATE EXTENSION IF NOT EXISTS pgcrypto;
CREATE TYPE inference_status AS ENUM ('success', 'error', 'timeout', 'blocked');
CREATE TABLE inference_log (
id BIGSERIAL PRIMARY KEY,
deployment_id UUID NOT NULL REFERENCES deployment(id),
operator_id UUID NOT NULL REFERENCES operator(id),
request_hash BYTEA NOT NULL, -- SHA-256 of canonical JSON input
response_hash BYTEA, -- SHA-256 of canonical JSON output; NULL on error
status inference_status NOT NULL,
duration_ms INTEGER NOT NULL,
input_tokens INTEGER,
output_tokens INTEGER,
model_version TEXT NOT NULL,
jurisdiction TEXT NOT NULL, -- ISO 3166-1 alpha-2 of serving location
created_at TIMESTAMPTZ NOT NULL DEFAULT NOW(),
UNIQUE(request_hash, deployment_id, created_at)
);
-- Index for audit queries
CREATE INDEX idx_inference_log_created_at ON inference_log (created_at);
CREATE INDEX idx_inference_log_operator ON inference_log (operator_id, created_at);Why Hashes?
- Non-repudiation: The hash proves a specific input was processed at a given time. You can later verify by re-hashing the original input (if you still have it).
- No weight leakage: The hash reveals nothing about the model's internal state. Even if logs are leaked, an attacker cannot reconstruct weights.
- GDPR-friendly: Hashed data is pseudonymous; you can delete the original input on request while keeping the log for audit.
Canonical JSON
To ensure consistent hashing, define a canonical form: sort keys alphabetically, use JSON without whitespace, and specify number precision. Example in Go:
import (
"crypto/sha256"
"encoding/json"
"sort"
)
func canonicalHash(v interface{}) []byte {
b, _ := json.Marshal(v) // Go's json.Marshal produces canonical keys
h := sha256.Sum256(b)
return h[:]
}Cross-Border Considerations
When inference crosses borders (e.g., model in Germany, user in France), the log stays in Germany. The audit authority in France can request a log extract, but they receive only hashes and metadata. To verify, the operator must provide the original input (which they may have stored separately under their own retention policy).
Jurisdiction Field
Always log the ISO code of the server that ran the inference. This satisfies Article 3(1) of the EU AI Act regarding geographic scope. If the model moves, the jurisdiction changes.
Querying for Audit
An auditor asks: "Show me all inferences by operator X between 2024-01-01 and 2024-01-31."
SELECT id, request_hash, response_hash, status, duration_ms, input_tokens, output_tokens, model_version, jurisdiction, created_at
FROM inference_log
WHERE operator_id = 'uuid-of-X'
AND created_at >= '2024-01-01'::timestamptz
AND created_at < '2024-02-01'::timestamptz
ORDER BY created_at;The auditor gets hashes. To verify integrity, they ask the operator to provide the original input for a specific hash. The operator can prove they had that input by re-hashing and matching.
Handling Errors and Blocked Requests
When inference fails (timeout, content filter, model error), log the error status and a hash of the error message (not the raw error). This prevents leaking internal stack traces.
-- Example: blocked request due to content policy
INSERT INTO inference_log (deployment_id, operator_id, request_hash, status, duration_ms, model_version, jurisdiction)
VALUES ('d1', 'op1', digest('{"prompt":"..."}'::text, 'sha256'), 'blocked', 12, 'v1.0', 'DE');Retention and Deletion
Logs must be retained for the life of the system plus 5 years (per Article 12). Use PostgreSQL table partitioning by month for efficient deletion. Example:
CREATE TABLE inference_log_y2024m01 PARTITION OF inference_log
FOR VALUES FROM ('2024-01-01') TO ('2024-02-01');When retention expires, drop the partition—not DELETE. This avoids vacuum bloat.
Integration with vLLM and llama.cpp
Both vLLM and llama.cpp support custom logging callbacks. Wire them to write to the inference_log table via a local HTTP endpoint or direct database connection.
vLLM Example (Python)
from vllm import LLM, SamplingParams
import hashlib, json, psycopg2
def log_inference(input_text, output_text, status, duration_ms, tokens):
conn = psycopg2.connect("dbname=audit")
cur = conn.cursor()
request_hash = hashlib.sha256(json.dumps({"prompt": input_text}, sort_keys=True).encode()).digest()
response_hash = hashlib.sha256(json.dumps({"output": output_text}, sort_keys=True).encode()).digest()
cur.execute("""
INSERT INTO inference_log (deployment_id, operator_id, request_hash, response_hash, status, duration_ms, input_tokens, output_tokens, model_version, jurisdiction)
VALUES (%s, %s, %s, %s, %s, %s, %s, %s, %s, %s)
""", (deployment_id, operator_id, request_hash, response_hash, status, duration_ms, tokens[0], tokens[1], model_version, jurisdiction))
conn.commit()
cur.close()
conn.close()Security Considerations
- HMAC for sensitive fields: If you must log user IDs or other PII, use HMAC with a server-side secret. The secret never leaves the jurisdiction.
- Separation of duties: The database that stores logs should be separate from the model serving database. Use a dedicated user with write-only access to the inference_log table.
- Encryption at rest: Use pgcrypto or disk-level encryption for the log database.
Compliance Checklist
- Article 12 (Record-keeping): Logs exist with timestamps, operator identity, and hashes of inputs/outputs.
- Article 13 (Transparency): The hash scheme is documented and auditable.
- Article 14 (Human oversight): Operators can retrieve logs per user request.
- Article 15 (Accuracy/robustness): Error logs capture failures without leaking internals.
Conclusion
Hashing inputs and outputs is the simplest way to satisfy EU AI Act audit requirements without exposing model weights. Use PostgreSQL with pgcrypto, partition by time, and never replicate logs across borders. This schema works with vLLM, llama.cpp, and any inference engine that supports custom logging. Your model stays sovereign; your auditors get what they need.
No fluff. Just logs.