Jurisdictional Inference Routing: Complying with the EU AI Act at the Proxy Level
The EU AI Act introduces a tiered compliance framework based on risk. For high-risk AI systems, data sovereignty and inference locality become legal requirements, not just architectural preferences. If your inference pipeline serves users across multiple jurisdictions, you need a mechanism to ensure that requests originating from the EU are processed within EU boundaries, using models and data that adhere to GDPR and the Act's transparency obligations.
A practical solution is to implement jurisdictional inference routing at the proxy level. This means inserting a lightweight decision layer — typically a reverse proxy or API gateway — that inspects each inference request, determines the user's jurisdiction, and routes the request to an appropriate inference endpoint that complies with local regulations.
Architecture Overview
The proxy sits in front of your model serving infrastructure. It receives all inference requests, extracts metadata, and applies routing rules before forwarding. The key components are:
- Geo-location extraction: Use a GeoIP database (e.g., MaxMind GeoLite2) on the proxy to map the requester's IP address to a country. For authenticated users, you can also use account-level jurisdiction metadata (e.g., billing address, declared residence).
- Request classification: Inspect the payload for data sensitivity. For example, if the input contains PII (names, emails, IDs), the request must be treated as high-risk and routed accordingly.
- Routing table: A configuration that maps jurisdiction + risk level to a specific inference endpoint. The table can be stored in a database or a YAML file, reloaded without restarting the proxy.
- Audit logging: Every routing decision is logged with timestamps, source IP (or user ID), target endpoint, and a hash of the request payload (for non-repudiation).
Proxy Implementation (Python + FastAPI Example)
Below is a simplified implementation using FastAPI as the proxy layer. It assumes you have two inference backends: one inside the EU (eu-inference.example.com) and one outside (global-inference.example.com).
from fastapi import FastAPI, Request
from fastapi.responses import JSONResponse
import httpx
import geoip2.database
import hashlib
import json
app = FastAPI()
# GeoIP reader (loaded at startup)
reader = geoip2.database.Reader('./GeoLite2-Country.mmdb')
# Routing table (could be loaded from config)
ROUTES = {
"EU": {
"high_risk": "https://eu-inference.example.com/v1/chat/completions",
"low_risk": "https://eu-inference.example.com/v1/chat/completions"
},
"default": {
"high_risk": "https://global-inference.example.com/v1/chat/completions",
"low_risk": "https://global-inference.example.com/v1/chat/completions"
}
}
# Simple PII detection (very basic — use a proper library in production)
def contains_pii(text: str) -> bool:
patterns = [r'\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Z|a-z]{2,}\b', # email
r'\b\d{3}-\d{2}-\d{4}\b', # SSN
r'\b\d{16}\b'] # credit card
import re
for pat in patterns:
if re.search(pat, text):
return True
return False
@app.post("/v1/chat/completions")
async def proxy_inference(request: Request):
body = await request.json()
user_ip = request.client.host
# 1. Determine jurisdiction
try:
response = reader.country(user_ip)
country = response.country.iso_code
except:
country = "UNKNOWN"
# EU countries list (simplified)
eu_countries = {"AT", "BE", "BG", "HR", "CY", "CZ", "DK", "EE", "FI", "FR", "DE", "GR", "HU", "IE", "IT", "LV", "LT", "LU", "MT", "NL", "PL", "PT", "RO", "SK", "SI", "ES", "SE"}
jurisdiction = "EU" if country in eu_countries else "default"
# 2. Classify risk based on PII in input
messages = body.get("messages", [])
full_text = " ".join([m.get("content", "") for m in messages])
is_high_risk = contains_pii(full_text)
risk_level = "high_risk" if is_high_risk else "low_risk"
# 3. Select endpoint
endpoint = ROUTES[jurisdiction][risk_level]
# 4. Audit log
payload_hash = hashlib.sha256(json.dumps(body, sort_keys=True).encode()).hexdigest()
# In production, write to a structured log or database
print(f"AUDIT: ip={user_ip} jurisdiction={jurisdiction} risk={risk_level} target={endpoint} hash={payload_hash}")
# 5. Forward request (using httpx async client)
async with httpx.AsyncClient() as client:
resp = await client.post(endpoint, json=body, timeout=30.0)
return JSONResponse(content=resp.json(), status_code=resp.status_code)Handling Model Compliance
Even if you route requests correctly, the model itself must comply with the EU AI Act. For high-risk applications, you need:
- Model cards: Each model version should have a documented card listing training data sources, intended use, limitations, and bias evaluations. The proxy can append a model version header to the response for traceability.
- Explainability: If the model is used for decisions affecting individuals (e.g., credit scoring, hiring), you must provide explanations. This is hard to retrofit; prefer using inherently interpretable models or implementing post-hoc explanation methods like LIME or SHAP.
- Human oversight: For high-risk systems, the Act requires human review. You can implement a fallback where the proxy flags certain requests for manual approval before inference.
Audit Trail Requirements
The EU AI Act mandates record-keeping for high-risk AI systems. Your proxy should log:
- Timestamp of request
- User identifier (pseudonymized if possible)
- Input data hash (not the raw data, to avoid storing PII)
- Model version and endpoint used
- Output hash (optional, for reproducibility)
- Routing decision rationale (jurisdiction, risk level)
- Any errors or timeouts
Store these logs in an append-only store (e.g., a blockchain-based ledger or a WORM storage) to prevent tampering. Retention period should align with the Act's requirements (typically 5 years after the last use).
Geographic Load Balancing and Failover
If your EU inference endpoint goes down, you might be tempted to fall back to a non-EU endpoint. However, for EU-originating high-risk requests, this would violate data sovereignty. Instead, implement a circuit breaker that fails open only for low-risk requests, or queuing mechanism that holds requests until the EU endpoint recovers.
Caveats and Real-World Considerations
- VPNs and IP spoofing: GeoIP is not foolproof. Users behind VPNs may appear to be in a different jurisdiction. For authenticated users, prefer account-level jurisdiction data. For unauthenticated, you can accept the GeoIP location and log the discrepancy.
- Performance overhead: The proxy adds latency. A common pattern is that the GeoIP lookup and routing decision take under 5ms, but the PII scanning can be heavier. Use a fast regex engine or a dedicated PII detection service.
- Regulatory changes: The EU AI Act is still evolving. Your routing rules should be configurable without code changes. Consider using a rules engine (e.g., Drools, Open Policy Agent) to manage complex conditions.
Conclusion
Jurisdictional inference routing at the proxy level is a pragmatic first step toward EU AI Act compliance. It decouples regulatory logic from model serving, allowing you to enforce data sovereignty without modifying your inference code. Combined with model cards, audit trails, and human oversight, it forms a solid foundation for a compliant AI infrastructure.
Remember: compliance is not a one-time checkbox. Your proxy routing must be continuously updated as regulations and model deployments change. Treat it as a critical piece of infrastructure, with version-controlled configurations and automated testing.