Jurisdictional Inference Routing: Complying with the EU AI Act at the Proxy Level

by

The EU AI Act introduces a tiered compliance framework based on risk. For high-risk AI systems, data sovereignty and inference locality become legal requirements, not just architectural preferences. If your inference pipeline serves users across multiple jurisdictions, you need a mechanism to ensure that requests originating from the EU are processed within EU boundaries, using models and data that adhere to GDPR and the Act's transparency obligations.

A practical solution is to implement jurisdictional inference routing at the proxy level. This means inserting a lightweight decision layer — typically a reverse proxy or API gateway — that inspects each inference request, determines the user's jurisdiction, and routes the request to an appropriate inference endpoint that complies with local regulations.

Architecture Overview

The proxy sits in front of your model serving infrastructure. It receives all inference requests, extracts metadata, and applies routing rules before forwarding. The key components are:

  • Geo-location extraction: Use a GeoIP database (e.g., MaxMind GeoLite2) on the proxy to map the requester's IP address to a country. For authenticated users, you can also use account-level jurisdiction metadata (e.g., billing address, declared residence).
  • Request classification: Inspect the payload for data sensitivity. For example, if the input contains PII (names, emails, IDs), the request must be treated as high-risk and routed accordingly.
  • Routing table: A configuration that maps jurisdiction + risk level to a specific inference endpoint. The table can be stored in a database or a YAML file, reloaded without restarting the proxy.
  • Audit logging: Every routing decision is logged with timestamps, source IP (or user ID), target endpoint, and a hash of the request payload (for non-repudiation).

Proxy Implementation (Python + FastAPI Example)

Below is a simplified implementation using FastAPI as the proxy layer. It assumes you have two inference backends: one inside the EU (eu-inference.example.com) and one outside (global-inference.example.com).

from fastapi import FastAPI, Request
from fastapi.responses import JSONResponse
import httpx
import geoip2.database
import hashlib
import json

app = FastAPI()

# GeoIP reader (loaded at startup)
reader = geoip2.database.Reader('./GeoLite2-Country.mmdb')

# Routing table (could be loaded from config)
ROUTES = {
    "EU": {
        "high_risk": "https://eu-inference.example.com/v1/chat/completions",
        "low_risk": "https://eu-inference.example.com/v1/chat/completions"
    },
    "default": {
        "high_risk": "https://global-inference.example.com/v1/chat/completions",
        "low_risk": "https://global-inference.example.com/v1/chat/completions"
    }
}

# Simple PII detection (very basic — use a proper library in production)
def contains_pii(text: str) -> bool:
    patterns = [r'\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Z|a-z]{2,}\b',  # email
                r'\b\d{3}-\d{2}-\d{4}\b',  # SSN
                r'\b\d{16}\b']  # credit card
    import re
    for pat in patterns:
        if re.search(pat, text):
            return True
    return False

@app.post("/v1/chat/completions")
async def proxy_inference(request: Request):
    body = await request.json()
    user_ip = request.client.host

    # 1. Determine jurisdiction
    try:
        response = reader.country(user_ip)
        country = response.country.iso_code
    except:
        country = "UNKNOWN"

    # EU countries list (simplified)
    eu_countries = {"AT", "BE", "BG", "HR", "CY", "CZ", "DK", "EE", "FI", "FR", "DE", "GR", "HU", "IE", "IT", "LV", "LT", "LU", "MT", "NL", "PL", "PT", "RO", "SK", "SI", "ES", "SE"}
    jurisdiction = "EU" if country in eu_countries else "default"

    # 2. Classify risk based on PII in input
    messages = body.get("messages", [])
    full_text = " ".join([m.get("content", "") for m in messages])
    is_high_risk = contains_pii(full_text)
    risk_level = "high_risk" if is_high_risk else "low_risk"

    # 3. Select endpoint
    endpoint = ROUTES[jurisdiction][risk_level]

    # 4. Audit log
    payload_hash = hashlib.sha256(json.dumps(body, sort_keys=True).encode()).hexdigest()
    # In production, write to a structured log or database
    print(f"AUDIT: ip={user_ip} jurisdiction={jurisdiction} risk={risk_level} target={endpoint} hash={payload_hash}")

    # 5. Forward request (using httpx async client)
    async with httpx.AsyncClient() as client:
        resp = await client.post(endpoint, json=body, timeout=30.0)
        return JSONResponse(content=resp.json(), status_code=resp.status_code)

Handling Model Compliance

Even if you route requests correctly, the model itself must comply with the EU AI Act. For high-risk applications, you need:

  • Model cards: Each model version should have a documented card listing training data sources, intended use, limitations, and bias evaluations. The proxy can append a model version header to the response for traceability.
  • Explainability: If the model is used for decisions affecting individuals (e.g., credit scoring, hiring), you must provide explanations. This is hard to retrofit; prefer using inherently interpretable models or implementing post-hoc explanation methods like LIME or SHAP.
  • Human oversight: For high-risk systems, the Act requires human review. You can implement a fallback where the proxy flags certain requests for manual approval before inference.

Audit Trail Requirements

The EU AI Act mandates record-keeping for high-risk AI systems. Your proxy should log:

  • Timestamp of request
  • User identifier (pseudonymized if possible)
  • Input data hash (not the raw data, to avoid storing PII)
  • Model version and endpoint used
  • Output hash (optional, for reproducibility)
  • Routing decision rationale (jurisdiction, risk level)
  • Any errors or timeouts

Store these logs in an append-only store (e.g., a blockchain-based ledger or a WORM storage) to prevent tampering. Retention period should align with the Act's requirements (typically 5 years after the last use).

Geographic Load Balancing and Failover

If your EU inference endpoint goes down, you might be tempted to fall back to a non-EU endpoint. However, for EU-originating high-risk requests, this would violate data sovereignty. Instead, implement a circuit breaker that fails open only for low-risk requests, or queuing mechanism that holds requests until the EU endpoint recovers.

Caveats and Real-World Considerations

  • VPNs and IP spoofing: GeoIP is not foolproof. Users behind VPNs may appear to be in a different jurisdiction. For authenticated users, prefer account-level jurisdiction data. For unauthenticated, you can accept the GeoIP location and log the discrepancy.
  • Performance overhead: The proxy adds latency. A common pattern is that the GeoIP lookup and routing decision take under 5ms, but the PII scanning can be heavier. Use a fast regex engine or a dedicated PII detection service.
  • Regulatory changes: The EU AI Act is still evolving. Your routing rules should be configurable without code changes. Consider using a rules engine (e.g., Drools, Open Policy Agent) to manage complex conditions.

Conclusion

Jurisdictional inference routing at the proxy level is a pragmatic first step toward EU AI Act compliance. It decouples regulatory logic from model serving, allowing you to enforce data sovereignty without modifying your inference code. Combined with model cards, audit trails, and human oversight, it forms a solid foundation for a compliant AI infrastructure.

Remember: compliance is not a one-time checkbox. Your proxy routing must be continuously updated as regulations and model deployments change. Treat it as a critical piece of infrastructure, with version-controlled configurations and automated testing.

#audit#compliance#data-sovereignty#eu-ai-act#geo-routing#inference-routing#proxy-layer
Share — X / Twitter · LinkedIn · HN · Email
Damir Radulić
Founder of RiNET. On the Croatian internet since 1996 (Kvarner Net). In Amsterdam now, building autonomous AI infrastructure that runs on Monday morning when nobody's watching — sovereign stacks, agent swarms, LoRA fine-tuning, civic-intelligence platforms.

Related