WireGuard Mesh for a Sovereign AI Stack: Keeping the Data Plane Off the Public Internet
Why your inference cluster shouldn't depend on public IPs
WireGuard Mesh for a Sovereign AI Stack: Keeping the Data Plane Off the Public Internet
If you're running a self-hosted AI stack — LLM inference, embedding pipelines, vector databases — you've probably accepted that traffic between services crosses the public internet. Maybe you authenticated with API keys, maybe you used TLS. But the data plane still flows through someone else's switches.
That's a sovereignty problem. Under the EU AI Act, data that passes through public infrastructure can trigger compliance obligations you didn't sign up for. More importantly, it's a performance and reliability problem: public internet routing adds latency, jitter, and unpredictable throughput.
A WireGuard mesh solves this. Not a traditional hub-and-spoke VPN with a single gateway that becomes a bottleneck and single point of failure. A true mesh: every node talks to every other node directly, over encrypted WireGuard tunnels, using private IPs that never touch the public internet.
Why Not a Traditional VPN?
Most teams reach for OpenVPN or WireGuard in a hub-spoke setup. You run one VPN server, all clients connect to it, and traffic routes through that central node. That works for remote access, but it's wrong for a distributed AI stack.
- Latency: Every packet between node A and node B goes through the hub. That adds a hop and doubles transit time.
- Bandwidth bottleneck: The hub's NIC or CPU becomes the ceiling for all inter-node traffic. With LLM inference, you're pushing gigabytes of model weights and embedding vectors. That hub will melt.
- Single point of failure: Hub goes down, the entire cluster loses connectivity. No graceful degradation.
A mesh avoids all of this. Each node maintains direct WireGuard tunnels to every other node. Traffic goes point-to-point, encrypted, with no intermediary. WireGuard's kernel implementation handles this with minimal CPU overhead — we're talking single-digit microseconds per packet.
Anatomy of a WireGuard Mesh for AI
Let's be concrete. Suppose you have three nodes:
node-0: Inference server running vLLM (0.6.x) with two A100snode-1: Embedding server running BGE-M3 on llama.cppnode-2: Qdrant (1.9.x) vector database + Postgres (16) with pgvector
Each node has a public IP (or sits behind NAT with a public endpoint). You want them to communicate over private 10.0.x.x addresses.
Configuration
On each node, create /etc/wireguard/wg0.conf. Here's node-0's config:
[Interface]
Address = 10.0.0.1/24
PrivateKey = <node-0-private-key>
ListenPort = 51820
# node-1
[Peer]
PublicKey = <node-1-public-key>
Endpoint = node-1.example.com:51820
AllowedIPs = 10.0.0.2/32
PersistentKeepalive = 25
# node-2
[Peer]
PublicKey = <node-2-public-key>
Endpoint = node-2.example.com:51820
AllowedIPs = 10.0.0.3/32
PersistentKeepalive = 25Node-1 gets 10.0.0.2, node-2 gets 10.0.0.3. Each node lists the other two as peers. PersistentKeepalive ensures NAT traversal stays alive.
Start the interface:
systemctl enable --now wg-quick@wg0That's it. Now ping 10.0.0.2 from node-0 reaches node-1 directly, encrypted, with no VPN gateway.
Routing the AI Data Plane
With the mesh up, you reconfigure your services to bind to the WireGuard interface instead of 0.0.0.0 or the public IP.
vLLM Inference
On node-0, start vLLM with --host 10.0.0.1. Your client on node-2 (or anywhere in the mesh) can call:
curl -X POST http://10.0.0.1:8000/v1/completions -H "Content-Type: application/json" -d '{"model": "mistral-7b", "prompt": "Hello"}'No TLS needed if you trust the mesh (WireGuard already encrypts), but you can layer mTLS for defense in depth.
Embedding Pipeline
On node-1, run llama.cpp's embedding server with --host 10.0.0.2. Your inference server can call it:
import requests
response = requests.post("http://10.0.0.2:8080/embedding", json={"input": "your text"})Vector Database
On node-2, Qdrant listens on 10.0.0.3:6333. Postgres with pgvector also binds to 10.0.0.3. Embeddings flow from node-1 to node-2 over the mesh, never leaving your private network.
Performance Considerations
WireGuard lives in the kernel. On modern Linux (5.6+), it's a crypto-ready tunnel that uses ChaCha20-Poly1305 — fast on CPUs with AES-NI, but also fast without. You'll see line-rate throughput on 10GbE NICs with 5-10% CPU overhead.
For AI workloads, the critical metric is latency. A direct WireGuard tunnel adds roughly 0.1-0.3ms per packet. Compare that to a hub-spoke VPN where you add that twice (A→hub→B), plus the hub's processing. In practice, your inference latency is dominated by GPU compute and memory bandwidth, not the network. But for embedding batches or vector search, every microsecond counts.
MTU Tuning
WireGuard's default MTU is 1420 bytes (1500 minus 80 for headers). If you're moving large model chunks or embedding vectors, consider jumbo frames if your underlying network supports them. Set MTU = 9000 in the [Interface] section — but only if all nodes and the physical switch support it.
Mesh Automation with systemd-networking
Manually editing configs on three nodes is fine. On twenty, it's not. Use a tool like wg-meshconf or a simple Ansible playbook to generate configs from a YAML inventory.
Example inventory (inventory.yaml):
nodes:
- name: node-0
wg_ip: 10.0.0.1
public_endpoint: node-0.example.com:51820
- name: node-1
wg_ip: 10.0.0.2
public_endpoint: node-1.example.com:51820
- name: node-2
wg_ip: 10.0.0.3
public_endpoint: node-2.example.com:51820Ansible task (pseudo):
- name: Generate WireGuard config for each node
template:
src: wg0.conf.j2
dest: /etc/wireguard/wg0.conf
loop: "{{ nodes }}"With Jinja2, you iterate over all peers except the current node. This scales to N nodes with O(N) configs.
NAT Traversal and Dynamic IPs
If nodes are behind NAT (e.g., home lab, colo with private IP), WireGuard's PersistentKeepalive keeps the tunnel up. But if public IPs change, you need dynamic endpoint resolution. Options:
- Use a DDNS domain for each node.
- Run a small coordination service (like
netmakerortailscale— but those add complexity). - For a small cluster, just update the configs manually when IPs change. It's a
systemctl restart wg-quick@wg0per node.
Security Beyond Encryption
WireGuard provides authenticated encryption. But you still need to control which services are reachable over the mesh. By default, any node can reach any other node's WireGuard IP on any port. That's fine for a trusted cluster, but you can add iptables or nftables rules to restrict inter-node traffic.
Example: only allow TCP/8000 from node-2 to node-0 (inference API):
nft add rule inet filter input iifname wg0 ip saddr 10.0.0.3 tcp dport 8000 accept
nft add rule inet filter input iifname wg0 dropReal-World Example: Replacing a SaaS Embedding API
Before the mesh, you might have called OpenAI's embedding API over the public internet. That means your raw text — potentially PII or business data — left your network. With the mesh, you run BGE-M3 on node-1, and your application on node-2 sends text over the WireGuard tunnel. No data ever touches a public IP. The EU AI Act's transparency obligations for high-risk systems still apply to the model itself, but the data transit is fully under your control.
The Catch: Key Distribution
WireGuard's security relies on pre-shared public keys. If an attacker compromises a node and steals its private key, they can impersonate that node. Mitigations:
- Store private keys in a hardware security module (HSM) or TPM if available.
- Rotate keys periodically.
- Use
wg-quick'sSaveConfig = trueonly if you trust the filesystem.
For most self-hosted stacks, the risk is low — your nodes are already hardened. But don't ignore it.
Final Take
A WireGuard mesh is the simplest, highest-performance way to build a private data plane for a sovereign AI stack. It eliminates public internet transit, reduces latency, and removes single points of failure. You get kernel-speed encryption, zero configuration overhead beyond the initial setup, and the flexibility to route any service — inference, embeddings, vector search — over private IPs.
If you're building a self-hosted AI pipeline and you haven't moved the data plane off the public internet, start today. Two hours of config work, and your traffic stops leaking through the open net.