No Public Endpoints: How We WireGuard-Mesh Every Service in a Sovereign AI Stack

Zero-trust networking for self-hosted LLM inference and agent infrastructure

by
No Public Endpoints: How We WireGuard-Mesh Every Service in a Sovereign AI Stack

No Public Endpoints: How We WireGuard-Mesh Every Service in a Sovereign AI Stack

Every public endpoint is a liability. In a sovereign AI stack—where you run your own LLM inference, vector database, Postgres control plane, and agent swarms—the attack surface is already large. Why make it worse by exposing ports to the internet?

We run everything behind a single WireGuard endpoint. Every service binds to a private IP (10.0.0.0/8) and talks only over the mesh. There are no open TCP ports on the public internet except the WireGuard UDP port itself. Here's exactly how we do it.

The Goal: Flat, Encrypted, Zero-Trust Network

  • Every host (bare metal, VM, container) gets a WireGuard interface with a static IP.
  • All inter-service traffic uses those private IPs.
  • No firewall rules for "allow port X from Y"—if you're on the mesh, you can reach any service (we layer application auth on top).
  • The only public listener is WireGuard on UDP 51820.

WireGuard Mesh Topology

We use a full mesh: every node has a config that lists every other node as a peer. For 10 nodes, that's 9 peers per config. We generate these with a script.

Node layout example:

Role Hostname WireGuard IP Public IP
Router gw 10.0.0.1 203.0.113.10
LLM inference vllm-1 10.0.0.10 (none)
Vector DB qdrant-1 10.0.0.20 (none)
Postgres pg-1 10.0.0.30 (none)
Agent orchestrator orchestrator 10.0.0.40 (none)
Embedding pipeline embed-1 10.0.0.50 (none)

The gateway node (gw) has a public IP and forwards nothing except WireGuard. All other nodes have no public IP at all—they reach the internet through gw's NAT (via iptables masquerade), but that's only for package updates and pulling Docker images.

Generating the Mesh Config

We use a simple Python script that reads a YAML inventory and writes per-host WireGuard configs.

# inventory.yaml
nodes:
  - name: gw
    wg_ip: 10.0.0.1
    public_ip: 203.0.113.10
    listen_port: 51820
    private_key: "<base64>"
  - name: vllm-1
    wg_ip: 10.0.0.10
    public_ip: null
    listen_port: 51820
    private_key: "<base64>"
  # ... more nodes

The script generates one file per node. For vllm-1, it looks like:

[Interface]
PrivateKey = <vllm-1 private key>
Address = 10.0.0.10/32
ListenPort = 51820

[Peer]
PublicKey = <gw public key>
Endpoint = 203.0.113.10:51820
AllowedIPs = 10.0.0.0/8
PersistentKeepalive = 25

[Peer]
PublicKey = <qdrant-1 public key>
Endpoint = 10.0.0.20:51820
AllowedIPs = 10.0.0.20/32

[Peer]
PublicKey = <pg-1 public key>
Endpoint = 10.0.0.30:51820
AllowedIPs = 10.0.0.30/32

# ... one Peer block per node

Note: Only the gateway node's Endpoint uses a public IP. All other nodes use WireGuard's internal IP as the endpoint—this is the mesh magic. WireGuard will route through the gateway if needed, but once the tunnel is up, traffic between two non-gateway nodes goes direct (if they can reach each other) or via the gateway (if behind NAT). We set PersistentKeepalive on the gateway peer to keep NAT mappings alive.

Deploying the Config

We use Ansible to push configs and restart WireGuard:

- name: Install WireGuard
  apt:
    name: wireguard
    state: present

- name: Push config
  copy:
    src: "wg-confs/{{ inventory_hostname }}.conf"
    dest: /etc/wireguard/wg0.conf
    mode: 0600
  notify: restart wg-quick

- name: Enable and start
  systemd:
    name: wg-quick@wg0
    enabled: yes
    state: started

We also enable IP forwarding on each node and add a masquerade rule on the gateway:

iptables -t nat -A POSTROUTING -o eth0 -j MASQUERADE

Now every node can reach the internet through the gateway, but no node exposes any TCP/UDP port to the internet except the gateway's WireGuard UDP port.

Binding Services to WireGuard IPs

Every service binds to its WireGuard IP. Examples:

Postgres (/etc/postgresql/16/main/postgresql.conf):

listen_addresses = '10.0.0.30'
port = 5432

pgbouncer (/etc/pgbouncer/pgbouncer.ini):

[databases]
* = host=10.0.0.30 port=5432

[pgbouncer]
listen_addr = 10.0.0.30
listen_port = 6432

vLLM (start with --host 10.0.0.10):

python -m vllm.entrypoints.openai.api_server \
  --model meta-llama/Llama-3.1-8B-Instruct \
  --host 10.0.0.10 --port 8000 \
  --tensor-parallel-size 2

Qdrant (config.yaml):

service:
  host: 10.0.0.20
  http_port: 6333
  grpc_port: 6334

Embedding pipeline (BGE-M3 via FastAPI):

uvicorn.run(app, host="10.0.0.50", port=8080)

Agent orchestrator (Docker compose with network host or explicit IP):

services:
  orchestrator:
    network_mode: host
    environment:
      - LLM_ENDPOINT=http://10.0.0.10:8000/v1
      - QDRANT_HOST=10.0.0.20
      - POSTGRES_HOST=10.0.0.30

DNS via /etc/hosts (or Consul)

We keep it simple: each node has an /etc/hosts file mapping hostnames to WireGuard IPs.

10.0.0.1   gw
10.0.0.10  vllm-1
10.0.0.20  qdrant-1
10.0.0.30  pg-1
10.0.0.40  orchestrator
10.0.0.50  embed-1

Ansible pushes this too. If you have more nodes, consider Consul with a WireGuard-constrained advertise address, but for <50 nodes, /etc/hosts is fine.

Exposing APIs to the Outside World (Carefully)

Sometimes you need to expose an API—say, a chat interface for internal users. We never expose the service directly. Instead, we run a reverse proxy (Caddy or nginx) on the gateway, bound to the public IP, and proxy to the WireGuard IP.

Caddyfile:

chat.example.com {
    reverse_proxy 10.0.0.40:8080
    tls internal {
        on_demand
    }
}

That's it. Only the gateway has a public listener, and only on ports 80/443 (plus 22 for SSH, but we use SSH over WireGuard too).

Why This Works for Sovereign AI

  • Data sovereignty: All traffic stays within your encrypted mesh. No third-party VPN, no cloud networking.
  • EU AI Act compliance: No data leaves your infrastructure. If you need to log access, you can audit every packet on the gateway.
  • Simplicity: One port to firewall. One protocol to understand. No complex SDN.
  • Performance: WireGuard is in-kernel on Linux. Overhead is negligible—we push 900+ Mbps per tunnel on modern hardware.

Gotchas

  • MTU: WireGuard adds 60 bytes of overhead. Set MTU to 1420 on the tunnel interface and 1500 on the physical. Otherwise, you'll see weird packet drops with large payloads (e.g., embedding vectors).
  • PersistentKeepalive: Set to 25 seconds on the gateway peer for nodes behind NAT. If a node changes IP (e.g., roaming laptop), it won't reconnect until the keepalive fires.
  • Key rotation: We rotate keys quarterly. Ansible handles this: generate new keys, push configs, restart WireGuard. Brief disruption (~1 second).
  • Split horizon DNS: If you also run public DNS for your domain, make sure internal clients resolve to WireGuard IPs, not public ones. We use a separate internal.example.com zone.

Alternatives Considered

  • Tailscale / ZeroTier: Great products, but they use a coordination server. For a sovereign stack, we wanted no external dependency. WireGuard is pure open source and auditable.
  • IPsec / OpenVPN: Too complex. WireGuard's crypto is modern (Curve25519, ChaCha20, Poly1305) and the config is 10 lines.
  • VLANs: Requires managed switches and doesn't encrypt traffic. WireGuard encrypts everything.

The Result

After deploying this mesh across 12 nodes (3 inference, 2 vector DBs, 2 Postgres, 2 embedding, 2 orchestrators, 1 gateway), we have zero public endpoints. The only way to reach any service is through the WireGuard tunnel. We sleep better.

If you're building a self-hosted AI stack, start with the network. WireGuard mesh is the foundation that makes everything else simpler: no TLS certs for internal services, no firewall rules for inter-service communication, no fear of accidental exposure.

Your stack is sovereign when your network is sovereign. WireGuard gives you that with 50 lines of config.

#infrastructure#networking#self-hosted-ai#wireguard#zero-trust
Share — X / Twitter · LinkedIn · HN · Email
Damir Radulić
Founder of RiNET. On the Croatian internet since 1996 (Kvarner Net). In Amsterdam now, building autonomous AI infrastructure that runs on Monday morning when nobody's watching — sovereign stacks, agent swarms, LoRA fine-tuning, civic-intelligence platforms.