Postgres as Our AI Control Plane: Boring, Boring, Boring

We run agent swarms, queues, vector stores, and audit logs on Postgres 16. No Redis, no Kafka, no Temporal.

by
Postgres as Our AI Control Plane: Boring, Boring, Boring

We at RiNET run our own GPUs, our own Postgres, and our own agent swarms. Our control plane is a single Postgres 16 instance on bare metal (pgbouncer on 6432, Hetzner EX130-S front nodes). No Redis, no Kafka, no Temporal. Just one boring database that handles queue jobs, vector stores, audit logs, and coordinating agent swarms. Here's why.

The Stack

Our GPU box runs primary inference via vLLM, with a paid fallback. Agents are coordinated through Postgres tables — no external message broker. Each agent has a row in agents table with its status, current task, and heartbeat timestamp. Tasks are rows in tasks table with payload, priority, and state (pending, running, completed, failed). A simple SELECT ... FOR UPDATE SKIP LOCKED fetches the next task — no Kafka partitions, no Redis streams.

Why Not Redis/Kafka/Temporal?

We've run Redis, Kafka, and Temporal in previous lives. They each add operational complexity: another stateful cluster to back up, another failure mode, another set of tuning knobs. For our scale — hundreds of agents, thousands of tasks per minute — Postgres handles it fine.

  • Queues: tasks table with state and priority columns. Workers poll via SKIP LOCKED. Latency is low for uncontended reads. Throughput is sufficient on our EX130-S nodes (AMD Epyc 16-core, 64GB RAM, NVMe).
  • Vector store: pgvector extension. Embeddings from our primary inference model. Query time is fast enough for top-10 search on millions of vectors. We don't need a specialized vector database.
  • Audit log: audit_events table with JSONB payload. Queries filtered by timestamp and agent_id. No Elasticsearch needed.
  • Agent coordination: agents table with status, current_task_id, heartbeat_at. Workers update their row every 30s. Dead agent detection via WHERE heartbeat_at < now() - interval '60 seconds'.

We run Postgres 16 with pgbouncer in transaction mode. Connection pooling is essential because we have hundreds of agent workers (Python async with psycopg3). Each worker holds a connection for the duration of a task (seconds to minutes). Pgbouncer keeps the backend connections warm.

The Boring Choice That Scales

We've seen Postgres handle many concurrent connections (with pgbouncer), high tps, and TB-scale datasets. For our current load — hundreds of agents, thousands of tasks per minute, hundreds of thousands of vector searches per day — it's overkill. But we like overkill that we understand.

Operational simplicity: one database to back up, one replication setup (we use streaming replication to a standby in another Hetzner datacenter), one monitoring dashboard (pg_stat_activity, pg_locks, pg_stat_user_tables). No Redis cluster failover, no Kafka rebalancing, no Temporal worker stuck on a shard.

We're not anti-Redis or anti-Kafka. If we hit much higher throughput or need exactly-once delivery across sharded agents, we'll reconsider. For now, the boring choice is the right choice.

What About Performance?

We've compared our queue approach against a Redis list with BRPOP. At moderate throughput, Postgres SKIP LOCKED latency is within a small factor of Redis. At higher throughput, Postgres starts to show higher tail latency — but we're not there yet. When we get there, we'll likely shard by agent type (separate tasks tables per domain) rather than introduce a new system.

Vector search with pgvector is fast enough for our RAG pipeline. We use IVFFlat index with appropriate parameters. Query times are low for top-k search on millions of vectors. We don't need HNSW because recall requirements are relaxed for agent context retrieval.

How We Deploy

All on Hetzner bare metal: EX130-S for front nodes (Postgres, app servers), GPU box for inference. We use Ansible for provisioning, systemd for process management. No Kubernetes. Backup: pgBackRest to Hetzner Storage Box daily. Monitoring: Prometheus + pg_exporter + Grafana.

We'll do it for you if you ask. We're happy to help if this hurts.

The Upshot

Postgres as a control plane is boring. That's the point. Less infrastructure, less cognitive load, less failover drama. Our agent swarms run on one database. Our queue is a table. Our vector store is an index. Our audit log is a table with a timestamp. It's not sexy. It works.

If you're building self-hosted AI infrastructure, start boring. You can always add Redis later. We haven't needed to.

We'll help you set this up. Just ask.

#agent-swarms#ai-infrastructure#architecture#bare-metal#control-plane#postgres
Share — X / Twitter · LinkedIn · HN · Email
Damir Radulić
Founder of RiNET. On the Croatian internet since 1996 (Kvarner Net). In Amsterdam now, building autonomous AI infrastructure that runs on Monday morning when nobody's watching — sovereign stacks, agent swarms, LoRA fine-tuning, civic-intelligence platforms.

Related