Multi-tenant from day one: one engine, many portals, zero forks
How to build an on-prem platform that serves many clients without duplicating code or data
Multi-tenant from day one: one engine, many portals, zero forks
Every time you fork a repo for a new client, you create a maintenance nightmare. I've seen teams with ten forks, each with its own database, its own inference server, its own deployment pipeline. The result: duplicated effort, security patches applied unevenly, and a growing backlog of feature parity work.
There is a better way. A single Postgres-backed inference engine, with row-level security and per-tenant embeddings, can serve hundreds of tenants without operational chaos. This is the architecture we use at The Sovereign Stack for every on-prem deployment.
The problem with per-tenant forks
Forking seems easy at first. You copy the repo, change the logo, deploy. But soon you face:
- Feature drift: Tenant A gets RAG support, Tenant B still runs on the old prompt template.
- Security debt: You fix a prompt injection vulnerability in one fork but forget the other three.
- Resource waste: Each fork runs its own vLLM instance, its own Qdrant cluster, its own Postgres. Idle memory everywhere.
- Upgrade pain: Moving from llama.cpp v0.2 to v0.3 means touching every fork.
Multi-tenant design eliminates these problems by design. One engine, one database, one inference cluster, many isolated portals.
Core principle: data isolation at the storage layer
The secret is to push tenant isolation as deep as possible — into the database. Don't trust the application layer to filter rows correctly. Use row-level security (RLS) in Postgres so that even if a bug leaks data, the database refuses to serve it.
Here's the schema pattern:
CREATE TABLE tenants (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
name TEXT NOT NULL,
api_key_hash TEXT NOT NULL
);
ALTER TABLE tenants ENABLE ROW LEVEL SECURITY;
CREATE TABLE documents (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
tenant_id UUID NOT NULL REFERENCES tenants(id) ON DELETE CASCADE,
content TEXT NOT NULL,
embedding vector(768),
created_at TIMESTAMPTZ DEFAULT now()
);
ALTER TABLE documents ENABLE ROW LEVEL SECURITY;
CREATE POLICY tenant_isolation ON documents
USING (tenant_id = current_setting('app.current_tenant_id')::UUID);The application sets app.current_tenant_id at connection pool initialization (via pgbouncer auth query or session variable). Every query automatically filters by tenant. No WHERE tenant_id = ? in application code — it's enforced by the database.
One engine, many portals
The "engine" is a single inference service (vLLM for GPU, llama.cpp for CPU) loaded with one base model. The "portal" is a lightweight web app that authenticates users and sets the tenant context. The engine doesn't know about tenants — it just runs inference. The portal handles:
- Authentication and tenant resolution
- Setting the Postgres session variable
- Routing requests to the engine with a tenant header for logging
This means you can deploy one vLLM instance serving all tenants. Throughput scales with GPU memory, not with tenant count. A single A100 can handle dozens of concurrent requests across many tenants.
Per-tenant embeddings with pgvector
Embeddings are the trickiest part. You can't share a single embedding index across tenants because the vector space is meaningless if you mix data. The solution: partitioned vector indexes.
With pgvector 0.7.0+, you can create indexes on partitioned tables:
CREATE TABLE embeddings (
id UUID DEFAULT gen_random_uuid(),
tenant_id UUID NOT NULL,
document_id UUID NOT NULL,
embedding vector(768),
PRIMARY KEY (tenant_id, id)
) PARTITION BY LIST (tenant_id);
CREATE INDEX idx_embeddings_tenant_ivfflat ON embeddings
USING ivfflat (embedding vector_cosine_ops)
WITH (lists = 100);Alternatively, use separate schemas per tenant. Both approaches work. The key is that each tenant's vectors are physically separate, so ANN search doesn't leak data.
Authentication and API keys
Each tenant gets an API key stored as a bcrypt hash. The portal validates the key and sets the tenant context. For on-prem deployments, you can also use client certificates or VPN-based trust.
# FastAPI middleware example
@app.middleware("http")
async def tenant_middleware(request: Request, call_next):
api_key = request.headers.get("X-API-Key")
tenant = get_tenant_by_key(api_key)
if not tenant:
return JSONResponse(status_code=401, content={"error": "invalid key"})
# Set Postgres session variable
async with database.connect() as conn:
await conn.execute(
f"SET app.current_tenant_id = '{tenant.id}'"
)
response = await call_next(request)
return responseRate limiting and resource isolation
One tenant should not starve another. Use a rate limiter per tenant, not per IP. Redis is fine for this, but a Postgres advisory lock with a token bucket works too for smaller deployments.
For GPU memory, use vLLM's per-request scheduling. vLLM 0.6.0+ supports request-level priorities via the priority header. Map tenant tiers to priorities: gold tenants get high, free tier gets low.
Deployment topology
A single-machine setup can handle 10–50 tenants:
- Postgres 16 + pgvector 0.7.0
- pgbouncer for connection pooling
- vLLM or llama.cpp server
- A lightweight web server (nginx + uvicorn)
- systemd units for each service
[Unit]
Description=Inference Engine
After=network.target postgresql.service
[Service]
ExecStart=/usr/local/bin/vllm serve /models/llama-3.1-8b --port 8000 --max-model-len 8192 --gpu-memory-utilization 0.9
Restart=always
User=vllm
[Install]
WantedBy=multi-user.targetFor >100 tenants, scale horizontally: multiple vLLM replicas behind a load balancer, each with a copy of the model. Postgres can stay single-node with proper indexing, or use streaming replication for read replicas.
Migration path from forks
If you already have forks, don't panic. The migration is straightforward:
- Create a new multi-tenant database with the schema above.
- For each tenant, copy their data into the new database with their
tenant_id. - Deploy the unified engine and point all tenants to it.
- Keep old forks running in read-only mode during transition.
- Update DNS and API endpoints.
You can do this incrementally, one tenant per week.
Why this matters for data sovereignty
Multi-tenant doesn't mean multi-jurisdiction. For EU AI Act compliance, you need to know where data lives and who can access it. With a single engine, you can still deploy per-region: one engine in Frankfurt, one in Ireland. Each engine is multi-tenant within its region. This gives you the operational benefits of multi-tenant while satisfying data residency requirements.
The bottom line
Multi-tenant architecture is not just for SaaS. It's for any platform that serves multiple clients. Start with Postgres RLS, add pgvector for embeddings, and run a single inference engine. You'll reduce operational overhead, improve security, and make upgrades trivial.
No more forks. One engine. Many portals. Zero maintenance debt.