Embeddings Are the Crown Jewels: Back Up Vectors Before Code
Why your vector database is more valuable than your source code, and how to protect it.
I've seen teams spend months curating a custom embedding pipeline, only to lose the vectors in a cluster failure. The code was safe in git. The training scripts were versioned. But the embeddings? Gone. They had to re-embed millions of documents, praying the model hadn't drifted. Don't be that team.
Embeddings are the crown jewels of any retrieval-augmented generation (RAG) system or semantic search pipeline. They encode your domain knowledge, your data distribution, and your business logic. Code is cheap to rewrite; embeddings are expensive to regenerate. Here's how to treat them accordingly.
Why Embeddings Are More Valuable Than Code
Code is deterministic. Given the same input, it produces the same output. Embeddings are the output of a complex, often non-deterministic pipeline: tokenization, model inference, normalization, quantization. If you lose your vectors, you must:
- Re-run the embedding model on every source document.
- Hope the model hasn't changed (new version, different weights).
- Hope the tokenizer hasn't changed (sentencepiece update, BPE merge changes).
- Hope the normalization parameters haven't drifted.
Even a minor change in the model can shift embeddings enough to break nearest-neighbor search quality. In production, I've seen cosine similarity drop measurably after a model update, causing recall to fall significantly. That's a regression you can't patch with a hotfix.
What to Back Up
A complete embedding backup includes:
- Vector data: the float arrays (or quantized integers) with their IDs.
- Metadata: text snippets, timestamps, source URLs, any fields used in filtering.
- Index structures: IVF centroids, HNSW graphs, product quantization codebooks. Rebuilding an index on millions of vectors can take hours.
- Model snapshot: the exact model weights and tokenizer config used to generate the embeddings.
- Normalization parameters: mean/std if you applied any post-processing.
Without the model snapshot, your backup is a pile of meaningless numbers. You can't regenerate them if the model changes.
Backup Strategy: Incremental + Full
Don't rely on a single nightly dump. Implement a two-tier strategy:
Incremental Backups (Every N Minutes)
Capture new or changed vectors since the last backup. Most vector databases support change tracking via write-ahead logs (WAL) or CDC streams.
For example, with pgvector (PostgreSQL extension), you can use pg_dump with --data-only on the embeddings table, but that's a full dump. Better: use logical replication to stream changes to a standby or a separate backup store.
-- Create a publication for the embeddings table
CREATE PUBLICATION emb_backup FOR TABLE embeddings;Then subscribe on the backup instance:
CREATE SUBSCRIPTION emb_sub CONNECTION 'host=primary dbname=vectors' PUBLICATION emb_backup;This gives you near-real-time replication. For a simpler approach, log vector IDs and timestamps, and periodically batch-export new ones:
import psycopg2
import numpy as np
conn = psycopg2.connect("dbname=vectors")
cur = conn.cursor()
cur.execute("""
SELECT id, embedding, metadata, updated_at
FROM embeddings
WHERE updated_at > %s
""", (last_backup_time,))
rows = cur.fetchall()
# Save to Parquet or numpy arrays
for row in rows:
vec_id, vec, meta, ts = row
np.save(f"backups/vec_{vec_id}.npy", np.array(vec))Full Backups (Daily)
Once a day, take a full snapshot of the vector database. For Weaviate, use their export APIs. For Qdrant, snapshot the collection:
curl -X POST 'http://localhost:6333/collections/my_collection/snapshots'For Milvus, use backup tool:
milvus-backup create -n daily-backup-2025-03-15Store full backups in cold storage (S3 Glacier, GCS Nearline) with a retention policy of 30-90 days. Incremental backups go to hot storage for quick restore.
Restore Procedure
Test your restore at least once a month. Here's a script that restores a pgvector backup:
#!/bin/bash
# Restore full backup from S3
aws s3 cp s3://my-backups/vectors/2025-03-15/full_dump.sql /tmp/
psql -h localhost -U admin -d vectors < /tmp/full_dump.sql
# Replay incremental backups in order
for inc in $(ls /tmp/incremental/*.sql | sort); do
psql -h localhost -U admin -d vectors < $inc
done
# Rebuild index (critical step!)
psql -h localhost -U admin -d vectors -c "REINDEX INDEX embeddings_idx;"For HNSW indexes, rebuilding can take hours. Use REINDEX CONCURRENTLY to avoid locking writes during restore.
Automating with a Backup Daemon
I run a small Python daemon that monitors the vector database and triggers backups. Here's a simplified version:
import time
import subprocess
from datetime import datetime
BACKUP_INTERVAL = 3600 # 1 hour
FULL_BACKUP_HOUR = 3 # 3 AM
def incremental_backup():
ts = datetime.utcnow().isoformat()
subprocess.run([
"pg_dump", "--data-only", "--table=embeddings",
"--file", f"/backups/inc_{ts}.sql",
"dbname=vectors"
])
# Upload to S3
subprocess.run(["aws", "s3", "cp", f"/backups/inc_{ts}.sql",
f"s3://backups/vectors/incremental/"])
def full_backup():
ts = datetime.utcnow().strftime("%Y-%m-%d")
subprocess.run([
"pg_dump", "--file", f"/backups/full_{ts}.sql",
"dbname=vectors"
])
subprocess.run(["aws", "s3", "cp", f"/backups/full_{ts}.sql",
f"s3://backups/vectors/full/"])
while True:
now = datetime.utcnow()
if now.hour == FULL_BACKUP_HOUR and now.minute == 0:
full_backup()
else:
incremental_backup()
time.sleep(BACKUP_INTERVAL)Quantization and Storage Costs
Raw float32 vectors are expensive. A 768-dim vector takes 3 KB. For millions of vectors, that's tens of gigabytes just for the vectors. Backups multiply that. Use quantization to reduce storage:
- Binary quantization: convert each float to a single bit (sign). 768 bits = 96 bytes per vector. High compression. Works well for cosine similarity.
- Product quantization (PQ): split vector into subvectors, quantize each to a centroid. Typical compression 4x-8x with minimal recall loss.
- Scalar quantization (int8): map float32 to int8. 4x compression, negligible recall drop.
Store quantized vectors in backups, but keep the original for training or fine-tuning. Example: convert to binary before backup:
def quantize_binary(vec):
return (vec > 0).astype(np.int8)
binary_vec = quantize_binary(original_float32_vec)
np.save("backup_binary.npy", binary_vec)On restore, you can either use the binary vectors directly (if your index supports it) or dequantize back to float32 (lossy).
Real-World Lesson: The 3 AM Index Rebuild
We once lost a large vector index due to a corrupted HNSW graph. The backup was hours old. Restoring the data was fast, but rebuilding the index took many hours. Users saw degraded search quality during that window.
Now we back up the index structure separately. For Qdrant, snapshots include the HNSW graph. For pgvector, we schedule REINDEX during low traffic and back up the index file.
Final Checklist
- Backup vectors daily (full) + hourly (incremental).
- Store model snapshot and tokenizer alongside vectors.
- Test restore monthly, including index rebuild.
- Use quantization to reduce backup size.
- Monitor backup freshness with alerts.
- Document the restore procedure in your runbook.
Your code can be rewritten. Your embeddings cannot. Treat them like the crown jewels they are.