Building a Sovereign Embedding Pipeline: BGE-M3, pgvector, and Qdrant Side-by-Side

What we kept after comparing three approaches on the same corpus

by
Building a Sovereign Embedding Pipeline: BGE-M3, pgvector, and Qdrant Side-by-Side

When you decide to own your data and inference, the first question is: which vector database? The second: which embedding model? This article explores a pipeline using BGE-M3 for embeddings, stored in both pgvector (inside Postgres) and Qdrant. We discuss the principles, trade-offs, and patterns observed, without presenting fabricated measurements.

The Setup

Model: BGE-M3

BGE-M3 is a multilingual embedding model that outputs 1024-dimensional vectors. It supports dense, sparse, and multi-vector representations. We used the dense-only mode for simplicity, quantized to 4-bit via llama.cpp. Inference runs on CPU.

Corpus

A corpus of English technical documents (manuals, API docs, changelogs) split into chunks of 512 tokens with 128 token overlap. Each chunk yielded one embedding.

pgvector: The Postgres Way

pgvector runs as an extension inside Postgres. We created a table with a vector column and an IVFFlat index.

Query Performance

Top-10 cosine similarity search latency varies with cache state. Recall is reasonable but may degrade with larger index sizes.

The Good

  • No extra service to manage. Backups, replication, and access control are Postgres-native.
  • Metadata filtering is trivial — just add WHERE clauses.
  • pgvector supports HNSW indexes (as of later versions), but IVFFlat is simpler.

The Bad

  • IVFFlat recall can degrade with index size.
  • Index build time scales linearly. HNSW would help but increases memory pressure.
  • No built-in hybrid search (dense + sparse). BGE-M3’s multi-vector output can’t be fully exploited.

Qdrant: Purpose-Built Vector DB

Qdrant runs as a standalone binary with a gRPC API. It uses HNSW indexing by default.

Query Performance

Top-10 cosine similarity search is fast, with near-perfect recall when HNSW parameters are tuned.

The Good

  • Fast with tuned HNSW parameters.
  • Built-in hybrid search: Qdrant natively supports storing dense and sparse vectors per point. This allows using BGE-M3’s sparse output for keyword boosting.
  • Payload filtering is efficient with indexed fields.

The Bad

  • Another service to monitor, back up, and tune. Resource usage is non-trivial.
  • No transactional guarantees across Postgres and Qdrant. You need to handle consistency yourself.
  • Backup/restore is slower than pg_dump.

The Comparison

Metric pgvector (IVFFlat) Qdrant (HNSW)
Insert rate Lower Higher
Query latency (cold) Higher Lower
Query latency (warm) Higher Lower
Recall Good Better
Memory (idle) Lower Higher
Disk usage Lower Higher
Operational overhead Low (Postgres extension) Medium (separate service)

What We Kept

We kept both. Not as an either/or — as a layered system.

Tier 1: Postgres + pgvector for Metadata-Grounded Queries

When the query includes filters like WHERE metadata->>'domain' = 'networking', pgvector shines. The combination of vector search and SQL filtering in one query avoids round-trips. IVFFlat is sufficient for many use cases where recall requirements are lenient. For high-stakes retrieval, we fall back to Qdrant.

Qdrant handles the bulk of real-time user-facing search. Its hybrid search with BGE-M3’s sparse vectors improves relevance for keyword-heavy queries. We replicate embeddings to both stores via a background worker.

Why Not One?

  • pgvector alone: recall drops and latency spikes under load. No sparse support.
  • Qdrant alone: operational complexity and no easy way to join with relational data. Metadata filtering is powerful but not SQL.

Operational Notes

Consistency Between Stores

We use a Postgres NOTIFY/LISTEN pattern: when a new chunk is inserted, a trigger sends a notification. A daemon picks it up, embeds with BGE-M3, and writes to both pgvector and Qdrant. If one write fails, the daemon retries with exponential backoff. We accept eventual consistency.

BGE-M3 Inference

We run llama.cpp’s embedding server with the BGE-M3 Q4_K_M model. It listens on a Unix socket. Both the ingestion daemon and a separate API server call it.

Backup Strategy

  • pgvector: standard pg_dump + WAL archiving.
  • Qdrant: snapshot API + periodic rsync of the storage directory.

Lessons Learned

  1. Don’t over-index. IVFFlat with lists = sqrt(rows) is a good starting point. Tuning HNSW in pgvector may not be necessary if IVFFlat is enough.
  2. Hybrid search matters. BGE-M3’s sparse vectors improve recall on exact-match queries. Qdrant makes it trivial; pgvector doesn’t support it.
  3. Operation cost is real. Qdrant needs memory and monitoring. If your team is small, pgvector may be the better default despite lower performance.
  4. Benchmark on your data. Recall numbers differ from public benchmarks because your corpus is domain-specific. Always test.

The Stack Today

  • Ingestion: Python for chunking, Go daemon for orchestration.
  • Embedding: llama.cpp server with BGE-M3 Q4_K_M.
  • Vector stores: Postgres + pgvector (IVFFlat) and Qdrant (HNSW).
  • Search API: FastAPI that routes queries to Qdrant by default, with a fallback to pgvector for filtered queries.

We run it on bare-metal servers in a colo facility. No cloud, no SaaS. That’s sovereignty.

Final Word

If you need a simple vector search with SQL integration, start with pgvector. If you need speed, recall, and hybrid search at scale, add Qdrant. Don’t force one to do everything. Our pipeline uses the strengths of both, and the overhead of running two stores is justified by the reliability and performance we get.

Build for your data, not for benchmarks.

#architecture#bge-m3#embeddings#pgvector#qdrant
Share — X / Twitter · LinkedIn · HN · Email
Damir Radulić
Founder of RiNET. On the Croatian internet since 1996 (Kvarner Net). In Amsterdam now, building autonomous AI infrastructure that runs on Monday morning when nobody's watching — sovereign stacks, agent swarms, LoRA fine-tuning, civic-intelligence platforms.

Related