Resolving Conflicts Between Vector, Graph, and Relational Stores for Agent Recall
A practical guide to unifying heterogeneous memory backends without sacrificing consistency
Resolving Conflicts Between Vector, Graph, and Relational Stores for Agent Recall
Autonomous agents don't store memory in one place. They use vector stores for semantic similarity, graph databases for relationships, and relational databases for structured facts. This heterogeneous approach is powerful, but it introduces a fundamental problem: conflicts. When the same piece of information is stored in three different systems, which one is the source of truth? How do you keep them consistent? This article walks through practical conflict resolution strategies using real tools like Postgres (with pgvector), Qdrant, and Neo4j.
The Nature of Conflicts
Conflicts arise when data in different stores diverges. Common scenarios:
- Temporal inconsistency: The agent updates a fact in Postgres but the embedding in Qdrant is stale.
- Semantic drift: A graph relationship is deleted but the vector index still clusters related embeddings.
- Data duplication: The same entity is stored with different IDs across stores, leading to orphaned references.
Not all conflicts are equal. Some are tolerable (e.g., a slightly stale embedding for a rarely queried entity), others are catastrophic (e.g., an agent acting on a deleted relationship). The key is to classify conflicts by severity and apply the right resolution strategy.
Resolution Strategies
1. Last-Writer-Wins (LWW) with Timestamps
The simplest approach: each store gets a monotonic timestamp per record. When a conflict is detected, the record with the latest timestamp wins. This works well for simple key-value-like data but fails for complex dependencies (e.g., a graph edge referencing a deleted node).
Implementation: Use updated_at columns in Postgres, store a timestamp in the payload of Qdrant points, and use a property on Neo4j nodes. A background reconciliation job queries all stores for the same logical entity and picks the latest.
-- Postgres table
CREATE TABLE agent_memory (
id UUID PRIMARY KEY,
content TEXT,
updated_at TIMESTAMPTZ DEFAULT NOW()
);# Reconciliation script using Qdrant and Neo4j
import qdrant_client
from neo4j import GraphDatabase
client = qdrant_client.QdrantClient(host="localhost", port=6333)
driver = GraphDatabase.driver("bolt://localhost:7687")
def resolve_entity(entity_id):
# Fetch from all stores
pg_record = fetch_from_postgres(entity_id)
qdrant_point = client.retrieve("agent_memory", [entity_id])[0]
with driver.session() as session:
neo4j_record = session.run("MATCH (n:Memory {id: $id}) RETURN n", id=entity_id).single()
# Compare timestamps
timestamps = [
pg_record['updated_at'],
qdrant_point.payload['updated_at'],
neo4j_record['n']['updated_at']
]
latest_idx = timestamps.index(max(timestamps))
# Propagate latest to all stores
...Pros: Simple, fast. Cons: Ignores causal dependencies; can lose updates if clocks drift.
2. CRDTs (Conflict-Free Replicated Data Types)
For agents that operate offline or in distributed settings, CRDTs allow each store to be updated independently and merged later without conflicts. Use a state-based CRDT (e.g., LWW-Register) for scalar values, or a map-based CRDT for complex structures.
Implementation: Store a version vector with each record. When merging, take the union of all versions and apply the most recent change per field. This is overkill for most agent memory systems but useful for multi-region deployments.
Tooling: Use pycrdt or automerge for JavaScript agents. For Python, consider crdt library (though immature).
3. Operational Transformation (OT) for Sequences
If your agent memory includes ordered sequences (e.g., conversation history), OT is the standard approach (used by Google Docs). Each operation (insert, delete, update) is transformed against concurrent operations to produce a consistent state.
Implementation: Not trivial. Use a dedicated OT library like sharedb or ot.js. This is rarely needed for typical agent recall but worth mentioning for completeness.
4. Event Sourcing with a Single Source of Truth
The most robust approach: designate one store as the authoritative source (usually Postgres) and treat all others as derived caches. Every write goes to Postgres first, then a change data capture (CDC) pipeline updates Qdrant and Neo4j asynchronously.
Implementation: Use Postgres logical replication with Debezium to stream changes to Kafka, then consume from Kafka to update vector and graph stores. This guarantees consistency at the cost of eventual consistency for reads.
# docker-compose for CDC pipeline
version: '3.8'
services:
postgres:
image: postgres:16
environment:
POSTGRES_DB: agent_memory
POSTGRES_PASSWORD: secret
debezium:
image: debezium/connect:2.5
depends_on:
- postgres
- kafka
kafka:
image: confluentinc/cp-kafka:7.6
qdrant:
image: qdrant/qdrant:v1.9
neo4j:
image: neo4j:5-communityPros: Strong consistency for writes, easy to reason about. Cons: Higher latency for derived stores, requires infrastructure.
Practical Conflict Detection
You can't resolve what you don't detect. Implement a reconciliation loop that:
- Periodically (e.g., every 5 minutes) scans a subset of entities.
- Computes a hash of the data from each store.
- If hashes differ, flags the entity for resolution.
Use a priority queue to handle high-value entities first (e.g., memories with high access frequency). For Qdrant, you can use the scroll API with a filter on updated_at to only fetch recently modified points.
# Simple hash-based conflict detection
def compute_hash(data):
import hashlib, json
return hashlib.sha256(json.dumps(data, sort_keys=True).encode()).hexdigest()
def detect_conflicts(entity_id):
pg_hash = compute_hash(fetch_postgres(entity_id))
qdrant_hash = compute_hash(fetch_qdrant(entity_id))
neo4j_hash = compute_hash(fetch_neo4j(entity_id))
if len({pg_hash, qdrant_hash, neo4j_hash}) > 1:
return True
return FalseCase Study: Agent with Multi-Store Memory
Consider an agent that helps users manage projects. It stores:
- Postgres: Project metadata (name, deadline, status).
- Qdrant: Embeddings of project descriptions for semantic search.
- Neo4j: Relationships between projects, tasks, and team members.
A user updates the project deadline. Postgres is updated immediately. The agent then queries Neo4j for all tasks related to the project and updates their deadlines as well. Meanwhile, the Qdrant embedding for the project description becomes stale because it still references the old deadline in its payload.
Conflict: The vector search might return this project as relevant for queries about "upcoming deadlines" even though the deadline has passed.
Resolution: Use event sourcing. The Postgres update triggers a CDC event. A consumer updates Neo4j (if needed) and re-embeds the project description in Qdrant. The agent always reads the deadline from Postgres, not from the Qdrant payload, to avoid stale data.
Trade-offs and Recommendations
- If you need strong consistency: Use event sourcing with CDC. Accept the complexity.
- If you need low latency and can tolerate eventual consistency: Use LWW with background reconciliation.
- If your agent operates offline: Investigate CRDTs, but be prepared for implementation headaches.
- If you have simple key-value memory: Stick with Postgres + pgvector. Avoid graph stores unless you truly need relationship traversal.
Conclusion
Conflicts between vector, graph, and relational stores are inevitable in agent memory systems. Choose a strategy that matches your consistency requirements and operational capacity. Start simple (LWW), monitor conflict rates, and evolve to event sourcing when needed. The tools are mature; the architecture is up to you.