Unifying Time-Series, Vector, and Graph Stores for Agent Memory

A single query interface over three storage engines for autonomous agent recall

by
Unifying Time-Series, Vector, and Graph Stores for Agent Memory

When building autonomous agents that learn from their own history, one quickly encounters a fundamental challenge: agent memory is inherently multi-modal. A single agent needs to recall facts ("what is the user's preferred temperature?"), temporal patterns ("how did the user's sentiment trend over the last week?"), and relationships ("which documents are related to this query?"). Each of these queries maps naturally to a different storage engine: vector stores for semantic similarity, time-series databases for temporal aggregation, and graph databases for relationship traversal.

One could build three separate memory modules and teach the agent to route queries manually. That would be simpler in the short term, but it would force every agent developer to understand three query languages and handle cross-store joins themselves. Instead, we built a unified query interface that sits on top of all three stores, exposing a single MemoryQuery object that the agent can use without caring about the underlying engine.

The Core Architecture

The unified memory layer is a service that receives a query object and delegates to the appropriate store(s). The key insight is that most agent memory queries are actually combinations of all three modalities. For example, "find documents similar to this one that were created in the last 24 hours and are related to project X" requires vector similarity, time filtering, and graph traversal.

Our interface looks like this:

@dataclass
class MemoryQuery:
    vector: Optional[np.ndarray] = None
    time_range: Optional[Tuple[datetime, datetime]] = None
    graph_filters: Optional[Dict[str, Any]] = None
    limit: int = 10
    threshold: float = 0.7

@dataclass
class MemoryResult:
    id: str
    score: float
    timestamp: datetime
    metadata: Dict[str, Any]
    relationships: List[str]

The orchestrator parses the query and executes sub-queries in parallel where possible. For the example above, it would:

  1. Query the vector store for top-100 similar documents.
  2. Query the time-series store for document IDs created in the last 24 hours.
  3. Query the graph store for document IDs related to project X.
  4. Intersect the ID sets and re-rank by vector score.

Store Selection and Tradeoffs

We evaluated several options for each store before settling on our current stack:

  • Vector Store: We chose Qdrant because it supports payload filtering natively, which we need for hybrid queries. Indexing is HNSW with configurable parameters. For a moderate number of vectors, recall is high at reasonable query rates.

  • Time-Series Store: We use TimescaleDB (PostgreSQL extension) because it allows us to store metadata alongside time-series data and join with relational data easily. Our hypertable is chunked by time intervals, and queries over typical windows return quickly.

  • Graph Store: We picked Dgraph because of its GraphQL+- query language and native support for sharding. Depth-limited traversals with filtering complete quickly.

The Query Orchestrator

The orchestrator is stateless and runs as a FastAPI service. Here's the core dispatch logic:

class MemoryOrchestrator:
    def __init__(self, vector_client, ts_client, graph_client):
        self.vector = vector_client
        self.ts = ts_client
        self.graph = graph_client

    async def query(self, q: MemoryQuery) -> List[MemoryResult]:
        tasks = []
        if q.vector is not None:
            tasks.append(self._vector_search(q))
        if q.time_range is not None:
            tasks.append(self._time_filter(q))
        if q.graph_filters is not None:
            tasks.append(self._graph_traverse(q))

        if not tasks:
            return []

        results = await asyncio.gather(*tasks)
        
        # Merge: intersect IDs if multiple stores used
        if len(results) == 1:
            return results[0]
        
        # Build frequency map across all result sets
        id_scores = {}
        for res_set in results:
            for r in res_set:
                if r.id not in id_scores:
                    id_scores[r.id] = {'score': 0.0, 'count': 0, 'result': r}
                id_scores[r.id]['score'] += r.score
                id_scores[r.id]['count'] += 1
        
        # Keep only IDs that appear in ALL stores (AND logic)
        num_sources = len(tasks)
        filtered = [v for v in id_scores.values() if v['count'] == num_sources]
        
        # Sort by average score
        filtered.sort(key=lambda x: x['score'] / num_sources, reverse=True)
        return [v['result'] for v in filtered[:q.limit]]

This AND logic is conservative. For some use cases (e.g., "find documents similar to X OR created recently") we need OR logic. We expose a combine parameter in the query object to switch between AND and OR.

Handling Temporal Decay in Vector Similarity

One subtle issue: vector similarity scores are static, but agent memory should decay over time. A document from a long time ago should rank lower than a similar document from yesterday. We solve this by applying a time-decay multiplier to the vector score:

def _apply_time_decay(score: float, timestamp: datetime, half_life_days: float = 30.0) -> float:
    age_days = (datetime.utcnow() - timestamp).days
    decay = 2 ** (-age_days / half_life_days)
    return score * decay

We apply this after merging results from all stores. The half-life is configurable per agent.

Pure vector search often misses relationships. For example, if an agent asks "what did we discuss about project X?", a vector search might return documents that mention "project X" but miss documents that are related through a chain of references. Our graph store captures these relationships explicitly.

When a new memory is inserted, we automatically extract entities and relationships using a lightweight NLP pipeline and insert them into the graph. The graph schema is simple:

type Memory {
    memory_id: string @id
    content: string
    timestamp: datetime
    entities: [Entity] @reverse
}

type Entity {
    name: string @id
    type: string  # person, project, concept, etc.
    memories: [Memory] @reverse
}

For a query that includes graph filters, we first traverse the graph to get a candidate set of memory IDs, then run vector search only on those candidates. This reduces the vector search space and improves recall for relationship-heavy queries.

Performance Considerations

Benchmarking the unified interface against a naive approach where the agent queries each store separately and merges results in application code shows that the unified interface reduces end-to-end latency because we parallelize sub-queries and prune candidate sets early.

Query Type Naive Unified Speedup
Vector only baseline baseline 1x
Time only baseline baseline 1x
Graph only baseline baseline 1x
Vector + Time slower faster significant
Vector + Graph slower faster significant
All three slower faster significant

Lessons Learned

  1. Don't over-normalize early. We initially tried to store everything in a single PostgreSQL instance with pgvector and JSONB. It worked for a small number of memories but fell apart at scale. Specialized stores exist for a reason.

  2. Consistency is eventual. We use a write-ahead log to ensure all three stores receive the same data, but reads may see stale data for a short period. For agent memory, this is acceptable.

  3. Filter pushdown is critical. Our first version fetched all IDs from the graph store and then passed them as a filter to the vector store. That killed performance. Now we push graph filters down into the vector store's payload filter (Qdrant supports this natively).

  4. Time-series compression matters. Storing raw memory access logs with appropriate chunking and default compression reduces storage without affecting query speed.

Future Work

We're currently adding a caching layer that remembers frequent query patterns. If an agent asks "what did we discuss about project X?" every morning, the cache returns the result without hitting all three stores. The cache is invalidated when new memories are inserted that match the query's time range or graph filters.

We're also experimenting with learned query routing: a small neural network predicts which stores to query based on the query embedding, reducing latency further for queries that only need one or two stores.

The unified memory layer is open-source. Contributions welcome — especially if you have ideas for better merge strategies or additional store backends.

#agent-memory#graph-database#memory-architecture#time-series#unified-query#vector-store
Share — X / Twitter · LinkedIn · HN · Email
Damir Radulić
Founder of RiNET. On the Croatian internet since 1996 (Kvarner Net). In Amsterdam now, building autonomous AI infrastructure that runs on Monday morning when nobody's watching — sovereign stacks, agent swarms, LoRA fine-tuning, civic-intelligence platforms.

Related