Agent Memory That Learns
State of the art long-term memory for your agents.
$5 of free credit when you sign up with GitHub
Benchmarks
Retrieval accuracy
All benchmarksCoding agents
Full resultsWhy Hindsight?
AI agents forget everything between sessions. Every conversation starts from zero—no context about who you are, what you've discussed, or what the assistant has learned. This isn't just an implementation detail; it fundamentally limits what AI Agents can do.
The problem is harder than it looks:
- Simple vector search isn't enough — "What did Alice do last spring?" requires temporal reasoning, not just semantic similarity
- Facts get disconnected — Knowing "Alice works at Google" and "Google is in Mountain View" should let you answer "Where does Alice work?" even if you never stored that directly
- AI Agents need to consolidate knowledge — A coding assistant that remembers "the user prefers functional programming" should consolidate this into an observation and weigh it when making recommendations
- Context matters — The same information means different things to different memory banks with different personalities
Hindsight solves these problems with a memory system designed specifically for AI agents.
How It Works
retain store what happenedrecall search it backreflect reason over itMemory Types
Hindsight does not store conversations. It extracts what was said into typed facts and then builds on them:
- World fact — an objective claim it was told. "Alice works at Google."
- Experience fact — something the bank itself did. "I recommended Python to Bob."
- Observation — a belief consolidated from many facts, with its evidence and its history. "User was a React enthusiast, has now switched to Vue."
- Mental model — a curated summary you write for a question you ask often.
- Knowledge page — a living document the bank writes about itself.
Facts are not a list. Each is linked to the entities it mentions and to the other facts that share them, which is what makes "where does Alice work?" answerable from two facts that were never stored together — the graph at the top of this page is one bank's.
Multi-Strategy Retrieval
recall() runs four searches in parallel, because no single one handles
every question:
- Semantic — meaning rather than wording, so "Alice's job" finds "Alice works as a software engineer"
- Keyword (BM25) — the names and technical terms an embedding blurs together; five pluggable Postgres backends, including one that works on a Citus cluster
- Graph — entity links, so a fact reachable in two hops comes back even when it shares no words with the query
- Temporal — time expressions parsed into a window, then filled by relevance and spread across the range, so "what happened in 2023?" is not all from December
Then the part that matters more than any single arm:
- Fused by rank, not score — a memory several strategies agree on wins, and no arm's scoring scale can dominate the others
- Re-ranked by a cross-encoder that reads the query and the memory together
- Cut to a token budget, not a top-k — you say how much context you can afford and Hindsight fills it, because agents budget in tokens and not in result counts
See Recall for the strategies in full.
Observation Consolidation
A background worker keeps turning raw facts into durable beliefs:
- Deduplicated — overlapping facts merge into one observation instead of piling up as repeats
- Evidence-grounded — each observation points at the memories that support it, with exact quotes and a proof count
- Refined, not overwritten — new evidence updates an observation and its history is kept, so you can see a belief change
- Freshness-aware — when memories have landed but not yet been consolidated,
reflecttreats the affected observations as stale and checks them against raw facts before trusting them
Knowledge Pages
Observations answer one question at a time. A knowledge page is a living document the bank writes about itself — "What are the components here?", "What's our error-handling convention?" — rewritten incrementally as consolidation produces new knowledge in its scope:
- A wiki, not a blob — pages live in a tree of folders, browsable and searchable, each answering one question
- Built from observations — synthesized from consolidated beliefs rather than raw conversation, so a page is not a transcript summary
- Never self-citing — a page never reads another page, so they cannot cite each other into a feedback loop
- Real files when you want them —
hindsight fs mountprojects the tree onto disk as ordinary markdown, sogrep, an editor or an agent's file tools all work with no SDK
See Knowledge Pages for the full model.
Reflect
recall() returns memories. reflect() returns an answer, by running an
agentic loop over the bank rather than a single query:
- It gathers its own evidence — the agent decides what it needs and calls its own tools, up to ten rounds, and cannot answer before it has retrieved something
- It checks sources in priority order — mental models and knowledge pages first, then observations, and only then raw facts, so it reads the distilled answer before the transcript
- It cites what it used — and only IDs it actually retrieved can be cited
What makes two banks answer the same question differently is their configuration:
- Mission — the bank's identity in plain language, which tells it what to prioritise. "I am a research assistant specializing in ML. I prefer simplicity over cutting-edge."
- Directives — hard rules it must never break. "Never recommend specific stocks."
- Disposition — skepticism, literalism and empathy on a 1–5 scale, shaping how it interprets what it finds
These shape reflect only. recall returns the same memories whoever is asking.
See Reflect for the loop in detail.
Architecture Deep Dive
Everything above, end to end and in motion: what a document turns into on the way in, what each of the three operations touches, and what the worker changes behind them. Play it, or step through it at your own pace.
Integrations
Browse all supported integrations in the Integrations Hub.
Next Steps
Getting Started
- Quick Start — Install and get up and running in 60 seconds
- RAG vs Hindsight — See how Hindsight differs from traditional RAG with real examples
Core Concepts
- Retain — How memories are stored with multi-dimensional facts
- Recall — How the 4-way parallel search retrieves memories
- Reflect — How mission, directives, and disposition shape reasoning
API Methods
- Retain — Store information in memory banks
- Recall — Search and retrieve memories
- Reflect — Agentic reasoning with memory
- Mental Models — User-curated summaries for common queries
- Memory Banks — Configure mission, directives, and disposition
- Documents — Manage document sources
- Operations — Monitor async tasks
Deployment
- Server Setup — Deploy with Docker Compose, Helm, or pip