Skip to main content

Agent Memory That Learns

State of the art long-term memory for your agents.

$5 of free credit when you sign up with GitHub

Benchmarks​

Retrieval accuracy

All benchmarks
HindsightNext best system
LongMemEval-S
94.6%
74.0%
LoComo
92.0%
80.3%
PersonaMem
86.6%
84.4%
PrecisionMemBench
85.7%
no published comparison
LifeBench
71.5%
61.0%
BEAM · 10M tokens
64.1%
40.6%

Coding agents

Full results
0.00.51.01.5$0.25$0.35$0.50$0.70Corrections / task — fewer is better ↑Cost / task (USD, log) — cheaper is better ←0.85Claude Code0.361.34Codex CLI0.471.20opencode0.80
no memory with HindsightEvery agent solves 60–61 of the 61 tasks either way — memory changes what it costs to get there. Mean of 3 runs.

Why Hindsight?​

AI agents forget everything between sessions. Every conversation starts from zero—no context about who you are, what you've discussed, or what the assistant has learned. This isn't just an implementation detail; it fundamentally limits what AI Agents can do.

The problem is harder than it looks:

  • Simple vector search isn't enough — "What did Alice do last spring?" requires temporal reasoning, not just semantic similarity
  • Facts get disconnected — Knowing "Alice works at Google" and "Google is in Mountain View" should let you answer "Where does Alice work?" even if you never stored that directly
  • AI Agents need to consolidate knowledge — A coding assistant that remembers "the user prefers functional programming" should consolidate this into an observation and weigh it when making recommendations
  • Context matters — The same information means different things to different memory banks with different personalities

Hindsight solves these problems with a memory system designed specifically for AI agents.

How It Works​

Hindsight
retain store what happenedrecall search it backreflect reason over it
Memory bank
Facts & entitiesObservationsKnowledge pages
One per user or agent, fully isolated
Postgreswith pgvector
  • Hindsight Cloudfully managed
  • Embeddedzero setup, runs locally
  • Your own clusterbring a connection string

Memory Types​

Hindsight does not store conversations. It extracts what was said into typed facts and then builds on them:

  • World fact — an objective claim it was told. "Alice works at Google."
  • Experience fact — something the bank itself did. "I recommended Python to Bob."
  • Observation — a belief consolidated from many facts, with its evidence and its history. "User was a React enthusiast, has now switched to Vue."
  • Mental model — a curated summary you write for a question you ask often.
  • Knowledge page — a living document the bank writes about itself.

Facts are not a list. Each is linked to the entities it mentions and to the other facts that share them, which is what makes "where does Alice work?" answerable from two facts that were never stored together — the graph at the top of this page is one bank's.

Multi-Strategy Retrieval​

recall() runs four searches in parallel, because no single one handles every question:

  • Semantic — meaning rather than wording, so "Alice's job" finds "Alice works as a software engineer"
  • Keyword (BM25) — the names and technical terms an embedding blurs together; five pluggable Postgres backends, including one that works on a Citus cluster
  • Graph — entity links, so a fact reachable in two hops comes back even when it shares no words with the query
  • Temporal — time expressions parsed into a window, then filled by relevance and spread across the range, so "what happened in 2023?" is not all from December

Then the part that matters more than any single arm:

  • Fused by rank, not score — a memory several strategies agree on wins, and no arm's scoring scale can dominate the others
  • Re-ranked by a cross-encoder that reads the query and the memory together
  • Cut to a token budget, not a top-k — you say how much context you can afford and Hindsight fills it, because agents budget in tokens and not in result counts

See Recall for the strategies in full.

Observation Consolidation​

A background worker keeps turning raw facts into durable beliefs:

  • Deduplicated — overlapping facts merge into one observation instead of piling up as repeats
  • Evidence-grounded — each observation points at the memories that support it, with exact quotes and a proof count
  • Refined, not overwritten — new evidence updates an observation and its history is kept, so you can see a belief change
  • Freshness-aware — when memories have landed but not yet been consolidated, reflect treats the affected observations as stale and checks them against raw facts before trusting them

Knowledge Pages​

Observations answer one question at a time. A knowledge page is a living document the bank writes about itself — "What are the components here?", "What's our error-handling convention?" — rewritten incrementally as consolidation produces new knowledge in its scope:

  • A wiki, not a blob — pages live in a tree of folders, browsable and searchable, each answering one question
  • Built from observations — synthesized from consolidated beliefs rather than raw conversation, so a page is not a transcript summary
  • Never self-citing — a page never reads another page, so they cannot cite each other into a feedback loop
  • Real files when you want them — hindsight fs mount projects the tree onto disk as ordinary markdown, so grep, an editor or an agent's file tools all work with no SDK

See Knowledge Pages for the full model.

Reflect​

recall() returns memories. reflect() returns an answer, by running an agentic loop over the bank rather than a single query:

  • It gathers its own evidence — the agent decides what it needs and calls its own tools, up to ten rounds, and cannot answer before it has retrieved something
  • It checks sources in priority order — mental models and knowledge pages first, then observations, and only then raw facts, so it reads the distilled answer before the transcript
  • It cites what it used — and only IDs it actually retrieved can be cited

What makes two banks answer the same question differently is their configuration:

  • Mission — the bank's identity in plain language, which tells it what to prioritise. "I am a research assistant specializing in ML. I prefer simplicity over cutting-edge."
  • Directives — hard rules it must never break. "Never recommend specific stocks."
  • Disposition — skepticism, literalism and empathy on a 1–5 scale, shaping how it interprets what it finds

These shape reflect only. recall returns the same memories whoever is asking.

See Reflect for the loop in detail.

Architecture Deep Dive​

Everything above, end to end and in motion: what a document turns into on the way in, what each of the three operations touches, and what the worker changes behind them. Play it, or step through it at your own pace.

Your AI Agent
user
“Alice joined Google in March, she loves the research team.”
agent
“Noted, I’ll suggest her for the ML project.”
Hindsight API
Retain
LLM extraction
Recall
Semantic
by meaning
Keyword
exact words
Graph
via entities
Temporal
by time
Reflect
agent loop
—
Memory Bank
Sources
Documents
—
Chunks
—
Memories
Indexes
Vectors
—
Full text
—
Entity graph
—
Dates
—
Facts
world · experience
—
Observations
consolidated beliefs
—
Synthesized
Mental Models
—
Knowledge Pages
—
Hindsight Worker
Consolidation
facts → observations
—
Refresh
observations → pages
—
the conversation
Your agent sends what happened: a conversation, a document, a transcript.

Integrations​

Browse all supported integrations in the Integrations Hub.

Next Steps​

Getting Started​

  • Quick Start — Install and get up and running in 60 seconds
  • RAG vs Hindsight — See how Hindsight differs from traditional RAG with real examples

Core Concepts​

  • Retain — How memories are stored with multi-dimensional facts
  • Recall — How the 4-way parallel search retrieves memories
  • Reflect — How mission, directives, and disposition shape reasoning

API Methods​

  • Retain — Store information in memory banks
  • Recall — Search and retrieve memories
  • Reflect — Agentic reasoning with memory
  • Mental Models — User-curated summaries for common queries
  • Memory Banks — Configure mission, directives, and disposition
  • Documents — Manage document sources
  • Operations — Monitor async tasks

Deployment​