Skip to main content

What's new in Hindsight 0.10.2

· 7 min read
Nicolò Boschi
Hindsight Team

Hindsight 0.10.2 lets you change a bank's id without stopping your clients, gives each tag scope its own consolidation rules, and lets each mental model pick how hard its refresh works. It also streams exports so large banks stop running out of memory, and fixes a set of hang, consistency and Oracle bugs. Everyone on 0.10.x should upgrade — see the changelog for the full list of fixes.

Hindsight 0.10.2 highlights: bank aliases, consolidation strategies per scope, a refresh budget per mental model, extraction tests on screenshots and files, operations, and reliability fixes

Bank aliases​

A bank answers to its old id and a new alias at once, so callers move over one at a time

Renaming a bank used to mean stopping every client, renaming, and repointing them all at once. Now a bank can answer to extra ids:

POST   /v1/default/banks/{bank_id}/aliases     {"alias": "new-id"}
GET /v1/default/banks/{bank_id}/aliases
DELETE /v1/default/banks/{bank_id}/aliases/{alias}

Adding an alias is instant and copies nothing. Both ids work, so you can move callers across a few at a time and remove the old one when nothing uses it.

Once the migration is done, you can make an alias the id the bank is shown under (PATCH .../aliases/{alias} with {"primary": true}, or hindsight bank alias primary <bank> <alias>). The real id stays visible beside it, and auth, metering, exports and logs keep using it.

Consolidation strategies per scope​

One fact is consolidated in full detail for a user scope and as a general trend for the company scope

One bank can now hold user, team and company memories and consolidate each under its own rules. For example, tell the company-wide scope to record only general trends while per-user scopes keep full detail:

"consolidation_strategies": [
{
"scopes": [{ "tags": ["company:*"] }],
"observations_mission": "Record only generalized, industry-level trends. Name no specific company.",
"max_observations_per_scope": 20
}
]

Set it per bank or globally with HINDSIGHT_API_CONSOLIDATION_STRATEGIES. The first strategy that matches a scope wins; anything it leaves unset falls back to the bank's own settings. The control plane has a tabbed editor under Bank → Configuration → Observations, with a preview of which existing scopes each rule would catch before you save. observation_scope_limits is deprecated but still honored.

Consolidation also now picks work fairly across scope groups, so one busy scope no longer starves the others.

A refresh budget per mental model​

A mental model&#39;s trigger sets the refresh budget, and reflect ignores the bank&#39;s reflect defaults

Mental-model refreshes were quietly running on the low budget, which halves how many steps reflect may take — too little for writing a whole page. Each mental model now has a refresh budget (low / mid / high); unset means mid. Set it in the trigger editor's Advanced panel, or as a bank default through the knowledge-page default trigger.

Refreshes also no longer read reflect_default_options. Those settings are for answering questions; tuning them used to silently change how pages were written too.

A failed refresh now settles instead of retrying in a loop and running up the LLM bill, and a model only ever runs one refresh at a time.

Test extraction on screenshots and files​

The Extraction Tester reads text and an image and shows which file each fact came from

The Extraction Tester in the control plane (and the dry-run extract endpoint) now takes the same mix of text, image and file blocks as retain. You can see what a bank would pull out of a screenshot or a PDF before storing anything, and each fact shows which file it came from.

Attachments also travel further: observations now show the attachments of the facts behind them, and reflect's based_on carries each cited memory with its document, tags, metadata and attachments. A screenshot retained in a KB article now comes back with the answer that relied on it.

Reflect can tell sources apart​

Reflect sees that one fact comes from the handbook and one from chat, and answers from the handbook

A handbook rule and a chat message that contradicts it look the same once they're facts — and reflect tended to pick the newer one, which is usually the chat. Reflect now sees each document's own metadata, so you can state the precedence in the bank's reflect_mission (e.g. "prefer documents with source=handbook"). No new setting; banks that store no metadata pay nothing extra.

Operations​

Exports read documents in batches and stream attachments in chunks, keeping peak memory near 25 MiB

  • Streamed exports. Bank and document exports are written piece by piece instead of built in memory. Exporting a bank with 256 MiB of attachments went from ~545 MiB of memory to ~25 MiB. Whole-bank export archives are now cleaned up by retention too.
  • CUDA for ONNX embeddings. Set HINDSIGHT_API_EMBEDDINGS_ONNX_DEVICE=cuda (and optionally HINDSIGHT_API_EMBEDDINGS_ONNX_CUDA_DEVICE_ID) to run in-process embeddings on a GPU. The default stays cpu; a sample image recipe lives in docker/docker-compose/cuda-onnx/.
  • Curation hooks for extensions. Operation validators can now approve or reject a memory edit, invalidation or revert, and get told what was re-embedded once it commits.
  • pg0 passes URL query parameters through to PostgreSQL.
  • Helm Ingress now routes data API requests under /v1 correctly.

Reliability fixes​

  • LLM timeouts are a real deadline. A provider that trickles its response can no longer hold a worker slot forever. In-process reranking got the same wall-clock limit.
  • Consolidation drops a batch whose source facts were edited mid-call, instead of writing observations from stale facts. Each observation source is also stored and shown once.
  • A failed migration releases its lock, so the next start can retry instead of waiting.
  • Recall on a bank deleted by another process returns 404, instead of an empty result or a 500.
  • OpenAI with GPT-6 works again, including over the Responses API.
  • Oracle: the remaining PostgreSQL-only retain paths, bank-archive import, and several JSON and timestamp edge cases now work.
  • MCP list tools use the same paging limits as HTTP.

One schema change touches a large table: the entity name index is now scoped by bank, which makes fuzzy entity matching stay inside one bank. Migrations run on startup, so allow extra time on large deployments.


See the changelog for the full list.