What's new in Hindsight 0.10.2
Hindsight 0.10.2 lets you change a bank's id without stopping your clients, gives each tag scope its own consolidation rules, and lets each mental model pick how hard its refresh works. It also streams exports so large banks stop running out of memory, and fixes a set of hang, consistency and Oracle bugs. Everyone on 0.10.x should upgrade — see the changelog for the full list of fixes.

- Bank aliases: Give a bank a second id, move callers over at your own pace.
- Consolidation strategies per scope: Different observation rules for user, team and company scopes in one bank.
- A refresh budget per mental model: Refreshes default to a
midbudget and stop inheriting reflect defaults. - Test extraction on screenshots and files: The Extraction Tester takes images and files, and attachments now reach observations and reflect.
- Reflect can tell sources apart: A document's metadata now reaches reflect.
- Operations: Streamed exports, CUDA for ONNX embeddings, curation hooks for extensions.
- Reliability fixes: No more hung workers, safer consolidation, Oracle parity.
Bank aliases
Renaming a bank used to mean stopping every client, renaming, and repointing them all at once. Now a bank can answer to extra ids:
POST /v1/default/banks/{bank_id}/aliases {"alias": "new-id"}
GET /v1/default/banks/{bank_id}/aliases
DELETE /v1/default/banks/{bank_id}/aliases/{alias}
Adding an alias is instant and copies nothing. Both ids work, so you can move callers across a few at a time and remove the old one when nothing uses it.
Once the migration is done, you can make an alias the id the bank is shown under (PATCH .../aliases/{alias} with {"primary": true}, or hindsight bank alias primary <bank> <alias>). The real id stays visible beside it, and auth, metering, exports and logs keep using it.
Consolidation strategies per scope
One bank can now hold user, team and company memories and consolidate each under its own rules. For example, tell the company-wide scope to record only general trends while per-user scopes keep full detail:
"consolidation_strategies": [
{
"scopes": [{ "tags": ["company:*"] }],
"observations_mission": "Record only generalized, industry-level trends. Name no specific company.",
"max_observations_per_scope": 20
}
]
Set it per bank or globally with HINDSIGHT_API_CONSOLIDATION_STRATEGIES. The first strategy that matches a scope wins; anything it leaves unset falls back to the bank's own settings. The control plane has a tabbed editor under Bank → Configuration → Observations, with a preview of which existing scopes each rule would catch before you save. observation_scope_limits is deprecated but still honored.
Consolidation also now picks work fairly across scope groups, so one busy scope no longer starves the others.
A refresh budget per mental model
Mental-model refreshes were quietly running on the low budget, which halves how many steps reflect may take — too little for writing a whole page. Each mental model now has a refresh budget (low / mid / high); unset means mid. Set it in the trigger editor's Advanced panel, or as a bank default through the knowledge-page default trigger.
Refreshes also no longer read reflect_default_options. Those settings are for answering questions; tuning them used to silently change how pages were written too.
A failed refresh now settles instead of retrying in a loop and running up the LLM bill, and a model only ever runs one refresh at a time.
Test extraction on screenshots and files
The Extraction Tester in the control plane (and the dry-run extract endpoint) now takes the same mix of text, image and file blocks as retain. You can see what a bank would pull out of a screenshot or a PDF before storing anything, and each fact shows which file it came from.
Attachments also travel further: observations now show the attachments of the facts behind them, and reflect's based_on carries each cited memory with its document, tags, metadata and attachments. A screenshot retained in a KB article now comes back with the answer that relied on it.
Reflect can tell sources apart
A handbook rule and a chat message that contradicts it look the same once they're facts — and reflect tended to pick the newer one, which is usually the chat. Reflect now sees each document's own metadata, so you can state the precedence in the bank's reflect_mission (e.g. "prefer documents with source=handbook"). No new setting; banks that store no metadata pay nothing extra.
Operations
- Streamed exports. Bank and document exports are written piece by piece instead of built in memory. Exporting a bank with 256 MiB of attachments went from ~545 MiB of memory to ~25 MiB. Whole-bank export archives are now cleaned up by retention too.
- CUDA for ONNX embeddings. Set
HINDSIGHT_API_EMBEDDINGS_ONNX_DEVICE=cuda(and optionallyHINDSIGHT_API_EMBEDDINGS_ONNX_CUDA_DEVICE_ID) to run in-process embeddings on a GPU. The default stayscpu; a sample image recipe lives indocker/docker-compose/cuda-onnx/. - Curation hooks for extensions. Operation validators can now approve or reject a memory edit, invalidation or revert, and get told what was re-embedded once it commits.
- pg0 passes URL query parameters through to PostgreSQL.
- Helm Ingress now routes data API requests under
/v1correctly.
Reliability fixes
- LLM timeouts are a real deadline. A provider that trickles its response can no longer hold a worker slot forever. In-process reranking got the same wall-clock limit.
- Consolidation drops a batch whose source facts were edited mid-call, instead of writing observations from stale facts. Each observation source is also stored and shown once.
- A failed migration releases its lock, so the next start can retry instead of waiting.
- Recall on a bank deleted by another process returns 404, instead of an empty result or a 500.
- OpenAI with GPT-6 works again, including over the Responses API.
- Oracle: the remaining PostgreSQL-only retain paths, bank-archive import, and several JSON and timestamp edge cases now work.
- MCP list tools use the same paging limits as HTTP.
One schema change touches a large table: the entity name index is now scoped by bank, which makes fuzzy entity matching stay inside one bank. Migrations run on startup, so allow extra time on large deployments.
See the changelog for the full list.
