NeverAsk: agents with self-editing context and a memory that knows what to keep
The problem
Agents re-ask for details they were already given, cite obsolete plans, and forget preferences between sessions. The workarounds are expensive: one study found Zep's memory graph consuming 600k+ tokens per conversation, and Zep memories took hours to become queryable; LangMem's search hit p95 ~60 seconds. Memory layers either bloat or go stale because every turn is persisted with hand-designed rules and nothing decides what is worth keeping.
The idea
Why now
The UW/Meta Context Language Models paper (arXiv 2609.37725, Sept 29) showed that models editing their own context as a file beat the strongest hand-designed context-management baselines by 11.4% accuracy with 21.5% fewer FLOPs on BrowseComp-Plus. Meanwhile the decision-model wave (Strands Decider 2B, Oct 1) makes the persist/skip judgment cost milliseconds instead of a full LLM call. Self-editing context plus a cheap write gate is a pairing nobody has productized yet.
What it combines
Context Language Models (self-editing in-task context that outperforms hand-designed compaction) + open agent memory systems like Hindsight (durable cross-session fact stores) + Strands Decider 2B (a calibrated gate deciding what is worth persisting). In-task self-management and cross-session memory are usually built as separate products; the write-gate is the missing piece that keeps the combination cheap instead of bloating to 600k tokens.
MVP
Weekend scope: REST sidecar with POST /remember (decision-model persist/skip gate, SQLite fact store with provenance + TTL) and GET /recall (session-scoped retrieval); a CLM-style system prompt template for in-context self-editing. Deliberately skip: graph relationships, entity resolution, analytics UI.
Distribution
B2B2C: embed in customer-support platforms (Intercom, Zendesk, Freshdesk) and CRM-adjacent SaaS that already own the support conversation. Price per resolved conversation; the wedge is a drop-in memory sidecar that upgrades an existing bot without re-platforming it.
Why it wins
Mem0 is a managed platform with per-token pricing and an add-only write path; Zep/Graphiti builds heavy entity graphs with hours-long availability lag; LangMem has ~60s search latency; Letta/MemGPT needs complex paging. None lets the model manage its own context, and none cheaply gates writes with a calibrated decision. NeverAsk is self-hostable, Apache-licensed components throughout, with per-fact provenance built for regulated support teams.
Risks
Stale or wrong memories will poison agent behavior and erode trust. De-risk: every fact ships with provenance (which turn created it), a TTL, and a one-call correction endpoint, and the MVP defaults to high-confidence gates so it under-remembers rather than over-remembers.
Build it with
- Context Language Models (Meta)In-task context self-management: model edits its own context as a file
- Hindsight: open-source agent memory with independently reproduced SOTAOpen-source agent memory with reproducible fact extraction
- Strands Decider 2B~115ms calibrated persist/skip gate keeping memory flat
- Perplexity pplx-decider-v1-27bAlternative larger decision model for higher-stakes write judgments
Repo to start from
neverask: memory sidecar for support agents combining CLM-style self-editing context with decision-model write gating and provenance-tracked durable facts.
Evidence
- UW-Meta paper: language models that edit their own context (AI Weekly)
- Context Language Models: what the paper changes for agent builders (dev.to)
- Agent memory research: chunking and caching comparison (GitHub, huyxdang/agent-memory)
- AI Agent Memory: Synap's promise, tested (Volanea)
- Mem0 paper (arXiv 2504.19413)
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.