Radar / Ideas / NeverAsk: agents with self-editing…

NeverAsk: agents with self-editing context and a memory that knows what to keep

daily ideamoderateJEV confidence 0.562026-10-03
OutcomeSupport and sales agents that stop re-asking customers for information they already gave, with memory costs that stay flat as history grows.

The problem

Agents re-ask for details they were already given, cite obsolete plans, and forget preferences between sessions. The workarounds are expensive: one study found Zep's memory graph consuming 600k+ tokens per conversation, and Zep memories took hours to become queryable; LangMem's search hit p95 ~60 seconds. Memory layers either bloat or go stale because every turn is persisted with hand-designed rules and nothing decides what is worth keeping.

The idea

A memory sidecar API for customer-facing agents. During a task, the agent manages its own live context CLM-style: the context is a file it can rewrite, compact, and restructure with code tools, beating hand-written summarization rules. Between sessions, a cheap decision model (Strands Decider 2B, ~115ms) answers one question per candidate fact: is this worth persisting? Only yeses are stored, so memory stays flat instead of ballooning. Each stored fact carries provenance and a TTL, and a correction endpoint lets humans fix wrong memories in one call.

Why now

The UW/Meta Context Language Models paper (arXiv 2609.37725, Sept 29) showed that models editing their own context as a file beat the strongest hand-designed context-management baselines by 11.4% accuracy with 21.5% fewer FLOPs on BrowseComp-Plus. Meanwhile the decision-model wave (Strands Decider 2B, Oct 1) makes the persist/skip judgment cost milliseconds instead of a full LLM call. Self-editing context plus a cheap write gate is a pairing nobody has productized yet.

What it combines

Context Language Models (self-editing in-task context that outperforms hand-designed compaction) + open agent memory systems like Hindsight (durable cross-session fact stores) + Strands Decider 2B (a calibrated gate deciding what is worth persisting). In-task self-management and cross-session memory are usually built as separate products; the write-gate is the missing piece that keeps the combination cheap instead of bloating to 600k tokens.

MVP

Weekend scope: REST sidecar with POST /remember (decision-model persist/skip gate, SQLite fact store with provenance + TTL) and GET /recall (session-scoped retrieval); a CLM-style system prompt template for in-context self-editing. Deliberately skip: graph relationships, entity resolution, analytics UI.

Distribution

B2B2C: embed in customer-support platforms (Intercom, Zendesk, Freshdesk) and CRM-adjacent SaaS that already own the support conversation. Price per resolved conversation; the wedge is a drop-in memory sidecar that upgrades an existing bot without re-platforming it.

Why it wins

Mem0 is a managed platform with per-token pricing and an add-only write path; Zep/Graphiti builds heavy entity graphs with hours-long availability lag; LangMem has ~60s search latency; Letta/MemGPT needs complex paging. None lets the model manage its own context, and none cheaply gates writes with a calibrated decision. NeverAsk is self-hostable, Apache-licensed components throughout, with per-fact provenance built for regulated support teams.

Risks

Stale or wrong memories will poison agent behavior and erode trust. De-risk: every fact ships with provenance (which turn created it), a TTL, and a one-call correction endpoint, and the MVP defaults to high-confidence gates so it under-remembers rather than over-remembers.

Build it with

Repo to start from

neverask: memory sidecar for support agents combining CLM-style self-editing context with decision-model write gating and provenance-tracked durable facts.

Evidence

Get the week's best AI launches, plus 3 ideas worth building

One email every Saturday. Ranked by traction, not hype. Free.