AI Agent Memory Engineering: The MEMENTO Framework
AI Agent Memory Engineering: The MEMENTO Framework
Build safer AI agent memory with MEMENTO: recall, freshness, privacy, evaluation, ownership, correction, and deletion.

Scope note: MEMENTO is an editorial design and review framework introduced in this article. It is not an official industry standard, vendor product, or guarantee that any memory store will be accurate by default.
What is AI agent memory?
A context window is a temporary working surface. Agent memory is a selective representation created from previous interactions or actions and reused later. Reliable memory is not a transcript archive and not simply a larger vector database. It requires rules for what may be stored, why it is useful, how it is retrieved, when it becomes stale, who can inspect it, and how it can be corrected or deleted.
The goal is useful continuity without turning every past interaction into permanent truth. A writing assistant may remember a verified preference for concise introductions. It usually should not preserve every draft sentence or a private detail mentioned during brainstorming.
Why long-term memory is difficult
Several recent benchmarks explore long-term conversational memory, but their datasets, tasks, and evaluation methods differ. LoCoMo tests multi-session and multimodal conversational memory. LongMemEval evaluates information extraction, multi-session reasoning, temporal updates, and abstention. BEAM studies very long conversations and long-context limitations.
These benchmarks do not prove that one architecture solves memory for every product. They do show why more context alone is not enough. A system still needs selection, provenance, freshness, contradiction handling, privacy controls, and an ability to say that evidence is missing.
The MEMENTO framework
| Layer | Question | Production output |
|---|---|---|
| M — Map | What continuity does the agent actually need? | Memory boundary and risk tier |
| E — Extract | Which signals deserve durable storage? | Candidate claims with provenance |
| M — Model | How should memories be represented? | Episodic, semantic, procedural, or profile records |
| E — Evaluate | Does memory improve the target task? | Recall, precision, freshness, latency, and abstention scores |
| N — Normalize | How are contradictions and duplicates resolved? | Canonical facts, versions, confidence, and links |
| T — Time | When should a memory weaken, expire, or require review? | Freshness policy, decay, review, and rollback |
| O — Ownership | Who can use, correct, export, or delete it? | Consent, access policy, audit trail, and deletion workflow |
M — Map the memory boundary
Divide memory into three practical categories. Working memory holds details needed only for the current task. Session memory preserves information for a short operational period. Long-term memory stores durable facts, preferences, procedures, or relationships that materially improve future work.
Map each class to a risk tier. Low-risk preferences may be saved automatically. Medium-risk business facts may need confidence thresholds and review. High-risk personal, financial, medical, identity, or access-related information should be minimized, isolated, or excluded from general-purpose memory.
This complements the AI Agent Identity and Zero-Trust guide: identity controls who may act, while memory controls what the agent is allowed to carry forward. For infrastructure and data-location trade-offs, compare the Sovereign AI Infrastructure guide.
E — Extract claims, not transcripts
A durable memory should be a compact claim with evidence rather than a raw copy of a conversation. A useful record contains the claim, source interaction, timestamp, subject, scope, confidence, sensitivity class, and update status. This lets the system distinguish what the user said, what the agent inferred, and what a tool verified.
Use a two-pass admission pipeline. First, the agent proposes candidate facts with source evidence. Second, a separate policy layer checks usefulness, explicitness, sensitivity, confidence, scope, and expected future value. If the answer is unclear, keep the information in session memory or ask for confirmation.
M — Model the right type of memory
| Memory type | Example | Typical policy |
|---|---|---|
| Episodic | The user rejected a design on Tuesday | Shorter retention and clear event date |
| Semantic | The site uses Blogger | Source, confidence, and update check |
| Procedural | An editorial checklist for publishing | Version with workflow changes |
| Profile | A verified preference for concise summaries | User-visible correction and review |
Do not force every memory into one undifferentiated vector index. Different memory types need different retention, access, ranking, and deletion policies. Indexing, retrieval, and reading are separate failure points: a fact that was never indexed cannot be fixed by better retrieval.
E — Evaluate memory with task metrics
Retrieval similarity is not memory quality. A semantically similar record may be outdated, contradictory, or irrelevant to the task. Evaluate both retrieval and the outcome:
| Metric | What it reveals | Warning signal |
|---|---|---|
| Recall | Whether relevant memories are found | Important history is missed |
| Precision | Whether retrieved memories are useful | Noise or irrelevant personalization |
| Freshness | Whether memory still reflects reality | Old preferences override new instructions |
| Abstention | Whether the agent can say “I don’t know” | Confident answers without evidence |
| Cost and latency | Whether memory is practical to run | Users cannot afford or tolerate it |
Build a private golden set with normal recall, updates, stale facts, conflicting facts, deletion requests, cross-session preferences, and cases where the correct answer is to abstain. The AI Agent Observability playbook explains how to connect memory retrieval to trace and outcome evidence.
N — Normalize contradictions
The most dangerous memory can be a true fact that is no longer true. When a new claim conflicts with an old one, compare timestamps, source quality, scope, and explicitness. Mark the old record as superseded rather than silently blending incompatible facts. Preserve a version link so an investigator can understand why the current memory exists.
If evidence remains ambiguous, keep both claims with uncertainty and ask a clarifying question. A low-confidence preference may personalize a draft, but it should not control a financial action or access decision.
T — Make time explicit
Attach a freshness policy to every memory type. A temporary project constraint may expire after a release. A travel preference may be reviewed after a year. A verified business policy may remain active until its source system changes. Decay is not only deleting old vectors; it is reducing the authority of information whose freshness has not been revalidated.
When deletion is requested, behavior depends on storage, backups, caches, logs, derived indexes, legal retention duties, and third-party services. Document what can be deleted immediately, what requires propagation, and what may require a documented retention exception.
O — Treat memory as an ownership and security surface
Persistent memory carries state across sessions, so users and teams need practical controls. They should be able to inspect important memories, correct them, request deletion, export what the system stores where feasible, and understand when a durable memory influenced an answer.
Keep memory scoped to a legitimate purpose, minimize sensitive data, restrict access by user or workspace, and connect memory events to production traces. A standardized tool connection such as Model Context Protocol (MCP) can make data available to an agent, but it does not decide which memories may cross a boundary. That policy must be enforced by a control layer that understands identity, purpose, and scope.
Do not use recalled memory as the sole basis for a high-impact decision involving health, finance, employment, legal status, identity, or access. Require verification, a second control, or human review as appropriate.
Memory admission policy
- Identify the claim. Express the proposed memory as a small, testable statement.
- Check provenance. Retain the source event, timestamp, scope, and evidence used to create the claim.
- Classify sensitivity. Apply a stricter policy or reject personal, confidential, security-related, and high-impact information.
- Score usefulness and stability. Prefer information likely to help future tasks and unlikely to become obsolete quickly.
- Admit with expiry or review. Durable preferences may be reviewed; temporary plans should expire or require confirmation.
Production memory record
| Field | Purpose | Example |
|---|---|---|
| Claim | The smallest reusable fact or preference | Prefers concise weekly summaries |
| Scope | Who, workspace, project, or agent may use it | User plus Project Alpha only |
| Provenance | Where and when the claim came from | Conversation ID, turn, timestamp |
| Freshness | When it should be reviewed or expire | Review in 30 days |
| Control | How a person can inspect, correct, or delete it | User-visible edit and delete action |
Correction, contradiction, and deletion workflow
Memory quality is defined as much by how the system forgets as by what it recalls. When a new statement conflicts with an old one, link the versions, preserve provenance, mark the older claim as superseded, and ask for clarification if the difference could change an important action. These memory events should also be connected to the AI Agent Observability workflow so retrieval, correction, and deletion can be investigated.
- Correction: update the claim after checking new evidence and scope.
- Contradiction: preserve versions and surface uncertainty instead of blending facts.
- Deletion: propagate the request across controlled stores, indexes, caches, and downstream copies where feasible.
- Incident: freeze automated writes if poisoned or private memories appear repeatedly, then review extraction and access.
After deletion, record only the minimum audit event needed to prove that the request was handled; do not retain the deleted content in the audit record.
30-day implementation plan
- Days 1–5: Map. List tasks, memory classes, risk tiers, owners, retention expectations, and information the system must never store.
- Days 6–12: Extract and model. Create a structured schema with claim, source, timestamp, scope, confidence, sensitivity, and status.
- Days 13–19: Retrieve and normalize. Add metadata filters, time-aware ranking, contradiction handling, and a clarification rule.
- Days 20–25: Evaluate. Test normal recall, updates, stale facts, conflicting facts, deletion, abstention, latency, and cost.
- Days 26–30: Govern and release. Add inspect, correction, deletion, export, audit, and rollback workflows before expanding authority.
Frequently asked questions
Is agent memory the same as RAG?
No. RAG retrieves external knowledge at query time, while memory preserves selected information from previous interactions or actions. They can share infrastructure, but memory adds ownership, updates, contradiction handling, and deletion.
Do larger context windows eliminate memory systems?
No. Larger windows help with short-term access to more text, but they do not decide what matters, what changed, who may see it, or how to correct and delete it.
Which memory should a small team implement first?
Start with narrow semantic or profile memory for low-risk, high-value preferences and constraints. Add episodic and procedural memory only when a real workflow demonstrates the need.
Should memories expire automatically?
Many should. Expiration depends on memory type, risk, source, and expected rate of change. High-value records may need revalidation rather than silent deletion.
How can an agent avoid incorrect memory?
Use conservative extraction, source evidence, timestamps, confidence, verification for sensitive facts, contradiction tests, and abstention. The agent should be able to say that a memory is uncertain or outdated.
Conclusion
The value of agent memory is not the amount of history stored. It is the quality of the claims, the discipline used to update them, and the controls that let people inspect, correct, and remove them.
Use MEMENTO as a design review: map the boundary, extract claims, model memory types, evaluate outcomes, normalize contradictions, make time explicit, and assign ownership. A small, reviewable memory system is usually safer and more useful than a larger one that nobody can explain or govern.
Sources and further reading
- LoCoMo — Evaluating Very Long-Term Conversational Memory of LLM Agents.
- LongMemEval — Benchmarking Chat Assistants on Long-Term Interactive Memory.
- BEAM — Beyond a Million Tokens.
- A Survey on Long-Term Memory Security in LLM Agents.
- Mem0: AI Agent Memory Progress Report. Vendor-produced context; interpret separately from independent benchmarks.
Join the conversation