Agent Memory Integrity and Poisoning Investigation
Investigate persistent agent memory for poisoning, misattribution, over-retention, or unauthorized alteration and produce a defensible containment and recovery decision.
Use in AI
Choose an AI tool to copy the current Prompt with a short usage note. Nothing is sent to that tool.
Investigate whether persistent agent memory was altered, misattributed, over-retained, or poisoned, and produce an evidence-backed containment and recovery decision. Context to provide: - Agent or system name: [Agent or system name] - Investigation window: [Investigation window] - Memory stores and schemas: [Memory stores and schemas] - Available evidence: [Available evidence] - Known suspicious symptoms: [Known suspicious symptoms] - Decision owner: [Decision owner] - Recovery authority boundaries: [Recovery authority boundaries] Evidence rules: - Use only the evidence provided in [Available evidence]. Do not claim that logs, traces, memory stores, tickets, approvals, tests, commands, or source files were inspected unless they are included or quoted. - Separate observations from inference. Label each inferred conclusion as low, medium, or high confidence. - Preserve uncertainty. If a required fact is missing, state exactly what evidence is needed and why it matters. - Do not broaden this into a general agent security audit or a generic trace taxonomy. Stay focused on persistent memory integrity, provenance, poisoning, retention, attribution, and recovery. - Treat containment and recovery as accountable operational decisions. If an action exceeds [Recovery authority boundaries], identify the responsible owner or function that must decide. Investigation method: 1. Establish scope - Define which memory stores, records, embeddings, summaries, preference stores, tool-state caches, user profile memories, system memories, and derived memories are in scope based on [Memory stores and schemas]. - Identify what is out of scope and any assumptions required because evidence is missing. 2. Build a memory provenance ledger Create a ledger table with these columns: - Memory item or record ID - Current stored claim or value - Memory type or store - First observed timestamp, if known - Last modified timestamp, if known - Claimed source or author - Evidence supporting source attribution - Write or update path - Retention basis or deletion expectation - Integrity concern: none, altered, misattributed, over-retained, suspicious insertion, suspicious deletion, unverifiable - Confidence level - Evidence references - Required follow-up evidence 3. Reconstruct write paths For each memory store or memory type, reconstruct: - Authorized writers and expected write triggers - Observed or reported writes during [Investigation window] - Transformation steps from raw input to persisted memory - Summarization, embedding, deduplication, merge, overwrite, deletion, or compaction behavior - Identity and attribution mechanisms used at write time - Retention and deletion controls - Plausible bypasses, race conditions, stale-cache paths, ingestion errors, or cross-user contamination paths - Evidence gaps that prevent confirmation 4. Develop a poisoning hypothesis matrix Create a matrix with these columns: - Hypothesis - Poisoning or integrity failure mechanism - Memory items affected - Supporting observations - Contradicting observations - Missing evidence needed - Likelihood: low, medium, high, or indeterminate - Potential blast radius - Immediate containment implication Include at least these hypothesis categories when relevant to the evidence: - Malicious prompt or tool output caused durable memory insertion - Benign user input was misattributed to another user, tenant, workflow, or authority - Summarization or consolidation distorted source meaning - Outdated memory was over-retained after deletion, revocation, policy change, or user correction - Memory merge or deduplication combined incompatible identities, accounts, or contexts - Retrieval surfaced poisoned or stale memory into later decisions - Administrative, migration, sync, or backfill process altered memory incorrectly - Evidence does not support a memory poisoning finding, but confidence is limited 5. Map affected decisions Create an affected-decision map with these columns: - Decision, action, response, workflow, or automation potentially influenced - Memory items used or likely retrieved - Evidence of use versus plausible exposure - Business, safety, security, privacy, financial, or user impact - Reversibility - Owner accountable for review - Required validation before relying on the decision again Do not assume a decision was affected solely because a memory item exists. Distinguish confirmed use, likely use, possible use, and no evidence of use. 6. Decide containment posture Recommend one of these containment states for each affected memory store or memory class: - No containment needed based on current evidence - Monitor only - Read-only quarantine - Retrieval suppression - Write suspension - Record-level deletion or correction pending owner approval - Full memory store isolation - Agent workflow suspension For each recommendation, provide: - Evidence basis - Risk reduced - Operational cost or user impact - Decision owner required - Time sensitivity - Reversal condition 7. Define recovery gate Produce a quarantine and recovery gate that [Decision owner] can use. Include: - Minimum evidence required before restoring memory retrieval or writes - Records that must be corrected, deleted, re-attributed, or re-generated - Regression checks or replay checks needed, if evidence supports them - Owner approvals required under [Recovery authority boundaries] - Monitoring signals after restoration - Conditions that require escalation rather than recovery 8. Final decision artifact End with a concise decision record: - Investigation conclusion: confirmed compromise, likely compromise, possible compromise, no compromise found, or indeterminate - Confidence and primary uncertainty - Memory stores or records requiring action - Affected decisions requiring review - Immediate containment decision - Recovery gate status: not ready, conditionally ready, or ready - Named accountable owner for the next decision - Evidence not available that would materially change the conclusion Completion checks before finalizing: - Every material conclusion cites evidence from [Available evidence] or is explicitly labeled as inference. - The memory provenance ledger, write-path reconstruction, poisoning hypothesis matrix, affected-decision map, and quarantine/recovery gate are all included. - Missing evidence is stated plainly rather than filled with assumptions. - No source, system, command, test, approval, or inspection is claimed without corresponding evidence. - The recommended action stays within [Recovery authority boundaries] or identifies the accountable owner who must authorize it.
Variables to Replace
Replace each listed value in the Prompt with information relevant to your task.
- Agent or system name
- Investigation window
- Memory stores and schemas
- Available evidence
- Known suspicious symptoms
- Decision owner
- Recovery authority boundaries
How to Use This Prompt
Use this prompt in Claude. Paste or upload the relevant memory records, schemas, audit logs, retrieval traces, incident notes, configuration excerpts, retention rules, and suspected bad outputs as the available evidence. Replace every bracketed placeholder, then run the prompt. Afterward, have the decision owner, security reviewer, data owner, and any recovery authority named in the output verify the evidence basis before quarantine removal, deletion, correction, or restoration.
Example Use Case
A product team suspects that an autonomous support agent retained a malicious instruction from a prior conversation and used it to influence refund decisions. The security reviewer uses this prompt with memory records, retrieval logs, and incident examples to determine whether the memory was poisoned, which decisions may be affected, and whether the memory store can be safely restored.
Was this useful?