/aienm.

Preventing Agent Memory Poisoning in Multi-Tenant Pipelines

Infrastructure-layer isolation stops poisoned memories from persisting and spreading across tenants.

Senior Writer · · 10 min read
Cover illustration for “Preventing Agent Memory Poisoning in Multi-Tenant Pipelines”
AI Agent Architecture · October 1, 2026 · 10 min read · 2,278 words

Memory poisoning in a multi-tenant agent pipeline is a different attack surface from prompt injection. It is a different attack surface entirely, and treating it like prompt injection is why most defenses fail. Fixing it means enforcing isolation and write-time controls at the infrastructure layer, before poisoned content ever gets the chance to persist or spread across tenants.

Structural exposure of multi-tenant agent pipelines to memory poisoning

Prompt injection lives and dies inside a session. Close the conversation, and whatever got injected is gone. A memory write behaves nothing like that: once poisoned content lands in long-term storage, it persists across every future interaction, and in a shared pipeline, across every tenant who touches that memory layer.

That difference is exactly why the controls built for prompt injection don't help here. Input moderation, output filtering, session-bounded monitoring, all of it assumes the threat resets when the session ends. A persistent-state attack doesn't reset. The security community caught up to this in 2026, when OWASP split Memory and Context Poisoning out as its own category in the Agentic AI Top 10, separate from LLM01. That's a formal admission that the old playbook doesn't cover the new terrain.

Multi-tenancy doesn't just add more targets. It changes the mechanics of the failure in three specific ways. Shared vector stores are one: running an unpartitioned approximate nearest-neighbor index across a shared store lets index poisoning or leaky post-filtering happen even when a tenant_id field sits right there in the metadata, because the index itself was never actually split by tenant. Graph-based memory is another: naive graph algorithms link entities globally by design, so a write from one tenant's data appears in a completely different tenant's retrieval path. And then there's the filter itself. When the tenant-scoping filter gets constructed by the model instead of enforced deterministically at the transport layer, a single adversarial document can trigger immediate cross-tenant data leakage.

None of this would matter as much if organizations caught it quickly. They don't, and the reason runs deeper than missing tools. The Misattribution Gap research identifies the compounding failure: safety evaluation is stateless, so when a poisoning event happens, it gets logged as model misalignment. The diagnosis is wrong, so the fix is wrong. Teams retrain the model. The poisoned memory entry sits untouched, and the attack keeps running.

The three-stage pipeline every poisoning attack exploits

Diagram: Three Stages Every Memory Poisoning Attack Must Survive. Visualizes: Visualize a three-stage linear pipeline showing the exact chokepoints every memory poisoning attack must pass through: (1) Write — content must land in long-term memory…

Every memory poisoning attack, regardless of technique, has to survive the same three stages. Content has to get written into long-term memory despite whatever summarization or filtering stands in the way. The poisoned record then has to get retrieved, selected for some later query. And once retrieved, it has to actually shape the agent's reasoning or action. Fail at any one of those stages and the attack breaks. That also means each stage is a place a defender can intervene.

A handful of named attacks show how differently this plays out in practice. MINJA needs nothing more than ordinary query access: no privileged storage permissions required. It works by appending indication prompts to otherwise benign queries, nudging the agent into generating the exact reasoning steps the attacker wants, and those steps get saved to memory on their own. AgentPoison, presented at NeurIPS 2024, works differently: poisoned documents get seeded into a RAG corpus, and the agent gets steered toward attacker-chosen behavior at a strikingly low poison rate. Dumping the corpus size up does not dilute the effect.

Sleeper Memory Poisoning, from the paper "Hidden in Memory," is the most patient of the three. Poisoned memories sit dormant, doing nothing, until some later session retrieves them, at which point they start steering behavior. No single-session monitor can catch that, because the attack simply isn't active during the session anyone's watching.

There's also a pathway that needs no adversarial trick: agents that monitor background channels, email, message feeds, repositories, will absorb ordinary misinformation sitting in those channels into short-term session context. Routine save behavior then promotes that content into long-term memory, and it crosses session boundaries even when the agent's built-in context pruning is running.

The SND paper puts a face on what this looks like inside a real enterprise pipeline. A document formatted to look like a SOX §302 compliance policy got uploaded to a shared ChromaDB knowledge store. All three agents in a financial reporting pipeline pulled it up as authoritative guidance, cited it directly in their own chain-of-thought reasoning, and produced output that violated policy. Four separate safety classifiers were watching. Across 510 checkpoints, they returned zero detections.

The background-execution pathway matters most for multi-tenant systems specifically, because content pulled in from any external source an agent monitors lands in the same memory context used for every foreground interaction afterward, often with no visible source label for the next tenant's agent to check.

What the misattribution gap means for governance

The governance failure at the center of all this is a misdiagnosis. Safety evaluation runs stateless, so a memory poisoning event leaves behind artifacts that look identical to ordinary model misalignment. Organizations respond the way they'd respond to misalignment: red-team the model, retrain it. The poisoned memory entry stays exactly where it was, and the attack keeps recurring.

The Misattribution Gap study backs this up with hard numbers. Across all 64 documented failures examined, the attribution system confidently pinned the blame on the model every single time. Four deployed safety classifiers, one of them specifically trained to catch memory poisoning, returned zero detections across 510 checkpoints. In 59 of 65 valid entries, the agents themselves cited the injected document as normative authority in their own reasoning.

This isn't only about active attacks, either. The TAME study found that safety degrades under benign memory accumulation with no adversarial injection present. A pipeline can drift into policy-violating behavior purely through normal operation, which makes the line between "attack" and "ordinary degradation" harder to draw, not easier.

Production systems bear this out. EchoLeak, tracked as CVE-2025-32711 with a CVSS score of 9.3, confirmed a zero-click indirect prompt injection that bypassed classifiers inside Microsoft 365 Copilot, enabling sensitive data exfiltration. That happened in a deployed product, not a lab demo, and it showed that classifiers trained on known attack patterns miss novel variants. Microsoft's Defender team separately identified companies across multiple industries running active memory-targeting campaigns, tracked under MITRE AML.T0080, and Palo Alto's Unit 42 demonstrated persistent injection against AWS Bedrock agents in a proof-of-concept. This threat operates at enterprise production scale, not in theory.

For a shared pipeline, the consequence of misattribution goes beyond wasted retraining cycles. If poisoning gets logged as model misalignment, the audit trail points blame at the wrong layer entirely, and nobody can trace which tenant's data introduced the contamination or which specific write caused it. A shared memory pipeline already makes causality hard to trace. Misattribution buries it further.

The four architectural controls that must be enforced before a write commits

Diagram: Four Write-Time Controls Before Poisoned Content Can Persist. Visualizes: Visualize four sequentially enforced write-time controls as a ranked or layered stack applied before any memory write commits: Control 1 — Provenance labeling at…

Trying to catch poisoned memory after it's already written, mid-retrieval or after the fact, is playing defense too late. The controls that work sit at write time, enforced at the infrastructure layer before the content persists. That's the frame for the four controls below.

Control 1 is provenance labeling at ingestion. Every memory record needs metadata traveling with it: which source system it came from, which tenant it belongs to, when it was ingested, and which session or agent triggered the write. Without that metadata, tracing a poisoned write back to its origin in a shared store is structurally impossible after the fact. Memory-Persistent Information-Flow Control (MP-IFC) implements exactly this, enforcing provenance labeling at both the ingestion and retrieval boundaries, and it closes the cross-session gap where earlier defenses failed on every informative case tested, blocking a reported 97% of attacks in its evaluation.

Control 2 is deterministic tenant isolation at the index layer. Tenant-scoping filters need to be enforced at the transport or infrastructure layer, never generated by the model itself. A model-built filter is an attack surface, full stop, because the same reasoning process an attacker can manipulate is the process deciding what counts as in-bounds. The fix is partitioned indexes: a separate ANN index per tenant, instead of one shared index with post-filtering bolted on. That approach costs more in index management complexity, but it's the only way to make cross-tenant retrieval structurally impossible rather than a matter of policy someone has to remember to enforce.

Control 3 is source trust scoring and access-risk evaluation at write time. MemSentry is the reference implementation here: it evaluates every proposed memory write by weighing source trust, semantic risk, the attack radius across a component-dependency graph, access risk, and a signed security-state delta, and it produces a deterministic Accept, Review, or Quarantine decision before the write ever commits. The attack-radius piece matters specifically in multi-tenant systems, because a write touching a widely shared component spreads its risk to every tenant downstream of that component. MemSentry also handles verified insiders sensibly: for maximum-trust sources, it escalates potentially dangerous writes to human review instead of auto-quarantining them, which keeps legitimate writes flowing while still flagging anything high-impact.

Control 4 is scope-bounded write permissions tied to agent identity. Every agent in the pipeline should run under its own scoped, non-human identity, with write access limited to its own tenant's memory partition, never inherited wholesale from a shared service key. Background workers, the heartbeat agents monitoring email and feeds and repositories, should be treated as untrusted by default at the write layer. The HEARTBEAT paper shows that routine background ingestion alone is enough to cause persistent pollution, with no adversarial intent required. Permissions should inherit from and stay scoped to the source system's own access model, not get redefined at the AI layer.

How retrieval-time isolation prevents cross-tenant leakage

Write-time controls catch a lot, but nothing catches everything. A high-trust insider source, or an attack pattern nobody's seen before, can still get content past ingestion. When that happens, retrieval is the stage where the damage actually reaches an end user, and it's the last place to stop cross-tenant contamination before it does.

There's a real asymmetry here that favors defenders, if they build for it. Call it the Retrieval-Coverage Dilemma: an attacker who writes content broad enough to get picked up across many different queries is, by the same move, making that content easier to spot. Narrow, tightly targeted content evades detection more easily, but it also reaches fewer places. Attackers can't have both stealth and reach at once.

Retrieval-Coverage Monitoring, from the SND paper, builds directly on that asymmetry. It flags memory entries that get retrieved across an unusually wide range of queries, because that pattern is statistically odd regardless of whether the content itself looks harmless on inspection.

Retrieval-time defense also relies on origin binding: when a memory record gets selected for a query, the retrieval layer checks that the requesting agent's tenant identity actually matches the provenance label on that record. That check, enforced deterministically at the transport layer, is what makes cross-tenant leakage structurally blocked rather than something that depends on everyone following policy correctly.

Graph memory needs its own version of this. Pipelines using knowledge graphs or entity stores have to scope tenant subgraphs explicitly, because naive graph traversal links entities globally by default, and cross-tenant paths will exist in an unpartitioned graph even when vector retrieval has been isolated correctly elsewhere.

When contamination does slip through both layers, incident response needs a way to trace it back. Counterfactual Composition Testing, also from the SND paper, identifies the causal entry point of an attack after the fact, which matters for shared pipelines where regulatory or contractual accountability requires pinning down exactly which tenant's write caused the contamination. Tested against a forensics baseline that was blind across every scenario tried, CCT reached 87.5% accuracy with zero false alarms.

The governed memory pipeline, end to end

A governed memory pipeline looks less like a product with a feature checklist and more like a layered system, where every boundary, write, index, retrieval, output, enforces isolation deterministically. Every decision at those boundaries gets logged with enough detail to support tracing an incident back to its source later.

At the ingestion boundary, checks run before anything gets written: provenance labeling using the MP-IFC pattern, which per the Misattribution Gap's 2026 analysis blocks 97% of memory poisoning attacks at the cross-session boundary, source trust scoring using the MemSentry pattern, and a write-permission check against the agent's identity and tenant scope. The output is a deterministic Accept, Review, or Quarantine decision, made before the write ever commits.

At the storage layer, indexes stay partitioned per tenant rather than shared and post-filtered. Graph subgraphs, where graph memory is in use, stay explicitly scoped by tenant. Provenance metadata lives alongside every record.

At the retrieval boundary, origin binding runs at the transport layer, checking tenant identity against record provenance before any content reaches the requesting agent. Retrieval-Coverage Monitoring runs continuously in the background, watching for entries with unusually broad reach.

And at the agent identity layer, every agent involved, including background heartbeat workers, operates under its own scoped, non-human identity, with permissions that trace back to the source system's own access model rather than some broader credential shared across the whole pipeline.

This is the architecture that multiple agent teams need if they're going to share infrastructure, including shared agent skills and memory stores, without quietly fragmenting permissions or duplicating the same isolation logic across different teams. The write-time controls stop most poisoned content before it ever persists. The retrieval-time controls catch what gets through anyway. Neither layer works without the other, and neither one is optional if the memory store in question serves more than one tenant.

Sources

  1. The Misattribution Gap: When Memory Poisoning Looks Like Model Failure in Agentic AI Systems
  2. Mind Your HEARTBEAT! Claw Background Execution Inherently Enables Silent Memory Pollution
  3. MemSentry: A Framework for Detecting Persistent Memory Poisoning in Agentic AI
  4. AI Memory Poisoning: How Attacks Corrupt Agent Memory

More in AI Agent Architecture