ResearchPod Summary
Memory-augmented LLM agents are increasingly vulnerable to memory injection attacks, where adversaries plant malicious records that divert the agent's behavior from the user's original intent. Existing defenses either rely on expensive, repeated LLM auditing or struggle with information redundancy in multi-turn contexts, where task-irrelevant data obscures attack signals. This paper asks: Can we develop a lightweight, effective defense that filters malicious memories by focusing on the relationship between initial user intent and agent behavior across turns?
To address this, the authors propose Memory Intent-Aware Neural Denoising (MIND). The method is built on the observation that benign and poisoned trajectories show distinct patterns in how the agent's behavior relates to the initial user intent. MIND employs an Information Bottleneck (IB) to compress interaction trajectories into a latent space, effectively filtering out task-irrelevant and repetitive information while preserving intent-relevant signals. A lightweight, multi-hyperplane classifier then identifies and rejects malicious memories based on these denoised representations. This approach avoids the high latency of LLM-based reasoning while maintaining better performance than simple record-level filters.
Extensive experiments across multiple LLM backbones (including DeepSeek-V4 and GPT-4o-mini) demonstrate that MIND significantly improves the security-utility trade-off. On the ReAct-StrategyQA benchmark, MIND reduced the mean attack success rate (ASR) by over 55% compared to undefended agents, while maintaining task accuracy and inference latency comparable to those of undefended systems. The results suggest that by focusing on intent-behavior alignment, the model can effectively detect malicious injections even in long-horizon tasks where traditional methods fail due to information noise.
As LLM agents are increasingly deployed for long-horizon tasks like scientific research and software engineering, they rely heavily on external memory systems. This makes them prime targets for memory injection. MIND provides a scalable, efficient defense mechanism that does not require the heavy computational cost of repeatedly querying an LLM to audit every retrieved memory, making it a practical solution for real-world agentic systems.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.