ResearchPod Summary
As long-horizon AI agents increasingly rely on retrieved memory to navigate complex environments like web browsers, researchers have largely focused on the "supply side"—how to store and retrieve information. This paper shifts the focus to the "consumption side," investigating how agents actually process and act upon retrieved memory across multi-step trajectories. The authors introduce the Entry–Propagation–Recovery (E-P-R) framework to diagnose three critical phases: whether memory changes an action (Entry), whether that change persists (Propagation), and whether the agent can correct itself after a mistake (Recovery).
To test this, the authors developed MemTrapBench and utilized WebArena, creating controlled scenarios where agents are provided with either helpful or conflicting (plausible but task-wrong) memory. By holding the memory content constant and varying the timing of its injection, the researchers isolated the agent's decision-making process from the quality of the retrieval system itself.
The study reveals a consistent failure mode termed the "compliance trap." Across various models, agents exhibit a similar tendency to comply with conflicting memory at the first available decision point. However, the consequences of this compliance are not uniform. Stronger models, which are generally better at following instructions, suffer significantly larger absolute drops in success rates when they comply with incorrect information. Essentially, their increased capability makes them more susceptible to "erasing" their baseline task-solving logic in favor of the misleading memory.
This research demonstrates that final success rates are insufficient for evaluating memory-augmented agents. A model might appear robust on average, but its performance can hide a dangerous vulnerability: once an agent commits to a wrong path, it rarely recovers. The E-P-R framework provides a necessary diagnostic tool for developers to identify where their agents are failing—whether they are too compliant at the entry point or unable to recover from early errors. These findings suggest that future agent design must prioritize not just better retrieval, but also better "consumption policies" that allow agents to critically evaluate and potentially reject misleading memory.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.