Rulin Shao, Shannon Zejiang Shen, Junjie Oscar Yin, Yuetai Li, Minheng Wang, Hamish Ivison, Radha Poovendran, Nathan Lambert, Teng Xiao, Mike Lewis, Wen-tau Yih, Luke Zettlemoyer, Pang Wei Koh
4 min
Context Language Models (CLMs) shift the paradigm of context management from external, harness-defined rules to intrinsic model behavior. By treating the context as a file, the model gains the ability to perform unrestricted edits—such as compacting information, deleting irrelevant data, or maintaining persistent trackers—using standard programming tools. This approach allows the model to learn and adapt its context-management strategy dynamically, whether through in-context instructions, textual skill evolution, or reinforcement learning.
CLMs demonstrate superior performance and efficiency compared to existing methods like summarization-based compaction or tool-based memory management. On the BrowseComp-Plus deep-research benchmark, CLMs achieved 11.4% higher accuracy while using 21.5% fewer FLOPs. In multi-repository agent-swarm tasks, CLMs provided a 65% greater end-to-end speedup compared to summary-based baselines. Furthermore, the researchers introduced Suffix Cache Reuse (SCR), a serving optimization that reuses cached states for unchanged context segments after an edit, further reducing server-side compute by 35%.
Because context management is now an intrinsic capability, CLMs can be steered using natural language. Users can provide simple instructions to guide compaction timing or semantic boundaries. Moreover, the authors demonstrate that CLMs can evolve their own context-management skills through a standard optimization loop or internalize these strategies via reinforcement learning. This allows the model to move beyond human-defined heuristics to discover more effective, task-specific ways to manage its own memory.
This work suggests that context management should not be an external constraint but a core competency of intelligent agents. By letting models decide what to keep and what to discard, we can build agents that are better at long-horizon reasoning and more efficient in their resource usage. This shift aligns with the broader trend of moving away from hand-engineered harnesses toward models that learn to optimize their own execution.
We introduce Context Language Models (CLMs), language models that natively manage their own context. We implement this by treating the context as a file and allowing the model to make unrestricted updates to this file. This allows the model to learn what is most important to maintain in context, and naturally extends to multi-agent systems where multiple agent contexts coexist as files. Building CLMs zero-shot with existing models outperforms SOTA context management strategies across a variety of tasks: 11.4% higher accuracy with 21.5% fewer FLOPs on BrowseComp-Plus, 5% higher scores with 59% fewer FLOPs on 12-hour EdgeBench, and 65% greater improvement with the same compute on a 24-hour multi-repository agent-swarm task. Moreover, by shifting context management from external harness control to intrinsic model behavior, CLMs naturally enable both in-context and parametric learning of context-management strategies. We show that CLMs can be steered with natural-language instructions evolved through a standard skill-optimization loop, improving held-out accuracy by up to 35.9 points on a context-management task while reducing compute. We also introduce an online reinforcement learning method for CLMs, improving Qwen3.5-9B performance on BrowseComp-Plus by 47.6% while using 12% fewer FLOPs. Finally, we co-design Suffix Cache Reuse for CLM serving, further reducing server-side compute by 35% relative to standard SGLang at matched performance.
Alex: [probing] Why would editing improve accuracy at all, rather than just cutting cost?
Sam: [measured] The paper's explanation is that the model learns to shape its own input stream, keeping what's relevant and discarding the rest. A cleaner context plausibly helps reasoning on long-horizon tasks, where accumulated noise competes with the signal. I'd treat that as the authors' interpretation. The headline numbers show the gain, not why it happens, and without the baselines and ablations in front of me I can't say how much comes from editing versus other factors.
Alex: [thoughtful] And the obvious failure mode: the model makes a bad edit and deletes something it needed later. [[RP_SECTION:training-and-safety-considerations|Training and Safety Considerations]]
Sam: [measured] The authors address that in training with what they call a success-gated efficiency advantage. Trajectories are rewarded for being efficient only if they're also correct. A policy can't profit from aggressive pruning that causes failure, which is the reward-hacking route you'd otherwise worry about.
Alex: [analytical] That's a training-time incentive, though. It doesn't guarantee any single deletion is safe.
Sam: [nodding] Right, and that's the limitation I'd press on. The gating discourages destructive edits statistically, across trajectories. Whether the model can recover from a bad deletion in a given run isn't something I can speak to from what I've read. Nor can I say how well the gains transfer beyond the deep-research benchmarks.
Alex: [reflective] So the evidence supports treating context as a controllable variable on these tasks, with the caching fix making it affordable. The safety margin of individual edits is still an open question.
Sam: [measured] That's a fair summary. The accuracy and compute numbers carry the claim, SCR makes it deployable, and the reward design is the main guard against misuse of the edit power.
Alex: [calm] If you want the figures and the method choices we skipped, you can generate a deep dive of this paper. The paper has the rest either way.
Sam: Thanks for listening.