Interactive world models are increasingly capable of generating environments and acting within them, yet deliberately editing an existing executable world remains underexplored. We formulate world editing as intervening on an existing world while preserving properties that should remain unchanged, and introduce intervention depth as an axis describing how strongly an edit couples world entities, dynamics, and systems. We instantiate this capability through industry-grade game modding and introduce IGMWorld, together with IGMBench, a benchmark of 110 tasks and over 1.1K executable state and behavioral criteria across Minecraft and Terraria. The tasks span property, entity, dynamics, and system interventions and are evaluated through deterministic executability, behavioral, preservation, and visual checks. Frontier coding agents already exhibit substantial world-editing capability: the strongest configuration solves 78.2% of tasks under a strict task-level criterion, while criterion-level performance reaches 94.8%. Reliability generally decreases with intervention depth, and this pattern persists even among tasks with similar numbers of evaluation criteria. Most failed edits still build and load successfully, suggesting that the main difficulty is making the edited world behave as requested. Visual consistency remains a separate weakness, with all evaluated configurations below 50% joint visual pass rate. These results show that world editing is a distinct capability from world generation and interaction, and that executable games provide a practical testbed for studying it.
Alex: Frontier coding agents can now modify complex, commercial-grade game codebases to add new mechanics, with the best configuration reaching 78.2% success on strict system-level tasks. A multi-institutional team led by Max Ku built a benchmark around this, treating world editing as surgical intervention on executable logic.
Sam: That's a high number, but I'm more interested in how the agents fail. If one is asked to add something like an economy overhaul, does it fail because the code doesn't compile, or because it breaks the game's internal logic once it's running?
Alex: Mostly the second. Most failures occur after the build stage. The code compiles and loads, but the resulting world doesn't behave as requested. And the difficulty scales with what the authors call intervention depth.
Sam: So the taxonomy, L1 through L4, isn't really about how much code changes. It's about coupling. How do they measure that?
Alex: The framework is called IGMWorld, and it categorizes edits by how many systems have to coordinate. L1 is a simple parameter tweak. L4 requires synchronizing multiple interacting subsystems.
Sam: So L1 is a stitch and L4 is a transplant. An L4 edit might touch player state, UI, and game rules at once. Does the data show coupling is the real bottleneck, rather than just more criteria to satisfy?
Alex: It holds up. Even when you control for the number of evaluation criteria, deeper interventions remain significantly harder. The failure structure also shifts as you go up the levels, from simple semantic errors toward interaction failures.
Sam: That fits why behavioral failures dominate. If you add a new combat mechanic, writing the function correctly isn't enough. It has to work with the existing damage calculations and entity states.
Alex: Right, and that shows up in the metrics. Agents often satisfy many individual criteria, the criterion pass rate, but still fail the world success rate, because the edit doesn't preserve the integrity of the whole.
Sam: Is that a lack of regression testing, or are the agents struggling with state across a long horizon?
Alex: The pattern points to a structural limitation. Agents tend to treat world editing as a sequence of independent patches, but the game state is tightly coupled. A change in one subsystem triggers ripple effects the agent doesn't anticipate.
Sam: So it's close to a planning problem. Without a way to simulate the consequences, it's guessing. Do agents with better internal models of the world do better?
Alex: The paper's evidence is about process rather than internal models. Agents that run iterative edit-build-inspect loops outperform those attempting single-shot generation. Probing the runtime and adjusting based on logs appears to be the main differentiator for successful interventions.
Sam: So much of the capability sits in the debugging loop, not the generation step. That would also fit the cost picture, where agents spending more tokens on log inspection are the ones clearing the system-level tasks.
Alex: That's the pattern, though I'd read it as an association. The most successful configurations treat the game as a dynamic system to be queried, not a static codebase to be rewritten.
Sam: What about the visual side? If an agent adds a new item, does it look like it belongs?
Alex: That's a separate bottleneck. Even the strongest functional agents struggle to match the host world's style, with joint visual pass rates below 50%. They measure it with TPIPS, Text-conditioned Perceptual Image Patch Similarity, which compares generated assets against the game's existing visual style.
Sam: So it's a perceptual distance measure. Does that support saying logic and aesthetics are decoupled? Or could it just be that the visual threshold is harsher?
Alex: That's a fair caveat. The pass rates stay low even for agents that excel at the underlying logic, which is consistent with two distinct skills. But the benchmark can't fully separate that from how the visual criteria are set.
Sam: And what does the benchmark leave out?
Alex: It prioritizes executable state and behavioral criteria. Audio, animation, and narrative aren't explicitly evaluated, since they're harder to capture with deterministic checks. The evaluation is also single-trial, so it characterizes configuration-level performance but gives no clear estimate of success probability across repeated, varied runs.
Sam: That limits how far the rankings can be pushed. What do the authors suggest for extending it?
Alex: Stronger regression testing and broader runtime coverage. They're particularly interested in long-horizon and multiplayer scenarios, where interaction-dependent failures are most likely to surface. Those would test whether agents can maintain coherence over time, which is where the current evidence suggests they're weakest.
Sam: If you want the figures and the method choices we skipped, you can generate a deep dive of this paper. The paper has the rest either way.
Alex: Thanks for listening.