Max Ku, Nok-Kan Law, Yu-Chien Tang, Shih-Ying Yeh, Ping Nie, Andy Zheng, Tat Hei Lai, Fei-Yueh Chen, Nikko Yu, Wei-Chieh Sun, Suzy Huang, Chiao-Wei Hsu, Chih-Chuan Huang, Chak-Wing Mak, Ho Yin Sam Ng, Edisy Kin Wai Chan, Min-Hung Chen, Ho Kei Cheng
5 min
Abstract
Interactive world models are increasingly capable of generating environments and acting within them, yet deliberately editing an existing executable world remains underexplored. We formulate world editing as intervening on an existing world while preserving properties that should remain unchanged, and introduce intervention depth as an axis describing how strongly an edit couples world entities, dynamics, and systems. We instantiate this capability through industry-grade game modding and introduce IGMWorld, together with IGMBench, a benchmark of 110 tasks and over 1.1K executable state and behavioral criteria across Minecraft and Terraria. The tasks span property, entity, dynamics, and system interventions and are evaluated through deterministic executability, behavioral, preservation, and visual checks. Frontier coding agents already exhibit substantial world-editing capability: the strongest configuration solves 78.2% of tasks under a strict task-level criterion, while criterion-level performance reaches 94.8%. Reliability generally decreases with intervention depth, and this pattern persists even among tasks with similar numbers of evaluation criteria. Most failed edits still build and load successfully, suggesting that the main difficulty is making the edited world behave as requested. Visual consistency remains a separate weakness, with all evaluated configurations below 50% joint visual pass rate. These results show that world editing is a distinct capability from world generation and interaction, and that executable games provide a practical testbed for studying it.
Sam: So much of the capability sits in the debugging loop, not the generation step. That would also fit the cost picture, where agents spending more tokens on log inspection are the ones clearing the system-level tasks.
Alex: That's the pattern, though I'd read it as an association. The most successful configurations treat the game as a dynamic system to be queried, not a static codebase to be rewritten.
Sam: What about the visual side? If an agent adds a new item, does it look like it belongs?
Alex: That's a separate bottleneck. Even the strongest functional agents struggle to match the host world's style, with joint visual pass rates below 50%. They measure it with TPIPS, Text-conditioned Perceptual Image Patch Similarity, which compares generated assets against the game's existing visual style.
Sam: So it's a perceptual distance measure. Does that support saying logic and aesthetics are decoupled? Or could it just be that the visual threshold is harsher?
Alex: That's a fair caveat. The pass rates stay low even for agents that excel at the underlying logic, which is consistent with two distinct skills. But the benchmark can't fully separate that from how the visual criteria are set.
Sam: And what does the benchmark leave out?
Alex: It prioritizes executable state and behavioral criteria. Audio, animation, and narrative aren't explicitly evaluated, since they're harder to capture with deterministic checks. The evaluation is also single-trial, so it characterizes configuration-level performance but gives no clear estimate of success probability across repeated, varied runs.
Sam: That limits how far the rankings can be pushed. What do the authors suggest for extending it?
Alex: Stronger regression testing and broader runtime coverage. They're particularly interested in long-horizon and multiplayer scenarios, where interaction-dependent failures are most likely to surface. Those would test whether agents can maintain coherence over time, which is where the current evidence suggests they're weakest.
Sam: If you want the figures and the method choices we skipped, you can generate a deep dive of this paper. The paper has the rest either way.
Alex: Thanks for listening.