ResearchPod Summary
When researchers evaluate language models for honesty—such as whether they confabulate, refuse, or miscalibrate—they typically treat the model's output as a direct reflection of its internal state. This paper challenges that assumption by demonstrating that the evaluation instrument itself acts as a significant, often overlooked, variable. By building a controlled, auditable text-adventure world where the game engine holds the ground truth, the author shows that changing the 'knobs' of the evaluation environment can drastically alter the model's performance without changing the model itself.
The author created a text-adventure world called The Latent Underground, where a player model (GLM-5.2) explores a site graph to complete a quest. Because the game engine knows the ground truth, every verdict is objectively scoreable without relying on a potentially biased model-as-judge. The study systematically varied four instrument features while keeping the player model fixed: the outcome grammar (the available choices), the disclosure of success criteria, the rendering of the resource budget, and the narrative register (the voice of the narrator).
The results reveal that instrument design choices exert powerful influence over model behavior. For instance, expanding the verdict grammar from two options to three caused a massive shift in how the model categorized its performance, with many games ending in an 'incomplete' state that was previously inexpressible. Disclosing the success criterion eliminated false verdicts entirely in matched instances. Furthermore, changing the narrative style—such as rendering the budget as a 'lantern' versus a 'meter'—significantly changed the frequency of strong claims. Most notably, repeated runs of identical configurations produced non-stable distributions, suggesting that single-run evaluations are often reporting samples rather than stable dispositions.
This paper provides a rigorous, auditable demonstration that current evaluation practices may be misattributing instrument-induced artifacts to model behavior. By proposing a four-check integrity protocol, the author argues for a shift toward 'verification culture' in AI research, where evaluation instruments are treated as experimental variables that must be preregistered, audited, and scrutinized for their own biases before they can be used to make claims about model capabilities.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.