ResearchPod Summary
{ "core_finding": "Information about a verified, causally relevant internal state can be preserved in generated text even when the final answer remains unchanged, as demonstrated by a controlled proof of concept in both modular feed-forward networks and transformer models.", "caveats": "The study relies on small, purpose-built architectures with supervised intermediate states and trusted instrumentation, leaving the discovery of natural causal states in large, unconstrained language models as an open problem.", "markdown": "## Introduction and Motivation\n\nA language model's output does not automatically provide verifiable evidence regarding the internal computations that generated it. While modern systems produce complex answers, plans, and chain-of-thought reasoning, two executions can reach the exact same conclusion through entirely different internal paths. This limitation becomes critical in high-stakes auditing and scalable oversight, where an evaluator needs to know not just that an answer is acceptable, but which internal factors actually drove the behavior. This paper introduces and explores computational provenance: the question of whether generated text can carry detectable evidence of which causally relevant internal state occurred, even across executions that share the same prompt, final answer, semantic content, and sampling randomness.\n\n## Controlled Architecture and Pathway Design\n\nTo study computational provenance under controlled conditions, the author implements a specific arithmetic task across two model architectures: a small modular feed-forward network and a transformer-based model. Both architectures are trained on inputs consisting of four numbers, mandating an internal computation pathway through two discrete intermediate states, and , before producing a final modulo-8 answer . \n\nBecause of the modulo-8 operation, two different internal states that differ by 8 map to the exact same final answer. This setup allows researchers to run the same prompt twice: once naturally, and once after intentionally substituting an alternative state shifted by 8. Trusted instrumentation records these states into cryptographically protected receipts using message authentication codes (HMACs). Only after verifying the state can its value determine a subtle statistical pattern in the model's generated text.\n\n## Text Generation and Signal Detection\n\nThe generated output consists of a short textual report with a fixed numerical content but variable wording across twenty-four word positions. For each verified value of , specific wording alternatives are designated as favored and slightly biased during generation. Because the natural and alternative executions use identical random draws at each word position, any resulting difference in wording stems purely from the state-dependent preferences. A fixed statistical scoring rule evaluates the generated text against the sixteen possible patterns without requiring a trained classifier.\n\n## Main Results and Evaluation\n\nBoth the feed-forward and transformer systems successfully passed all 128 matched pairs in both their public and separately sealed protected end-to-end evaluations, with the detector reliably recovering the signal associated with the authenticated internal state. The required causal computation also reproduced consistently across five independently trained feed-forward models and three independently trained transformers. However, in a separate answer-only transformer experiment where models were trained without intermediate supervision, linear probes failed to recover naturally learned intermediate states, highlighting the challenge of applying these techniques to unconstrained models.\n\n## Key Terms and Definitions\n\n- Computational provenance — The property whereby generated text carries detectable, verifiable evidence of the specific causally relevant internal states that occurred during its production.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.