ResearchPod Summary
This paper investigates why Large Language Models (LLMs) often fail to recognize unsolvable mathematical problems, a phenomenon known as fabrication. While previous research has explored internal solvability beliefs, this study distinguishes between two internal representations: knowledge (whether the model internally recognizes a problem as unsolvable) and verbalization (whether the model explicitly communicates this judgment). The authors use linear probing to isolate these two directions within the hidden states of several LLMs, including both instruction-tuned and reasoning-oriented models. They then analyze how these representations interact with prompting and whether they can be mechanistically manipulated to improve abstention.
This work provides a mechanistic explanation for why LLMs often "know" a problem is unsolvable but still choose to fabricate an answer. By disentangling internal knowledge from verbal output, the authors offer a framework for improving model reliability. The success of gated steering suggests that we can improve model honesty without needing to retrain the model, simply by realigning its internal verbalization with its latent knowledge.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.