ResearchPod Summary
Large Language Models (LLMs) often produce correct final answers through flawed or unfaithful reasoning, a phenomenon known as a "lucky guess." Existing methods like Self-Consistency (SC) rely on majority voting of final answers, which fails to detect these logical inconsistencies. This paper asks: How can we reliably quantify uncertainty in LLM reasoning, filter out unfaithful "lucky guesses," and assess the robustness of reasoning topologies?
The authors propose GRAPHEVAL, a framework that treats LLM reasoning as a graph-based structure. First, they use a Decomposer LLM to convert raw text reasoning into a Directed Acyclic Graph (DAG) of atomic facts and causal dependencies. They then introduce the Graph Reasoning Coherence Score (GRCS), a metric based on Semantic-Structural Graph Edit Distance (SS-GED) that measures the consensus within a model's reasoning manifold. Finally, they introduce Graph Self-Consistency (GSC), a decoding strategy that selects the "medoid" reasoning path—the most central and supported path in the graph—rather than simply picking the most frequent final answer.
This work shifts the evaluation of LLM trustworthiness from final-answer accuracy to holistic reasoning fidelity. By providing a way to quantify when a model is "guessing" versus "reasoning," GRAPHEVAL offers a path toward more reliable deployment of LLMs in mission-critical environments where the logical process is as important as the final result.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.