ResearchPod Summary
Multi-agent large language model systems are designed to improve reasoning by dividing tasks among specialized agents, but the presence of multiple agents does not guarantee coherent reasoning or correct outputs. This study asks what discourse-level patterns distinguish accurate multi-agent decisions from inaccurate ones, and whether those insights can guide a diagnostic-to-redesign loop to improve performance.
The researchers used a five-agent Toulmin-structured multi-agent debate system for automated essay scoring on a diverse benchmark dataset. Epistemic Network Analysis was applied to model the conversational structure of the debates. Rather than just tracking output accuracy, the study coded agent turns using the Conversational Argument Coding Scheme to map how argumentative moves like assertions, justifications, and challenges co-occurred during accurate versus inaccurate scoring decisions.
Initial results revealed that accurate scoring decisions were characterized by rubric-grounded justification, agreement, and elaboration. In contrast, inaccurate decisions featured prolonged proposition-challenge-response exchanges disconnected from the rubric. Based on these insights, the researchers revised the agent prompts to enforce explicit rubric mapping and restrict unanchored score proposals. This intervention improved exact scoring accuracy from 27.78% to 40.28% and successfully aligned the discourse of inaccurate debates with the rubric-grounded pattern of accurate ones.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.