ResearchPod Summary
Multimodal Large Language Models (MLLMs) have shown promise in interpreting engineering imagery, yet it remains unclear if they can perform the complex, multi-step reasoning required in professional architecture and civil engineering (AEC). Existing benchmarks often focus on simple tasks like symbol recognition or document recall, which do not test whether a model can synthesize visual evidence with domain-specific engineering principles.
To address this, the authors introduce MMArch, a benchmark featuring 1,212 short-answer items derived from peer-reviewed AEC literature. The construction process uses a decoupled planner-writer pipeline to prevent the model from tailoring questions to answers, and employs a rigorous multi-stage validation process—including blind adversarial audits and expert review—to ensure that each item requires genuine reasoning rather than textual or visual shortcuts.
The evaluation of 18 leading open-weight and proprietary MLLMs reveals a significant performance gap. While human experts achieved an accuracy of 94.6%, the top-performing proprietary model (GPT-5.5) reached only 51.7%, and the best open-source model attained approximately 30%.
Error analysis indicates that models are generally capable of locating visual evidence but fail when required to perform 'compositional reasoning'—the process of combining multiple pieces of evidence across different figures and applying a governing engineering principle to reach a conclusion. Simple prompting strategies, such as chain-of-thought, provided only marginal improvements, suggesting that the current limitation lies in fundamental domain knowledge and compositional logic rather than a lack of explicit reasoning traces.
Automating the design-analysis-review cycle in AEC could provide substantial economic value, but current AI systems remain unreliable for tasks requiring professional judgment. By isolating the specific failure modes of MLLMs in this domain, MMArch provides a diagnostic tool for researchers to improve models beyond simple pattern matching. The benchmark highlights that future progress in specialized multimodal reasoning will likely require models that can better integrate domain-specific knowledge with complex, multi-hop visual synthesis.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.