ResearchPod Summary
Creating and editing presentation slides is a ubiquitous professional task that requires reasoning over multimodal content, including text, images, tables, and complex layouts. While general-purpose computer-use agents are advancing, existing benchmarks often lack the depth required to evaluate realistic, iterative slide editing. PPT-Eval addresses this gap by providing a comprehensive testbed for GUI-based interaction with PowerPoint Online.
PPT-Eval consists of 120 tasks across 12 diverse PowerPoint files, categorized by difficulty (easy, medium, and hard). Unlike API-bound benchmarks that limit agents to programmatic manipulation, PPT-Eval uses a sandboxed GUI environment, allowing agents to utilize the full feature set of PowerPoint, such as design tools, animations, and transitions. To handle the open-ended nature of slide editing, the authors developed a tree-structured rubric system that provides partial credit for intermediate steps, penalizes extraneous changes, and generates natural language feedback. This evaluation framework shows a strong correlation (Kendall’s τ-b = 0.77) with human expert judgments.
Experimental results indicate that current frontier models, such as Claude-4.5-Opus, achieve only a 45% success rate and an average partial score of 57% on the benchmark. These models significantly lag behind the human baseline (80% success rate). Interestingly, while API-based agents currently outperform GUI-based agents, they struggle with tasks requiring native GUI features like advanced design layouts or specific animation triggers, highlighting the need for improved GUI-based reasoning capabilities.
PPT-Eval provides a rigorous, interpretable, and scalable way to measure the progress of computer-use agents in a high-stakes professional domain. By moving beyond binary success metrics to nuanced partial credit, the benchmark offers a clearer picture of where agents fail and where they show promise, ultimately helping researchers identify the specific bottlenecks in current agentic workflows.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.