ResearchPod Summary
As voice agents transition from simple text-based interfaces to full-duplex speech-to-speech (S2S) systems, they are increasingly deployed in enterprise customer service and as personal assistants. However, existing benchmarks often focus on narrow database manipulation or static turn-taking, failing to account for the complexity of daily life. The authors introduce DuplexWorld, a comprehensive benchmark designed to evaluate voice agents across six distinct domains: banking, insurance, travel, healthcare, logistics, and a novel pathfinding navigation task.
DuplexWorld evaluates agents using 156 authored scenarios and 11 conversation types, totaling over 350 hours of interaction. The benchmark employs a unified evaluation harness that measures performance across three pillars: conversational dynamics (e.g., turn-taking, floor management), agentic capability (e.g., task completion, tool usage), and speech naturalness. Unlike previous benchmarks, DuplexWorld includes a pathfinding world where the agent must guide a user through a synthetic urban grid, requiring the agent to manage information asymmetry and real-time spatial reasoning.
Evaluation of five leading commercial S2S systems shows that current agents are far from reliable. The authors found that conversational fluency does not guarantee task success; for instance, models that excel at natural turn-taking often fail to complete the underlying analytical tasks. Reliability is a significant issue, with Pass@3 scores (the probability of succeeding at least once in three attempts) remaining low across all models. Furthermore, the study demonstrates that agent performance varies drastically depending on the domain, even when the underlying task logic remains similar, suggesting that models are not yet robust enough for seamless integration into daily life.
This paper highlights a critical gap between the perceived conversational capability of modern voice agents and their actual utility in performing multi-step, grounded tasks. By providing a unified, multi-pillar evaluation framework, DuplexWorld sets a new standard for assessing whether voice agents can truly function as reliable, autonomous assistants in complex, real-world environments.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.