ResearchPod Summary
Agent self-evolution—the process by which an agent updates its persistent state (e.g., skills, memory, or prompts) based on experience to improve performance on future tasks—is difficult to measure reliably. Existing benchmarks often suffer from data contamination, lack of task diversity in economically valuable domains, or ambiguous train-test splits that make it unclear if performance gains are due to genuine learning or memorization.
To address this, the authors introduce GDPevo, an evolution-native benchmark focused on enterprise workflows (CRM, ERP, finance, healthcare, legal, and data-centric tasks). The core innovation is "rule hybridization," a process that decomposes complex business workflows into atomic, domain-specific rules. These rules are distributed across training tasks and recombined in held-out test tasks. This design forces agents to learn and generalize these rules rather than relying on pre-existing world knowledge.
The benchmark is generated by a fully automated pipeline that enables rapid scaling and resistance to data contamination. The pipeline follows three stages:
This automation allowed the authors to scale from 120 tasks (V1) to 240 tasks (V2) in just two days.
The authors evaluated four agents under four supervision types (no-evolution, few-shot, reflect, and self). Key findings include:
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.