ResearchPod Summary
How can multi-turn tool-using agents effectively learn from their own execution history to improve long-horizon tool selection without suffering from training-deployment mismatches or incorrect preference signals?
The authors introduce ToolGraph, a structured orchestration layer that represents domain tools as a weighted dependency graph. This graph is constructed from tool schemas and empirical transition probabilities derived from successful agent rollouts. To further improve the agent, they implement a self-evolution pipeline using Direct Preference Optimization (DPO). A critical innovation is the construction of preference pairs at trajectory divergence points—the specific turns where a successful and a failed trajectory deviate. By ensuring these pairs are trained under the exact same ToolGraph context used during inference, the authors eliminate the prompt-mismatch issues that often plague agent fine-tuning.
ToolGraph significantly enhances performance on the tau2-bench benchmark, raising the weighted average reward from 0.304 to 0.338. Integrating DPO further boosts this to 0.355, with the most notable gains occurring in the airline and retail domains. The study also provides key diagnostic insights: in the telecom domain, performance is primarily bottlenecked by agents exhausting their step budget rather than by poor action accuracy. Furthermore, the authors identify that maintaining positive reward signals during DPO training is a more reliable indicator of model health than standard metrics like accuracy or preference margin.
This work demonstrates that structural guidance (ToolGraph) and parameter-level learning (DPO) are most effective when tightly coupled. By using an agent's own successful trajectories to build both a planning graph and a preference dataset, the system creates a self-improving loop that is robust to the complexities of multi-turn, tool-intensive environments. The findings suggest that for complex agentic tasks, the quality of the training context is just as important as the learning algorithm itself.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.