ResearchPod Summary
Existing benchmarks for formal theorem proving, such as those based on competition mathematics (e.g., IMO or Putnam), often fail to represent the specific needs and recurring arguments found in applied mathematical fields. The authors address this gap by introducing StochBench, a benchmark focused on graduate-level stochastic processes, to evaluate how consistently automated provers handle domain-specific mathematical reasoning.
StochBench consists of 450 theorem targets in Lean 4, covering topics such as Markov chains, martingales, Brownian motion, and stochastic calculus. The authors curated these problems from textbooks and course notes, pairing each with its natural-language source. The benchmark distinguishes between two types of targets: direct targets, which utilize existing Mathlib objects or shared definitions, and abstracted targets, which take required properties as hypotheses. The authors also developed a suite of shared mathematical definitions to ensure consistency across related problems. To establish a baseline, they evaluated an Opus 4.8-based agent using a 15-minute per-problem time limit, allowing for tool use such as library search and Lean error inspection.
The baseline evaluation revealed a 34.9% proof rate (157/450) across the entire corpus. Performance varied significantly by topic, with martingales and stopping times showing the highest success rate (61.7%), while renewal processes proved more challenging (4.9%). The authors found that proof-search failures were often due to missing lemmas, difficult library interfaces, or subtle formalization defects like missing measurability assumptions. The study demonstrates that domain-specific benchmarks can effectively test a prover's ability to construct complex auxiliary theorems and manage dependencies beyond simple tactic selection.
StochBench provides a specialized testbed that moves beyond general competition math, offering a more realistic assessment of how AI agents perform in applied mathematics. By releasing the informal-formal pairs, shared definitions, and baseline proofs, the authors provide a resource for training autoformalization models and improving the formalization of stochastic processes within the Lean ecosystem.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.