ResearchPod Summary
Standard backtesting methods for quantile forecasts often rely on fixed-horizon assumptions that are ill-suited for modern, continuous data streams where regimes shift and data drifts. Furthermore, existing tests often ignore the fact that calibration is information-dependent—a model might appear calibrated to an auditor with limited data but fail when scrutinized by an auditor with access to richer, contextual features. The authors propose a sequential, game-theoretic "testing-by-betting" framework that is anytime-valid, meaning it maintains statistical rigor even under continuous monitoring and optional stopping.
In this framework, the auditor (the "skeptic") bets on the occurrence of "hits" (whether the outcome falls below the predicted quantile). By using a predictable feature dictionary—such as calendar indicators, prices, or past forecast errors—the auditor can construct contextual bets. If the forecaster is miscalibrated, these bets accumulate wealth, providing evidence against the null hypothesis of calibration. This method requires no i.i.d. assumptions and remains valid under arbitrary, non-stationary data streams.
The study establishes a hierarchy of calibration nulls, where the coarseness of the auditor's information set determines the difficulty of the testing problem. A key finding is that validity transfers from coarser to richer information sets, but power does not; an auditor with coarse information may be "blind" to predictable, feature-specific errors that a more informed auditor could easily detect.
Empirical validation using the Chronos-2 time series forecaster demonstrates that while marginal audits (which ignore specific features) often fail to reject the null, feature-aware audits—specifically those incorporating promotion and day-of-week data—consistently detect significant miscalibration. The learned weights in the contextual betting strategy provide an interpretable diagnostic, highlighting exactly which features are associated with the forecaster's systematic errors.
As black-box models are increasingly deployed for high-stakes sequential decisions like inventory planning and financial risk management, the ability to continuously monitor their reliability is critical. This work provides a robust, interpretable, and mathematically sound way to audit these models in real-time. By allowing auditors to tailor their monitoring to the specific features relevant to their operational context, this framework helps practitioners identify not just that a model is failing, but where and why it is failing, enabling more informed interventions.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.