ResearchPod Summary
Visual anomaly detection in unified settings—where a single model must handle diverse object categories—is often hindered by the high-dimensional, anisotropic nature of foundation model embeddings. Traditional Energy-Based Models (EBMs) struggle to learn stable density estimates in these spaces because standard Markov Chain Monte Carlo (MCMC) sampling fails when data features are strongly correlated. This paper investigates whether geometric reparameterization can stabilize EBM training in high-dimensional token spaces without resorting to information-losing dimensionality reduction.
The authors introduce ReFP-AD, a two-stage framework. First, they use a rectified flow to learn an optimal transport map that transforms standardized DINOv2 tokens into an isotropic, well-conditioned latent space. This transformation is guided by specific geometric diagnostics—such as anisotropy, correlation, and tail-heaviness—to ensure the resulting space is conducive to stable sampling. Second, they train an unconstrained EBM in this preconditioned space using persistent contrastive divergence with preconditioned Stochastic Gradient Langevin Dynamics (pSGLD). Anomaly scores are subsequently derived from the gradient norm of the learned energy landscape.
ReFP-AD demonstrates state-of-the-art performance on the MVTec-AD and VisA benchmarks under a strict unified protocol. It achieves 98.6% Image AUROC on MVTec-AD and 97.3% on VisA, outperforming prior unified EBM baselines by up to 10.8%. The authors show that the geometric reparameterization is critical for maintaining stable MCMC trajectories, allowing the model to leverage the full richness of foundation model tokens rather than collapsing them into low-dimensional bottlenecks.
This work provides a principled way to bridge the gap between powerful, high-dimensional foundation model representations and generative density estimation. By treating the instability of EBM training as a geometric problem rather than an architectural one, ReFP-AD offers a scalable path for unified anomaly detection that maintains high semantic resolution, which is essential for detecting subtle defects across heterogeneous industrial datasets.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.