A significant gap exists between theory and practice in deep learning. Generalization and approximation error bounds are often derived for simplified models or are too loose to be informative. Many rely on the manifold hypothesis and on geometric regularity such as intrinsic dimension, curvature, and reach. Progress requires insight into data-manifold geometry and suitable benchmarks, yet existing options are polarized: analytic manifolds with known geometry but limited applicability, or real-world datasets where geometry is only coarsely estimable. We introduce a benchmarking framework for studying data geometry. We repurpose and extend dSprites and COIL-20 with additional transformation dimensions and dense, axis-aligned sampling, and pair them with finite-difference estimators that recover curvature, reach, and volume at near-ground-truth accuracy in a regime where general-purpose estimators are unreliable or difficult to deploy. The framework is intended as a controlled testbed, useful as a calibration environment for geometric estimators and a sandbox for probing theoretical assumptions. To illustrate its use, we present two application studies, namely assessing the scaling behavior of the bounds of Genovese et al. and Fefferman et al., and tracking the layer-wise geometry of a $β$-VAE, highlighting the behavior of current bounds and the value of controlled benchmarks for guiding and validating future theory. A reference implementation is available at https://github.com/koulakis/manifold-microscope.
Alex: Welcome to another episode of ResearchPod. Today, we're looking at a new paper that tries to solve a persistent, invisible problem in deep learning. Sam, what exactly are we talking about?
Sam: We're discussing a new benchmarking framework for data geometry. The puzzle is this: while modern AI models are genuinely successful, we still don't understand how they "fit" the data they learn from. The theoretical guarantees we use to trust these models rely on geometric constants — specific measurements about the shape of data — that are simply impossible to calculate from real-world datasets. That leaves a significant gap between the math on paper and what actually happens in practice.
Alex: So if we can't measure the geometry of real-world data, how can we ever prove that our AI models are working the way we think they are?
Sam: Exactly. Real-world datasets are messy, and that messiness hides their own structure. To address this, the authors built what you might call a "geometric microscope" — a controlled testbed using synthetic datasets, meaning data they constructed themselves, where they know the exact shape in advance. Because they designed the data, they can calculate its geometric properties with high precision, and then check whether the tools researchers normally use actually give the right answers.
Alex: That's a clever move. Build something you already know the answer to, and use it to test whether your measuring tools are trustworthy. But you mentioned "geometric properties" — what specifically are they measuring?
Sam: The two key ones are curvature and reach. Imagine a complex sculpture hidden in a dark room. Curvature tells you how sharply the surface of that sculpture bends at any given point — a flat wall has no curvature, but a tight corner has a lot. Reach is a bit more subtle. It's a measure of how thin or fragile the shape is — specifically, how close any part of the surface comes to folding back on itself. A shape with low reach is either very thin, or has parts that nearly touch each other, which makes it much harder to analyze accurately.
Alex: So if an AI is trying to learn the pattern hidden in that sculpture, it needs to know both of those things to avoid making mistakes?
Sam: Precisely. If the model doesn't account for how curved the data is, it might try to approximate a sharp bend with a straight line — which introduces errors. The authors address this by using what are called finite-difference estimators. Think of these as tools that figure out the steepness of a curve by comparing tiny neighboring points on the grid, the same way you might estimate the slope of a hill by measuring how much the ground rises over a very short distance. That approach turns abstract geometric properties into something you can actually compute and verify.
Alex: And the framework itself acts as a calibration tool — like a ruler you know is accurate, so you can use it to check whether your other, less certain instruments are giving you the right readings?
Sam: That's a perfect way to put it. The authors use this known-geometry testbed to do two things. First, they check whether the theoretical error bounds from previous research actually hold up when you run real experiments. Second, they use it to track how a specific type of AI model — called a beta-VAE — reshapes the geometry of data as it passes through successive layers of its internal network.
Alex: So instead of just guessing why a model performs well or poorly, you can actually watch the geometric transformations happening layer by layer. What does that reveal?
Sam: It reveals whether the model is doing what the theory predicts. If the geometry changes in ways that match the theoretical expectations, that's evidence the theory is sound. If it doesn't, that's a signal the theory needs revision. Either outcome is useful — it's the difference between having a map you can trust and navigating blind.
Alex: That makes sense. Now, every tool has limits. What are the boundaries of this approach?
Sam: The framework currently works well in low-dimensional spaces — roughly where the intrinsic dimension of the data is around four or five. When the data becomes higher-dimensional or has irregular shapes that don't fit clean patterns, the approach struggles. There's also an issue the authors call rasterization artifacts. Think of trying to measure a smooth curve on a screen made of large, blocky pixels. If your measurement grid is too coarse relative to the underlying shape, you end up measuring the jagged edges of the pixels rather than the smooth curve beneath. That creates noise that obscures the true geometry.
Alex: It's like trying to draw a circle using square tiles — at some point, you're just seeing the tiles. So what does the future of this work look like?
Sam: The logical next step is extending the framework beyond flat, standard surfaces. Future versions could handle what are called non-Euclidean topologies — shapes that curve or twist in ways a standard grid can't easily capture, like the surface of a sphere or a torus. That would open the door to what the authors describe as geometric debugging for much larger, more powerful models — the kind used today for generating images or text.
Alex: So the trajectory is from a precise but limited calibration tool toward something that can handle the full complexity of modern AI. And the underlying value is transparency — being able to see, rather than assume, how a neural network is processing information.
Sam: That's exactly it. By turning abstract geometric theory into concrete, reproducible experiments, this framework gives researchers a way to verify their assumptions rather than simply hoping they hold. That kind of rigorous foundation is what the field needs as these models grow more powerful and more consequential.
Alex: A clearer view inside the black box. Thanks for walking me through it, Sam. And thanks to everyone listening — this has been ResearchPod.