Regression discontinuity and kink designs are typically analyzed through mean effects, even when treatment changes the shape of the entire outcome distribution. To address this, we introduce distributional discontinuity designs, a framework for estimating causal effects for a scalar outcome at the boundary of a discontinuity in treatment assignment. Our estimand is the Wasserstein distance between limiting conditional outcome distributions; a single scale-interpretable measure of distribution shift. We show that this weakly bounds the average treatment effect, where equality holds if and only if the treatment effect is purely additive; thus, departure from equality measures effect heterogeneity. To further encode effect heterogeneity we show that the Wasserstein distance admits an orthogonal decomposition into squared differences in $L$-moments, thereby quantifying the contribution from location, scale, skewness, and higher-order shape components to the overall distributional distance. Next, we extend this framework to distributional kink designs by evaluating the Wasserstein derivative at a policy kink; this describes the flow of probability mass through the kink. In the case of fuzzy kink designs, we derive new identification results. Finally, we apply our methods on real data by re-analyzing two natural experiments to compare our distributional effects to traditional causal estimands.
Alex: Welcome to another episode of ResearchPod. Sam, what paper are we diving into today?
Sam: This is about a paper called "Distributional Discontinuity Design" by Kyle Schindl and Larry Wasserman. It introduces a new way to measure causal effects in studies where treatment assignment jumps at a clear cutoff point. The key idea is that traditional methods only look at average changes, which can hide important differences across the full range of outcomes.
Alex: So this paper is basically saying that averages might miss the real story when a treatment changes how outcomes spread out?
Sam: Exactly. Imagine two groups of people just on either side of a cutoff for getting a treatment—like a scholarship if your test score hits a certain mark. Researchers often just compare the average scores and might conclude there's no effect if those averages match. But the paper shows a case where the average stays the same, yet the spread of scores doubles on the treatment side—meaning some people do much better, others worse, but it balances out in the middle.
Alex: Huh. That makes sense—it's like judging a game by the score tie, ignoring that one team dominated the edges.
Sam: Right, and that's the puzzle they tackle. They propose measuring the full shift in the outcome spread using the minimal amount of "movement" needed to turn the untreated group's outcome pile into the treated one's—like the least dirt you'd shift between two sand piles to make them identical shapes. In math terms, that's the Wasserstein distance between the distributions right at the cutoff. It gives one clean number that bounds the average effect and flags when outcomes vary across the range.
Alex: So this Wasserstein number captures the full rearrangement needed, and it always stays at least as large as the simple average difference between the groups?
Sam: Yes. The average treatment effect—think of it as the straightforward gap in typical outcomes between treated and untreated at the cutoff—is always smaller in absolute size than or equal to the Wasserstein measure. Equality holds only if the treatment adds the same fixed amount to every outcome, like shifting the whole pile up by a constant without changing its shape. Otherwise, the Wasserstein picks up extra distance from variations across the range. That's because it squares those shifts at different points in the spread and averages them, so ups and downs don't cancel out.
Alex: So it breaks down the difference at each point in the spread, like from the lowest outcomes to the highest?
Sam: Exactly. Picture plotting a curve of how much the treatment changes outcomes at the bottom 10% level, the middle 50%, the top 20%, and so on—that's the quantile effect curve. The average is just the net area under that curve, where ups and downs can cancel. But the Wasserstein shows the total mismatch. You can even split it into positive and negative parts to see overall direction.
Alex: So once they've got these breakdowns, how do they actually pull the distributions from the data right at that cutoff?
Sam: They start by estimating the cumulative distribution function on each side separately—like plotting the proportion of outcomes below a certain value, using only data just left of the cutoff for the untreated limit, and just right for treated. To smooth it accurately near the edge, they fit a low-degree polynomial curve, weighted so points closer to the cutoff matter more; that's local polynomial regression. The constant term gives the limiting CDF value. From there, they invert it to get the quantile functions—what outcome sits at the 10th percentile, 50th, and so on. They add a bias correction to subtract off systematic error from the polynomial.
Alex: Weights fading away from the cutoff makes sense for local accuracy. But testing if the Wasserstein is truly zero—that quadratic nature complicates things?
Sam: Yes, near no effect, the usual methods break because it's like the square of a mean. They test by expanding the quantile process into orthogonal basis functions via Karhunen-Loève, turning the scaled squared Wasserstein into a weighted sum of chi-squares. They estimate eigenvalues of the covariance, simulate the null many times with truncation, and get critical values.
Alex: A conservative Monte-Carlo test grounded in that expansion. Practical for spotting hidden shifts averages miss.
Alex: That sounds solid for the main test. But what if samples are small—do they have a backup way to check if the Wasserstein is zero?
Sam: Yes, they offer a conservative test that errs on the safe side. It estimates just the mean and variance of the squared Wasserstein under no effect, using bounds from prior work, and sets a threshold that's guaranteed not to falsely reject too often.
Alex: So it's like building a fence around the no-effect zone that's a bit wide on purpose, to avoid calling a shift when there isn't one?
Sam: Precisely. For confidence intervals around the squared Wasserstein, they propose two conservative approaches. One uses bands around the quantile shifts, squaring the extremes to bound the integral safely. The other widens a normal-style interval with extra padding based on outcome spread. Both ensure coverage at least as good as claimed.
Alex: Okay, practical tools for sharp cases. But real policies often aren't all-or-nothing—sometimes crossing a cutoff just raises treatment odds?
Sam: That's the fuzzy case. Here, treatment probability jumps at the cutoff, but some below still get it, some above don't—like an incentive that sways most but not all. The key group is compliers: people who'd skip treatment below but take it above. They identify complier distributions by differencing weighted outcomes—treated below versus untreated above, normalized by propensity jumps—estimated via local polynomials near the cutoff. The fuzzy Wasserstein follows the same formula on those, bounding the complier average effect just like sharp designs.
Alex: So it extends cleanly, isolating true responders amid fuzziness.
Alex: That fuzzy extension handles partial compliance well. But what if there's no sharp jump in treatment probability—just a change in how steeply it rises past the cutoff?
Sam: That's a kink design. Imagine the chance of getting treatment as a line that bends at the cutoff, steeper on one side—like a road ramp that suddenly gets inclines instead of flat. Researchers measure effects from that slope change, focusing on compliers whose treatment shifts more because they're responsive to the kink. The paper defines a counterfactual where treatment nudges proportionally to each person's responsiveness, normalized so a unit shift matches the average kink size. For Wasserstein, that identifies the derivative of the distance itself at the kink—think a sensitivity measure for how it changes under small nudges.
Alex: So the counterfactual weights units by how much their treatment path bends at the cutoff, capturing who drives the observed slope jump?
Sam: Exactly. With smoothness assumptions, it decomposes into L-moment derivative differences, so you see instant changes in center, spread, or skew from the kink. Local polynomials estimate those slopes, with variance scaling accordingly. Simulations show conservative intervals narrower than bootstrap for modest samples.
Alex: So with all these tools in place—what does this look like in actual studies?
Sam: A good example comes from re-analyzing elections where the winner is decided by a tiny vote margin. The average vote share for the winner's party jumps in the next election. The Wasserstein measure comes in close, and the breakdown shows most of that from a simple location shift, with the rest from slight shape changes—plus a dominance score pointing to benefits across the outcome range.
Alex: Right, so means and full shift align there, suggesting even effects.
Sam: Contrast that with a study on Swedish grants that kick in more steeply past a certain out-migration rate. The average employment slope shows no change. Yet the Wasserstein derivative has a confidence interval touching zero, and the L-moment split reveals almost nothing from the mean—mostly higher-order terms, hinting at outliers where a few places used grants very differently.
Alex: Huh—that flags something means miss, like uneven impacts in the tails.
Sam: Exactly. Overall, this frames the Wasserstein as a benchmark that always bounds the average effect from above, equal only for pure additive shifts. When it's meaningfully larger, you know heterogeneity matters. The L-moment table acts like a dashboard, answering if effects come from center moves, wider spreads, tilts, or tails—practical for reports. While focused here on discontinuities, the logic fits randomized trials or differences over time. One limit is sticking to single outcomes; multivariate would need new identification.
Alex: Makes sense—this gives a fuller causal picture, especially for policies where spreads or skew matter as much as averages. Thanks, Sam—appreciate the clear breakdown. That's it for this look at distributional effects in discontinuity designs. Thanks for listening to ResearchPod.