Shuoqing Deng, Gaoyue Guo, Dominykas Norgilas
11 min
This paper tackles a key challenge in constrained optimal transport: ensuring stability when we impose supermartingale constraints on couplings between probability measures on the real line. Optimal transport (OT) finds the cheapest way to move mass from one distribution μ to another ν, but real-world applications—like financial pricing—demand extra structure. Here, supermartingale optimal transport (SOT) restricts couplings π so that conditional expectations E[Y|X=x] ≤ x almost surely, reflecting scenarios where value can't increase on average (e.g., no-arbitrage pricing). The paper proves that the optimal cost functional remains continuous under perturbations of μ and ν, plus bonus results like approximation of optimal plans and a monotonicity principle. Why care? Stability is crucial for numerics, asymptotics, and approximations in finance and stochastic control, extending famous martingale OT (MOT) results where E[Y|X=x] = x.
A coupling π ∈ Π(μ,ν) is supermartingale if it disintegrates as π(dx,dy) = μ(dx) π_x(dy) with ∫ y π_x(dy) ≤ x for μ-a.e. x. Feasibility (Π_S(μ,ν) ≠ ∅) holds iff μ ≼_cd ν in decreasing convex order: ν has higher means and is "less risky" to the right, but allows mass loss to the left. This generalizes classical OT (no constraint) and MOT (equality constraint, requiring convex order μ ≼_c ν). Intuition: supermartingales model downward drifts, like asset prices under risk.
Standard OT costs c(x,y), but weak OT lets costs C(x, π_x) depend on the full conditional law π_x ∈ P_r (r-moment measures). WSOT minimizes ∫ C(x, π_x) μ(dx) over π ∈ Π_S(μ,ν). This flexibility captures path-dependent pricing in finance. The paper focuses on Polish spaces (separable complete metric) for Borel measures with finite r-moments.
Theorem 2.1 (Approximation): If (μ_k, ν_k) → (μ,ν) in W_r (Wasserstein-r distance) with μ_k ≼_cd ν_k, then for any π ∈ Π_S(μ,ν), there exist π_k ∈ Π_S(μ_k, ν_k) with AW_r(π_k, π) → 0. Here, adapted Wasserstein (AW_r) metrises couplings by matching marginals and nested W_r on conditionals—perfect for dynamic constraints.
This implies stability of V_C^S(μ,ν): the value function is continuous at (μ,ν). Proof outline: localize around barycenters, correct drifts, glue pieces, adjust marginals—handling strict supermartingale parts cleverly.
Theorem 2.6: Optimal π are C-supermartingale monotone: if π, π' optimal, their average (1-t)π + tπ' is suboptimal for t∈(0,1) unless π=π' a.e. This uniqueness-like property aids computation. Derived from stability + lower semicontinuity of C.
We investigate stability properties of weak supermartingale optimal transport (WSOT) problems on $\mathbb{R}$. For probability measures $μ,ν\in\mathcal{P}_r$ satisfying $μ\leq_{cd} ν$ (equivalently, $Π_S(μ,ν)\neq\emptyset$), we consider supermartingale couplings $π=μ(d x)π_x(d y)$ and the weak transport functional \[ V_S^C(μ,ν) := \inf_{π\inΠ_S(μ,ν)} \int_\mathbb{R} C(x,π_x)\,μ(d x), \] for some appropriate cost function $C:\mathbb{R}\times\mathcal{P}_r\to\mathbb{R}$. Our first main contribution is an approximation result in adapted Wasserstein distance: under $W_r$-convergence of marginals $(μ^k,ν^k)\to(μ,ν)$ with $μ^k\leq_{cd} ν^k$, any $π\inΠ_S(μ,ν)$ can be approximated by $π^k\inΠ_S(μ^k,ν^k)$ such that $A\mathcal{W}_r(π^k,π)\to0$. As a consequence, we obtain the continuity of the functional $(μ,ν) \mapsto V_S^C(μ,ν)$, and the monotonicity principle for WSOT.
Alex: The martingale case? That's when the average future price has to exactly match the starting price, right? Like a fair game with no drift?
Sam: Yes, martingale optimal transport enforces exact equality in those averages, used for pricing where no drift is assumed. Here, supermartingales allow the average to be less or equal, fitting downward drifts in markets. This adjustment with left mass ensures the gluing preserves that.
Alex: And that feeds into stability of the overall cost function? So prices don't jump around with noisy data?
Sam: Precisely—Theorem 2.3 shows the weak supermartingale optimal transport cost, V_C^S, is continuous in the marginals under Wasserstein convergence, as long as the cost C is continuous and convex. For minimizers, they converge too, even in adapted Wasserstein if C is strictly convex. It relies on the approximation to bound values from above and lower semicontinuity arguments for the rest.
Alex: So optimal strategies stay close even as data refines.
Sam: A byproduct is supermartingale C-monotonicity: optimal pairings sit on sets where no reshuffle of distributions—respecting mass and average constraints—lowers the total cost. It's like a no-better-swap rule, necessary and sufficient for optimality, extending to classical supermartingale transport.
Alex: That grounds hedging stability in finance, then. No discontinuous price swings from data noise.
Alex: Right, that no-better-swap idea for C-monotonicity seems key. How exactly do they prove it's necessary and sufficient for optimality?
Sam: They split the proof into sufficiency first—if a coupling follows the no-better-swap rule with respect to some set, it's optimal. For simple cases with finite lumps of mass at distinct points, they compare it to any competitor that matches the total mass and uses martingale averages exactly equal to starting points; the rule ensures the cost can't be beaten. Then, using the stability result from Theorem 2.3, they extend this to general continuous distributions by approximating with finite versions.
Alex: So the finite case acts like a building block, and stability glues it to the full picture?
Sam: Yes. They handle continuity of the cost function with tools like Lusin's theorem to approximate without losing the property. For the necessity side—that optimal couplings must obey the rule—they adapt an abstract theorem for inequality constraints, replacing equalities with "less or equal."
Alex: Inequality constraints—like the supermartingale averages being at most the starting value?
Sam: Precisely. The abstract result uses duality and measurable selection to show any optimizer concentrates on a no-better-competitor set. They define competitors that preserve mass and satisfy the inequalities via functions testing the constraints, then construct a better plan if it weren't monotone, reaching a contradiction.
Alex: Huh. That duality argument sounds like a minimax check—no room for improvement.
Sam: It is. They embed WSOT into a linear-constrained problem over sub-probabilities, proving equivalence under convexity, so the monotone property carries over.
Alex: Now, circling back to the approximation proof in Theorem 2.1, they outline reducing to the strict supermartingale part right of a cutoff point x° where excess tomorrow's mass starts. What's that strict supermartingale part?
Sam: The irreducible decomposition splits the coupling into an identical diagonal part, pure martingale lumps that can't be simplified further, and one supermartingale component left of x° with averages at most starting, and martingale to the right. They show nearby marginals μ_k, ν_k decompose similarly, converging component-wise in Wasserstein, so approximate each piece separately—diagonal trivially, martingales via prior results, and the supermartingale core with the localisation and barycentre steps we discussed. By additivity of adapted Wasserstein over disjoint parts, the whole approximates tightly.
Alex: So that decomposition lets them tackle pieces separately... but walk me through the first step in localizing the tricky supermartingale part. How do they bound it without losing the constraint?
Sam: They begin with truncation: for each starting point x, they clip the pairing kernel far out at distance R, swapping distant tails for point masses at -R and R to keep mass balanced. This creates a new coupling close to the original in adapted Wasserstein distance as R increases, and it stays supermartingale because the clips respect the average drop. Next comes affine contraction: they shrink the paired future points toward x by factor alpha less than 1, like gently pulling a rubber band back to center without snapping it. This pulls the average even lower, ensuring the inequality holds firmly. They choose a compact interval K inside the support where today's distribution has most mass. On K, the truncated-and-shrunk kernel lands in another safe compact zone L inside tomorrow's support.
Alex: Shrinking toward x... that enforces the "at most x" average by design, right?
Sam: Yes. But to match nearby noisy marginals μ_k and ν_k, they prepare an intermediate target by taking the convex hull adjustment—blending to enforce decreasing convex order, meaning tomorrow's distribution stochastically dominates today's in a way that allows safe pairings. That's like ensuring tomorrow's possible outcomes are skewed low enough overall, compared to today—one distribution beats another if every convex decreasing test function, like expected payoffs from put options, gives a lower or equal value for the first.
Alex: Then localize further: restrict to subregions left and right of K, couple the main part closely via prior lemmas, but correct kernels where averages exceed x by mixing in mass from the far left—low values that drag the mean down just enough, like adding heavy weights to one side of a seesaw.
Sam: Precisely; the mix ratio c_k fixes exactly the excess, keeping the kernel a probability and the whole sub-coupling close. For leftovers—unpaired remnants—they verify μ_k^{rem} ≤_{cd} ν_k^{rem} using put potentials, which quantify cumulative mass via integrals like expected upside from thresholds. Strassen's theorem glues a supermartingale coupling on remnants. Finally, compose with a martingale lift to hit exact ν_k, adding negligible distance since the mismatch is small.
Alex: That chains everything: truncation localizes, contraction and left mix enforce inequality, orders enable gluing, and the lift perfects marginals. A tight approximation without jumps.
Alex: For the decomposed parts, like turning noisy marginals into matching pieces—how do they do that?
Sam: They use the cumulative distribution functions—think of them as running totals of mass from left to right, like scanning a histogram to see how much probability sits below each point. For each piece in the decomposition, defined by boundaries where today's and tomorrow's masses match or differ, they integrate over those intervals to create sliced versions μ_k^n and ν_k^n that mimic the limit's structure. This preserves the convex decreasing order, ensuring safe pairings are possible, much like sorting puzzle sections by their edge shapes before approximating. For the main off-diagonal part J, they pull out η_k and v_k directly from the approximating coupling π_k, integrating kernels over J to match the leftover mass exactly.
Alex: Fair points on the limits—this is one-dimensional, relying on convergence of put potentials to glue remnants. Strict convexity needed for optimizer convergence in adapted Wasserstein.
Sam: Precisely. The paper establishes that the weak supermartingale optimal transport cost is continuous in the marginals under Wasserstein convergence, provided the cost function is continuous and convex. Optimal couplings converge too, in adapted Wasserstein distance if strict convexity holds. This completes a full stability theory for WSOT on the real line, extending martingale results. For finance, it means hedging strategies won't jump with noisy market data, supporting stable computational solvers for multidimensional WSOT and robust model-free hedging under perturbations. The paper suggests no discontinuous jumps in optimal costs from noisy distributions—a meaningful advance for pricing.
Alex: That's a solid foundation. Thanks, Sam—this clarifies how stability anchors practical use.
Sam: My pleasure, Alex. Thanks for listening to ResearchPod.