❌

Normal view

Sliced Wasserstein Regression

1 January 2026 at 00:00
While statistical modeling of distributional data has gained increased attention, the case of multivariate distributions has been somewhat neglected despite its relevance in various applications. This is because the Wasserstein distance, commonly used in distributional data analysis, poses challenges for multivariate distributions. A promising alternative is the sliced Wasserstein distance, which offers a computationally simpler solution. We propose distributional regression models with multivariate distributions as responses paired with Euclidean vector predictors. The foundation of our methodology is a slicing transform from the multivariate distribution space to the sliced distribution space for which we establish a theoretical framework, with the Radon transform as a prominent example. We introduce and study the asymptotic properties of sample-based estimators for two regression approaches, one based on utilizing the sliced Wasserstein distance directly in the multivariate distribution space, and a second approach based on a new slice-wise distance, employing a univariate distribution regression for each slice. Both global and local Fréchet regression methods are deployed for these approaches and illustrated in simulations and through applications. These include the joint distribution of excess winter death rates and winter temperature anomalies in European countries as a function of base winter temperature, and also data from finance.

End-to-End Deep Learning for Predicting Metric Space-Valued Outputs

1 January 2026 at 00:00
Many modern applications involve predicting structured, non-Euclidean outputs such as probability distributions, networks, and symmetric positive-definite matrices. These outputs are naturally modeled as elements of general metric spaces, where classical regression techniques that rely on vector space structure no longer apply. We introduce E2M (End-to-End Metric regression), a deep learning framework for predicting metric space-valued outputs. E2M performs prediction via weighted Fréchet means over training outputs, where the weights are learned by a neural network conditioned on the input. This construction provides a principled mechanism for geometry-aware prediction that avoids surrogate embeddings and restrictive parametric assumptions, while fully preserving the intrinsic geometry of the output space. We establish theoretical guarantees, including a universal approximation theorem that characterizes the expressive capacity of the model and a convergence analysis of the entropy-regularized training objective. Through extensive simulations involving probability distributions, networks, and symmetric positive-definite matrices, we show that E2M consistently achieves state-of-the-art performance, with its advantages becoming more pronounced at larger sample sizes. Applications to human mortality distributions and New York City taxi networks further demonstrate the flexibility and practical utility of this framework.

Global Fr{\'{e}}chet Manifold Learning for Random Objects, With Application to Low-Dimensional Wasserstein Representations of Distributional Data

1 January 2026 at 00:00
We study manifold learning with multidimensional scaling for samples of metric space valued data. By adopting a global version of ISOMAP we obtain low-dimensional Euclidean representations. A key innovation is that we demonstrate that global Fréchet regression can be utilized for mapping the elements of a convex set in the Euclidean representation space back to the metric space where the objects reside. We refer to this approach as Fréchet manifold learning and showcase it with one-dimensional distributions as random objects, equipped with the Wasserstein metric, which is an important special case of our general approach. The resulting low-dimensional representations mimic the parametric representation in a parametric family of distributions but are entirely learned from the data without postulating any parametric model. These Wasserstein representations of distributional data can be viewed as an empirical parametrization of a sample of distributions. The utility of these representations rests on the map from the low-dimensional Euclidean representation space to the space of distributions, which is obtained with global Fréchet regression. We illustrate the proposed approach with distributional data for baby names, bike rentals and age pyramids and further demonstrate how it can be applied for a novel distributional regression method that features one-dimensional distributions as predictors.
❌