The Sinkhorn treatment effect is a new entropic optimal transport measure of divergence between counterfactual distributions that admits first- and second-order pathwise differentiability, debiased estimators, and asymptotically valid tests for distributional treatment effects.
hub
Computational optima l transport
17 Pith papers cite this work, alongside 40 external citations. Polarity classification is still indexing.
abstract
Optimal transport (OT) theory can be informally described using the words of the French mathematician Gaspard Monge (1746-1818): A worker with a shovel in hand has to move a large pile of sand lying on a construction site. The goal of the worker is to erect with all that sand a target pile with a prescribed shape (for example, that of a giant sand castle). Naturally, the worker wishes to minimize her total effort, quantified for instance as the total distance or time spent carrying shovelfuls of sand. Mathematicians interested in OT cast that problem as that of comparing two probability distributions, two different piles of sand of the same volume. They consider all of the many possible ways to morph, transport or reshape the first pile into the second, and associate a "global" cost to every such transport, using the "local" consideration of how much it costs to move a grain of sand from one place to another. Recent years have witnessed the spread of OT in several fields, thanks to the emergence of approximate solvers that can scale to sizes and dimensions that are relevant to data sciences. Thanks to this newfound scalability, OT is being increasingly used to unlock various problems in imaging sciences (such as color or texture processing), computer vision and graphics (for shape manipulation) or machine learning (for regression, classification and density fitting). This short book reviews OT with a bias toward numerical methods and their applications in data sciences, and sheds lights on the theoretical properties of OT that make it particularly useful for some of these applications.
hub tools
citation-role summary
citation-polarity summary
representative citing papers
Modern text encoders resist second-order collapse under mean pooling because token embeddings concentrate tightly within texts, and this resistance correlates with stronger downstream performance.
Defines Hausdorff-style and Wasserstein-style metrics on C-sets, proving the latter are convex relaxations of the former and computable as linear programs.
Retrieval from out-of-domain foundation models enables personalization of a lightweight transformer for stress detection, yielding +3.92% accuracy and +4.76% F1 gains on WESAD without user labels.
Optimal-transport couplings are used to construct confidence intervals that reduce coverage error relative to classical quantile-based intervals, with consistency theory and data-driven hyperparameters.
ARC is a self-tuning joint state-covariance estimator using adaptive robust loss and block-coordinate descent that recovers inlier covariance and matches baseline accuracy in simulations and UWB experiments without manual tuning.
Layer-resolved OT detects source-disengagement hallucinations in NMT but achieves only 57% balanced accuracy on summarization because content misrepresentation can occur with correct attention.
MGP uses a MERGE-based Markovian process from linguistic minimalism to discover and combine atomic building blocks into exact symbolic regression models, avoiding bloat when a suitable lexicon is provided.
An interior-point method is introduced to compute dynamical quantum optimal transport geodesics on density matrices, shown to approximate some quantum chemistry problems after parameter tuning.
Decoding alignment metrics can remain high and unchanged even when encoding manifold topology is causally altered, so they do not imply similar function or computation across neural populations.
SPIN performs bidirectional domain transfer in SBI to retain parameter mutual information from unlabeled real observations, improving real-world posterior inference under increasing misspecification.
Derives MSIP algorithm from MMD gradient flows for weighted quantization, extending mean shift and relating to preconditioned gradient descent and Lloyd's clustering.
D-CLIPSE reaches near-centralized multi-robot localization accuracy and consistency by consensus on shared states only, plus passive listening of preintegrated odometry exchanges.
AURA is an adaptive uncertainty-aware refinement method for auditing LLM-as-a-judge pairwise decisions that learns human-consistency signals through selective human verification on uncertain cases.
Dual-stream EEG decoder separates identity and orientation to support 3D reconstruction from neural signals via circular regression and conditioned diffusion.
PCA scatterplots misleadingly indicate clusters in Kuehneotherium teeth data, whereas t-SNE and persistent homology detect a ring-like one-dimensional manifold, backed by a generative model of uniform sampling from a unit circle whose cosine distances follow an arcsine distribution.
Documents a practical PyTorch implementation of batched Sinkhorn iterations for the entropy-regularized Wasserstein loss introduced by Cuturi.
citing papers explorer
-
Sinkhorn Treatment Effects: A Causal Optimal Transport Measure
The Sinkhorn treatment effect is a new entropic optimal transport measure of divergence between counterfactual distributions that admits first- and second-order pathwise differentiability, debiased estimators, and asymptotically valid tests for distributional treatment effects.
-
Why Mean Pooling Works: Quantifying Second-Order Collapse in Text Embeddings
Modern text encoders resist second-order collapse under mean pooling because token embeddings concentrate tightly within texts, and this resistance correlates with stronger downstream performance.
-
Hausdorff and Wasserstein metrics on graphs and other structured data
Defines Hausdorff-style and Wasserstein-style metrics on C-sets, proving the latter are convex relaxations of the former and computable as linear programs.
-
Retrieval-Augmented Personalization with Foundation Models for Wearable Stress Detection
Retrieval from out-of-domain foundation models enables personalization of a lightweight transformer for stress detection, yielding +3.92% accuracy and +4.76% F1 gains on WESAD without user labels.
-
An Optimal Transportation Approach for Improved Confidence Intervals
Optimal-transport couplings are used to construct confidence intervals that reduce coverage error relative to classical quantile-based intervals, with consistency theory and data-driven hyperparameters.
-
ARC: Adaptive Robust Joint State and Covariance Estimation
ARC is a self-tuning joint state-covariance estimator using adaptive robust loss and block-coordinate descent that recovers inlier covariance and matches baseline accuracy in simulations and UWB experiments without manual tuning.
-
Layer-Resolved Optimal Transport for Hallucination Detection in NMT and Abstractive Summarization
Layer-resolved OT detects source-disengagement hallucinations in NMT but achieves only 57% balanced accuracy on summarization because content misrepresentation can occur with correct attention.
-
Minimalist Genetic Programming
MGP uses a MERGE-based Markovian process from linguistic minimalism to discover and combine atomic building blocks into exact symbolic regression models, avoiding bloat when a suitable lexicon is provided.
-
An algorithm for dynamical quantum optimal transport with applications to quantum chemistry
An interior-point method is introduced to compute dynamical quantum optimal transport geodesics on density matrices, shown to approximate some quantum chemistry problems after parameter tuning.
-
Decoding Alignment without Encoding Alignment: A critique of similarity analysis in neuroscience
Decoding alignment metrics can remain high and unchanged even when encoding manifold topology is causally altered, so they do not imply similar function or computation across neural populations.
-
Information-Preserving Domain Transfer with Unlabeled Data in Misspecified Simulation-Based Inference
SPIN performs bidirectional domain transfer in SBI to retain parameter mutual information from unlabeled real observations, improving real-world posterior inference under increasing misspecification.
-
Weighted quantization using MMD: From mean field to mean shift via gradient flows
Derives MSIP algorithm from MMD gradient flows for weighted quantization, extending mean shift and relating to preconditioned gradient descent and Lloyd's clustering.
-
D-CLIPSE: Distributed Consensus-based Localization with Passive Listening on Shared State Exchange
D-CLIPSE reaches near-centralized multi-robot localization accuracy and consistency by consensus on shared states only, plus passive listening of preintegrated odometry exchanges.
-
AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing
AURA is an adaptive uncertainty-aware refinement method for auditing LLM-as-a-judge pairwise decisions that learns human-consistency signals through selective human verification on uncertain cases.
-
Dual-Stream EEG Decoding for 3D Visual Perception
Dual-stream EEG decoder separates identity and orientation to support 3D reconstruction from neural signals via circular regression and conditioned diffusion.
-
Beyond Explained Variance: A Cautionary Tale of PCA
PCA scatterplots misleadingly indicate clusters in Kuehneotherium teeth data, whereas t-SNE and persistent homology detect a ring-like one-dimensional manifold, backed by a generative model of uniform sampling from a unit circle whose cosine distances follow an arcsine distribution.
-
Implementation of batched Sinkhorn iterations for entropy-regularized Wasserstein loss
Documents a practical PyTorch implementation of batched Sinkhorn iterations for the entropy-regularized Wasserstein loss introduced by Cuturi.