pith. sign in

arxiv: 2606.11646 · v1 · pith:K5V3PEUYnew · submitted 2026-06-10 · 💻 cs.LG · q-bio.QM· stat.ML

Tree-Structured Orthonormal Decomposition of the Aitchison Simplex

classification 💻 cs.LG q-bio.QMstat.ML
keywords aitchisonorthonormalstructuretreecoordinatedatadecompositionfeatures
0
0 comments X
read the original abstract

Compositional data -- vectors encoding relative proportions -- arise across scientific domains, including ecology, geochemistry, and genomics. The features in these data often come with known hierarchical structure (e.g., taxonomies, phylogenies, ontologies), yet existing methods either ignore this structure, discard the intrinsic Aitchison geometry, are designed for binary trees, or yield incomplete coordinate systems. We describe PolyILR, a canonical orthonormal decomposition of the Aitchison tangent space aligned with any tree topology. Our construction defines a weighted local geometry at each internal node capturing full branching structure, then lifts these to a global orthonormal basis where every coordinate corresponds to a specific tree location. On microbiome and single-cell benchmarks, PolyILR yields stable, interpretable features and enables inference at multiscale tree resolution. We also establish a novel theoretical connection to softmax classifiers, suggesting possible applications to probabilistic modeling.

This paper has not been read by Pith yet.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.