Edge Rewrite
// HTMLRewriter · presentation

This page was redesigned at the edge.

Cloudflare fetched the original article and streamed it through HTMLRewriter to apply an entirely new visual system without rebuilding the source page.

Jump to content

Compositional data

From Wikipedia, the free encyclopedia

In statistics, compositional data are quantitative descriptions of the parts of some whole, conveying relative information. Mathematically, compositional data is represented by points on a simplex. Measurements involving probabilities, proportions, percentages, and ppm can all be thought of as compositional data.

All analyses of real data deals with noisy measurements. The simpler case is in empirical studies such as questionnaire surveys, where each respondent's total allocation is fixed and one must deal solely with the effects of closure, odds and odds ratios serve as invariant statistics. In such cases, associations are determined solely by the magnitude of the odds ratio relative to 1. But in compositional data containing noise, the noise is present in the total sum and is distributed asymmetrically among the individual components. Since noise and signal are inseparable, the data can only be dichotomized into zeros and noise-contaminated components.[1] Therefore, without explicit noise filtering, even statistics based on odds remain mathematically ill-defined. Indeed, odds and their ratios can only be calculated across samples in the absolute absence of any noise other than that induced by the compositionality itself (except for normalization).

Ternary plot

[edit]

Compositional data in three variables can be plotted via ternary plots. The use of a barycentric plot on three variables graphically depicts the ratios of the three variables as positions in an equilateral triangle.

The problem of noise

[edit]

Under conditions where Poisson noise, overdispersion, or additive instrumental errors occur, the process of composition (scaling by the total sum) propagates these distortions asymmetrically across all components, making the value of each component difficult to compare due to noise contamination. As a simple degree-of-freedom check or limit analysis ( or ) demonstrates, compositional data is fundamentally closed and isolated within a single row (each sample), illustrating the inherent difficulty of direct cross-sample comparisons.[2] In the presence of such noise, any attempt to use the central limit theorem and Chebyshev's inequality to define a strict boundary between signal and noise fails, as the full joint distribution remains unknown and the threshold parameter remains purely an arbitrary matter of human interpretation.[3]

To demonstrate the baseline invariance under pure conditions, consider the ratio between two components, and , within a given sample under purely compositional effects. Letting denote the scaling factor applied to maintain a constant row total, the closed compositional values are expressed as and . Their ratio then becomes , wherein cancels out completely. Similarly, the odds ratio between two distinct samples (or conditions), expressed as , remains entirely unaffected by sample-specific closure factors. Under these pure conditions, the relative order of the components derived from the observed percentage data deterministically reflects the true physical ordering, allowing for the exact calculation of the odds ratio as a definitive value for the study. The resulting odds ratio can be classified into three distinct categories based on its magnitude relative to 1: a negative association (< 1), complete independence (= 1), and a positive association (> 1).

Simplicial sample space

[edit]
An illustration of the Aitchison simplex. Here, there are 3 parts, represent values of different proportions. A, B, C, D and E are 5 different compositions within the simplex. A, B and C are all equivalent and D and E are equivalent.

In general, John Aitchison defined compositional data to be proportions of some whole in 1982.[4] In particular, a compositional data point (or composition for short) can be represented by a real vector with positive components. The sample space of compositional data is a simplex:

The only information is given by the ratios between components, so the information of a composition is preserved under multiplication by any positive constant. Therefore, the sample space of compositional data can always be assumed to be a standard simplex, i.e. . In this context, normalization to the standard simplex is called closure and is denoted by :

where D is the number of parts (components) and denotes a row vector.

Aitchison geometry

[edit]

The simplex can be given the structure of a vector space in several different ways. The following vector space structure is called Aitchison geometry or the Aitchison simplex and has the following operations:

Perturbation (vector addition)
Powering (scalar multiplication)
Inner product

Endowed with those operations, the Aitchison simplex forms a -dimensional Euclidean inner product space. The uniform composition is the zero vector.

Orthonormal bases

[edit]

Since the Aitchison simplex forms a finite dimensional Hilbert space, it is possible to construct orthonormal bases in the simplex. Every composition can be decomposed as follows

where forms an orthonormal basis in the simplex.[5] The values are the (orthonormal and Cartesian) coordinates of with respect to the given basis. They are called isometric log-ratio coordinates .

Linear transformations

[edit]

There are three well-characterized isomorphisms that transform from the Aitchison simplex to real space. All of these transforms satisfy linearity as given below.

Logarithm of the odds ratio or Additive log ratio transform

[edit]

The logarithm of the odds ratio is also known as the additive log ratio (alr). The additive log ratio (alr) transform is an isomorphism where . This is given by

The choice of denominator component is arbitrary, and could be any specified component. This transform is commonly used in chemistry with measurements such as pH. In addition, this is the transform most commonly used for multinomial logistic regression. The alr transform is not an isometry, meaning that distances on transformed values will not be equivalent to distances on the original compositions in the simplex.

Center log ratio transform

[edit]

The center log ratio (clr) transform is both an isomorphism and an isometry where

Where is the geometric mean of . The inverse of this function is also known as the softmax function.

Covariance singularity and rank deficiency under linear dependency

[edit]

By construction, the CLR transformation enforces a strict linear dependency constraint on the system:

Geometrically, this constraint confines the transformed data to a -dimensional hyperplane embedded within the -dimensional real space . This structural collinearity induces an absolute rank deficiency in the variance-covariance structure, reducing its true rank to .

As a direct consequence, the resulting covariance matrix of the CLR-transformed values is strictly singular, and its determinant is zero:

Because a singular matrix mathematically lacks an inverse ( does not exist), standard multivariate procedures that strictly rely on the inverse covariance matrix—such as the Mahalanobis distance, Linear Discriminant Analysis (LDA), and maximum likelihood estimations—are structurally unsupported. Resorting to the Moore-Penrose pseudo-inverse () serves as a numerical workaround but does not resolve the underlying dimensional contradiction.

Isometric log ratio transform

[edit]

The isometric log ratio (ilr) transform is both an isomorphism and an isometry where

There are multiple ways to construct orthonormal bases, including using the Gram–Schmidt orthogonalization or singular-value decomposition of clr transformed data. Another alternative is to construct log contrasts from a bifurcating tree. If given a bifurcating tree, a basis from the internal nodes in the tree can be constructed.

A representation of a tree in terms of its orthogonal components. l represents an internal node, an element of the orthonormal basis. This is a precursor to using the tree as a scaffold for the ilr transform

Each vector in the basis would be determined as follows

The elements within each vector are given as follows

where are the respective number of tips in the corresponding subtrees shown in the figure. It can be shown that the resulting basis is orthonormal[6]

Once the basis is built, the ilr transform can be calculated as follows

where each element in the ilr transformed data is of the following form

where and are the set of values corresponding to the tips in the subtrees and

Problems of ILR

[edit]

ILR assumes statistical homogeneity and geometric isometry within the simplex , allowing a projection into a -dimensional Euclidean space while preserving uniform distances and angles. But in any real-world measurement system (even under idealized, error-free conditions), the unconstrained variables in the positive real space exhibit fundamentally asymmetric and independent variance structures, leading to disparate component-specific Coefficients of Variation (). When the closure operator is applied to enforce the constant-sum constraint, the resulting "closure strain" is allocated across the components in a highly non-linear and asymmetric fashion, governed by the non-uniform Jacobian of the mapping.[7] The result is the generation of artifacts on real-world data when using ILR to generate covariance matrices and distance metrics.[8]

ILR proponents also claim usefulness on synthetic or artificial datasets where variable relationships are artificially isolated from physical covariation. However, given these restrictions, odds and odds ratios themselves are sufficient.

Examples

[edit]
  • In chemistry, compositions can be expressed as molar concentrations of each component. As the sum of all concentrations is not determined, the whole composition of D parts is needed and thus expressed as a vector of D molar concentrations. These compositions can be translated into weight per cent multiplying each component by the appropriated constant.
  • In demography, a town may be a compositional data point in a sample of towns; a town in which 35% of the people are Christians, 55% are Muslims, 6% are Jews, and the remaining 4% are others would correspond to the quadruple [0.35, 0.55, 0.06, 0.04]. A data set would correspond to a list of towns.
  • In geology, a rock composed of different minerals may be a compositional data point in a sample of rocks; a rock of which 10% is the first mineral, 30% is the second, and the remaining 60% is the third would correspond to the triple [0.1, 0.3, 0.6]. A data set would contain one such triple for each rock in a sample of rocks.
  • In high throughput DNA sequencing and RNA sequencing, data obtained are typically transformed to relative abundances, rendering them compositional.
  • In probability and statistics, a partition of the sampling space into disjoint events is described by the probabilities assigned to such events. The vector of D probabilities can be considered as a composition of D parts. As they add to one, one probability can be suppressed and the composition is completely determined.
  • In chemometrics, for the classification of petroleum oils.[9]
  • In a survey, the proportions of people positively answering some different items can be expressed as percentages. As the total amount is identified as 100, the compositional vector of D components can be defined using only D  1 components, assuming that the remaining component is the percentage needed for the whole vector to add to 100.

Examples of further analyses

[edit]

In biology, the relative abundances of specific sequences ("reads") in a sequencing result is used as a noisy estimate of the relative abundances of these actual sequences. For example, the amount of RNAs in different cells could be used to identify what cell type they are, or to understand the specific way a cell responds to a stimulus. Transformation of these abundances given the sequencing depth is essential to adjust for variable sampling efficiency and different variances.[10] Some of the better methods are based on CLR.[11]

See also

[edit]

Notes

[edit]
  1. Lin, H.; Peddada, S. D. (2020). "Analysis of compositions of microbiomes with bias correction". Nature Communications. 11 (1): 3514.
  2. Itagaki, Tatsuki; Kobayashi, Hirokazu; Sakata, Ken-Ichiro; Miyamoto, Ikuya; Hasebe, Akira; Kitagawa, Yoshimasa (2024). "Compositional Data and Microbiota Analysis: Imagination and Reality". Microorganisms. 12 (7): 1484. doi:10.3390/microorganisms12071484. PMC 11279367. PMID 39065253.
  3. Tchebichef, P. L. (1867). "Des valeurs moyennes". Journal de Mathématiques Pures et Appliquées. 12: 177–184.
  4. Aitchison, John (1982). "The Statistical Analysis of Compositional Data". Journal of the Royal Statistical Society. Series B (Methodological). 44 (2): 139–177. doi:10.1111/j.2517-6161.1982.tb01195.x.
  5. Egozcue et al.
  6. Egozcue & Pawlowsky-Glahn 2005
  7. Ohta, T. (2011). Limitations of log-ratio transformations in anisotropic manifolds. Journal of Compositional Data, 23(2), 115-132.
  8. Ohta, T., Arai, H. & Noda, A. Identification of the Unchanging Reference Component of Compositional Data from the Properties of the Coefficient of Variation. Math Geosci 43, 421–434 (2011). https://doi.org/10.1007/s11004-011-9332-y
  9. Olea, Ricardo A.; Martín-Fernández, Josep A.; Craddock, William H. (2021). "Multivariate classification of the crude oil petroleum systems in southeast Texas, USA, using conventional and compositional analysis of biomarkers". In Advances in Compositional Data Analysis—Festschrift in honor of Vera-Pawlowsky-Glahn, Filzmoser, P., Hron, K., Palarea-Albaladejo, J., Martín-Fernández, J.A., editors. Springer: 303−327.
  10. Ahlmann-Eltze, C; Huber, W (May 2023). "Comparison of transformations for single-cell RNA-seq data". Nature Methods. 20 (5): 665–672. doi:10.1038/s41592-023-01814-1. PMC 10172138. PMID 37037999.
  11. Booeshaghi, A. Sina; Hallgrímsdóttir, Ingileif B.; Gálvez-Merchán, Ángel; Pachter, Lior (2026-06-22). "Normalization for sampled count data". pp. 2022–05.06.490859. bioRxiv 10.1101/2022.05.06.490859.

References

[edit]
[edit]