Edge Rewrite
// HTMLRewriter · presentation

This page was redesigned at the edge.

Cloudflare fetched the original article and streamed it through HTMLRewriter to apply an entirely new visual system without rebuilding the source page.

// request.cf · coarse context

A page that knows where it met you.

Only coarse request metadata is shown. This demo does not display or persist visitor IP addresses.

Country
US
Cloudflare location
CMH
Connection
HTTP/2
Language
Not provided

Ray ID: a233ff715d7d1709

Jump to content

Generalized pairwise comparisons

From Wikipedia, the free encyclopedia

Generalized pairwise comparisons (GPC) is a non-parametric statistical methodology for comparing treatments or interventions in clinical trials using one or more outcomes simultaneously. The method is an extension of classical pairwise comparison procedures, including the Wilcoxon rank-sum test and the Mann–Whitney U test, allowing comparisons to incorporate outcomes of different types, such as survival time, binary events and continuous measurements.[1]

GPC compares every participant in one treatment group with every participant in the other treatment group. Each comparison is classified as favorable, unfavorable or neutral according to pre-specified clinical criteria. The results of all pairwise comparisons are then combined into summary measures of treatment effect.[1]By allowing flexible and clinically meaningful comparison rules, GPC can simultaneously account for multiple outcomes that contribute to the overall assessment of treatment benefit. In many clinical trials, treatments may improve one outcome while worsening another, making the combined interpretation of conventional analyses based on a single endpoint and traditional composite endpoint difficult. GPC provides a framework for incorporating multiple outcomes while preserving their relative clinical importance through hierarchical prioritization or relative weighting.[2]

History

[edit source]

Generalized pairwise comparisons were introduced by Marc Buyse in 2010 as an extension of rank-based pairwise comparison methods such as the Wilcoxon rank-sum test and the Mann–Whitney U test. The methodology was developed to accommodate multiple prioritized outcomes, clinically meaningful thresholds and outcomes of different types within a unified analytical framework.[1][2]

Methodology

[edit source]

Pairwise comparison methods have a long history in statistics. The Wilcoxon rank-sum test, introduced by Frank Wilcoxon in 1945, and the closely related Mann–Whitney U test developed two years later, compare two independent groups without assuming normally distributed data.[3][4]Although widely used, classical rank-based methods were developed primarily for analyzing a single outcome. Nowadays, clinical trials increasingly evaluate treatments using multiple outcomes, including measures of efficacy, safety, quality of life and survival.[5]

Conventional approaches typically analyze these outcomes separately or combine them into composite endpoints. Separate analyses may produce conflicting conclusions, while composite endpoints require combining events often of speculative clinical importance into a single variable.[6]Subsequent methodological developments extended the approach to accommodate censored survival data, clinically meaningful thresholds, competing risks and stratified analyses.[7][8]

Generalized pairwise comparisons evaluate treatment effect by comparing every participant receiving one treatment with every participant receiving the comparator treatment. For two treatment groups containing n and m participants, the analysis considers all n × m possible patient pairs.[1]For each pair, outcomes are compared according to predefined decision rules. If the participant receiving the experimental treatment has a better outcome than the participant receiving the control treatment, the comparison is considered favorable. If the opposite is true, the comparison is unfavorable. When neither participant can be considered to have a better outcome according to the predefined criteria, the comparison is classified as neutral.[1][9]Unlike conventional rank tests, GPC is not restricted to a single outcome. Multiple outcomes may be incorporated using two principal approaches:[1][10]

  • In a weighted approach, each outcome contributes to the comparison according to a predetermined weight reflecting its relative importance.
  • More commonly, GPC uses a hierarchical or prioritized approach. Outcomes are ranked according to their clinical importance before the analysis begins. Each patient pair is first compared using the highest-priority outcome. Only when that comparison is neutral or inconclusive is the next outcome considered. The process continues until either a favorable or unfavorable comparison is identified or all outcomes have been evaluated.

Clinical relevance thresholds may also be incorporated. Rather than considering any numerical difference between patients to represent a treatment benefit, it is possible to specify a minimum clinically important difference that must be exceeded before a comparison is judged favorable or unfavorable. Differences smaller than the threshold are treated as neutral.[1]Because GPC is based on pairwise comparisons rather than assumptions about the distribution of outcome variables, it generally requires fewer distributional assumptions than many parametric statistical methods. However, appropriate interpretation depends on the predefined comparison rules, outcome priorities and estimand selected before the analysis.[10]

Measures of treatment effect

[edit source]

Several measures of treatment effect can be derived from generalized pairwise comparisons, all of which are based on the proportion of favorable and unfavorable patient pairs.[1]

  • The Net Treatment Benefit is calculated as the difference between the probability of a favorable comparison and the probability of an unfavorable comparison.[1]
  • The probabilistic index estimates the probability that a randomly selected participant receiving the experimental treatment has a more favorable outcome than a randomly selected participant receiving the control treatment, with ties contributing equally to both groups.[11]
  • The win ratio expresses the ratio of favorable to unfavorable patient pairs after excluding ties.[12]
  • Closely related measures include the win odds and success odds.[10]

Although these measures are mathematically related, they differ in interpretation and statistical properties. The choice of estimand depends on the objectives of the analysis and the study design.[13]

Limitations

[edit source]

The methodology has several recognized limitations. The interpretation of results depends on decisions made before the analysis, including the choice of outcomes, their order of priority and any thresholds defining clinically meaningful differences. Different specifications may lead to different estimates of treatment effect, making pre-specification a critical aspect of study design.[10]

References

[edit source]
  1. 1 2 3 4 5 6 7 8 9 Buyse, Marc (2010). "Generalized pairwise comparisons of prioritized outcomes in the two-sample problem". Statistics in Medicine. 29 (30): 3245–3257. doi:10.1002/sim.3923. PMID 21170918.
  2. 1 2 Buyse, Marc; Verbeeck, Johan; Saad, Everardo D.; De Backer, Mickaël; Deltuvaite-Thomas, Vaiva; Molenberghs, Geert (2025). Handbook of Generalized Pairwise Comparisons: Methods for Patient-Centric Analysis. Chapman & Hall/CRC Handbooks of Modern Statistical Methods (1st ed.). Chapman & Hall/CRC. p. 584. ISBN 978-1032488059.
  3. Wilcoxon, Frank (1945). "Individual Comparisons by Ranking Methods". Biometrics Bulletin. 1 (6): 80–83. doi:10.2307/3001968. JSTOR 3001968.
  4. Mann, H.B.; Whitney, D.R. (1947). "On a Test of Whether One of Two Random Variables is Stochastically Larger Than the Other". Annals of Mathematical Statistics. 18 (1): 50–60. doi:10.1214/aoms/1177730491.
  5. De Backer, Mickaël; Sengar, Manju; Mathews, Vikram (2024). "Design of a clinical trial using generalized pairwise comparisons to test a less intensive treatment regimen". Clinical Trials. 21 (2): 180–188. doi:10.1177/17407745231206465. PMC 11195000. PMID 37877379.
  6. Verbeeck, Johan; De Backer, Mickaël; Verwerft, Jan (2023). "Generalized Pairwise Comparisons to Assess Treatment Effects: JACC Review Topic of the Week". Journal of the American College of Cardiology. 82 (13): 1360–1372. doi:10.1016/j.jacc.2023.06.047. hdl:1942/42083. PMID 37730293.
  7. Ozenne, Brice; Budtz-Jørgensen, Esben; Péron, Julien (2021). "The asymptotic distribution of the Net Benefit estimator in presence of right-censoring". Statistical Methods in Medical Research. 30 (11): 2399–2412. doi:10.1177/09622802211037067. PMID 34633267.
  8. Deltuvaite-Thomas, Vaiva; Verbeeck, Johan; Burzykowski, Tomasz (2023). "Generalized pairwise comparisons for censored data: An overview". Biometrical Journal. 65 (2) 2100354. doi:10.1002/bimj.202100354. PMID 36127290.
  9. Deltuvaite-Thomas, Vaiva (2023). "Generalized pairwise comparisons for censored data: An overview". Biometrical Journal. 65 (2) 2100354. doi:10.1002/bimj.202100354. PMID 36127290. {{cite journal}}: Unknown parameter |DUPLICATE_article-number= ignored (help)
  10. 1 2 3 4 Verbeeck, Johan (2023). "Generalized Pairwise Comparisons to Assess Treatment Effects: JACC Review Topic of the Week". Journal of the American College of Cardiology. 82 (13): 1360–1372. doi:10.1016/j.jacc.2023.06.047. hdl:1942/42083. PMID 37730293.
  11. Verbeeck, Johan; Deltuvaite-Thomas, Vaiva; Berckmoes, Ben (2021). "Unbiasedness and efficiency of non-parametric and UMVUE estimators of the probabilistic index and related statistics". Statistical Methods in Medical Research. 30 (3): 747–768. doi:10.1177/0962280220966629. PMID 33256560.
  12. Pocock, Stuart J.; Ariti, Cono A.; Collier, Timothy J.; Wang, Duolao (2012). "The win ratio: a new approach to the analysis of composite endpoints in clinical trials based on clinical priorities". European Heart Journal. 33 (2): 176–182. doi:10.1093/eurheartj/ehr352. PMID 21900289.
  13. Ozenne, Brice; Budtz-Jørgensen, Esben; Péron, Julien (2021). "The asymptotic distribution of the Net Benefit estimator in presence of right-censoring". Statistical Methods in Medical Research. 30 (11): 2399–2412. doi:10.1177/09622802211037067. PMID 34633267.