Win statistics
The win statistics are a family of flexible, nonparametric statistics used to quantify effect size in comparisons of hierarchical composite outcomes between two groups. The win statistics – comprising the win ratio, the win odds, and the net treatment benefit (also called the net benefit and the win difference) – are primarily used as measures of effect size in the generalized pairwise comparisons (GPC) framework, a generalisation of the theory underpinning the Finkelstein-Schoenfeld test.
The Finkelstein-Schoenfeld test and, by extension, the GPC framework were developed to address limitations in composite outcome modelling when using conventional statistical methods. Composite outcomes are typically analyzed using survival analysis on a time-to-first composite event, but doing so gives equal weighting to all outcomes and typically results in comparisons driven by the more frequent, less clinically relevant events (e.g., hospitalizations instead of deaths). To remedy this, the GPC framework considers an ordered hierarchy of outcomes that readily incorporates comparisons of recurrent events and continuous outcomes.
Methodology
[edit]Outcome hierarchy
[edit]The first step in calculating the win statistics is to define a hierarchy of outcomes ranked in decreasing order of clinical relevance. This hierarchy dictates the sequence in which outcomes are considered when evaluating the data, with greater priority given to the more clinically important outcomes.
Pairwise comparison procedure
[edit]Forming pairs
[edit]Win statistics are built on subject-level comparisons of between-treatment patient pairs. Two approaches exist for forming these pairs: matched and unmatched.

In a matched comparison, pairs are formed by matching each patient in the treatment arm to one patient in the control arm; Pocock et al. recommended that this matching be performed on baseline risk[1]. If there are patients in the treatment arm and in the control arm, this results in pairs, with any "excess" patients in either arm excluded from the analysis.
In an unmatched comparison, every patient in the treatment arm is paired with every patient in the control arm. If there are patients in the treatment arm and in the control arm, this produces total pairs. The construction of all possible between-arm patient pairs forms the basis for Buyse's GPC framework.
Pair-level comparisons
[edit]Each pair is then evaluated sequentially against the outcome hierarchy. If one patient in the pair has a more favorable result than the other, that patient is declared the "winner", the pair is deemed untied, and the comparison terminates; if neither demonstrates a better result, the pair proceeds onto the next outcome in the hierarchy. Subsequent outcomes are considered similarly, with this process repeating until the pair is untied or the hierarchy exhausted.
In theory, one has infinite flexibility to define what constitutes a favorable outcome; in practice, almost all comparisons fall into one of three comparison types: time-to-first event, recurrent event count, or a generic quantitative comparison. The two event-based outcomes consider comparisons over the pair's shared follow-up, the period over which both patients are fully observed, with events occurring outside this period not considered when performing the comparison.
All three approaches allow the use of "margins of win", whereby the difference between patients' outcomes must exceed some minimally clinical relevant threshold to declare a "winner".
Note that "wins" for the control arm are sometimes referred to as "losses" for the active arm.
Time-to-first event
[edit]
Within each pair, patients are compared on the time until they first experience an event: if both patients experience the event, the one with the longer time-to-event is deemed the "winner"; if one patient experiences the event within the shared follow-up window and the other does not, the patient without the event is deemed the "winner"; if neither experiences an event during their shared follow-up period, the comparison is deemed indeterminate.
Events may also be considered successes (e.g., complete response in oncology studies), in which case the approach described above is reversed.
Recurrent event count
[edit]
Patients are compared on the total number of repeat events experienced during their shared follow-up: if one patient experiences more events than the other, the patient with the fewer events is declared the "winner"; otherwise, the comparison is said to be indeterminate. As with time-to-first outcomes, these events may be considered positives, in which case the patient with the greater number of events is deemed the "winner".
Quantitative
[edit]Generic quantitative comparisons involve comparing patients in each pair with respect to some single quantitative variable (e.g., a quality of life score), where the patient with the higher value is deemed the "winner". The direction of comparisons may be reversed if lower scores are considered more favorable.
Calculation of the win statistics
[edit]After all pairs have been evaluated, the total number of "winners" in each arm is determined. Pairs that are not untied are deemed "tied" for the purpose of calculating the win statistics.
Win ratio
[edit]The win ratio is calculated as the ratio of the number of "winners" in the treatment arm to the number of "winners" in the control arm:
The win ratio can be interpreted as the odds of having a "better" outcome, conditional on being untied by the hierarchy.
Win odds
[edit]The win odds are calculated as the number of "winners" in the treatment arm plus half the number of ties, divided by the number of "winners" in the control arm plus half the number of ties:
Net treatment benefit
[edit]The net treatment benefit is defined as the difference between the number of "winners" in the two arms divided by the total number of patient pairs, calculated as follows:
Example
[edit]To demonstrate, consider a two-arm clinical trial comparing a new treatment () with a control (), with respect to the following hierarchy of outcomes assessed 12 months after randomization:
- Time-to-death
- Number of hospitalizations
- Quality of life score (higher scores indicate better quality of life)
Let denote the patients in the treatment arm and the patients in the control arm. An unmatched analysis would utilize all patients' data, yielding pairwise comparisons; a matched approach would give comparisons.
Starting at the top of the hierarchy, patient pairs are first compared according to their time-to-death over their shared follow-up period. Without loss of generality, consider a pair comprising patients and .
For component (1), a "favorable" outcome is defined as being observed to survive beyond the other patient in the pair. If was censored 9 months after randomization while was followed up until their death at 12 months post-randomization, the comparison would be declared a tie because the event happened outside the pair's shared follow-up period (the first 6 months after randomization during which was observed). If instead had been censored at 24 months, 's death would occur within their shared follow-up period and the comparison would be declared a "win" for the active arm.
For component (2), a "favorable" outcome is defined as experiencing fewer hospitalizations than the other patient in the pair over the pair's shared follow-up period.
To illustrate, assume and were tied on component (1); was hospitalized at 3 and 5 months and later censored at 6 months; and that was hospitalized at 4, 8, and 10 months, before being administratively censored at 12 months. The pair's shared follow-up period is then the first 6 months of follow-up (the point at which was censored); during this period, experienced 2 hospitalizations (at 3 and 5 months), while experienced just 1 (at 4 months). As such, this comparison would then be declared a "win" for the control group.
The subset of pairs who remain tied after comparisons on components (1) and (2) then proceed to comparisons of component (3). For component (3), a "favorable" outcome is defined as having a greater quality of life score.
Suppose and were tied on components (1) and (2) and continued on to a comparison of (3). Assume 's quality of life score was 5 and C1's was 4: in this instance, has a higher score and is declared the "winner".
We might have introduced a "margin of win", stipulating that the difference between the patients' values must be greater than the some threshold to declare a "winner": had we used a margin of two points in the quality of life score, this pair would remain because the difference did not exceed 2 points. Here, without any remaining outcomes left to consider, no further attempts are made to untie the pair and they are considered "tied".
Variance estimation and confidence intervals
[edit]Variance estimation differs significantly between the matched and unmatched analyses due to the patient-level clustering in the latter.
Matched approach
[edit]To estimate the variance in a matched analysis, patient pairs that were not untied by the hierarchy are first discarded, leaving only the "winners" in each arm[1]. The probability of a given winner belonging to the treatment arm, , is then:
Where and denote the number of winners in the active and control arms, respectively. The variance of follows from the Binomial distribution and the bounds of its 95% confidence interval, and , are asymptotically given by:
The confidence intervals for the matched win ratio (WR) and net treatment benefit (NTB) are then given by:
, for
The variance of the win odds is not well defined under this approach.
Unmatched approach
[edit]In 2023, Dong et al.[2] formally described the relationship between the variance estimators for the unmatched win ratio, win odds, and net treatment benefit:
Deriving the variance of one of the win ratio, win odds, or net treatment benefit is then sufficient to find all three.
U-statistic
[edit]In 2016, Dong et al.[3] and Bebu and Lachin[4] developed U-statistic-based asymptotic variance estimators for the win ratio. Define kernel functions , which takes value 1 if patient in the active arm "wins" over patient in the control arm (and 0 otherwise), and , which takes value 1 if patient in the control arm "wins" over patient in the active arm (and 0 otherwise). The U-statistics and are then unbiased estimators of the win and loss probabilities:
Where denotes the number of patients in the active arm and the number of patients in the control arm. It then follows that and are asymptotically normally distributed:
Where and are equal. By the delta method, the variance of the log win ratio is asymptotically given by:
The components of the variance-covariance matrix, , , , are defined as:
It is the definition of the and terms that vary betwen the Dong and Bebu and Lachin approaches.
Bebu and Lachin
[edit]Define , , , and , the patient-specific win and loss counts, as:
It then follows that:
Dong
[edit]Let be defined as:
Then:
Yu and Ganju
[edit]In 2021, Yu and Ganju[5] developed an approximate variance estimator, valid under the null hypothesis that the win ratio is 1, that is a function of the treatment allocation ratio , the number of patients in each arm and , and the number of ties :
, where
The primary use of Yu and Ganju's variance estimator is in sample size calculation: Dong and Bebu and Lachin's estimators remain the preferred approaches for real-world application due to their superior accuracy, particularly when there is evidence of a treatment effect.
Accounting for covariate imbalance
[edit]While matched methods directly address the issue of "unfair" comparisons between patients with different risk profiles, covariate imbalance in the unmatched approach can lead to a loss of precision and biased estimates. Several approaches have been developed to address this, to varying degrees of abstraction.
Stratification
[edit]In 2018, Dong et al. introduced the stratified win statistics,[6] in which patients are divided into distinct strata, stratum-specific win statistics calculated, and pooled estimates produced. This is a simple approach to dealing with confounding, but it is generally limited to dealing with confounding by a small number of categorical variables. Dong et al. described three separate approaches to weighting the stratum-specific win statistics, but recommended the Mantel-Haenszel-style weights for real-world use. Dong et al. also note the potential use of Cochran's Q test to examine the homogeneity of the win statistics across strata.
Inverse probability weighting
[edit]In 2025, Wang et al.[7] introduced an unmatched variance estimator for inverse probability of treatment weighting, building Dong et al.'s work on inverse probability of censoring weighting.[8] Inverse probability of treatment weighting is easily incorporated into matched analyses, although confidence intervals must be estimated via bootstrapping).
Covariate adjustment
[edit]Compared to inverse probability weighting, formal covariate adjustment typically provides more stable estimates, more favorable small-sample properties, and the ability to leverage the prognostic value of covariates to improve estimate precision. However, GPC methods, being non-parametric, do not naturally lend themselves to this form of adjustment: all existing methods introduce parametric modelling techniques into the estimation process.
Ordinal model
[edit]The ordinal approach to covariate adjustment, first described by Hazewinkel et al.,[9] uses ordinal regression to relate the pairs' outcomes, coded as three-level ordinal variables (typically, ), to the within-pair differences in covariate values (for a pair composed of patient in the active arm and in the control arm, this is then , where is patient 's covariate vector). The inverse logits of the two model cutpoints then represent the adjusted loss probability and the adjusted loss or tie probability, respectively, which can be used to construct the adjusted win statistics.
Since each patient contributes to multiple pairs in the model, within-patient clustering must be accounted for in the variance derivation; this is typically achieved using two-way cluster robust sandwich estimators, accounting for clustering in both the control and active arms, or via bootstrapping. Because each pair is necessarily unique, the two-way cluster robust variance matrix decomposes into the sum of the one-way cluster robust variance matrices minus White's heteroskedasticity-consistent variance matrix.[10]
Generalized ordinal logistic regression can be used to relax the proportionality assumption across the thresholds.
Probability index models
[edit]Probability index models, as defined by Thas et al.,[11] are conceptually similar to the ordinal approach, but consider both between-arm and the combinatorially unique within-arm comparisons. Logistic regression models are used to relate pairwise outcomes, coded as three-level ordinal variables (strictly, ), to the within-pair differences in covariate values and a pseudo-treatment indicator (for a pair comprising patients and , the pseudo-treatment indicator is defined as , where if patient is a member of the active arm and otherwise).
In 2023, Song et al.[12] showed that the exponent of the resulting treatment coefficient gives a conditional estimate of the win odds; the win ratio and net treatment benefit cannot be recovered from this approach. Coefficient estimates can be used to quantify the prognostic effect of the different covariates on the outcome.
Proportional win fraction regression models
[edit]In 2021, Mao and Wang[13] introduced the proportional win fraction regression model, a generalization of probability index models. Proportional win fraction regression considers both between-pair and within-pair comparisons and uses logistic regression with (in the binary treatment case) the same pseudo-treatment and pseudo-covariate values as defined above. The model additionally imposes a time-invariance assumption on the win ratio – analogous to the proportional hazards assumption in Cox's proportional hazards model – such that the coefficient estimates are then independent of the censoring distribution. The model residual process at follow-up time (although time invariance is necessary for valid inference) is defined as:
Where if patient had a better outcome than patient , and otherwise, denotes the vector of coefficient estimates (including that of the treamtent), and and defined as before.[13] Notably, pairs that do not get untied do not contribute to this residual and are effectively ignored by the model. The term gives the expected win ratio at time for the pair comprising patients and ; more usefully, gives the estimated win ratio resulting from a 1-point increase in the th covariate.
This model reduces exactly to the Cox proportional hazards model when using a hierarchy consisting of a single time-to-first event comparison.[13] The authors suggest using inverse probability of censoring weights to relax the proportionality assumption.
Randomisation-based
[edit]The randomisation-based approach, first developed by Koch et al.[14] and later framed in the context of the win statistics by Gasparyan et al.,[15] provides an approach to univariate adjustment. This approach directly models the Mann-Whitney U probability (that is, the probability of a win for the treatment arm plus half the probability of a tie): Hazewinkel et al.[9] showed how this can be used to recover the conditional win ratio, win odds, and net treatment benefit using basic linear regression techniques, as well as extending the method beyond a single numerical covariate.
Other approaches
[edit]Other approaches include those described by Cao et al.,[16] which are limited to ordinal outcomes.
Limitations
[edit]Interpretation
[edit]The flexibility of the win statistics also increases their scope for misuse. GPC methods are only applicable when a clinically meaningful hierarchy of events can be defined, in that there is a clear distinction between the clinical importance of different components.[17]
The win statistics can also be dominated by common, less clinically relevant events or comparisons on continuous measures if the more important outcomes included earlier in the hierarchy fail to untie many pairs. It is recommended that cumulative win statistics be estimated at each level of the hierarchy to recognize instances of this, and that interpretations include careful discussion around the contribution of each component to the overall win and loss counts.[18] In such cases, it may be preferable to analyse the "softer" outcomes independently using parametric methods.
The win ratio has been criticized for its failure to consider ties: a scenario in which 75% of comparisons end in a win for the treatment arm, 25% in a win for the control arm, and 0% in ties yields the same win ratio () as a scenario in which 3% of comparisons end in a win for the treatment arm, 1% for the control arm, and 96% in ties. Both the win odds and net treatment benefit are attenuated by ties, although similarly lopsided scenarios can result in equal estimates of the net treatment benefit. Pocock et al.[19] recommend reporting both the win ratio and the net treatment benefit to fully characterize the distribution of wins and losses.
Although the net treatment benefit has been described as an absolute measure of effect, this is not strictly true[19]: for a hierarchy consisting of a single outcome, the net treatment benefit can be expressed as the difference between the concordance and discordance probabilities and thus reduces to a special case of Somers' D statistic, which is neither relative nor absolute (although indeed it does reduce to the absolute risk difference in the case of a single, binary outcome).[20]
Methodology
[edit]Although methods for covariate adjustment now exist, concerns have been raised regarding the compatibility of multivariate data-generating mechanisms with the assumptions underpinning the approaches to covariate adjustment.[21]
By restricting pairwise comparisons to the subset of patients who remain at-risk, censoring diminishes the contribution of future events to the test statistic in Gehan's Wilcoxon test. Such dependence on censoring is undesirable: early events are afforded undue weight and dominate the win and loss counts, while differential censoring can lead to considerable bias.[22] As direct extensions of Gehan's Wilcoxon test, GPC methods that incorporate longitudinal endpoints are potentially subject to these biases. However, in 2016, David Oakes demonstrated that the win statistics were independent of censoring under the proportional hazards assumption.[23] In the same paper, Oakes also described an integration-based approach to estimating the win statistics in the hypothetical scenario where no patients are lost to follow-up. As a result of this dependence on censoring, the win statistics can also vary over the course of follow-up.
In 2026, Valerie Fu described the non-monotonicity of the win statistics, whereby a treatment with positive marginal effects on each component in the hierarchy can nonetheless result in a negative overall win statistic.[24] Fu showed that conditioning on earlier ties reweights subsequent comparisons and thus introduces the scope for non-monotonicity.
For unmatched analyses, these methods can become computationally intensive and execute with quadratic run time in the worst case (if margins of win – which break the transitivity of comparisons – or recurrent event comparisons are specified), making unmatched analyses impractical for large datasets. Sorting-based procedures may be used to evaluate comparisons in special cases, reducing the computational complexity to .
Additional criticisms include the inability to quantify patient-level measures of effect when using win statistics.[25]
History
[edit]The theory underpinning the win statistics was first described by Lemuel Moyé, Barry R. Davis, and C. Morton Hawkins in 1992.[26] Moyé et al. proposed an extension to Gehan's Wilcoxon test – itself a generalization of the Wilcoxon rank sum test to account for censoring in survival analysis – that first considers pairwise comparisons of survival and then, among patients who survived the follow-up period, on a single quantitative measure presumed to be collected at the end of follow-up.
In 1999, Dianne Finkelstein and David Schoenfeld generalized this approach further with the Finkelstein-Schoenfeld test, permitting larger, more flexible hierarchies and constraining longitudinal comparisons to the shared follow-up period.[27] In 2010, Marc Buyse formalised these methods into the GPC framework and proposed the net treatment benefit as a measure of effect for the Finkelstein-Schoenfeld test;[20] in 2012, Stuart Pocock, Cono Ariti, Tim Collier, and Duolao Wang published the win ratio as another, while simultaneously introducing the concept of matched analyses.[1] In 2020, Gaohong Dong and colleagues introduced the win odds.[28]
In 2016, Dong et al.[3] and Bebu and Lachin[4] developed U-statistic-based variance estimators. In 2018, Dong et al. introduced the stratified win statistics.[6] In 2021, Dong et al.[8] introduced inverse probability of censoring weighting to the win statistics; in 2025, Wang et al. extended this to further permit inverse probability of treatment weighting.[7]
Covariate adjustment in the GPC framework had been explored as early as 2012 in the form of probability index models by Thas et al.,[11] but it was not until 2023 that this method was formally extended to the win odds.[12] Mao and Wang generalized the probability index models with the introduction of proportional win fraction regression in 2021.[13] Koch et al.'s 1998 paper on a parametric, randomization-based approach to covariate adjustment in non-parametric comparisons was later framed in the context of the win statistics by Gasparyan et al. in 2021.[14][15] In 2026, Hazewinkel et al. generalized the randomization-based approach and introduced the ordinal regression method;[9] also in 2026, Cao et al. developed an additional approach to covariate adjustment for the win odds.[16]
Further developments include the introduction of hypothetical estimands absent loss to follow-up by Oakes in 2016 and the invent of formal sample size formulae.[23][5]
Use of the win statistics remains largely confined to cardiology, with a review of fourteen general medical and cardiology journals between January 2022 and July 2024 identifying 61 articles from 36 cardiovascular trials reporting them.[19]
Software implementations
[edit]Various user-written packages exist for calculating the unmatched win statistics, but no built-in packages exist:
References
[edit]- 1 2 3 Pocock, S. J.; Ariti, C. A.; Collier, T. J.; Wang, D. (2012-01-02). "The win ratio: a new approach to the analysis of composite endpoints in clinical trials based on clinical priorities". European Heart Journal. 33 (2): 176–182. doi:10.1093/eurheartj/ehr352. ISSN 0195-668X. PMID 21900289.
- ↑ Dong, Gaohong; Huang, Bo; Verbeeck, Johan; Cui, Ying; Song, James; Gamalo-Siebers, Margaret; Wang, Duolao; Hoaglin, David C.; Seifu, Yodit; Mütze, Tobias; Kolassa, John (January 2023). "Win statistics (win ratio, win odds, and net benefit) can complement one another to show the strength of the treatment effect on time-to-event outcomes". Pharmaceutical Statistics. 22 (1): 20–33. doi:10.1002/pst.2251. ISSN 1539-1604. PMID 35757986.
- 1 2 Dong, Gaohong; Li, Di; Ballerstedt, Steffen; Vandemeulebroecke, Marc (September 2016). "A generalized analytic solution to the win ratio to analyze a composite endpoint considering the clinical importance order among components". Pharmaceutical Statistics. 15 (5): 430–437. doi:10.1002/pst.1763. ISSN 1539-1604. PMID 27485522.
- 1 2 Bebu, Ionut; Lachin, John M. (2016-01-01). "Large sample inference for a win ratio analysis of a composite outcome based on prioritized components". Biostatistics. 17 (1): 178–187. doi:10.1093/biostatistics/kxv032. ISSN 1468-4357. PMC 4679075. PMID 26353896.
- 1 2 Yu, Ron Xiaolong; Ganju, Jitendra (2022-03-15). "Sample size formula for a win ratio endpoint". Statistics in Medicine. 41 (6): 950–963. doi:10.1002/sim.9297. ISSN 0277-6715. PMID 35084052.
- 1 2 Dong, Gaohong; Qiu, Junshan; Wang, Duolao; Vandemeulebroecke, Marc (2018-07-04). "The stratified win ratio". Journal of Biopharmaceutical Statistics. 28 (4): 778–796. doi:10.1080/10543406.2017.1397007. ISSN 1054-3406. PMID 29172988.
- 1 2 Wang, Duolao; Zheng, Sirui; Cui, Ying; He, Nengjie; Chen, Tao; Huang, Bo (2025-01-02). "Adjusted win ratio using the inverse probability of treatment weighting". Journal of Biopharmaceutical Statistics. 35 (1): 21–36. doi:10.1080/10543406.2023.2275759. ISSN 1054-3406. PMID 37947400.
- 1 2 Dong, Gaohong; Mao, Lu; Huang, Bo; Gamalo-Siebers, Margaret; Wang, Jiuzhou; Yu, Guanglei; Hoaglin, David C. (2020). "The inverse-probability-of-censoring weighting (IPCW) adjusted win ratio statistic: An unbiased estimator in the presence of independent censoring". Journal of Biopharmaceutical Statistics. 30 (5): 882–899. doi:10.1080/10543406.2020.1757692. PMC 7538385. PMID 32552451.
- 1 2 3 Hazewinkel, Audinga-Dea; Gregson, John; Bartlett, Jonathan W.; Gasparyan, Samvel B.; Wright, David; Pocock, Stuart (2026-03-31), Covariate adjustment for hierarchical outcomes and the win ratio: how to do it and is it worthwhile?, doi:10.64898/2026.03.30.26347966, retrieved 2026-08-26
- ↑ Cameron, A. Colin; Gelbach, Jonah B.; Miller, Douglas L. (April 2011). "Robust Inference With Multiway Clustering". Journal of Business & Economic Statistics. 29 (2): 238–249. doi:10.1198/jbes.2010.07136. ISSN 0735-0015.
- 1 2 Thas, Olivier; Neve, Jan De; Clement, Lieven; Ottoy, Jean-Pierre (2012-09-01). "Probabilistic Index Models". Journal of the Royal Statistical Society Series B: Statistical Methodology. 74 (4): 623–671. doi:10.1111/j.1467-9868.2011.01020.x. ISSN 1369-7412.
- 1 2 Song, James; Verbeeck, Johan; Huang, Bo; Hoaglin, David C.; Gamalo-Siebers, Margaret; Seifu, Yodit; Wang, Duolao; Cooner, Freda; Dong, Gaohong (2023-03-04). "The win odds: statistical inference and regression". Journal of Biopharmaceutical Statistics. 33 (2): 140–150. doi:10.1080/10543406.2022.2089156. ISSN 1054-3406.
- 1 2 3 4 Mao, Lu; Wang, Tuo (December 2021). "A class of proportional win-fractions regression models for composite outcomes". Biometrics. 77 (4): 1265–1275. doi:10.1111/biom.13382. ISSN 0006-341X. PMC 7988303. PMID 32974905.
- 1 2 Koch, Gary G.; Tangen, Catherine M.; Jung, Jin-Whan; Amara, Ingrid A. (1998-08-15). "Issues for covariance analysis of dichotomous and ordered categorical data from randomized clinical trials and non-parametric strategies for addressing them". Statistics in Medicine. 17 (15–16): 1863–1892. doi:10.1002/(SICI)1097-0258(19980815/30)17:15/16<1863::AID-SIM989>3.0.CO;2-M. ISSN 0277-6715. PMID 9749453.
- 1 2 Gasparyan, Samvel B; Folkvaljon, Folke; Bengtsson, Olof; Buenconsejo, Joan; Koch, Gary G (February 2021). "Adjusted win ratio with stratification: Calculation methods and interpretation". Statistical Methods in Medical Research. 30 (2): 580–611. doi:10.1177/0962280220942558. ISSN 0962-2802. PMID 32726191.
- 1 2 Cao, Zhiqiang; Zuo, Scott; Baumann, Mary Ryan; Plourde, Kendra; Heagerty, Patrick; Tong, Guangyu; Li, Fan (2026-06-11), Covariate-adjusted win statistics in randomized clinical trials with ordinal outcomes, arXiv, doi:10.48550/arXiv.2508.20349, arXiv:2508.20349, retrieved 2026-08-29
- ↑ Butler, Javed; Stockbridge, Norman; Packer, Milton (2024-05-14). "Win Ratio: A Seductive But Potentially Misleading Method for Evaluating Evidence from Clinical Trials". Circulation. 149 (20): 1546–1548. doi:10.1161/CIRCULATIONAHA.123.067786. ISSN 0009-7322. PMID 38739696.
- ↑ Pocock, Stuart J.; Ferreira, João Pedro; Collier, Timothy J.; Angermann, Christiane E.; Biegus, Jan; Collins, Sean P.; Kosiborod, Mikhail; Nassif, Michael E.; Ponikowski, Piotr; Psotka, Mitchell A.; Teerlink, John R.; Tromp, Jasper; Gregson, John; Blatchford, Jonathan P.; Zeller, Cordula (2023-05-01). "The Win Ratio Method in Heart Failure Trials: Lessons Learnt from EMPULSE". European Journal of Heart Failure. 25 (5): 632–641. doi:10.1002/ejhf.2853. ISSN 1388-9842. PMC 10330107. PMID 37038330.
- 1 2 3 Gregson, John; Taylor, Dylan; Owen, Ruth; Collier, Tim; J. Cohen, David; Pocock, Stuart (2025-06-03). "Hierarchical Composite Outcomes and Win Ratio Methods in Cardiovascular Trials: A Review and Consequent Guidance". Circulation. 151 (22): 1606–1619. doi:10.1161/CIRCULATIONAHA.124.070251. ISSN 0009-7322. PMID 40455842.
- 1 2 Buyse, Marc (2010-12-30). "Generalized pairwise comparisons of prioritized outcomes in the two-sample problem". Statistics in Medicine. 29 (30): 3245–3257. doi:10.1002/sim.3923. ISSN 0277-6715. PMID 21170918.
- ↑ Thas, Olivier; Neve, Jan De; Clement, Lieven; Ottoy, Jean-Pierre (2012-09-01). "Probabilistic Index Models". Journal of the Royal Statistical Society Series B: Statistical Methodology. 74 (4): 623–671. doi:10.1111/j.1467-9868.2011.01020.x. ISSN 1369-7412.
- ↑ Prentice, R. L.; Marek, P. (December 1979). "A Qualitative Discrepancy between Censored Data Rank Tests". Biometrics. 35 (4): 861–867. doi:10.2307/2530120. JSTOR 2530120. PMID 393312.
- 1 2 Oakes, D. (September 2016). "On the win-ratio statistic in clinical trials with multiple types of event". Biometrika. 103 (3): 742–745. doi:10.1093/biomet/asw026. ISSN 0006-3444.
- ↑ Fu, Valerie R. (May 2026). "When Better is Worse: A Paradox of the Win Ratio and Net Treatment Benefit". Statistics in Medicine. 45 (10–12) e70580. doi:10.1002/sim.70580. ISSN 0277-6715. PMID 42082167.
- ↑ "Navigating the Landscape of Hierarchical Multi-Component Strategies: GPC, DOOR, and MOST". arxiv.org. Retrieved 2026-08-26.
- ↑ Moyé, Lemuel A.; Davis, Barry R.; Hawkins, C. Morton (January 1992). "Analysis of a clinical trial involving a combined mortality and adherence dependent interval censored endpoint". Statistics in Medicine. 11 (13): 1705–1717. doi:10.1002/sim.4780111305. ISSN 0277-6715. PMID 1485054.
- ↑ Finkelstein, Dianne M.; Schoenfeld, David A. (1999-06-15). "Combining mortality and longitudinal measures in clinical trials". Statistics in Medicine. 18 (11): 1341–1354. doi:10.1002/(SICI)1097-0258(19990615)18:11<1341::AID-SIM129>3.0.CO;2-7. ISSN 0277-6715. PMID 10399200.
- ↑ Dong, Gaohong; Hoaglin, David C.; Qiu, Junshan; Matsouaka, Roland A.; Chang, Yu-Wei; Wang, Jiuzhou; Vandemeulebroecke, Marc (2020-01-02). "The Win Ratio: On Interpretation and Handling of Ties". Statistics in Biopharmaceutical Research. 12 (1): 99–106. doi:10.1080/19466315.2019.1575279. ISSN 1946-6315.
- ↑ Gregson, John; Ferreira, João Pedro; Collier, Tim (2023-09-01). "winratiotest: A command for implementing the win ratio and stratified win ratio in Stata". The Stata Journal. 23 (3). SAGE Publications: 835–850. doi:10.1177/1536867X231196480. ISSN 1536-867X.
- ↑ Cui, Ying; Huang, Bo (2025-07-02), WINS: The R WINS Package, retrieved 2026-08-27
- ↑ Wang, Lu Mao and Tuo (2021-11-26), WR: Win Ratio Analysis of Composite Time-to-Event Outcomes, retrieved 2026-08-27
- ↑ Duarte, Kevin; Ferreira, Joao Pedro (2020-11-23), WinRatio: Win Ratio for Prioritized Outcomes and 95% Confidence Interval, retrieved 2026-08-27