Cramér–von Mises criterion
This article may be too technical for most readers to understand. (July 2023) |

In statistics the Cramér–von Mises criterion is a criterion used for judging the goodness of fit of a cumulative distribution function (CDF) compared to a given empirical distribution function , or for comparing two empirical distributions. It is also used as a part of other algorithms, such as minimum distance estimation. It is defined as , where
In one-sample applications is the theoretical distribution and is the empirically observed distribution. Alternatively the two distributions can both be empirically estimated ones; this is called the two-sample case.
The criterion is named after Harald Cramér and Richard Edler von Mises who first proposed it in 1928–1930. [1][2] The generalization to two samples is due to Anderson. [3]
The Cramér–von Mises test is an alternative to the Kolmogorov–Smirnov test (1933).[4]
Cramér–von Mises test (one sample)
[edit]Let be the observed values, in increasing order. Then the test statistic is[3]: 1153 [5]
If this value is larger than the tabulated value, then the hypothesis that the data came from the distribution can be rejected.
Watson test
[edit]A modified version of the Cramér–von Mises test is the Watson test[6] which uses the statistic U2, where[5]
where
Cramér–von Mises test (two samples)
[edit]Let and be the observed values in the first and second sample respectively, in increasing order. Within the combined sample of size , let be the ranks of the xs in the combined sample, and let be the ranks of the ys in the combined sample. Anderson[3]: 1149 shows that
where U is defined as
If the value of T is larger than the tabulated values,[3]: 1154–1159 the hypothesis that the two samples come from the same distribution can be rejected. (Some books[specify] give critical values for U, which is more convenient, as it avoids the need to compute T via the expression above. The conclusion will be the same.)
The above assumes there are no duplicates in the , , and sequences. So is unique, and its rank is in the sorted list . If there are duplicates, and through are a run of identical values in the sorted list, then one common approach is the midrank[7] method: assign each duplicate a "rank" of . In the above equations, in the expressions and , duplicates can modify all four variables , , , and .
Cramér distance
[edit]For two distributions on the real line with cumulative distribution functions and and finite first moment, the Cramér distance is
a metric on the space of such distributions.[8] Note that some sources define the Cramér distance as , but this fails the triangle inequality and so cannot be properly defined as a distance. The Cramér distance is the one-dimensional case of the energy distance via the relationship ,[9] and when represents a single observation with cumulative distribution , is equivalent to the continuous ranked probability score, a strictly proper scoring rule.[10]

Under the probability integral transform (PIT), the plot of the empirical distribution of the transformed values and the uniform distribution on creates a PIT reliability diagram. The Cramér distance between these two distributions equals , the square root of the criterion, and serves as a numerical score of the calibration error of . This may also be referred to as the Root Mean Square Calibration Error (RMSCE).
For a deterministic (point) forecast at , the PIT degenerates to a Bernoulli random variable on with success probability , so in the population limit the Cramér distance between the PIT CDF and the uniform distribution evaluates in closed form to
This quantity is minimized at (the unbiased case) with value , establishing a calibration-error floor that no point forecast can fall below regardless of how accurate its central value is. In contrast, a well-calibrated probabilistic forecast can approach 0. Similarly, this quantity is maximized at the bias extremes with value .
References
[edit]- ↑ Cramér, H. (1928). "On the Composition of Elementary Errors". Scandinavian Actuarial Journal. 1928 (1): 13–74. doi:10.1080/03461238.1928.10416862.
- ↑ von Mises, R. E. (1928). Wahrscheinlichkeit, Statistik und Wahrheit. Julius Springer.
- 1 2 3 4 Anderson, T. W. (1962). "On the Distribution of the Two-Sample Cramer–von Mises Criterion". Annals of Mathematical Statistics. 33 (3). Institute of Mathematical Statistics: 1148–1159. doi:10.1214/aoms/1177704477. ISSN 0003-4851.
- ↑ A.N. Kolmogorov, "Sulla determinizione empirica di una legge di distribuzione" Giorn. Ist. Ital. Attuari, 4 (1933) pp. 83–91
- 1 2 Pearson, E.S., Hartley, H.O. (1972) Biometrika Tables for Statisticians, Volume 2, CUP. ISBN 0-521-06937-8 (page 118 and Table 54)
- ↑ Watson, G.S. (1961) "Goodness-Of-Fit Tests on a Circle", Biometrika, 48 (1/2), 109-114 JSTOR 2333135
- ↑ Ruymgaart, F. H., (1980) "A unified approach to the asymptotic distribution theory of certain midrank statistics". In: Statistique non Parametrique Asymptotique, 1±18, J. P. Raoult (Ed.), Lecture Notes on Mathematics, No. 821, Springer, Berlin.
- ↑ Bellemare, M. G.; Danihelka, I.; Dabney, W.; Mohamed, S.; Lakshminarayanan, B.; Hoyer, S.; Munos, R. (2017). "The Cramer Distance as a Solution to Biased Wasserstein Gradients". arXiv:1705.10743.
- ↑ Székely, G. J.; Rizzo, M. L. (2013). "Energy statistics: A class of statistics based on distances". Journal of Statistical Planning and Inference. 143 (8): 1249–1272. doi:10.1016/j.jspi.2013.03.018.
- ↑ Gneiting, T.; Raftery, A. E. (2007). "Strictly proper scoring rules, prediction, and estimation". Journal of the American Statistical Association. 102 (477): 359–378. doi:10.1198/016214506000001437.
- M. A. Stephens (1986). "Tests Based on EDF Statistics". In D'Agostino, R.B.; Stephens, M.A. (eds.). Goodness-of-Fit Techniques. New York: Marcel Dekker. ISBN 0-8247-7487-6.
Further reading
[edit]- Xiao, Y.; A. Gordon; A. Yakovlev (January 2007). "A C++ Program for the Cramér–von Mises Two-Sample Test" (PDF). Journal of Statistical Software. 17 (8). doi:10.18637/jss.v017.i08. ISSN 1548-7660. OCLC 42456366. S2CID 54098783. Retrieved June 12, 2009.