Tetrachoric correlation
The tetrachoric correlation coefficient is a measure of association between two dichotomous variables. It is based on the assumption that the observed binary variables arise by dichotomizing two underlying continuous variables that follow a normal distribution. The dichotomization may result from measurement design, categorization, or other practical considerations.[1]
The coefficient can be interpreted as the Pearson correlation coefficient that would be obtained between the two underlying continuous, normally distributed variables if their values were observable.
Calculation
[edit]Suppose that the observed values of two dichotomous variables, X and Y, are summarized in the following 2 × 2 contingency table, where the entries denote cell frequencies:
| y = 1 | y = 0 | total | |
| x = 1 | |||
| x = 0 | |||
| total |
Determining the exact value of the tetrachoric correlation coefficient from sample data requires fairly numerically complex calculations.[2] One must find using the following equation:
- ,
where and .
In 2005, Bonett and Price proposed a simplified formula that yields an estimate with good properties[1]:
- ,
where: is the number π, and is determined by the following formula:
and is an estimate of the odds ratio using the contingency table cell counts increased by 0.5:
- .
See also
[edit]References
[edit]- 1 2 Douglas G. Bonett, Robert M. Price (2005-06-01). "Inferential Methods for the Tetrachoric Correlation Coefficient". Journal of Educational and Behavioral Statistics. 30 (2): 213–225. doi:10.3102/10769986030002213. ISSN 1076-9986. Retrieved 2025-02-08.
- ↑ Bernard Harris (2014). Tetrachoric Correlation Coefficient. John Wiley & Sons, Ltd. doi:10.1002/9781118445112.stat00385. ISBN 978-1-118-44511-2. Retrieved 2025-02-08.
Bibliography
[edit]- Why so many Correlation Coefficients
- Bruce M. King, Edward W. Minium, Statistical Reasoning in Psychology and Education, 2003, p. 146.