Glossary · Test methods
Reliability: How reliably does a test measure?
Reliability describes how consistently a test measures, meaning how free a result is from random measurement fluctuations. Put simply, would the same test give a similar result if what is being measured had not changed?
Key points
- Reliability is the consistency of a measurement, not whether the test measures the right thing.
- Important forms include internal consistency, test-retest reliability, and interrater reliability.
- A test can be reliable and still have low validity. It then consistently measures the wrong thing.
What reliability means
Every measurement contains some random variation. Reliability describes how small this random component is, meaning how consistently a test measures. A reliable test gives similar values under the same conditions. This is a necessary foundation because if a result changes substantially even though what is being measured has not changed, it is hard to base conclusions on it.
Three forms you should distinguish
Internal consistency asks how well the items on a scale fit together, meaning whether they measure the same characteristic in a similar direction. A common measure is Cronbach's alpha, which Lee Cronbach described in 1951. Test-retest reliability asks how similar the results are when the same person completes the test again after some time, assuming nothing has changed. Interrater reliability asks how closely different raters agree when assessing the same thing. These forms address different questions and are not interchangeable. The international COSMIN framework therefore treats them separately.
Why Cronbach's alpha alone is not enough
Cronbach's alpha is often used as evidence of quality. David Streiner explained in 2003 why that is too simplistic. Alpha measures internal consistency and depends, among other things, on the number of items. A high value does not show that a test measures the right thing, and a very high value can even suggest that items are too similar to one another. General rules of thumb, such as a specific threshold above which a test is supposedly good, are of little use without considering the purpose of the measurement and its context. We therefore do not give such a general cutoff here.
Common confusion
Reliability and validity are often treated as the same thing, but they are different. One way to picture this is a scale that always reads two kilograms too high. It is very reliable because it repeats the same result. But it is not valid because it is systematically wrong. In the same way, a questionnaire can measure consistently while missing the actual characteristic of interest. High reliability is a requirement for validity, but it does not prove validity.
What this means for self-tests on medtests.net
Reliability explains why a result can vary somewhat from day to day. The Big Five Test , for example, measures characteristics that are not fixed, and borderline scores near the middle can shift more easily when you repeat the test. The same applies to the Depression Test with the PHQ-9 and the Anxiety Test with the GAD-7, where a score reflects your experiences during a time window that can change. With the Dissociation Test with the DES-II, your answers describe experiences that can vary over time. Such fluctuations are not an error, they are part of every measurement. A single result is best read as a snapshot.
What the term does not mean
A reliable test is not automatically an accurate test. High reliability proves neither validity nor a diagnosis. A high or low score remains a clue that a qualified professional can assess. A self-test provides orientation and does not replace medical, psychotherapeutic, or diagnostic evaluation.
Sources
- Cronbach, L. J. (1951). Coefficient alpha and the internal structure of tests. Psychometrika, 16(3), 297-334. DOI: 10.1007/BF02310555
- Streiner, D. L. (2003). Starting at the beginning: an introduction to coefficient alpha and internal consistency. Journal of Personality Assessment, 80(1), 99-103. DOI: 10.1207/S15327752JPA8001_18
- Mokkink, L. B. et al. (2010). The COSMIN study reached international consensus on taxonomy, terminology, and definitions of measurement properties. Journal of Clinical Epidemiology, 63(7), 737-745. DOI: 10.1016/j.jclinepi.2010.02.006
Related terms and matching tests
More terms in the glossary
Matching self-tests