Glossary · Test methods

Reliability: How reliably does a test measure?

Reliability describes how consistently a test measures, meaning how free a result is from random measurement fluctuations. Put simply, would the same test give a similar result if what is being measured had not changed?

ReliabilityTest-retestinternal consistency

Key points

  • Reliability is the consistency of a measurement, not whether the test measures the right thing.
  • Important forms include internal consistency, test-retest reliability, and interrater reliability.
  • A test can be reliable and still have low validity. It then consistently measures the wrong thing.

What reliability means

Every measurement contains some random variation. Reliability describes how small this random component is, meaning how consistently a test measures. A reliable test gives similar values under the same conditions. This is a necessary foundation because if a result changes substantially even though what is being measured has not changed, it is hard to base conclusions on it.

Three forms you should distinguish

Internal consistency asks how well the items on a scale fit together, meaning whether they measure the same characteristic in a similar direction. A common measure is Cronbach's alpha, which Lee Cronbach described in 1951. Test-retest reliability asks how similar the results are when the same person completes the test again after some time, assuming nothing has changed. Interrater reliability asks how closely different raters agree when assessing the same thing. These forms address different questions and are not interchangeable. The international COSMIN framework therefore treats them separately.

Why Cronbach's alpha alone is not enough

Cronbach's alpha is often used as evidence of quality. David Streiner explained in 2003 why that is too simplistic. Alpha measures internal consistency and depends, among other things, on the number of items. A high value does not show that a test measures the right thing, and a very high value can even suggest that items are too similar to one another. General rules of thumb, such as a specific threshold above which a test is supposedly good, are of little use without considering the purpose of the measurement and its context. We therefore do not give such a general cutoff here.

Common confusion

Reliability and validity are often treated as the same thing, but they are different. One way to picture this is a scale that always reads two kilograms too high. It is very reliable because it repeats the same result. But it is not valid because it is systematically wrong. In the same way, a questionnaire can measure consistently while missing the actual characteristic of interest. High reliability is a requirement for validity, but it does not prove validity.

What this means for self-tests on medtests.net

Reliability explains why a result can vary somewhat from day to day. The Big Five Test , for example, measures characteristics that are not fixed, and borderline scores near the middle can shift more easily when you repeat the test. The same applies to the Depression Test with the PHQ-9 and the Anxiety Test with the GAD-7, where a score reflects your experiences during a time window that can change. With the Dissociation Test with the DES-II, your answers describe experiences that can vary over time. Such fluctuations are not an error, they are part of every measurement. A single result is best read as a snapshot.

What the term does not mean

A reliable test is not automatically an accurate test. High reliability proves neither validity nor a diagnosis. A high or low score remains a clue that a qualified professional can assess. A self-test provides orientation and does not replace medical, psychotherapeutic, or diagnostic evaluation.

Sources

  • Cronbach, L. J. (1951). Coefficient alpha and the internal structure of tests. Psychometrika, 16(3), 297-334. DOI: 10.1007/BF02310555
  • Streiner, D. L. (2003). Starting at the beginning: an introduction to coefficient alpha and internal consistency. Journal of Personality Assessment, 80(1), 99-103. DOI: 10.1207/S15327752JPA8001_18
  • Mokkink, L. B. et al. (2010). The COSMIN study reached international consensus on taxonomy, terminology, and definitions of measurement properties. Journal of Clinical Epidemiology, 63(7), 737-745. DOI: 10.1016/j.jclinepi.2010.02.006

Frequently asked questions

What does reliability mean?

Reliability describes how consistently a test measures, meaning how little the result is affected by random measurement variation. Put simply, would the same test produce a similar result under the same conditions if what is being measured had not changed?

What is the difference between reliability and validity?

Reliability refers to the consistency of the measurement. Validity asks whether the test measures the intended characteristic. A test can consistently measure the same thing and still miss the characteristic it is supposed to measure. High reliability is a requirement for validity, but it does not prove validity.

Does a high Cronbach's alpha mean that a test is good?

No. Cronbach's alpha is a measure of internal consistency, meaning how consistently the items relate to one another. A high value alone does not show whether the test measures the intended characteristic, and general thresholds without considering purpose and context provide little information.

What is test-retest reliability?

Test-retest reliability describes how similar the results are when the same person takes the test again after some time, assuming that what is being measured has not changed. It is different from internal consistency, which refers to a single measurement time point.