Glossary · Test methods

Sensitivity and specificity explained simply

Sensitivity and specificity describe how well a test distinguishes between people with and without the target condition. Sensitivity refers to people who have the condition, while specificity refers to people who do not. Both measures depend on the chosen cutoff.

2x2 tablefalse positivefalse negative

Key points

  • Sensitivity: the proportion of affected people whom the test correctly identifies as positive.
  • Specificity: the proportion of unaffected people whom the test correctly identifies as negative.
  • No test achieves both perfectly at the same time. False positive and false negative results remain possible.

How the two measures differ

Both terms come from the evaluation of diagnostic tests. Altman and Bland described them briefly and clearly in the BMJ in 1994. Sensitivity looks only at the group of people who actually have the condition and asks what proportion the test correctly identifies as positive. Specificity looks only at the group of people who actually do not have the condition and asks what proportion the test correctly identifies as negative. A test can perform strongly on one measure and less strongly on the other.

High sensitivity is useful when you want to miss as few affected people as possible. High specificity is useful when you want to incorrectly flag as few unaffected people as possible. Achieving both at the same time is difficult because changing the cutoff to improve one measure usually worsens the other.

The 2x2 table

A 2x2 table shows this most clearly. It compares the test result with the actual condition and produces four cells.

2x2 table: test result compared with actual condition
Condition present Condition not present
Test positive true positivecorrectly identified as positive false positivepositive even though not affected
Test negative false negativenegative even though affected true negativecorrectly identified as negative

Sensitivity comes from the left column: true positives divided by everyone who is actually affected. Specificity comes from the right column: true negatives divided by everyone who is actually unaffected. False positive and false negative results are the two ways a test can be wrong.

A simplified calculation example

The following numbers are made up and are used only for illustration. They do not come from a study. Suppose a group contains 100 people who are actually affected and 900 who are not. A test correctly identifies 90 of the 100 affected people as positive, so it misses 10 (false negatives). Sensitivity is then 90 out of 100, or 90 percent. Of the 900 unaffected people, the test correctly identifies 810 as negative and incorrectly flags 90 as positive (false positives). Specificity is then 810 out of 900, or 90 percent.

The other side of the calculation is more interesting. Among all positive results, 90 are true positives and 90 are false positives, so only half are true positives. This is not because the test is poor, but because there are far more unaffected than affected people. This is where the frequency of a condition matters, known as prevalence.

Common confusion

Sensitivity and specificity are often confused with the question of how likely a single positive result is to be correct. That is a different measure, the positive predictive value, and it also depends on prevalence. A test with high sensitivity and specificity can still produce many false positives in a group where the condition is very rare. Sensitivity and specificity are properties of the test at a given cutoff. By themselves, they do not tell you what an individual result means for you.

What this means for self-tests on medtests.net

These measures explain why we phrase results cautiously. In the Eating Disorder Test with the SCOFF, we note that a low score does not reliably rule out an eating disorder. This reflects the possibility of false negative results. In the Autism Test with the AQ-10 and the Depression Test with the PHQ-9, the reverse also applies. A positive result is a clue, not proof, because false positive results also occur. We provide specific performance measures for an instrument only in relation to the relevant validation study and the group studied there.

What the terms do not mean

High sensitivity does not mean that a positive result proves a condition. High specificity does not mean that a negative result rules out a condition. No self-test is 100 percent reliable, and none of these measures turns it into a diagnosis. A self-test provides orientation and does not replace medical, psychotherapeutic, or diagnostic evaluation.

Sources

  • Altman, D. G. & Bland, J. M. (1994). Diagnostic tests 1: sensitivity and specificity. BMJ, 308(6943), 1552. DOI: 10.1136/bmj.308.6943.1552
  • Altman, D. G. & Bland, J. M. (1994). Diagnostic tests 2: predictive values. BMJ, 309(6947), 102. DOI: 10.1136/bmj.309.6947.102
  • Wilson, J. M. G. & Jungner, G. (1968). Principles and Practice of Screening for Disease. WHO Public Health Papers No. 34. iris.who.int

Frequently asked questions

What is the difference between sensitivity and specificity?

Sensitivity refers to people with the target condition. It shows what proportion of them are correctly identified by the test as positive. Specificity refers to people without the condition. It shows what proportion of them are correctly identified as negative.

What is a false-positive result?

A false-positive result occurs when the test is positive even though the condition is not present. The opposite is a false-negative result, in which the test is negative even though the condition is present. Both can occur with any test.

Does a positive screening result mean that I have the condition?

No. A positive screening result is a sign that further assessment may be useful, not proof. How likely a positive result is to be a true positive also depends on how common the condition is in the group being studied.

Why does no test have 100 percent sensitivity and specificity?

Sensitivity and specificity depend in part on the chosen cutoff and often involve a tradeoff. A cutoff that identifies more people with the condition usually also incorrectly flags more people without the condition as positive, and vice versa. In practice, there is therefore almost always some proportion of false-positive and false-negative results.