Distractor analysis is the post-administration examination of how each incorrect option in a multiple-choice item performed. It sits inside item analysis but answers a tighter question: did each distractor do its job? An option that no one chooses is dead weight; an option that strong candidates choose is suspect; an option whose endorsement rate matches the key's is a sign of two-key ambiguity. The procedure is the standard evidence base for keeping, revising, or replacing options after pretest or operational use.
Two statistics carry most of the analysis. Endorsement rate, the proportion of candidates who selected the option, tells you whether the distractor attracted any responses at all. Option discrimination (the point-biserial or biserial correlation between selecting the option and total test score, or the simpler upper-lower group difference) tells you who selected it. A well-functioning distractor has a positive endorsement rate and a negative discrimination index: weaker candidates pick it more than stronger ones. The key shows the mirror pattern: high endorsement, positive discrimination.
Haberman, Liu and Sinharay (2019) lay out the canonical interpretation grid for international language assessments. The pattern below assumes a four-option item with a moderate p-value.
| Pattern | What it means | Action |
|---|---|---|
| Distractor endorsed by < 5% | Non-functional; effectively reduces the item to fewer options | Revise or replace |
| Distractor endorsed by ≥ 5% with negative discrimination | Working as intended | Keep |
| Distractor with positive discrimination | Stronger candidates prefer it; possible miskey, ambiguity, or content flaw | Investigate before reuse |
| Key with low or negative discrimination | Item likely flawed at the source: wrong key, ambiguous stem, or construct-irrelevant cue | Pull from operational use |
| Two options with similar endorsement rates near the top | Two-key item; both arguably correct | Rewrite |
The 5% threshold for non-functional distractors traces to Haladyna and Downing (1993), who reviewed teacher-built tests and found roughly half of all distractors failed it, meaning many four-option items operated as three- or two-option items in practice. Subsequent work (Rodriguez, 2005; Gierl et al., 2017) converged on three plausible options as the empirical sweet spot, since writing a fourth working distractor is costly and frequently fails.
In a classical test theory pipeline, distractor analysis runs alongside item difficulty and item discrimination computation. For each item, the output is a small option table:
| Option | % choosing | Upper-third % | Lower-third % | Point-biserial |
|---|---|---|---|---|
| A | 14 | 6 | 24 | -0.18 |
| B (key) | 62 | 86 | 32 | 0.42 |
| C | 19 | 6 | 32 | -0.21 |
| D | 5 | 2 | 12 | -0.09 |
Option D is on the threshold: endorsed but only marginally discriminating. The reviewer would either rewrite it to look more like a plausible competitor or replace it. Option A and C are functioning. The key has acceptable discrimination.
Under IRT approaches the question recasts as how does the probability of selecting each option change with ability? Nominal response models and the Rasch-based partial credit framework let analysts plot one trace line per option. A working distractor's trace line peaks among low-ability candidates and falls as ability rises; the key's trace line rises monotonically. Bechger, Maris and Hsiao's (2010) work on detecting unexpected response patterns under Rasch extends these diagnostics to systematic item flaws.
Language items often fail diagnostically rather than statistically. A reading-comprehension distractor may pull responses from candidates with strong general English but weak inference skill, or capture L1 transfer rather than the intended construct. Distractor analysis is where these failures surface. When a distractor in a reading item shows positive discrimination, the usual cause is that it matches the surface form of the passage closely enough to reward word-spotting: the item is testing scanning, not comprehension. When a distractor fails to attract any low-ability candidates, the cause is usually transparent wrongness, whether grammatical mismatch with the stem, length asymmetry, or content far outside the passage.
Operational test programs build distractor analysis into a pretest cycle: items are trialled, options are rewritten or rejected on the basis of statistics from a pretest sample, and only items meeting endorsement and discrimination thresholds enter the live form. Cambridge English's pretesting workflow, in which items are calibrated against anchored items and entered into a banked pool, depends on this analysis step to keep reliability stable across forms.
Distractor statistics describe how options did function on the sample they were trialled on; they do not by themselves explain why. A non-functional distractor in pretest data on a high-ability sample may function on the operational target population, and vice versa. Analysts pair the numbers with qualitative review (reading the option against the stem and key) before acting. Endorsement-rate floors also work poorly on small samples: the 5% rule presumes hundreds of responses, not a class of thirty.