The key is the correct option in a multiple-choice item — the response that scores. The other options are distractors; together with the stem they form the item. Answer key is the same word used for the full list of correct options across a test, the document used in scoring.
A key earns its place by two properties: it is unambiguously correct against the construct, and it is not betrayed by surface features that let test-wise candidates pick it without engaging the construct. The first is content; the second is craft. Most operational scoring failures trace to one or the other — a defensible alternative answer, or a giveaway in option length, syntactic form, or specificity.
| Property | Why it matters |
|---|---|
| Singular correctness | No defensible second answer under the construct or in plausible candidate readings |
| Surface parallelism with distractors | Same length range, same grammatical form, same level of specificity |
| Position rotation across the test | Keys distributed roughly evenly across A/B/C/D positions; humans favour middle positions when rushed |
| Construct-aligned | The reasoning needed to select it is the reasoning the test wants to measure |
| Defensible to a content expert | A second reviewer trained in the construct would mark it correct without hesitation |
Haladyna, Downing and Rodriguez (2002) consolidate the empirical evidence: keys that are systematically longer, more qualified, or more grammatically complete than distractors are picked at higher rates by candidates who have not engaged the construct. The asymmetry is what test-wiseness training exploits.
Two failure modes recur:
Both are caught by pretest. Operational test programs reject items where two raters disagree on the key, and rewrite items whose key wins on surface features.
A miskey is an item operationally scored against the wrong option. Diagnosed when the supposed key shows low or negative point-biserial in distractor analysis while a distractor shows positive — strong candidates are rejecting the listed key for an option the writers labelled incorrect. Standard responses: rescore against the empirically correct option, double-key (accept both), or drop the item. Miskeys are most damaging on high-stakes single-form tests, which is why operational programs run independent key-verification panels before live administration.
Random placement of the key across positions is the textbook prescription, but Bar-Hillel (2015) and follow-up work show that human writers, left to themselves, place keys in middle positions disproportionately and that uniformly random placement reduces guessing yield. Some test programs balance positions by design — equal counts of A, B, C, D keys per section — though this can introduce its own pattern that test-wise candidates exploit. Pseudo-random placement with position-frequency monitoring is the operational compromise.
Language items add construct-correctness wrinkles other content domains do not face:
Each requires the writer to verify the key against the passage with a content-blind reviewer who has only the question and options, to catch keys that depend on knowledge the passage does not actually deliver.