A cloze test is a reading passage with words systematically deleted at regular intervals (typically every 5th, 6th, or 7th word), which the test-taker must restore. Developed by Wilson Taylor (1953) as a readability measure, it was adopted into language testing as a measure of integrative language proficiency.
The term "cloze" comes from "closure," the Gestalt psychological principle that humans naturally fill in gaps to perceive wholes.
Words are deleted at a predetermined interval regardless of what the words are. Every 7th word is removed, whether it is a function word, a content word, or a proper noun. This is Taylor's original procedure and is what "cloze test" technically refers to.
The mechanical deletion ensures the test samples a representative cross-section of the text's linguistic features, including grammar, vocabulary, discourse markers, and collocations, without the test writer selecting what to test.
Specific words are chosen for deletion based on what the test writer wants to assess, such as particular grammar structures, vocabulary items, or discourse features. This is often called a "modified cloze" or "gap-fill" exercise. It sacrifices the representative sampling of the standard cloze for targeted assessment of specific language areas.
Strictly speaking, a selective deletion task is not a true cloze test (Hughes 2003; Brown & Abeywickrama 2010), though the terms are often conflated in practice.
A variant where the second half of every second word is deleted. The first sentence is left intact. Developed by Klein-Braley & Raatz (1984) as a more efficient alternative to traditional cloze. Advantages: shorter, higher reliability per minute of testing time, less controversial scoring.
Exact-word scoring. Only the original word is accepted. "She went to the ___" accepts only the specific original word. This is stricter but more reliable; no rater judgment needed.
Acceptable-word scoring. Any contextually appropriate word is accepted. This is more valid (it credits genuine comprehension) but introduces scorer variability and requires an answer key that anticipates alternatives.
Research (Oller 1979; Brown 2002) shows that both methods rank test-takers in the same order; the correlation between exact and acceptable scoring is very high (typically r > .95). Exact-word scoring is therefore generally preferred for its practicality, despite appearing harsh.
This is one of the most debated questions in language testing.
The integrative argument. Oller (1979) argued that cloze tests measure a global "expectancy grammar": the ability to predict and process language using all available linguistic and contextual cues simultaneously. This made cloze a powerful proficiency measure because it tapped multiple skills at once.
The criticism. Alderson (1979, 1980) demonstrated that most cloze items can be answered using local context only (the immediately surrounding words), not the broader discourse understanding Oller claimed. Different deletion rates produce different tests measuring different things. The construct being measured is unstable.
The current consensus. Cloze tests primarily measure lower-level reading processes, including vocabulary knowledge, grammatical competence, and local cohesion processing, and are less effective at measuring higher-level skills like main idea comprehension, inference, or critical evaluation (Alderson 2000; Hughes 2003). They are useful but limited.
The cloze test represents an important moment in the history of language testing: the shift from discrete-point testing (testing one item at a time in isolation) to integrative testing (testing multiple skills through extended text processing). Even though the cloze test's limitations are now well documented, the underlying principle, that language ability is best measured through integrated tasks rather than isolated items, shaped the development of modern communicative testing.
In practical terms, cloze-type tasks remain common in language classrooms and tests. Understanding what they can and cannot measure helps teachers choose them appropriately: as useful tools for quick assessment of reading-level vocabulary and grammar processing, not as comprehensive measures of language proficiency.