Pretesting is the trialling of test items on a sample of candidates before operational use, to estimate item parameters, surface defects, and decide what enters the live form. It is the empirical step between item authoring and operational administration: items the writers believe in are exposed to actual candidates, statistics are computed, and only items that meet thresholds for difficulty, discrimination, distractor functioning and absence of DIF move forward. Without pretesting, item quality rests on author intuition, which the empirical literature consistently shows is poorly calibrated.
The term overlaps with field testing and field trial. The distinction in practice:
| Stage | Purpose | Sample | Data collected |
|---|---|---|---|
| Pretest | Estimate item parameters; cull weak items | Smaller, representative of operational population | Item statistics, distractor analysis, candidate comments |
| Field trial | Verify operational performance under live conditions | Larger, full operational simulation | Section timings, full form statistics, administrator feedback |
Pretesting feeds field testing; field testing feeds operational launch. Lower-stakes tests sometimes fold the two together; high-stakes language tests (Cambridge Main Suite, TOEFL iBT, IELTS) keep them sequenced.
Each pretested item produces a row of statistics:
Items failing on any axis are rewritten and re-pretested or rejected outright. A typical operational programme rejects 30–50% of authored items at this stage.
Three patterns dominate:
The frame is calibration, not authoring. Pretesting tells you which items work on the trial sample under the trial conditions; it cannot retroactively fix items written against the wrong specification or items that target a construct the test does not aim to measure. Pretest data is necessary, not sufficient. Items that pass pretest can still fail operationally if:
For these reasons operational programmes pair pretest statistics with content review by subject matter experts and post-administration monitoring on live forms.
Operational language tests pretest at scale. Cambridge English's pretesting cycle administers candidate pretests through participating centres worldwide; items receive Rasch calibration against an anchor set already on the bank's scale, are reviewed item-by-item for content, and enter the bank with parameters attached. Items in the bank can be drawn into operational forms with target difficulty and content profiles, and forms can be equated across sittings via shared anchors. Reed (2013) treats this end-to-end pretest-to-bank pipeline as the architectural feature distinguishing standardised language tests from classroom assessments.