Anchor items are common items administered with two or more test forms so that the forms can be put on the same score scale. They make scores from different forms (different administrations, different item sets, different candidate cohorts) comparable. Without anchors, two forms of the same test give two unrelated scores; with them, the scores live on a shared scale and a single proficiency report is meaningful across sittings.
The mechanism is statistical. Item difficulties estimated on the anchor set bridge between the two cohorts: differences in item performance attributable to the people are separated from differences attributable to the items. Both classical test theory and IRT equating methods rest on this separation. The canonical language-testing example is Cambridge English's banked-item architecture, in which items are pretested, calibrated against an anchor on a common Rasch scale, then drawn into operational forms with target difficulty profiles.
Internal vs. external anchors
| Type | Where the anchor sits | Counted in the score? |
|---|
| Internal | Embedded inside the operational test, indistinguishable to candidates | Yes; anchor responses score |
| External | Administered as a separate section identified to the analyst, sometimes labelled as a trial section | No; anchor responses are excluded from the reported score |
Internal anchors are more common in language tests because they preserve test-taking conditions: candidates cannot detect anchor items and adjust effort. External anchors are more common where item exposure is a security concern and the anchor needs to be retired faster than the operational pool.
What makes a good anchor
Kolen and Brennan's (2014) treatment of equating defines three properties an anchor set should meet:
- Mini-test property. The anchor should mirror the operational form in content coverage and difficulty distribution. A reading anchor of all easy literal-comprehension items will not equate forms whose operational sections include inference and function items.
- Length sufficiency. Anchors carry equating error proportional to their length; rules of thumb run from 20% of the operational test for similar-cohort designs to 30%+ for non-equivalent groups.
- Stable item parameters. The anchor items should not show DIF across the two cohorts. An item whose difficulty has drifted between administrations cannot serve as a fixed reference point.
Anchor items that fail these properties produce equating that is biased rather than noisy: every reported score on the new form drifts in the same direction.
Equating designs
Three designs use anchors:
- Common-items non-equivalent groups (CINEG). Two cohorts take different forms; an anchor is shared between the forms. The anchor carries the cohort difference. This is the workhorse of large-scale language testing: IELTS, TOEFL iBT and Cambridge Main Suite all rely on variants.
- Single-group with counterbalancing. One cohort takes both forms in randomised order; the anchor is implicit in the full overlap. Rare in operational language testing because of fatigue and exposure problems.
- Random groups. Forms are spiralled across candidates from a single sitting; no formal anchor needed. Common at the pretest stage, less so for live equating across sittings.
Calibration in an item bank
The standard pipeline runs anchors twice. New items are pretested alongside operational items that already sit on the bank's calibrated scale; those operational items act as anchors that place the new items on the scale. Items are then released into the operational pool with known parameters. Each subsequent administration carries its own anchor back to the scale, so drift is monitored over time. When anchor items begin to show DIF or systematic difficulty change (exposure, curriculum shift, copying), they are retired and replaced.
Common failure modes
- Anchor too easy or too hard for the cohort. If most candidates pass or fail the anchor, the equating estimate has wide standard errors at the relevant ability range.
- Curricular drift. An anchor written for a 2010 cohort may be culturally or topically dated for a 2025 cohort; what was a stable reference becomes a moving one.
- Exposure. Repeated use leaks anchor content; coaching companies cluster around recognisable items. Operational programs replace exposed anchors before the leakage shows up in DIF analyses.
- Construct underrepresentation. An anchor that omits an entire skill area equates well on what it covers but cannot detect drift in what it does not.
References
- Kolen, M. J., & Brennan, R. L. (2014). Test Equating, Scaling, and Linking: Methods and Practices (3rd ed.). Springer.
- Cambridge ESOL. (2011). Research Notes Issue 43: Item banking and pretesting. Cambridge English Language Assessment.
- Liu, J., Curley, E., & Low, A. (2014). Test score equating using discrete anchor items versus passage-based anchor items: A case study using SAT data. ETS Research Report Series, RR-14-29.
- Holland, P. W., & Dorans, N. J. (2006). Linking and equating. In R. L. Brennan (Ed.), Educational Measurement (4th ed., pp. 187–220). Praeger.