The halo effect is a rater bias in which judgements on one rating criterion contaminate judgements on others, producing scores that cluster more tightly across criteria than the underlying performance warrants. In analytic scoring of writing or speaking, a candidate strong on fluency tends to be rated equally strong on accuracy and lexical resource even when the performance evidence diverges; a candidate with a heavy accent receives lower scores across criteria including those that have nothing to do with pronunciation. The result is a flattened score profile: the analytic rubric collects multi-criterion data on paper but yields effectively one-dimensional scores.
The construct cost is that diagnostic feedback collapses. An analytic rubric exists to tell candidates which dimension of their performance is weak; a halo-affected score cannot. Reliability rises spuriously (raters agree because they are all collapsing the same way) while validity falls, since the score correlates with general impression rather than the specific criterion.
Three sources are well attested:
Eckes (2011, 2015) and Knoch, Read and von Randow (2007) document each in operational language-rating contexts.
Many-facet Rasch measurement is the standard diagnostic. The model expects candidate-by-criterion residuals (the rated profile a candidate gets across criteria) to vary in the way the rating scale predicts. Halo shows up as systematically flattened candidate profiles: residuals are smaller than expected, criterion scores correlate more tightly than the model expects given the rubric structure. Bechger, Maris and Hsiao (2010) formalise the test for unexpectedly flat patterns. Gross between-criteria correlations from inter-rater data give a coarser indicator, but distinguish poorly between halo and a genuinely correlated underlying construct.
| Mechanism | What it does |
|---|---|
| Criterion-specific descriptors | Forces raters to attend to one dimension at a time; operational rubrics put descriptors in fully separated columns |
| Rater training with anchored samples | Calibrates each rater to the criterion-specific construct via worked examples, not just rubric exposure |
| Sequential rating | Rate criterion 1 across all candidates, then criterion 2; separates judgements in time and reduces global-impression bleed |
| Double rating + adjudication | Two raters independently; large discrepancies trigger third rating. Catches halo-flattened raters as outliers |
| Rater monitoring | MFRM-based ongoing check for unusual flatness; raters showing chronic halo are recalibrated or retired |
The IELTS Speaking rating model uses a four-criterion analytic rubric (Fluency and Coherence, Lexical Resource, Grammatical Range and Accuracy, Pronunciation) with criterion-by-criterion training and routine MFRM monitoring of rater patterns. The architectural choice of keeping criteria narrow and behaviourally specified is itself a halo countermeasure.
The hard analytic problem is distinguishing rater halo from genuine inter-criterion correlation in the construct. Speaking performance criteria do correlate: a candidate strong on grammatical range often is strong on lexical resource, because both draw on shared underlying L2 proficiency. The methodological question is whether observed flatness exceeds the construct-level correlation. In-Gee Han (2020) shows that the order in which criteria appear on the rubric (whether fluency or accuracy is read first) itself modulates halo, evidence that the bias has a procedural component independent of the construct correlation.
| Bias | Pattern |
|---|---|
| Halo effect | Criteria collapse together within candidate |
| Severity / leniency | Rater consistently scores lower / higher than peers across all candidates |
| Central tendency | Rater compresses scores toward the middle of the scale, avoiding extremes |
| Range restriction | Rater uses only a subset of available bands, regardless of performance spread |
| Rater drift | Rater's criteria shift across a long rating session |
Many-facet Rasch designs estimate severity, central tendency, and halo simultaneously, which is why operational language tests with high-stakes scoring (IELTS, Pearson Test of English, Cambridge Main Suite) have converged on it as the rater-monitoring framework.