Construct-irrelevant variance (CIV) is the portion of test-score variance attributable to factors that are not part of the construct the test claims to measure. The term was formalised by Samuel Messick in his unified theory of validity (Messick 1989) and stands as one of the two primary threats to construct validity, alongside construct underrepresentation.
Where construct underrepresentation means the test leaves out parts of the construct it should cover, CIV means the test picks up signal it should not. Both invalidate score interpretation. Messick's framing reshaped the field by treating validity not as a property of the test but as a property of score interpretations and uses, which makes both threats actionable in development rather than only diagnosable after the fact.
Several CIV sources recur in reading and writing assessment:
CIV is invisible if you only look at total scores. A test can be reliable, internally consistent, and predictive of outcomes while still measuring partly the wrong thing. The diagnostic moves are differential item functioning (DIF) analysis, which compares item-level performance across demographic groups; expert review against the construct definition; and triangulation of test scores against independent measures of the construct. AI-generated items add a fresh CIV concern: the generator's stylistic biases may introduce variance correlated with training-data demographics rather than with the construct.
In test design the lesson is upstream, not downstream. CIV that is built into the specification cannot be fully removed by item analysis. A construct definition tight enough to name what is not part of the construct, paired with a TLU domain description that pins authentic task characteristics, is the cleanest defence.