Item difficulty (also called the facility value or p-value) is the proportion of test takers who answer an item correctly. Despite its name, a higher value means an easier item.
The value ranges from 0 (no one answered correctly) to 1 (everyone answered correctly). It is one of the two core statistics in item analysis, alongside item discrimination.
| P-Value | Interpretation |
|---|---|
| 0.90–1.00 | Very easy: nearly all candidates correct |
| 0.70–0.89 | Easy |
| 0.30–0.69 | Moderate: optimal range for most items |
| 0.10–0.29 | Difficult |
| 0.00–0.09 | Very difficult: investigate for possible flaws |
The ideal difficulty depends on the test's purpose:
Extremely easy and extremely difficult items cannot discriminate well. If everyone gets an item right (p = 1.0) or everyone gets it wrong (p = 0.0), the item tells us nothing about individual differences. Items in the moderate range (0.30–0.70) have the most room to differentiate between stronger and weaker candidates.
However, an item with p = 0.85 can still have acceptable discrimination if the 15% who got it wrong are consistently from the low-scoring group.
In CTT, item difficulty is sample-dependent: the same item will have different p-values when administered to groups of different ability levels. This is a fundamental limitation: an item is not inherently "difficult" or "easy" in isolation; it is difficult for a particular group.
Item Response Theory (IRT) addresses this limitation by modelling item difficulty as a parameter independent of the sample, but CTT's simplicity makes p-values practical and widely used in classroom and institutional testing contexts.