A learner corpus is a systematic, electronic collection of texts produced by second or foreign language learners. Unlike native-speaker corpora, learner corpora document Interlanguage, the developing language systems of L2 users, enabling large-scale analysis of error patterns, developmental features, and the effects of L1 background, proficiency level, and task type on learner production.
| Corpus | Full name | Size | Features |
|---|---|---|---|
| ICLE | International Corpus of Learner English | ~5.7 million words | Argumentative essays from advanced university learners; 25+ L1 backgrounds; directed by Sylviane Granger (UCLouvain, 1990-) |
| EFCAMDAT | EF-Cambridge Open Language Database | ~83 million words | Written submissions from EF online learners; 180+ nationalities; all proficiency levels |
| CLC | Cambridge Learner Corpus | ~55 million words | Exam scripts from Cambridge English exams; error-tagged |
| LINDSEI | Louvain International Database of Spoken English Interlanguage | ~1 million words | Spoken interviews from advanced learners; 11 L1 backgrounds |
| ICNALE | International Corpus Network of Asian Learners of English | ~2.3 million words | Written and spoken data from 10 Asian countries |
Learner corpus research follows the principles of Corpus Linguistics but adds learner-specific metadata: