Automated writing evaluation (AWE) is software that analyses a piece of writing and returns a score, feedback, or both, using natural language processing rather than a human reader. The label spans two purposes that pull in different directions: summative automated essay scoring (AES), which produces a single grade for a test, and formative feedback, which flags errors and suggests revisions while a learner is still drafting. The same underlying engine can do either, but the stakes, the validity demands, and the classroom value differ sharply.
The field is older than the chatbots it now sits beside. Ellis Page built Project Essay Grade (PEG) in 1966, predicting teacher scores from surface countables — essay length, word length, punctuation — through linear regression. Commercial systems arrived in the late 1990s: Vantage Learning's IntelliMetric began development in 1996 and scored essays commercially from 1998; Educational Testing Service's e-rater, led by Jill Burstein, went live in February 1999. The Intelligent Essay Assessor, from a team under Thomas Landauer, instead used latent semantic analysis to judge content against reference texts. e-rater powers Criterion, ETS's classroom-facing tool.
A separate strand targets learners directly rather than testing bodies. Grammarly and Cambridge's Write & Improve give draft-stage feedback to language learners, and Pigai scores English writing for millions of students across China. The generative-AI shift has reshaped both strands: large language models now produce fluent, targeted commentary on content and organisation that earlier feature-based engines could not, blurring the line between AWE and a conversational writing assistant of the kind surveyed under AI in Language Teaching.
Classic AES does not read. It measures proxies, features the system is programmed to count, and weights them to approximate human ratings on a training set. Lexical sophistication, sentence and essay length, error rates for spelling and grammar, and discourse markers all feed the model. The score is a statistical prediction of what a human would have given, not a judgement of meaning. This is the engine's core constraint, captured bluntly in the critical literature: the machine notices only what it was built to notice, and it cannot weigh logic, accuracy, or whether the argument is true. Generative models loosen this by attending to meaning, but they introduce their own opacity and inconsistency.
For formative feedback, the evidence is real but modest. Mark Warschauer and Paige Ware's framing of the classroom research agenda found that students revise most where AWE is strongest (mechanical errors, garbled sentences, capitalisation) and that effects depend on how a teacher integrates the tool, not on the software alone. Matthew Stevenson and Aek Phakiti's critical review concluded that AWE feedback produces modest gains on the quality of the immediate text but offers little evidence of transfer to general writing improvement. Jiangping Li, Stephanie Link, and Volker Hegelheimer found teachers warmer toward AWE than students, with learner attitudes tracking proficiency, teacher stance, and how the software is folded into instruction. The recurring lesson matches the broader AI finding: AWE works best alongside teacher feedback, not in place of it.
The sharpest critique is that AES rewards surface features and is gameable, so the score can drift from the construct it claims to measure. Les Perelman's BABEL Generator demonstrated this by producing deliberately incoherent prose stuffed with long words and the right proxies, which earned high marks from several engines. Criterion's reported accuracy on the errors it does target is limited, identifying only a fraction of preposition and article errors, so feedback on meaning, content, and pragmatics is weaker still than its surface coverage. The National Council of Teachers of English's 2013 position statement on machine scoring argues that computers cannot judge logic, evidence, ethical stance, or argument, that scoring narrows instruction toward what the machine can count, that it is easily gamed, and that it disadvantages multilingual writers and under-resourced schools. Overreliance is a further risk: learners may revise to satisfy the algorithm, and the resulting washback pushes toward formulaic, feature-heavy prose rather than genuine communication.
Treat AWE as a first-pass filter, not a verdict. Let it absorb the mechanical, high-frequency error correction (spelling, agreement, repeated slips) so teacher attention goes to argument, evidence, and voice that no engine reads well. Anchor it inside a drafting cycle, where its strength at flagging revisable surface errors fits the logic of process-oriented writing. Be explicit with learners that the score reflects countable features, and show them the gap by discussing what a strong essay does that the tool misses, which doubles as critical-AI training. Keep AES out of high-stakes writing assessment as a sole rater; pair it with human marking and audit for bias against multilingual and dialect-varied writing.