Writing test design is the principled construction of prompts, tasks, and scoring procedures that elicit and measure a defined writing construct. Unlike a reading or listening test, where a candidate selects among given options, a writing test asks the candidate to produce an extended performance, then asks a human or machine to judge it. This makes writing a performance assessment, and it moves the central design problem from item-and-distractor engineering to the harder terrain of task elicitation and rating consistency. The foundational synthesis is Sara Cushing Weigle's Assessing Writing (2002).
Everything starts with deciding what the test is meant to measure. A construct narrowly defined as control of grammar and mechanics yields very different tasks from one that includes the ability to organise an argument, adapt to an audience, or synthesise sources. Weigle frames the construct against a target language use domain: the real-world writing the score is supposed to predict. A test for university admission should sample the academic writing students will actually do; a test for a customer-service role should sample workplace correspondence. The danger is construct underrepresentation, sampling too thin a slice of writing to support the inferences users want to draw, and construct-irrelevant variance, where something other than writing ability (topic knowledge, handwriting, typing speed) leaks into the score. A defensible test pins the construct in a specification before any prompt is written.
A prompt has to elicit comparable performances from every candidate while leaving room for the targeted writing operations to show. Topic choice is the most under-managed confound: a prompt that rewards prior knowledge measures what candidates know rather than how well they write. Effective prompts are accessible to all candidates without advantaging any, specify audience and purpose clearly enough to constrain register, and avoid culturally loaded content that would disadvantage some test-takers in an international population.
The major design fork is independent versus integrated tasks. An independent task asks the candidate to write from their own knowledge and opinion in response to a prompt. An integrated task supplies source material (a reading passage, a lecture, a chart) and asks the candidate to write in response, as in the reading-into-writing and listening-into-writing tasks of the TOEFL iBT. Integrated tasks raise authenticity, since most real academic and professional writing responds to sources rather than springing from a cold prompt, and they reduce the topic-knowledge confound by giving everyone the same content. Their cost is construct contamination: a low score may reflect weak reading or listening rather than weak writing, muddying the inference the test is meant to license.
Three families dominate. Holistic scoring assigns a single score for overall quality against a rating scale of band descriptors. It is fast and reflects a reader's global impression, but it hides why a script earned its score and can let one salient feature dominate. Analytic scoring rates several dimensions separately. The most influential analytic instrument is the Jacobs et al. ESL Composition Profile (1981), which scores five weighted dimensions on a 100-point scale: content (30), language use (25), organisation (20), vocabulary (20), and mechanics (5). Analytic scoring yields a diagnostic profile and, as Weigle notes, tends to deliver stronger reliability and construct validity, at the cost of being slower and risking a halo effect across categories. Primary-trait scoring judges a script against the single rhetorical feature the task was designed to elicit (for example, persuasiveness in a persuasive task), which is tightly valid for that task but hard to generalise. The choice between holistic and analytic is a practicality-versus-information trade-off, and Weigle's summary is that holistic wins on practicality and authenticity while analytic wins on reliability and validity.
A writing score is only as good as the agreement behind it. Because human judgement varies, design has to engineer consistency through rater training, clear descriptors, benchmark scripts, and double-marking with adjudication of discrepancies. Inter-rater reliability estimates the resulting agreement, but a simple correlation hides systematic rater effects. The many-facet Rasch model, operationalised in software such as FACETS, models candidate ability, task difficulty, and rater severity (along with rater-topic and rater-category interactions) on a common scale, so a candidate's estimate can be adjusted for the leniency or harshness of whoever happened to mark them. Research with these models repeatedly finds that significant severity differences survive even extensive training, which is why severity is corrected statistically rather than assumed away.
The default operational format is the timed impromptu task: a single draft on an unfamiliar prompt under time pressure, prized for standardisation and cheap administration. Its validity is contested. As Weigle and others argue, writing produced in one sitting on an unseen topic poorly represents how writing is actually done and taught, where planning, feedback, and multiple drafts are the norm, and a single sample generalises weakly to the wider universe of genres, audiences, and purposes a candidate will face. Portfolio assessment answers this by collecting multiple pieces written across contexts under realistic conditions, with opportunities to revise on feedback, raising authenticity. The trade-off is practicality and reliability: portfolios are costly to manage and hard to standardise, which is why they remain largely a classroom tool and rarely anchor large-scale, high-stakes programmes.
Automated essay scoring systems such as ETS's e-rater score from measurable surface features: essay length, lexical diversity, syntactic complexity, grammar, and mechanics. They deliver instant, perfectly consistent marking at scale, which is their appeal. The validity critique is sharp. Critics argue the features these systems measure are not the writing construct; word and sentence length are proxies, not the thing itself. The systems are gameable, demonstrated most vividly by Les Perelman's BABEL Generator, which produces incoherent prose engineered to score highly, and they are blind to argument quality, originality, and audience awareness. The National Council of Teachers of English has formally opposed machine scoring of high-stakes student writing on these grounds. The recurring design hazard is construct narrowing: optimising a test for what the machine can count quietly redefines writing as the sum of those countable features.
Treat any writing test as a sample, and ask what universe of writing it claims to generalise to before trusting its score. Prefer prompts that constrain audience and purpose and that no candidate can answer from prior knowledge alone, and consider integrated tasks where source-based writing is the real target, accepting that they blur the reading-writing boundary. Score analytically when diagnostic feedback or defensible reliability matters, holistically when throughput dominates and a single global judgement suffices. Invest in rater training but do not expect it to erase severity differences. When the stakes are high, the strong washback consequence of a single timed draft pushes classrooms toward one-sitting essay drilling, so weigh that backwash against the practicality that made the format attractive in the first place.