Speaking test design is the principled construction of tasks, rating scales, and delivery procedures that elicit and measure a defined speaking construct. It is a form of performance assessment: unlike a reading or listening test, where the candidate selects among fixed options, a speaking test produces a sample of behaviour that a human (or increasingly a machine) must observe and rate against criteria. That structure introduces a chain of design surfaces: the construct, the task, the rating scale, the rater, and the delivery mode. Each of these can leak unwanted variance into the score. The two standard monographs are Sari Luoma's Assessing Speaking (Cambridge University Press, 2004) and Glenn Fulcher's Testing Second Language Speaking (Pearson, 2003).
What a speaking test claims to measure has expanded over four decades. Early constructs were largely about the individual speaker's linguistic resources: pronunciation, grammatical accuracy, vocabulary range, fluency. Communicative models added strategic and pragmatic competence, the ability to get a message across and to do so appropriately.
The sharpest recent development is interactional competence, the recognition that in talk between people, performance is jointly produced rather than owned by one speaker. Ana Maria Ducasse and Annie Brown's study of raters watching paired interactions (Language Testing, 2009) identified three features raters actually attend to: nonverbal interpersonal communication (gaze, body language), interactive listening (showing attention and engagement), and interactional management (handling topics and turns). Michelle May's work on co-constructed paired interaction sharpened the construct further and exposed its central tension, treated in the critique below. The construct decision is load-bearing: a test claiming to measure interactional competence but scoring only one speaker's linguistic output does not warrant its own claim, the same construct-validity logic that governs reading and listening tests.
Speaking tasks trade off between standardisation and authenticity, and each elicits a different slice of the construct.
| Task type | What it elicits | Trade-off |
|---|---|---|
| Interview (OPI-style) | Range across topics and functions | Asymmetric: examiner controls the talk |
| Paired / group | Peer-to-peer interactional competence | Score contaminated by the partner's level |
| Monologue / long turn | Sustained, planned production | Little interaction; closer to a presentation |
| Role-play | Pragmatic and functional ability in a scenario | Performance may not transfer to real use |
| Integrated (read/listen-then-speak) | Speaking grounded in source input | Confounds speaking with reading/listening |
A specification grounded in a TLU domain is what disciplines the choice: a test of academic seminar participation needs interactive tasks; a test of conference-presentation ability needs long turns. The mismatch between task and target domain is the most common design failure.
A speaking score is only as good as the scale behind it. The first decision is holistic versus analytic. A holistic scale gives a single overall judgement; it is fast and captures the gestalt of a performance but hides why a score was awarded and gives weak diagnostic feedback. An analytic scale rates separate criteria (fluency, accuracy, range, pronunciation, interaction), giving richer feedback and more consistent rating at the cost of speed and the risk that raters cannot actually attend to five things at once. The trade-offs are the general ones treated under analytic versus holistic scoring, here applied to live oral performance.
How the scale is built matters as much as its type. Fulcher distinguishes a measurement-driven approach, where descriptors are written top-down from a theory or an existing framework such as the CEFR, from a performance-data-driven approach, where descriptors are derived bottom-up from close analysis of actual recorded performances so that each band descriptor traces back to observable language. Fulcher's later "performance decision trees" (with Davidson and Kemp, 2011) push the data-driven logic further, replacing vague impressionistic wording with concrete behavioural decisions a rater can verify. Data-driven scales tend to describe real performance more faithfully; framework-driven scales tend to be more transferable across tests.
Two human facets beyond the candidate move the score. Raters differ in severity, in consistency, and in how they weight criteria, even after training, which reduces but never eliminates the variation. Interlocutors, the examiners or partners who talk with the candidate, shape the very performance being rated: their questions, accommodation, rapport, and language level change what the candidate produces. McNamara and Lumley (1997) showed that perceptions of interlocutor competence, candidate–interlocutor rapport, and even audio quality function as measurable facets affecting ratings.
The standard tool for untangling these is the many-facet Rasch model. By treating candidate ability, rater severity, task difficulty, and interlocutor effect as separate facets on a common scale, it estimates each independently and can adjust a candidate's score for the severity of the rater who happened to assess them. It also flags raters whose pattern of judgements is internally inconsistent (high "misfit"), which ordinary inter-rater reliability correlations cannot do.
A speaking test can be direct (live, face-to-face, with a human interlocutor) or semi-direct (the candidate responds to recorded or on-screen prompts, with answers captured for later rating, as in computer-delivered tests). Semi-direct delivery scales cheaply, standardises the prompt perfectly, and removes interlocutor variability. The cost is construct coverage: a candidate talking to a screen cannot demonstrate genuine turn-taking, negotiation, or interactive listening, so any test claiming interactional competence is weakened by the semi-direct mode. The two are not interchangeable measures of the same thing, which is the central tension of the computerised variants.
Because speaking tests sit at the top of high-stakes systems, their format drives classroom practice. A test built on monologue tasks pushes teachers toward presentation drilling; a test with genuine paired interaction pushes them toward communicative pairwork. This washback is a design lever as much as a side effect, and aligning task formats with the speaking the curriculum wants to develop is one of the few ways a test improves teaching rather than narrowing it.
The interlocutor and interaction effects are not noise to be cleaned up; they sit at the heart of a construct problem. In any interactive speaking test, performance is co-constructed, produced jointly by candidate and partner, yet scores must be assigned to individuals. May's research found that when paired interaction was asymmetric (one partner dominating), raters could not cleanly separate each candidate's contribution and struggled to award individual interactional-competence scores. A strong partner can scaffold a weaker one into a better performance, and a passive partner can starve a capable candidate of the chance to show their range. The fairness of an individual score is then hostage to who the candidate was paired with.
Rater reliability, even at its best, has a ceiling. Training improves consistency but cannot make raters identical, and severity differences remain large enough that an unadjusted score can swing a band on the luck of the rater draw. Many-facet Rasch adjustment helps only where the design supplies enough overlapping ratings to estimate the facets, which many operational tests cannot afford.
The deepest objection is that the move to interactional competence creates a measurement target that classical test theory was not built for. Rating a co-constructed, jointly owned performance as if it were an individual trait is a category problem, not just a precision problem. The field's response, scoring scales that explicitly include interaction features (after Ducasse and Brown) alongside discourse-analytic study of what actually happens in test talk, is ongoing rather than settled.
Teachers preparing learners for speaking tests should rehearse the interaction, not just the language. In paired formats, candidates can be coached to keep a struggling partner in the conversation and to claim their share of the floor, since both protect the individual score. Practising the specific task types and the rating criteria, especially the interaction descriptors that learners rarely see, turns opaque assessment into something learners can target.