The Oral Proficiency Interview (OPI) is a structured, face-to-face speaking test in which a trained tester elicits an oral sample through conversation and rates it against a fixed proficiency scale. The best-known version is the ACTFL OPI, administered as a roughly 20-to-30-minute telephone or in-person interview and used across US universities, government agencies, and certification bodies. Its enduring influence on how the field thinks about oral assessment is matched only by the depth of the critique it has attracted.
The OPI grew out of US government language testing. In the mid-1950s the Foreign Service Institute (FSI), under its dean Henry Lee Smith, developed a speaking test and a proficiency scale (the work is associated with Claudia Wilds, in consultation with Harvard's John B. Carroll). The scale ran from 0 (no functional ability) to 5 (educated native speaker). In 1968 several federal agencies wrote formal descriptions of these levels across the four skills, and the scale was later revised under the Interagency Language Roundtable (ILR) to add the intermediate "plus" levels. In 1982 the American Council on the Teaching of Foreign Languages (ACTFL) adapted the government instrument for academic use, developing the ACTFL Proficiency Guidelines (Novice, Intermediate, Advanced, Superior, with sublevels) that the ACTFL OPI rates against.
The interview follows a deliberate four-phase shape designed to find the ceiling of a candidate's sustainable performance:
A role-play can be inserted as a level check or probe to elicit functions that ordinary conversation does not reach, such as managing a complication or making a complex request. The tester moves between checks and probes iteratively, then assigns a single global rating on the ACTFL scale.
The OPI is a performance test scored holistically against criterion-referenced descriptors. A rating is not a percentage or a sum of subscores but a judgement that the candidate's sustained performance matches the functional, content, accuracy, and text-type profile of one band. Functions ("can narrate and describe in past, present, and future"), content domains, and text type (words → sentences → paragraphs → extended discourse) are the anchors. Raters are certified through ACTFL training, and a second rating is typically used for reliability.
To meet large-scale demand, ACTFL introduced the OPIc (Oral Proficiency Interview – Computer) in 2006. It replaces the live tester with an avatar named Ava who delivers recorded prompts; the candidate's spoken responses are recorded for later human rating. A background survey and a self-assessment at the start route the candidate to one of several test forms targeting a proficiency range, so the prompts roughly match the candidate's level and interests. OPIc standardises delivery perfectly and scales cheaply, but in removing the human interlocutor it removes genuine interaction, a delivery-mode trade-off that the critiques below sharpen.
The OPI's prominence made it the most scrutinised oral test in applied linguistics, and the scrutiny has been damaging.
The construct-validity critique (Bachman 1988). Lyle Bachman's "Problems in Examining the Validity of the ACTFL Oral Proficiency Interview" (Studies in Second Language Acquisition, 10(2), 149–164) argued that the OPI cannot be properly validated as a measure of communicative language ability for two reasons. First, it confounds traits with test methods: the abilities it claims to measure are tangled up with the elicitation procedures used to measure them, so a score reflects both the candidate's ability and the way it was elicited, with no way to separate them. Second, it yields a single global rating that has no theoretical or empirical basis in a model of language ability. His prescription, distinguish abilities from methods and build validity evidence during test development, set the agenda for a decade of research.
The discourse-analytic critique: the OPI is not a conversation. A second line of attack came from researchers who recorded and analysed what actually happens in OPIs. Leo van Lier's "Reeling, Writhing, Drawling, Stretching, and Fainting in Coils: Oral Proficiency Interviews as Conversation" (TESOL Quarterly, 1989, 23(3), 489–508) showed that the OPI's claim to elicit conversational ability is undermined by its own structure. The tester controls topic, turn allocation, and pacing; the relationship is asymmetric in a way ordinary conversation is not. The candidate rarely gets to initiate, change topic, or take the kind of conversational risks that reveal genuine interactional ability. What the OPI elicits is competence in being interviewed, its own speech genre, not competence in conversation.
Marysia Johnson's The Art of Non-Conversation (Yale University Press, 2001) extended this with a full discourse analysis, asking directly what kind of speech event the OPI is, and concluding it is neither everyday conversation nor a neutral elicitation but a distinct, tester-dominated genre. Because the turn-taking and topic control are the tester's, the candidate cannot display the interactional management that real talk requires, so the test under-samples the very competence its label promises. Johnson grounds an alternative in Vygotskian sociocultural theory, treating spoken interaction as jointly constructed rather than a fixed individual trait. Agnes Lazaraton's conversation-analytic studies documented the concrete mechanism: testers routinely accommodate, supplying scaffolding, rephrasing, and support that varies from candidate to candidate, which both helps weaker candidates and undermines the assumption that each is being measured under standard conditions.
The cumulative force of these critiques is that the OPI measures something real and useful, structured, ratable oral performance, but that the something is narrower and more interview-specific than its claim to assess general oral proficiency or conversational ability implies. Defenders counter that the OPI's strong inter-rater agreement, decades of operational refinement, and practical utility for placement and certification justify its survival, and it remains widely used. The standoff is a textbook case of construct validity (does the test measure what it claims?) pulling against reliability and practicality (does it measure something consistently and usefully?).
Candidates preparing for an OPI benefit from understanding its architecture: the tester is deliberately pushing toward a ceiling, so the probes are meant to be hard and a candidate need not "pass" every question. Practising the elicited functions, narrating across time frames, handling a complication in a role-play, sustaining paragraph-length discourse, targets exactly what the bands describe. Because the format rewards initiative and extended turns, learners should be trained to volunteer detail and develop their answers rather than wait to be questioned, which is also the part of the OPI that least resembles real conversation and therefore needs the most explicit rehearsal.