Conversations are wonderfully unpredictable. One person asks a question. The other responds. The conversation changes direction. New questions emerge. Ideas develop naturally.
That's exactly what makes conversation such a valuable way to assess speaking. It’s also exactly what has made it so difficult to standardize.
For decades, language testing has faced a tradeoff. Human interviews can capture authentic interaction, but every conversation unfolds a little differently. Computer-delivered speaking tests create consistent testing conditions, but they typically replace dialogue with a series of independent monologues.
A new paper from the Duolingo English Test researcher team, recently published in Language Testing, explores whether advances in AI and psychometrics can finally bridge that divide.
Standardizing a conversation requires solving two new problems
Large language models have made conversational AI commonplace. But building an adaptive speaking task for a high-stakes standardized assessment requires solving a very different problem.
A standardized test must ensure that every score means the same thing, regardless of which questions a test taker receives. That becomes much harder once different people begin having different conversations.
The researchers identified two challenges that needed to be solved simultaneously:
- How should the system decide which follow-up question to ask next?
- How can scores remain comparable if every conversation is different?
These questions sit at the intersection of assessment science, psychometrics, applied linguistics, and artificial intelligence.
Teaching AI to ask better follow-up questions
Adaptive speaking assessment is a testing approach that changes follow-up questions based on a test taker's responses while maintaining comparable scores across different conversations.
Adaptive conversations require the system to evaluate each response before deciding what to ask next. Every follow-up question depends on what came before.
After every response, the system evaluates how completely the test taker answered the current question using question-specific task completion rubrics created and reviewed by assessment experts. Rather than simply analyzing vocabulary or grammar, the AI identifies which key ideas the response addressed and which remain unexplored.
That information guides the next question. If a test taker has already discussed one aspect of a topic, the system selects a follow-up that explores a different idea. If their response suggests a higher level of proficiency, the conversation can become more cognitively demanding.
In effect, the conversation adapts as a human interviewer might—while following the same underlying decision process for every test taker.
Different conversations can still produce comparable scores
Adaptive conversations introduce another challenge: If two test takers receive different follow-up questions, how can their scores still be fairly compared?
The answer comes from psychometrics.
Rather than treating every prompt as equally difficult, the researchers used Item Response Theory (IRT) to account for differences among prompts. This allows the scoring system to separate question difficulty from test taker ability, ensuring that scores remain comparable even when conversations differ.
The results showed that this approach improved test-retest reliability compared with conventional scoring. In fact, incorporating IRT produced an improvement equivalent to increasing the effective test length by roughly 31%—without asking test takers additional questions.
The conversations also sounded more like conversations
The researchers also examined whether the new task actually elicited different language.
Using multidimensional discourse analysis, they compared responses from Interactive Speaking with responses from a traditional monologic speaking task.
The adaptive task encouraged language that was more conversational, reflecting the kinds of interactive language people use in real discussions rather than isolated speeches. Measuring speaking through conversation allows test takers to demonstrate abilities that monologic tasks simply cannot elicit.
A new direction for speaking assessment
This new paper demonstrates that advances in generative AI can support one of language assessment's longstanding goals: combining interactive, dialogic exchange with the consistency required for large-scale standardized testing.
Rather than choosing between human interviews and scripted monologues, assessment developers now have evidence that adaptive, computer-delivered conversations can provide a third path that balances interactivity, fairness, and standardization.