Can AI Replace the Moderator’s Instinct for Follow-Up Questions?
Unmoderated usability testing has gained traction because it lets research teams gather data asynchronously, from a larger and more geographically diverse participant pool, at a lower cost. Participants can complete tasks on their own devices, in their own environments, without scheduling constraints. The trade-off has always been the absence of a human moderator who can adapt on the fly: probe unexpected behavior, clarify vague responses, and keep participants on track.
In a traditional unmoderated session, every participant receives identical, static instructions and a fixed questionnaire. Researchers must anticipate all relevant angles in advance, and the depth of insight is capped by how well participants articulate their own experience. Generative AI—particularly large language models—offers a potential bridge: using LLMs to converse with participants in real time, generate contextual follow-up questions, and extract more detailed feedback without needing a human in the loop.
This is an appealing vision of human-centered AI: the human remains the source of feedback, while the model handles the interactive probing. But there are significant unknowns. To address some of them, UXtweak research conducted a case study evaluating whether AI can generate follow-up questions that are meaningful and lead to valuable participant answers.
Testing GPT-4 in a Real Usability Scenario
The study focused on fundamental principles rather than on any particular commercial tool, since models and prompts change rapidly. The researchers used GPT-4 (the most current OpenAI model at the time), which they found handled complex prompts better than the newer GPT-4o in their circumstances. The test involved a usability session on an e-commerce prototype with a typical product-purchase flow. Three experimental conditions were compared:
- A static questionnaire with three pre-defined questions (Q1, Q2, Q3), serving as an AI-free control. Q1 was open-ended; Q2 and Q3 were direct follow-ups about usability issues and dislikes.
- Seed question Q1 followed by up to three generative AI follow-ups, replacing Q2 and Q3 entirely.
- All three pre-defined questions (Q1, Q2, Q3) each used as a seed for its own GPT-4-generated follow-up.
Informativeness and Emotional Response
The researchers measured the informativeness of each response: how useful it was for uncovering new usability issues. The results showed a significant drop in informativeness between seed questions and their AI-generated follow-ups. The follow-ups rarely surfaced a brand-new issue; they did, however, encourage participants to elaborate on points already raised.
Participant sentiment told another part of the story. Initial answers were neutral, but responses shifted negative as sessions progressed. This is partly expected for Q2 and Q3, which explicitly probed for usability problems and dislikes. But AI-generated follow-ups drew even more negatively than their seed questions—and often for different reasons. Frustration was a recurring theme in how participants interacted with the AI prompts.
Much of that frustration stemmed from redundancy. Participants frequently felt they were re-explaining things they had already said. The static control questionnaire saw 27–28% repeated answers (likely because Q2 and Q3 overlapped with themes already covered by the open-ended Q1). The AI-generated follow-ups performed only marginally better, at 21% repetition—a modest improvement against a baseline that had no capacity for context-awareness whatsoever.
Worse, when AI follow-ups were appended to every pre-defined question, the repetition rate climbed to 35%, and participants rated the questions as significantly less reasonable. Statements like “I already said that” and “The obvious AI questions ignored my previous responses” were common in the answers.
This is particularly telling because the GPT-4 prompt had full access to all available context from the seed question, its answers, and earlier follow-ups. The high prevalence of repetition within a single, self-contained conversation thread demonstrates that many generated questions were not sufficiently distinct—they lacked the direction needed to justify their existence.
Balancing Gains Against User Experience
The study’s takeaways for anyone considering AI in usability testing are mixed. On the positive side, generative AI can polish and deepen participant answers with contextual follow-ups, and the qualitative richness of the collected data can improve. On the negative side, the ability to uncover truly new issues is limited, and the risk of frustrating participants with repetitive or generic questions is real. When participants become annoyed with the questioning itself, their natural behavior and the relevance of their feedback can suffer, overshadowing any gains in elaboration.
Successes:
- Generative AI (GPT-4) is capable of refining and deepening answers through contextual questioning.
- Qualitative data can become more detailed.
Challenges:
- AI follow-ups struggle to surface issues beyond the scope of the pre-defined baseline questions.
- Repetitive or ill-directed questions can frustrate participants and compromise feedback quality.
Practical Caveats for AI-Assisted Testing
These findings highlight several concerns for practitioners, whether you are evaluating an off-the-shelf AI testing tool, writing your own prompts, or deploying your own model. Context is critical; the model had full conversation context in this experiment yet still produced redundant follow-ups. Accuracy of tone and direction is equally important. An AI that asks questions participants perceive as repetitive or irrelevant can introduce a negative bias that colors the entire session.
Researchers should also be aware of the systemic nature of AI models. Because they are neural-network black boxes, their raw outputs are not transparently traceable to specific rules; a prompt that performs well today may degrade with the next model update. Any AI-powered testing approach requires ongoing validation on actual participant responses rather than a one-time evaluation. In this study, the AI’s strongest role was in extracting slightly richer elaboration of known issues—not in discovering new terrain. That alone suggests current AI follow-up generation may be best treated as a supportive tool, not a standalone replacement for a well-designed static questionnaire or, where depth is critical, a trained human moderator.
Context Is the Missing Ingredient
Across the failures observed in our study, one root cause stood out: the AI lacked proper context. Nearly every problematic follow-up question GPT-4 generated could be traced back to a missing or misunderstood piece of situational information. For practitioners evaluating third-party tools or building their own AI-driven moderation systems, the following taxonomy serves as a practical checklist. Use it to judge whether a model can ask context-aware questions before letting it interact with real participants.
- General Usability Testing Context. AI questions should respect basic usability research principles. They should not be leading, should not solicit design suggestions, and should not ask participants to predict behavior in purely hypothetical scenarios. These errors appeared in our study despite the principle being obvious.
- Usability Testing Goal Context. Every study has specific objectives tied to design stage, business goals, or the features under evaluation. Follow-up questions are a drain on participant time and should stay on topic. In our prototype test featuring placeholder product photos, questions probing opinions of those fake products were useless and wasteful.
- User Task Context. The nature of your tasks—goal-driven or exploratory—must shape the questions that follow. Open tasks invite motivational questions. But when a participant is following a clear directive, asking why they placed a specific required item into the cart makes both the AI and the researcher look incompetent.
- Design Context. Without detailed knowledge of the tested artifact—prototype, mockup, website, or app—follow-up questions can become unanswerable or insulting. An AI asking why a participant believed a fact that was prominently displayed in the UI fails on both counts. Design context also helps focus questions on the most relevant interface aspects.
- Interaction Context. Where design context describes what participants could see and do, interaction context captures what they actually did, including consequences. Video of the session and audio of think-aloud protocols belong here. This context allows the AI to build on participant actions, such as probing why a task was not completed even when the participant believes they succeeded.
- Previous Question Context. Participants naturally form associations between questions they hear in sequence. A skilled human moderator recognizes when an earlier answer already covered a planned question and moves on. AI models should do the same to prevent the session from descending into repetitive questioning.
- Question Intent Context. Participants often answer the letter of a question, not its spirit. Asking an open question and receiving a technically valid but off-target reply is common. The AI needs to recognize when it failed to retrieve the intended information and re-ask from a fresh angle.
When evaluating a third-party tool, ask whether you can feed all these contextual layers in explicitly. The model cannot invent what it was never given.
If AI does not have an implicit or explicit source of context, the best it can do is make biased and untransparent guesses that can result in irrelevant, repetitive, and frustrating questions.
Supplying context is not the same as the AI respecting it. In our study, even when conversation history was provided within a question group, repetition remained a persistent issue. The most direct way to test any model's contextual sensitivity is to converse with it in ways that rely heavily on context—natural conversation does this already, so this should not be hard. Pay attention to which context types the model handles and which it drops.
Multiple Contexts Create Conflict
The hardest problem is that these context types mix. A human moderator may deviate from best practice—asking a narrow question instead of an open-ended one—when research goals demand it, consciously accepting the tradeoff. Our study showed that when the AI paired generically open-ended follow-ups with equally open-ended seed questions, without shifting perspective, the result was repetition, irrelevance, and frustration.
The ability to resolve such contextual conflicts appropriately is a meaningful yardstick for any AI follow-up generator. With so many combinatorial possibilities, judging the AI's sensitivity at every intersection is the real challenge. Researcher control also matters: decisions grounded in the researcher's vision should never be delegated wholesale.
The fine-tuning of the AI models to achieve an ability to resolve various types of contextual conflict appropriately could be seen as a reliable metric by which the quality of the AI generator of follow-up questions could be measured.
A hybrid approach is the pragmatic path forward. Static, pre-scripted questions carry the weight of researcher intent, while AI-driven questions add flexibility where context permits. Complementary strengths and weaknesses may unlock insights that neither format captures alone.
Participant Perceptions and Ethical Grounding
Contextual sensitivity matters beyond question quality. The wider AI backlash, driven by valid worries about utility, ethics, data privacy, and environmental impact, means some usability participants will be skeptical or even openly hostile toward an AI moderator. Tools must be introduced explicitly as reasonable and helpful, not as replacements for human judgment.
Ethical research principles still govern. Participant data must be collected and processed with full consent, and sensitive data must not be used to train AI models without permission. An AI that ignores context will not only produce bad questions—it may produce ethically compromised research.
Where the Opportunity Actually Lies
Will AI erase the difference between moderated and unmoderated testing? Not yet. The evidence is encouraging: when AI follow-up questions work, participants open up, and essential details surface. For any researcher who has ever analyzed vague feedback and wished for one clarifying question, an automated follow-up is an alluring prospect.
But reverting the very process that resolves ambiguity is dangerous. Blindly adding AI introduces its own biases precisely because relevance depends on all the context types outlined above. Replacing human oversight with an unexamined model simply trades one blind spot for another.
Humans + AI = Better Insights
The balanced recommendation is steady integration over replacement. UX researchers and designers should keep learning how to deploy AI as a discovery partner, using this taxonomy as an early guide to its known weak points, monitoring requirements, and improvement opportunities. The viable future of unmoderated testing lies in making AI genuinely context-aware, not in making it autonomous.




