Simulating the Impossible: The Wizard of Oz Method in UX Research

The gap between a promising concept and a usable product is often where good ideas go to die. The Nintendo Power Glove is a cautionary example: it sold over a million units after its late 1989 release, yet was discontinued less than a year later. Developed in about eight weeks, the device proved cumbersome and unintuitive. Users struggled with syncing the glove to specific games, a process that required coding moves into preset buttons and memorizing which button triggered which action. The software built for it sold poorly, and its compatibility with existing popular titles was minimal. With hindsight, the idea was ahead of its time, but the execution failed because no one had verified that people could actually use it as intended.

The Nintendo Power Glove
The Nintendo Power Glove: A Fistful of Frustration. (Image source: ACMI) (Large preview)

Traditional UX research methods like surveys and interviews would not have exposed those usability flaws. They rely on self-reporting rather than observed interaction. What was needed—and what remains a valuable tool today—is a way to test an unbuilt system with real users. The Wizard of Oz (WOZ) method does exactly that. A human operator, hidden from the participant, simulates the behavior of a fully functional system. The user believes they are interacting with a real product, but their experience is being orchestrated in real time by a researcher pulling the strings.

The name comes from Frank L. Baum’s classic tale, in which the omnipotent Wizard is revealed to be an ordinary man manipulating events from behind a curtain. Similarly, WOZ emulates technologies that may not yet exist—or may never be built—in a way that convincingly simulates a working tool. This allows teams to explore user needs, validate concepts early, and reduce development risk before committing engineering resources to a flawed idea.

For the Power Glove, researchers could have had participants wear a mock device—even a cardboard glove with controls drawn on it—and simulate the experience of coding moves and switching between games. That simple test would have revealed the illogical requirement of asking everyday users to program hardware to work with their games, the frustration of reprogramming it each time the game changed, and the awkwardness of the physical control layout.

The origins of the method trace back decades. Jeff Kelley credits himself with coining the term "Wizard of Oz" in 1980 for his dissertation research. However, Paula Roe attributes the technique to Don Norman and Allan Munro, who reportedly used it in 1973 to test an automated travel assistant for airports. Both narratives converge on IBM's later use of the method to study a speech-to-text application called The Listening Typewriter.

Wizard of Oz testing: The listening typewriter IBM 1984
Wizard of Oz testing: The listening typewriter IBM 1984. (Image source: CDS) (Large preview)

This article examines the core principles of WOZ, explores advanced applications from practical experience, and demonstrates its value through real-world cases, including its growing relevance to agentic AI. For UX practitioners, it is another way to unlock user insights and build human-centered products grounded in observed behavior.

Inside The Wizard Of Oz Method

The Wizard of Oz (WOZ) method tests product concepts by letting users interact with what appears to be an autonomous system while a human operator, hidden from view, actually controls the responses. Users experience a realistic, dynamic interface before any real backend exists, enabling researchers to evaluate novel ideas without building them first.

Who Does What

A WOZ study needs a small cast of roles working in concert:

  • The User — the participant who believes they are using a real, functional system.
  • The Facilitator — the researcher guiding the user through tasks and observing their behavior.
  • The Wizard — the person manipulating the system in real time, generating responses to user inputs.
  • The Observer — an optional additional researcher who watches without direct interaction, providing a secondary perspective.

Making The Illusion Credible

The research setup must support a convincing simulation. For a voice-controlled smart speaker study, a physical mock-up and preset scenarios like "Play my favorite music" or "Dim the living room lights" give the wizard clear cues for what to trigger remotely. For screen-based chatbot testing, users type commands while a product team member responds through collaborative tools like Figma/Figjam, Miro, or Mural.

Several factors keep the pretense intact:

  • Timely responses — hesitating or phrasing replies unnaturally will reveal the human operator and break the illusion.
  • Consistent logic — the system's behavior must stay plausible. If a user asks for the weather in one city, the wizard should provide consistently sensible information.
  • Handling surprises — users will go off-script. The wizard needs the flexibility to respond credibly while preserving the sense of a coherent system.

Debriefing participants after the session is essential, with a clear explanation of the method and why it was used. Standard data privacy and participant comfort practices apply.

Where WOZ Fits

The method occupies an unusual place among research tools:

  • Unlike usability testing, which examines existing interfaces, WOZ tests concepts before any meaningful development happens.
  • Unlike A/B testing, which compares design variants, WOZ explores entirely new functionality that would lack context if shown to users directly.
  • Unlike static prototyping, WOZ gives a dynamic, interactive experience that reveals real-time user behavior with a simulated system.

This approach is especially useful for testing novel interactions or complex systems where building a functional prototype is premature or too costly. It lets the team answer fundamental questions about user needs and expectations before committing substantial engineering effort.

Time Savings Versus Crude Prototypes

Paper prototypes and static mockups are fast and cheap, but both lack dynamic responsiveness. Paper cannot simulate complex flows and mockups cannot support personalized or context-aware output. WOZ justifies its overhead when the concept under study involves genuinely interactive behavior that static artifacts can't approximate. Simulating a live-feeling system, even human-powered, exposes usability flaws and conceptual problems earlier and more completely than static wireframes, preventing expensive rework later.

Advanced Practice

WOZ works especially well in iterative cycles. Early rounds validate the concept and surface user reactions, while later passes refine the simulated functionality against the collected findings. If an initial session reveals confusion around a particular flow, the simulation can be adjusted and retested in a follow-up study.

When a system is too complex for one person to simulate on the fly, break the interaction into smaller components. Different team members can handle separate aspects or coordinate their responses sequentially. Clear communication protocols and explicit role definitions keep the experience seamless.

Researchers should define metrics aligned with their goals, not just observe. Tracking hesitation, confusion, or task completion time, then pairing those measures with qualitative findings, adds rigor to the insights.

WOZ also compounds its value when paired with other methods. Pre-study interviews inform the simulated experience with real user mental models; post-study surveys can gauge broader reactions, such as trust or perceived usefulness of a system users just explored.

When To Skip It

A human wizard cannot perfectly reproduce precise system behavior, so WOZ works best at the early formative stage where broad direction is still the goal. In detailed design phases, testing against a richer wireframe or conventional prototype yields better data.

Highly complex systems whose outputs are extremely varied, based on sophisticated real-time calculations, or genuinely unpredictable are a poor fit. A human will struggle to maintain consistent, convincing behavior, and flagging performance compromises research validity.

Preparing The Wizard

A well-trained wizard makes or breaks the study. Essential preparation includes:

  • Clear understanding of the research goals and the questions being asked.
  • Discipline to keep responses consistent across sessions.
  • Familiarity with likely user paths and common deviations.
  • Commitment to remaining neutral, never leading participants or injecting opinions.
  • Established protocols for unexpected actions, including fallback responses or fast consultation lines to the facilitator.

Dry runs are indispensable. Team members and volunteers should practice both participating and devising challenging inputs to stump the wizard. A prepared, generic failure message — such as "I'm sorry, I am unable to perform that task at this time" — can save a session when users go far off course while flagging edge cases for product design.

The Debrief As Research

The post-session conversation is a rich data source in itself. After explaining that the experience was simulated, the researcher should probe the reasons behind observed behavior. Questions such as "Why did you try that?" or "What were you expecting to happen when you clicked that button?" reveal the user's mental model. Investigating moments of confusion, frustration, or delight in detail identifies the highest-priority areas for redesign.

Inside the Wizard of Oz Method: Two Field Studies

The practical value of the Wizard of Oz (WOZ) method becomes clear when it is put to work on real research problems. Two projects — one probing user mental models of agentic AI in enterprise HR, another testing voice controls in a car cabin — show how simulating unbuilt functionality can shape design decisions early.

Probing Mental Models of Agentic AI

When a team began exploring agentic AI for enterprise HR software, they faced a familiar problem with emerging technology: users had no frame of reference. Agentic AI differs from generative AI in that it operates autonomously — understanding intent, planning multi-step tasks, and adapting with minimal human intervention. Survey and interview respondents were intrigued but could not grasp what such a system would actually do, even with analogies to familiar tools.

Building a full prototype was impractical. The underlying algorithms and integrations were too complex and immature, and the risk of building on false assumptions was high. WOZ offered a path around that blockage.

How the Simulation Ran

HR employees sat at a web interface they believed was powered by an intelligent AI assistant. They could ask for help with realistic tasks, such as “draft a personalized onboarding plan for a new marketing hire” or “identify employees who might benefit from proactive well-being resources based on recent activity.” A designer behind the scenes played the wizard, assembling responses from pre-written templates and scenario details to mimic plausible agentic output.

Users were encouraged to interact naturally — to ask follow-ups and test the system’s perceived limits. A user might ask whether the system could schedule team introductions; the wizard would reply that it could propose meeting times automatically, again simulated. Each session ended with a structured debrief: the simulation was disclosed, then open questions explored first reactions, expectations (“Why did you expect that?”), trust and control, and how users believed the “AI” worked. One participant who already understood agentic AI offered an additional expert perspective.

Key Findings

  • Overestimation of capabilities. Some users assumed the system could parse highly ambiguous requests without explicit instruction, highlighting a need to communicate scope and limits clearly.
  • Trust and control. Users were excited by time savings but anxious about losing control over sensitive HR processes. Designs needed transparency into decision-making and room for human oversight.
  • Value in proactive assistance. Proactive suggestions (like flagging burnout risk) were welcomed only when the system explained its reasoning and let a human review and approve actions.
  • Tangible examples matter. Abstract explanation fell flat; simulated interactions with concrete tasks built genuine understanding.

Design Consequences

The findings drove four design choices: the UI had to expose the AI’s reasoning and source data; approval workflows had to be built in for critical actions; development focused on a few specific high-value use cases rather than a general-purpose agent; and onboarding would show clear, concrete examples of capability.

Simulating Voice Control in a Car Cabin

A second project used WOZ to study voice interaction for in-car functions, specifically naturalness and efficiency of commands for climate control, navigation, and media playback. Researchers set up a cabin simulator with a microphone and speakers. The wizard sat in an adjacent room, listened to commands, and triggered the corresponding actions via visual display changes and audio feedback.

Because the “speech recognition” was actually human-powered, the team could isolate problems of command phrasing and interaction style rather than speech engine accuracy. The setup surfaced ambiguous phrasing, frustration points, and user preferences before any speech recognition technology was committed to.

Both cases show the method’s range across product types: one involving deep conceptual uncertainty around an AI system, the other a familiar interface modality in an environment with heavy technical constraints.

Why the Method Fits Emerging Tech

WOZ is often associated with early computer prototypes, but its core strength — simulating complex functionality with human effort — maps directly onto today’s hardest research questions.

For generative AI, a wizard can curate and present model output (text, images, code) in response to prompts, letting researchers study user judgments of quality and trust without a trained, integrated model. Recommendation systems can similarly be simulated by a human guessing at items based on stated preferences and observed behavior, collecting feedback on perceived accuracy before algorithms are written. Even autonomous systems can be studied this way: by simulating autonomous acts in specific scenarios, researchers can learn how much control users want and what explanations they need.

Immersive environments open further possibilities. In VR, a researcher tracking hand movements can trigger virtual events, testing gestural interactions for intuitiveness without full gesture recognition code. In AR, a wizard remotely controlling the appearance and position of virtual objects in the real world lets the team assess placement and relevance before committing to tracking and rendering pipelines.

Across all of this, the human-centered design principle holds: technology should conform to human needs, not the reverse. WOZ inherently focuses on user reactions and behavior, which anchors technological progress to human expectations. That alone justifies its continuing place in the research toolkit — it is a practical way to de-risk product development by validating assumptions and uncovering usability problems before costly rework becomes necessary.