When AI Predictions Become Policy

In 2024, an Air Canada customer asked a chatbot about bereavement fares. The bot confidently offered a refund policy that didn’t exist. The airline refused to honor it. A tribunal sided with the customer. The bot hadn’t decided anything; it had predicted an answer from patterns in training data. The company treated that prediction as policy.

That incident captures the central challenge of designing with AI today: probabilistic systems wrapped in deterministic interfaces. An AI offers a guess, the interface presents it as certainty, and the user — or organization — acts on it as fact.

Humans lean toward deterministic thinking. Assuming past actions determine future outcomes feels natural. Flip a coin 999 times and get heads each time; a deterministic mind assumes the coin is rigged. A probabilistic mind accepts that the 1000th flip might still land either way. That second mindset is harder to sustain, yet it is precisely what design teams need now. Products live in complex, nonlinear environments, and AI is amplifying that complexity. When teams treat AI output as the answer rather than one plausible answer among many, they build fragile experiences — and in high-stakes fields like medical diagnostics or financial forecasting, genuinely dangerous ones.

Designing probabilistically means using AI as a thinking partner, not a substitute for thought. It requires accounting for model bias, human sentiment, and perceived risk at every step.

Reading Outputs as Signals, Not Conclusions

Most AI questions do not yield binary answers. Asking “Do aliens exist?” produces something between plausible and uncertain: scientists consider extraterrestrial life likely, but without evidence, certainty is impossible. The response frames the question as a probability rather than resolving it.

Designers should read AI outputs the same way — as signals, not conclusions; possible outcomes to interpret within the context of product goals, user behavior, and business constraints.

Many respected digital products already operate this way. Netflix doesn’t know you will enjoy Superstore because you watched The Office; it estimates the likelihood and surfaces the title accordingly. The interface responds to a prediction.

Design decisions can adopt the same logic. Integrating behavioral analytics with research insights allows AI models to estimate the probability of specific outcomes, and those probabilities can anchor the design strategy. Suppose analytics suggest a 60% chance users will complete a purchase, versus a 90% chance. At 60%, the interface must persuade: testimonials, explanations, comparisons, and reassurance signals may tip the user toward action. At 90%, users are already motivated, and the design job shifts to removing friction so the purchase happens quickly. The same screen presents two very different design problems based on the probability at hand.

Comparison of two hair product ads showing the same model, with the simplified design on the right labeled 90% confidence and the text-heavy design on the left labeled 60% confidence.
Note: This is an oversimplification of the idea. Please be mindful of the intricate details of your product. (Large preview)

AI also can simulate outcomes from historical data and behavioral models before a direction is committed. The value depends heavily on prompt structure — the context defined, the hypothesis tested, user motivation considered, and the edge cases stressed.

One practical use: evaluating early designs through structured prompts when direct access to the target user group is not available. The following prompt evaluates a design from the perspective of neurodivergent users, and it works as a template — adapt the user group, criteria, and output format to your product, then use it as a discussion starter with the team rather than a final verdict:

Evaluate the [design file or weblink] for usability, accessibility, and content relevance from the perspective of neurodivergent users such as those with autism spectrum disorder, ADHD, learning disabilities, etc.
Please consider the following criteria:
  1. Is the layout and navigation intuitive for neurodivergent users?
  2. Is the language and content appropriate and engaging for neurodivergent users?
  3. Are there any barriers (technical, cognitive, or sensory) that this group might face when using the site?
  4. How well does the site meet the specific needs or goals of neurodivergent users?
Provide a SWOT analysis, probability score for successful use by neurodivergent users, and any recommendations for improvement.

Note: This is an oversimplification of the idea. Please be mindful of the intricate details of your product and make any appropriate changes.

Crucially, simulations do not replace experimentation. Because models train on historical data, they mirror past behavior more strongly than they anticipate change. Designing a voice interface for elderly users who struggle with touchscreens might yield a low engagement prediction — not because the idea lacks merit, but because the source data reflects different habits. Simulations should surface assumptions, never block innovation.

Skepticism Toward Skewed Probabilities

AI systems inherit the biases of their training data. At the AI Summit in France, India’s Prime Minister Narendra Modi illustrated this with an example: ask a model to generate an image of a person writing left-handed, and the output may still show right-handed writing. The cause is statistical — most people are right-handed, and training data mirrors that. The behavior may have improved over time, yet the incident remains instructive; it still appears occasionally with similar image models.

Output is not truth. It represents the most statistically likely outcome given available data. Designers must consider whether past data meaningfully predicts future behavior. If added context would sharpen the prediction, supply it. Lacking context, the output is only one possible answer dressed as the only one.

Promt, which reads: create an image of a person sitting in his chair facing his desk and writing with his left hand in his notebook, and the image created for it.
(Large preview)

Confidence scores demand the same scrutiny. Over-trusting a 90% prediction recreates the Air Canada failure. Dismissing a 40% signal can make a team miss genuine insight in noisy data. A high-confidence prediction may still be wrong; a low-confidence one may still carry value. Designers must weigh the alternatives, examine the specific case, and apply judgment to whatever the AI recommends.

Transparency makes that judgment possible. As AI increasingly informs decisions, users need visibility into how outputs are produced — the sources, reasoning, and summaries behind a recommendation. Black-box systems breed distrust; systems that reveal their reasoning encourage independent evaluation. This transparency is both good design and ethical practice, respecting the trust placed in these tools.

Thinking in probabilities often demands resisting snap answers. AI accelerates research and exposes patterns faster than ever, but its outputs are starting points, not final decisions.

Treat Design as a Portfolio of Bets

Design decisions do not produce binary outcomes. A feature launch, a new flow, or a content strategy can lead to a spectrum of results, and the designer’s job is to navigate that range. Probabilistic thinking treats each decision as a bet with varying odds, not a guaranteed win or loss. This shifts the goal from being "right" to being value-driven

Assume multiple valid solutions exist for any given problem, each with a different likelihood of success. A well-researched idea can still fail in the real world. This is where AI helps: it can act as a portfolio-thinking engine, generating multiple interpretations of a problem, highlighting risks in each path, and structuring recommendations by likelihood. Instead of asking whether an idea will succeed, ask the AI to estimate the likelihood and get a score. Use those signals to decide where to invest.

Design decisions should be optimized for likelihood, not certainty.

Make Uncertainty Visible

The Air Canada chatbot case illustrates what happens when a probabilistic system wears a deterministic interface. The bot generated plausible text as language models do, but the UI presented that prediction as fact. The user read that confidence as commitment — and legally, so did the tribunal. The problem was not the underlying prediction engine but an interface that erased the difference between a guess and a guarantee.

Designing for likelihood means letting the interface preserve a degree of uncertainty. Confidence indicators, ranges, and visible paths to human support should replace absolute statements. A delivery window of "Friday to Monday" tells the truth about variability without misleading anyone, whereas a specific timestamp that slips erodes trust every time. A face recognition feature that says "this looks like Pratik, is that right?" sets more honest expectations than one that just labels the photo with a name.

Communicating uncertainty does not weaken trust — it strengthens it.

Use Data as a Compass, Not a Map

Even a valid probability score is not the final answer. An 80% likelihood that users prefer a minimal checkout does not mean you simply build that checkout — it means you investigate. Ask why the model produced that prediction, what data influenced it, which assumptions it leans on, and what user behavior it is actually detecting. AI excels at finding patterns but rarely explains why those patterns exist. Understanding motivation is still a human-centered research task.

When AI goes wrong, the root cause is usually the data. Amazon’s experimental AI recruitment tool learned to downgrade resumes from women. Trained on roughly a decade of historical hiring decisions skewed toward male candidates, the model began penalizing resumes that included the word "women’s," as in "women’s chess club captain," and favoring language more commonly found on men’s resumes. The system was not intentionally biased — the data was. Amazon reportedly adjusted it, then shut the project down because they could not guarantee it would not surface other discriminatory patterns.

This is why designers must interrogate AI output critically. Evaluate the reliability of the models you depend on and ask what the training data might be hiding. A recommendation is only as good as the data behind it, and the only way to know what that data is hiding is to ask.

Hypothesis Filters and Continuous Loops

Experiment teams often frame testing as a way to validate a design decision. Probabilistic thinking reframes the goal: experiments should reduce uncertainty, not just confirm a chosen solution.

  • Traditional approach: Testing features to confirm success.
  • Probabilistic approach: Testing assumptions to reduce uncertainty.

Traditional A/B tests are expensive in engineering time, traffic allocation, and user exposure — especially when a losing variant runs against a significant chunk of the audience. AI simulations can model potential outcomes from historical and behavioral data before anything ships, filtering weaker ideas before they reach production. This becomes a continuous feedback loop.

Predict → Test → Learn → Adjust → Repeat!

Putting this in practice requires several shifts:

  • Shift the framing. Instead of "Will this feature succeed?" ask "What assumptions are we testing?" Define the hypothesis explicitly.

We believe [behavioral assumption] will impact [metric] because [reason]. We’ll know we are right when [evidence].

A concrete example: We believe simplifying the onboarding flow from 5 steps to 3 will increase completion rate because users experience decision fatigue when too many choices are presented. We’ll know we’re right when we see at least a 15% increase in step-to-step conversion with no drop in activation rate.

  • Use AI simulations first. Let AI predict which assumptions are worth testing, then use that learning to identify top candidates for real experiments.
  • Embrace multi-versions. Version A may resonate with high-intent users while version B works better for exploratory ones. Multiple live versions are fine and can be intentional.
  • Fail fast. Reward learning over success. Normalize smaller experiments instead of sweeping large changes — pick up a few probabilities and test them.
  • Visualize probability. Keep a probability table with each variant’s likelihood of success to track all live changes.

Keeping Humans in the Decision Loop

AI should augment human judgment, not replace it. Human-in-the-loop (HITL) is not a safety net — it is a refinement engine. Each accept, reject, or override is high-quality feedback that improves the model over time. Users are more willing to rely on AI when they understand how a suggestion was generated, can evaluate its implications, and can intervene easily.

Every accept, reject, or edit produces far more meaningful training data than passive analytics, closing the loop between real-world usage and model performance.

Simple accept/reject affordances suit low-risk suggestions that improve speed without real consequences. Git Hum Copilot works this way: developers accept, edit, or ignore inline suggestions, and authorship stays with the human. Gmail’s Smart Compose similarly presents predicted text as optional. As stakes climb into data, money, or human impact, preview and approval steps become essential, and explicit explanations help users calibrate trust.

Risk and fraud systems route decisions by probability score: low-risk proceeds automatically, medium-risk triggers additional verification, and high-risk escalates to a human reviewer. In healthcare, AI may flag anomalies or suggest a diagnosis, but the clinician retains final authority — tools that explain the details reinforce confidence without removing accountability.

Capture decisions with context, feed the outcomes into learning workflows, and log overrides for auditability. Track override rate, confidence accuracy, time-to-approval, and perceived trust. A high override rate is not a user failure; it signals that the design or the model needs attention.

(Large preview)

Users process uncertainty differently, and the design should account for those differences.

User typeRiskDesign goal
Overtrusting usersThey act too quickly and trust AI results easily./Show uncertainty more prominently.
Distrustful usersThey ignore AI entirely.Show historical accuracy or confidence levels.
Skeptical/balanced usersUses AI as a guide, not as a rule.Reinforce AI assistance and let them decide the sort of framing.

Poorly implemented HITL fails in predictable ways: human review devolves into a rubber stamp, workflows slow enough that users route around the safeguards, or feedback skews toward a narrow user subset. Those are design problems, not arguments for removing HITL. The goal is not to maximize human involvement but to focus it where uncertainty, impact, or ethics demand it. That requires clarity about who decides, when uncertainty matters, and how responsibility is shared between people and machines.

Resilience as a Design Target

Conversion-optimized interfaces are inherently fragile once they depend on a probabilistic engine. User intent shifts, models drift, and external conditions change—meaning that an interaction pattern tuned for today’s data distribution can quietly fail tomorrow. A resilient system instead asks a broader question: how does this behave over time, under stress, and in uncertainty?

Such a system adapts to new behaviors, fails safely rather than catastrophically, and remains explainable. It also avoids brittle, over-optimized interaction patterns and anticipates second-order effects. Extending measurement beyond last quarter’s numbers into subsequent quarters surfaces those shifts before they become problems.

Building for Shifting Probabilities

A common failure mode is designing as though conditions are stable. Recommendation feeds illustrate this well: an early version optimizes for engagement, engagement rises, and then users describe the feed as narrow, repetitive, or exhausting. A resilient counterpart rebalances continuously—introducing novelty, diversifying signals, and weighing long-term satisfaction against short-term clicks.

Designers should create interfaces that expect change, dynamic re-ranking, contextual explanations, and escape hatches from stale personalization loops, all of which help systems stay useful as probabilities shift.

Long-term Outcomes vs. Short-term Wins

Short-term conversion metrics often conceal downstream costs. Accelerated onboarding can reduce comprehension; maximizing notification click-through rates can erode trust; pure engagement optimization can produce unhealthy usage patterns. These second-order effects typically surface weeks or months later.

Duolingo’s hearts system demonstrates intentional friction against this pattern. Running out of hearts after repeated mistakes forces users to wait or practice older material. On paper, this reduces lessons per session. In practice, the team has publicly discussed how it supports long-term motivation and retention—the metric that matters for a learning app. Short-term engagement dips while long-term outcomes improve.

Meta made a comparable acknowledgment after optimizing for “time spent” produced unintended emotional and societal effects. Its stated pivot toward “meaningful social interactions” as a guiding metric—regardless of implementation success—underscores that optimizing the wrong metric at scale has real human cost.

Routine design review should therefore ask:

  • What behaviors are we unintentionally reinforcing?
  • Will this interaction remain healthy when repeated at scale?
  • Are we optimizing for ecosystem wellbeing or just the next click?

Planning for Uncertainty Like Scale

Teams plan for traffic spikes but rarely plan for uncertainty spikes. AI systems degrade, adversarial actors evolve, and external shocks can reshape user behavior overnight. Resilience assumes variability and prepares for it, especially when confidence drops.

The interface should have a defined answer for low AI confidence: graceful handoff to a human or fallback state, not silent failure. The experience must remain coherent even if AI assistance is removed entirely. Practical steps to build this in include:

  • Design for degrading confidence.
    Show fallback states, support manual overrides, and visualize uncertainty where it matters.
  • Measure long-term user health.
    Track satisfaction, retention quality, and unintended behavior—not just conversion.
  • Build adaptability in.
    Configure adjustable ranking rules, dynamic states, and continual experimentation across segments.
  • Model second-order effects early.
    Every optimization has downstream consequences; surface them before shipping.
  • Use a resilience checklist before launch.
    Define behavior under low AI confidence, safe fallbacks, and anticipated drifts.

The reframe is subtle but consequential: instead of asking whether a design works, ask how likely it is to work and what happens when it doesn’t. This changes hypothesis writing, AI output interpretation, experiment scoping, and the design of failure states. A concrete starting point is to name the assumption behind every accepted AI recommendation, find one place where probabilistic output is presented as certainty, fix that framing, and design the fallback before the happy path.

This shift is less about adopting new tools than about adopting a new posture. AI has not introduced uncertainty into the world; it has made the uncertainty that was always present impossible to ignore. Models can estimate, simulate, and recommend, but they cannot determine what matters, which users are overlooked, or which unconventional idea deserves defense against a model trained on yesterday’s data. Those judgment calls remain human responsibilities.
Think in ranges, not points. Test assumptions, not features. Design for adaptation, not perfection. When prediction is cheap and judgment is rare, the most valuable designer’s question is: What else might be true?

Smashing Editorial