Quality Inputs Beat Clever Outputs
Most discussions about machine learning and AI center on model architecture, compute, or the wow-factor of generated output. Ovetta Sampson, Director of User Experience Machine Learning at Google, argues the industry is looking at the wrong end of the pipeline. The real leverage, she contends, sits in something she calls “minimum viable data”—the smallest, most representative data set required to solve a problem without raising human engagement risks.
Having spent two decades in journalism before moving into tech leadership, Sampson approaches ML and AI from a human-centric angle. Her team at Google works to make machine learning accessible and useful beyond a niche technical audience. In a conversation with Figma, she outlined why the data feeding these systems deserves far more scrutiny than the models themselves.
The Ghosts in the Data
Sampson helped create the design industry’s first set of AI ethics principles, work that shaped her perspective on how people influence data. Her central warning concerns what happens when teams either omit people entirely or treat data points as abstractions disconnected from real lives. The result, she says, can be “traumatized data sets”—data carrying the residue of social, cultural, and economic inequities.
Historical examples illustrate the stakes. The math behind FICO credit scores dates to 1958, but women in the U.S. could not independently obtain mortgages or credit cards until the 1970s. The U.S. Census began in 1790, yet did not recognize LGBTQ individuals until 2021. These gaps are not footnotes; they are systematic exclusions baked into the data used to make consequential decisions.
“The worst thing you can do with ML and AI is be careless about the data—the people—you omit,” Sampson said. Minimum viable data is a directive to product builders to scrutinize what a product genuinely needs to make its business case, and to ask whether ML and AI are even the right instruments for the problem at hand. Her framing is blunt: “There is no AI and ML without data, and there is no data without people.”
Power Flows From the Input
On the current state of ML and AI, Sampson does not mince words about where quality originates. “The quality of an AI’s output depends 99.9% on the input—namely, the data,” she said. “The power is in the input.”
That orientation shifts the governing questions away from model tuning. Who decides what is good or bad? Who decides what enters the data set? Who determines how much data is needed to solve a given problem? For Sampson, these are the foundational inquiries product teams must confront before writing any training pipeline.
Start With the Problem, Not the Model
Practical guidance for builders starts with defining the problem before selecting a tool. Just because AI or ML can be applied does not mean it should be. Sampson’s checklist includes testing whether the data is both equitable and high quality, and whether it represents the true end users as well as the business objectives.
Her recommended reading for teams working through these issues includes Weapons of Math Destruction by Cathy O’Neil, Ghost Work by Mary L. Gray and Siddharth Suri, and the research paper “Everybody wants to do the model work, not the data work” by Nithya Sambasivan and colleagues. The through-line is consistent: model work gets celebrated, while data work—the unglamorous labor of curation and maintenance—gets ignored at great cost.
An Open Box and a Call to Advocacy
Sampson sees a meaningful shift in the public’s relationship to AI. For decades, ML and AI operated invisibly, shaping credit, hiring, and search without broad awareness. Generative AI has changed that dynamic by putting the technology directly in people’s hands.
“Now that more of us have been exposed to AI, we’ll realize our place in it,” she said. “The box is open, and now that people are aware, next comes advocacy.” Her advice for individuals is to recognize their importance to these industries—not merely as consumers of AI products but as participants whose data and feedback shape what gets built.
The path forward, in her view, is not more compute or more elaborate models. It is more attention to the input, the omitted voices, and the fundamental question of who the technology is actually solving problems for.



