Scouting for Bias: What Baseball Teaches Us About Hiring in Tech

Historical baseball scouting reports read less like talent evaluations and more like casting notes. Scouts described prospects in terms of their "good face," whether they had a "high butt," or if they resembled another player. One report on a recruit noted he would "be a real specimen with chance to have a Dave Parker body" with "very large hands." These notes were used to evaluate future major leaguers like Lloyd Moseby, Jim Abbott, and Derek Jeter—players who, by any measurable standard, were exceptional athletes.

But the same flawed lens produced wildly inaccurate predictions. Top prospect Adam Eaton, who never panned out as a pro, drew praise for his "cat-like reactions" and "old fashioned bull-dog" mentality. Conversely, Albert Pujols—eventually one of the best hitters in history—was flagged for having a "heavy, bulky body" and concerns that his "weight will become an issue in time." It didn't.

This pattern is instructive. Baseball has better performance data than almost any other industry, and it still took decades of quantifiable failure before teams embraced statistical analysis over subjective impression. The question for tech is uncomfortable: if baseball scouts who watched players every day couldn't reliably separate future stars from busts using their guts, how accurate are our own hiring and promotion instincts?

The Persistence of Appearance Bias

Even today, physical presence shapes who gets ahead in fields where it shouldn't matter. Observing senior engineers and VPs at major technology firms, a noticeable pattern emerges: men at the highest levels skew significantly taller than the population average. When successful men who are 6'0" or 6'1" report being the shortest person in important meetings, the selection pressure becomes hard to deny. In baseball, height provides a genuine physical advantage; in programming, there's no reason it should be a factor at all.

The gap between what we see in mental competitions and corporate hierarchies highlights the issue. Top chess players—a field that directly measures cognitive performance—hover right around average height: Magnus Carlsen, Viswanathan Anand, and Garry Kasparov stand at 5'8", 5'8", and 5'9", respectively. Elite go and shogi players like Lee Sedol and Yoshiharu Habu also fall in normal height ranges. The correlation between height and cognitive ability is weak (estimates range from 0 to 0.3), explaining at most 9% of variance even in the strongest case. This suggests the height segregation seen in executive suites reflects a halo effect rather than underlying merit.

How Gatekeeping Reinforces Bias

Systemic bias doesn't just affect who gets hired. It funnels into who gets promoted, who is given high-impact assignments, and who gets a seat at the table in the first place.

Genuine talent frequently slips through the cracks of biased systems. Legendary players like Mel Ott, Joe Morgan, and Kirby Puckett were all consistently overlooked because of their small stature—each breaking into professional baseball only through pure chance encounters with someone willing to look past appearances. The career of a friend who is now an engineering professor at a top Canadian university illustrates the same dynamic in academia. Frequently mistaken for a homeless person, an administrative assistant, or a professor's wife, she was once berated in front of her civil engineering class for asking what a "corn dog" was—a question on an exam that assumed cultural knowledge she simply didn't have. Her failure wasn't technical aptitude but a mismatch between expectation and reality.

Even the stories that seem to celebrate meritocracy reveal a dark undercurrent. One tech executive tells how he got into Carnegie Mellon University despite poor grades and SAT scores—when the admissions office rejected him, he walked the halls, convinced professors to meet him, and eventually earned an on-the-spot acceptance from the school's vice president. His conclusion: gatekeepers are just genuinely looking for talent and agency.

"I think one secret, at least when it comes to gatekeepers, is that they're usually just looking for high agency and talent."

The subtext of that story is telling. A young man showing persistence fits a familiar, favorable pattern. A girl from rural Canada with the top grades in her high school who dresses poorly because she's raising her younger brother while paying for college doesn't fit any expected pattern of an engineer—and gets treated accordingly. Both narratives might stem from genuine merit, but only one gets recognized as such.

Years of research bear this out. Names that sound "white" on resumes receive more callbacks than names that sound "black"—and even more for Asian-sounding names. In some studies, professors with white-sounding names receive better interpersonal evaluations based solely on their CVs than professors with black or Asian names.

The Difficulty of Measuring What Matters

Backing up one level, there's a critical mismatch between what information's needed and what information could theoretically be collected. Sports—particularly baseball—offered teams direct measurements of on-field performance, yet they ignored these numbers for a century. Distinguishing signal from noise wasn't the issue; refusing to look for signal was. Even when teams did adopt stats, the change came only after amateur analysts—using data publicly available to anyone—publicly demolished professional scouting departments. Initially dismissed, these hobbyists eventually outperformed entire scouting staffs of professional baseball organizations.

Deciding who to "hire" for a team was a high-stakes decision with millions of dollars on the line. Teams that invested in measurements gained a massive advantage. If that's true in baseball—the most easily quantified major sport in the United States, where outcomes can be tracked with precision and data collection has been mandated since the game's earliest decades—what does that say about tech?

In programming, individual performance is dramatically harder to measure than in sports. There's no batting average for engineers. In environments where performance evaluation is ambiguous, reliance on proxy signals—physical appearance, confidence, credentials, cultural background—becomes much higher. The evidence confirms this: informal polling of engineers about their colleagues reveals a striking split. While one cluster judges people by their output, another cluster rates engineers who are tall and confident highly—even when those same engineers produce broken or nonexistent systems. A third cluster emphasizes credentials: which school someone attended, which company's badge they carried, what title they hold. Within clusters, people agree on who does excellent work. But between clusters, the evaluations diverge dramatically.

Reading Performance From Short Samples

Even evaluators without bias face the challenge of sample size. Scouts saw Chipper Jones on one day and wrote:

"Was not aggressive w/bat. Did not drive ball from either side. Displayed non-chalant attitude at all times. He was a disappointment to me."

On another day, another scout's evaluation of the same player predicted:

"Definite ML prospect . . . superstar potential."

The scout who got it right didn't have access to some mysterious evaluative insight like the one who got it wrong. He simply happened to watch Jones during a hot streak instead of a slump. Performance for humans isn't static—it includes enormous variance—but our full-time employment review culture, structured around annual goal-setting cycles, misses or multiplies daily variance into erratic long-term signals.

Most interestingly, talent assessment likely contains the same biases, baked in through layers of promotion reviews, title assignments, and meeting invite lists. Promotions to senior levels often depend on visibility plus previous titles plus skill—forming a closed loop where already-advantaged individuals get yet more opportunities. And the feedback loops that might address these errors often loop incorrectly: junior engineers who seek promotion-suitable projects get told they're too inexperienced to work on them; moving to a better team gets harder if you've fallen behind.

Departing from this systemic pessimism, individual remedies exist—the equivalent of Bill Wight, the baseball scout who, because he evaluated players on directly observable performance, went on to sign "funny-looking" athletes who other teams dismissed and who rarely regretted it. In engineering, broadly unconventional companies such ones that ignored pedigree to hire interns with strong internal initiative, someone who mastered systems, or who launched a profile for a dying game he re-created in his spare time, outperform their bigger-name rivals. For a startup, evaluations were essentially universal: output. But at large companies, no single evaluation cluster dominated, which makes a clear path forward difficult—yet this is exactly the situation worth recalling for tech HR leaders.

Time could improve things. There's already anever-increasingpile ofevidencethat

we still handle talent horribly—even unfairly—compared with the statistics-driven methods available. Blind auditions, which successfully reduced bias in orchestras after research showed basic causal effects, remain one easy remedy worth more experiment than the current reliance on unstructured gut assessment. But as measurable performance measurement matters less, the danger that we're rewarding whoever appears to fit a script—tall, confidently male, alma mater in order—grows worse than baseball ever faced.