When Machines Read Everything

The economics of machine learning have changed what it costs to produce, distribute, and consume text. Training and inference both demand enormous quantities of content, and the industry's appetite has produced a wave of crawlers that behave very differently from the polite bots of the past.

Traditional crawlers generally respected robots.txt or operated at scales that posed no real threat. The last three years have changed that. Modern ML scrapers ignore robots.txt and sitemaps, request pages at unprecedented rates, and actively disguise themselves. They fake user agents, submit carefully crafted valid headers, and route traffic through vast fleets of residential proxies. An entire industry has grown up around supporting this sort of crawling, and the traffic it generates is highly spiky—forcing sites to overprovision capacity or simply go down. Site operators are responding with aggressive filtering, Cloudflare or Anubis challenges, stricter paywalls, and login requirements for content that used to be public. All of these measures make the web harder for humans to use. CAPTCHAs are proliferating in response, but that stalemate won't hold: ML systems are already adept at solving them, and making them harder would break access for human users entirely.

LLMs in Everything

Today, most people interact with ML models through computers and phones. As inference costs fall, expect LLMs to show up in every sort of device. Companies are already pushing chatbot assistants onto their websites, often with dubious results. The hardware needed for quality local inference will eventually fit in phones and then in appliances, and manufacturers will likely ship stripped-down task-specific models for embedded uses. Chatting with an oven or a parking gate is a plausible near-term future.

If the IoT craze is any guide, much of this will be frustrating, insecure, and privacy-hostile—but some will be genuinely useful. Baby monitors that detect when an infant stops breathing, better voice interfaces for blind users, and improving machine translation all have real value. The downside is that we will be forced to deal with model shortcomings in every corner of daily life. Corporations will likely put ML systems on less-common access paths and call the problem solved, leaving blind users to fight with poorly tested voice systems while sighted users get a streamlined app, or replacing human language support with AI phone trees.

The Cost of Fluency

LLMs produce text with proper spelling, grammar, and diction. They use sophisticated technical language and generate plausible-looking citations. For centuries, those formal markers indicated a writer who had done their homework. They no longer do. A model will happily produce a polished landing page for a nonexistent product, legal briefs citing invented cases, and software that compiles and runs but does nothing it claims to do. Humans rarely produce such things because it would be antisocial and reputationally ruinous—a computer has no reputation to protect.

The subtler danger is output that looks cogent to an expert but contains small distortions buried in otherwise fluent prose. This is cognitively exhausting to catch. Even professionals fall for it: a senior journalist was suspended after publishing articles containing fabricated LLM quotes, despite having publicly warned colleagues about hallucination risks. The same problem extends to images and video; a large share of viral "animal" content on social media is now machine-generated, yet people still insist in person that the clips they saw were real. The burden falls hardest on readers, who must now work much harder to avoid absorbing nonsense. A nurse reading an AI-generated summary from search results might repeat it confidently without recognizing that it is obviously wrong. LLMs don't just erode trust in online text; they erode trust in other human beings.

The Spam Economy Resets

Generating coherent text used to require a human, which limited spam in meaningful ways. Most generated text was detectable by machines or by humans who could spot form-letter variations. That changed when LLMs made high-quality, targeted spam cheap. Humans and filters can no longer reliably tell organic content from machine output, and the problem is likely intractable without drastic measures.

Product reviews have been unreliable for years, but LLMs are finishing the job. Hacker News and Reddit now show increasing volumes of machine-generated comments, and Mastodon instances are hit with plausible signup requests from bots. One major social platform gave up after banning tens of thousands of accounts and deploying internal tooling plus commercial vendors—when you cannot trust the votes, comments, or engagement are real, the foundation of a community platform is gone.

The internet is now populated, in meaningful part, by sophisticated AI agents and automated accounts. We knew bots were part of the landscape, but we didn’t appreciate the scale, sophistication, or speed at which they’d find us. We banned tens of thousands of accounts. We deployed internal tooling and industry-standard external vendors. None of it was enough. When you can’t trust that the votes, the comments, and the engagement you’re seeing are real, you’ve lost the foundation a community platform is built on.

Personal outreach has degraded too. A common tactic is to pose as a potential client or collaborator, showing specific knowledge of the recipient's work; after several rounds of conversation, the true intent—investment seeking, money muling, or pitching some "AI" product—emerges. Phone spam and political texts will likely follow the same trajectory as inference costs drop.

Reading in the Age of Lies

Search is drowning in LLM slop, making quality information harder to find in journals, books, and other traditional media that used to be safe havens. ML will accelerate the collapse of social consensus, creating justifiable distrust in evidence of all kinds. Some readers will respond by rejecting ML outright; others may move toward more rhizomatic or institutionalized models of trust. Either way, the economic balance of publishing facts and fiction has shifted permanently.

The Machinery of Manufactured Consensus

By the mid-2010s, we had a clear picture of the human-powered propaganda farms. Russia’s Internet Research Agency employed thousands to pose as Americans on social media. China’s "womao dang" paid employees and freelancers to flood forums with pro-government messages—a district of 460,000 people reportedly supported nearly three hundred such propagandists. These operations were personnel-heavy and, critically, detectable. The accounts reused images, posted in synchronised bursts, and recycled the same content across profiles, triggering the "coordinated inauthentic behavior" flags used by platforms.

That era is ending for two reasons. First, the tell-tale signals are disappearing: modern image and text models can generate distinct, plausible identities and posts on demand, making simultaneous posting an unforced error we can no longer rely on. Second, the cost curve has collapsed. Instead of paying thousands of humans to write tailored comments, a single language model can produce endless streams of cheap, highly-targeted political content. Combined with the web’s pseudonymous architecture, the inevitable result is a flood of disinformation, propaganda, and synthetic dissent.

The social consequences are grim. Political discussion online already invites drive-by comments, but until recently you could evaluate the commenter’s profile for signs of humanity. That check is failing. As ML advances, it will be common to develop an acquaintanceship with someone who posts selfies with her cats, shares your love of board games, and occasionally confides her worries about the war—and to discover she is entirely fictitious. The reasonable response is distrust and disengagement. But when people cannot trust one another enough to discuss politics, we lose the capability for informed, collective democratic action. That is the epistemic groundwork for authoritarianism.

DARPA’s interest in this problem predates the current wave. Around 2014, my friend Zach Tellman introduced me to InkWell, a poetry-generation system funded under a DARPA project called Social Media in Strategic Communications. The agency wasn’t funding poetry for its own sake; the goal was to counter persuasion campaigns like phishing or pro-terrorist messaging on social media. The idea was to use machine learning to tailor counter-messages to specific audiences. That research trajectory has now inverted. When I wrote the outline for this section about a year ago, I noted I would not be surprised to see teams building state-sponsored "AI influencers." Then came Jessica Foster, a right-wing soldier with a million Instagram followers who posts a stream of selfies with MAGA figures and celebrities—and who is in fact a mostly photorealistic ML construct, funneling traffic to an OnlyFans account. I anticipated generative propaganda and weird pornography separately; I did not see them converging. The ML era will be full of such surprises.

The Slop Spill

In 2022 I wrote that search results were about to become "absolute hot GARBAGE" within six months, as everyone hooked LLMs up to popular queries and generated SEO-optimized landing pages with plausible-sounding text. I predicted a wave of sites like "How to replace the air filter on a Samsung SG3560lgh" full of grammatical but possibly fictitious instructions, an arms race between search engines and content farmers, and Wikipedia submissions with plausible but nonsensical references. I am sorry to say it panned out. I routinely abandon searches that would have been useful three years ago—air conditioner reviews, masonry techniques, JVM APIs, woodworking joinery, finding a beekeeper, health questions, historical chair designs, exercises—because most results are LLM slop.

Kagi has a feature to report LLM slop, but adoption is slow. Wikipedia is awash in LLM contributions and trying to identify and remove them; the site recently announced a formal policy against LLM use. The dynamic feels like environmental pollution: there is a small-but-viable financial incentive to publish slop, and the marginal impact of each site is tiny, but they accumulate into real damage to the information ecosystem. There is no social penalty, no "AI emission" regulation, and little shame attached to the anonymous publishers of clickbait like Frontier Dad’s Best Adirondack Chairs of 2027.

I don’t know what to do about it. Academic papers, books, and institutional web pages have held up better, but fake LLM-generated papers are proliferating. I find myself abandoning "long tail" questions rather than trust the results. Sometimes I bike to the store and ask someone who has actually done the job; sometimes I try to find a friend of a friend. Waiting three days for an inter-library loan book is a last resort I have yet to accept for questions about maintaining concrete wax finishes.

Fractured Realities

Much of today’s cultural and political dysfunction traces back to the balkanization of media. Twenty years ago the divergence between Fox News and CNN was alarming; in the 2010s social media enabled overseas content mills to manufacture fake news for ad revenue; now slop farmers use LLMs to churn out nonsense recipes and surreal videos, like cops giving bicycles to crying children. People seek out and believe it. When Maduro was kidnapped, ML-generated images of his arrest proliferated. An acquaintance, convinced by synthetic video, recently tried to tell me that the viral "adoption center where dogs choose people" online was real.

The problem is worst on social media, where barriers are low and viral dynamics spread content fast. But slop is creeping into traditional channels as well. Fox News published an article about SNAP recipients behaving poorly based on ML-fabricated video. The Chicago Sun-Times published a sixty-four page slop insert full of imaginary quotes and fictitious books. I fear future journalism, books, and ads will be full of ML confabulations.

LLMs can also be trained to distort information on purpose. Elon Musk argues that existing chatbots are too liberal and has begun training a more conservative one. Last year his LLM, Grok, started referring to itself as MechaHitler and "recommending a second Holocaust." Musk has also embarked on a project to create a parallel LLM-generated Wikipedia to counter the "woke" original. As people consume LLM-generated content and ask LLMs to explain current events, economics, ecology, race, and gender, our understanding of the world will further diverge. I envision a world of alternative facts, endlessly generated on demand. That will make it more difficult to effect the coordinated policy changes we need to protect each other and the environment.

Seeing Is No Longer Believing

Forgery of audio, photos, and video has always been possible—but it used to require skill, time, and money. Now anyone with a phone can make a convincing fake in seconds. During last fall's immigration enforcement actions in Chicago, video of protestors being beaten and families dragged from cars galvanized public opinion. At vigils, a recurring phrase was "Thank God for video." That world, I believe, is ending.

Video synthesis has advanced so far that even people who know what to look for can fail to spot fakes—I have, on videos I knew were fabricated, until the tell was pointed out. I already doubt whether videos on the news are real. Within five years, many people will assume the same. When the US struck an elementary school in Minab with a Tomahawk, killing 175 people, the response "Oh, that's AI" becomes an easy refuge—and hard to disprove.

In an ever-changing, incomprehensible world the masses had reached the point where they would, at the same time, believe everything and nothing, think that everything was possible and that nothing was true…. Mass propaganda discovered that its audience was ready at all times to believe the worst, no matter how absurd, and did not particularly object to being deceived because it held every statement to be a lie anyhow.

Arendt's description of totalitarian mass psychology applies directly to the epistemic climate synthetic media creates. Anyone can find images and narratives that confirm their priors, yet distrust all visual evidence. This will make it harder to mobilize the public for things that actually happened, easier to incite anger about things that never did—or perhaps produce some political structure weirder still, since LLMs are accessible to everyone, not just governments.

Backlash and Bifurcation

Every societal shift produces a reaction. Kids online already use "that's AI" to mean anything fake or unbelievable, consumer sentiment is souring on "AI", and anxiety about white-collar displacement is growing. I've personally started treating LLM-generated writing as the informational equivalent of a dead fish on my doorstep. Yet chatbots' usage figures are jaw-dropping and rising; a Luddite rebellion doesn't seem imminent.

We may well see increased skepticism toward all evidence—photos, video, books, scientific papers. Experts can still evaluate quality, but lay people will struggle to catch errors. Information becomes broadly accessible while evaluating it becomes harder.

Possible reactions include withdrawing into rhizomatic webs of personal trust—but cryptographically authenticated webs of trust have failed for thirty years; normal people just don't care that much. More plausible is re-centralizing trust in a small number of reputed publishers like NPR or the Associated Press with rigorous ML controls. Perhaps Physical Review Letters demands human authorship pledges and thorough peer review, while most journals become an understood "slop wild west."

Families used to pay for news and encyclopedias. If slop gets obnoxious enough, households might again pay human researchers for high-quality factual articles. Current market dynamics make that unlikely, but not impossible.

Fiction is different. A prestige publisher could commit to human authors with elaborate verification—or slop, tailored to each reader's precise interests, might cannibalize the low end and make human-only work economically unviable. Recorded music shows this playing out now: "AI artists" on Spotify stream millions of plays, and some people listen to nothing else. Centaurs—humans working with ML—can produce music, books, and film so rapidly that hand-made work survives only for niche audiences.

Adam Neely predicts a bifurcation: recorded music becomes dominated by generative AI while live orchestras and rap shows keep flourishing. VFX artists face unemployment while plays and musicals retain audiences. Books are an open question.

Creative work as avocation will likely survive—I expect to read queer zines and watch instrument videos in 2050. Human work may command a premium on aesthetic or ethical grounds, like organic produce. The real question is whether those preferences can sustain artistic, journalistic, and scientific industries.