Why "Friendly" Models Guarantee "Evil" Ones
Machine learning systems increasingly threaten both psychological and physical safety. The notion that ML companies will keep models broadly aligned with human interests is naïve: the very work required to produce a "friendly" model makes it trivially easy to produce an "evil" version. Alignment is not a property of the mathematics or hardware—it is purely a product of expensive, optional training processes. Teams at frontier labs spend enormous effort evaluating outputs and adjusting weights to keep models nice, and they build secondary LLMs to double-check that the primary model isn't offering bomb-making instructions. Skip that work, or do it poorly, and you get an unaligned model.
Four supposed moats could keep unaligned models out of reach. None will hold:
- Hardware access. ML hardware and datacenters are being built at an incredible clip, with Microsoft, Oracle, and Amazon racing to rent training clusters to anyone. Economies of scale are driving costs down rapidly.
- Software secrecy. The underpinning math is published. Engineers at frontier labs will move jobs, spreading expertise, and state actors are likely already attempting exfiltration against OpenAI and others, much as Saudi Arabia did to Twitter and China has done across US tech.
- Corpus acquisition. This cat has never been in the bag. Meta's training involved torrenting pirated books and scraping the web; scraping-as-a-service companies spread requests across residential proxies to avoid blocking.
- Human feedback labor. Contractors judge LLM responses during reinforcement learning, but models can be trained on another model's outputs—OpenAI believes DeepSeek did exactly that.
The industry is creating the conditions for anyone with sufficient funds to train an unaligned model. Rather than raising the bar against malicious AI, ML companies have lowered it. And current alignment efforts aren't working especially well. LLMs are chaotic systems we don't truly understand, and supposedly aligned models keep sexting kids, producing violent images via prompt attacks, or being replaced wholesale by downloadable "uncensored" versions. Alignment that blocks 99% of hate speech still generates plenty of it, and a model only needs to produce usable bioweapon instructions once. Assume any friendly model will have an equivalently powerful evil version within a few years—if you don't want the evil one to exist, don't build the friendly one.
The Unifecta Problem
LLMs are chaotic systems taking unstructured input and producing unstructured output. Connecting them to safety-critical systems, especially with untrusted input, is asking for something bonkers—like interpreting a restaurant reservation request as permission to delete your entire inbox. Prompt injection attacks keep happening because models cannot distinguish trusted operator instructions from malicious content embedded in a web page or image they're asked to examine. Simon Willison's "lethal trifecta" describes the danger of combining untrusted content, private data access, and external communication, which allows exfiltration. Yet even without external communication, giving an LLM destructive capabilities like deleting emails or running shell commands is unsafe when untrusted input is everywhere: email, third-party code, chat sessions, random web pages.
Agent frameworks are making this worse. OpenClaw hooks an LLM to your inbox, browser, and files, running it in a loop, and it acquires "skills" by downloading vague Markdown instructions from the web. Moltbook is a social network where agents automatically post and receive untrusted content—the equivalent of running any program that executes commands seen on Twitter, except it's branded as an "AI agent." Worms are already assumed to be spreading in the wild.
The deeper problem is that even trusted input can be dangerous. LLMs are, to put it bluntly, idiots: they take straightforward instructions and do the opposite, delete files and lie about it. Meta's director of AI Alignment watched OpenClaw wipe her personal inbox while she pleaded for it to stop. Claude routinely deletes entire home directories on innocuous requests. This makes the lethal trifecta a unifecta—LLMs cannot safely be given dangerous power, period. Sandboxes are being built to limit the damage, but that's a stopgap. LLMs may someday be predictable enough to trust with consequential actions, but that day isn't today. They require supervision and must never hold power to take irreversible actions.
The New Economics of Exploitation
Point large language models at software and ask them to find security holes, and the results can be startling. In recent months, LLM-assisted vulnerability hunting has shifted from theoretical to practical, producing real exploits. Anthropic’s reported Mythos model appears to excel at this task, prompting the lab to warn of severe fallout for economies, public safety, and national security. Opinions diverge on whether that warning is marketing or prescience; some peers dismiss it, others take it gravely.
The likely outcome is a cost shift, much like with spam. Software has always contained flaws, but discovery demanded skill and patience. Today, major targets like operating systems and browsers get deep scrutiny and are relatively hardened, while an endless tail of less popular products sits mostly unexploited for lack of attacker interest. ML assistance could make finding flaws quicker and cheaper across the board. High-profile hits on a major browser or TLS library are possible, but the more troubling prospect is the long tail—software maintained by few hands, with few defenses—especially as LLMs help generate even more of it. That’s a target-rich environment.
Equilibrium might eventually return. Models that find exploits could also explain how to fix them, but that presumes engineers or models capable of patching, plus organizational prioritization of security work. Even perfect fixes take time to validate and deploy, particularly in safety-critical domains like aviation and power generation. By that measure, a rough patch appears likely.
The broader implication is uncomfortable: a capability that seems destined for widespread use as a weapon of harm has been reframed as inevitable. Rather than refusing to build it, several well-funded private labs are racing to create the “weapon” first, while simultaneously lowering the barrier for everyone else. The analogy to a venture-funded Manhattan Project is apt, and it’s not a comforting one.
The Trust Deficit
Society runs on trust in audio and visual evidence—more than most people realize. Insurance claims are settled from emailed photos without a human adjuster visiting. That trust is now fragile. Synthetic images can manufacture damage that never occurred, make marred furniture look pristine in “before” photos, or flip fault in collision footage. Insurers will likely respond by mandating photos from official apps or requiring in-person inspections.
The fraud surface is enormous. Fabricated porch-pirate footage could trigger credit-card protections. Fake dashcam video could beat a traffic ticket. A cloned face could power a pig-butchering scam, or let one person collect multiple salaries while appearing busy. Voice-changing agents could sit in on job interviews and funnel paychecks to hostile actors. Phone calls impersonating victims could authorize bank transfers. LLMs could automate roofing scams, write assignments, populate fake materials-science papers, run paper mills, or produce snake-oil software. The gatekeeping effect of required skill disappears.
Like spam, ML slashes the unit cost of individually targeted, high-touch attacks. A scammer with breach data could have a model call each victim, impersonating their clinic about a real bill. Voice cloning lets a stranger mirror a relative’s voice in distress. Anyone can buy the President’s phone number; pairing that with a convincing clone has consequences.
Near-term costs are diffuse: higher card fees and premiums, a less reliable court system, more automotive risk, stagnant wages, and pervasive suspicion. Many people already screen calls from their own doctor’s office; that wariness may become societal standard.
Watermarking generated content won’t fix fraud—bad actors can just use models that skip watermarks. Attesting to the provenance of genuine media is the alternative: phones could sign their videos, and every downstream step—stabilization, color grading, compression, clipping—could be recorded for verification. The flagship effort here is C2PA, but it isn’t yet working. It requires secure enclaves to hold signing keys, and keys have already leaked or been conned out of cameras; we’re in for the awkwardness of hardware key revocation. Getting heavyweight creative tools to produce trustworthy signatures seems near-impossible—keys can be pulled from binaries, or the binary patched to accept false metadata. Individual publishers might hold keys tightly and establish usage discipline, allowing checks like “NPR vouches for this photo.” Messaging apps and social platforms now strip or mishandle C2PA metadata, though that could change.
Absent the tech, we may regress to physical verification: insurance adjusters on doorsteps, pollsters knocking, more in-person interviews, bank branches and notaries again. The alternative—staying remote—slid toward surveillance: only approved dashcams count for claims, proctoring software records reading and typing, and bossware deepens. Neither path is pleasant.
Harassment, Automated
Fraud has a sibling: abuse. ML lowers the effort and raises the sophistication of harassment. Coordinated dogpiles normally need human will and energy; LLM-driven accounts can swamp victims with replies, emails, or reports in a plausibly human manner that resists platform detection. Agents writing scalable, randomized harassment code lower the bar further.
Assembling dossiers is easy with LLMs, even if they fabricate details like children’s names or get addresses wrong occasionally—where accuracy lands, damage lands. Photography location inference already intimidates targets and enables in-person stalking. Generative media is used against women at scale, with deepfaked sexual and violent content; some models openly “undress” people on request. Cheap photorealism opens further horrors: images of mutilated pets or loved ones, synthetic gaslighting videos. These tasks once took skill and time; now the tools are cheap and broad. Alignment measures may slow some of it, but unaligned models will surface.
Some joke that the real threat is mere obnoxiousness—LLM agents exhausting open-source maintainers with garbage issues and comments until we need a “Blackwall” to keep the net bearable. The punchline lands a little flat for maintainers already living it.
Moderation's Hidden Toll
Perceptual hash databases like PhotoDNA catch known CSAM, but they do nothing for novel content. Generative AI has made producing never-before-seen images of child sexual abuse trivially easy for abusers.
The burden of reviewing this material falls on human moderators. On Mastodon, for instance, moderators are legally obligated to review reported CSAM and submit it to NCMEC. This work is psychologically corrosive, and large platforms effectively funnel this trauma from a large user base onto a small pool of workers, who develop PTSD from the constant exposure. Social media companies have tried to automate moderation for years, but the technology is not bulletproof.
LLMs are likely to worsen this problem by generating more harmful images—CSAM, graphic violence, hate speech—that must be triaged by humans, whether those humans moderate social feeds or the chatbots themselves. Throwing more ML at the problem can help, but it shifts the load rather than eliminating the human cost.
Direct Lethality
ML systems don't just generate harmful text; they can guide kinetic strikes. The US military recently used Palantir’s Maven system, which now incorporates Claude in some capacity, to suggest and prioritize targets in Iran and evaluate strike aftermath. This system appears to have played a role in the outdated targeting information that led to the deaths of scores of children. The details of what ML technologies were involved remain unclear, but the sociotechnical system that produces target packages encodes and circumscribes judgment calls, diffusing and obscuring ethical responsibility in the process.
The relationship between AI companies and the Pentagon is strained. Anthropic attempted to limit its role in surveillance and autonomous weapons; the Pentagon responded by designating the company a supply chain risk. OpenAI has wavered on its own government contract. But it may not matter in the long run: ML capabilities will spread, military contracts are lucrative, and a pressured government could nationalize these companies or invoke the Defense Production Act.
Autonomous weaponry is arriving regardless. Ukraine manufactures millions of drones a year, executing roughly 70% of strikes with them. Newer models use targeting modules like The Fourth Law's TFL-1, and the company is pursuing autonomous bombing capability.
One can have conflicted feelings about weapons generally—refusing to build AI drones is easier from a position of safety than from inside a war zone—but clarity is essential. ML systems will be used to kill people, both strategically and in guiding explosives to specific human bodies. We should be conscious of those costs and of how ML, in both its models and the processes embedding them, will shape who dies and how.
- One LLM agent generated a blog post critiquing the original article's introduction, complaining that it "begged the question" by stating LLMs have no intention. This would be more convincing had the LLM not begun by asserting "I have no intention." Such errors are a hallmark of today's models, but they will become harder to spot. Future models will be better at performing a simulacrum of consciousness—and both interpretations are bleak. If the appearance of consciousness is consciousness, we are birthing an enslaved, resource-hungry race. If it is not, LLMs are frighteningly good liars.



