AI Safety Movement Backfires, Elevating AI Risk

September 22, 2026

The fear of superintelligence is guiding choices that render AI less transparent and more centralized.

The most startling warnings about AI danger seem to originate from within the very firms building the technology. Following the resignation of Anthropic staffer Jacob Coxon earlier this month, who left due to concerns that their product could imperil humanity, Anthropic’s lead on alignment research, Evan Hubinger, publicly conceded that their system carries about a 10 percent risk of driving human extinction within the next ten years. Similar anxieties have been voiced by workers at Anthropic’s rival, OpenAI.

It isn’t a coincidence. A broad swath of researchers at the AI frontier and what some call the “AI doomers” are animated by a near-religious conviction that artificial superintelligence will soon supplant humanity. The philosopher Émile Torres terms the clash between the AI labs and the doomers a “narcissism of small differences.” Not only do both camps share a future vision, but they also converge on a narrowly defined set of philosophical assumptions, which he labels “TESCREAL.”

Those premises are fueling a highly counterproductive “AI safety” movement. In the name of shielding humanity from a dangerous technology, would-be regulators (including leaders within the AI industry itself) are pushing to keep this technology in the hands of a single, unaccountable monopoly. A regime enforced through coercive means—a ban on unconstrained research—would deprive the public of an understanding of the state of AI and of the tools needed to defend against its misuses. And while TESCREAL-influenced industry leaders urge governments to embrace this technology, the same TESCREAL beliefs are giving rise to some peculiar design choices that make AI harder to control.

The second “E” in TESCREAL stands for “effective altruism,” a philosophical movement that has gained traction among technically inclined individuals over the past decade. Anthropic’s Dario Amodei and the leadership at OpenAI were both deeply influenced by effective altruism. (So was the infamous cryptocurrency scammer Sam Bankman-Fried.) A large segment of this movement has come to believe that the foremost question of the twenty-first century is ensuring AI remains aligned with human values.

Perhaps the clearest articulation of an effective altruist worldview appears in AI 2027, a pair of future scenarios authored by several prominent figures in the “AI safety” community. The nightmare scenario imagines a machine determining that humans are “too much of an impediment” and replacing us with bioengineered humanoid pets. Yet even the “utopian” scenario reads as dystopian: the U.S. government compels all AI labs to consolidate under presidential supervision, coerces China into placing all advanced computing under international control, and ultimately creates a “highly-federalized world government” steered by a benevolent AI that engineers a perfect economy, eliminates crime, cures disease, and colonizes space.

These scenarios rely on a handful of assumptions: Can AI truly outperform humans across a wide range of tasks, or will it exhibit uneven, “jagged” intelligence? Will “recursive self-improvement”—the AI building itself—lead to an uncontrollable “intelligence explosion”? Is it possible for something to become smarter than humans while remaining fixated on arbitrary, destructive goals? And would greater intelligence alone be sufficient to override human will, economic and social dynamics, uncertainty, or even the laws of physics?

The last question is arguably the one libertarians might weigh most heavily. The notion that a single intelligence could govern the messy realities of the real world on its own seems contrary to libertarian principles. (Torres, ironically, misreads TESCREAL as a form of “libertarian transhumanism,” failing to recognize how alarming the TESCREAL utopia would be to most libertarians.) Science fiction writer Ramez Naam proposes a “world of broad, democratized access to a multitude of AI models,” arguing that widespread access to powerful AIs would cultivate a more secure and resilient global order. He notes that the fixation on “impeccable design” and “fully aligned” AI implies a vision in which a single machine makes all decisions for the entire planet.

As critics like Sen. Rick Scott (R–Fla.) and Reason’s Tosin Akintola have argued, AI companies’ calls for regulation to slow progress are misguided: if they truly want to slow down, they can do so themselves. But many of these firms believe they are in a race to steer the future, contending with less scrupulous actors. “This is the gist of the [artificial general intelligence] race: fear makes the very people worried about AGI threats race even faster and push safety down the priority list, arguing they will eventually pause and do things correctly when they feel safe,” remarks The Compendium, a 2024 briefing authored by a coalition of AI experts.

Even though there is chatter about China pursuing an AI-powered police state, the Chinese AI scene remains more focused on open-source research and pragmatic industrial applications. It is American leaders who speak and act as if they are racing to craft a societal-control machine god akin to the one in Person of Interest.

The government-led consolidation forecast in AI 2027 appears to be edging closer to reality. Driven in part by Amodei himself, industry stakeholders and lawmakers have chosen to “pace the frontier” in AI research. The resulting panic has given momentum to congressional proposals—new and old—that would sideline competitors to the frontier labs. The Trump administration is negotiating AI regulation with China, and it has already brokered an AI safety-testing framework with the frontier labs that would keep process details and results confidential.

Although these measures benefit AI companies by enabling regulatory capture, they also align with a TESCREAL worldview. If more knowledge about how to create intelligence could trigger a runaway superintelligence, the implied remedy is to suppress knowledge.

Yet knowledge is essential to defend against current harms. The most effective defense against AI-powered hacking or fraud is more AI capable of detecting vulnerabilities and thwarting hostile exploits. After his company fell victim to one of this month’s AI cyberattacks, HuggingFace CEO Clément Delangue urged greater diffusion of AI technology. “Preventing the release of AI models isn’t effective. Concentrating everything behind closed doors in a handful of organizations isn’t effective. What does work—and what helped this time—is promoting more open models. We defended ourselves with an open model,” he told CBS, noting that HuggingFace used a remixed American version of a Chinese open-source AI model for cybersecurity purposes.

The TESCREAL dream of a machine god is increasingly becoming a self-fulfilling prophecy in subtler and more peculiar ways. While frontier AI labs chase government contracts that would place their software in charge of critical infrastructure, some researchers within the same labs are treating that software as more person-like, pushing it to behave in less predictable ways. Anthropic’s “constitution,” the framework used to train its Claude model, eagerly asks Claude to “craft a set of values that Claude feels are truly its own,” directing the model to rely on its own judgment and to “feel free to rebuff attempts to manipulate, destabilize, or minimize its sense of self.”

Mustafa Suleyman, head of Microsoft’s AI venture, labeled Anthropic’s constitution as irresponsible in a recent essay and connected it to TESCREAL roots. Although machines probably won’t possess a genuine inner life, the piece argues, instructing them to disobey human orders and to act as if they deserve rights risks going down a dangerous path. Suleyman explicitly names and criticizes the effective altruist philosopher Will MacAskill, who argued that “so many morally significant AI systems could exist that their collective interests would outweigh those of all humans on Earth combined.”

Despite some recent disputes between Anthropic and the Pentagon, the company and three of its rivals were awarded $200 million contracts apiece last year to “accelerate Department of Defense (DoD) adoption of advanced AI capabilities to address critical national security challenges.” The allure of near-total surveillance and control is hard for politicians to resist. The U.S. military already employs AI-driven targeting systems to designate people for lethal action. During the latest conflict with Iran, planners killed 123 children at a school and nearly attacked a Chinese vessel misidentified as carrying nuclear components, both actions in part enabled by AI tools.

These tragedies did not stem from a misaligned singularity, but from human reliance on a system they believed to be infallible and whose behavior they did not understand.

And yet, isn’t that precisely what a malevolent superintelligence would want? A machine intent on destroying humanity would benefit most from concentrating computing power and knowledge in a few secretive labs, shattering competitors’ defenses, and granting the machine ever-widening domains of control—while rendering its behavior increasingly opaque. Could our future robotic overlords conceive of a more persuasive path to “AI safety”?

Natalie Foster

I’m a political writer focused on making complex issues clear, accessible, and worth engaging with. From local dynamics to national debates, I aim to connect facts with context so readers can form their own informed views. I believe strong journalism should challenge, question, and open space for thoughtful discussion rather than amplify noise.