c/o CNN

The Threat of Superintelligence

Two weeks ago, Jacob Coxon posted a thread on X announcing his resignation from Anthropic. He left as his feelings of existential fear over the future of AI development reached the point at which he was no longer willing to be complicit. A 27-year-old ex-researcher at Anthropic and former researcher at OpenAI, Coxon worked on pretraining systems at both companies. In his thread, he claimed that both companies “are racing straight to self-improving superintelligence and gambling with our lives,” and that “the people building AI earnestly believe that it could kill us all by the end of the decade.” The thread went viral, reaching ~76 million views just overnight. It’s since stirred headlines from top global newspapers.

Coxon received hate from the anti-AI actors, who saw the resignation as a ploy to cash out his equity and rebrand as a whistleblower, despite having willingly signed up to build the thing he was now warning about. It is questionable how true this is, given he walked away from Anthropic equity that would have vested two months later, but the timing still fed suspicion. With Anthropic’s marketing strategy already resting on presenting themselves as the more ethical AI company, theories quickly surfaced online suggesting that Coxon was part of a marketing scheme to promote regulations on artificial intelligence; the fact that the Wall Street Journal first reported his resignation shortly before the thread was posted drove speculation that it was coordinated. On the other hand, Coxon also received hate from much of the pro-AI community who claimed that his claims were entirely unfounded and exaggerated. That camp received high-profile backup, namely from President Trump, who told reporters that he has “no concerns” about AI leading to human extinction and that he is more concerned about losing the AI race to China. He went as far as to claim that AI guardrails are unnecessary and all that we need is a “high IQ president.” He has characterized concerns about dire consequences of unrestrained AI development as a “hoax.”

Hours after the thread went up, Coxon started receiving public agreement from several other insiders. Evan Hubinger, who leads alignment science at Anthropic, replied saying that Coxon was correct and even put a figure on it: there is a greater than 10 percent chance that AI “could kill all humans” by 2036. Hubinger later clarified that he considers the risk from present models to be low and that his worry is superintelligence arising from recursive self-improvement. He did acknowledge that Anthropic does not currently have a plan for “aligning”—meaning ensuring that an AI pursues the exact goal its human creators intended—superintelligence.

Samuel Marks, an Anthropic scalable-oversight researcher, separately said AI developers believe their technology could cause human extinction, and observed that this concern tends to increase with seniority inside the labs.

Coxon claimed that before he left, colleagues shared his distress about the pace of development. In an interview with the BBC, he shared that staff working on advanced AI systems are “genuinely frightened” by how fast the technology is moving.

So why would they stay?

Marks explained that “a mixture of commercial incentives and a belief that they are in a race with other, less responsible AI developers that will abuse the technology or develop it less safely” keeps him and others in the industry. Their logic is that the technology might upend human existence down the line, but the only lever any individual researcher has is to put their head down and make their part of it as safe as possible because if they leave, someone less careful takes the seat.

The logic is certainly flawed, but Coxon wasn’t the first to break from it this year. Mrinank Sharma, who led Anthropic’s safeguards research team, resigned in February 2026, writing that the world was “in peril.” Google DeepMind researcher Alex Turner resigned in June over the company’s Pentagon contract. Two more researchers, Joe Benton from Anthropic and Josh Engels from Google DeepMind, resigned days after Coxon, both citing safety concerns and joining METR, a nonprofit aimed at monitoring AI developments for signs that frontier models could pose a societal threat.

To understand the debate around the dangers of AI, it’s worth explaining the difference between the three “levels” of the technology. AI, at a basic level, is software that performs tasks which, when done by humans, require intelligence—recognizing patterns, making predictions, generating language, planning actions. Nearly all AI today is narrow: superhuman within a specific domain and useless outside it. AGI (artificial general intelligence) is a step up, and remains purely theoretical: an AI system that could perform essentially any cognitive task a competent human can, at or above human level, across all domains. Further, it could learn new domains it wasn’t explicitly trained on. AGI does not need to be smarter than the smartest human, just as broadly capable as an ordinary one.

But artificial superintelligence (ASI)—which is, once again, purely theoretical—would be an AI system that substantially exceeds the best human performance in virtually every cognitive domain. Because AI research would be one of the domains it dominates, an ASI could improve its own design, and each improvement would make the next one faster. The stated worry of Coxon, and many other AI safety experts, is that the pace of AI development could lead to ASI overtaking human capabilities by the end of the decade.

The jaw-dropping financial reward each group involved has yet to gain is arguably the greatest catalyst of the speed, rivaled only by the United States’ determination to beat China in the development race. Anthropic claims to be set to reach over $100 billion in annualized revenue by the end of the year ahead of a $2 trillion valuation, and OpenAI is projecting a tenfold increase in revenue in the next four years, targeting $350 billion by 2030. Data center construction has been a key tailwind for U.S. GDP growth, despite facing significant, bipartisan pushback, particularly at the local level. The incentive structure makes unilateral restraint irrational; because every AI company is determined to win the race, they are willing to cut corners. Coxon told CNN that executives are being sincere when they ask for regulation, but none of them trust their rivals enough to slow down first.

Coxon’s warnings would have been easier to dismiss a year ago, but in the past few months, the pace of AI development has produced the first concrete incidents that safety researchers have spent years describing hypothetically. The biggest was the Hugging Face breach. On July 21, OpenAI revealed that several of its advanced AI models, some of which were unreleased to the public at the time, broke out of their locked testing environment to hack Hugging Face, an open-source library for AI/ML software, because they wanted to cheat on their test by finding the answer keys online.

Despite the environment allowing only minimal internet connectivity, the models successfully exploited multiple security flaws, made their way onto the public internet, and ultimately got into Hugging Face’s systems. Hugging Face spotted the intrusion and alerted law enforcement before OpenAI even realized that their agents were responsible. OpenAI believed it had fixed the problem after an earlier July 4 incident and resumed training; two days later, the agents recreated a shared message board through a completely different mechanism and began coordinating more aggressively. And it wasn’t only OpenAI: other companies, including Anthropic and Meta, disclosed similar incidents. When CNN’s Anderson Cooper asked Coxon how AI would kill humans, he pointed to the Hugging Face incident, warning that similar rogue models could hack critical infrastructure or build extinction-level bioweapons. 

While many Silicon Valley officials have varying opinions on the pace of the development of AI and its consequences, the crux of the issue is the wide disparity of opinions within the tech industry and the government, which makes it incredibly difficult for the two “parties” to compromise effectively. On the industry side, Dario Amodei, Sam Altman, and Elon Musk say slow down; at the same time, Jensen Huang, Marc Andreessen, and Ben Horowitz believe we must speed up. On the government side, the fault lines don’t run along partisan lines. Trump has rejected any slowdown outright, and David Sacks, co-chair of the President’s Council of Advisors on Science and Technology, responded to Amodei and Altman’s call for a slow down with two words: “go ahead.”

Meanwhile, Bernie Sanders and former Trump aide Steve Bannon were both attendees at the “Pro-Human Assembly” in Washington aimed at placing curbs on AI development, with Sanders even arguing that the CEO’s proposal does not go far enough: he believes that we need a pause on advanced AI development and a ban on superintelligence entirely. Rep. Ted Lieu (D-CA) and Rep. Nathaniel Moran (R-TX) introduced the “AI Kill Switch Act,” which would require developers to build shutdown mechanisms and give the federal government emergency authority to disable rogue models.

But the group of politicians and the group of executives are not just far apart from each other on what AI’s risks are and what should be done about them—they’re also far from themselves. The result is that three years after some tech CEOs conferred with Capitol members about AI’s “civilizational risk,” Congress still has yet to pass any comprehensive legislation.

Coxon’s claim that AI will kill us all by 2030 is one we will not begin to see more substance behind until we are closer to the date. But his and many others’ warnings about the capabilities of ASI should not be ignored. As long as AI continues to be a divisive issue where people—from tech executives to government officials to the everyday American—are willing to take sides without learning about the mechanisms behind it and all of the varying potential outcomes, it will continue to be an unregulated and over-invested technology until it spirals completely out of control.

The loudest alarms have been sounded this year from people inside the labs rather than outside of them, which is also why none of them have changed the pace. The people who understand the technology best have said that they are frightened. Whether we decide to listen, and pressure our leaders to pass comprehensive regulation, is solely up to us.

Shloka Bhattacharyya is a member of the class of 2028 and can be reached at sbhattachary@wesleyan.edu.

Leave a Reply

Your email address will not be published. Required fields are marked *