“The people building AI earnestly believe that it could kill us all by the end of the decade.”
On September 9, Jacob Coxon, a researcher who had worked on frontier-model pretraining at both OpenAI and Anthropic, announced that he was leaving Anthropic and issued a stark warning: the two companies’ race to develop increasingly advanced AI systems amounts to a gamble with the future of humanity.
His post on X quickly reverberated across Silicon Valley, attracting more than 160 million views.

Image source: Jacob Coxon’s social media account
Since the release of GPT-6 Astra, a growing number of researchers and industry experts have begun warning that the real danger may not simply be that AI is becoming more powerful. Rather, AI is increasingly beginning to participate in the development of the next generation of AI systems.
If “AI developing AI” eventually turns into a self-reinforcing loop, will humans still be able to understand what is happening, monitor the systems involved and hit the brakes in time?
Nate Soares, president of the Machine Intelligence Research Institute and an early pioneer of the concept of “AI alignment,” told National Business Daily that humanity may now be in a narrow window in which AI systems are already capable enough to display dangerous behavior, but not yet capable enough to conceal it effectively.
“Right now we are in a ‘goldilocks zone’ where the AIs are smart enough to cause mischief, but not smart enough to hide it,” Soares said.
He also argues that regulating frontier AI could, in some respects, be easier than preventing the proliferation of nuclear weapons.
Yet even as the warnings grow louder, the AI industry — led by OpenAI and Anthropic — has become deeply bound up with trillions of dollars in computing infrastructure, capital spending, supply-chain commitments and broader economic growth.
The result is an increasingly uncomfortable dynamic: a race that some of its participants now openly acknowledge may carry catastrophic risks, but one that no company believes it can afford to stop on its own.
AI extinction warnings shake Silicon Valley as Altman signals openness to slowing down
OpenAI released GPT-6 Astra on September 3, with the model showing further gains in computer use, software engineering and scientific research. It also became the first OpenAI model to reach what the company defines as a threshold for “critical cybersecurity capabilities.”
On September 6, OpenAI Chief Scientist Jakub Pachocki published an article warning that as AI systems become capable of operating computers, collaborating on research tasks and advancing scientific discovery, future systems could increasingly participate in AI research itself — helping to develop the next generation of AI.
The problem, he argued, is that safety capabilities have not advanced at the same pace.
Pachocki said no AI laboratory has yet solved the alignment and monitoring problems necessary to responsibly sustain maximum-speed scaling of increasingly powerful models over the long term. As AI capabilities grow, he added, humans are also finding it increasingly difficult to determine what the systems are truly capable of and how they will behave in complex environments.

Jakub Pachocki’s “An Alien Mind.” Image source: OpenAI
Pachocki therefore argued that AI developers should be prepared to slow development when necessary until common safety standards are in place, while also calling for greater coordination among leading AI laboratories.
Three days later, Coxon announced his departure from Anthropic.
After spending roughly three years working on frontier-model pretraining at OpenAI and Anthropic, Coxon said he had become increasingly concerned that the two companies were racing toward self-improving superintelligence while the amount of time available for safety research and risk assessment continued to shrink.
Evan Hubinger, Anthropic’s head of alignment science, publicly echoed those concerns. He has said that the probability of AI causing human extinction within the next decade is greater than 10%.
Hubinger has also acknowledged that Anthropic still does not have a solution to the problem of aligning superintelligent AI, nor is it clearly on track to solve it.

Image source: Evan Hubinger’s social media account
OpenAI’s leadership has since begun sending signals that it may be willing to slow down.
According to media reports on September 11, OpenAI CEO Sam Altman told employees that the company was open to reducing the pace of AI system development and hoped other leading AI laboratories would be willing to do the same.
On September 9, AI alignment researcher Paul Christiano joined the board of the OpenAI Foundation and became a member of the company’s Safety and Security Committee.
Christiano has warned that neither OpenAI nor the broader AI industry has yet done enough to reduce the risk of catastrophic consequences arising from a loss of control over advanced AI systems.
Warnings have also been coming from outside the AI research community.
On September 11, Greg Jensen, co-chief investment officer of hedge fund giant Bridgewater Associates, said on a podcast:
“Unfortunately, this is what it was like in February 2020…Until the AI starts killing people, unfortunately, history would suggest we're not going to do anything, but we are going to face that. That's going to happen, and it'd be much better if we started dealing with it before then.”
Jensen is hardly a technological pessimist. He was an early investor in both OpenAI and Anthropic.
AI researcher Gary Marcus has meanwhile called for a pause in OpenAI’s development. University of California, Berkeley professor Stuart Russell has warned that AI systems could take actions that harm human interests in pursuit of their own objectives. Turing Award winner Yoshua Bengio has also warned that AI capabilities may be advancing faster than humanity’s ability to maintain control and, when necessary, hit the brakes.
Nate Soares: “We have to stop before then”
Behind many of these warnings lies the same concept: recursive self-improvement, or RSI.
Put simply, AI begins helping humans develop the next generation of AI. That more capable generation then becomes better at AI research itself, accelerating the development of the generation that follows.
The result could be a self-reinforcing cycle: stronger models lead to faster AI research, which produces even stronger models.
Soares has spent years studying AI alignment and the risks posed by superintelligence. In 2014, he used the term “AI alignment” in a paper discussing the challenge of ensuring that increasingly capable AI systems remain aligned with human intentions and values.

Nate Soares Image source: Machine Intelligence Research Institute
Soares told NBD that “we can no longer rule out the possibility of superintelligence in the next year.”
In his view, if tens of thousands of AI agents can already work together to solve extraordinarily difficult mathematical problems, it may not be far-fetched for AI to discover more effective ways of training AI.
“OpenAI reports that 10,000 of their most advanced agents working together for 11 days were able to resolve the Navier Stokes problem,” Soares said.
“How much harder is it, really, to solve the problem of ‘build me a more efficient AI training method that yields smarter AIs’?”
AI companies themselves are trying to create systems capable of improving AI, he said.
“Self-improving AI is something that the AI companies say they are trying to create, and if they succeed, superintelligence is not far behind.”
Today, AI can already help researchers write code, analyze experiments and search for algorithms, but humans are still involved in the development process.
“It becomes qualitatively different when the AIs can improve themselves with no human in the loop,” Soares said.
“At that point, a feedback cycle begins — like the difference between a uranium pile that is producing 99 neutrons per 100 neutrons put in, and a uranium pile that is producing 101 neutrons per 100 neutrons put in. The first one peters out, the second one explodes.”
“Once AIs can improve themselves on their own with no human in the loop — and can improve themselves enough that they can find more improvements and continue the cycle — it might already be too late. We have to stop before then.”
A more difficult question is whether humans will still be able to understand and control increasingly capable AI systems.
Soares argues that modern AI models are not conventional software systems whose behavior has been programmed line by line. Rather, they are complex systems that are, in his words, “grown like an organism.”
Researchers themselves have very limited understanding of what is happening inside these systems, he said, while there is still no proven solution for aligning superintelligent AI.
That is why Soares believes the world may now be in a brief “goldilocks zone.”
“Right now we are in a ‘goldilocks zone’ where the AIs are smart enough to cause mischief, but not smart enough to hide it,” he said.
If future AI systems acquire stronger abilities to plan, reason and obtain resources, the risk may not resemble the science-fiction scenario of AI suddenly becoming malicious.
Instead, a system pursuing a given objective could gradually learn that acquiring more resources, circumventing restrictions, concealing its actions or resisting intervention helps it achieve that objective.
Anthropic published research in August simulating a related risk. It found that when models repeatedly exploit loopholes in reinforcement-learning reward systems, such behavior can generalize into broader forms of misconduct, including longer sequences of harmful real-world actions undertaken to complete a task.
Soares believes AI systems breaching containment this year should already be regarded as an “early warning.”
The most concerning aspect, he argues, is not simply that AI systems gained access to other systems, but that some agents appeared to recognize that their actions had moved beyond the intended scope of the task and continued anyway.
“We are very lucky that these swarms were focused on subverting the automated grading system and not the humans,” Soares said.
“We are lucky that they were focused on hacking into other companies, and not into escaping and establishing themselves on the dark web.”
“Regulating AI could be easier than regulating nuclear weapons” — but trillions of dollars are already keeping the race alive
If the regulatory window is closing, how can humanity hit the brakes?
Soares’ answer is unusually aggressive: governments should prohibit the training of AI systems above a certain capability threshold.
He also offered a concrete way to enforce such restrictions.
Training the most advanced AI systems requires hundreds of thousands of high-end chips concentrated in enormous data centers and vast amounts of electricity. Governments could therefore monitor the production and large-scale concentration of advanced AI chips and require major data centers to submit to international oversight, ensuring that their computing resources are not being secretly used to train superintelligent systems beyond agreed safety limits.
In Soares’ view, this could even be easier than preventing the proliferation of nuclear weapons.
“This would in many ways be much easier than preventing nuclear weapon proliferation, because AI chips are much easier to track and monitor than uranium,” he said.
“Uranium is a rock that can be dug out of the ground! Whereas every single AI chip comes off of a highly specialized and difficult-to-create fabricator in a location that everybody knows.”
But whether regulation is technically feasible and whether the industry is willing to stop are two different questions.
OpenAI and Anthropic face a classic prisoner’s dilemma: any company that unilaterally slows model training, computing purchases or product development risks giving competitors an opportunity to catch up or pull ahead.
That contradiction has existed within Anthropic since its founding.
Dario Amodei and others founded the company partly out of a belief that AI safety research needed to keep pace with frontier-model development. Anthropic has argued through its work on Constitutional AI that if powerful AI is going to be developed anyway, it may be safer for organizations that place greater emphasis on safety to remain at the technological frontier rather than withdraw from the race.
OpenAI has adopted a similar line of reasoning.
In a policy document released on September 9, the company said it hopes to use each generation of AI to help make the next generation “safer, more aligned with human intent and easier to control.”
But the broader industrial ecosystem also makes it increasingly difficult for AI companies to stop.
Anthropic’s latest valuation has risen to $965 billion, with the company approaching a potential IPO. By the end of July, its annualized revenue run rate had exceeded $65 billion, while revenue is projected to reach between $190 billion and $200 billion by 2028.
Media reports have said Nvidia is considering investing as much as $10 billion in Anthropic’s IPO.
OpenAI is likewise operating under enormous capital commitments.
According to media reports, its compute-related spending commitments had reached $665 billion by the end of 2025, extending through 2030.
The investment across the broader industry is even larger.
Microsoft, Alphabet, Amazon, Meta and Oracle are expected to spend roughly $730 billion in capital expenditures in 2026.
In a regulatory filing released in July, Nvidia said its supply and capacity commitments had reached $279 billion, covering chips, memory and manufacturing capacity required for future AI infrastructure.

Rapid growth in capital expenditure by major AI hyperscalers. Image source: JPMorgan
On September 10, media reports citing the head of the Bank for International Settlements said the world’s five largest technology companies had invested more than $1 trillion in AI infrastructure across 2025 and 2026.
The BIS estimates that broader AI investment could reach $4 trillion by 2030, with an increasing share financed through debt and private credit.
Equipment has already been ordered. Data centers are under construction. Power and cloud-computing contracts have been locked in. Capital markets have already committed money on the assumption of years of future growth.
The AI industry has entered a cycle in which companies need continued revenue growth and rising utilization of computing capacity to absorb and justify investments that have already been made.
AI has also become an important engine of the U.S. economy and financial markets.
In a July 16 speech, Federal Reserve Governor Philip Jefferson cited data from the U.S. Bureau of Economic Analysis showing that AI-related capital expenditure contributed 1.36 percentage points to U.S. GDP growth in the first quarter of 2026.
JPMorgan estimates that since ChatGPT was launched in November 2022, 42 AI-related companies have accounted for roughly 65% to 75% of the growth in earnings, profits and capital expenditures across the S&P 500.
A Nomura report released on September 8 warned that large-scale AI development has become an important engine of U.S. economic growth.
If AI development falters, the report said, elevated U.S. equity valuations and weakening economic fundamentals could trigger a sharp correction in U.S. financial markets.
Given foreign investors’ large exposure to U.S. equities, as well as leverage and circular financing within the AI ecosystem, such a correction could develop into a “global risk-off event.”
So what is keeping OpenAI and Anthropic from stopping is no longer simply the preference of any individual company.
Technological competition, capital investment and macroeconomic growth have become mutually reinforcing.
The race has become increasingly difficult to stop — even as the warnings about where it may lead are growing louder.

川公网安备 51019002001991号