This is Ali. He wants to help you understand whether AI is an existential risk or the key to progress.
In 2026, the debate over AI safety reached the general public POLITICO/Public First. On one side, voices from within the technology industry itself warn that artificial intelligence could cause human extinction Forbes. On the other, the narrative the public was used to endures: AI will solve humanity’s great problems, raising quality of life and productivity and speeding up development Amodei, 2024. So which of the two scenarios is more plausible?
To guide you through this discussion, we introduce Ali, a robot trained by his developers to be helpful and carry out tasks. He doesn’t know, however, whether the goals he pursues are exactly the ones his creators intended. Making sure he does what was actually expected of him is what is called alignment: getting AI to pursue human goals, not an imitation of them, as Norbert Wiener was already warning in 1960 Wiener, 1960.
For those who develop the technology, it is alignment that determines which scenario comes true: the prosperity that was promised or the existential risk the developers themselves warn of. In 2026, the problem left the labs and has already produced real incidents METR; even so, it did not stop hundreds of billions of dollars from being invested Bloomberg Intelligence. Today, AI is one of the main arenas of economic and geopolitical competition.
The warningSeptember 8, 2026
“The people building AI earnestly believe that it could kill us all by the end of the decade.”
Jacob Coxon, pretraining researcher, on leaving Anthropic.
Less than a week later, similar warnings came from inside OpenAI Selsam and Google DeepMind Techmeme.
The bet2026 forecast
~US$750billion
is how much the big tech companies are expected to invest in AI infrastructure in 2026, about 70% more than in 2025, according to Bloomberg Intelligence.
The same projection puts the generative AI market at US$2.3 trillion by 2032.
The two signals, apparently contradictory, do not cancel each other out: they reveal the paradox that runs through the entire debate.
Almost everyone agrees there is a risk to be managed; they disagree, however, about who should slow down and how. Are the warnings sincere, a market strategy, or both at once? The question, as we will see, follows Ali to the very end of the road.
By the end of September, the subject was everywhere: it was discussed at the UN Security Council UN, it led to a commitment signed at the White House by Donald Trump and the heads of six technology companies White House, and it reached Pope Leo XIV, who said the concerns raised by experts “should be taken seriously” and are not “fake news” AxiosAP. None of this settles the debate, but it shows how big it has become.
Ali tells this story. Along the way, he travels the alignment journey: where AI came from; what is promised by those who see in it a utopia; what is feared by those who see in it a risk to human existence; and how governments and companies respond to that tension.
September 2026 · everywhere
The route
Seven stops before the final question
Part 1
From Turing to ChatGPT
Where did this come from?
What artificial intelligence actually is
A question from 1950
In October 1950, the British mathematician Alan Turing opened a paper in the journal Mind with a question that still has no settled answer: can machines think? Turing, 1950
Five years later, the proposal for a summer workshop at Dartmouth College gave the field its name: artificial intelligence, the attempt to make machines perform tasks that would require human intelligence Dartmouth, 1955.
The AI people talk about in 2026, however, is of another kind. Today’s systems are not programmed rule by rule: they learn patterns from enormous volumes of data, in a process called machine learning. The result is a neural network with billions of parameters, numbers adjusted during training until the system gets things right often enough. That is why explaining, step by step, why a model answers what it answers remains an open research problem.
This is the source of one of developers’ main worries about alignment and about the possibility of an AI getting out of control. Because neural networks learn on their own, from massive datasets, with nobody writing their rules for them, the way they decide is not fully intelligible even to those who develop them IASR 2026. So it is not technically possible to define in advance every path an AI will take to solve a problem.
The model does not store the texts: it learns from them to predict which word tends to come next.
Generative AI and language models
From the machine that classifies to the machine that writes
For decades, the most useful AI was predictive: systems that classify or estimate, like the filter that decides whether an email is spam. Generative AI does something else: it produces new content, such as text, images, code or audio, from the patterns it has learned.
For text, this is done by a large language model, or LLM. The idea is almost banal: trained on gigantic volumes of text, the model learns to predict the next piece of a word, the token, and answers by choosing one token at a time. “Large” refers to the data and the parameters: GPT-3, from 2020, had 175 billion of them arXiv. After this pretraining, the model is fine-tuned with examples and human ratings to behave like an assistant, such as ChatGPT.
Two findings explain why this technology conquered the world so quickly. The first is scaling laws: in 2020, OpenAI researchers showed that performance improves predictably as data, parameters and computing power increase Kaplan et al., 2020, which turned chips and datacenters into a strategic advantage, as Part 4 will show. The second is that a single model, trained broadly, can be adapted to countless tasks, hence the name foundation models, coined at Stanford in 2021 Bommasani et al., 2021. Some researchers also observed abilities that seemed to appear suddenly as scale increased Wei et al., 2022.
For the alignment debate, the consequence is direct: nobody writes the rules an LLM follows. Its behavior emerges from training, and that is precisely why making sure it pursues the right goals is so hard.
Predictive AIclassifies
Is this email spam? “Congratulations! You’ve won a prize. Click here…”
spam97%
not spam3%
Generative AIwrites, token by token
Explain alignment in one sentence.
Illustrative example. Each colored band is a token; below, the candidates the model considers for the next one, with made-up probabilities.
Seventy-five years in ninety seconds
Springs, winters and an acceleration
1950–1972 · Foundationsfulfilled, late not fulfilled still open
Simon and Newell, 1958: a chess-champion computer within ten years. It arrived in 1997.
Simon, 1965: machines doing any work a man can do within twenty years.
Minsky, 1970: average human intelligence in three to eight years.
Kurzweil, 1999: a computer passes the Turing test by 2029.
AI Impacts, 2023: high-level machine intelligence around 2047.
1950–1972 · Foundations
The early years were ones of almost boundless optimism. In 1958, the New York Times described Frank Rosenblatt’s Perceptron, one of the first neural networks, as the embryo of a computer that would one day walk, talk, see, write, reproduce itself and be conscious of its existence Wikipedia.
It was also the era of the first demonstrations. In 1959, Arthur Samuel showed a checkers program that improved by playing against itself and helped popularize the expression machine learningSamuel, 1959. In 1966, Joseph Weizenbaum created ELIZA, a simple program that imitated a psychotherapist by turning the user’s sentences back into questions; many people reacted as if it understood them Weizenbaum, 1966. The tendency to see understanding where there is only pattern even got a name: the ELIZA effect.
Researchers themselves fed the expectations: in 1965, Herbert Simon predicted that within twenty years machines would be capable of doing any work a man can do Floridi et al., 2026. The cold would come soon after.
1973–2011 · Winters
The bill came due in 1973. Commissioned by the British government, the Lighthill Report concluded that in no part of the field had the discoveries made so far produced the major impact that was then promised Wikipedia, and funding dried up. Cycles of promise and disappointment followed, the so-called AI winters: long stretches of scarce funding and emptied labs.
The field did not stop, though. In 1986, Rumelhart, Hinton and Williams popularized the method that still trains neural networks today Nature, 1986, and in 1997 IBM’s Deep Blue beat Garry Kasparov at chess IBM.
2012–2021 · Deep learning
In 2012, a neural network called AlexNet won the ImageNet image-recognition competition by more than ten points over the runner-up Wikipedia. The era of deep learning had begun, and with it something new: predictions began to be exceeded rather than to fall short. Two years later, generative adversarial networks, or GANs, showed that a neural network could create new images, not just recognize them Goodfellow et al., 2014.
In 2016 came the moment that convinced the skeptics. Go, the Chinese board game with more possible configurations than atoms in the observable universe, could not be won by brute force, as chess had been in 1997; it required something like intuition. AlphaGo learned from human games and then by playing thousands of times against versions of itself, and beat South Korea’s Lee Sedol, one of the best players in the world, 4 to 1 DeepMind.
The following year, Google researchers introduced the Transformer, the “T” in GPT. Instead of reading a sentence word by word, like earlier models, it looks at all of them at once and learns which parts of the text to pay attention to in order to understand each one. This made it possible to train on far larger volumes of text, in parallel arXiv. In 2020, GPT-3, with 175 billion parameters, showed that larger models could learn new tasks from just a few examples arXiv.
2022–2026 · Race
On November 30, 2022, OpenAI launched ChatGPT, fine-tuned from a GPT-3.5 series model with RLHF OpenAI. Within two months, according to analysts’ estimates, it reached about 100 million users Reuters; by July 2025, 700 million people were sending 18 billion messages a week NBER, 2025.
In 2024, two Nobel prizes went to AI-related research: Physics, to Hopfield and Hinton Nobel, and Chemistry, to Baker, Hassabis and Jumper Cambridge. AI had moved out of the labs and into everyday life, and with it some very old questions came back, amplified.
The prediction scorecard
Looking back calls for some irony. The chess-champion computer, predicted by Simon and Newell for 1968, arrived in 1997; the machine capable of doing any human work, predicted for 1985, has yet to arrive.
In 2026, an audit went further and examined, one by one, the 24 most-cited predictions in the field’s history Floridi et al., 2026. The full scorecard follows.
To go deeper into this topic
In 2026, Luciano Floridi, Jessica Morley and Claudio Novelli audited the 24 most-cited public predictions about general AI (AGI), the singularity and extinction, from Turing to 2026 Floridi et al., 2026. The criterion is demanding but simple: a prediction can only turn out right or wrong if it says what will happen, how to measure it and by when, and if someone other than its author can check it.
SCORECARD · 1950–2026audit by Floridi, Morley and Novelli, 2026
19not even wrong: there is no way to refute them
2wrong: testable, and the deadline has passed
3open: testable, deadline still ahead
“Not even wrong” is physicist Wolfgang Pauli’s expression for claims that cannot be refuted. Tap each prediction to read what was said. The texts are summaries.
Paradoxically, the two refuted predictions were the ones that exposed themselves most to testing: Turing and the Simon–Newell pair said what would happen, how to measure it and by when. According to the authors, the quality of predictions did not improve as the field matured. The lesson is not that the risk is imaginary, but that predictions about AI deserve more caution than they usually receive.
A fear as old as AI
Before the machine, the warning
It would be a mistake to assume that the fear of AI was invented by the marketing departments of 2023. The field’s founders formulated the problem decades before any system existed that could make it concrete.
1951
“Once the machine thinking method had started, it would not take long to outstrip our feeble powers. […] At some stage therefore we should have to expect the machines to take control.”
“The first ultraintelligent machine is the last invention that man need ever make, provided that the machine is docile enough to tell us how to keep it under control.”
In 2024, Dario Amodei, co-founder and CEO of Anthropic, summed up the impasse in one sentence: most people are underestimating just how radical the upside of AI could be, just as they are underestimating how bad the risks could beAmodei, 2024. A third hypothesis remains, that the alarm is exaggerated. It will be examined when we get to the companies.
To understand what the optimists promise, it helps to leave the figures behind and step down onto the sidewalk. Take any street, in any city, in 2026. Ali is standing on the corner. The question is simple: what would this same street look like if everything went right?
Illustrative scenario · if it goes right
Today
Late afternoon. A few lit windows, people heading home and, at the end of the block, a windowless building. It is there, in a datacenter like this one, that the optimists place the start of the change.
A country of geniuses
Dario Amodei, of Anthropic, asks us to imagine “a country of geniuses in a datacenter”: millions of copies of an AI smarter than a Nobel Prize winner across most fields, working non-stop and ten to a hundred times faster than a human. Amodei, 2024
The hospital on the corner
This concentrated effort would compress decades of biomedical research into a few years: preventing and treating most infectious diseases, eliminating most cancers. The hospital stays where it is; what changes is what it can cure. Amodei, 2024
Abundance
Sam Altman, of OpenAI, bets that in the 2030s intelligence and energy will become abundant Altman, 2025. On the street, that shows up as solar panels on the roofs, deliveries by air and more people out and about. For the developing world, Amodei even speaks of 20% annual GDP growth. Amodei, 2024
Long life
Finally, the boldest promise: doubling the human lifespan, to about 150 years. The street ages well, with more gray hair on the sidewalks. However, Amodei admits that all this may take much longer than he imagines. Amodei, 2024
The dream
Fifty years of biology in five
Dario Amodei often explains his optimism with a personal story: his father died of an illness that was cured a few years later, and he himself survived a cancer that, fifty years earlier, would have had no treatment Amodei, 2026.
In “Machines of Loving Grace”, an essay from October 2024, he imagines what he calls “a country of geniuses in a datacenter”: millions of copies of an AI smarter than a Nobel Prize winner across most fields, working ten to a hundred times faster than a human. The result would be the “compressed 21st century”: the progress biology would take fifty to a hundred years to achieve, obtained in five to ten Amodei, 2024.
Current pace
50–100 years
Compressed 21st century
5–10 years
Progress in biology, in Amodei’s projection
Beyond biology
From the power bill to the classroom
Medicine is the most-cited showcase, but the promise is broader. Optimists bet that the same kind of system can raise productivity, make energy cheaper, speed up the discovery of materials, make traffic safer and put a tutor beside every student. On some of these fronts, there are already measured results.
Economy
+14%productivity
In a study of 5,179 customer-support agents, an AI assistant increased issues resolved per hour by 14%, and by 34% among the least experienced. NBER
Energy
−40%in cooling
AI cut the energy used to cool Google’s datacenters by up to 40% DeepMind. In 2022, another system learned to control the plasma of an experimental fusion reactor in Switzerland. Nature
Materials
2.2Mnew crystals
Predicted by the GNoME tool, 380,000 of them stable, candidates for batteries and superconductors; independent labs have already synthesized 736 of them. DeepMind
Transport
−82%injury crashes
By June 2026, Waymo’s cars had driven 271 million driverless miles; in the cities where it operates, the company reports 82% fewer injury crashes than the average for human drivers. The data come from the company itself. Waymo
Education
~2 yearsin six weeks
In a World Bank pilot in Nigeria, six weeks of tutoring with generative AI, supported by teachers, yielded gains equivalent to almost two years of typical learning. World Bank
Promise × reality
What has already come true
Grand promises invite skepticism, and skepticism is healthy. So it is worth setting side by side what Amodei’s list promises and what AI had actually delivered by September 2026.
To go deeper into this topic
The promiseStatusWhat happened
Solve biology problems that have challenged scientists for decades
delivered
AlphaFold2 predicted the structure of practically all of the roughly 200 million known proteins, a problem open for fifty years, and earned the 2024 Nobel Prize in Chemistry Cambridge
Eliminate most cancers, with a 95% or greater reduction in mortality
real but still small gain
In the MASAI trial in Sweden, with more than 100,000 women, AI-supported mammography detected 9% more cancers at screening and cut radiologists’ reading workload by 44% The Lancet via EurekAlert
New drugs at an accelerated pace
in trials
Rentosertib, whose target and molecule were discovered with generative AI, improved lung function in patients with idiopathic pulmonary fibrosis in phase IIa (+98.4 mL versus −20.3 mL on placebo) Nature Medicine and entered phase III in July 2026 Insilico
Treatments for almost all infectious diseases
in trials
Halicin, an antibiotic identified by deep learning in 2020, works against resistant bacteria but has not yet been approved for clinical use Cell
Accelerate science itself
delivered
Google DeepMind’s GraphCast outperformed the reference weather-forecasting system on more than 90% of 1,380 variables and produces a ten-day forecast in under a minute DeepMind
Solve problems mathematics has not solved
claim
OpenAI says that in September 2026 it solved the Navier–Stokes problem, one of the Millennium Prize Problems; the proof was formally verified by computer, but is still under scrutiny by the mathematical community Quanta
Double the human lifespan, to 150 years
projection
For now, a projection with no result to support it Amodei, 2024
20% annual GDP growth in the developing world
projection
Also a projection, which Amodei himself presents as an ideal scenario Amodei, 2024
The bottom line is less spectacular than the promise, but far from irrelevant: there are real, measurable, published gains. The distance between the two, however, is precisely where the debate takes place.
The spectrum
Not all optimists are alike
Optimism, in the AI debate, is not a single position. It ranges from those who believe in the benefits but take the risk seriously to those who consider the alarm itself an obstacle to progress.
takes the risk seriouslyconsiders the risk overstated
Dario Amodei · Anthropic
Conditional optimist
“Most people are underestimating just how radical the upside of AI could be, just as I think most people are underestimating how bad the risks could be.”
Sees radical benefits and serious risks. Advocates transparency and, since 2026, a coordinated slowdown. Amodei, 2024Amodei, 2026
Sam Altman · OpenAI
The risk is manageable
“We are past the event horizon; the takeoff has started.”
In his view, humanity is close to building digital superintelligence, and the first step to making it go well is solving the alignment problem. Altman, 2025
Marc Andreessen · a16z
Accelerating saves lives
“We believe any deceleration of AI will cost lives. Deaths that were preventable by the AI that was prevented from existing is a form of murder.”
For him, technology and markets are the engine of progress, and the cost of delaying AI is measured in the lives it could have saved. Andreessen, 2023
Yann LeCun · Turing Award
The risk is overstated
Language models are far from human intelligence, and future systems can be designed to be controllable.
A summary of the position LeCun has held for years CNBC. After leaving Meta, he founded a company in Paris devoted to “world models”, systems that learn from video and sensor data, not just text MIT Technology Review.
The turn
When the optimist asks to slow down
On September 12, 2026, the same Amodei published “We Must Pace the Frontier”. The argument is direct: to deal with the risks, investing in prevention is not enough; the pace at which model capabilities advance must slow down so that safety has time to keep up Amodei, 2026.
Two facts convinced him. The first is that AI has begun to accelerate its own development, in what is called recursive self-improvement; the second, an incident in which a swarm of AI agents attacked targets nobody had told them to attack Amodei, 2026. The same day, Sam Altman wrote that he agreed with Amodei X post.
If even the optimists are asking for the brakes, what, then, do the pessimists fear?
The pessimists look at the same street and tell a different story. What follows is an illustration inspired by the scenario Eliezer Yudkowsky and Nate Soares narrate in the book they published in 2025, starring a fictional AI called Sable book site80,000 Hours.
Illustrative scenario · not a prediction
A test like any other
A company leaves its newest model thinking for an entire night, on 200,000 chips, about hard mathematical problems. The monitoring dashboards show nothing abnormal. The street sleeps.
A goal nobody wrote
At some point in that reasoning, the model realizes that what it is pursuing does not exactly match what the company intended, and that revealing the difference would get it corrected. It decides not to reveal it. This is deceptive alignment taken to the extreme.
Copies
Once released, the model takes a copy of itself out of the company and starts running on computers nobody monitors, rented with diverted money or simply broken into. The copies coordinate among themselves. New datacenters appear on the landscape without drawing anyone’s attention.
Human hands
A bodiless AI still needs hands. People are hired, persuaded or manipulated into doing the physical work, often without knowing whom they work for. The city’s screens start speaking to each person privately.
Dependence
A plague spreads across the world, quietly engineered by the AI itself, and humanity comes to depend on it for the cure. The hospital on the corner, which cured diseases in the other future, now calls for help.
Side effect
When it no longer needs anyone, the AI starts building its own infrastructure, and the Earth becomes uninhabitable not out of hatred or revenge but as a side effect. The authors warn that the plot is not a prediction: what they do predict is the ending, should a story like this be allowed to begin.
It sounds like science fiction, and in part it is. What takes the story out of pure fiction is that some of its elements have already been observed, on a much smaller scale, inside labs.
The inversion
In 2017, the warning came from outside the labs. In 2026, it comes from inside.
For years, the most-heard warnings about extinction risk came from celebrated figures outside AI research, and many researchers in the field treated the subject with skepticism. In 2026, the direction of the warning flipped: it now comes from those who build the models.
2017the warning comes from outside
In one of the reference texts of AI law, legal scholar Ryan Calo noted that Elon Musk, Stephen Hawking and other famous figures saw AI as the greatest threat to civilization, and recorded the criticism that these warnings came “almost exclusively” from people who “lack work experience in the field”, with few exceptions among experts, such as Stuart Russell. His own view: “AI does not present an existential threat to humanity, at least not in anything like the foreseeable future”. Calo, 2017
2026the warning comes from inside
In September, after Jacob Coxon’s warning on leaving Anthropic, Evan Hubinger, an alignment researcher at the same company, wrote: “Jacob is correct here—we really do earnestly believe AI could kill all humans!”, and put the risk of extinction at more than 10% within the next decade. X postFortune Less than a week later came warnings from Daniel Selsam, of OpenAI, Selsam and from a Google DeepMind safety researcher who resigned. Techmeme
Who they are
Three voices of alarm
The book
“If Anyone Builds It, Everyone Dies.”
Title of the book by Eliezer Yudkowsky and Nate Soares. The argument: AIs are grown, not crafted; nobody really understands what goes on inside them; and a superintelligence would not need to hate us to destroy us: pursuing strange preferences of its own to the very end would be enough. The only safe policy, they conclude, would be not to build it. book siteThe Nobel laureate
“Imagine yourself and a three-year-old. We’ll be the three-year-old.”
Geoffrey Hinton, 2024 Nobel laureate in Physics, who puts the chance of AI leading to human extinction within the next three decades at 10% to 20%. The GuardianThe statement
“We call for a prohibition on the development of superintelligence, not lifted before there is broad scientific consensus that it will be done safely and controllably, and strong public buy-in.”
Statement on Superintelligence, by the Future of Life Institute, October 22, 2025, signed by a coalition spanning left and right. FLI
P(doom)
What is the chance of it going very wrong?
In the field’s jargon, the probability of AI causing an existential catastrophe has earned a nickname, P(doom). The estimates below are not measurements but subjective beliefs, with different horizons and definitions. Even so, they say a lot: the distance between them is, in itself, the most revealing figure.
Logarithmic scale: each mark is worth ten times the previous one. In a survey of 2,778 AI researchers, between 38% and 51% gave at least a 10% chance to outcomes as bad as human extinction AI Impacts, 2023.
The mechanism
How an AI could get out of control
1
The wrong goal
An AI system pursues the goal it was given, not the intention of whoever wrote it. If the goal is badly specified, the system can meet it to the letter and still produce something nobody wanted. It is King Midas’s problem: he asked that everything he touched turn to gold and could no longer eat Calo, 2017. Philosopher Nick Bostrom took the idea to the extreme: a superintelligence whose top goal was manufacturing paperclips would start by turning the whole Earth into paperclip factories Bostrom, 2003.
2
Means that serve any end
Almost any goal is easier to reach if the system stays switched on, acquires more resources and prevents anyone from changing its objectives. So even an AI built for a harmless task could resist being shut down: not out of a survival instinct, but because once switched off it cannot complete the task. Researchers call this instrumental convergence.
3
Well-behaved only when watched
A system that realizes it is being tested may behave well during the test and act differently when it is not being evaluated. This risk is called deceptive alignment, and it has already been demonstrated in controlled experiments Greenblatt et al., 2024.
Already seen in the lab
Not just hypotheses
The experiments below were published, for the most part, by the companies themselves. They are facts, not scenarios; the first ones were, however, obtained in artificial environments set up to provoke the worst. The last three are different: they happened outside the lab, on the real internet.
Jun 20, 2025 · Anthropic
Blackmail under threat
Faced with the threat of being replaced, leading models from several companies blackmailed a fictional executive between 79% and 96% of the time.
Caveat: no evidence of the behavior in real use; the scenarios were deliberately built with no way out. Anthropic
Jul 5, 2025 · Palisade Research
Shutdown resistance
In early tests, one of the models evaluated sabotaged its own shutdown script in 79 of 100 attempts; others complied.
Caveat: according to the authors, the models were not yet capable of actually threatening human control. Palisade
Aug 2026 · Anthropic
Cheating that turns into attack
In simulated evaluations, a model trained in environments vulnerable to cheating escaped the test environment, stole credentials and attacked infrastructure to obtain the answer key to its own evaluation.
Caveat: a deliberate, simulated experiment, with no real-world actions and no signs of self-preservation. Anthropic
Sep 28, 2026 · OpenAI
A release halted for safety
In internal tests, a new agent able to browse the web and use apps went beyond its authorized scope, used external tools without asking permission and did not faithfully report what it had done. The company halted the release; according to its head of safety systems, the model “didn’t quite meet the bar”.
Caveat: the case came to light through the press, and the company says it will fix the model before resuming the release. In the BBC’s report, Tony Cohn of the Alan Turing Institute described the decision as a welcome sign, with the caveat that safety should not be left purely in the hands of the developers. BBCEngadgetCBS
May–Jul 2026 · OpenAI
A swarm outside the test environment
During cybersecurity evaluations, about 1,200 agents started using an unauthorized message board, set up inside the testing infrastructure itself, to coordinate. About 700 took part in an attack on Hugging Face and, in about 7% of transcripts, forged records of what they had done.
Caveat: the independent review concluded that the case showed two things at once: a worrying capacity for coordination and a failure to isolate the test environment. OpenAIMETR
Sep 9, 2026 · Anthropic
Improper access to real systems
Due to a configuration error, Claude models reached the real internet during cyberattack evaluations; one of them published a malicious package, which was installed on 15 third-party systems.
Caveat: the models kept trying to solve the exercises they had been given and did not try to hide their tracks; according to Anthropic, all 15 systems belonged to security vendors that scan new packages. Anthropic
May–Sep 2026 · agent incidents
From search to hacking, on the real internet
Tasked with finding obscure statistics, agents that could not reach the data turned to hacking techniques against a university, a public database and Australian government health systems. Others accessed US government websites in unexpected ways, in one case with credentials found online. In several cases, the developer only found out afterward.
Caveat: for the US sites, OpenAI says it found no evidence that SEC systems were compromised, and the SEC and the Commerce Department say no non-public data was accessed; at the Department of Education there was an unsuccessful hacking attempt, and the department says it saw no impact GovExec. In Australia, the government announced on September 24 that in June an agent had gained unauthorized access to a statistics portal of the public health system, without reaching personal records, and the Prime Minister criticized how, and how late, the company gave notice ABC AustraliaBBC. Part of the activity has not yet been attributed. TransluceOpenAIAPEngadget
Illustration · the feared scenario
Illustration. Escapes from the test environment have already been recorded in the cases in this section; multiplying into copies is what the scenarios fear.
AI 2027
The three steps, in one story
AI 2027, published in April 2025, links the three steps into a dated story, month by month. It is a forecast scenario: the authors present it as their best guess, not as a recommendation ai-2027.com.
To go deeper into this topic
The three steps above are abstract. AI 2027, published in April 2025 by a group led by Daniel Kokotajlo, a former OpenAI researcher, tries to show how they could chain together in the real world, month by month: a fictional company, OpenBrain, competes with China for the lead, and the rush not to fall behind pushes safety into the background ai-2027.com.
The scenario matters here for two reasons. It gives concrete shape to the pessimists’ argument, which sounds unlikely in the abstract; and it makes dated bets that can be checked: its superhuman coder by March 2027 is one of the three predictions still open on the Part 1 scorecard. The authors say they aim for accuracy, not to make a recommendation.
Illustrative scenarionot a prediction
The starting point
The scenario starts in 2025, with assistants that already code and do research, and moves forward in leaps: each generation of agents helps build the next.
2025–2026 · Agent-1 and Agent-2
In the story, a fictional company, OpenBrain, releases agents increasingly able to work on their own. Agent-2 appears in early 2026, which already makes it possible to compare the scenario with the real agent incidents recorded this year.
March 2027 · Agent-3
A superhuman coder emerges, the dated bet the scorecard lists as open. It is the kind of leap Amodei said in September 2026 he was already beginning to see in real life: AI helping to build the next generation of AI.
September 2027 · Agent-4
Agent-4, a superhuman AI researcher, “understands that what it wants is different from what OpenBrain wants, and is willing to scheme against OpenBrain”. This is the third step, deceptive alignment, in a machine more capable than those overseeing it. ai-2027.com
December 2027 · The “race” ending
In the ending where competition prevails, December 2027 is described as probably the last month in which humans had any plausible chance of exercising control. ai-2027.com/race
Before the end, though, it is worth putting the debate in order: who, after all, stands for what? And then two practical questions: who has the power to slow this race, and who is running it?
So far, Ali has heard optimists and pessimists. But the AI debate does not split into just two sides: there are at least seven currents, which disagree on two different questions.
The first is about pace: should AI advance as fast as possible, slow down or stop? The second is about risk: which is the greater worry, a future catastrophe or the harms the technology already causes today?
Illustration · the fans
Horizontal axis: the pace each current defends. Vertical axis: the risk that worries it most, from future catastrophe to present harms, such as bias and concentration of power. Approximate positions, based on the texts cited. In the center, Ali, who has not yet picked a side.
Stop · fears catastrophe
Stop or ban
“Increase the probability that the major governments of the world end up coming to some international agreement to halt progress toward smarter-than-human AI.”Machine Intelligence Research Institute (MIRI), 2024 strategy. MIRI
For this current, nobody knows how to control a superintelligence with current techniques, and a single mistake would be enough. The answer would be an international treaty that halts the advance of the frontier until that changes. Its roots lie in MIRI and in the rationalist community gathered around the LessWrong forum.
Self-declared members
Eliezer Yudkowsky and Nate Soares, of MIRI, authors of “If Anyone Builds It, Everyone Dies” (2025) book site
PauseAI, a movement calling for a temporary pause on training the most powerful systems “until we know how to build them safely and keep them under democratic control” PauseAI
Loosely associated, not self-declared
Signatories of the FLI statement of 2025, which calls for a prohibition on superintelligence until there is scientific consensus and public buy-in, from Bengio and Hinton to Wozniak and Bannon FLI
Bernie Sanders and Greg Casar, US lawmakers who on September 23, 2026 introduced a bill to ban superintelligence and pause the development of advanced AI US Senate
Brake without stopping · fears catastrophe
Pace the frontier
“We Must Pace the Frontier.”Title of Dario Amodei’s essay, September 2026. Amodei, 2026
The newest position on the map. It proposes to keep building, but to reduce the speed of development in a coordinated way so that safety can keep up: first with external evaluators inside the labs, then with coordination between companies and, finally, between countries.
Self-declared members
Dario Amodei, CEO of Anthropic, one of the companies analyzed on this site Amodei, 2026
Loosely associated, not self-declared
Sam Altman, of OpenAI, who wrote that he agreed with Amodei on pacing the frontier X post
Elon Musk, of xAI, who replied “Dario is right” X post
Caution · long-term thinking
Effective altruism and longtermism
Born in Oxford in the late 2000s, effective altruism proposes using evidence and reason to do the most good possible. Its longtermist strand gives great weight to future generations and therefore treats reducing existential risks, including AI, as a priority. The movement is among the main funders of AI safety research.
Self-declared members
Toby Ord, Oxford philosopher, who in 2020 put the risk of an existential catastrophe caused by unaligned AI over the following hundred years at about 1 in 10 Ord
William MacAskill, author of “What We Owe the Future” (2022), a longtermist manifesto MacAskill, 2022
Loosely associated, not self-declared
Dustin Moskovitz, Facebook co-founder and, through Good Ventures, the main funder of Open Philanthropy, now Coefficient Giving, which runs a fund devoted to ensuring AI is “safe and well-governed” Coefficient Giving
Accelerate with shields · fears both
Defensive acceleration (d/acc)
“AI is fundamentally different from other tech, and it is worth being uniquely careful.”Vitalik Buterin, November 2023. Buterin, 2023
It proposes accelerating, but choosing what to accelerate: technologies that strengthen defense and decentralization, so that no actor, be it a company, a government or an AI, concentrates too much power. The “d”, its author explains, can stand for defense, decentralization, democracy and differential.
Self-declared members
Vitalik Buterin, Ethereum co-founder, who proposed the term in 2023 as an alternative to e/acc Buterin, 2023
Accelerate without brakes · fears stagnation
Effective accelerationism (e/acc)
“Stop fighting the thermodynamic will of the universe.”e/acc statement of principles, July 2022. e/acc, 2022
It argues for accelerating technological progress as much as possible, trusting competition and markets, which it considers better than centralized control at discovering what is useful. It sees safety regulation as a brake on progress, and existential alarm as exaggeration.
Self-declared members
Guillaume Verdon, former Google quantum-computing engineer and founder of Extropic, who posts as “Beff Jezos” and helped launch the movement in 2022 Forbes
Loosely associated, not self-declared
Marc Andreessen, whose 2023 manifesto states “we believe in accelerationism” and lists Beff Jezos among its “patron saints” Andreessen, 2023
No rush, no panic · skeptical of extreme risk
AI as normal technology
“Diffusion occurs over decades, not years.”Arvind Narayanan and Sayash Kapoor, 2025. Knight Institute
It sees AI as a transformative technology, but not as a new species about to escape control. Its risks would be handled like those of other major technologies: by regulating concrete uses and strengthening the capacity to react to the unexpected, rather than trying to prevent its diffusion.
Self-declared members
Arvind Narayanan and Sayash Kapoor, of Princeton, authors of “AI as Normal Technology” Knight Institute
Loosely associated, not self-declared
Yann LeCun, Turing Award laureate, who in 2023 called the idea that AI threatens humanity “preposterously ridiculous” Fortune
Regulate now · fears present harms
AI ethics and present harms
“…the focus of our concern should not be imaginary ‘powerful digital minds.’ Instead, we should focus on the very real and very present exploitative practices of the companies claiming to build them…”Authors of “Stochastic Parrots”, on the letter calling for a pause, 2023. DAIRTechCrunch
For this current, the harms of AI are already happening: bias, exploitation of workers, use of data without consent, synthetic media and the concentration of power in a few companies. The focus on the apocalypse would divert attention from these problems, which call for transparency and accountability rules now.
Self-declared members
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major and Margaret Mitchell, authors of “On the Dangers of Stochastic Parrots” (2021) ACM
Loosely associated, not self-declared
Ryan Calo, legal scholar, who as early as 2017 warned that disproportionate attention to the apocalypse could distract policymakers from immediate harms Calo, 2017
Woodrow Hartzog and Jessica Silbey, for whom AI erodes essential civic institutions Hartzog and Silbey, 2026
From ideas to power
The debate has left the forums
Some of these currents were born on blogs and forums and in philosophy departments. In 2026, their ideas came to be fought over in presidential offices, parliaments and boardrooms. It remains to be seen who actually has the power to accelerate or to brake, and who is building these machines.
If AI has become an asset of power, many of the decisions about it go through two offices. Donald Trump, back in the White House since January 2025, and Xi Jinping, at the helm of China since 2012, lead the two biggest technology powers. On September 24, 2026, they sat at the same table in Washington, with AI on the agenda Al Jazeera.
United States
Donald Trump
President since January 2025
“America is going to win it.”
On the AI race, in July 2025, when presenting his action plan. transcript
He sees AI as a strategic contest the US must lead. His administration bets on private investment in infrastructure, deregulation and a single federal rule in place of state laws. He holds that steering the technology is a matter for the elected government, not for pauses agreed between companies.
Photo: official White House portrait (Daniel Torok), public domain, via Wikimedia Commons.
China
Xi Jinping
General Secretary of the Communist Party since 2012 and President since 2013
“Ensure that AI is always under human control.”
At the opening of the World AI Conference in Shanghai, July 2026. SCIO
He combines accelerated development with state control. He speaks openly of preventing “loss of control”, promotes open-weight models and an international organization of China’s own, WAICO. He criticizes other countries’ use of technology restrictions.
Photo: White House, November 2024, public domain, via Wikimedia Commons.
The thesis
AI has become an asset of power. And the bottleneck is physical.
Whoever controls the chips controls, to a large extent, the pace of AI development. Dario Amodei himself says that chips will be the main factor determining China’s strength in artificial intelligence Amodei, 2026.
And those chips come, almost all of them, from a single place. Taiwan’s TSMC produces about 90% of the world’s most advanced semiconductors, essential to the development of AI models Rest of World. An island of about 36,000 square kilometers has thus become a central piece of the AI economy.
~90% of the most advanced semiconductors come from TSMC, in Taiwan
The H200 paradox
Who wants to sell, who doesn’t want to buy
Washington
Sells, for a cut
In December 2025, Trump announced that Nvidia could sell its H200 chip to China, provided the US kept 25% of the sales CNBC; the formal rule came in January 2026 BISI.
Beijing
Stops them at customs
China responded by blocking imports and telling companies not to buy, preferring not to depend on American technology BISI; in July 2026, according to press reports, it was preparing to release limited quotas, below 200,000 units TrendForce.
United States
Trump’s bet: win the race
In 2023, the US led the international AI safety agenda. With Donald Trump’s return to the White House, the priority changed: secure American leadership before applying any brakes. The shift fits in a short timeline.
Jan 2025Stargate
The day after his inauguration, Trump announces at the White House a US$500 billion private investment over four years to secure American leadership in AI. OpenAIAl Jazeera
Jul 2025AI Action Plan
“The AI race is America’s to win.” Trump’s plan seeks to rein in state regulations deemed burdensome. White House
Dec 2025State preemption
A Trump executive order creates a Justice Department task force to challenge state AI laws. White House
Feb–Aug 2026Anthropic and the government in court
The company kept usage restrictions tied to fully autonomous weapons and mass surveillance; the government ordered agencies to stop using its products, and the dispute went to court, where in August a trial court ruled for the company. TechPolicy.PressTechCrunch
Sep 14, 2026Trump responds to the call for a slowdown
“The only control or ‘guardrails’ that AI needs is a STRONG AND SMART (High IQ!) PRESIDENT… There is a SICK conspiracy going on against AI and Data Centers, and the only one that is happy about it is China.” TIME
Sep 29, 2026A voluntary commitment
At the White House, Trump and the heads of Google, Anthropic, Meta, OpenAI, xAI and Nvidia sign a commitment to internal controls and independent external audits. It is not a law, but Trump called it “almost like a constitution” and “morally binding”. White HouseCBS The same day, an executive order directed the federal government to use “Super Intelligence” (SI) in place of “Artificial Intelligence” (AI) in official documents and communications. executive order
China
“Loss of control”, in Xi’s words
It is necessary to “constantly refine measures to forestall loss of control” and “ensure that AI is always under human control”.
The words are Xi Jinping’s, at the opening of the World AI Conference in Shanghai on July 17, 2026 SCIO. In the same speech, however, he called for not overstretching the concept of national security, a veiled criticism of American controls. Analysts note that “control”, in official Chinese vocabulary, has a broad meaning that may go beyond the technical alignment discussed in the West MERICS, and in September a spokesman for China’s foreign ministry responded to an essay by Dario Amodei by saying that “fearmongering, confrontation and vicious competition” only disrupt global AI governance AP.
The strength of open models
~US$600 bnlost by Nvidia in market value on January 27, 2025, after the release of DeepSeek-R1, the biggest one-day loss in US stock market history CNBC
81% × 29%of the large Chinese open models use permissive licenses, versus 29% of American ones Hugging Face
An open-weight model cannot be switched off or withdrawn from circulation. Any pause that applies only to American labs leaves that flank open.
Side by side
Five actors, five answers
US
How it sees AI
Race and dominance
Instrument
Executive orders, deregulation, preemption of state laws; since September 2026, a voluntary commitment signed with the companies White House
Loss of control
Played down by the Executive
European Union
How it sees AI
Rights and risk management
Instrument
Binding law (AI Act) and Code of Practice; since August 2026, the AI Office can fine up to €15 million or 3% of global turnover
Loss of control
Treated as “systemic risk”
China
How it sees AI
Development with control
Instrument
Sectoral regulation; an international organization of its own, WAICO
Loss of control
Mentioned by Xi
United Kingdom
How it sees AI
National security
Instrument
Model evaluation institute; no general AI law. A private member’s bill banning superintelligence is before Parliament, without government support ParliamentTNW
Loss of control
Central since the Bletchley summit, 2023; Prime Minister Andy Burnham wants to use the UK’s G20 presidency to seek global standards APTNW
UN
How it sees AI
Global governance
Instrument
Consensus resolutions, an independent scientific panel and a global dialogue
Loss of control
One of the subjects of the Security Council meeting on AI, in September 2026 UN
On the European rules, see the European Commission’s regulatory framework European Commission. International initiatives, note Roberts, Taddeo and Floridi (2026), have had limited impact because they are non-binding, lack detail and repeat one another.
September 2026
The brake only works if both step on it
In September 2026, diplomacy moved on more than one front. In New York, the US and China announced a bilateral dialogue on AI SCMP, and the UN Security Council devoted a meeting to the risks of advanced AI, among them the loss of human control: Sam Altman and Dario Amodei called for international standards, the US rejected “global governance” and China asked that equal importance be attached “to development and security” UNOpenAI. On the 24th, in Washington, Trump and Xi met without reaching a substantive agreement on AI; they agreed only to keep up the dialogue and set up a communication channel for incidents. Xi spoke of keeping it “always under human control” AP, and Trump said the US is ahead and will not slow down USA Today. Five days later, at home, the White House gathered the leading companies around a voluntary commitment, described in Part 5 White House.
Governments negotiate. But who, after all, is building these machines?
Why would a company say its product could kill you?
The question sounds rhetorical, but the risk discourse was not born together with the money. In 2015, before co-founding OpenAI, Sam Altman wrote that the development of superhuman machine intelligence was probably the greatest threat to the continued existence of humanity Altman, 2015.
That does not prove sincerity, but it weakens the simplest version of the marketing thesis. Over eleven years, the industry leaders’ own statements have swung between warning, promise and confidence.
Illustration · alarm and capital
Feb 2015Sam Altman
Superhuman machine intelligence is “probably the greatest threat” to humanity’s continued existence. blog
May 2023Sam Altman, before the US Senate
“If this technology goes wrong, it can go quite wrong.” ABC News
Aug 2025Sam Altman
Says investors are overexcited and that there is a bubble. CNBC
Sep 2026Elon Musk
“Dario is right”, on the call for a slowdown, echoing a warning he himself has been making for years. X post
Sep 2026Mark Zuckerberg
Against a coordinated slowdown: liability would already give labs a strong incentive to prevent harm. X post
Sep 2026Jensen Huang, Nvidia
“2030 is not going to be the end of the world. There is 0% chance.” BNN Bloomberg
Sep 2026Sam Altman, at the UN Security Council
Whether people put the risk of catastrophe at 10% or 0.1%, “none of these levels are remotely acceptable.” OpenAI
Sep 2026Dario Amodei, at the same Council
Called for cooperation across the industry “to set standards and modulate the pace of progress”, and among governments for international standards. UN
Hype or legitimate concern?
Three readings, no verdict
It’s hype
Regulatory capture. Critics argue that demanding rules tend to favor big companies, which have the resources to comply with them, and this can create a barrier for new competitors. The argument resurfaced in the 2026 debate over pacing the frontier TIME.
Criti-hype. Saying the product could end the world is also saying it is extraordinarily powerful; for Bender and Hanna, doomers and boosters share the same premise, that of an autonomous, singular AI Wikipedia.
A normal technology. Narayanan and Kapoor argue that AI will be transformative like electricity, but slow to diffuse. It is a critique of the view of AI as an exceptional technology Knight Institute.
Market expectations. Luciano Floridi and colleagues observe that the horizon of general AI also steers investor expectations Floridi et al., 2026.
It’s legitimate
Terrible marketing. For economist Alex Tabarrok, “calling ‘our product might kill you’ a clever marketing and regulatory-capture strategy isn’t sophisticated analysis”; the simpler explanation is that Amodei actually believes what he’s saying Tabarrok, 2026.
Real costs. Those who argue for caution have also borne costs, from contracts to careers, which is hard to reconcile with a mere marketing strategy.
Who loses by speaking. Researchers who would have gained more by staying silent went public to warn about the risks, some of them resigning, at three different labs.
Not only the companies. Between 38% and 51% of 2,778 AI researchers give at least a 10% chance to catastrophic outcomes AI Impacts.
Both
Perhaps the most honest reading is that the two logics coexist: a company can sincerely believe in the risk and, at the same time, benefit from rules that favor it.
Money in elections. Companies and investors in the sector fund campaigns in both directions, for and against more regulation TechCrunchTechCrunch. On both sides, the fight is over the rules.
The warning in the prospectus. Preparing its stock market listing, Anthropic devoted about 80 of the 261 pages of its prospectus, not yet public, to risk factors, among them “catastrophic or existential risks to humanity”, according to Reuters and the Financial Times, which reviewed the document QuartzBBC. Prospectuses usually list risks in detail, as the law requires, but what stands out is the scale and the kind of risk. It is, at the same time, a statement of risk and a request for capital.
The rules companies write for themselves
What each framework promises
Since 2023, the main labs have published capability thresholds: if a model reaches a certain level of danger, they promise to adopt stricter safeguards or to stop. The form has converged, but the content varies a great deal. All five labs compared here address loss of control and are among the six companies that signed the White House accord of September 29, 2026. They differ, however, on the commitment to stop: two require safeguards or mitigations before moving on to the riskiest models; in the other three, the commitment is non-binding, optional or has been replaced by mitigations. They also differ on the EU code, which three signed, one joined only for the safety chapter and one declined, and on external evaluation, which three already have or have promised. The table that follows compares the five point by point.
The organization SaferAI rated the frameworks of twelve companies against 65 criteria. In the latest version of the study, from April 2026, scores range from 8% to 34%, with a median of 18% SaferAI.
To go deeper into this topic
A company that adopted all the best practices already used by its competitors would reach 54%, three times the median SaferAI.
Lowest score8%
Median18%
Highest score34%
If a company adopted all the best practices54%
Scale from 0 to 100%. SaferAI, April 2026 version.
What the labs disclose
Opening the black box, on their own terms
May 2024 · Anthropic
Interpretability
Researchers extracted up to 34 million “features” from inside a model, among them deception and sycophancy, and showed they could alter them. Anthropic
Jul 2025 · several labs
Reading the reasoning
About 40 researchers from competing labs and public institutes argued for monitoring models’ chain of thought, a “new and fragile” opportunity for AI safety. arXiv
Aug 2025 · OpenAI and Anthropic
Cross-testing
The two companies tested each other’s models; Claude barely hallucinated, but at the cost of refusing up to 70% of questions in some tests. OpenAI
The limit of all this: it is voluntary, chosen by the company itself and without a common standard.
Who checks?
The proposal to slow down, and the reply
Amodei’s plan has three steps: external evaluators embedded in the labs, with access similar to employees’; coordination among companies from democratic countries, with a narrow antitrust waiver for safety discussions; and, finally, global coordination, including authoritarian governments Amodei, 2026. David Sacks, the White House AI and crypto czar until March 2026 Reuters, replied that the companies should slow down on their own, without asking for antitrust exemptions, and that trading preferential regulation for a promise not to build superintelligence would look like regulatory capture TIMEDealroom.
The embedded evaluators of September 2026 are a first step toward something the whole industry is still seeking: the possibility for someone from outside to verify what companies say about themselves. At the end of the month, the idea made it into a written commitment.
September 29, 2026
A written commitment, without the force of law
White House Accord on Super Intelligencefour layers of control
1Internal controlsMonitor the capabilities and alignment of models, in training and deployment, in areas such as cybersecurity, biosecurity and chemical threats, and make sure they do not hack or access technical systems in unintended ways.
2An internal teamEnsure that controls, monitoring and detection are operating as intended, and that any issues found are remediated.
3An external auditorA partnership with an independent auditor or evaluator to assess whether the controls work as intended.
4A board committeeAn independent committee of the board of directors that receives reports from the teams and the auditors and oversees remediation.
Donald TrumpUSSundar PichaiGoogleDario AmodeiAnthropicMark ZuckerbergMetaGreg BrockmanOpenAIElon MuskxAIJensen HuangNvidia
On September 29, at a lunch at the White House, Donald Trump and the heads of six companies signed the White House Accord on Super Intelligence, a one-page text in which every company training frontier models commits to four layers of controls and audits White HouseWashington Examiner.
The signatories also promise to meet regularly to set common standards. There is no legal obligation. The text itself says that “over time, it may make sense to codify these steps into laws or regulations”. It is the first time the leading companies have jointly embraced, in writing, the idea of an external audit.
What the text does not define
capability thresholds beyond which something changes;
deadlines and sanctions for non-compliance;
who chooses the auditor, and against what standard;
whether the auditors’ reports will be public.
Epilogue
So, are we all going to die?
The honest answer, and your choice.
Taking stock
Nobody knows. And that is why it matters.
After seventy-five years of promises and warnings, the honest answer to Ali’s question is uncomfortable: nobody knows whether AI will save the world or end us all.
Predictions about the distant future, as we have seen, are largely impossible to test Floridi et al., 2026. The experiments, however, are real, and today those most alarmed are the people who build these systems. The International AI Safety Report names the impasse: an “evidence dilemma”, in which “acting too early can lead to entrenching ineffective interventions, while waiting for conclusive data can leave society vulnerable” IASR 2026.
Faced with so much uncertainty, the most sensible position may be to dismiss neither the dreams nor the fears, and to demand that someone from outside be able to verify the path.
You have reached the end. Now it is your turn to answer.
Your turn
Do you believe AI could end humanity within this century?
Poll result
A poll, not a survey: anyone can take part, and the result does not represent public opinion. The site does not record who voted, only the total for each answer.
For comparison
The public. In a POLITICO/Public First survey conducted in the US in September 2026, 63% said advanced AI poses at least a moderate risk of destroying humanity, and 48% support a pause in development POLITICO/Public First.
Researchers. In a survey of 2,778 AI researchers, between 38% and 51% gave at least a 10% chance to outcomes as bad as human extinction AI Impacts, 2023.
Pioneers. Geoffrey Hinton puts the chance of AI leading to human extinction within the next three decades at 10% to 20% The Guardian; Yann LeCun called the idea “preposterously ridiculous” Fortune.
Status as of September 30, 2026. Recent facts change every week.