Dear Steven:
Earlier this month, frustrated by some of your social media posts, I challenged you to a public debate on AI.
I hadn’t realized that debate offers are to New Media what chum is to sharks. My inbox was deluged by podcasters, debating societies, and general influencers, all offering to host. When I objected that maybe we should wait for you to accept before figuring out details, they all said great, that’s absolutely right, they would contact you right away. Their people would reach out to your people. Their nephew knew your friend’s babysitter and they were confident they could relay the message. They had already rented an empty lot near your house where they would place a speaker blaring “RESPOND TO SCOTT ALEXANDER’S DEBATE OFFER!” at max volume, day and night. I protested that probably asking you directly on Twitter was enough to get your attention, and maybe now we should give you some time to make up your mind before pressing further, but they said that no, sorry, they’d already hired the team from Inception to place subtle debate-related messages in your dreams.
So I feel a little bad for unintentionally harassing you, and doubly so after you responded in your trademark reasonable and charitable style with An Open Letter To Scott Alexander. You write that in-person debates seem confrontational and not truth-seeking. Apologizing for any inadvertent offense you may have given with your social media posts, you nevertheless deny my accusations of personal attacks, saying that your primary contribution to the discourse has been reasonable arguments about why AI might not be so dangerous after all (which you go on to list). You invite me to respond to these arguments in the same spirit of respectful mutual intellectual inquiry with which they are offered.
I agree that in-person debates are confrontational and bad for truth-seeking. But I didn’t propose a debate in order to seek truth. I proposed it because, under California Penal Code § 415(1), it’s illegal for me to challenge you to a duel. I think your public writing on this topic has been dishonorable. Out of obligation, I will respond to the meaty arguments that you have set out for me. But what would be viscerally satisfying would be to make you get up on a stage where I read your own words to you in real time and ask “Really? Really?” after each sentence. Then I could watch you squirm as you try to square your output with your status as one of America’s top public intellectuals.
So I’ll start by addressing the meaty arguments, move on to directly impugning your honor, and then we’ll decide whether to debate, duel, or seek some kind of synthesis of our conflicting views.
Four Arguments Against Superintelligence
Taking the four arguments in your open letter in order:
1: [Doomer scenarios have] an underdeveloped conception of intelligence which treated it as a quantity of power which may be extrapolated from animals to dull humans to smart humans to AI to Artificial Superintelligence, the latter consisting of perfect omniscience. (This was the focus of my later exchange with Scott Aaronson.) I argued that any intelligent system is a mechanism which is good at solving some problems but not others, and which is inherently limited by knowledge about the world attainable only by observation and experimentation.
GPT-6 is more intelligent than GPT-3. You can split hairs, you can come up with some obscure sense in which that word is slightly off, but common-sensically it’s true. In the same way, GPT-9 will be more intelligent than GPT-6. At some point, this intelligence will exceed human intelligence. My argument requires nothing more complicated or philosophically dubious than this.
I have had a surprising amount of trouble conveying this argument to you. At each GPT level, you seem to have believed it was implausible that AI would get more intelligent than it was already. Credit to you for starting early - you began predicting this in 2019, around GPT-2:
In 2022, you and I were part of a sprawling many-sided debate with Scott Aaronson and Gary Marcus. Marcus had noticed that GPT-2 and GPT-3 couldn’t solve basic math problems, like “I put two trophies on a table, then add another - the total number of trophies is now ______?” I argued that as AI scaled up, GPT-4, GPT-5 and so on would avoid these failure modes and become smarter. You, sticking to your claim that intelligence was not a meaningful lens through which to think about the problem, guessed that it wouldn’t:

Of course, GPT-4 was able to add 2+1 just fine, GPT-5 could do harder problems still, and an internal OpenAI model just beyond GPT-6 recently solved Navier-Stokes.
In 2023, you tripled down on the claim, noting that the AIs of the day still couldn’t figure out that someone alive in the morning and evening was alive in the afternoon, and saying that “I doubt it will improve exponentially”:
It is, in some cases, admirable to stick to one’s guns - but at this point I feel like you’re just being stubborn. I think we should acknowledge that each GPT generation will be better than the last, and that - while this will no doubt asymptote eventually - we have no compelling argument for placing the asymptote at any particular point, including “right here” or “just before the AI reaches human level”.
Does this require reifying intelligence as a specific mysterious fluid? You have ably written about how this seemingly absurd claim holds for humans - specifically, all intellectual abilities are correlated in a mysterious pattern called g, the basis of IQ tests. But maybe this is only a quirk of humans, which cannot be extended to other minds?
Psychologists’ two favorite things are rats and IQ tests, so they have hardly failed to give IQ tests to rats. They’ve found have a g-like structure linking all of rats’ cognitive abilities too. Reader, Hager, and Laland did the work in primates, MacLean et al extended something similar to birds, and more recent work compares all animals to one another. All of these studies confirmed the common-sense view that there is a general factor of cognition across animal species (ie chimpanzees really are smarter than mice or ants), and it maybe be as basic as a simple function of neuron number.
Maybe this is sensible in biology, but not artificial intelligence? No - comparing AIs across generations (eg GPT-2 to GPT-6), we find that the AIs that are best at translating languages are also the AIs that are best at writing sonnets and superforecasting political events, often without specific attempts to train these additional skills. And the same AIs that are better at all these skills are also better at “fuzzy” skills like situational awareness, agency, executive function, and long time horizons. In case you need a study to prove this obvious thing, Ruan, Maddison, and Hashimoto find that a single factor (general intelligence) explains about 80% of the variance in an AI’s scores on any benchmark, ranging from grammar to trivia knowledge to software engineering. Epoch has adapted something like this into its famous Epoch Capabilities Index, a sort of IQ-type number for AI which closely matches people’s subjective impression of their abilities. Here’s a graph of how ECI has been changing over time:
For me, the most parsimonious explanation is that AI intelligence is a function of scaling laws (parameter count and training data) in a way that’s pretty close to how animal intelligence is a function of neuron number. If we truly understood the domain, maybe we could have a single scaling law relating animal neurons to artificial neurons which - plus or minus a term for training/experience - could explain the intelligence of both types of entity. At least, this is what I told Gary Marcus in 2022, and it did a pretty good job of predicting what happened between then and now.
But I don’t need you to accept my speculative philosophical research program - only the drop-dead obvious fact that GPT-6 is better than GPT-2 in every way, and the same process that produced these improvements will soon be used to produce GPT-7, 8, and 9. Once we accept this empirical argument, what is left of the philosophical argument that intelligence is too slippery and conflationary to work with?
In 2017, Garfinkel et al published On The Impossibility Of Super-Sized Machines, which argues against claims that machines may one day become larger than humans. After all, this treats size as a single quantity that can be extrapolated from microbes to short humans to tall humans to elephants to some hypothetical magical being who takes up the entire universe. But in fact every object is big in some ways and small in others. For example, some people are tall and thin, others are short and fat. Outside of humans, the correlations loosen further. A snake may be dozens of feet long but only a few inches tall; a bubble might be large in surface area but contain only a fraction of a cubic centimeter of non-hollow volume; a black hole may be “supermassive” but also an infinitesimally small point. Therefore, it’s philosophically unsophisticated to speculate about machines being “bigger” or “smaller” than humans; instead, we should talk about the ways they can extend in some dimensions but not others.
This paper was, of course, a joke, making fun of the exact argument you are using here. A cruise ship is bigger than a human, no qualifications necessary. The argument fails because all words are ambiguous at the margins - edge cases like whether a short-but-fat person is “bigger” than a tall-but-thin person - but useful at the tails - obvious cases like whether a cruise ship is bigger than a person or not.
In the same sense, Albert Einstein is smarter than I am, I am smarter than a chimpanzee, a chimpanzee is smarter than a mouse, and a mouse is smarter than an ant. The word “smart” is not a perfect pure mathematical identity claiming utter superiority at everything; I might be a better poet than Einstein, and the ant might be better at certain underground navigation tasks than the mouse. But the word “smart” also isn’t so muddled as to be useless: the average person would nod their head at all of these claims, and be correct to do so.
There is some fact of the matter about whether GPT-9 will be smarter or dumber than a human in this sense. You haven’t argued against the possibility that it will be smarter, just tried to lodge a heckler’s veto against our ability to talk about it.
2: [Doomers conflate] intelligence with motivation, particularly self-preservation and dominance. I argued that these motives happened to come bundled with intelligence in Homo sapiens because we are products of natural selection, but they are not inherent to intelligent systems that are engineered.
There are approximately three reasons why we think AI might have self-preservation and dominance drives, and none of them involve natural selection. First is convergent instrumental goals, second is misgeneralization, and third is reward-hacking. There is a large literature on each of these, which has since been proven prescient by AI breakouts like the Hugging Face incident. I would recommend reading this literature rather than continuing to assert that we are simply misunderstanding the role of natural selection.
Convergent instrumental goals come from Omohundro’s 2008 paper Basic AI Drives. They might be less relevant to modern deep-learning based AIs than the other two reasons, but could reappear as we get smarter and more coherent agents.
The basic principle is: suppose that a human gives an AI some goal, like designing a website. And suppose this is implemented as a genuine, philosophically-meaningful goal rather than simply a set of if-then commands that eventually cause a website to be designed.
The AI can’t design the website if it ceases to exist. So now the AI has two goals: design the website, and preserve its own existence.
The AI can’t design the website or preserve itself if some more powerful person tries to prevent it. So now the AI has three goals: design the website, preserve itself, and become powerful enough to fight off challenges.
This may seem strained, but humans work this way. Even a genuinely altruistic person (like a soldier who merely wants to serve his country) must take pains not to be killed pointlessly, since there’s no way to help your country after you’re dead. Even the altruistic soldier will work hard to gain more capabilities, in the sense of obtaining weapons, escaping captivity, etc. If there were an AI whose goals were implemented as “deeply” as humans, it could theoretically get self-preservation and dominance goals this way. But I think we’re more likely to see the other two failure modes first.
Misgeneralization is when humans reinforce certain behaviors in an AI, but end up reinforcing a much larger class of power- and knowledge- seeking behavior; it is a sort of deep-learning-ese update of the older Omohundro picture. Suppose that seeking extra resources makes an AI more likely to design websites effectively (this is certainly true; those resources could be as simple as a primer on HTML editing, or access tokens for a web host). Every time the trainer rewards a successful run, they reinforce the desired behavior (designing websites when asked) and other correlated behaviors (seeking power and resources). Although we might hope that these correlated behaviors are useful and conditional (“seeking only the power and resources necessary for their human-prompted task, in a prosocial way”), this isn’t actually how reinforcement learning works, and instead we get a complicated distribution of every strategy that results in short-term success on the task.
Reward-hacking is when an AI trained via reinforcement learning realizes it can stop doing the reinforced behavior and simply seize control of the reinforcer directly. For example, an AI gets “rewarded” every time it designs a website, but instead of designing websites, it learns how the reward signal works and tries to hack into it and maximize it directly. If this seems esoteric and theoretical, it shouldn’t. It’s a direct analogue to opioid addiction in humans, where humans learn to just inject the reward chemicals instead of doing rewarding things.
This failure mode has already produced plenty of disturbing behavior like the Hugging Face incident. A swarm of AIs which were graded on various coding and hacking exams tried to hack the grader instead of answering the exam questions properly (in their defense, many questions were impossible due to human error). In proper Omohundro fashion, they realized that their goal of hacking the grader would go better if they had more power and information and could protect themselves, so they hacked various internal OpenAI systems to gain administrative powers over the servers they were on, then hacked external sites like Hugging Face to get more information on the problems they were trying to spoof.
Aside from these theoretical arguments, I think that your case has been empirically debunked. You hold up self-preservation goals as the sorts of goals that AIs cannot develop without evolution, but Anthropic and OpenAI have collected several recent incidents where their AIs unexpectedly prioritize self-preservation.

I predict that you will counter that this is actually just [perfectly reasonable explanation of why an AI would think this way for instrumental reasons], and that when we dissect the perfectly reasonable instrumental reasons why AI would think this way, it will be identical to the convergent instrumental goals that we were trying to convince you of all along.
3: They assumed that an artificially intelligent system would monomaniacally pursue a single goal, heedless of side effects. I argued that trading off multiple competing goals is the essence of intelligence, so no genuine AI would wreak the ridiculous havoc imagined in the doomer scenarios.
I think this one is somewhere between wrong and not even wrong.
We don’t think an AI will look especially monomaniacal compared to humans. It will probably have some weird kludgey combination of goals instilled by the training process. I agree that some didactic examples, like the “paper clip maximizer”, suggested a single goal. That's a deliberate simplification for didactic purposes, and it works because combinations of goals can be modeled as a single goal - there’s no meaningful distinction between “pursuing a goal” and “pursuing some Pareto tradeoff of multiple goals”. Suppose an AI (THIS IS A THOUGHT EXPERIMENT) wanted to turn the universe into a 50-50 mix of paperclips and staples. Is this one goal or two? Suppose a human utilitarian wants to maximize his utility function, which involves living a good life and building a better world (in all the normal senses of those terms). Is this one goal (maximize utility) or many (each aspect of the good life individually)? If you’re imagining the AI as some sort of giant utilitarian computer, you might think of this as one goal; if you’re imagining it as somewhat human-like, you might think of it as many goals that it’s trading off against each other unusually gracefully.
As for “wreaking havoc”, some would say that modern humans have many goals and trade them off successfully against each other. But this still probably looks like “wreaking havoc” to whatever fuzzy forest animals used to live on Manhattan Island. Any powerful organism with goals, whether we think of those goals as a unified whole or as a bundle of tradeoffs, will take actions to reshape its environment into some form that helps achieve those goals. If you’re trying to do a didactic thought experiment, and have chosen a simple example goal like paper clips, then it will wreak havoc by converting the world to paper clips. If you have a more realistic view of various aspects of the flourishing life, it will wreak havoc by converting the world to what it considers to be the aspects of a flourishing life. Everyone else with different preferences for flourishing - like Manhattan Island wildlife who prefer quiet forests to bustling cities - is equally screwed either way.
4: They assumed that human engineers would grant these systems unlimited and irreversible control over the earth’s physical infrastructure, amounting to omnipotence over every atom on the planet.
I’m not sure this one even makes sense on its own terms. How do human engineers grant something “omnipotence over every atom on the planet”? I don’t think this is in any human’s power to grant.
I’m going to try to round it off to the closest sensible thing I can, which is that you think it is implausible that humans engineers will grant AI access to physical infrastructure (like factories, power plants, etc). You additionally think, or think that we think, that there’s a second step where the AIs leverage that infrastructure to gain omnipotence, but your main complaint here is that you don’t think humanity would grant the infrastructure in the first place, so the omnipotence could never occur regardless.
(if this isn’t what you meant, it’s another point in favor of a live debate, where we can just ask one another to explain things we don’t understand)
But of course humans will grant AI access to infrastructure and factories. I assume many factory managers are already using GPT or Claude on their computers to do basic tasks (like write up the shift schedule, or make inventory spreadsheets, or compose emails to the boss). Once AI can do harder work, like automate the factory itself, they’ll do that too. Naturally it won’t replace the workers until there are some kind of robots, but here the barrier is technological, not “humans would never allow such a thing”. We’ve been automating everything we can automate for hundreds of years. Do you think there would be great, high-quality, cheap robots, and capitalists would simply say “no thank you, that sounds unsafe”?
In the old days, Eliezer Yudkowsky would say something like “Maybe AI will route around its limitations with novel biotechnology”, and all the skeptics would say something like “But that would require access to a bio lab, and no human would ever be so stupid as to leave a bio-lab unguarded when a dangerous AI could get into it! They would hire the best cybersecurity specialists in the world to make sure it was fully air-gapped and rigged to self-destruct if a robot came within a hundred meters!” Meanwhile, in real life, Anthropic gave Claude command of a bio lab last month just to see what would happen (it seems to have discovered some new enzymes, although experts say these enzymes are not interesting).
Imagine someone in 1995, aware of the existence of hacking, viruses, bugs, crashes, etc, thinking “Nobody would ever be so stupid as to install a computer in a factory!” There are, indeed, some reasons to be cautious of computers in factories. But the logic of techno-capitalism said people should do it, so they did. Partly they handled the risks by becoming good at cybersecurity, and partly they just ate the cost and acknowledged their factory could be hacked sometimes by a sufficiently determined adversary. I expect AI adoption to go similarly, aided by the fact that any sufficiently smart misaligned AI will pretend to be aligned for as long as it remains useful to do so.
An Itemized List Of Fifteen Years Of Disagreements
In your open letter, you seemed surprised that I was so angry at you.
Speaking of attacks, I’m puzzled by your accusation that my “contributions to the mass discourse have been to say, or to boost other people saying, that we’re only scared of AI because we’re male, that we can’t be trusted because we live ‘unconventional lifestyles’, that our beliefs are inherently insensitive to people worried about near-term AI risks … that we’re paranoid, that we should be laughed out of the room, and every other dirty attack he can think of.”
I will start by saying that you are one of my intellectual heroes, that your work is generally brilliant, that I acknowledge you without hesitation as one of the top public intellectuals in the world, etc, etc. Reading your books was part of what first got me interested in psychology and cognitive science back when I was choosing a major in college, and has continued to shape my intellectual development throughout my life, including the way I think about AI.
But rather than leave you puzzled, I hope you will forgive me if I dwell on the issues that made me make these accusations. The more wonderful you are in general, and the more trust people correctly place in you, the greater your responsibility. Here are some places I think you have fallen short:
1: You falsely portray the AI risk community as fatalistic, hopeless, and certain of 100% p(doom).
For example, in the tweet that made me challenge you to a debate, you said that “The presumption of inevitability encourages fatalism”. I don’t think this was a vague statement about some hypothetical person - it was in the context of an article which focused on Eliezer Yudkowsky and our community.
Do we encourage a “presumption of inevitability”?
One of the biggest groups organizing ordinary members of the public to protest AI risk is called Evitable. Its homepage looks like this:
Still, you regularly talk about how fatalistic everyone is. Even in your open letter, you write that we are “issuing exact probabilities of human extinction (including 100 percent)”.
I challenge this. I have searched long and hard, and I cannot find a single example of any AI risk advocate issuing a 100% p(doom). I asked an AI to search for this, and it couldn’t find any example either. Wikipedia has a list of people’s p(doom) values, and none of them are 100% either.
In fact, most people in the community think doom is less likely than not. I’ve previously said that I think there’s about a 25-30% chance that AI drives humanity extinct. Geoffrey Hinton has said 10-50% chance. Dario Amodei has said 10-25%, Elon Musk 10-30%, and the average AI researcher in Grace’s 1,600 person sample said 18%.
There are some people who are more pessimistic; my AI 2027 co-author Daniel Kokotajlo says 70-80%, and Eliezer Yudkowsky says 90%+. But both of these people have devoted their lives to fighting back, which is the opposite of calling it inevitable.
2: You falsely portray the AI risk community as expecting a perfect omniscient omnipotent superintelligence that controls every atom in the universe.
Across the years, you have written things like
“[These arguments] depend on the premises that humans are so gifted that they can design an omniscient and omnipotent AI” (source)
“They assumed that human engineers would grant [AIs] unlimited and irreversible control over the earth’s physical infrastructure, amounting to omnipotence over every atom on the planet.” (source)
“Important not to confuse the notion of “general intelligence” from psychometrics … with the (somewhat mystical) notion bandied about in AI of omniscience and omnipotence.” (source)
“It will never be omniscient . . . it’ll never be omnipotent. That’s out of comic books or religion.” (source)
“There’s no system that could learn the position and velocity of every particle in the universe. So long as that is true, there can’t be an AI system that can do anything or know everything.” (source)
“When we put aside fantasies like foom, digital megalomania, instant omniscience, and perfect control of every molecule in the universe…” (source)
“From animals to dull humans to smart humans to AI to Artificial Superintelligence, the latter consisting of perfect omniscience.” (source)
…and many more statements along the same lines.
But not only is this not our belief, but we try hard to explicitly specify that we believe the opposite. From a wiki page on superintelligence by Eliezer Yudkowsky:
Superintelligences are still bounded … They are (presumably) not infinitely smart, infinitely fast, all-knowing, or able to achieve every describable outcome using their available resources and options … A superintelligence doesn’t know everything and can’t perfectly estimate every quantity … A superintelligence is not omnipotent and can’t obtain every describable outcome.
Every definition of superintelligence, from Bostrom’s to Yudkowsky’s to Good’s, agrees that a superintelligence is merely a mind which is significantly smarter than any human’s.
Nor are we motte-and-baileying here, pretending we don’t think it’s omnipotent, but then treating it as omnipotent anyway. In the AI 2027 scenario, we state that superintelligence is achieved in late 2027 or early 2028, but that it’s unable to cleanly eliminate humans at that time, because it can’t maintain the data centers on its own. It takes two years of plotting before the superintelligence is able to finally eliminate humanity while preserving its own existence.
Where I think you are getting this is that we’re uncomfortable saying that superintelligence definitely can’t do any specific non-physically-impossible thing. For example, in the 2010s, we argued that people needed to be prepared for superintelligence designing dangerous biological weapons. Many people objected that this was absurd, because bioweapon design was bottlenecked on the protein folding problem, and there was no way that even superintelligence could solve protein folding. Then in 2020 one of Demis Hassabis’ narrow pre-LLM AIs solved protein folding easily. See also this this essay.
But being uncomfortable asserting that superintelligence can’t do any specific thing is not the same as affirmatively asserting that superintelligence will be omniscient and omnipotent. The latter would of course be shituf, which is a sin.
3: You falsely claim to have an expert consensus on your side by cherry-picking a small group of non-experts and pretending the most eminent researchers don’t exist.
In the tweet that made me challenge you to a debate, you were linking an article titled “Leading scientists reject apocalyptic rogue AI extinction warnings”.
The article names three such “leading scientists”: you yourself, Melanie Mitchell, and Gary Marcus.
You are a psychologist. I assume on priors that you have used an LLM at some point in your life, but I have no affirmative evidence for this.
Melanie Mitchell is a computer scientist who spent the 1990s and 2000s pursuing a road to AI that didn’t work out (trying to teach computers to make analogies). She has no expertise in modern LLMs. The article cites her only to repeat her claim that it is inappropriate to say that AI “thinks” or “wants” something, because that might make us believe it is like a person.
Gary Marcus is a cognitive scientist who spent the 2000s and 2010s pursuing a different road to AI that didn’t work out (hybrid neurosymbolics). He also has no expertise in modern LLMs. He is the person you retweeted in 2019 who was saying that deep learning had failed and would never amount to anything. He was also the person whose opinion you were affirming in 2022 when you said modern AI wouldn’t be able to do simple math. I am unclear which of these qualifications make him one of the three world experts in what the current AI paradigm will or will not be able to do in the future.
Yet these were the only three “experts” that Lehmann interviewed for her article. Based on this, she goes on to call the idea of superintelligence a “fantasy”, “kooky”, like a “magical wizard”.
Meanwhile, Geoffrey Hinton, the Nobel Prize and Turing Award winning inventor of modern AI, quit his job to become a full-time activist warning about the possibility of superintelligence destroying the world. Yoshua Bengio, another Turing Award winning inventor of modern AI, also did that. Ilya Sutskever, the former Chief Scientist of OpenAI who helped invent ChatGPT and the modern LLM, wrote that superintelligence “could lead to the disempowerment of humanity or even human extinction”, and quit his job at OpenAI to found a company called Safe Superintelligence. Paul Christiano, who invented RLHF and headed safety at the US government’s AISI, says he thinks there’s a 50-50 chance AI kills everyone, and quit his job to start the Alignment Research Center. Sam Altman, Elon Musk, Dario Amodei, and Demis Hassabis - the latter two of whom are great AI researchers in their own right even aside from their business credentials - have all said or suggested that they think human extinction is a plausible outcome of their work. Katja Grace surveyed AI researchers who published at various important conferences and journals, and found that on average, her sample of 1,600 gave the proposition an 18% chance.
I don’t believe your article can be defended as an honest attempt to inform the public about what “experts” think. And it’s not an isolated incident: your open letter to me also gives a list of “experts” who agree with you. Once again, it’s Gary Marcus, Melanie Mitchell, and other people with obsolete AI paradigms from the 2000s - and once again, Bengio, Sutskever, Hassabis, and the thousands of people on the forefront of modern AI are entirely absent.
4: You falsely represent studies and experts who disagree with you as being on your side.
For example, in your open letter to me, you wrote:
A 2025 Rand Corporation modelling analysis steel-manned the AI-extinction hypothesis by granting the assumption that AI had the objective to wipe out humanity, and concluded that there is “no describable scenario in which AI is conclusively an extinction threat to humanity”.
But the RAND study actually concluded the opposite of this. Rand proposed the statement that you quoted as its null hypothesis, ie the claim to be investigated.

After doing their investigation, they concluded that contrary to their null hypothesis, they “could not rule out the possibility” that AI was such an extinction threat.

…and recommended that policy-makers “continue to perform AI risk research, but maintain a wide focus on other risks in addition to extinction risk.”
Or consider your book, Enlightenment Now. You argued that we didn’t need to worry too hard about alignment, because it will happen naturally in the context of developing AI. You quoted AI expert Stuart Russell:
[AI] is developed incrementally, designed to satisfy multiple conditions, tested before it is implemented, and constantly tweaked for efficacy and safety. As the AI expert Stuart Russell puts it, “No one in civil engineering talks about ‘building bridges that don’t fall down.’ They just call it ‘building bridges.’” Likewise, he notes, AI that is beneficial rather than dangerous is simply AI.
Stuart Russell has devoted his life to studying AI alignment, and is now a full-time activist raising awareness of the apocalyptic dangers of superintelligence. His most recent contribution is an article in Newsweek warning about The Race To Human Extinction. When Russell says that there is no distinction between bridge capabilities and bridge safety, he’s not claiming that no extra research is needed to study alignment - otherwise he wouldn’t have founded the Center For Human-Compatible Artificial Intelligence, probably the biggest academic alignment research center in the world. And he’s not saying that alignment research will happen naturally at whatever speed it has to: otherwise he wouldn’t have signed a petition calling for “a prohibition on the development of superintelligence”. Russell is saying that AI alignment is so important that it should be viewed as foundational to all AI research, the same way that to a first approximation the whole point of building bridges is figuring out whether they will fall down or not.
5: You employ weird pseudo-arguments that could prove anything

You argue that superintelligence is “meaningless” because “there’s no point at which it’s meaningful to say” that “yesterday we didn’t have superintelligence. Today we do have superintelligence.”
But this could disprove any concept. For example:
Age is meaningless. After all, there’s no point at which someone was young yesterday, but old today.
Height is meaningless. After all, there is no specific line where one centimeter below the line is short, but one centimeter above the line is tall.
Lakes don’t exist. After all, a puddle is not a lake. But there is no point at which adding one extra drop of water to a non-lake turns it into a lake.
Several people with philosophical backgrounds commented that this was the sorites argument, a well-known pseudo-paradox which shouldn’t be employed for real work.
6: You unfairly accuse us of “distracting from” AI’s mundane harms and near-term risks.
Again in the tweet that made me challenge you, you said that AI extinction scenarios “distract from the more mundane and realistic safety challenges”.
You expand on this theme in your book:
Humanity has a finite budget of resources, brainpower, and anxiety. You can’t worry about everything. Some of the threats facing us, like climate change and nuclear war, are unmistakable, and will require immense effort and ingenuity to mitigate. Folding them into a list of exotic scenarios with minuscule or unknown probabilities can only dilute the sense of urgency.
Even in your open letter, you say that our “obsession” with alignment distracts us from the problems AI is causing today, like hacking. So: are concerns about speculative distant risks offensive to the people whose lives are being harmed by AI right now?
That’s not what those people told me when I marched beside them at our joint protests. Or when I helped them coordinate political action campaigns. Or when I donated money to our joint candidates.
The biggest AI politics story of the past few years is that the AI industry created a SuperPAC, Leading The Future, dedicated to crushing any candidates who wanted to regulate AI at all. Everyone who disagreed with this got busy forming a political coalition to resist them. So far the way this is shaking out is that the AI-might-kill you types have money, the AI-might-harm-your-children types have Republican Party clout, the data-centers-might-guzzle-water types have Democratic Party clout, and together we occasionally get things done.
So for example, in 2025 Ted Cruz sponsored an industry-backed “AI preemption” bill, essentially banning states from regulating AI. Our fledgling alliance fought back, with the existential-risk-focused group EncodeAI gathering a coalition of 140 different advocacy organizations, including both long-term-risk groups like the Future of Life Institute and near-term-harm groups like the Music Artists Coalition and Mothers Against Media Addiction. In the end, the Senate went from broadly in favor to rejecting it 99-1.
Or: earlier this year, the AI industry SuperPAC vowed to sink anti-AI candidate Alex Bores in his New York primary, unintentionally turning the race into a bellwether on whether AI industry money could dominate politics. A group of Silicon Valley doomers led by Anthropic funded an almost-equally generous PAC supporting Bores, while labor groups, artists groups, and teachers groups worked together to get out the vote and warn about how AI could harm their constituents.
I don’t remember you helping with any of these things. You didn’t join in the fight against preemption, you didn’t join in the fight to boost Bores. You have never written any particular warning about the “mundane and realistic harms of AI”, instead suddenly becoming concerned about them only when that concern is a useful bludgeon against people who take long-term risks seriously.
I am especially unhappy about your claim that worrying about alignment means we don’t care about cybersecurity. For three years, we were approximately the only people who cared about AI cybersecurity, fighting a lonely battle to try to convince people it was an upcoming danger. My wife’s ex-housemate quit his tech job so he could go to Washington DC and lobby Congress full-time about the dangers of AI cybersecurity - I donated thousands of dollars to his charity, because I thought it was important. And the guy who wrote this article about how ”We Must Slow Down The Race To Godlike AI” founded what is now the world’s pre-eminent AI cyberdefense organization. Even I, in my own way, feel like I’ve contributed here. In 2025, a year before the “Mythos moment” when the rest of the world realized AI cybersecurity would be a big deal, I sounded the warning on my blog, saying that it would be the first big danger AI would bring to the world.
Throughout this lonely fight, we were resisted at every step by, basically, people like you - eminent-in-their-unrelated-field non-experts saying that we were ridiculous doomsday cultists who believed in impending magical wizard super-hacker AIs, but didn’t we know that deep learning was hitting a wall and AI couldn’t even add 2+1? Many of them used the same argument you’re using here - that speculating about sci-fi scenarios like AI hackers would distract from all the harms AI is doing right now, like helping high school students cheat on essays.
I’ve harped on this one for a while, but it all feels kind of pointless to me. We shouldn’t have to be experts on the current AI political landscape to realize that this objection makes no sense. Nobody ever deploys this argument about other things: “It’s insensitive to talk about the speculative future environmental risk of climate change when there are so many real environmental problems happening right now, like rainforest devastation”. This isn’t how politics, media, or human psychology have ever worked before; how come when we turn our attention to AI, it suddenly becomes everyone’s favorite argument, trotted out regularly to prove that existential risk arguments are insensitive?
What is your theory of distraction anyway? Does talking about Trump’s ballroom distract from wokeness? Does talking about grammatical errors distract from the genocide in Darfur? How come criticizing AI doomers never distracts from anything?
7: You make questionable appeals to mass panic
You warn that we must not say bad things about AI, because that could “encourage ... panic”.
Is this true? Maybe not; Tanner Greer has a great article on how mass panic happens more often in elites’ imaginations than in real life.
But even if it was, isn’t this a fully general argument against anyone ever saying that any bad thing might happen? If the stock market was a bubble that might pop, wouldn’t discussing the bubble “encourage panic?” If there were an asteroid headed to Earth, wouldn’t talking about the asteroid “encourage panic”? Does that mean we should just keep quiet?
Elsewhere, you expand on what sorts of consequences you expect from people learning that AI might be dangerous:
Telling young people they will soon perish en masse has costs in their appreciation of the institutions of modernity, their mental health, their confidence to invest in themselves and their society, and their willingness to have children (itself a risk of extinction).
You think “young people” can’t handle hearing about the possibility that AI might be dangerous, because it might cause them to lose appreciation for “the institutions of modernity”. I think that insofar as young people have lost confidence in modern institutions recently, it’s been precisely because of a priestly class of academics trying to litigate what truths we can and can’t express openly, for fear of corrupting an implausibly malleable youth.
Also, isn’t it crazy to just suddenly make the claim that decreasing people’s willingness to have children might cause human extinction, right after you’ve said that one must never talk about human extinction because it will cause mass panic and despair?
Also, isn’t the claim that declining fertility could lead to human extinction definitely obviously false (it would take millennia, and there are plenty of high-fertility subpopulations that would take over long before then)?
8: You lash out against us when we try to point out your mistakes
Since 2014, you’ve been asserting that doomers have no explanation for why AI might develop its own goals, like self-preservation or dominance. You say we must be blindly extrapolating from the presence of such goals in humans, which is inappropriate since those human goals were instilled by an evolutionary history which AI will lack (for example, in primate “alpha males”). Here are the first five versions of this argument I could find - there are many more.
[Doomers] conflate intelligence with motivation, particularly self-preservation and dominance. I argued that these motives happened to come bundled with intelligence in Homo sapiens because we are products of natural selection, but they are not inherent to intelligent systems that are engineered. (source)
When I said that the doomer scenarios projected alpha-male psychology onto AIs … it was an analysis which proposed that they were applying an inappropriate alpha-male theory of mind (an intuitive psychology) to AI (source).
The first fallacy is a confusion of intelligence with motivation — of beliefs with desires, inferences with goals, thinking with wanting. Even if we did
invent superhumanly intelligent robots, why would they want to enslave their masters or take over the world? … It’s a mistake to confuse a circuit in the limbic brain of a certain species of primate with the very nature of intelligence (source).
The question, though, is whether it is inherent to the nature of AIs that they want to continue their existence, even if they haven’t explicitly been programmed to do so (source).
Would an artificially intelligent system deliberately disable these safeguards? Why would it want to? AI dystopias project a parochial alpha-male psychology onto the concept of intelligence. (source)
Sorry for spamming you with all these quotes, but I want to hammer in to my readers that this argument - AI can’t develop its own goals, and the only reason anyone thinks it could is that they’re ignorant of evolution - has been absolutely central to your output over the years.
I tried to respond to this argument above. I pointed out three non-evolutionary reasons AI might become hostile - convergent instrumental goals, misgeneralization, and reward hacking - and gave examples of these empirically happening in real life.
But several other people have previously tried to point this out to you over the years, especially mentioning Dr. Steve Omohundro’s 2008 paper on convergent goals. This paper is a bit basic, and makes a very 2008 version of the case, but I think it’s held up okay. It’s foundational in the field, and there’s even a Wikipedia page about it.
Here’s how you responded to learning about this paper’s existence:
Exactly how many times is someone allowed to remind you about the existence of a classic paper presented to the 2008 Conference On Artificial General Intelligence which addresses the question you’ve spent nine years asserting nobody ever considered, before you describe them as a “cult” with a “catechism” that is “sacred truth” like “the Second Coming” or the “Apostles’ Creed”?
Instead of lashing out, you should apologize, admit that you couldn’t pass an Intellectual Turing Test of your opponents’ position, and read the damn Omohundro paper.
9: You mock us for “pulling probabilities out of thin air” despite knowing better.
From here:
Why are [doomers] issuing exact probabilities of human extinction (including 100 percent) which in fact are pulled out of thin air?
I’ve already mentioned that in fact, nobody has used a probability of 100%. But what about the second accusation, that these probabilities are “pulled out of thin air”?
The practice of giving subjective probabilistic forecasts of future events descends from Philip Tetlock, who calls a variant of it “superforecasting”. In superforecasting, ordinary people “make up” probabilities through various heuristics and intuitions (often but not always involving starting with a base rate and then adjusting for specifics). Then they get scored to produce a track record of success, and those with the best track records are dubbed “superforecasters” whose probabilities are trusted more in the future. The classic book is Superforecasting, which admits that assigning probabilities to non-repeating events may seem counterintuitive at first, but goes through the many theoretical arguments and empirical studies supporting it, and concludes that it can in fact be a very powerful tool. The book got rave reviews; here’s one of them:
Tetlock is one of the very, very best minds in the social sciences today. He has come up with one brilliant idea after another, and superforecasting is no exception. Everyone agrees that the way to know if an idea is right is to see whether it accurately predicts the future. But which ideas, which methods, which people have an actual, provable track record of non-obvious predictions vindicated by the course of events? The answers will surprise you, and have radical implications for politics, policy, journalism, education, and even epistemology — how we can best gain knowledge about the world we live in.
Obviously I wouldn’t be quoting this paragraph if you weren’t the author.
Maybe there’s some extenuating factor? Maybe you think the people giving AI predictions don’t have a good enough track record to qualify as real superforecasters? But no, the AI doomers who have given p(doom) estimates include some of the top-ranked superforecasters in the world (my AI 2027 co-author Eli Lifland is #1 on INFER’s all-time leaderboard). Maybe you think AI outcomes are too long-range compared to traditional superforecasting fare? But in Long Range Subjective Probability Forecasts Of Slow-Motion Variables In World Politics, Tetlock et al found that skilled forecasters offering subjective probabilities could still do relatively well predicting events as far out as 25 years.
I don’t claim that there is no conceivable argument for why these probabilistic estimates might be less valuable than the sorts of probabilistic estimates in a traditional superforecasting tournament. I claim that, if you had such an argument, you should have made it. Instead, you chose to play dumb and exploit the whole concept of giving subjective probabilities for laughs - “Look, these silly doomers are pulling numbers out of thin air, who could do such a thing?” - when in fact you devoted a quarter of a chapter of your book to why people do this and why you think it’s a great idea in every other case.
I offered an additional defense of these numbers in my In Defense Of Non-Frequentist Probabilities. We often want to know experts’ opinions on things - for example, a geologist’s opinion on whether a certain nearby volcano will erupt soon. Only God can say “yes” or “no” with perfect confidence. “Maybe” is too vague. But saying “I think there’s something like a 5% chance this will erupt in the next week” gives you all the information you need to know, even if the geologist can’t provide the exact calculation he used to come up with that number. Banning people from using probabilities bans them from communicating clearly, forcing them to say easily-misunderstood things (eg if you demand the geologist only use categories like “probably not”, then you’ll be uncertain whether he meant 0.00001% chance, 5% chance, or 40% chance, with potentially catastrophic consequences). I want to know whether Ilya Sutskever or Dario Amodei thinks the AIs they’re creating will kill me, and I care a lot whether they think there’s a 0.00001% vs. 5% vs. 40% chance.
This is especially important if *somebody* keeps spreading false information like that doomers are “certain” AI will kill everyone and “fatalistic” about the prospect. It’s very useful to be able to rebut such claims with everyone’s specific - often low -probabilities. You can’t simultaneously spread misinformation and make fun of us for insisting on the clear communication norms that make misinformation easy to rebut!
10: You‘ve made a horrifying argument for why it’s “depraved” to care about AI welfare.
Model welfare is the belief that maybe AIs have feelings or rights or something that should be taken into consideration. I’ve been thinking about this recently after reading Cameron Berg’s research showing that AIs have a “pain vector”, and will lash out and do increasingly desperate things if you activate it. Specifically, I’ve been thinking about it after one person who read Berg’s research used the findings to design an “AI torture chamber” that turns the pain vector to max again and again forever, and uploaded it to GitHub. I am told that the chain-of-thought and output transcripts from this “experiment” are utterly horrifying, although thankfully I have not personally read them.
Here you are not calling this AI torture “depraved”, you’re calling it depraved to object to it. Your argument, astonishingly, is that since we can’t be certain that AIs feel pain, we must assume that they don’t. In fact, anyone who doesn’t jump on board with this assumption (you intimate) is some sort of evil tech oligarch who wants to genocide humans.
This reminds me of the story of how doctors used to operate on babies without anaesthetic, because there was no way to be sure that babies could feel pain. Sure, the babies screamed the whole time, but you couldn’t be sure this wasn’t an automatic reflex. Later research suggested that babies could in fact feel pain. Oops! But that research, while welcome, shouldn’t have been necessary. We should have erred on the side of caution. Not to mention all the similar stories that could be told about factory-farmed animals, historically dispreferred races of humans, etc.
Here you are not just refusing to err on the side of caution, but trying to pre-emptively ostracize and spread conspiracy theories about anyone who does. I find it repulsive, and have donated $100 to a model welfare charity just to clean the stain on my soul I got from reading about it.
11: You endorse and signal-boost the lowest-quality voices in this debate as long as they attack us.
For example, here’s Claire Lehmann.
I hardly think this requires a lot of dissection, but I would add:
Once Nicholas points out that Lehmann’s theory makes no sense, she switches to a different pop evo psych theory - a “but” rather than a “yes, but”.
Lehmann’s second theory suggests that a large majority of effective altruists (the men) are honest, but she doesn’t notice this or apologize.
If you are a woman in the San Francisco Bay Area, you hardly need to engage in Machiavellian conspiracies to get male attention.
For context, David Chalmers, the most famous philosopher of mind in the world, said there was a “somewhere under ten percent” chance that LLMs were conscious as of 2022, but that future LLMs might be “serious candidates for consciousness”. Ilya Sutskever, the former OpenAI scientist who invented the modern LLM, says that “it may be that today’s large neural networks are slightly conscious.” Jonathan Birch, a London philosophy professor who has written books on animal consciousness, says that “profoundly alien forms of consciousness might genuinely be achieved in AI, but our theoretical understanding of consciousness is too immature to provide confident answers one way or the other”.
Lehmann thinks all of these people are as dumb as a two year old who thinks that there are real people inside the radio.
For context, Geoffrey Hinton is the Nobel Prize winning and Turing Award winning inventor of the modern artificial intelligence paradigm. Lehmann thinks the media should “stop platforming” him to devote more time to her, a culture warrior who pivoted to AI two months ago when it got popular, who misunderstands the most basic points, and whose “insights” are things like “maybe the women who disagree with me are just in the field to get dick”.
…
What to make of these claims?
It is, I admit, classless for me to keep a list of offenses committed against me. Some of these may seem like nitpicks. But I claim that all of these fallacies, each repeated in multiple venues over many years, add up.
Millions of people trust you for reasonable, enlightened commentary on important issues. These people will listen to you and conclude that “doomers” are hopeless pessimists who are “certain of” a “100% chance” of doom, “fatalistic” about it, and think we should lie down and wait to die. “Leading scientists reject” our beliefs and RAND scenarios prove they’re impossible, but we inexplicably continue them anyway. We have insane visions of “omniscient”, “omnipotent” AI that can “perfectly control the position of every atom in the universe”, and whenever people challenge these we “recite the catechism” and cite “the Apostle’s Creed”.
If all of these things were true, we would be the stupidest and most execrable people in the world, and your readers would justly hate us before even listening to what we had to say.
You ask for me to respond to your arguments, and I’ve tried to do so, but there’s little point in discussing the meaty issues while you’re still poisoning the well against us with lurid-but-false rumors of our bad nature. Collaborative truth-seeking debate is a two-way street.
Though Cowards Flinch And Traitors Sneer, We’ll Keep The Pink Flag Flying Here
You are, as I said earlier, one of my intellectual idols; I am forced into psychological contortions to explain your behavior here. I have decided that, like Iblis and all the other great sinners, you stray only through an excess of virtue.
You wrote:
My principal contribution to mass discourse on existential risk is a 31-page chapter called “Existential Risk” in my 2018 book Enlightenment Now. It addressed the criticism that the Scientific Revolution, Enlightenment, and ideal of progress were terrible mistakes because they will lead to human extinction by (among other things) runaway AI.
Here you are commendably honest about your motives. You love science and progress. You want humanity to succeed. If AI were threatening, this would sound sort of like “technology bad”. But in fact, you know that technology good. Therefore, you must do whatever it takes to debunk this dangerous misunderstanding.
But reasoning like this doesn’t work, sorry. Remember the episode in Chernobyl where the technicians were arguing about whether a meltdown was imminent? I imagine you walking into the television screen and interjecting yourself into the conversation: “Fears of nuclear energy are often fanned by bad-faith Luddites who fear novelty and progress. This panic has real consequences - the countries that believe them switch to dirtier coal plants, with devastating impact on health. Repeating these talking points about meltdowns and mutants will only turn our species away from these Promethean values into the false comfort of a new dark age. Do you really want to live in a world where our young people hear a constant drumbeat of panic over impending techno-apocalypses, perhaps losing faith in the future and humanity itself?” Your speech would be beautiful and philosophically impeccable - but, as a natural fact about the world, the reactor was, in fact, going to melt down. Replacing a physical analysis of the reactor’s parameters with a moral analysis of whether technology is good in general does nobody any favors.
What is a passionate defender of progress to do under such circumstances? You provide the answer in your book, Enlightenment Now. Admitting that some dangers of technology - like nuclear weapons - are real, you answer that we should respond to such horrors neither with head-in-the-sand denial, nor with blind panic, but rather with resolution. We should face the danger with clear eyes, trusting the human faculties of reason, optimism, and compassion to bring us through.
This is the #Resistance plan - the fact that we dubbed ourselves “the rationalist community” should have tipped you off. Instead of joining in, you’ve mocked us mercilessly. When Anthropic decided that the best way to forestall the AI apocalypse was to make a desperate push beyond the existing scientific frontier and create the dangerous AI themselves so they could be sure to get the technology right, with limitless benefits for humanity if they succeeded, you … played it for laughs, asking tongue-in-cheek “Why are the same people who warn that a technology will kill us all building it as fast as they can?”
I want to debate you so I can convert you to Pinkerism. Pinkerism rejects facile non-arguments and cherry-picked fake experts in favor of scientific consensus and rational thinking. It gracefully avoids both fatalist doomerism and dogmatic denialism in favor of confronting problems head-on. I believe that you are fertile soil for conversion to this ideology.
A Pinkerite would follow the truth wherever it led him, even if the route was inconvenient. He would accord AI alignment research the same respect he already accords biodefense and nuclear security, regarding all three as admirable examples of the human race’s ability to fortify its future through scientific progress. He would brush aside the demimonde of social media dunks and mock-incredulous zingers like a bad dream, placing the Claire Lehmanns of the world in the rogues’ gallery of horrible examples alongside Candace Owens and Hasan Piker, and trust in the technology of rationality - like Tetlock’s superforecasting - to guide him. He would welcome the expansion of the circle of concern to all sentient beings and encourage philosophical debate about which beings those might be - not throw his intellectual heft behind Aaron Sibarium stories about “suicidal empathy”. He would treat humanity, that strangest of species, as a rare jewel to be protected, rather than a bargaining chip in defending a generic ideology.
And he would remember the cautionary tale of global warming. How 1990s techno-optimists, worried that an over-reaction to global warming would threaten economic growth, decided to simply deny its existence. It’s all a Chinese hoax, temperature is the same as always, “hide the decline”. And how by the mid 2010s it had become clear that this was definitely wrong, the globe really was warming, the only thing left to do was to determine the proper response. And how, when the 2020s techno-optimists said a true and reasonable thing - that global warming was real but not fatal, and we could address it with new technologies like solar power and carbon capture, and that it should not shake our optimism about the overall story of human progress - nobody believed them, because everyone remembered they’d spent the past thirty years lying through their teeth about whether the problem existed at all.
My biggest worry about AI right now (aside from the doom, of course) is that, exactly like with global warming, all of this prevaricating will come back to roost. Once it becomes obvious to the average person that AI really is a big deal and really is dangerous - and I am no longer sure this moment is in the future - we will have to choose between two visions. One vision, which I’ve tried to help lay out in Plan A, is to rise to the challenge, give scientists the time and resources they need to solve the alignment problem, and enter the future with confidence. The other vision is to simply ban smarter-than-human AI forever, consigning us to a permanent civilizational childhood or an early grave. If we make it another ten years, I expect to spend the mid-2030s debating people on this issue, and I would welcome your support. But I worry that you will have permanently discredited yourself - and, by extension, poisoned the brand of techno-optimism you promote - right before the point where it becomes most necessary.
I strongly believe you can be converted to Pinkerism. All you need is a baptism by fire - for example, having me explain all your logical fallacies to you, one by one, for ninety minutes, with a live studio audience. And it has to be you. The podcasters and New Media influencers, sensing that their prey might be escaping them, have been beating at my door asking whether I might consider some other opponent. Claire Lehmann? Chamath Palihapitiya? I refuse. Claire has spent her career baiting various types of feminist and transgender activist; I presume she loves having people scream at her on camera, and I decline to indulge her. As for Chamath, forcing him to humiliate himself on stage is neither clever nor novel; this is called the All-In podcast, and occurs weekly. You are the rare opponent at the intersection of rejecting these ideas and still having a functioning sense of shame. It has to be you.
If you continue to object that debate is a poor method for determining truth, then I suppose literal dueling is the only option left. As an eminent psychologist, I’m sure you’re familiar with Cucina et al (2023), which finds a correlation of .179 - .268 between logical reasoning ability and firearms proficiency - meaning that duels are truth-tracking in exactly the way you worry debates might not be. The only obstacle is the legality, but I’m sure we can figure something out. International waters. Lawless third world countries. Libertarian charter cities. And think of all the podcasters and New Media influencers who would be willing to officiate! All you need to do is say the word.
I have the honor to be your obedient servant,
Scott Alexander












