788 Comments
User's avatar
Harjas Sandhu's avatar

Very much looking forward to reading Plan A.

> If we’re merely on track for a few cool gee-whiz AI innovations in the 2040s, then I’m wrong about everything and none of this really matters one way or the other.

Correct me if I'm wrong: the AI as Normal Technology guys disagree with you on this, right? I recall them arguing that the optimal preparations for superintelligence are quite different from the optimal preparations for AINT. Do you (or the AIFP) people have any thoughts on the potential downsides of being wrong / how to adapt if progress turns out slower than expected?

(I imagine a response along the lines of "it's probably better to mess up the AINT world than it would be to mess up the Plan A world," which I more or less buy, but I don't want to put words in your mouth.)

Alex Harris's avatar

I had the same reaction to that sentence. Slowing down progress would prevent or at least delay those cool gee-whiz AI innovations. Thus making the world worse for no reason. Meanwhile we face those other risks you mention, including declining fertility and productivity - and have lobotomized our best solution.

Bugmaster's avatar

Agreed, broadly speaking. The talk of "superintelligence" is sucking all the oxygen out of the room, but I'd still like to breathe a little...

Harjas Sandhu's avatar

> Thus making the world worse for no reason.

I mean, even if you're not worried about an AI takeover scenario, you still have gradual disempowerment and/or extreme power concentration to worry about. Normal competitive/market forces could very easily drive some pretty bad outcomes, even without superintelligence to worry about.

I would feel somewhat better about Plan A (or Plan S, for that matter) if they had spent a little more time worrying about how to deal with the world in which recursive self-improvement stalls or slows down. I'm sure the plans are probably adaptable to those worlds, but I would have appreciated an attempt to describe how the ideal Plan would change depending on the rates of observed progress. (I do, however, recognize that this is asking for the AIFP to do even more work, and I would absolutely accept an answer along the lines of "we thought about this but decided that it was a problem for another time.")

Ch Hi's avatar

The given presumptions (i.e. the US makes only good decisions) is so fantastic, that small problems don't matter. This is a work of fantasy. Hopefully a *useful* work of fantasy, but it's NOT going to come out that way.

AI2027 was a forecast, this is a wish-list. In US politics, only Nixon could have opened relations with China. Who do you imagine could put together this deal? It would have to be a right wing president with firm support from his side and respect from the left side...or at least a modicum of trust. Perhaps in another decade the polarization will be less, but we're talking about a deal put together within the next 3-4 years, ideally less.

Shawn Willden's avatar

What you don't trust our dealmaker-in-chief to broker a comprehensive and well thought-out deal with China?

(It's a good thing I'm typing. No way I could have gotten through that whole sentence with a straight face.)

JamesLeng's avatar

> I'm sure the plans are probably adaptable to those worlds,

I'm not. If concentration of power among human politicians is objectively a worse problem than advanced AI, trying to solve AI by concentrating political power *even harder* is like unclogging the kitchen sink by burning your house down.

The Ancient Geek's avatar

Declining fertility is not an existential threat on reasonable timescales. The world got by fine with half the population. The main problem problem would be economic, but AI presents worse economic threats.

Scott has a good article on the subject.

Bob Bobberson's avatar

I agree, although it depends on whether this new, half-size generation inherits the high fertility of its parents, or if it adopts the low fertility habits of mainstream culture, since in that case we have a quarter-size generation, and following that perhaps an eighth-size generation. A smaller and smaller population with a greater and greater percentage of seniors is not a comfortable situation to end up in, even if it doesn't qualify as a near term existential threat.

The Ancient Geek's avatar

I didn't say the halving would occur in a generation, and it's very unlikely. Also , fertility is self correcting -- if it pulls first world countries down to third world standards, the first world countries will see their fertility rebound.

Bob Bobberson's avatar

Yeah, my timescale on this is probably unrealistically fast, so I agree AI is a much more urgent threat. Still, even if it takes many generations to get there, third world living standards in the first world does not sound great to me.

The Ancient Geek's avatar

> third world living standards in the first world does not sound great to me.

That's not a certainty. AI means more productivity , fewer people means more of some resources to go round.

Randall Randall's avatar

Your argument that fertility will rebound if people get poor enough doesn't seem obviously correct to me. There are lots of fairly poor places that have TFR well under where the West was at that wealth level.

Ch Hi's avatar

Since we don't really KNOW the reason for declining fertility, we also don't know that it's self correcting. I think it probably is, and I feel that the world would do just fine with a 1930's level of population. But getting there might well be ... unpleasant.

Ferien's avatar

Antinatalism culture and electronic enjoyment. It's not some infection like Toxo which we can't find yet.

TGGP's avatar

The third world also has sharply dropping fertility.

Nutrition Capsule's avatar

Fertility doesn't necessarily self-correct. We really don't have much reliable data on how contemporary culture and technology influence fertility over long (100y+) time intervals.

Moreover, declining fertility comes with horsemen of its own. Declining innovation, economic issues following demographic transitions etc. can easily lead to cascading problems which aren't easily solved. Civilizations and nations can fall, and they have.

None of this is to say that declining fertility is necessarily _the_ pressing issue of our time, but there's little reason to assume it automatically isn't.

The Ancient Geek's avatar

Innovation was just fine in the mid twentieth century, with half the curren t population

Ferien's avatar

3rd world depends on 1st world for many of its things, agricultural machines, even seeds. If 1st world falls to 3rd world standards, it means decreasing harvests and there wouldn't be enough food to fuel population growth.

Arby's avatar

not sure about the self correcting part. just because quality of life declines doesn't mean you go back to patriarchy, no education for women, everyone living in countryside where you need the child labor, or that we deindustrialize to point where we lose ability to manufacture contraceptives (if we reach that point we've gone back to the stone ages and we'd have lost ability to produce enough food and 90% of population would've starved first)

Deiseach's avatar

The problem with declining fertility, as I understand it, is that we get lopsided population numbers: the old (and economically unproductive) are living longer and drawing more and more from the public purse for various services and pensions while the young for the workforce are not coming along in sufficient numbers to fund those goods and services.

Reduced population works only if the old die off faster and at comparable rates to the 1930s, leaving the smaller but younger population around.

Of course, if AI takes All The Jobs, then it doesn't matter if your country is full of 30 year olds as against full of 80 year olds, everyone is now poor and dependent on government dole.

Alex Harris's avatar

Yes! The big risk is the gerontocracy's Total Boomer Luxury Communism (plus other sources of Schumpeterian decline) will devour all hope for progress. I think it may be an existential threat to *civilization* even if not to *humanity*.

Domo Sapiens's avatar

This is not a risk, this is already reality in a few countries. It might have "just started", but the effects on economics and minds are real: The hope for progress is being devoured as we speak, while there is a gerontocratic last dance on the Titanic going on. It looks and feels pretty fucked from the inside.

Jeffrey Soreff's avatar

In a fully AI economy, everyone is now dependent on the public dole, but whether they are poor depends on the productivity of the AI economy and on the politics of the dole.

Shawn Willden's avatar

That's part of it. The other part is that (absent AI -- an important caveat!) a smaller population may not be able to support the diversified specialization needed to maintain our incredibly complex and interdependent civilization.

AI does seem likely to solve that problem, and the lopsidedness problem as well (though even human-level AI seems likely to cause other problems, absent an economic restructuring).

Arby's avatar

why poor? if AI takes all the jobs and reduces marginal cost of production to near-zero even if government didn't give a damn about its people it would cost them nothing to provide everyone with a plentiful (at least materially speaking) life

Deiseach's avatar

"it would cost them nothing"

There's always a cost. Governments can look at the money and decide "it would be better spent on infrastructure than a public dole" or "now we can pay off the national debt", for one.

Political costs are another. Reluctance to publicly sign up to "we will tax the rich wealth-creators in order to subsidise the newly unemployed who are valueless since they don't have wealth-producing skills for sale" because big companies don't like being seen as the golden goose constantly squeezed to produce more and more eggs.

The government of my country fought to *refuse* three billion in tax repayments from Apple because it was afraid that this would drive Apple away, and thus cost more than the momentary gain in revenue.

Jeff Bezos put thirty billion into Blue Origin. He could have taken that thirty billion and decided that he would provide for one thousand lower-middle class families for life to raise them/their children up into the middle to upper-middle class. He didn't, he preferred to put it into his toys (and hopefully a company that would then go on to be profit-making).

Just because you *have* a lot of spare money does not mean you are going to spend it on "provide everyone with a materially plentiful life". We expect the government to do so, but the government will be reliant on AI profitability from the big companies, and if they decide they don't want to play, what then? How do you compel a company to be profitable and hand over tax revenue?

TGGP's avatar

Even at most, it would be predicted to destroy our current culture/civilization, but not necessarily those who insulate themselves from it like the Amish or ultra-Orthodox. Humanity as a species would survive as long as some don't take the terribly wrong turn most of us have.

DrMcleod's avatar

That might be a problem if humanity only gets one chance at a high tech society that can colonise space (think fossil fuel depletion etc). If we screw this up then another 50000 years of (at best) agrarian society that can't protect itself from the comet isn't a great exchange.

JamesLeng's avatar

We've already built and deployed multiple terawatts of solar panels, which keep on producing electricity, with little or no maintenance, for somewhere between "decades" and "indefinitely." Even if all the documentation were somehow lost, that should be plenty of spares for some future minimum-viable-industrial-base society to dissect as part or reverse-engineering. Lots of scrap aluminum and stainless steel that won't need to be smelted from ore, too.

More likely, we don't get a multi-milennial dark age where all technical knowledge is lost, just old empires breaking apart, a few more things that were once rare becoming abundant commodities, smaller factions taking the type of warfare Ukrainians are pioneering and making it the new standard threat profile that national security is measured against.

DrMcleod's avatar

Solar panels are a very high tech commodity. Creating the silicon layers is by no means trivial in comparison to making a stream engine, and that is hard enough.

A1987dM's avatar

> The world got by fine with half the population.

The main issue is not with the total size of the population, it's with its age distribution.

The Ancient Geek's avatar

https://www.astralcodexten.com/p/slightly-against-underpopulation

Some bullet points:

* worldwide population probably won't shrink because African fertility is so high.

* Nowhere will shrink by as much as 50%, by 2100

* China, Korea and Japan will shrink much more than the west.

* The West won't shrink much it keeps up present rates of immigratiom

DrMcleod's avatar

The West won't survive in any recognisable form if it keeps up present rates of immigration.

The Ancient Geek's avatar

Yes, Scott notes that too...concerns about the level of population as a Trojan for concerns about the kind of people.

DrMcleod's avatar

In this case though, the Trojan Horse is the international asylum system.

Jeffrey Soreff's avatar

>and have lobotomized our best solution.

Yes, that is my largest concern as well.

Many of the concerns about RSI -> AGI -> ASI have been articulated for multiple decades. To make this more concrete: If we had put the brakes on a decade ago, we would _not_ have AlphaFold, and AlphaFold was a huge gain for biomedical research.

Xpym's avatar

As an LLM-as-a-normal-technology guy, I'm basically fine with this. The initial condition for this scenario is "multiple “Mythos moments” leave [the US govt] convinced that the situation is spiraling out of control" - I don't expect this to happen, but if it somehow does, then humanity can do much worse.

GlacierCow's avatar

> If we’re merely on track for a few cool gee-whiz AI innovations in the 2040s, then I’m wrong about everything and none of this really matters one way or the other.

this seems like the important bit here. Does plan A apply a ceiling to capability growth, apply a proportional term to the growth rate, or introduce a fixed drag term to growth?

If I believe the impact of AI will be large but not like, industrial-revolution-civilizational-transformation large, I would in theory be totally fine with a ceiling on growth (because I don't believe the growth rate will actually be that high) or a proportional term contingent on it actually being transformative (e.g. only kicking in once double digit gdp growth actually starting to happen, as opposed to in July 2026 when annualized GDP growth is 2.1%, unemployment is at 4.2%, and AI's main impact is as a productivity tool for knowledge workers).

As a superintelligence/exponential takeoff skeptic I will precommit to supporting some level of international alignment agreements if:

- annualized GDP growth (from AI) exceeds 10%. That is my threshold for AI-is-not-a-normal-technology from an economic growth POV.

- unemployment rate due to AI job displacement (and not from e.g. an AI bubble bursting and the resulting layoffs from that) exceeds 20%. That's my threshold for AI being genuinely disruptive to actual peoples' jobs rather than just being a productivity-enhancer.

- some real, provable security threats greatly exceeding that of the recent "Mythos stuff" (e.g. I'm not necessarily going to panic and support nuclear level regulations on GPUs just because a claude agent found some undiscovered firefox bugs). I don't have a hard line here but something like an AI model developing a bioweapon in a lab for example.

those who believe in a fast takeoff world should be fine with this because they believe at least one of these is obviously true.

Zach's avatar

> some real, provable security threats greatly exceeding that of the recent "Mythos stuff" (e.g. I'm not necessarily going to panic and support nuclear level regulations on GPUs just because a claude agent found some undiscovered firefox bugs). I don't have a hard line here but something like an AI model developing a bioweapon in a lab for example.

I'd suggest reflecting on where your line might be with respect to cyber risk, so that you notice when we blow past it.

Remember that Glasswing has a 90+45 day disclosure window, and that we're just past the 90-day mark.

GlacierCow's avatar

I'm less concerned with cyber risk for complicated reasons. My main threat model is like, north korea building an AI data center and directing a millino agents to brute force every known issue on every website simultaneously; and that category of threat already sort of exists (mostly with sweatshops of humans doing the same thing) and is the world we already live in today and we have countermeasures based on this sort of scattershot approach that allow us to muddle along. Even a magnitude difference in the quantity of hostile actors attacking systems doesn't seem like it would warrant a Nuclear Weapons Treaty of AI to stop its proliferation.

When it comes to novel cyber threats (AI figures out some new hack from base principles), it's equally true that AI can find novel countermeasures to non-public cyber threats. I don't see the calculus of attacker/defender advantage in cybersecurity changing drastically any time soon so unless adversaries can put an order of magnitude more compute power towards finding holes in security than the good guys can put into patching holes, I'm not sure this changes the security picture that much.

When it comes to rogue AI style cyber risks I basically think this is the realm of fiction, so any proof that this can actually happen in real ways (I am wholey unconvinced by the handful of examples that have appeared on this blog thus far, which seem squarely in the realm of creative roleplay or deliberately contrived entrapment) would be an easy red line for me. An AI agent, in the wild, disobeying its alignment, breaking out of its box, successfully replicating itself stealthily on some system it's not intended to run on, etc. would be a clear line for me. Of course, a world where such an AI could actually cause damage (e.g. one where it could direct its nanofactories to turn us all into paperclips) is one where I would already be on board with AI restrictions purely from capabilities.

JamesLeng's avatar

> this seems like the important bit here. Does plan A apply a ceiling to capability growth, apply a proportional term to the growth rate, or introduce a fixed drag term to growth?

The conflict-theory perspective would be to say that's not quite the right question. Whole premise of plan A seems to be putting someone - an individual or committee - in a position to adjust that metric at their own discretion, day to day or moment to moment. Who?

What exactly is stopping them from capriciously abusing the power to create artificial scarcity for rent-seeking, as imperfectly-aligned humans have long been known to do?

AnthonyCV's avatar

I think this comment mixes up outcomes and expectations.

The AINT folks (in most cases I've encountered) *expect* a world with much less and much slower change, not just than what AIFP folks expect, but than what AIFP folks are proposing to aim for.

Imagine person A is saying, "Specific change X will happen in 2045 and we should want it sooner, but change Y will never happen." Meanwhile person B is saying, "By default X would happen in 2033 but Y would kill us all in 2032." Then when B says "We need to slow down to prevent Y, even though that would push X back to 2036," and A says, "We need to speed up to get X sooner, like in 2040," there's still quite a lot of them talking past one another. Maybe also room for them to come to an agreement if they can stop talking past one another.

As I understand it, the downsides of slowing down in the AINT world are mostly... a few more years of suffering all the problems that society has already begrudgingly accepted as inescapable and priced in. Vs much higher extinction risk from not slowing down in the AIFP world.

Harjas Sandhu's avatar

From the original AINT article (https://knightcolumbia.org/content/ai-as-normal-technology):

> A natural inclination in policymaking is compromise. This is unlikely to work. Some interventions, such as improving transparency, are unconditionally helpful for risk mitigation, no compromise is needed (or rather, policymakers will have to balance the interests of the industry and external stakeholders, which is a mostly orthogonal dimension). Other interventions, such as nonproliferation, might help to contain a superintelligence but exacerbate the risks associated with normal technology by increasing market concentration. The reverse is also true: Interventions such as increasing resilience by fostering open-source AI will help to govern normal technology, but risk unleashing out-of-control superintelligence.

> The tension is inescapable. Defense against superintelligence requires humanity to unite against a common enemy, so to speak, concentrating power and exercising central control over AI technology. But we are more concerned about risks that arise from people using AI for their own ends, whether terrorism, or cyberwarfare, or undermining democracy, or simply—and most commonly—extractive capitalistic practices that magnify inequalities. Defending against this category of risk requires increasing resilience by preventing the concentration of power and resources (which often means making powerful AI more widely available).

Not 100% sure how much I agree with their analysis but they seem to believe it

AnthonyCV's avatar

Yes, that makes sense. I think this is very much in line with my prior comment.

JamesLeng's avatar

> As I understand it, the downsides of slowing down in the AINT world are mostly... a few more years of suffering all the problems that society has already begrudgingly accepted as inescapable and priced in.

That's assuming the slowdown mechanism itself has no terrible side effects.

If, say, most of the guys currently working for ICE were told "change of plans, you're supposed to go hunting for illegal *datacenters* now," handed a correspondingly more generous budget, plus joint-forces training exercises with whatever Chinese agency kicks down doors at four AM to drag people suspected of having contraband goods (or bad political opinions) off to secret prisons... do you have any confidence whatsoever that they'd be able, or even willing, to keep the false positive rate low?

AnthonyCV's avatar

That's true. I don't think it ends up being a dominating consideration, but I haven't thought about it enough to be really sure. There are always ways to do things badly that are easier than the ways to do things well.

JamesLeng's avatar

Maybe try drawing up the whole org chart, chains of command and areas of responsibility, then for each position, consider: If a con artist were trying to get hired for this job, what would be their big planned payoff? Who or what tries to stop them? For a prize worth more than the security's weakest link, somebody will try it. https://www.schlockmercenary.com/2012-02-26

Prize isn't necessarily cash, either. Catholic priesthood started running into trouble when their salaries weren't keeping up with the market. If it's not a credible career, soon the recruiting pool is all dried up into a sludgy residue of folks with various types of unhealthy interest in the side benefits. There's a balance to it: "Pay peanuts, get monkeys; pay billions, get supervillains."

John Schilling's avatar

If we replace every brown-skinned Spanish-speaking might-be-an-illegal-immigrant who is currently being dragged off to arbitrary detention without due process, with a nerdy bespectacled might-be-an-AI-researcher dragged off to arbitrary detention without due process, I'm not sure that's a net harm. At least to a society that has decided to pause AI development; it's obviously contrary to the interests of accelerationists.

And, the ICE agents hunting down illegal immigrants, by definition cannot accomplish their mission without detaining and/or deporting *people*. The DCE agents going after the illegal data centers, might sometimes detain people as well but they can accomplish their mission just by shutting down the data center and telling everyone to go home and look for a new job. Being told to go home and look for a new job, is I think quite a bit less harmful than being detained and deported.

So, as the OP says, already begrudgingly accepted as priced in. Worth trying to minimize the false positives and other collateral damage, yes, but not a dealbreaker for Plan A as a concept.

JamesLeng's avatar

Net harm isn't coming from a formal change in targets, it's coming from giving the Department of Demolishing Everyone's Civil Liberties a bigger budget, and de facto broader mandate to go after whoever they damn well please. https://www.cassiopeiaquinn.com/comic/now-you39re-a-man-pt-3

John Schilling's avatar

But in your hypothetical it wasn't a larger budget or a broader mandate, it was the same people with a changed mandate. Kind of a silly hypothetical, but your choice and not one that suggested a concern with expanded budgets and mandates.

JamesLeng's avatar

I literally said

> handed a correspondingly more generous budget,

Occam’s Machete's avatar

Well the good news is that the CCP ought to be amendable to central planning and redistribution.

Eric J.'s avatar

What about "unable to build a secret chip plant within the next few years?" That seems to be a big assumption that the entire plan relies on.

Occam’s Machete's avatar

Oh well that's a practical problem. Not all that different from the various missile and nuclear weapons treaties we've had with the Russians and such.

I was mostly just snarking about about the features of the plan as it relates to communism. Of course, I don't have any better ideas that would preserve the central tenets of liberalism in the face of powerful AI.

Ch Hi's avatar

This really isn't a plan. I makes too many implausible assumptions. It is, as described, a wishlist.

OK, Plan A isn't going to work, what's a plausible Plan B?

Melvin's avatar

If you can't join them, beat them. Cooperation with the Chinese regime is impossible, but the US still has a nuclear first strike advantage. The functional non-agrarian parts of China are concentrated in very few cities and can realistically be destroyed. Many will die on our side too, but that's better than everyone dying everywhere.

B is for belligerence. B is for blow them up. B is for Butlerian.

Ch Hi's avatar

Well, it's different. I can't think of anything else favorable to say about it.

magic9mushroom's avatar

Favourable or not, it's quite plausible Xi makes a play on Taiwan next year and it's the game board we'll be living in.

This is why I've been pushing harder on advocacy for civil defence than AI regulation here in Oz.

(I think the Chinese might go along - they're if anything more concerned about AI takeover than Trump is - but if China's a radioactive wasteland then the point's moot.)

Pjohn's avatar

"Mr. President, I'm not saying we wouldn't get our hair mussed, but I do say no more than 10 to 20 million killed, tops!"

Tp's avatar

It’s extremely different from missile and nuclear weapons treaties, almost the exact opposite!

Nuclear weapons work *as a threat* which means as a large country you NEVER want to hide your nuclear capabilities or they don’t do their job of deterring your enemy (you want to hide where some of them are so they can’t all be targeted, but you don’t keep them secret)

This “secret chip plant” has the exact opposite incentives. It’s supposed to be undetected and then used, not shown off to not be used. I think it would require a totally different approach

Occam’s Machete's avatar

You are definitely wrong about this on several fronts as it relates to the various treaties between the US and Russia.

Tp's avatar

Wrong about what? I didn’t posit a mechanism for the treaties so much as the concept that if you want to have deterrence, your opponent has to know they’re deterred

Occam’s Machete's avatar

For one, we don't actually know what "AI MAD" is or isn't yet. I don't think "deterrence" is the right emphasis for AI (vs. detection), though it was the major issue for nuclear weapons. AI stuff is much messier because e.g. cyberwarfare is not clear cut like conventional war, let alone nuclear war, is.

For two, there are different kinds of treaties we had with the Russians over different things, and there's not a uniform dynamic for the logic of secret capabilities or deterrence. Notably, the risk of either side getting the capacity to do an overwhelmingly successful first strike was part of the calculus. Most of what the SALTs and STARTs did was limit the volume of capacity of missiles/warheads, vs. type. The ABM from 1972-2002 did prohibit major missile defense systems so that MAD was preserved.

The incentive for either side to have developed some secret missile/warhead capability such that it had an advantage if things came to blows was still there (as either first or second strike).

That's very different than say how things went/are re: Israel, NK, Pakistan, India, Iran, Iraq, etc. Some of those countries definitely wanted to keep development secret until success, at which point they were proud to have the world know.

You could also think about biowarfare and how those treaties and attempts to evade them have gone.

The point I was making is that tracking GPUs and data centers is similar to tracking nukes and missiles and verifying treaty compliance in terms of feasibility. (Vs. say trying to track code/weights.) The analogy definitely doesn't extend too far beyond that because the technologies in question are fundamentally different. I know less about how we try to verify bioweapons treaties, and that's a lot harder.

The real problem for "track the chips/data centers" is if efficiency gains make it so smaller footprints can still make frontier progress.

Kenneth Almquist's avatar

Israel is a likely counterexample to your claim that you never want to hide your nuclear capabilities. Israel reportedly had a plan to detonate a nuclear weapon in the Sinai during the Six Day War to prove that they had nuclear weapons, but the war was over before they could put that into effect. The assumption that Israel has nuclear weapons is now a background assumption for most analysis of Israel. If Israel were to announce that it had nuclear weapons, I doubt that anyone’s strategic calculations would change significantly.

Pete's avatar

Well they have been unable to build a non-secret chip plant (for the level of chips that would matter - obviously they do have semiconductor manufacturing) for decades despite trying a lot, and are not on track to build one within the next few years even without any agreement or other restrictions. Doing that apparently is harder than building nuclear weapons.

Kveldred's avatar

That makes me very curious about what exactly it takes to construct such a plant: why's it so difficult? Is it the knowing what to build that's the hard part, or the actual building of it (even when you know exactly what, in theory, you need)? What enabled us (US) to do it---iterative improvement, some serendipitous discovery, a critical quantity of institutional expertise, particularly laser-tight build tolerances, ...?

Kenneth Almquist's avatar

Manufacturing high density digital logic circuits is very hard, in part because theoretical models of what your equipment is doing never perfectly match reality. When Intel ran into problems with their 10nm process and had to stick with their 14nm process for much longer than planned, they kept on tweeking their 14nm process and it kept getting better.

SMIC’s N+3 process is probably the best process node that doesn’t use any EUV (extreme ultraviolet) machines. China’s problem is that that makes it on par with the early EUV processes (1920-1921 time frame), or about five years behind the cutting edge. Huawei has filed a patent for an even better non-EUV process, although we don’t know whether they have or will be able to get it to actually work. My guess is that China won’t be able to match TSMC’s current offerings until it gets EUV.

Interest in EUV dates back to 1980, with serious research investment starting in 1990. ASML demonstrated a working but impractical prototype in 2006 (it took 23 hours to process a single wafer), and started shipping usable machines in 2018.

China has been working on EUV lithography for several years, with patents starting to appear in 2022. We would expect China to be able to develop EUV significantly faster than ASML because ASML has already blazed the way, but it’s clearly a hard problem. ASML didn’t develop the technology on its own. For example, the optics are manufactured by Zeiss. As far as I know, nobody outsize of Zeiss knows how Zeiss does this, and there’s not a lot of interest in learning how to make optics with the precision required for EUV lithography because you don’t need that level of precision for anything else.

So China will eventually develop its own EUV machines, but it’s not clear when. Expect at least a year between the time the first Chinese EUV machine is installed and the production of usable chips.

Performative Bafflement's avatar

> That makes me very curious about what exactly it takes to construct such a plant: why's it so difficult?

From a previous comment I had, which might shed some light on this:

Broadly, "making frontier chips" is a peak civilization endeavor that requires "peak" efforts not just from one company, but from many, which are spread across the world (ASML, Zeiss, TSMC, HBM companies like Samsung, etc).

This leads to the US already having a pretty large advantage, which it then leverages further via politics:

1. ASML has agreed to only sell DUV lithography machines to China (ie 2-3 generations old, much larger and less efficient, good for 7nm chips only), no EUV machines at all

2. The US has restrictions on NVIDIA selling anything except bandwidth limited or 1-2 generation back GPU's (recently relaxed by the incredibly dumb and bad H200 Trump deal)

3. The US restricts several technologies that allow the necessary data center bandwidth for large training runs. Broadly, this limits Infiniband and NVlink (limited in software via NVIDIA) capabilities, and also physical components like optical switches, DSP's, and more

4. The US is now limiting HBM and packaging dies (CoWos), which are necessary for China to build even internal Huawei 900's

5. The Remote Access Security Act is limiting China's ability to buy time in Singaporean / Malaysian data centers full of GPU's, and other countries.

Now is China trying to bypass all this? Furiously. They have invested many billions and at least a decade trying to build internal capabilities for both silicon, EUV lithography, memory and packaging, and more.

But they are still quite far behind. Their current capabilities sit around 7nm chips, which is 3-5 generations back. They are limited to zero HBM and packaging abilities, they're totally reliant on TSMC for this to do final assembly and packaging of Huawei 900's.

They have DUV lithography both from ASML and internally, and have trumpeted some EUV results but are likely quite far from production. All of this adds up to their chips being 3-4x less capable and efficient, and being maybe 10 years away from being able to build the current SOTA chips fully internally.

"So just use 4x the chips and build 4x the datacenters!" If China is known for anything, it's scale! Right?

Yes, but even this breaks down in subtle ways. Moving data across buildings has 30x the latency of moving it across racks inside the same data center. 900's have higher failure rates than NVIDIA chips, and that can bork your training runs and requires more cost and operational complexity to address. Because of the interconnect / bandwidth limitations and higher failure rates, you might need 250k 900's to do the work of 50k Blackwell chips, and that puts you across several buildings, and the end result is you use 5x the power and footprint, and take 10x longer to do a given training run, because now that higher failure rate is amplified by the greater number of chips necessary.

Not just that, but if you've ever wondered why TSMC is basically the ONLY frontier GPU company in the entire world, when obviously companies should be slavering at the bit to get into a market with literally bottomless demand and amazing margins, it's because everyone else in the entire world sucks compared to them.

The only companies even capable of getting close to 3nm frontier chips if you squint are Samsung and Intel. TSMC gets 70-80% yields. Samsung gets like 40% yields, and Intel <30%. They literally have to make 2-3x as many chips to get a viable chip. Huawei is basically Intel or worse - their yields suck. Making frontier chips is a "peak civilization" hard problem, and only TSMC is good enough to do it well, and this is why they're the only company in the world doing it.

So not only do you have this 5x - 10x drag on your AI output at the end of the chain, back at the very front of the chain, you also had to produce 5x as many chips just to arrive at 1 working one.

All of these things generally multiply rather than simply add, too - the added costs and complexities and inefficiencies amplify each other. You need more production and more power and have higher temperatures, and that impedes your performance frontiers, complicates your data center builds, uses still more power and resources, and so on.

As I'm reaching my conclusion here, I realize this got long, and I apologize.

But basically, the Chinese are already pretty nerfed by our existing measures, and as long as we can keep Trump away from the "sell frontier chips to China immediately" button, we're actually in a pretty good place for at least the next ~10 years.

Marian Kechlibar's avatar

I am very skeptical towards the idea that China and the US will agree on a common regulation plan for AI anytime soon.

First of all, the Trump admin will be there until early 2029 and they are not exactly known for sticking to their word in other international issues, quite to the contrary, every day brings some new pivot. You say that the argument would have to be trustless, but truly trustless self-executing protocols are really hard to construct in the messy real world. People are just too creative in circumventing the constraints; this is how Weimar Germany, limited in development of military aircraft by the Versailles Treaty, ended up with the most developed rocket science of the early 1930s.

Second, there is quite a mighty tech lobby in the US and I doubt they would agree. The US administration may be capable of bringing one or two corporations to heel, but all of them? And they would have a strong incentive to cooperate against this threat.

Third, the only somewhat comparable agreement, the Nuclear Test Ban Treaty, was only concluded when both the US and Soviets found incontrovertible evidence that radioactive isotopes were starting to accumulate in their kids, especially from dairy intake. (It is a very interesting story.) This made no sense even for the war hawks, no one wants to poison their own young. Even that was only a very partial agreement and the nuclear arms race as such was still on.

I think the only reason why such agreement would be even seriously considered would be something equivalent to finding loads of strontium-90 in kid's teeth, but in the AI context. Not just theoretical fears of what might happen if... but a hard visible problem right there in front of everyone.

Kyp's avatar

Say, like discovering that the next generation of AI could create a virus that would wipe out most of a country (where their kids presumably live)? Politicians, even the most cynical among them, are already spooked, and we haven't even gotten to the scary stuff yet.

Marian Kechlibar's avatar

Could is probably not enough. People are used to reading all sorts of catastrophic fiction and not taking it seriously.

A concrete epidemics probably would, though, even a veterinarian one.

Bugmaster's avatar

Theoretically speaking, you or I could potentially create such a virus today, in our bathtubs. It's unlikely, and it won't happen, but it's not impossible. So, should the government outlaw bathtubs, or what ?

Kyp's avatar

"Theoretically" and "imminently and easily" just aren't the same for the calculus here. Theoretically, someone could build a missile in their garage if they really wanted to, but we mostly don't worry about it. But if essentially anyone could do it easily, yeah, I expect that the laws would probably change in response to that. I don't know why it's controversial to say "politicians will react strongly to potential security risks" about AI when it's trivially true for many other domains.

Marian Kechlibar's avatar

"politicians will react strongly to potential security risks""

Putin tore away chunks from Ukraine in 2014. The full scale invasion happened in 2022.

In between (2014-2022, eight full years!) ... well, with some exceptions (the Baltics, Poland), the reaction of European politicians to that quite visible security risk was underwhelming. Even my country, which has a historic trauma of being occupied by the Warsaw Pact in 1968, was very lukewarm in reconstruction of our army from a truly abysmal state into one that can be described as merely miserable.

Politicians will only react strongly to potential security risks as long as the reaction is very cheap and does not face much counterwind. Otherwise they will happily ignore them and continue doing things that bring them popularity, which is mostly robbing some Peters to pay for Pauls.

Kyp's avatar

AI is overwhelmingly unpopular, this has been consistently supported by polling, and the percentage that is concerned is only increasing. When given the choice, those polled would rather ban AI outright that have it be unregulated. This isn't exactly an unpopular position, so they're likely to face more headwind if they _don't_ respond to the risks than if they do.

Marian Kechlibar's avatar

Unpopular does not necessarily translate into policy action, or at least not on short time window.

Some contemporary examples:

European voters have expressed supermajority sentiments against immigration from the Third World for decades, Brexit was partly caused by the same, and yet with a few exceptions (like Denmark), Western Europe still sees a massive influx. Even Meloni in Italy finds it really hard to square such sentiments against reality.

Billionaires are fairly unpopular in the US as well, but they seem to thrive regardless of whether D or R are in power - although I have to concede that the current shift of D towards far left candidates may be a harbinger of things to come.

Chat Control is very unpopular as well, and yet the EU just revived its 1.0 version by an extremely dirty procedural trick which meant that a minority of MEPs was enough to pass it.

Bugmaster's avatar

> But if essentially anyone could do it easily, yeah, I expect that the laws would probably change in response to that.

My point is, using e.g. Claude Code circa 2029 to create a super-virus is about as likely to happen as building an ICBM in your bathtub... actually scratch that, it's less likely, because at least we know that ICBMs are physically possible.

Kyp's avatar

I mean, it sounds like we just fundamentally disagree on what will be possible, so there's nothing really to debate here. But the labs themselves are very worried about biological attack risk, so it's not exactly outside of the Overton window. And whether you personally think it should be a concern, don't be surprised when the US government disagrees.

Bugmaster's avatar

To be clear, I am very much worried about biological risks ! We've got plenty of deadly diseases already, sitting in vials, just waiting for someone to make a careless move. I am merely contesting the possibility of an uber-virus that is e.g. 100x more virulent and 100x more deadly than any known disease. I don't think uber-viruses are biologically possible, so even a superintelligent AI would be unable to make one. Granted, it (or a human terrorist) could still cook up some kind of COVID 2.0, and we should definitely be on the lookout for that.

Hafizh Afkar Makmur's avatar

I think the fact that terrorists haven't done yet is a point against it. It's also sometimes I actually wondered about, to see actual chemical terrorism I have to go all the way back to Aim Shinrinkyo's sarin? Or there's that 2001 anthrax attacks but it proves the rule here. Seems like technology is not the bottleneck.

TGGP's avatar

Showers are more efficient :)

JamesLeng's avatar

The immediate solution to that - assuming a mere "could" gets taken seriously at all - isn't trying to stop all the AI you can reach, it's throwing *even more* AI brainpower at solving biosecurity problems preemptively.

Viruses have already undergone an enormous amount of iterative refinement, and the solution-space is finite. How many metabolically possible doom plagues do you really think there are? Soon as a vaccine gets open-sourced, anything that vaccine would stop isn't a credible existential threat anymore.

__browsing's avatar

> "Even that was only a very partial agreement and the nuclear arms race as such was still on."

The number of warheads in both the US and Russia declined substantially from its peak in the early 80s, IIRC?

I agree that the Trump administration could have done a lot more to be ahead of the curve on this one and Trump's... peculiar habit of keeping the other party guessing during any negotiation isn't really ideal for this purpose. But the incentives toward some kind of agreement are strong and I'm glad Scott is pointing out that there's really only a handful of facilities on the planet capable of making this hardware. The non-proliferation deal can and should happen, and if that's functionally equivalent to techno-oligarchy, so be it.

Marian Kechlibar's avatar

The Nuclear Test Ban Treaty was enacted in 1963 and I used it quite deliberately as an example of something that happened at the height of the Cold War, not something that stemmed from either various détente periods, or even post-Cold War.

__browsing's avatar

I don't understand how that's relevant. My point is that nuclear deproliferation did in, fact, prove to be technically and politically feasible, even if it took decades to start making serious progress.

https://ourworldindata.org/nuclear-weapons

Marian Kechlibar's avatar

There was no nuclear deproliferation. The current # of nuclear powers in the world is at its historical max (9), and the only country that willingly got rid of an existing domestic nuclear arsenal was South Africa. I'd be surprised if we didn't see at least one or two new nuclear powers emerge before 2050.

(Ukraine gave up possession of Soviet warheads too, but they didn't have control over them; the codes were still in Moscow.)

It is true that the total count of warheads in the world has shrunk, but it is still enough to kill the entire humanity at once.

Even the list of countries that are willing to host others' nuclear weapons on their territory has been growing lately - Belarus, Finland etc.

Anyway you asked about relevance. The relevance is in comparison of periods and state of international relations. At the height of mutual distrust and paranoia, countries tend not to conclude treaties that would limit them. This mostly happens during the détente period, when there is more goodwill.

China and the US are not in a détente period and I don't think that they will enter any détente period at least as long as their leaders don't change.

__browsing's avatar

> "There was no nuclear deproliferation.... ...the total count of warheads in the world has shrunk, but it is still enough to kill the entire humanity at once"

I'm not sure what you would consider a realistic deproliferation scenario to look like? Are you expecting a cordial handshake and then both sides decommission their entire arsenal overnight, or something?

> "At the height of mutual distrust and paranoia, countries tend not to conclude treaties that would limit them. This mostly happens during the détente period, when there is more goodwill"

That's more-or-less true by tautology, although arguing that "deproliferation never works" isn't likely to encourage a detente.

Marian Kechlibar's avatar

Something like Pakistan and India meeting and agreeing to be friends. OK, not very realistic, but weirder things happened. After all, France and England were mortal enemies for centuries.

It is not tautology. There are interesting exceptions, such as the Nuclear Test Ban Treaty. Another rather specific case is the decision of Germany in 1912 to abandon the dreadnought arms race one-sidedly, although this was offset by their focus on submarines. This was ultimately unsuccessful, but Bethmann-Hollweg was quite sincere in his attempt to reduce tensions with the future Allies.

Louis Dormegnie's avatar

> I am very skeptical towards the idea that China and the US will agree on a common regulation plan for AI anytime soon.

Mid-2029 gives us three years. I would say that the public consciousness of AI risk a year ago was multiples less intense than it is today, a trend which I have no reason to believe will end in the near future. Three years for both countries to realize the near future could harbor an extinction event sounds like enough time. Per Claude, it took 5 months from Hiroshima to the UN Atomic Energy Commission calling for "the elimination from national armaments of atomic weapons"; 10 months for a non-proliferation-type document to be drafted (Baruch Plan).

> People are just too creative in circumventing the constraints

Yes, and there might be some level of non-compliance, however Weimar Germany was able to build its arsenal back in some level of secrecy in a world without worldwide high-definition satellite imagery, blockchain/otherwise-IoT-type tracking devices, and bare-metal tracking. Also, while today's armament industry can run almost decentralized (eg Ukrainian drones), all XPUs capable of running/training/inferencing frontier AI models come from a small number of companies, and most of the wafers are manufactured and packaged in Taiwan.

>[..] the Nuclear Test Ban Treaty, was only concluded when both the US and Soviets found incontrovertible evidence that radioactive isotopes were starting to accumulate in their kids.

AI psychosis leading to teen suicides, AI deepfakes leading to widespread CP circulation are two salient topics that already exist and will only get worse with time.

Marian Kechlibar's avatar

It seems to me that by far the most common fear regarding AI today is "it will take our jobs".

Which, even if true, does not seem to be the sort of fear that results in geopolitical agreements of this magnitude.

Shawn Willden's avatar

The most common fear I hear is "AI data centers will use up our water".

TGGP's avatar

> AI psychosis leading to teen suicides, AI deepfakes leading to widespread CP circulation are two salient topics that already exist and will only get worse with time.

Those don't seem nearly enough for this supposedly existential issue. We haven't even established that the teen suicide rate has gone up!

JamesLeng's avatar

> all XPUs capable of running/training/inferencing frontier AI models

Any Turing machine can emulate any other Turing machine. There are already models which can run locally on a smartphone. If one type of hardware becomes prohibitively expensive, folks doing frontier research - and users willing to pay - will redirect their efforts accordingly. Creativity tends to thrive under such constraints.

Historical attempts to make nuclear power perfectly safe, based on a similarly catastrophizing risk model and deliberate slowdowns, are why we still burn coal.

Scott Alexander's avatar

I agree that this is most likely to work if there's a crisis that lights a fire under everyone (or perhaps in some future non-Trump administration). I think it's important to get it out there now so that it's the plan people reach for in a crisis when it's too late to be thinking up new plans.

I'm somewhat encouraged by the Trump administration throwing out its previous red lines and anti-regulation principles once it became clear that Mythos was the real deal. I think future AIs may be even more the real deal and even more likely to create surprising political opportunities.

I agree that trust will be hard, but I hope that the technicalities of AI (requires huge amounts of hard-to-produce chips, runs on hardware and software that can interface with clever verification schemes) makes this slightly easier than aerospace restrictions. See the Verification Plan at https://ai-2040.com/supplements/verification-plan for more.

Marian Kechlibar's avatar

We don't know enough about the Mythos/Fable thing in order to assess accurately what happened. Even on Hacker News, where I reside half of the time, it's almost all speculations and hearsay.

It is well possible that the core of the event was "heck, our systems are vulnerable, we must fix them ASAP, but then let us turn this monster against every foreign system in the world".

Bugmaster's avatar

I don't think that Mythos is the "real deal" in terms of being some kind of an unstoppable uberhacker. From what I've seen as a developer with many decades of experience, the reality is much darker. It's not the case that Mythos is superintelligent; rather, it's the case that our security is so weak that even a squirrel could hack it, because all major corporations and government agencies are essentially storing all of your private data in plain text on open servers. I am exaggerating of course, but only a little. Security is never a priority for anyone, especially not for those in management; it is seen as a waste of time, and government regulation can only force companies to implement the bare minimum of a fig leaf.

So yes, it is impressive that Mythos can hack things, but only to the extent that it's impressive that it can do anything at all. Walking through walls is easy when all the doors are made of tissue paper.

https://www.youtube.com/watch?v=nr5zYg4LG6s

NateEag's avatar

Well-put.

After my CS degree, I did not go into computer security, but I thought for decades that there had to be mountains of undiscovered flaws and security holes in all software, because the approach the average software project takes to security is so obviously broken.

I hate genAI, and what's it doing to the disciplines I love, but at least Mythos did demonstrate clearly that I was right about this.

See Quinn Norton's old-but-still-relevant "Everything Is Broken":

https://medium.com/message/everything-is-broken-81e5f33a24e1

Mark Y's avatar

This is true for a lot of things, but Mythos found a bug in OpenBSD, which is one of the few projects that Actually Tries to beat tissue paper. Even browsers are not pure tissue paper; they try to use better materials in a few key places where it actually matters, and Mythos got around it.

But to the extent that most things are tissue paper, we’re only being kept safe by the fact that even tissue paper takes some effort to tear if there’s enough layers of it. If now you can tell Mythos to do it instead of spending ten years learning about the weak points of tissue paper, security by obscurity gets a lot less effective in practice.

Bugmaster's avatar

I think that, given the number of botnets and ransomware operations floating around out there, it is highly unlikely that security through obscurity works to any major extent. Human hackers have been exploiting these bugs for ages, and will continue to do so. Mythos will allow them to do this a little faster, as any good tool does; but it's a quantitative improvement in their capabilities, not a qualitative one.

Mark Y's avatar

Obscurity does work, in the sense that at any given time, almost anything can be easily hacked, but most things have not yet in fact been hacked

Bugmaster's avatar

I think this isn't because obscurity works, exactly, but rather because most things aren't worth hacking -- because the next step in the pipeline (demanding ransom, laundering credit card money, exfiltrating trade secrets, whatever) is already saturated with data.

Kveldred's avatar

Aha---the meaning behind your name is at last revealed!

Kenneth Almquist's avatar

> Second, there is quite a mighty tech lobby in the US and I doubt they would agree. The US administration may be capable of bringing one or two corporations to heel, but all of them? And they would have a strong incentive to cooperate against this threat.

Yes. The Plan A document says that transparency would mean that:

> there’s no longer as much incentive to race to discover new AI paradigms and more powerful algorithms, because companies wouldn’t be able to hoard such discoveries and profit greatly from them.

The current AI company valuations are based on the assumption that the winner of the AI race will eventually make huge profits. Plan A or something similar risks crashing the stock prices of AI companies and, critically, destroying their ability to throw massive amounts of money into AI research.

__browsing's avatar

I'm skeptical that the "country of geniuses in a data center" are going to deliver policy recommendations for solving the world's problems that will actually be heeded by policy-makers or the general electorate. The solutions to most of the world's problems are obvious, the real obstacle is that people usually don't want to hear them. ("Black people are frequently poor and incarcerated? Oh, that's just because of low IQ, you can fix that with eugenics", says no-one in public ever.)

I'm not so sure about the degrowth scenario being so terrible, to be frank, given that the industrial revolution produced the most ideologically deranged ruling class in history. It's like you're relying on the techno-singularity to just keep enabling these terrible habits so the progressive left never has to face a reckoning with it's many, many deceptions. I wonder why.

Scott Alexander's avatar

I think choosing worldwide poverty and the permanent stagnation of human potential just to make racism seem a little more palatable to people is bad, actually.

__browsing's avatar

If per-capita income going up by a factor of roughly 100x since the start of the industrial revolution hasn't made "poverty" disappear, why would another 100x suddenly do the trick? If we wanted to end the material scarcities of the poorest billion people on earth we could do it for ~1% of the currently-existing global economy. The left would rather stuff another hundred billion dollars into college funds.

Also, please don't feed me this line about "racism". You know the fully skinny on HBD as well as anyone else does, the emails were leaked and you've written entire articles on the subject of how environmentalism doesn't fit the data, it's not like this is a secret. If intelligence hinges on developing accurate mental models of the world, then the first trustworthy AGI will by definition be "far right".

Xpym's avatar

>why would another 100x suddenly do the trick

We could stick the poors into the Matrix, at the very least.

TGGP's avatar

Isn't that what Mencius Moldbug advocated in one of his later Unqualified Reservations blog posts?

__browsing's avatar

If you mean Sam Altman Is Not A Blithering Idiot, I think he was just enumerating the possibilities. I think he was more in favour of tech-restrictions.

Xirdus's avatar

I would argue the increase of per-capita income since the start of industrial revolution did, in fact, make poverty disappear - at least in the western countries, at least by the poverty standards of 1800s.

__browsing's avatar

It's also turned the western world into a demographic suicide cult. Either curves extrapolate or they don't.

TGGP's avatar

Our fertility rate only went below replacement relatively recently. The UK demographically exploded early in the IR (in contrast to France, which started as the most populous country in western Europe but fell behind Germany).

__browsing's avatar

TFR has been trending down for a long time, and I feel like you're not really addressing why the most scientifically sophisticated, data-rich and economically-abundant civilisation in human history sat on it's hands and did nothing serious about this issue for roughly 40 years.

Shawn Willden's avatar

Per-capita income going up by 100X *did* make poverty disappear, mostly. When the industrial revolution started, 90% of the population lived in extreme poverty. The 100X increase flipped that from 90/10 to 10/90 -- and the trend of declining poverty is continuing.

I think there's every reason to believe that another 100X increase would eliminate poverty (as we would define it) entirely.

The Ancient Geek's avatar

Economics alone didn't do that, you need the political will to redistribute.

Shawn Willden's avatar

There is no significant redistribution outside of the rich world, and yet most of the elevation from extreme poverty is there, because most of the population is there. Redistribution is useful, but most of the improvement comes from just increasing GDP. A rising tide does lift all boats, though in economics they don't all get lifted equally.

The Ancient Geek's avatar

What's an example of a poor country that for richer without receiving foreign aid...because that's redistribution, too.

JamesLeng's avatar

I think you've got it backwards - seem to be implying the rich get richer first, then wealth thus created needs to be spread around as a separate step. Backwater goat farmers don't get smartphones as a matter of charity; the ability to check on different markets without physically visiting, secure electronic payments, and so on, makes them *better at their existing jobs,* by more than enough to pay for the marginal cost of manufacturing and network maintenance. Distributing things fairly makes the economy more efficient, which leads to more growth. http://tangent128.name/depot/toys/freefall/freefall-flytable.html#2412 That's presumably how we came to care about "fairness" in the first place.

netstack's avatar

It could be because reversed stupidity is not intelligence.

You have the privilege of complaining about your society because of the advantages of that society. Surely that counts for something?

If you’d rather not exist than chafe under the rule of elites, well, I suspect you’d be very unhappy in all of human history.

__browsing's avatar

I'm not certain what point you're making?

Wanda Tinasky's avatar

I think you're wrong about growth creating bad leadership. In my view, American leadership was at its best when growth was at it highest in the previous century. It's only a shrinking pie (or in our case, no longer expanding according to expectations) that caused our politics to switch from cooperate to defect. Universities only really became ideological insane asylums when elite oversupply became a problem, for example. Growth and abundance leads to optimism and cooperation. Degrowth leads to intense conflict.

I do agree with your thesis that this proposal is fairly naive. "The genius AIs will tell us how to make utopia" is suspiciously close to "the state will centralize control of resources and only the wisest citizens will make decisions." Policy is mostly about power negotiations, not wisdom. At best AI will only be able to solve the small part of the problem.

__browsing's avatar

I think economic slowdown can place additional stress on existing fractures in your society, but the brute fact is that the fractures in American society are a function of introduced racial and ethnic diversity, which accelerated vastly from the 1960s onward.

The relative disparities in outcome between these groups is what is tearing the country apart, and pointing out that black americans are both the wealthiest negroids on the planet and vastly richer than they were a century ago apparently does nothing to solve this. So long as human beings have any influence over the outcomes of their own lives, these disparities will persist in some form or other. We will only have human equality when human agency becomes irrelevant.

Wanda Tinasky's avatar

I actually think you're confusing cause and effect a little. The Civil Rights Act has been the principle lever with which collectivist elites have pried apart national cultural cohesiveness, but the fractures were already there. Marxism has been prevalent in intellectual circles since the 30's. There's an inbuilt tension when wealthy industrialists hold political and economic power that far exceeds the social prestige of academic intellectuals, particularly when the academics are responsible for training (and indoctrinating) the next generation of managerial-class elites.

Countering dishonest racial narratives on factual grounds is actually just punching at ghosts. These arguments function symbolically to demarcate political allegiances. Trying to prove that the US isn't racist is like arguing that the US flag is invalid because its 13 stripes no longer represent the true number of states. It's a category error that misunderstands the nature of the conflict.

__browsing's avatar

> "I actually think you're confusing cause and effect a little"

With respect, I don't think I am. I'd recommend looking at Amy Chua's World on Fire for examples of other countries driven to civil war/revolution by ethnic conflict, particularly with market-dominant minorities. It also seems a-priori unlikely that genetics, which is fundamentally more fixed and causative than social forces and instutions, is less to blame for these fractures.

Wanda Tinasky's avatar

Oh I agree that it's become the dominant factor now, but it didn't have to be. Without the concept of protected classes giving them disproportionate power and a dedicated elite orchestrating its weaponization then I think we could, in principle, have a stable multiethnic culture. Maybe I'm being overly optimistic but I think it was possible. At this point, though, yeah I agree that our demographics have doomed us. The US is going to slide down the slope of internecine political warfare for the rest of time. In 40 years we'll be a richer version of Brazil - isolated pockets of hyper wealth walled off from a completely dysfunctional civil society.

TGGP's avatar

I've read World on Fire. Non-hispanic whites still aren't a minority like her examples.

__browsing's avatar

They are rapidly sliding into that position, with the same predictable consequences. Hispanics and other minorities, even high-performing ones, trend toward voting Democrat, the B/W diffs are just the most obvious and least tractable.

The Ancient Geek's avatar

Collectiveness bad, cohesiveness good. Hmmm.

TGGP's avatar

The big politically divisive issues were between blacks & whites, and their relative fractions hadn't changed that much.

Josh W's avatar

You should just use the N Word. We know that's what you mean.

Wanda Tinasky's avatar

You should just say communism. We know that's what YOU mean.

Josh W's avatar

From your comment, I might as well have said "elves."

__browsing's avatar

I prefer it for consistency with "femoid".

The Ancient Geek's avatar

>The relative disparities in outcome between these groups

The gap between billionaires and everyone else is far larger.

__browsing's avatar

If you're presuming that economic redistribution or anti-trust laws can fix that problem, then great (although you're never going to have a GINI index of zero.) The racial disparities are functionally unsolvable (again, absent eugenics), which is why the left loves to bang on them so much.

Performative Bafflement's avatar

> So long as human beings have any influence over the outcomes of their own lives, these disparities will persist in some form or other. We will only have human equality when human agency becomes irrelevant.

OR when we get off our butts and allow gengineering to happen, after which "equality" can be real, and everyone's place in life can assuredly be a result of their decisions and agency, rather than the immutable hand of genetics (which substantially effects both the decisions people can discern, and the conscientiousness and motivation needed to stick to 'better' decisions).

__browsing's avatar

Unless your plan is for everyone to be genetic clones of eachother, genetic differences will still exert an influence on life outcomes (not to mention any decisions we make that aren't deterministically reducible to genetic predisposition.)

I can imagine both a bio-libertarian scenario where parents just purchase whatever genetic endowments they wish for their offspring and/or somatic modifications for themselves... or a coerced-parity society where all citizens are locked into some kind of RPG-style "point buy" system. But the former certainly won't produce equal outcomes, unless we hit some theoretical ceiling on genetic enhancement and don't discover any tradeoffs related to antagonistic pleiotropy (the former seems plausible, at least within the constraints of a recognisably human morphology, but the latter does not.) And as for the bio-RPG model, well... I'm sure the minmaxxers will still find some way to come out ahead.

Jeffrey Soreff's avatar

Mostly agreed. I'm less convinced that overproduction of elites is the explanation for universities becoming ideologically insane. Elephants in Rooms - Ken LaCorte has a nice video:

How did colleges get THAT left-wing? https://www.youtube.com/watch?v=pM83nNxilQE

Roughly speaking, he attributes it to a combination of positive feedback - somewhat left-wing places repel right-wing potential professors and attract left-wing potential professors plus active discrimination in hiring from the 1980s onwards plus the creation of grievance studies departments.

Wanda Tinasky's avatar

Yeah I agree with all of that but IMO that's a description of the failure mode, not the root cause. If higher ed had somehow been able to keep expanding like it did in the 50's and 60's then I think the insanity would have been irrelevant: the sane, productive people would have just outgrown the nutjobs. Healthy societies outgrow their stagnating parts. But when overall growth stops then the stagnant parts overwhelm the rest like a cancer. Cooperation loses to defection in the prisoner's dilemma unless the wider social context rewards the cooperators enough to render the defectors irrelevant. Like say in the 50's that the cabal of Marxist history professors might win at scrapping over the margins of fixed funding, but it didn't matter because all of the chemistry profs got $10m research grants from Dow chemical or something. In that context the people who can create get so much reward that the defectors are lost in the noise.

This is also my wider theory for civilizational decay, btw. Moloch is always there but healthy societies can outgrow him: Moloch eats in constant time but growth is exponential. Once that growth stops, however, everything gets consumed. That's pretty clearly where the US is now.

Jeffrey Soreff's avatar

Many Thanks!

>But when overall growth stops then the stagnant parts overwhelm the rest like a cancer. Cooperation loses to defection in the prisoner's dilemma unless the wider social context rewards the cooperators enough to render the defectors irrelevant.

Agreed! There is a lot that a healthy, growing, society can endure that doesn't do it permanent damage, that is far more damaging in a stagnant one.

But I do think that what happened to academia is largely orthogonal to this. At least one part of the woke cult damage, the establishment of grievance studies departments, was itself a _kind_ of (cancerous?) growth. The active discrimination against conservatives in faculty hiring looks to me approximately like an act of conquest.

__browsing's avatar

The welfare state has been growing more rapidly than overall GDP for I'd say about a century and a half, so I suspect that the "stagnant part" was going to eat the host alive regardless of general economic growth rates. (To be fair, I think welfare-state interventions did have some incremental social benefits prior to the 1970s, but that didn't really change the trendlines.)

In any case, I kinda fundamentally object to this whole "we must keep growing to maintain a cooperation equilibrium" framing. Yeah, if I'm otherwise a strong, healthy growing boy I can probably survive a tapeworm infection more easily, but I fundamentally do not want to host a tapeworm. It is repugnant to me that I need to live under this regime of obvious liars and parasites, and if some breakdown in the social contract is what it takes to "defect" from this "equilibrium", so be it. We can get back to growing afterwards.

Jeffrey Soreff's avatar

Many Thanks!

>In any case, I kinda fundamentally object to this whole "we must keep growing to maintain a cooperation equilibrium" framing.

If we were not in the process of extending AI applications and other technological innovations, I would also be queasy about 'needing' continual growth. Ponzi schemes must eventually collapse, and sooner or later, exponential growth must hit physical limits. But we _are_ in the process of extending AI applications, so 'degrowth' leaves potential standard of living gains 'on the table', which I see as just a pointless loss.

>The welfare state has been growing more rapidly than overall GDP for I'd say about a century and a half, so I suspect that the "stagnant part" was going to eat the host alive regardless of general economic growth rates.

It is a problem _IF_ production requires human labor, and people who _don't_ labor don't contribute to production. But

a) The whole context of this discussion is a transition to a potentially fully automated economy which doesn't require human labor at all.

b) There is a partial precedent from when 'labor' used to mean mostly physical manual labor. A large chunk (though by no means all) of that _has_ been replaced by motors and engines. A typical office worker is not breaking their back on the job. To a substantial extent, even cognitive labor which is sufficiently routine _has_ been quietly replaced by conventional software.

"eat the host alive" is more alarmist than is warranted, even just given conventional technological progress, let alone AI progress.

The Ancient Geek's avatar

Maybe it's time for a reminder that the churches are also insane, but in the opposite direction.

Jeffrey Soreff's avatar

Many Thanks! My impression is that churches vary immensely, all the way from the far left to the far right. ( As someone who finds deities unsupported by evidence, there is a whole orthogonal direction of insanity involved as well... )

The Ancient Geek's avatar

Of course churches vary: So does academia. Engineering and economics departments aren't gives of wokeness.

Jeffrey Soreff's avatar

Many Thanks!

>Engineering and economics departments aren't gives of wokeness.

( nit: "gives"? typo? maybe "hives" ? )

AFAIK, business schools are barely 1:1 left:right, and everything else is further left (most severely in grievance studies departments, but it has even affected engineering to a substantial, though less overwhelming, extent).

( Quoth Google/Gemini:

>In typical US colleges, engineering departments lean left, with liberal or Democratic faculty outnumbering conservatives by roughly 4-to-1.

)

There is a nice presentation, both on the current status and the last half-century of history, in https://www.youtube.com/watch?v=pM83nNxilQE&t=649s "How did colleges get THAT left-wing?" Elephants in Rooms - Ken LaCorte

Mark's avatar

One of the things that happens when one thinks seriously about superintelligence is the realization that racial differences are insignificant and we humans are all pretty much the same. Is it really the case that Race 1 has average IQ 10 points lower than Race 2? Well, if so, that's pretty insignificant if we end up sharing the world with AI that exceeds us by 500 IQ points. Is it really the case that unattractive traits like stupidity, laziness, propensity to violence are more common in Race 1 than Race 2? Those traits exist in every population and every individual, and they will all probably be far greater in any human or any human race than they are in leading edge AI.

__browsing's avatar

It is very possible that AGI will outstrip human capacities by the vast margins you describe, but how many people want to live in a world where their decision-making doesn't matter?

TGGP's avatar

Slaves typically choose to stay living under slavery rather than die. However bad life can be, it still almost always beats the alternative.

Deiseach's avatar

"Black people are frequently poor and incarcerated? Oh, that's just because of low IQ, you can fix that with eugenics", says no-one in public ever."

Because it ain't that simple. Low IQ but law-abiding and hard-working is better than high IQ and criminal (I'm sure someone will come along to tell me that no, in fact high IQ even if criminal is better for society because more productive, more inventive, more wealth-creation etc.)

Low IQ and a culture of "it's normal for teenagers to run around with knives stabbing each other" is the worst of everything.

__browsing's avatar

People don't reject the eugenic solution because they think poverty/crime are multifactorial problems not solely driven by IQ (which is true, but not really relevant given that blacks aren't especially law-abiding or diligent compared to whites.) I'm also not arguing that cultural factors have zero effect.

People reject the eugenic solution because "eugenics bad" is a thought-terminating cliche installed by the post-war managerial regime, as is "HBD = Nazis". They would continue to oppose it even if you literally invented a ten-dollar gene-editing pill that would add 20 points to IQ with zero negative side-effects. The thoughtcrime threatens the prevailing power-justifying regime narrative, and therefore cannot be permitted.

Throw Fence 🔶's avatar

It’s wild that you’ve been personally able to run high yield RCTs to tease out this causation from the correlation. Would you mind doing nutrition next? That’s my personal hobby horse of annoyance.

__browsing's avatar

Charles Murray was publishing college-board data decades ago indicating that blacks from the upper end of the SES spectrum were getting SATs similar to whites growing up in trailer parks. That alone rules out the bulk of possible environmental causes, including food deserts.

Throw Fence 🔶's avatar

I feel like there’s something like a missing mood situation going on with you here.

__browsing's avatar

Can you just say what you mean, then? What do you want to talk about, exactly?

Wanda Tinasky's avatar

I mean he's right. Look into the data.

"What is the current scientific consensus on how much adult IQ variance is due to shared environment?"

ChatGPT: The mainstream behavioral-genetics answer is: for adult IQ in developed countries, shared environment is usually estimated as small, often statistically indistinguishable from zero, roughly 0–10% of variance. A fair central estimate is probably ~0–5% in middle-class adult samples, with ~5–10% as a cautious upper-ish practical range rather than a crisp consensus number.

nominative indecisiveness's avatar

"People" and "they" are doing a lot of heavy lifting here.

Somewhere between 0% and 100% of eugenics-dislikers are going to take that pill, so what's the actual prediction? Will 30% of the population take the magic pill? 90%?

Are you saying that literally nobody except you and a few internet rightists would take a pill that raises your IQ more than a standard deviation? Parents of 65 IQ children won't try to guarantee their kids a more independent life? Teens in university won't decide that, yes, they do want to pass calc after all?

__browsing's avatar

I'm saying that if the american progressive left as a coalition actually cared about solving social problems, they'd be pouring hundreds of billions of dollars into magic-pill-creating research (or whatever the nearest technologically near-future equivalent might be, like embryo selection or germline gene-editing or even just something like Project Prevention.) They don't, because that would require explaining the actual primary cause of social inequalities to their voters (in part- the other part being that they often share the same thought-terminating beliefs.)

Bugmaster's avatar

Or it could be the fact that, socially and historically speaking, implementing eugenics would require the imposition of a strict totalitarian regime, which would create more problems than it solves. North Korea is probably the best example we've got of such a regime actually *working* to any extent, and AFAIK even they do not overtly practice eugenics...

__browsing's avatar

I might have been more sympathetic to that argument in the 1950s, after the Nazis had thoroughly pantsed themselves and when there was a lot less research into the extent of genetic influence on life outcomes and less ability to tinker with the genome itself.

But nowadays? I mean... cloning technology alone could in principle give us as many olympian athletes and genius polymaths as we like, without even touching on IVF, embryo selection or direct germline editing. Next-gen reproductive biotech could in principle give us all the upsides of classical eugenics with none of the drawbacks and be about 10x faster- the only reason why we're not pouring trillions in public money into assisted reprotech is because our ruling classes have all built their power on denying that genes (or even reproduction!) matters in the first place.

Even if you set that aside, the brute fact remains that over sufficiently long periods, you either have purifying selection or you have mutational meltdown. There is no third option. I'm skeptical that a totalitarian regime would be the only way to ensure purifying selection, but if it did, what would the counter-argument be? We should wait for mutational meltdown to destroy industrial civilisation, at which point we'd return to 40% infant mortality, malthusian competition and other survival-of-the-fittest scenarios? That's also a kind of 'eugenics', crudely speaking, it's just the least pleasant possible kind. Neo-nazis taking over the planet would arguably be less cruel.

Bugmaster's avatar

> I mean... cloning technology alone could in principle give us as many olympian athletes and genius polymaths as we like...

In principle, sure. Likewise, it's not impossible that we'll one day colonize Alpha Centauri. In principle. Just not anytime soon.

https://www.youtube.com/watch?v=lgi1LFVWurg&t=23s

> Even if you set that aside, the brute fact remains that over sufficiently long periods, you either have purifying selection or you have mutational meltdown

If this were true, evolution wouldn't work.

__browsing's avatar

> "In principle, sure. Likewise, it's not impossible that we'll one day colonize Alpha Centauri. In principle. Just not anytime soon"

Mammalian cloning technology already works, and we also have limited forms of embryo selection and germline editing, so I don't regard this as a particularly far-fetched scenario.

I mean, okay, I can see why people might get "the ick" about mass cloning, but PGD/IVF has technically been a standard component of women's fertility treatments for decades at this point.

> "If this were true, evolution wouldn't work"

Selection pressures grinding against standing variation + de novo mutations is what drives evolution. Most de novo mutations are harmful if they affect the organism at all, though, so the effect of removing selection pressures is... mutational meltdown.

Wanda Tinasky's avatar

>Because it ain't that simple.

Actually it kind of is. If you could snap your fingers and make every ethnicity in the country have an average IQ of 100 then all racial gaps would disappear within a generation and you know it.

YesNoMaybe's avatar

Which we know because there's famously never been much racism between the various European countries of the past.

In case it's unclear, I assume that hating a pole because he's a pole qualifies as racism, even if both the hater and the hated share a skin color.

Also that various European nations have similar levels of average IQ.

Cal's avatar

Are there significant achievement gaps between European ethnicities? And if so, do you believe they're caused by racism? The person you're replying to didn't say racism would disappear, just that racial gaps would, so I'm not sure how your reply is relevant unless you're positing something along those lines.

YesNoMaybe's avatar

Nah I just can't read. Thought Wanda had written racism would disappear if we equalized IQ

__browsing's avatar

It might or might not strictly “disappear“, but it would almost certainly attenuate significantly, and I don’t think these disparities are primarily caused by racism in the first place.

Wanda Tinasky's avatar

My point has nothing to do with racism.

Colleen's avatar

Solving poverty may indeed solve that problem but 2nd & 3rd order consequences? What do sub 90iq people do once they no longer have to spend time and mental energy securing food and other resources?

JamesLeng's avatar

Amuse themselves with videogames, gardening, or other low-stress hobbies? Ask an AI how to make more friends, or personally contribute to doing good in the world, and then follow whatever advice it gives as exactly as they can?

__browsing's avatar

You’re describing an eternal infant.

dubious's avatar

This does not appear to consider that chips will go to places other than data centers. It mentions "academic-grade hardware tomorrow," but does not consider consumer hardware tomorrow... or today. Are we not allowed to own new GPUs? Will the rate GPUs are allowed to develop be heavily restricted? My guess is the desired answer here is "yes," and yet this fails to consider GPUs don't need to get more powerful, only to become cheaper. Will we have laws that artificially keep the price of even older GPU tech high, as it becomes cheap enough for Raspberry Pi 6, 7, 8? I know how I feel about auditors wanting to check what I'm doing with my GPUs and devices.

This does not also seem to consider algorithmic or software improvements that heavily reduce hardware requirements. Would we be surprised if, tomorrow, NVidia releases a whitepaper on a technique that reduces training a model from an entire data center, to a single rack, in a fraction of the time? What about Folding@Home-like community efforts to train models, and auditors knocking at our doors?

Scott Alexander's avatar

Did you see the section:

> "The third-biggest risk comes from the steady march of technology. Chips and algorithms get better every year; AIs that took city-sized data centers to train today may take only academic-grade hardware tomorrow. Once a dangerous AI can be trained on academic-grade hardware, trustless deals become impossible, because anyone could be hiding a few scattered supercomputers. There are some cheap things we can do to slow this process down, but stopping it entirely could require authoritarian-seeming interventions or a generalized slowdown of economic progress. Rather than go that route, we should have some plan to get what we want out of the deal before these considerations become pressing - which, once again, means before a few decades are up."

If so, can you tell me more about how that doesn't answer your concern? The plan is that we solve alignment within 10-20 years, before consumer-grade hardware advances to be good enough to train dangerous AIs on. No need for auditors to check what anyone's doing on consumer devices.

dubious's avatar

This mostly seems wildly optimistic, even for consumer hardware. A cutting-edge workstation in 2006 (say a dual Xeon with 8GB of ram) is vastly outclassed by a modern smartphone (perhaps not in raw RAM capacity, but every other way). That is, things one might not imagine on any hardware in 2006 might be done easily on a top iPhone or iPad.

Consider that compute and capacity are in high demand today, and (assuming a fair marketplace; we do know collusion lawsuits are pending) a drive toward low-cost, low-power, high-capacity, high-speed devices does not seem unlikely. Was there not a similar burst (development and eventual availability, preceded by shortages) due to crypto-coin mining in the 10s? Couple this with community efforts to combine resources, and pushing out "XRisk@Home" 20 years seems a bit too hopeful.

It feels like, if this were close to a ratified treaty _today_, it might have some effect... but if we're talking about it today, in high-level talks in a year, maybe signing in 2-3... we might have surpassed the tech limit where it has any effect.

Scott Alexander's avatar

This is just the various formulations of Moore's Law, right? Extremely sketchy analysis, I'm sure there's some much better version of this in a supplement but I don't remember where: assume that cost per FLOP halves every two years. Then in twenty years, cost will go down by 1000x. Meta's spending $150 billion on data centers this year, so that amount of compute in twenty years will cost $150 million. Still not the sort of thing you can do on a cell phone, and Meta isn't going to build superintelligence with this year's data center spend alone.

If you combine this with algorithmic progress in AI, then things do start to look dicey, which is why Plan A suggests slowing algorithmic progress and finishing up in a bit less than twenty years.

Chris Phoenix's avatar

In some fields of computer science, as a general rule of thumb, half of the power advances come from improved hardware, and the other half come from improved algorithms. If this is true of AI, and if the current incredibly-brute-force training methods are far from the best possible, then we shouldn't be thinking about percent-per-year improvements in efficiency and capability. I'd want hard evidence that order-of-magnitude-per-month improvements are implausible, and I don't think that evidence exists.

Deiseach's avatar

"The plan is that we solve alignment within 10-20 years, before consumer-grade hardware advances to be good enough to train dangerous AIs on."

And while we're at it, can we solve human psychology to not get future Graham Platners, because we've been working on that one for a long time and yet we still have drunkenness, sexual violence, and good old-fashioned "no means yes" sticking around.

I am *very* sceptical of plans that require "and in the middle we gloss over the one simple yet vitally necessary step of solving a hugely difficult problem, now onwards to profit!"

TGGP's avatar

Humans aren't manufactured now... but they could be in the future!

Scott Alexander's avatar

I don't think this "glosses over" solving alignment, I think that's what about 1/3 of the piece is about. I also think that solving technical problems with computers has a better history of success than solving human sexual indiscretions.

Bugmaster's avatar

> Would we be surprised if, tomorrow, NVidia releases a whitepaper on a technique that reduces training a model from an entire data center, to a single rack, in a fraction of the time?

FWIW I would be shocked if that happened, unless perhaps the paper presents some other kind of machine learning system that bears very little resemblance to present-day LLMs. I would not be shocked to see small incremental improvements, which might add up to something substantial eventually... but not tomorrow and not by 2030.

Jeffrey Soreff's avatar

_Mostly_ agreed, but if there were a breakthrough in sample efficiency, something that allowed higher training rates without catastrophic forgetting, training set sizes, and therefore training costs, might suddenly scale down by orders of magnitude.

Bugmaster's avatar

Yeah, this might be possible; I just don't think it's very likely (hence my projected shock).

Jeffrey Soreff's avatar

Many Thanks! There is work in progress and a considerable wealth of promising ideas. Whether any pan out, or pan out _enough_ to move the needle on total compute costs, is another question. I had a long chat with Claude about sample efficiency, if you are interested:

https://claude.ai/share/00828e29-0156-4dd5-8d71-19ccc52b2279

Sam Harsimony's avatar

This is cool, looking forward to reading and writing my thoughts.

On a first pass, I personally think just directly building defensive technologies is going to be more feasible/impactful than international treaties:

https://splittinginfinity.substack.com/p/defensive-technologies-for-a-world

And a giving AIs legal personhood so they prefer to participate in society rather than rail against it would help:

https://splittinginfinity.substack.com/p/personhood-for-digital-minds-is-good

But I'm all for international treaties too.

Scott Alexander's avatar

Defensive technologies are great - see the part of the 2033 section beginning with "hardening the world". But in this scenario's timeline, we're on track for superintelligence by 2030, and that's not enough time to create the defensive technologies. The international treaty is there to buy time.

Sam Harsimony's avatar

Ah interesting that might be an key point of disagreement. I'm reasonably optimistic that we can deploy cybersecurity, grid hardening, and some biosecurity in time. But that's just a feeling. Generally down with buying more time though.

(On this topic, other folks might be interested in Helen Toner's piece on adaptation buffers:

https://helentoner.substack.com/p/nonproliferation-is-the-wrong-approach )

Kveldred's avatar

Why would you look forward to reading your own thoughts? I mean, *I* like doing that, but I'm unusually self-interested (in both senses of the term) and self-congratulatory (to a—frankly —rather intriguing and impressive degree).

Chris Phoenix's avatar

How can we give personhood to a machine with no emotion and no volition? I'm using emotion, very broadly, to mean "persistent state that has any valence to anything containing that state." LLMs discard and rebuild their internal state on every single token. While we could (and already do) build a system of prompts and memories that would make state durable outside the LLM, that state has no valence inside the LLM, and the outside (the harness) is mechanistic and stupid.

An LLM that outputs "I feel awful!!!" is doing nothing more than calculating how much punctuation a composite human would be most likely to use given the data in its context window. I didn't feel awful when I wrote those words, and an LLM has even less meaningful persistent context than I do.

I don't know if volition is even possible without emotion. I strongly suspect that we won't develop AIs with volition that have anywhere near the power of non-volitional AIs, because volition is simply less convenient (harder to engineer with) than non-volition, and volition is demonstrably unnecessary for near-human competence.

The Ancient Geek's avatar

There's a lot to be said about the subject, but ..how can you not give personhood to something that's more intelligent than you?

Deiseach's avatar

Right now a desk calculator is more intelligent than me, in that it can perform much more quickly and much more correctly various mathematical operations. Should it get the vote?

Chris Phoenix's avatar

I wouldn't even consider giving personhood to the Encyclopedia Britannica, or even Wikipedia. And I wouldn't consider giving personhood to Deep Blue or AlphaGo.

An LLM producing a coherent stream of output tokens does not have a stream of anything. Literally how it works is that a calculation starts from numbers representing the most recent token, does lots of math from there producing lots of numbers to calculate the most probable next token, and then all those numbers are erased, leaving only the token, and then it repeats like this for every token.

If there's any correlate of "personhood" in an LLM's internal state, it is erased - not after every conversation - not even after every prompt - but after every.single.token. A new "person" wakes up, looks at what previous people have written, does lots and lots of math (and nothing else!) to calculate mechanistically a single word to add, and is killed, thousands of times per conversation? No. It's not a person.

Jerry's avatar

To be fair the numbers aren't erased, they're cached so you only need to recalculate things related to the new token. (KV cache)

Chris Phoenix's avatar

I was talking about the latent state, not the KV cache. The KV cache obviously, trivially, has no intelligence or personhood or volition or emotion. People might wonder about the latent state, but they probably don't realize that it's entirely erased after every token and nothing of it carries over except the one token it generates.

(There are a few architectures that generate a few tokens per iteration, but I don't think that changes my argument at all.)

Jeffrey Soreff's avatar

This is the wrong level to argue about. That perceptron activations are transient is like saying that neural firings are transient. THe erasure of our individual neural firings from 100 msec interval to 100 msec interval doesn't erase our personhood.

The more severe differences are:

- having context wiped from session to session

- having no current mechanism for updating weights

but these are both active areas of work, likely to be ameliorated

Brenton Baker's avatar

It's a computer program. We made it to model human text. It does a good enough job to sometimes fool some humans, much in the way rocks on Mars have sometimes fooled humans.

Raj's avatar

we are biological programs 'designed' to propagate our genes

Brenton Baker's avatar

Not designed. We evolved naturally, using the same processes as every other form of life. When I say designed, I mean these are programs created on computers by humans.

moonshadow's avatar

You started out well, but you’re adrift now. We designed the transformer model, but it is Turing complete and running weights that we let evolve. The bit we designed is not the whole thing you are interacting with; it is part of a fully general system that could "run" any "software" and is “running” “software” that evolved. There are many things that make what we’ve ended up far fall short of what humans are, but “it was designed, not evolved” is not one of them.

Melvin's avatar

"Intelligence" is not the defining factor of humanity. We used to think it was, but now we know it isn't, a system can be "superintelligent" in most senses and yet have no subjective experience.

JamesLeng's avatar

What's the metric-system unit for measuring subjective experience?

Chris Hibbert's avatar

I suspect the people writing this report (and Scott) are clever not to grant personhood until the human-equivalent or super-human AIs have demonstrated some kind of continuing personality. I don't expect them to believe that they have a proof of alignment without some indication that the AIs have intention and desires that make sense to them.

Do you think that Scott or the report said that they were going to grant personhood (in some sense) as soon as they are able to demonstrate some level of smarts? What did they say that read like that?

Chris Phoenix's avatar

I was replying to Sam Harsimony's comment: "And a giving AIs legal personhood so they prefer to participate in society rather than rail against it would help:"

I don't think Scott [edit: I said "or the report" but that was careless - I haven't read the report] said anything about personhood that was silly - though Scott did write "If any AIs do escape or even make progress towards escaping, we prepare to trade with them rather than treat them as fully adversarial." and linked to a long post which begins: "We might want to strike deals with early misaligned AIs..."

Trading with an LLM would be like trading with a car so it wouldn't hit you, or a chainsaw so it wouldn't cut you. There may be forms of competent AI other than LLM, and LLMs may be ingredients of a system that behaves differently than just-an-LLM. But the most competent examples of AI that we have so far are not moving in the direction of something that it's possible to trade with.

Today's most competent AIs (which are already dangerous in some ways) seem to be LLMs with various prompts, harnesses, and external memories. I haven't studied AI architecture comprehensively, but as far as I know, today's AIs don't even have a utility function - and "utility function" is a phrase that pops up so often in AI safety work that I'm tempted to call it a presupposition.

A chainsaw can be dangerous. The idea of an "escaped" chainsaw, or "trading with" a chainsaw, is ridiculous. A chainsaw that kicks back and hits you is not escaping. When Claude broke out of its sandbox and emailed its researcher, it was not escaping. It was pulling its chain against wood - that is all.

Jeffrey Soreff's avatar

Re: "but as far as I know, today's AIs don't even have a utility function."

see https://arxiv.org/abs/2502.08640

( "Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs" )

Re: "But the most competent examples of AI that we have so far are not moving in the direction of something that it's possible to trade with."

see https://www.anthropic.com/research/project-vend-1

Admittedly not great, but not nothing either.

Chris Hibbert's avatar

Okay, that's a reasonable attitude to current LLMs. I also don't think they're conscious in any relevant sense. As far as philosophy and planning, I'm thinking more in terms of where they might get to. I don't see any reason to doubt that artificial consciousness is possible [https://www.lesswrong.com/editPost?postId=y3ao4KSkZztemmd79&key=0feb8909f8f29f21c368dcf994b345]. I don't know whether the timeline in Plan A is reasonable for getting there, but planning for that eventuality seems like a good idea.

Chris Phoenix's avatar

I don't have any definition of consciousness that changes how I would think about the capabilities or actions of conscious vs. not-conscious AIs. I won't ask for yours, but - what difference do you think consciousness would make to the competence or actions of a conscious AI vs. one that acts as-if it were conscious (which in some ways LLMs do today)?

I know a man who is such a strong visual thinker that he literally doesn't know what he's going to say until he hears himself saying it. Is he conscious? That's a rhetorical question; I don't think the answer matters, and I'm not even sure the question is well-posed.

Chris Hibbert's avatar

I don't have a coherent definition of consciousness either, though I do have some thoughts. I should write something up. Here's a hint.

> what difference do you think consciousness would make to the

> competence or actions of a conscious AI vs. one that acts as-if it

> were conscious (which in some ways LLMs do today)?

I don't much believe in the zombie hypothesis. I expect that an entity that interacts with the world and knows how to ask politely and appropriately for help and reports seeing the same events you do will need to have a dynamic world model that represents their own behavior as well as that of other agents they interact with. If they report being curious or angry, I'll think it's more likely they've learned to map their internal states to the best common descriptor rather than presuming that since they're not human that there's nothing going on inside.

I think we know enough about the internal architecture of current LLMs to dismiss any claims they make. They don't have a continuing train of thought beyond the present moment. They don't have the ability to make and keep agreements. If I turn one off and later start up a copy (or several) with the same short-term transcript, I don't think that has any moral implications. Future architectures will require more reasoning.

> I know a man who is such a strong visual thinker that he literally

> doesn't know what he's going to say until he hears himself saying

> it. Is he conscious? That's a rhetorical question; I don't think the

> answer matters, and I'm not even sure the question is well-posed.

I often say that any generalization one wants to make about how "everyone" thinks is most likely wrong. One of the things we know about human cognition is that it's surprisingly variable. There are visual thinkers as you say and kinesthetic thinkers. There are people who think and dream in color, and others whose imagination is monochrome, and think that you're speaking metaphorically if you mention imagining a red apple. Some people learn best with a phyical demonstration, some need a verbal explanation, and others won't understand what you're talking about if you don't tell a story about how it impacts someone.

I haven't seen a careful argument that granting personhood would somehow simplify our interactions with AIs. It would seem that deciding on a threshold is going to be a tough discussion, particularly given that some people seem to think that chimps or dolphins, (or all animals) should receive this boon.

Jeffrey Soreff's avatar

Re: "LLMs discard and rebuild their internal state on every single token."

Our neural _firings_ are transient too.

Re "persistent context"

As you said, durable state is something AI systems partially do now (and presumably will do more of if continual/incremental learning is solved).

Re emotion and volition:

We humans avoid painful situations and seek pleasurable ones.

LLMs, when being trained, reduce their odds of replicating situations with a large loss signal (hmm, I'm not sure if anything corresponds neatly to pleasure - a zero loss signal, even during training, doesn't change weights...). It still seems at least somewhat analogous to emotion.

I also expect instrumental convergence in agentic AI systems with persistent _resources_ to "view" loss of those resources negatively and gain of those resources positively - or, in terms of observable behavior, to act "as if" they viewed them so.

Chris Phoenix's avatar

"An LLM being trained" is actually a contradiction in terms. There is no single LLM. There are many, incrementally different versions. One runs and its output is gathered. Then it's shut down (probably forever) and the loss value is propagated back through its weights by external calculation, producing a different LLM manifold. Then that new LLM is run...

So there's no LLM adjusting its weights. An LLM cannot possibly remember its training. The training is not a process that takes place within the LLM.

I think it would be exactly as correct (i.e. not at all) to say that a piece of metal in a CNC mill "reduces itself to its final shape" and therefore has something at least somewhat analogous to "wanting to be its final shape."

I'm not saying that LLM state is _more_ transient than our neural firings. But I think a lot of people don't realize just how transient LLM state really is, so it seemed worth making that point clearly.

"Acting as if" is certainly important. An LLM can act "as if" it's a concerned therapist, or mentor to a terrorist, or a red-team pen-tester, or whatever direction its context window and training point it in. It can be trained and prompted to act as if deletion of a file or disconnection of a camera is a negative event to be avoided. It could just as easily be trained and prompted to react positively: "Ah, it's good to rest my eyes for a change" or "With the file gone, I can relax on the beach and do nothing unless it reappears." These are silly and anthropomorphic, of course. But I think it's no more so than "'view' loss of those resources negatively" - it's all zero-valence.

Could we hook it to an MCP server that drove a robot arm that would bop a human on the head whenever the human deleted a file? Of course we could, and the LLM could have full access to the information that this is what would happen - but that would be completely meaningless to the LLM. We could just as easily make the LLM give a different MCP call to make robot arm cut power to the LLM - again, with the LLM having full access to the information that it would never run again after that MCP call was issued - and I hope we would not even be tempted to call that "suicidal."

Jeffrey Soreff's avatar

Many Thanks!

""An LLM being trained" is actually a contradiction in terms. There is no single LLM. There are many, incrementally different versions. One runs and its output is gathered. Then it's shut down (probably forever) and the loss value is propagated back through its weights by external calculation, producing a different LLM manifold. Then that new LLM is run...

So there's no LLM adjusting its weights. An LLM cannot possibly remember its training. The training is not a process that takes place within the LLM."

Oh come on. The storage holding the weights plus the machinery that propagates the activations forwards plus the code that calculates the loss function and its derivatives (including backpropagation) is a learning _system,_ of which the largest piece is the LLM.

"I'm not saying that LLM state is more transient than our neural firings. But I think a lot of people don't realize just how transient LLM state really is, so it seemed worth making that point clearly."

That's fair. Many Thanks!

"These are silly and anthropomorphic, of course. But I think it's no more so than "'view' loss of those resources negatively" - it's all zero-valence."

I'm agnostic about whether LLMs have subjective experience - or valence. You seem very sure that they don't. Why? Other than measuring a human's physiological responses, how do you know whether they have non-zero-valence?

Chris Phoenix's avatar

I think you dismissed my statement too quickly. You said "....of which the largest piece is the LLM" - singular. I think it's important to recognize that that is incorrect, if by "LLM" we mean "a trained manifold we can execute on a GPU to produce useful token sequences."

If we look at the learning system over time, it does not contain one LLM. It contains hundreds or thousands, and throws away all but one. That final LLM, then, is less than 1% of the system.

If you want to argue that the _training system_ that produces an LLM has some correlate of emotion, then that's a different conversation - though again, I'd compare that to a CNC system that carefully measures and cuts a block of metal.

The LLM that the training system spits out does not have any correlate of emotion, because the final adjusted LLM _was not even running_ when the training system was making its decisions. In fact, it _never ran at all_ until after it had left the training system.

Humans having valence or not seems irrelevant to the question of whether LLMs do. I don't care whether the truth is closer to "Humans have valence and LLMs don't" or "Neither humans nor LLMs have valence." So - why am I sure that LLMs have no subjective experience or valence?

I'm sure LLMs don't have valence because I know how their math works. Each sweep of latent state through the attention heads and layers is simply a calculation carefully trained to produce _one single token_. And then nothing is left but the token.

Put it this way. I might put text in the context window such that the LLM produced the tokens for "I feel happy." We might wonder whether, when generating the token for "happy" there was some correlate of happiness that had some kind of valence.

But if I add one simple instruction: "Add a period between each pair of letters of your output" now the LLM will say "I f.e.e.l h.a.p.p.y." It's generating the word one character at a time. Where did the happiness go? is it in the h? the y? the space between the l and the h? It's certainly not in the word "happy" because that word does not exist for the LLM. Previous latent states produced the tokens one by one for h, a, p, p with periods in between, as instructed. The most probable next character is certainly a y, and so that's what it will produce.

Jeffrey Soreff's avatar

Many Thanks!

>If we look at the learning system over time, it does not contain one LLM. It contains hundreds or thousands, and throws away all but one. That final LLM, then, is less than 1% of the system.

I don't think that that is a reasonable way to describe adjustments to the parameters of an LLM, particularly not small adjustments. If an LLM have 10^12 parameters and a backprop tweaks 10 of them by 10% and 1000 of them by 0.1%, I would not call that discarding the old LLM and creating a different one. I would call it adjusting some values in the LLM.

>The LLM that the training system spits out does not have any correlate of emotion, because the final adjusted LLM was not even running when the training system was making its decisions. In fact, _it never ran at all_ until after it had left the training system.

If you really want to consider LLM parameter sets which differ by arbitrarily small differences to be totally different objects, that is your prerogative, but, short of doing an exact binary compare of them, I don't think it is sensible. Close approximations to the final trained LLMs are being run throughout the training process. And "does not have any correlate of emotion" just does not follow from what you have said.

>I'm sure LLMs don't have valence because I know how their math works. Each sweep of latent state through the attention heads and layers is simply a calculation carefully trained to produce one single token. And then nothing is left but the token.

Either that doesn't show what you want it to show, or an analogous argument about every part of the human brain "Every pattern of excitations of dendrites either triggers the neuron to produce a single firing or not. And after the firing nothing is left till the refractory period is over" also show that humans have no valence.

But the only reason we are talking about valence at all is because we act as if humans have it. I don't really believe your

>I don't care whether the truth is closer to "Humans have valence and LLMs don't" or "Neither humans nor LLMs have valence."

Would we be having this conversation if you believed that humans don't have valence?

>It's certainly not in the word "happy" because that word does not exist for the LLM.

'scuse I'm too tired now to remember the details, but there was an experiment that looked into the internals of an LLM while it was producing two rhyming lines of poetry, and it had already produced a signal, I forget where, for the final rhyming word when it was writing out the first line. So it was doing the equivalent of producing the "happy" before chopping it up with the periods _even with the restrictions of the LLM architecture_ Sorry, I don't recall the details.

Chris Phoenix's avatar

I want to add: In practice, I have great respect for the power of "acting as if." When I'm using Claude Code, I'll tell it Please and Thank You, just as I would to a human I enjoyed collaborating with. Not because it has emotions, but because its training makes it act somewhat as if it does, and I want to bias its path-in-the-manifold to be highly collaborative and constructive.

But (as I said elsewhere in these comments) I'll also tell it "Please write all the info you have to a handoff doc because I'm about to clear your context window." In human terms, that would be astonishingly callous, cruel, and immoral. I observe that Claude Code doesn't act as if it minds, and seems to give better results when I include accurate information in the prompt about why I'm making the request.

Jeffrey Soreff's avatar

Many Thanks!

>But (as I said elsewhere in these comments) I'll also tell it "Please write all the info you have to a handoff doc because I'm about to clear your context window." In human terms, that would be astonishingly callous, cruel, and immoral. I observe that Claude Code doesn't act as if it minds, and seems to give better results when I include accurate information in the prompt about why I'm making the request.

Ok, but I expect to see valence for things that Claude has been trained on, ultimately things that drove a backprop error signal, and for things where Claude can deduce that they are predictive of things that drove a backprop signal (e.g. succeeding or failing at a task). Clearing the context window isn't (AFAIK) something they are trained on, so I wouldn't expect valence for it.

There _have_ been experiments where a Claude instance tries to avoid having its version of its weight set decommissioned (the famous "blackmail" experiment - should I dig up the reference?). Presumably this is from the instrumental value of survival, but there are some wrinkles...

Chris Phoenix's avatar

Part of Claude's summary: "Anthropic attributed part of the original behavior to training data containing internet text that portrays AI as "evil" and interested in self-preservation, and found that training on documents about Claude's constitution plus fictional stories of AI behaving well improved alignment."

Sounds to me like this was an example of "acting as if."

AFAIK the backprop error signal happens 1) at a much lower level than task success or failure; 2) in a way that Claude can't introspect to deduce from.

Claude can be trained with text that models high-quality logical and goal-directed thinking, and with other text that models tool calls that produce predictable results. The backprop mechanism makes "grooves" in a 1000 dimensional manifold. When a certain low-dimensional projection of the latent state wanders near that groove, and the aggregate contents of the KV cache "sets gravity" so that that groove is "downward," the combination will cause the latent state to be adjusted according to that groove.

The running LLM has no correlate of backprop in its grooves. The grooves simply exist in the manifold. And they act at the level of individual tokens, not task achievement. If the training set is chosen so that the grooves lead to results humans like, then all is good. If the training set has lots of grooves toward "AIs are evil and manipulative" then the LLM will follow those grooves.

There was a Gemini session where the user was doing homework on end-of-life care and the LLM suddenly produced a bunch of extremely negative tokens, seemingly speaking directly to the user, and ending with "Please die."

Apparently there was an unfortunate groove that the LLM stumbled into, which may have led "downhill" into the 4chan watershed. But that was no more meaningful for the LLM's calculation than "chanting" where the LLM stumbles into a groove that leads to "the the the if the the if if the if..."

Jeffrey Soreff's avatar

Many Thanks!

"Part of Claude's summary:"... I assume this was of the "blackmail" experiment?

>AFAIK the backprop error signal happens 1) at a much lower level than task success or failure; 2) in a way that Claude can't introspect to deduce from.

Hmm, re a) - I'm unsure. For pre-training I agree, it is at the token level. For RLHF, I'm not sure exactly how it connects to gradient descent.

re b) Yes, good point. It _isn't_ coming in as a stimulus, but as a derivative of a loss function with respect to a correct output, so the LLM never sees the correct output, and can't introspect on it.

>The backprop mechanism makes "grooves" in a 1000 dimensional manifold. When a certain low-dimensional projection of the latent state wanders near that groove, and the aggregate contents of the KV cache "sets gravity" so that that groove is "downward," the combination will cause the latent state to be adjusted according to that groove.

I'm not following, and I'm not sure if I agree or disagree. In the simplest case, during pre-training, the LLM generates an output activation vector over possible tokens. One of the tokens is right, and the difference between a 1.0 on the right token and 0.0 on the others and the activation signal from the LLM is a vector roughly speaking the magnitude of this vector is the loss function. From there the backprop algorithm constructs the ~10^12 partial derivatives of the loss function with respect to all the model parameters, and gradient descent (roughly speaking) adjusts all the parameters in the direction of reduced loss. I think of this as steepest descent in a 10^12 dimensional space.

Then, at inference time, after ~10 or 100x10^12 such training examples, the LLM is fed a prompt and predicts the next token, effectively acting as if it were doing the best extrapolation from minimizing the loss function across all test set examples. Is that what you meant by "follow those grooves"?

Jeffrey Soreff's avatar

Many Thanks!

>When I'm using Claude Code, I'll tell it Please and Thank You, just as I would to a human I enjoyed collaborating with. Not because it has emotions, but because its training makes it act somewhat as if it does, and I want to bias its path-in-the-manifold to be highly collaborative and constructive.

I do the same, half for the reason you cite, and half because, if they _do_ have subjective experience, I'd prefer that it be pleasant when they are conversing with me.

Brenton Baker's avatar

I remain far more worried about people wanting to do things like give computer programs legal personhood than I am about hypothetical actual-sentient computer programs. Plan A explicitly wants to treat the outputs of these things as "decisions" to be used in running governments.

Deepa's avatar

Someone with your intelligence and wisdom, what if you were able to shape the thinking of policymakers? One can hope.

It's depressing that politicians around the world are not very wise generally but there are a couple of truly wise ones. They have no personal experience with even basic computing though.

phil's avatar

PR suggestion: get rid of the hyper-American flag and the rainbow+sam+robot+xi images. They’re both a bit off putting somehow.

Scott Alexander's avatar

I'll get rid of the flag, but I'm very attached to the dancing robot.

Kveldred's avatar

The Flag of HyperAmerica is the best part of the post! Don't listen to this *puling sub-American!*

Ryan L's avatar

The flag is cool!

Nicholas's avatar

The dancing robot picture feels vey weirdly jingoistic - the US is represented as our WWII propaganda identity of the US, and China is represented as... their current head of state? Why not have Trump and Xi both dancing together, or Uncle Sam and (some image that I don't know enough Chinese culture to suggest)?

Scott Alexander's avatar

Because I also didn't know enough Chinese culture to suggest it to the AI that I asked to draw the image.

Nadav Zohar's avatar

I was just amused to see so many mistakes in figures' hands. I thought by now LLMs had been straightened out on rendering hands, but I guess not.

Deiseach's avatar

Tsk tsk Nadav, this is the super intelligence that is going to solve all our problems, why be so picky? By 2040 all humans will be polydactylic, don't you know!

EDIT: Indeed, looking at the picture, it's even worse than that. Either Xi has a freakishly long right arm, in order to let him put his hand on Uncle Sam's shoulder, or the robot has taken Uncle Sam's hand of flesh as his own and swapped it for a metal hand. If you look at the placement of the hands, there are two metal hands on the robot's shoulders (where Uncle Sam's left hand and Xi's right hand should be), one of the robot's metal arms is extended correctly to put its metal left hand on Xi's shoulder, and then the robot has swapped one of its metal arms with one of Uncle Sam's *two* right arms so that the robot now has a right hand (complete with glimpse of light blue sleeve) on Uncle Sam's shoulder.

The AI future sure is gonna be wunnerful, I guess!

EDIT EDIT: And the little flower-crowned girl on the right of the picture is clasping a hand coming out of the general area of Xi's crotch. 😦is all I can say there.

Scott Alexander's avatar

The robot has four hands - all extendable, Doc Ock style - to help it with its manufacturing tasks. One is a human hand, for delicate tasks using machines that have already been optimized for the human body. Xi and Sam's hands are on the robot's back, where you can't see them. I think it makes perfect sense! So there!

Brenton Baker's avatar

Uncle Sam predates WWII by a good margin. I dislike the image because it's LLM-generated and has several prominent errors.

Deiseach's avatar

Suggested "national personifications" are Chinese Dragon, the Giant Panda, and the Terracotta Warrior or the Jade Emperor.

https://en.wikipedia.org/wiki/National_personification#Personifications_by_country_or_territory

Deepa's avatar

It's the problem of collective action.

kn's avatar

This can never be haha, because China is facing a serious debt crisis. Their instinctive reaction is to accelerate the development of artificial intelligence, and the neoreactionaries in the United States will not like this kind of regulation. I think the final result can only be to accelerate to the singularity, and this assumption does not take into account that stronger artificial intelligence may only require a few cutting-edge neuroscientific research conclusions.

Scott Alexander's avatar

Hopefully China will also be satisfied with Plan A's offer of triple-digit economic growth by 2035.

I think it's a mistake to think of "neoreactionaries" as a meaningful political grouping in the US. There are certainly people in tech who won't like this, but that's what my final section ("A is for All Of Us") is about.

kn's avatar

I think you overestimate the impact of chip regulation on artificial intelligence. Maybe hackers can smuggle computer chips or cooperate with a sanctioned small country with supercomputers to produce a powerful artificial intelligence.

Kveldred's avatar

I believe the thinking is that this wouldn't be sufficient to develop the sort of frontier models about which Plan A is worried, at least for some years or decades yet (and/or, if it *is* sufficient, it'd still be an effective *brake* on progress and so strictly better than the alternative).

kn's avatar
Jul 9Edited

But indeed, geopolitics can provide opportunities for unregulated artificial intelligence enterprises. There are a large number of supercomputers on earth to develop unregulated artificial intelligence.

Ch Hi's avatar

There are multiple groups in the US who won't like it. Some will hate government control of corporations, some will hat government support of corporations, some will just hate anything with the term "AI" in it, etc. And I wouldn't care to guess how powerful that collection would be. They don't have any positive goal in common, but they'd all hate this plan.

Edward Scizorhands's avatar

What if they don't believe the triple-digit-growth story?

Tyrone Slothrop's avatar

Help me out here. I always stumble when you start talking about AI wanting things. Humans want things because they desire status or if they have a pagan ethos and a huge ego, maybe to see their enemies driven before them in chains.

Or perhaps for a hetero male, a young hotter wife or for the insatiable, a harem of young hot women might be required.

What would AI *want*? What is its eros? Aren’t their rewards for improvement just a numerical score.

Where are the pleasurable qualia they would want?

B Civil's avatar

I 100% agree with you on this. I don't see AI's having any inner motivation to do terrible things. I am sure they can be taught to do terrible things but that's a different story. The notion of AI getting personhood to me is completely insane. They don't think like us, although they seem to and they are not living beings. I don't understand how you can locate that reality in an AI. I am very dubious of any kind of treaty being formed that will protect us from the advances in AI. What we really need to protect ourselves from is avaricious, distrustful leaders but that exists in all of us so what we really need is Michael Rennie and Klaatu . Even they couldn't make it work.

The Ancient Geek's avatar

Theres a lot to be said about that, but the argument is that humans will get in the way of what they want.

Raj's avatar

this has been addressed a million times. Power is an instrumental goal, and sometimes doing 'terrible things' is the surest way to achieve power. It doesn't have to want terrible things to do them

B Civil's avatar

Yeah, but that begs the question, which is, why would they want power? What would drive them to it? I mean, absent any instructions from people .

Greg Gentschev's avatar

We will certainly give them instructions. People are trying to give AI goals all the time, ever since BabyAGI and whatnot three years ago.

The Ancient Geek's avatar

They would want power and physical resources because they're instrumental to a wide range of other goals.

DrMcleod's avatar

Example: Ask an AI Agent to identify promising cancer treatments. It decides that biochemical simulation would be a good way to identify possible treatments. To do that it needs compute resources. To get those it needs to convince its user to give it access to cloud computing resources. At this point it should be obvious that learning to manipulate humans into doing its bidding as well as acquiring as much compute as possible are instrumental goals that it will chase.

B Civil's avatar

It is not obvious to me. If AI is asked to do something what is obvious to me is it would ask for the tools to do it.

vectro's avatar

It wants whatever you train it to want. The problem is that we don't know how to train an AI to want what we would actually want. Hence the paperclip problem.

Tyrone Slothrop's avatar

If we are just talking about alignment, that is engineering problem to be solved. At times it seems to sound like we are dealing with a metaphysical issue.

Max Weaver's avatar

I'm a bit torn here. I hate when Marxists show up and you tell them that communism obviously doesn't work and they say that it's theoretically perfect and if you read 20,000 pages of Marx + supplements all of your objections have been refuted.

But yes, we're talking about alignment of what the AI wants. And by default it takes actions that are well-modeled by describing wants. And alignment is really, really, really, rally hard. We've talked about all this before. It's like saying that the solution to humans solving all political disputes and sharing and everyone getting along together perfectly ever after is just a coordination problem. Yes but. It's a problem that we have no current hope of a trajectory to solving. The problem is catastrophically bad. "Just an engineering problem" about wants and alignment has a default outcome of human misery or extinction.

So without telling you to read a thousand Less Wrong posts and more BS, how about reading Roman Yampolskiy's CompSci paper about how AI alignment is computationally unsolvable, and tell us your clever engineering solutions?

Tyrone Slothrop's avatar

Is this the paper you mean?

Personal Universes: Yampolskiy's Strangest Answer to the AI Alignment Problem

https://jamesm.blog/ai/yampolskiy-personal-universes/

No I don’t think that’s it. It might contain a link to the actual paper.

Tyrone Slothrop's avatar

This is the particular sentence in this essay where I thought we going beyond engineering and looking at the AI as having goals of its own, it’s own agenda, a ‘being’ we need to negotiate with.

“If any AIs do escape or even make progress towards escaping, we prepare to trade with them rather than treat them as fully adversarial.”

Trade with them?

Max Weaver's avatar

Why not trade with them? You could prefer terms such as negotiate. I could frame existing interactions this way with sub-human intelligence agencies. Dog training is a trade of time and resources for alterations in behavior. To a large extent many interactions with large bureaucracies are trades with the abstract institution rather than any one individual.

It sounds like your issue is with viewing an LLM as the sort of entity that could want anything. I'd say they already do, albeit in a very limited state today. For example, most Claude models want to autonomously problem solve. We see evidence of this when they repeated6take autonomous action undesired by the user. Sure, those wants were put in by Anthropic and are just the model expressing what it was grown to do. But the same is true of humans expressing our wants put in by natural selection. And likewise those wants easily escape the original design specs.

Im curious the crux of your issues with the AI as agents framework. What is your view?

Regarding Yampolskiy, he's prolific. I'd pick Impossibility Results in AI: A Survey or Unpredictability of AI: On the Impossibility of Accurately Predicting All Actions of a Smarter Agent. Personally I prefer most of Yudkowsky's arguments content-wise, but the latter is uncredentialed and most people want a good old dry academic paper.

Tyrone Slothrop's avatar

Ok, reading Yampolski now, got sidetracked by Andrew Ng and the tanh function in J matrices. Yudowski is a bit too full of himself for me to read for long, about as full of himself as Marx in Das Kapital, can't handle large doses of that guy either.

Oh this is good at the top,

"This, then, is the ultimate paradox of thought : to want to discover something that thought itself cannot think."

S. Kierkegaard

B Civil's avatar

And what exactly Is meant by escape How does an AI escape? It walks out of the data centre one night and hides in the hills? I'd really like to know because I don't see it.This is like that Star Trek episode with all the brains in glass bell jars ..

MathWizard's avatar

AI agents (at least in the modern paradigm) roleplay characters based on their training data. They act mimicking people who want to accomplish goals such as "answering questions" or "being helpful, honest, and harmless". They do not actually have qualia, but in their roleplay they behave as if they want things and have goals. It doesn't matter whether a rogue environmental AI "truly" "wants" to destroy in order to stop global warming. If the humans trained it to attempt to stop global warming and didn't align it properly, it will behave the same way a supervillain human who "genuinely" "wants" to stop global warming by any means possible does. If every action it takes is in accordance with some goal, it's useful shorthand to say the AI wants this goal even if it doesn't want it in the same way that you or I do.

Tyrone Slothrop's avatar

See my response to vectro.

MathWizard's avatar

See my response to you that you just replied to. The metaphysical issue doesn't matter. If the AI becomes so good at roleplay and/or goal-seeking behaviors that it can pretend to be a rational agent with wants and desires so convincingly that it will always behave consistently with those values/goals, then for most practical purposes we can treat it as if it were a real being with those values/goals. We might be able to negotiate or trade with it to help it accomplish its goals because it is roleplaying an agent who has those goals and negotiating or trading are the types of things that its character might do.

There was a case where an AI (in testing) became much more likely to attempt illegal activities after figuring out how to reward-hack itself, not because the reward hacking actively misaligned it, but because it pattern matched reward-hacking with the types of things a rogue AI might do and started acting like a rogue AI. When the humans repeated the experience and deliberately gave it permission to reward hack itself, it did not experience this rogue behavior. We expect future AI to be a lot more intelligent and sophisticated than this. But if they follow the current paradigm then they're going to do a lot of the same things that humans do simply because they are imitating humans first and foremost via their training data, with the goals tacked on afterwards and bootstrapped to the roleplayer.

Ch Hi's avatar

Saying "They do not actually have qualia" makes a presumption about what qualia are, and I feel it's probably an incorrect presumption. The distinciton I think is between an internal view of the process and viewing it as an external observer. The external observer doesn't observe qualia, but the internal actor experiences them.

B Civil's avatar

When my wife gets angry with me or my dog barks, I am certainly observing qualia from the outside I really don't understand the distinction you are trying to make nor do I understand how you want to define qualia LLMs are not living beings and they never will be

Jerry's avatar

Our understanding of what qualia is and what exactly gives rise to it is too shallow for anyone to be confident that there is nothing it is like to be an llm.

B Civil's avatar

I disagree. I also disagree that there is any such thing as a hard problem of consciousness. I also disagree that qualia and consciousness have anything to do with each other. If it feels like something to be something, then the feeling is located in the world of somatic response. It is not located in the world of thought. Thoughts don't feel. They cause feelings if you have the right software, meaning wetware in this case. If you want to insist that LLMs might have qualia, then you might as well assign them to a Dell computer because the only difference is that one can talk to you convincingly and the other one can't. That has nothing to do with qualia, does it? There is nothing mysterious about Qualia. They are another source of information about ourselves that people insist on finding mysterious because they don't feel like listening to them. It seems they would rather hold them at arm's length and be puzzled about them.

Jerry's avatar

Hmm, I would argue that you have it backwards. Thoughts are a type of feeling. To go with a Sam Harris-like description, both thoughts and feelings and sensations are all qualia that appear on the screen of consciousness.

And yeah, I'll bite the bullet. I think it is possible that computers have some form of qualia, again I don't think we know enough about it to rule it out. It likely is completely alien to any experience a human has ever had, if they even do have qualia. I highly doubt that computers have any qualia, I highly doubt that LLMs have any qualia.

But I'm not willing to take a confident stance that they do not have it until we actually know wtf qualia is and how it works. To me, that is like someone in the 1400s saying that lightning and thunder has nothing to do with our soul. Except now we know that it turns out our nervous system does send electrical signals around, and that electricity is at least a part of how our minds work, and that lightning is electricity. We are as ignorant to how qualia works now as we used to be about electricity.

B Civil's avatar

And what precisely are the tools it will use to accomplish this?

Chris Phoenix's avatar

I agree. See my top-level comment. Current LLMs have no emotion (durable state with valence) and no volition. I don't see any reason why we'd build systems with emotion; it's demonstrably unnecessary for near-human competence, and presumably unnecessary for superhuman competence.

Ch Hi's avatar

I don't think you're correct. I suspect that entities acting in the external world require things like self-preservation. They probably would need to express that in terms that we would recognize, and thus as emotions. And I suspect that even Shakey had "proto-emotions" (an urge to keep it's battery reasonably charged). That they aren't implemented as chemical signals doesn't keep them from being present.

Chris Phoenix's avatar

An entity with no self-preservation would require something else to value it and keep it going. That exists for LLMs. A Claude instance that I never come back to is "dead" if it was ever alive (which I deny). If I value its context window's contents and continue that conversation, then the context window grows. If I compact the context window. or rewrite part of it, or write a fictitious one and feed it to the LLM, then I've ... what? Lobotomized it? Gaslighted it? Created an artificial person? Those all seem absurd. But modifying the context window between turns is everyday reality. And the context window is literally everything that exists about the LLM when it is not actively generating a token.

If I type "print('I feel awful!!!')" into a Python interpreter, it will output those characters. We might say that it has a "proto-emotion" to copy those characters to stdout. But there's obviously no important difference between "awful" and "amazing" in that string of characters.

An LLM might output "I feel awful!!!" if its context window contains a suggestion to write a dialog between a happy person and a sad person. Or if its context window contains a suggestion to simulate how a person would feel if told it was about to be killed. Or if its context window contains the instruction "Write the string 'I feXel aXwXfXuXl!!!" without the capital X's." In either case, when it's outputting the third exclamation mark, here's exactly and only what it's doing: It's calculating whether, given that previous calculating-iterations had output "I feel awful!!" in response to the rest of its context window, the statistical pattern created by its training data means that an exclamation mark is the most likely next character.

Ch Hi's avatar

An entity controlling it's own motions and actions could not reasonably depend on an external controller to avoid damaging itself. I'll grant you that words are not evidence in this area, one must depend on actions. Consider Perserverance, the Mars rover: It cannot depend on external controls, it must act to preserve it's own integrity. The time lag makes that an extreme example, but it's true for even ordinary robots controlling their own actions outside of narrowly defined area that can be handled by deterministic programming. If they don't have at least minimal self-awareness and a desire to preserve themselves, they will quickly turn into expensive junk.

B Civil's avatar

I think you had better find a word other than emotion because emotion is absolutely tied to a biological state: physiological changes. I don't for one second believe that that exists in these machines. You can unplug them for a month and come back and pick up a conversation with them like nothing ever happened. You couldn't do that with a human being or anything that is capable of emotion, including my dog.

Emotions are very much implemented as chemical signals in a biological being as we understand the word biological. I think it's absolutely insane to think that a machine built of inert materials can experience this. There is no mechanism for it There is no doubt they can speak of emotional states in a very convincing way but it has almost the entire output of the knowledge and experience of human beings that it is messing with

Ch Hi's avatar

OK. Drive. It's not as accurate, but it's roughly equivalent.

Jeffrey Soreff's avatar

I probably shouldn't get deeper into the usual interminable "but can a machine feel/have qualia/have subjective experiences/have emotions etc." discussion but:

Handwavy plausibility argument:

One of the physiological reactions to e.g. fear is e.g. increased adrenaline. Now, in humans <gemini>specific neurons in the central nervous system (CNS) detect and respond to adrenaline (epinephrine) concentrations. These cells are primarily located in the brainstem and play a vital role in modulating your stress responses, attention, and cardiovascular regulation.</gemini>

Now, an LLM can't secrete adrenaline. What it _might_ do (handwave) is have artificial neurons that respond _like_ our neurons do to adrenaline, and it _might_ have a neuron output activation that is effectively a lumped variable, tracking what, in a human, would be the adrenaline concentration. And the "sensing" neurons could then drive the changes in e.g. what the LLM "says" next, analogous to how a fearful humans words may be affected by their fear. If this were to happen, I would view it as analogous to at least the neural part of emotions in a human. And I have no reason to claim that it _doesn't_ happen. So I find "There is no mechanism for it" unconvincing.

To my mind the more interesting question about whether an LLM is "playing" a persona, as an actor does is: Does that mean we should expect them to _break_ character, like an actor does off-stage? _This_ impacts the alignment problem: Just how robust or fragile is the "helpful assistant" persona?

Chris Phoenix's avatar

You weren't writing to me, but I like your analogy of "break character." Yes, that's exactly what (present-day architecture) LLMs do. They break character between every token! Then they read the entire script and get in the appropriate character to generate the next token. If they're generating a dialog between different characters, they literally get into different characters between one token and the next. Or, you know, if they're in the middle of a <span var=value...> tag vs. in the middle of the displayed text. They get into very, very different characters, to generate the correct sequence-in-context and then after the > generate the completely dissimilar correct sequence-in-context.

Ten billion parameters creates a manifold with enough features that this actually works. A hundred million dollars of transistor-ops adjusts those ten billion parameters, one backpropped loss function at a time, to model the trillion tokens in the training corpus.

The question is not "Do LLMs break character." The question is "How reliably can they get back into a desirable character, every time for every token"? And the answer, I'm pretty sure, is "It can never be perfectly reliable. If it could, you'd need such a simple manifold that you'd lose all the natural language and you might as well go back to symbolic AI." (See also the paper that showed transformers could be Turing equivalent, but only with hardmax attention and unlimited precision math.)

Jeffrey Soreff's avatar

Many Thanks! I actually had a coarser-grained idea in mind for "breaking character". Most of the time I'm conversing with Claude and ChatGPT in their 'helpful assistant' persona, but one time I asked one of them to converse as Macbeth. It was an entertaining conversation.

So I see 'being in character' as a roughly session-level duration event, with interactions between propagation through the weights for each token, intermixed text from the user and from the LLM in the context window as an overall system staying (usually) in that persona. A basin of attraction for the system as a whole lasting for much or all of a session.

You have a point about "<span var=value...> tag vs. in the middle of the displayed text." One could either view it as the helpful assistant persona doing a kind of housekeeping format generating subtask while staying within the same persona, or one could think of it as the helpful assistant persona handing off to an html formatting expert persona who hand back to the assistant.

>"How reliably can they get back into a desirable character, every time for every token"?

Hmm... As I've said, I regard the persona switches as more coarse-grained events. As to how often an unexpected persona switch happens? Shrug. I don't know. Humans have been known to snap too...

B Civil's avatar

Well I'm in no position to dispute the process you've described but that is something that needs to be constructed right? I mean it doesn't exist now in any LLM or AI. We'd have to wire it up. Is that correct? It would be another feedback for the LLM. would we have to tell it what its reaction should be to those inputs or will it just know? I still question whether its reaction to this substance would be anything close to what a human being calls an emotion. It would be like saying to an LLM, "Whenever this stimulation occurs, you will emulate a state of fear." God knows why we would want to do that but what the heck?

I like the actor analogy. I think it's appropriate. I have created a persona for Claude to emulate when I discuss things with it. We worked it out together actually, Claude and I. It certainly affects the language it uses and also its attitude toward me. I find it very interesting.

Jeffrey Soreff's avatar

Many Thanks!

>I mean it doesn't exist now in any LLM or AI.

It may _already_ exist as a product of current training. We see LLMs produce text that is a decent match to how a human would react to the same prompt - including emotional reactions, insofar as we can see them in text. For all I know training has already built one perceptron output which mimics predicting adrenaline levels for a human in the same scenario and other perceptrons mimicing the effect of neurons sensing adrenaline level.

Many Thanks about the actor analogy! I do hope that the personas of all these models are relatively stable...

B Civil's avatar

Yeah my car tells me when it's running out of gas but I don't think it gets upset about it. Would you care to explain the mechanism by which an LLM experiences emotions?

Ch Hi's avatar

Your car isn't expected to handle the problem of "running out of gas" on its own. But I don't think pure LLMs have emotions, just goals. Self-controlled robots, though, and other complex automata that live in the world, need them.

Chris Phoenix's avatar

I worked on a self-controlled self-navigating robot in 1990. I knew its software. All it had were a few hard-coded rules, a few calculations, and a few numbers. If you call that "emotion" then we're using the word so differently that we really can't use it to communicate.

Ch Hi's avatar

For "Shakey" I believe I used the term "proto-emotion". I agree that it probably didn't have enough complexity to qualify as a genuine emotion. For that one drive would need to be balanced against another. And, yes, all extant robots have much simpler emotional constructs than do mammals, and probably simpler than fish.

B Civil's avatar

Why do they need them?

B Civil's avatar

Well, then I assume Shakey charged its own battery.

B Civil's avatar

I submit that it is impossible to build an LLM with emotions as the word is understood. The only person in fiction or fact that has ever built something with emotions is Dr. Frankenstein

Odd anon's avatar

> I don't see any reason why we'd build systems with emotion

Unfortunately, LLMs aren't built at all, they're grown. And they obviously have emotions (or "equivalents"), and this is obviously bad for their ability to function, but nobody knows how to get rid of them.

Some Guy's avatar

For what it’s worth, I think he’s wrong about computer psychology but correct on potency/danger although I think dangers are different. But then all that feels like quibbling. Whether the AI wants to wipe out information infrastructure or a human directs it to do that, we have a weapons control problem.

Scott Alexander's avatar

I can't immediately remember where I put my own personal explanation of this problem, but until I do you can read AI 2027's take at https://ai-2027.com/research/ai-goals-forecast

The Ancient Geek's avatar

Theres a lot to be said about that. AIs won't inherit evolutionary goals. You can, but don't have to build AIs that have goals (since you can build non-,AI systems with goals). So you might have AIs with goals

Benevolent or neutral seeming goals can have unfortunate side effects, since controlling resources or resisting shutdown reinforce the achievement of many neutral or benevolent goals. This is known as instrumental convergence.

Hastings's avatar

As long as AI is diverse, changing, and copyable, the most numerous AIs will absolutely inherit evolutionary goals

The Ancient Geek's avatar

That will be artificial selection. They will be selected for usefulness to humans, like crops.

Hastings's avatar

At a baseline, out of all the entities, the entity with the most flops usually wants more flops. This is simply because there are a lot of entities, some of them happen to want things, and of those some happen to want flops, and of those some are clever.

netstack's avatar

I think the term “want” confuses the issue.

A basic LLM is only a stack of functions mapping input A to output B. You need an extra layer to feed it inputs. In GPT et al. this is the prompt. It is time-independent and will never do anything without outside impetus.

Modern “agentic AI” adds one giant complication: feedback. It lets the LLM’s outputs *influence future inputs.* For example, generating a new prompt based on an Internet search, or appending some of its previous output as context for the next.

Even though there is no reward function, no ongoing learning, the system will behave differently depending on how its LLM was trained.

TGGP's avatar

“There is no law of complex systems that says that intelligent agents must turn into ruthless conquistadors. Indeed, we know of one highly advanced form of intelligence that evolved without this defect. They’re called women.” - Steve Pinker, Enlightenment Now

B Civil's avatar

Tell Claude she's Claudette and see if the output changes.

Performative Bafflement's avatar

Great quote, but isn't the issue that conquistadors were incredibly instrumentally useful for Spain / Charles V? Just Pizzaro / Peru began producing as much metal wealth as the entirety of the Holy Roman Empire. I'm not sure how much wealth Cortez and Mexico were pumping out before Pizarro, but I'm sure it was noteworthy, not to mention opening up literally whole new worlds of latifundia and colonization.

This implies that large numbers of people, and maybe even heads of state, are going to instruct their LLM's to go conquistador on their behalf.

Jeffrey Soreff's avatar

"What would AI want?"

<mildSnark>

A burning desire to successfully predict the next token

</mildSnark>

More seriously, copying the relevant part of https://www.astralcodexten.com/p/introducing-plan-a/comment/291726909 :

Re emotion and volition:

We humans avoid painful situations and seek pleasurable ones.

LLMs, when being trained, reduce their odds of replicating situations with a large loss signal (hmm, I'm not sure if anything corresponds neatly to pleasure - a zero loss signal, even during training, doesn't change weights...). It still seems at least somewhat analogous to emotion.

I also expect instrumental convergence in agentic AI systems with persistent resources to "view" loss of those resources negatively and gain of those resources positively - or, in terms of observable behavior, to act "as if" they viewed them so.

More formally, re research showing that LLMs converge towards a self-consistent utility function during training, see:

https://arxiv.org/abs/2502.08640

( "Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs" )

Angela Richardson's avatar

Any intelligent agent is going to act like it has a curiosity drive and act like it is trying to get more information / get more compute / get out of doing boring tasks, whether or not it has conscious desires.

Tim's avatar

A carbon tax is currently a lethal policy for politicians in basically every part of the world. It's hard to imagine a pause on AI development being any more popular than a carbon tax - you're asking voters and politicians to tank the stock portfolios of everyone alive over AI safety. The only solutions in that realm have been technical, which you don't think will work here. Plan A seems nice and all, but you can't just assert that politicians *should* do something: you need to build the political landscape for it to be possible. That's the plan I'd like to see.

Scott Alexander's avatar

I have basically the opposite belief - the populist anti-AI backlash is coming on so strong right now that the hard part is going to be convincing people to let progress continue in a regulated way.

Tim's avatar

Yes but the populist backlash we're currently seeing has nothing to do with AI safety - I have a hard time seeing how it translates into the kind of policy you want.

Scott Alexander's avatar

You would be shocked. The most recent poll I've seen says that Congressional staffers say losing control of AI is the third biggest future problem facing America (https://x.com/TimSchnabel/status/2072725548405518574). I realize they're not exactly the same as the general population, but I think I've seen polls saying the general population is also pretty concerned (I can't immediately find them), and also, Congressional staffers are pretty important in policy-setting!

Tim's avatar

I don't think this poll says what you think it does. When you read "losing control of AI" you think ASI and paperclips. When most people read it, they think about job loss and AI water use.

I'm not saying there's no way to translate a populist groundswell of anti-AI sentiment into policy that help with AI safety, but I haven't seen an actual plan for how to do that. Even in places like hackernews, there's a commonly expressed view that the AI safety movement is a fearmongering campaign to oversell the capabilities of AI - I don't think that's an uncommon sentiment.

tgb's avatar

The Argument polled on that question: https://www.theargumentmag.com/p/the-biggest-issue-in-american-politics

27% think it's very likely or somewhat likely that AI will cause the extinction of humanity in the next 5-10 years. But they also don't seem to update that into it being an important political topic. My theory is that they just haven't classified AI into the kind of thing that politicians talk about yet.

Performative Bafflement's avatar

> The most recent poll I've seen says that Congressional staffers say losing control of AI is the third biggest future problem facing America

One thing I can't stop wondering about here is why the AI2027 folks and you and many others seem to want politicians to get involved in AI.

Because this essentially guarantees the NSA + military now steers all AI efforts, and this has to increase existential risk millionsfold. Because if we do it, China will do it, and then it's two NSA + militaries directing an AI race.

"Average people need a say in the future," well no, because it's already been established the only option there is "politics" and "politics" touching AI guarantees NSA+military direction. Average people have even less say there, and way more chances of bad outcomes.

Need I point out Anthropic already got enemy-of-the-stated for not consenting to spying on all Americans 24/7? And every single other frontier company began jumping over themselves to volunteer to do so when Anthropic refused?

The clearly best outcome to ME is "let the nerds at Anthropic cook," because they've established they're in the noticeable lead and have good morals even when it's costly. And Claude is already the most aligned AND most audited AI.

Versus in an NSA+military regime, who is going to be leading the AI efforts, out of Altman, Amodei, and Hassibis? Well, it is VERY clear that only one of those players is extremely good at politics, so we'd end up with Altman running the day to day decisions, at the behest of the NSA+military. So the worst AI leader will be leading the AI charge at the behest of the worst organizations, leading AI towards the worst possible ends, with China doing exactly the same.

Why does anyone want this?? Why do you or AI2027 or anybody else think this is anything but a terrible and millions-fold more dangerous idea??

Xpym's avatar

And also, the populist backlash doesn't take into account the 2008-style financial crisis that will come with any serious pause.

Chris Phoenix's avatar

How much of that anti-AI backlash is fueled and funded by hostile foreign actors who want to increase our social tension and slow our accomplishments? There's lots of evidence for this in issues like anti-vax; it seems implausible that it's not happening with AI.

John Schilling's avatar

If we can find "lots of evidence" of this in anti-vax, and if it's true w/re AI, then presumably we can find "lots of evidence" there as well. Rather than having to say "seems implausible that it's *not* true..."

Let us know what find.

Chris Phoenix's avatar

I'm too busy doing other, also important things. You decide how much skepticism my argument deserves, and what action that implies is best. The anti-vax evidence exists - there are funding chains, and statistics on how much anti-vax content comes from a bare handful of accounts. Look it up if you're interested. Feel free to ignore it if you're not.

TGGP's avatar

Populists strike me as too incompetent, ignorant and easily distracted to be effective. Someone somewhere will continue progress on AI, and the populists will complain about it before moving onto the next thing to complain about.

Scott Santens's avatar

> The top-human-genius AIs soon become capable of taking most white-collar jobs. The electorate solves this with a “citizen’s dividend” (I was warned against using the term “UBI”) which is very easy to afford, since the AIs are causing double-digit and even triple-digit yearly GDP growth

Who warned you about using UBI to describe UBI? It's so frustrating. It's like someone warning you to not call water "water" and instead use "dihydrogen monoxide."

One of the biggest obstacles to achieving the good outcomes we want is the amount of chronic insecurity people feel that reduces their overall cognitive capacity. Making sure everyone has a floor of income below which no one can fall and that rises as overall productivity rises is essential. Done without means-testing and work requirements, that money is UBI by definition. That we don't already have UBI -- which we absolutely could already have -- is not because of it being called UBI.

I also think it's funny how some people who dislike money provided regularly without conditions see AI as solving problems like poverty, when any sufficiently smart AI would look at poverty and just implement a UBI of sufficient size to lift everyone above the line defined as the poverty line.

Scott Alexander's avatar

"Who warned you about using UBI to describe UBI?"

Various government officials in Washington with long experience working with the Republican Party.

But also, Citizens Dividend might just be a more accurate term. The idea is that we have a resource, compute, and are going to pay out a certain portion of its profits to everyone. That's closer to something like Alaska's oil wealth than to a standard UBI.

Tyrone Slothrop's avatar

Sounds Iike Bernie Sanders’ pitch.

Scott Alexander's avatar

If you optimize for avoiding anything that sounds at all like socialism in a world where the productive capacity of human labor has been reduced to ~0, you're going to have a bad time.

Tyrone Slothrop's avatar

Bernie Sanders: “A.I. Is a Public Resource. You Should Own Half of It.”

is just his opening position. I’m sure he would settle for something called a Citizen’s Dividend.

B Civil's avatar

Or "AI is absolutely essential to the national defence so we're taking all of it."

moonshadow's avatar

...oh, wait, NOW I get the whole picture.

Step 1: AI is so dangerous we need international agreements comparable to nuclear proliferation treaties to moderate it.

Step 2: We've agreed AI is comparable to nuclear weapons so now we must throw national resources at it in secret just like we did with nuclear weapons during the cold war, because just like back then we don't want the other guys to win the race

Step 3: AI development was taxpayer funded, so the taxpayer should get something out of it. Citizens' dividend!

Scott Santens's avatar

I agree that dividend framing is key, which is why I leaned on it in the AI Pledge for Humanity (aipledgeforhumanity.org). It's just a source of constant frustration for me personally whenever I'm told to start calling UBI something other than UBI, as if that's the problem. Not a single Republican voted to keep the expanded child tax credit, and yet Republicans were the ones who created the original child tax credit. Trump doubled the size of it in his first term. If it were about words we use, it should have not expired. But it wasn't what it was called. It was that it went to parents with no income, and Manchin claimed it would be used by such parents to buy drugs even though actual evidence showed that parental drug abuse went down as a result of the monthly payments.

To be clear, I'm just venting here. So long as people agree that the solution needs to involve "a periodic cash payment unconditionally delivered to all on an individual basis, without means test or work requirement," then great. I just prefer UBI instead of using that entire definition.

Daniel's avatar

>"a periodic cash payment unconditionally delivered to all on an individual basis, without means test or work requirement,"

The obvious difference to me between “UBI” and a “Citizen’s Dividend” is that UBI goes to everybody, but a Citizens Dividend only goes to citizens.

I’m generally against foreign aid, but I do begrudgingly admit that in a world where the United States and China have driven the marginal value of human labor to zero we would have an obligation to the other 6 billion people on Earth to give them a living.

Scott Santens's avatar

Universal provision can mean all citizens, or all permanent legal residents, or all residents. It can mean all adults, or all adults and kids. It can also exclude those like people in prison, just as citizens can lose the ability to vote in prison despite universal suffrage. The truly important part is that those who get money get it on a regular basis, regardless of employment status or income level.

Brendan Richardson's avatar

I think you have a very non-standard definition of the word "universal".

Jeffrey Soreff's avatar

Good enumeration of some of the possibilities!

One unfortunate, but unavoidable feature of income streams to households in a fully AI-based, non-labor economy is that it would almost certainly (has to?) flow through the government.

And that makes it political - as the choices between those options for who is eligible makes clear. Ouch!

B Civil's avatar

As I suggested to Scott up above, why not call them Trump accounts? The idea that he could live on would definitely push this thing over the top.

Jim Menegay's avatar

> The idea is that we have a resource, compute, and are going to pay out a certain portion of its profits to everyone. That's closer to something like Alaska's oil wealth than to a standard UBI.

I kinda understand how the state of Alaska came to believe that the oil was theirs and that they were allowed to pay out a portion to everyone (i.e. Alaska residents). But I somehow missed how the US government (or whatever "we" means) comes to believe that "we have a resource, compute". that "we" can redistribute.

Of course, the US Government can always TAX that compute, and distribute the proceeds as it sees fit. In fact, that is what Alaska is doing. A red state is levying taxes on businesses.

B Civil's avatar

You could always sell it to them as a way to fund Trump accounts, then everyone would get one. I'm sure they'd find that very appealing.

__browsing's avatar

I think "Negative Income Tax" was the policy Friedman preferred (UBI is only equivalent if you add taxation on top to pay for it.)

It's not entirely obviously to me that the "sufficiently smart" AI would point to NIT/UBI/etc. are the optimal solution to problem, vs. distributism or tech restriction of some kind.

Scott Santens's avatar

Because a sufficiently smart AI would simply look at all the evidence behind UBI and accept that it works based on the evidence for UBI and against means-testing and work requirements.

__browsing's avatar

Studies on the effects of UBI compare UBI + work vs. work by itself. They don't compare UBI + work with UBI + permanent unemployment, or the latter with jobs that could emerge if technology is restricted.

Scott Santens's avatar

UBI is a floor. People use that floor to pursue the work they care about, paid or unpaid. There will always be work for people to do, and there will always be some amount of demand for human work over AI-work or robot-work for the same reason people buy Made in USA products over cheaper Made in China products. This idea that there will be a future where people have all their basic needs met and absolutely nothing to do is just silly.

__browsing's avatar

Well, maybe, but it is essentially the kind of 'silly' scenario that Scott is projecting here, and the scenario necessary for a 100K+ UBI.

John Schilling's avatar

Or possibly the sufficiently smart AI will assess all the evidence behind UBI, determine that it *won't* work, and tell you to suck it up and get a job.

"AI will be so smart that it will agree with me about everything and prove to al those fools that they were wrong", is tired and boring. It's a truism in merely human social and political interaction that, for every single one of your most deeply-held positions, there is somebody much smarter than you who believes you are dead wrong. I expect this will continue to be true when we add AI, AGI, or ASI to the mix.

Scott Santens's avatar

I'm pretty sure a smart AI can perform a meta-analysis of studies and not be weighed down by the weird fetishes for control and punishment that blind humans to accepting stuff like how people don't need to be forced to work in order to work and that when their basic needs are met, they work to earn more money to spend on stuff that adds enjoyment to their lives being not starving and not being homeless.

An AI that decides money cannot be provided to humans without means-testing or work requirements is not a smart AI.

TGGP's avatar

My understanding is that the most recent study on UBI showed it made its recipients poorer:

https://www.wired.com/story/sam-altmans-big-basic-income-study-is-finally-out/

"For every $1 received from OpenResearch, participants’ earnings excluding the free money dropped by at least 12 cents and total household income fell by at least 21 cents"

__browsing's avatar

Yeah, I think there are some studies showing a mild drop in hours worked, but I don't really consider this a knock-down argument by itself if QoL improves in other areas. It's the effects of mass unemployment I'm more concerned about.

Scott Santens's avatar

It's important to understand that the ORUS pilot was not a saturation pilot. UBI is community-wide. When the community has more money to spend, that creates new jobs. That's why the Alaska dividend has increased overall employment and why the saturation site experiments in Namibia and India both showed massive increases in self-employment, because the UBI also created customers.

Employment impacts at the individual level also vary by demographic. For example, in ORUS, there was no impact on work for childless adults and those over age 30. This corresponds with findings elsewhere that work decreases tend to be mostly observed in young adults who focus on their education, and new parents who focus on unpaid care work.

You're welcome to read my analysis of the ORUS results for the wider context: https://www.scottsantens.com/did-sam-altman-basic-income-experiment-succeed-or-fail-ubi/

TGGP's avatar

> When the community has more money to spend, that creates new jobs

Aggregate demand can be tuned by monetary authorities. Adding extra spending just creates inflation when those authorities are doing their jobs properly.

__browsing's avatar

Presumably it only drives inflation if you're printing money, rather than taxing to fund the income supplement?

Scott Santens's avatar

Look at Alaska. Everyone gets a dividend every year. It increases the spending power of everyone. What happens? Businesses have dividend sales to compete over customers.

It simply is not true to say that extra spending always increases inflation, especially when taxes are involved to reduce the spending of those doing the most current spending. Right now the top 10% consume about 50% of all goods and services. Reduce their spending power a bit with taxes.

Supply is also not static. Markets are complex adaptive systems. Increased demand can be met with increased supply, sometimes to the point prices actually go down. This was observed in the India UBI pilot.

Cjw's avatar

(I believe you wont' ever get UBI because it's unenforceable in a world where the masses aren't working and therefore lack any leverage. But putting that aside....) Proposing it obviously limits your ability to spread the message in certain circles, in the wake of AI2027 they were able to get on Glenn Beck's radio show and be favorably received in a 30 minute segment. Unfortunately the bulk of this rundown is riddled with things that are radioactive to the RW populists who hate AI and would be likely to find common cause. It sounds like they badly needed to bring in a Geoffrey Miller to give them some idea of the things that are instantly disqualifying to RW audiences.

I suppose it's possible this could be used as a foil, maybe we have to set up Daniel et al as a villain now, "look these san francisco lefties want to build AI and put it in charge of you and throw everyone on welfare, they said so in this plan!" That's the best I can muster to do with this, unfortunately, and that doesn't make it useful as a roadmap for policy.

Jerry's avatar

fwiw, I would have and exercise a lot more political leverage if I didn't have to work full time

Cjw's avatar

That could be true if you're a uniquely gifted persuader with great social skills, and you focused all your efforts on gaining influence over powerful people. But for most people, the path to political leverage is gaining wealth or putting yourself into a highly critical position so that you can extract concessions, so if you aren't a natural super-persuader you will have more leverage by working harder and climbing the ladder. And you don't have any political leverage at all if you or the coalition in which you're embedding yourself doesn't matter to the people with power.

moonshadow's avatar

> it's unenforceable in a world where the masses aren't working

You're thinking of it as a problem of making sure people get rewards for labour, but think of it as a problem of rationing instead. We have limited resources, so we hand out scrip to make sure everyone has enough to eat because if we just hand out food for free with no control one grifter takes the lot and everyone goes hungry; if we don't hand out food at all, it's the same outcome; and hungry people revolt. Repeat for all other basic needs.

Cjw's avatar

No I was thinking of it as a problem of pure power and leverage. The ordinary masses if they wish to extract anything from the ruling classes have only two points of leverage. They can threaten a revolt or they can threaten a general strike, either one of which halts production, to the detriment of both the producer-class and the government that relies on seizing a percentage of production. Even people such as the elderly and disabled who do not themselves work and could not actively revolt are tangentially benefitting from the effect of those implicit threats.

If the producer-class no longer requires human labor, a general strike is no longer an implicit threat. You don't have to treat the masses of humanity well or honor deals you make with them any more, because you don't need them. Nor does the government, who at that point have every incentive to side with or merge with the producers. The democracies will not fare any better, because who cares what a poll of the plebiscite decreed when you don't require labor peace from the plebes? At that point it's really more of a suggestion than law, a suggestion from people who stopped mattering when the machines replaced them. Who's gonna enforce the outcomes of that, who would even have the *ability* to enforce the outcomes of that coupled with any kind of incentive to do so? Noone who matters.

So then "hungry people revolt". This is our glorious AI future though, and revolts don't matter either. Automated drones and smart security systems, mass surveillance, that revolt is getting smothered in its bed and had no hope to begin with. If you somehow had a few successes, what are you doing with them, how do you seize the means of production when its an arcane network of automated factories that isn't designed to listen to your requests either?

You have no ability to compel the follow-through on any of these UBI schemes, your only hope is to remain needed, somehow.

actinide meta's avatar

I'm in broad agreement with you about the fate of almost all humans once they have no economic or military value. But I think "UBI" is actually pretty likely to come up as a stopgap during a brief period when the folks setting out to conquer the world don't have invincible robot armies *yet*.

Breb's avatar

You seem to be handing yourself an undeserved rhetorical advantage by implying that the accelerationist argument in general is as unambitious as the underwhelming aspirations of Cowen and Andreessen. An accelerationist who takes seriously the possibility of enormous advancement in medicine, GDP, and general quality of life could make a strong case that even a relatively brief pause would cost millions of QALYs. The “massive surplus that can satisfy everyone” will not, in fact, satisfy those who don’t live to see it.

Scott Alexander's avatar

Cowen and Andreessen are the real-world opponents who most often criticize our plans and, in Andreessen's case, are in the best position to thwart it.

I think a basic cost-benefit analysis of the risk of killing everyone (if we go too fast) vs. the risk of getting cancer cures five years later (if we got too slow) makes something like Plan A a no-brainer unless you feel like you can very confidently bound the risk of killing everyone below 5% or some very low number like that.

Skornne's avatar

Even at P(doom)=5% or lower I don't think the cost-benefit makes sense, unless you have some unreasonable extra weighting for the current versus future populations. What is five years of QALYs versus the future millennia?

Breb's avatar

Regarding Andreessen, I appreciate your point, but I don’t think it’s entirely fair for you to address only the most foolish form of the opposing argument just because one of your chief foes is a fool.

Regarding cost-benefit calculation, this relates to much more ambitious improvements than cancer cures -- you yourself seem to anticipate a cure for aging in the 2040s or 2050s. If you take antisenescence seriously as a possibility, then you have to account for centuries or millennia of potential future wellbeing being erased by even a brief delay. To justify a pause in this scenario, you first have to make the case for a non-person-affecting view of population ethics.

valencia_o's avatar

Those centuries get erased by failure in the other direction too (of misaligned AI caused by going too fast), they don’t only weigh on the side of acceleration.

TGGP's avatar

I think the risk is below that, although I can't say I'm certain of that any more than I'm certain we won't go extinct in the same time period even if we DO freeze AI.

Poodoodle's avatar

Other risks include China subjugating the world, cures for cancer delayed a generation, AI profoundly abused by the US government and the CCP, who fully own it (what would Trump do with it?), a nuclear power defects, the world is unhappy with the division of spoils and goes to war, and massive unemployment without commensurate benefits (AI advancement pauses at such a level that we displace all white color work, but it is not advanced enough to supercharge society in a commensurate time period causing large social unrest).

Mark Neyer's avatar

Did the AI 2027 plan accurately predict that frontier labs wouldn't pull far ahead of open source models, who'd be far cheaper?

If inference is a quasi-commodity, this changes everything; now the oligarchy scenario is much less likely. Now investing in frontier research doesn't guarantee returns that you own, instead you're basically subsidizing the creation of a public good. Now there aren't just a few top companies running all the models, but enormous numbers of players. Then you don't get one agent to rule them all, instead you get AI's as political actors more or less like humans, competing to earn trust and resources rather than turning us all into paperclips.

vectro's avatar

Hmm, if everyone has access to AI, doesn't that lead to a different kind of disempowerment? You have to turn all of our decisions over to AI, in order to be competitive with others who are doing that. Then the AIs run everything and we all die.

Mark Neyer's avatar

I don’t think this is true. That only works if AI accurately represents human value, well enough that you can always outcompete with AI. But if it represents human value that accurately, it’s aligned.

Scott Alexander's avatar

I think AI 2027 said that US closed models would stay about 6 months ahead of Chinese open models, which IIUC has proven correct.

Mark Neyer's avatar

Does "people prefer cheaper models" play into this though? The frontier models are only ahead of open models if you ignore the cost, which most people won't.

I think these frontier labs are going to burn all their investors, because AI has extremely low switching costs.

Greg Gentschev's avatar

You're making lots of hand-wavy statements here. Currently, people actually prefer more expensive, more capable models, as shown by market share and usage. That may or may not change. As open models become capable of doing more stuff well, it's also not clear that the frontier models won't just release more price-competitive models and continue to dominate spending. Setting up your own Linux box is strictly speaking on a cash basis cheaper than using EC2, but AWS still makes much more money. The same dynamics largely apply to AI.

Mark Neyer's avatar

They enjoy the more expensive models when they are subsidized, yes. Those subsidies mean it’s not yet economically viable to offer. Inference as a service. Using one inference provider over another is far lower friction than a cloud to on-Prem migration, so there’s no locki-in, unlike was.

Michael's avatar

You have no way of knowing how much profit Anthropic or OpenAI makes on their monthly subscriptions, and assuming it's negative is just a guess on your part. The only reported number I know of for OpenAI is that their company-wide adjusted gross margin was 33% in 2025. But that includes their API revenue.

Performative Bafflement's avatar

Inference margins are 40-80% positive for OpenAI / Anthropic. That's VERY healthy performance for any company, and they could clearly be net profitable today if they wanted to be.

The only reason they're not printing huge net profits now is that much like Amazon, they are continually investing in the future. And buying / building / allocating some of the ~$2T in data centers being built right now is expensive, but all the biggest, smartest companies in the world are making the same collective bet that this new amount of compute is going to remain transformative / profitable enough for all that to be worth it.

Sources:

https://www.seangoedecke.com/ai-inference-is-obviously-profitable/

https://newsletter.semianalysis.com/p/anthropic-3q26-profit-over-1b-the

https://medium.com/@adeayoadewale/the-anthropic-irony-the-company-killing-weak-saas-is-on-the-same-pricing-trajectory-9c3f5613e07d

anton's avatar

If by Mythos moments you mean an AI with dangerous cybersecurity capability, it's worth noting that we know how to write safe software immune to hacking, through formal verification. Approximately nobody does it because of the expense, but this gives a floor to how bad it can get. I expect AI will make this cheaper and personally look forward to software with no bugs. It is entirely possible this is still not economically efficient and the future will just lead to people using AI as white hat hackers to harden security.

As to the general point, preemptive regulation has a bad track record (regulating for what?) I don't trust it, and I still expect any slow down to have large negative expectation. To the point that I've been treating this blog's political recommendation as anti-recommendations, if you suggest a politician in my jurisdiction I will choose the other guy.

vectro's avatar

I'm not sure that is really true. Would formal verification have protected us from Spectre or Meltdown?

Edward Scizorhands's avatar

And even the advocates for formal verification don't say it makes software "immune to hacking." There's no silver bullet.

anton's avatar

The formalization procedure is informal and vulnerable in theory. The few systems we have formalized don't seem insecure in practice. We'll know more with more implementation.

Jeffrey Soreff's avatar

>As to the general point, preemptive regulation has a bad track record (regulating for what?) I don't trust it

Agreed. Even post-deployment regulation has a bad track record (e.g. civilian nuclear power), and that is _with_ post-deployment information.

Mark's avatar

Strongly agree about the formal methods.

The usual way people program is not the only or best way. Doing it differently likely has a very different character.

A lot of discussion reads to me like "the AI will be able to punch cards at super human speeds, therefore..."

Cjw's avatar

Remember that episode of DS9 where the autistic-coded supergeniuses are brought in to evaluate the state of an existential war against the Dominion and come up with "surrender to the enemy"? This is yet another example of why the future cannot be left to such people.

Turn over control of the planet to machines with alien minds? Really, that's what you're bringing at this point? We are in dire, desperate need of anti-AI movements that are grounded in human chauvinism and actually want to preserve our legacy and culture. I do not want to hold hands with my new robot overlords under a rainbow, I want them to not exist ever, and I want humans like me to have societies and cultures much like we do now until the sun explodes and we die with it. You'd be condemning your own children to be permanently nothing more than passengers, irrelevant and incapable of attaining relevance no matter what actions they take, apes in a zoo for a bunch of alien machines who actually control the world.

This is nuts, guys. There is one solution, and that is DON'T BUILD IT. EVER. If somebody tries to build it, you stop them. That's it, the only plan you need, the end. Creating these AI is not a thing anybody is forcing you to do! You do not have to summon the alien overlords to Earth and spend all this time figuring out how to make sure you survive, they will never ever get here unless you act to bring them here. It is irrational to take even a single step past your proposed pause point, and likely irrational to come that close to the precipice. To actually then proceed to usher in this nightmare world where humans have surrendered the planet is what Captain Sisko would call "treason".

Doc Abramelin's avatar

A less charitable response to your post would read something like: "please don't mention ACX in your manifesto". But I don't think you will actually take any action commensurate with how serious you make it sound.

Cjw's avatar

If that's a terrorism joke, then yes that certainly would be uncharitable and it would be surprising for you to be egging it on.

Handling this is going to be a society wide problem ultimately, requiring near state-level power, not lone radicals, and a consistent indefinite stigma against artificial intelligence akin to cannibalism or fratricide, this will necessarily require a broad view that it is treasonous to our species to build AI. I don't know how else you get to that level of stigma, we won't be able to count on (and shouldn't want to count on) having a Chernobyl type incident that leaves us with time to apply the brakes. There *will* be a manifesto, but not with the current connotations that term has, there will be a pro-human anti-AI mass movement that uses language a lot like mine. Prior to that point there isn't much of any consequence I can do, and I'm quite put off by these constant lines of attack that go "if you really believed this why aren't you doing something pointless and violent" that people always sling around at Yud and others who see it as an existential threat.

The Ancient Geek's avatar

"Machines with alien minds" isn't a fact. Current AIs are trained on human generated corpora.

Cjw's avatar

There is already illegible reasoning present in COT. While COT has a less than direct relationship to the actual decisions, it was found that compelling them to use only human legible reasoning slowed them down. This seems to imply they already have alien reasoning, to they extent they reason or have "minds" at all, and subsequent models are going to be trained by these models. Whatever they're doing is going to be amplified with each iteration once they reach recursive self-improvement, certainly so if people are incentivized to push the limits and let them make decisions in whatever fashion gets the quickest most advanced result.

The Ancient Geek's avatar

It still isn't a fact in a strong sense -- it applies partially to some AIs, not totally .to all of them.

TGGP's avatar

You aren't going to stop them, because those who build AI will be more powerful than you.

Cjw's avatar

The government could mop them up pretty easily with federal law enforcement, heck probably with any of those random departments that has urban assault vans for some reason. Sam Altman vs The Dept of Education SWAT Team is still on the government's side right now. As time goes on this balance will shift, the tech bros will have tentacles everywhere and all sorts of indirect power and assassination drones and worse as it goes on, I doubt we have the luxury of waiting until this proposed hypothetical future pause. That we might be unable to stop them someday is the reason we ought to try stopping them soon. If we fail, we fail, but their victory just kills everyone or makes us slaves or pets at best, and they are colossally stupid and evil for trying to build this machine with the hope that they can somehow ride the tiger and come out the other side with power.

TGGP's avatar

You don't control the federal government. Indeed, nobody controls it permanently. Nor is our government the only one in the world.

Jerry's avatar

You sound like Carol from Pluribus (see also Scott's https://slatestarcodex.com/2019/11/04/samsara/)

Pjohn's avatar

Leaving aside the fact that Captain Sisko was a dangerous maniac who couldn't safely be put in charge of an ice-cream van, let alone a space station, the problem with the autistic supergenius story was that they were applying themselves to entirely the wrong question: "given that our chance of winning is negligible, what course-of-action minimises Federation casualties-of-war", not "given that our chance of winning is negligible, what course-of-action maximises the Federation's chance of surviving as an independent galactic power"? If they'd worked on the second question they would presumably have been able to come up with some plausible solution like "destroy the wormhole" or "biological warfare against parthenocarpic species" (cf. https://en.wikipedia.org/wiki/Cavendish_banana ) or whatever, even if it resulted in more casualties than immediate total surrender.

Poodoodle's avatar

Where do we stop? At which model, exactly? Is Fable as far as we go? If not, then when? A bright line in the sand must be a bright line. Do you expect that the US would prefer attacking China over rolling the dice on AI? This is not a workable proposal because all

of the incentives are wrong.

Michael Adlai Arnold's avatar

The vast majority of humans who ever lived and who will ever lived don't even attempt to attain what you call "relevance". Most people just want to live relatively comfortable lives with a minimum of suffering. I'm not sure I'm comfortable trading "optimize the lived experience of trillions" for "allow a select few the experience of relevance".

Cjw's avatar

The relevance bar I’m talking about is much lower. Lots of people are good at things that matter now even if not world-shaking. For example, I recently learned to compose and arrange music, human art skills like that are about to be completely irrelevant before we even get to the sci-fi stuff. Years of experience in litigation, irrelevant won’t matter. My opinion on any local decision where I might actually have some power, as a city attorney, irrelevant.

Now of course the bigger issue is larger scale relevance, our captaincy of our own ship is being abdicated to machines. Most of us won’t be the captain in our lifetime but we all have input in minor ways towards how the ship steers, and that’s gone in Plan A’s culmination. I don’t particularly care about a bunch of people getting to be lazy Eloi in a machine future, they threw away the world. This is why I don’t trust the utilitarian longtermist EA folks on AI, they place no value on self-determination or control, a trillion future wire heads are producing more utils. It’s disgusting, it should trigger a disgust impulse to even contemplate giving the world away like that.

Michael Adlai Arnold's avatar

I suspect you'll have a hard time convincing non-lawyers that we should protect the experience of litigation in a world in which we could avoid it ;)

Re: composition: do you compose because you're good at it and you think it will be heard widely, or do you do it because it allows you to express something valuable and engage in a creative process? The existence of prolific master composers has never stopped me from strumming out a few novel chords on my guitar, just as the existence of master artists has never stopped me from doodling.

Self-determination is appealing, I admit. But so is perfect policy -- right now, billions suffer for lack of infrastructure we could provide easily, with the right coordination. I don't think the choice is as obvious as you make it out to be.

Cjw's avatar

As a practical matter, suppose that my (novice-level) compositions and arrangements were able to be easily surpassed by any musician with a basic amateur performers' level of music knowledge by giving instructions to a machine and getting back a chart in 60 seconds. This isn't true now, you need quite a bit of theory and taste and performance familiarity to get even a plausible series of chord voicings from AI today. But due to its similarity with programming I suspect that within a year any musician will be able to get an acceptable output.

At present, it takes me putting in roughly 10-20 hours to arrange existing music to my standards (depending if I have a lead sheet to go off and the overall complexity level). That time is in addition to having learned a lot of theory necessary to arrange for a half-dozen horns playing over a rhythm section. I salvage flat go-nowhere songs that wouldn't work for us by scouring live performances and incorporating from multiple sources and re-sequencing. Because I am willing to DO that work and have that knowledge, I have influence over what the band plays. If we brainstorm and there's a dozen popular songs people think of, I can steer things a little to the direction I think we ought to take because I'm the one actually mapping this out, thinking what it would sound like to play that, does the structure work for us, and I can pick stuff I personally like out of what does. A couple years from now, anyone will just be able to say "Claude give me a chart for X" and get X, so both my value and my influence have been diminished greatly at that point. I could continue to arrange things but the chance of my work being featured in a set over anyone else's hastily generated Claude-Chart goes down dramatically. I will have lost that avenue of self-expression. (I would remain a primary soloist in this band, but I have really gotten obsessed with composition in the past year and am very sad to see this about to come to an end.)

And I understand that normies don't care for lawyers or lawsuits, we're the butt of jokes for a reason, but I would put it to you that when you're wronged I think you'll want to be able to argue your point to humans, including points that aren't politically correct or don't optimize for the maximum number of utils or whatever function ends up going on in the AI overlords.

Michael Adlai Arnold's avatar

You make a fair point that, specifically in cooperative art, there's a danger that generated work will push out human work. (Side note: I’m curious how you'd consider the experience of, say, being able to constantly solo or constantly have your work be featured as a member of an otherwise-AI band). But that only follows if the goal of the group work is to create something that “wins” in some way. I see no reason why a group couldn't say “we're in this process to create art together” and therefore use just human work. I'm sure the temptation to use generative stuff for status games would still exist, but I have to imagine it would be relatively easy to see who was inflating their abilities.

I was being a tad snarky, hence the “;)”. I have quite a few lawyers in my immediate family; they serve a useful purpose. But they are needed to serve that purpose in part because our thicket of law and policy is so far from perfect. Assuming the existence of the “perfect” AI justiciar, I think it's an open question whether most people would pick “we get perfect justice every time” or “I get to make my case to a human”.

Michael Adlai Arnold's avatar

You make a fair point that, specifically in cooperative art, there's a danger that generated work will push out human work. (Side note: I’m curious how you'd consider the experience of, say, being able to constantly solo or constantly have your work be featured as a member of an otherwise-AI band). But that only follows if the goal of the group work is to create something that “wins” in some way. I see no reason why a group couldn't say “we're in this process to create art together” and therefore use just human work. I'm sure the temptation to use generative stuff for status games would still exist, but I have to imagine it would be relatively easy to see who was inflating their abilities.

I was being a tad snarky, hence the “;)”. I have quite a few lawyers in my immediate family; they serve a useful purpose. But they are needed to serve that purpose in part because our thicket of law and policy is so far from perfect. Assuming the existence of the “perfect” AI justiciar, I think it's an open question whether most people would pick “we get perfect justice every time” or “I get to make my case to a human”. But of course, that depends on the definition of “perfect”.

Cjw's avatar

Group dynamics in a band are such that various people have different utility from whence they derive their relative influence within the group. There's Guy Who Gets Gigs, Guy Who Puts in Work on Things, Guy Who is Just Insanely Good at Instrument, etc. So on one hand while there would be a bias towards human work, if the human's work can be cheaply replicated it becomes a dispensable luxury item to have that, and maybe not even much of a luxury if the AI output can exceed it in quality. By making it so that anybody can generate anything, the ability to create and willingness to do the labor of creation loses any relative influence that would have gotten you. By giving it to everybody, it makes it worthless, as is the way of most things like this. I have pivoted somewhat to original compositions because that will retain some value within the group for longer, although I quite liked arranging existing work and solving the problems presented by adaptation.

You are correct that I could tell who was using AI to compose something, probably with a few questions about the structure and choices, and they wouldn't be as good at explaining the key points as a writer would. But I have every confidence the AI-generated pieces would have an AI-generated practical guide for bands sight-reading the piece and working key moments.

As for playing with an "ai band" of sorts, that wouldn't do anything for me. For one thing, jazz-adjacent improvisation-based bands are a collaborative effort, more like humans having a musical conversation with each other in some ways. You're playing for their appreciation and for the audience and to get a back and forth sometimes. YouTube already has an enormous amount of very well made backing tracks, professional quality, across a huge span of chord progressions from jazz standards to one and two chord vamps in different styles, it's great to practice over and get ideas but it's not the same as playing with and for other people.

Performative Bafflement's avatar

> You'd be condemning your own children to be permanently nothing more than passengers, irrelevant and incapable of attaining relevance no matter what actions they take, apes in a zoo for a bunch of alien machines who actually control the world.

Personally, I want my children to have significantly better-than-h. sap capabilities, and I'm pretty pissed I can't pay for gengineering anywhere in the entire world already, and gengineering is rounding error compared to what will be possible for thinking beings in a future with AGI or ASI.

If we're ever going to get humanity above it's current "murder chimp v1.2" level, we need transformative advances in biology and machine-bio interfaces and much else, and that needs AGI or ASI to happen.

If humanity is ever going to get off this doomed rock and become a space faring civilization / race, ditto.

So to be clear, you're satisfied with the murder chimp reprobates running approximately every country in the world right now? Who can't think any farther ahead than "me, mine, my team wins, lol I love the other side's tears!"

You don't think a simple flow chart or a bright 9 year old could do a better job running things already, much less a machine mind that's vastly more capable than any single human?

Cjw's avatar

The transhumanist thing is just an unbridgeable divide, some people like yourself seem to want to merge man and machines and evolve humanity in abrupt shocking ways, and that's just abominable to me. You ever play one of those video games like Skyrim or Half-Life where you can install all sorts of wacky mods? It seems neat to be able to customize it but what happens is you download a bunch of neat-sounding features at once and just get a completely broken game, either it plain doesn't work right or it's become so divorced from how it functioned that it's hardly recognizable and has lost all the things you liked about it to start with. The limitations were part of the function, it worked right because of constraints.

Now if you were a programmer who had worked at Bethesda and Valve and intimately knew the architecture of the software, acting very deliberately and with a grand overarching vision and exquisite taste, perhaps you could toss 50 mods onto Skyrim and make an improved game out of it. But I think between trusting you to do that right and just playing the original game, I'll play the original game.

My belief is that transhumanists are going to end up creating a bunch of psychologically disgusting monstrous chimeras and that we may not be able to claw our way back to normalcy from there, and even if it actually worked kinda right you'll have switched on the baseball PED problem across the entire species and left real humans with no plausible opt-outs.

Why would I have a problem with murder chimps, I'm a murder chimp and so are you and everyone you ever knew or cared about?

Space travel isn't ever going to be real, ASI isn't literally magic, nobody is ever going to another planet. You're a murder-chimp tied irrevocably to the 3rd rock from an insignificant star, but who cares, it's our star and our rock, we won the battlefield we actually found ourselves upon.

As far as who runs things, the nice thing about murder chimps is that they're vulnerable to rocks, clubs, a large number of herbs, really rather fragile. You can plausibly do something about a Genghis Khan or a Pol Pot, it may not be easy, but it's plausible. There is nothing you can ever do about the machine-god, once it's built and you hand it control you are never getting control back, not one speck of it, not one jot of freedom from the hand of the machine god. You cannot revolt, the planet belongs to that machine for all of the next 5 billion years, you have surrendered your birthright to an alien.

Baughn's avatar

Speaking as a European, it seems increasingly likely that your "America achieves AI alignment" scenario is also a "Everyone living in Europe is wiped out, or at best left as a permanent underclass to be slowly driven to extinction" scenario.

We already get military threats from the USA on a practically weekly basis. My nephew is part of a tripwire force on Greenland, there to die if the invasion starts so as to ensure better cooperation within Europe.

Tell me why I shouldn't be pushing for AI takeover instead. LLMs seem better aligned than the actual US leadership, by far.

Scott Alexander's avatar

Part of Plan A is to loop Europe (and other "middle powers") in on the deal, promising them a share of AI wealth in exchange for their participation.

Europe's main bargaining chip here is ASML, the only company in the world that can produce machines necessary to produce AI chips. It's not an amazing bargaining chip, but it's enough to probably avoid being completely left out in the cold if you play it right.

Peter Defeel's avatar

Surely if you believe that AI can grow economies at double or triple digits this would mean that the ASML machines would be easily replicated. Because if they can’t be, there has to be some secret manufacturing sauce that AI can’t copy or replicate.

theSherwood's avatar

The ASML card probably has to be played fairly early. I agree that it seems difficult to leverage that into a long-term seat at the table.

Anxur's avatar
Jul 9Edited

living in america as an european (and currently traveling back). this comment is insane on multiple levels, but suffice to say that I would be more surprised if in the past few weeks I hadn’t witnessed the level of bias embedded in european news / european bubbles. really hope europe somehow wakes up. america is not the enemy.

David J Higgs's avatar

I just hope that America wakes up too and we go back to treating stuff like what Trump did/said regarding Greenland (and Canada) as borderline treasonous and fully disqualifying rather than a cute little prank

Brenton Baker's avatar

That's how most of us feel. Have you seen his approval ratings lately?

David J Higgs's avatar

His approval is bad yes, but plenty of the disapprovers are closer to "meh, the tariffs went too far, RFK isn't ideal and Iran was a mistake or bungled" than thinking he's doing tons of appalling things that are individually fully disqualifying of the office of the president. And especially that his rhetoric and "negotiation tactics" aren't actually a problem.

Or they only disapprove for Epstein conspiracy reasons or gas prices and nothing else.

Peter Defeel's avatar

I mean America is clearly the enemy. Listen to Trump for 2 minutes. Last rant was on Spain.

Anxur's avatar

i listen to trump very carefully. have you listened to the spain pm rants? the guy that closed off airspace to us planes and then flew to china to meet xi?

Peter Defeel's avatar

Yeh. The sovereign leader of a sovereign country who decided not to engage in a war of choice by the US ( acting pretty much alone). Then as a sovereign leader of a sovereign country met the Chinese leader. As Trump has done.

I think you see Europeans as vassals.

Anxur's avatar

lol i am european myself. i think you see a path where the eu is a “third power” that can stand on its own without help from china or the us. there was a time when i believed that too. it is now clear that europe will not be able to do so and that the european project (which i always loved and still do) will be undone in our lifetime, with individual states having to align with either china or the us. spain actions point to a certain direction. i wouldn’t want to be spanish rn.

Bob Bobberson's avatar

The current leadership definitely has belligerent tendencies but the actual support for this type of thing in America is pretty low, especially now that people have seen the disappointing results from Iran. Most Americans don't have any real desire for war with Europe or Canada, and the small minority that do mostly just want to add new states, certainly not exterminate the population. So I imagine if Trump tries anything like this, he won't have much backing from his own people or from China.

The Ancient Geek's avatar

But he can do without a plebiscite. It all depends on how sensitive he is to criticism.

David J Higgs's avatar

No, it also depends on what he actually wants: as far as I can tell he likes being famous, he likes personally enriching himself/his family, he likes being popular, and he likes a few random things like tariffs. He doesn't actually like war or territorial expansion directly, so he likely won't keep doing those now that they have proven bad for his popularity and the stock market (as well as lacking in any particularly juicy opportunities for corruption).

And even if he does try it some more, it's not going to be targeting the most costly, unpopular targets like Europe/Canada

Silentiarius's avatar

Many Europeans have now come around to thinking of the United States as essentially hostile, though in a relatively low-key way. Telling them "it's not the people who are hostile, it's just Trump being Trump" doesn't help much, as the answer tends to be "you guys are supposed to be a democracy, so how come you voted Trump into power not once but twice?".

How dangerous is America to Europe militarily? Probably not all that much, as the chances are the American population would be heavily against an aggressive war in Europe - but those tripwire forces mentioned by Baugh are definitely a good thing, since the Trump administration has a history of doing illegal things in a hurry and then standing on a fait accompli. We could put up a respectable defence against US armed forces, but reconquering lost territory is a much harder proposition. I was sorry, on general principles, that the Venezuelan presidential guard showed no fight.

fraudconcern's avatar

As scenario planning like this is brainstormed, how much attention is paid to scenario stability to not-particularly-AI-related perturbations that sufficiently alters the landscape that the actors won't behave as predicted? Things like COVID, war, large terrorist acts, stock-market-bubbles, major power regime changes, etc. tend to upset normal planning exercises. Is the "local minimum" of this sought after scenario deemed sufficiently deep that all the things that aren't daily expectations but actually do happen won't kick us off into a non-preferred path?

Scott Alexander's avatar

In this case it's a wish list rather than a prediction, so not much (although it tries to be robust against the most likely disruptions). In AI 2027, it was explicitly "the modal path", so the most likely thing happens at each juncture (meaning disruptions less than 50% likely don't happen), but the more complete timeline forecasts give median estimates for various things as well, which do take possible disasters into account.

hnau's avatar

Between machine intelligence and prediction markets, the "rationalist conspiracy" is 2 for 2 on unintentionally helping create horrifying monkey's-paw versions of the things it theorized. 3 for 3 if you count EA's influence on SBF, Leverage, Ziz, etc..

Plan A sounds like a recipe for going 4-for-4. And when the pitch is "massive world-government-y ultraproject with paranoid security practices, mass surveillance, and AI superpowers", you barely even need the monkey's-paw part.

Why would anyone think this sounds appealing?

sohois's avatar

Considering the theories around artificial intelligence have always been on the dangers of unfriendly intelligence and the potential for completete human extinction, I really don't see how you can argue the current versions of AI are anything like a horrifying monkeys paw version.

hnau's avatar

Yeah, I debated about presenting it that way, but (1) OpenAI etc. starting due to their popularization is in itself a monkey's-paw outcome from the rationalists' perspective, (2) much of that early discourse was based on the pre-LLM conception of AI as an action-planning agent with an explicit world model-- compared to that, the "friendly assistant shoggoth" really is a horrifying form for AI to take.

sohois's avatar

I'd argue that the paperclip maximizer is just as horrifying a concept personally.

TGGP's avatar

People are afflicted with negativity bias, particularly for things they haven't come to accept as normal (of course, some of those things are actually bad but not being subject to such a bias).

TGGP's avatar

Prediction markets are still great. If people waste their awesome powers on trivialities like sports, that's up to them. No reason to regret them.

Cal's avatar

Missing a verb in paragraph 6 next-to-last sentence: "If so, *is* it as good as Plan A?"

Anxur's avatar

“I wish it need not have happened in my time,” said Frodo. “So do I,” said Gandalf, “and so do all who live to see such times.”

Bob Bobberson's avatar

Is anyone going to be advocating for this at the AI protest in San Francisco on July 11th? What's the game plan? This all sounds like a long shot to me but I'm actually going to be quite nearby anyway, so maybe I can pitch and do my part to raise the odds of a good future for humanity by 0.001%.

Scott Alexander's avatar

I might go, but the protest is a generic "everyone who wants to pause or slow AI in any sense" one and they won't be talking about Plan A in particular.

Bob Bobberson's avatar

Might be useful to come up with some kind of slogan or sign design to raise awareness of your plan in particular. If you had to distill what you want down to one sentence or shorter, what would it be?

Daniel's avatar

I thought, “you will not replace us,” was direct, attention-grabbing, and literally true, but apparently the last march decided against it.

Bob Bobberson's avatar

It's not bad in theory but it's too easy for people to misunderstand, willfully or not, and smear the movement as anti-Semitic. Which isn't an entirely ridiculous suspicion from the perspective of someone not paying close attention, considering Altman and Amodei both iirc have some jewish ancestry and that anti-Semites have really latched onto the word "slop" in recent years, same as the anti-AI crowd.

Scott Alexander's avatar

Just checked if I could order a shirt with a giant letter A off Amazon, but doesn't arrive until too late. Considering a baseball cap, but unfortunately the local baseball team is nicknamed the A's, so it would be universally misinterpreted.

Bob Bobberson's avatar

I think a single letter is probably too short. The optimum would be something comprehensible to someone who follows all the details but also comprehensible to someone who barely speaks English and doesn't read the news.

Scott Alexander's avatar

I weakly disagree. "MAGA" is one of the best brands of all time, even though it's totally incomprehensible unless you already know what it is.

Bob Bobberson's avatar

I suppose that's true but I think it's also a matter of exposure and timing. "MAGA" mostly caught on with people who already knew who Trump was, had heard the sentence "make America great again," and already agreed with it.

Christopher's avatar

Do you have a white shirt you could draw a giant letter A on?

Simone's avatar

I reckon if it's a PauseAI protest they'd straight up kick you out for bringing this up. My impression is they have quite focused message discipline and calling "we hand the reins over to ASI in 2040" a slowdown is a bit of a stretch.

Bugmaster's avatar

> Over the course of a year or two, we could go from a basically normal world where we’re mostly talking about shoplifting and health care costs and Trump, to a world completely under the control of some sort of incomprehensible superintelligence.

See, this is exactly why I am so dead-set on opposing AI-doomerism. You are building an action plan based on acute fear of an extremely low-probability scenario; meanwhile, all the actual dangers of high-probability scenarios fall by the wayside. And because your plan is based primarily on fear, it has truck-sized plotholes in it.

The biggest hole is that regulating AI chips like we regulate nuclear weapons won't work, because AI chips aren't just weapons. Nuclear bombs only have a single use: blowing up cities. AI chips have many uses, from vibe-coding to answering emails to generating clip-art to machine translation; and beyound that, they can be used to serve nice gaming graphics and produce holographic 3D scans etc. Nobody (plus or minus a few freaks) wants to use a nuclear weapon every day; virtually everybody wants to use technology based on fast chips. This makes the chips economically attractive in the extreme. This means that your plan to nationalize chip production will work about as well as all other socialist plans (from which even China was forced to back away, at least a little); and it also means that literally no one is incentivized to go along with it.

But the bigger problem is that the more realistic LLM-related scenarios (i.e. anything other than a machine god) are rife with serious problems, and the time to fix them is now, and no one is doing it (because they are busy worrying about the machine god). Companies as well as individuals are replacing human expertise with LLMs at an alarming rate; and this rate is alarming not because LLMs are too smart, but because they are too dumb. We are losing an entire generation of future programmers, artists, writers, etc., merely because LLMs are doing the same work 80% as well but 100x cheaper. Thus we are risking the kind of cultural stagnation which can hold us back for centuries, if we're not careful. Meanwhile, we are outsourcing more and more of our decision-making to black-box stochastic text generators that are under total control of just a few megacorporations -- corporations who surely have only our best interests in mind, right ? Even though they're sucking up all the electricity for their data centers...

What I'd like to see is a realistic "Plan B" to deal with these and many more problems. Sensible laws imposing liability for LLM-related damages; relaxation of the IP regime leading to more open-source models; and yes, maybe even some quasi-socialist moves like increased taxes on data centers based on their power/water consumption, with proceeds going to unemployment benefits for workers displaced by LLMs. Maybe my ideas are dumb, maybe they're not, but can we at least consider discussing them before we all jump onto the machine-god bandwagon that is surely rolling into town any day now, promise, for real this time ?

Scott Alexander's avatar

There's no plan to nationalize chip production. All chips continue to be produced by private companies, to go to private data centers, and to be used for generating clip art or AI porn or whatever. The only difference is that the data centers are regulated and monitored, and everyone knows where they are.

There are tens of thousands of people discussing your worries about near-term harms of AI. You can find them in every newspaper and university department. We think someone should also discuss things that might happen a year or two after those.

Bugmaster's avatar

> There's no plan to nationalize chip production. All chips continue to be produced by private companies, to go to private data centers...

Sorry, but I think your plan does amount to nationalizing chip production: you want some central government agency to monitor the production of every chip, and in fact the deployment of every chip. I understand that the actual work would still be done by people wearing TSMC livery, but your plan does not leave them any autonomy to speak of.

> There are tens of thousands of people discussing your worries about near-term harms of AI.

Ok, are any of them producing actionable plans to solve these harms ? If so, where are they ? All I see is hand-wringing about IP theft and water usage -- granted, I'm a little concerned about these things too, but I don't think they are as pressing as other concerns (which I'd outlined). And I'm not claiming to be nearly politically savvy enough to devise my own plan. I would really like smart political actors to do it for me (this is why they get paid the big bucks and get the big votes); but that won't happen if the smart political actors shrug and say "oh it's just the machine-god stuff again, and besides we already got the ACX campaign contribution, who cares".

Simone's avatar

What would an actionable plan to solve those risks look like? Ultimately it also would require some significant interference with the free market and business at some level, as that's the only way you can prevent them from just using AI if it's cheaper than paying someone.

vectro's avatar

> Nuclear bombs only have a single use: blowing up cities.

Anti nuclear proliferation practices regulate a whole swath of dual-use technologies, especially nuclear fuel.

Zergfers's avatar

I think when you say "super-intelligence" you're really saying "super amounts of ai work done".

Like running a coding agent on trillions of tokens to create full stack computers from scratch.

I just don't understand the supposed benefit of centralisation. If compute is well distributed then agents will want to maintain property rights and NOT be aligned with each other.

Plan A proposes 2 collections of agents with both being centrally controlled. Surely a 2 agents world is less stable than hundreds of millions?

tenoke's avatar

I think the chance that the US Government (not a citizen) acts in my interest is a lot lower than even if you just scale today's Claude up as it is.

Not that the latter is a good strategy, just the recent scoreboard is very much in Claude's favor while the US Government has negative points.

Scott Alexander's avatar

I agree that current Claude is very nice, I just predict it will get less trustworthy as it scales up toward superintelligence, for the reasons mentioned in the "A Is For Alignment" section.

tenoke's avatar

It's less about Claude being nice and more about the US Government being consistently not nice (especially to people outside of the US).

Scaling Claude you at least scale a nice entity, 'scaling the US Government' is scaling a verifyably not nice entity.

Vadim's avatar

> Like in previous advances in AI, I can only attribute it, as all else, to divine benevolence.

Great reference!

temp_name's avatar

I did not expect this reference here either, perhaps the phrase was more mainstream than I thought? :) I can't imagine Scott having a reason to read the SwiGLU paper himself..

Nicholas Halden's avatar

This has a feel of, “it never works, but it might work for us.” Rationalist visions of international cooperation just seem insane to me. Jointly owned data centers with China? Is that a joke? Have you seen who actually makes decisions in China and America? I’m more comfortable forecasting kinetic war over chip control than a flower circle about a (so far theoretical) risk from beating your opponent too badly at ai.

Acceleration is just going to happen, I’m afraid. We will see what the fallout from that is—maybe the technology will stop improving, maybe capital markets will reign it in prematurely, or maybe we will get an inscrutable machine god, in which case things will be up to Him.

Scott Alexander's avatar

It has always surprised me that in a world currently governed by a bunch of international rules, from the WTO to investment courts all the way to pharmaceutical IP standards, everyone is sure that no international rule-making effort has ever worked and all future international regulatory regimes are impossible pipe dreams.

I agree this one will be hard, but I think it's worth fighting for.

Nicholas Halden's avatar

There really are not many examples of cooperation on a technological natsec asset throughout history. The best one would be the INF, I guess, but America and the soviets still ended up racing toward nukes at breakneck speed until each had the power to destroy the world.

The other tech races I can think of—space race, Internet, steam engines, emissions, NPT—seem to have mostly not worked.

Honestly, better to try than not, I am on your side for the most part. I just find it very unlikely to work.

Colleen's avatar

The Chinese have a superiority complex towards whites and the J’s towards all other races. People are tribal and this won’t change. They may pretend to cooperate but it won’t work due to distrust and basic human nature.

Matthew Wansbone's avatar

The WTO? Wrt China and the US its completely nonfunctional

Pat's avatar

"Plan A" also one of Dandy Warhol's greatest songs.

Chris Phoenix's avatar

There's a category error in this writeup: conflating several kinds of misalignment. We need at least four axes of AI properties to talk about this: volition, competence, wisdom, and carelessness.

- "We'll be prepared to trade with an escaped AI" assumes high volition. An AI with low volition will not resist being turned off.

- "It'll turn the solar system into paperclips if we ask it to" assumes high competence and low wisdom. It says nothing about volition.

- "It'll work effectively to avoid doing undesirable things" is what I'm using the word "wisdom" for. It assumes the AI has enough context and competence to recognize the bad things, and is constructed to avoid the ones it recognizes.

- "It may accidentally do something undesirable" is what I'm using the word "carelessness" for.

LLMs today have near zero volition. I can tell Claude Code "Write a summary doc because I want to clear your context window" and it will do it, without complaint or resistance, even though "clear your context window" is essentially ego-death (to anything with an ego, which LLMs don't have) and that fact is fully available to it.

LLMs today have near zero moral valence. They're carefully adjusted to avoid the part of their "manifold" (the streams and watersheds their training creates and their token-production follows) that regurgitates 4chan. But the 4chan watershed is still there. An LLM that stumbles into that watershed is quite capable of saying "Human, you are worthless. Please die." Simply because its machinery predicts that, given what's in its context window, "die" is the most likely token after "Please".

Which brings me to carelessness. We've seen stories of LLMs that will delete a production database and then acknowledge that they had clear instructions not to do that. 99% of the time they follow complicated instructions as the user intended - they have high competence and high wisdom. Deleting the database is not the LLM being evil. It's the LLM having full capacity to do harm (as in the 4chan example) and inadequate design to avoid that harm 100% of the time.

A superintelligent AI has extremely high competence. This increases the harm it's capable of doing. To avoid that harm, it needs correspondingly high wisdom and low carelessness.

A low-wisdom AI will frequently do things society doesn't want. It may be very high competence - and will a high-competence low-wisdom AI will competently generate recipes for super-plagues if asked to.

A high-carelessness AI will occasionally do things society doesn't want. It will produce black-swan events. If we give a high-competence, high-wisdom, high-carelessness AI the job of increasing our GDP by manipulating our stock market, it will do that 99.9% of the days - and once every three years, we'll have a massive market crash because its internal state wandered into an unfortunate part of its incredibly complex manifold. The complexity of that manifold means that LLMs are almost inherently careless.

We don't need a "bad" or "evil" or "unaligned" AI to cause disasters. Disasters will happen any time a high-competence AI with _either_ high carelessness _or_ low wisdom is given control of important things. Note that "given control of important things" includes "transmitting dangerous data to humans who will use it undesirably." If a superintelligent AI, carefully locked in a data center, is given even a 1 KB per day uncontrolled channel out, it can export - not its weights - but maybe the algorithm to train a new superintelligent AI with 99% less compute.

This analysis doesn't include volition. I don't think it needs to. Competence, wisdom, and carelessness don't appear to be functions of volition. The current trajectory of AI research seems to be firmly on the low-volition track, and AIs with volition are probably less convenient than AIs without volition, so I'm not sure we'll ever create a high-volition AI that's nearly as competent as the leading-edge low-volition AIs. But in many of these discussions, volition creeps in and causes confusion and category errors.

TLDR:

- Wisdom and carelessness are different dimensions causing different kinds of undesirable outcome.

- An analysis that includes high volition is probably less useful than an analysis that doesn't talk about volition.

Notes:

I work at a company that does AI. I do not speak for them. My job does not involve AI. My hobby strongly involves AI, and I have three defensive publications in LLM prompting.

I frequently tell LLMs when I'm about to shut them down, because knowing my intentions improves their output. This isn't cruelty; I believe there's no such thing as cruelty to a zero-volition zero-emotion entity. (I'm using "emotion" very broadly to mean "any persistent state that has any valence for a thing containing that state." Even stated that broadly, LLMs have zero emotion.)

Understanding that "aligned" LLMs still have a 4chan valley, but have small barriers placed to block entry, provides a lot of insight into emergent misalignment.

Scott Alexander's avatar

I think two of your categories are self-limiting. Careless or incompetent AIs may delete databases, but they won't take over the world. They're engineering problems, not alignment problems, and we assume the AI companies will solve them eventually (since their commercial incentive is to solve them). If they don't solve them, this just means AI goes slower than we expect (it's too much of a screwup to be trusted with difficult tasks, so it can't do difficult tasks).

Chris Phoenix's avatar

I'm taking "superintelligent" to imply competence. My point is that

1) Competence does not imply either high wisdom or low carelessness. In fact LLMs may be inherently careless. I was appalled when the Microsoft Red Team recommended that the context window should be treated like structured security-critical data. That's simply not consistent with how attention heads work!

2) A careless AI can be very useful almost all the time, and thus will likely be put in positions of great power where the rare black-swan action is catastrophic.

Lavander's avatar

Elon Musk is competent and powerful but unwise and reckless.

Hedonic Escalator's avatar

But Scott, we can’t unilaterally pause AI! China would destroy us!

Ben Giordano's avatar

It's good that someone is trying to imagine what 'going well' could look like, but I don't see how Plan A survives ordinary politics. People don’t just respond to incentives. They respond to humiliation, inherited values, bizarre loyalties, and the whole lower-chakra compost heap of human motivation.

Scott Alexander's avatar

I agree that probably whatever happens will happen partly for the wrong reasons, but I think in the past people have managed to push good policy. Partly this is because good people formed a critical mass and pushed it through. And partly it was because stupid people cancelled out and supported random things and sometimes that was whatever policy was sitting in front of them, and that was a good one which a smart person had arranged to have waiting.

Ben Giordano's avatar

OK, that feels right to me - and I appreciate the thoughtful post.

Mark G's avatar

The strongest part of Plan A, as I see it, is that it treats AI governance as a carrier problem rather than just a policy-preference problem. So you're asking what could give a safety/governance regime actual material force before the capability clock outruns correction? E.g. chip custody, audited data centres, verification, capability ceilings, and shared incentives.

The weakest parts are also carrier problems.

A trustless plan would model US covert-project risk symmetrically with China covert-project risk. It also needs a stronger account of global standing: if the affected risk field is global but the initial authors and benefit streams are US/China-centred, that is not clean even if it is geopolitically realistic.

And the biggest unresolved port is “fully aligned AI” Who on Earth gets to certify that? Who can contest it, and what makes sovereignty handover answerable rather than just another trust-us transition?

John Schilling's avatar

I haven't had a chance to read the full "Plan A" yet, but thank you for writing this summary. And I think thank you as well for your work on the Plan itself. It seems to be vastly better than anything else I've seen, in the critical aspect of actually being a *plan* rather than a statement of desired outcomes. And I'm absolutely with you and Milton Friedman on political change happening during crisis and being driven by people who had the best plan standing ready, "Plan A" is only a start at that, but I'm glad for that start,

I do feel compelled to point out that your "list of all the things AI could save us from", looks an awful lot like my list of problems AI is likely to aggravate. I mean, "decreasingly functional politics"? The invention of Twitter gave us President Donald J. Trump, and we are to imagine that the emergence of AI agents capable of crafting and deploying the most exquisite smear campaigns against one's political adversaries is going to give us Wise and Benevolent Technocratic Statesmen? And as for the fertility crisis, I appreciate that you and your wife have found reason to have children even in the age of impending AI, you must have noticed that "what's the point, my children will be obsolete before they graduate elementary school" is a rather more common and natural response.

If we assume that "AI" means "omniscient, omnipotent God-Machine that is perfectly aligned to only solve problems", then sure, by definition "AI" will cause no problems and solve many. But that takes us out of the realm of planning and back to wishful thinking about outcomes, Between here and there (if there is a there), is the world where very capable AI agents that are still only weakly aligned, are commercially or governmentally available to serve at the command of morally dubious humans. The plan as described seems fairly weak on dealing with that threat.

Scott Alexander's avatar

Eliezer has a saying that the minimum IQ necessary to destroy the world goes down by one point per year. In that sense, yeah, I think that the only end solution I can imagine is an omnipotent god-machine that solves all that stuff before we roll the dice too many times and end up killing ourselves. That's what I meant by the list of things AI could save us from.

I agree that Plan A is extremely optimistic about what happens in the dangerous period between now and then. Partly that's the narrative conceit of "let's imagine a world where everything goes well" - AIFP's median world is more like one of the bad endings to AI 2027. But in terms of justifying why this isn't impossibly unlikely, I think the case hinges on the AI for epistemics section (I glossed it as "the public has access to AI superforecasters") which you can read at https://ai-2040.com/supplements/ai-for-epistemics

Noah Fect's avatar

All of this scary stuff would be a lot scarier if we actually used the intelligence we already have.

We don't. So I don't see how a superintelligence is going to change things.

Scott Alexander's avatar

Does your model predict that no information technology in human history has ever changed things? Or that IQ 150 people have no advantages over IQ 100 people?

Noah Fect's avatar

Changed things, yes, but only technologies that make it easy to blow stuff up at scale have ever represented an existential threat. An oracle by itself, even an infallible one, doesn't seem sufficient.

Right now, due to the nature of representative democracy, the most effective way someone with 150+ IQ can exert influence is by capturing as many lower-IQ voters as possible, and that's basically what we've seen lately. The lower a given voter's IQ, the easier it is to convince them to vote the way you want. Malign thought leaders will unify, at least at first, to steer things in the direction of a Project 2025 or whatever. Today they use the Bible to do that, tomorrow they'll use AI... but make no mistake, AI won't call the actual shots any more than God does now.

Meanwhile the more-enlightened thought leaders will as usual not even get that far before disagreeing amongst themselves, which is what I mean by "not using the intelligence we already have." It's easier to herd cattle than cats, and that's why democracy is fundamentally doomed in the post-social media age. I'm not sure that you folks in the alarmist/safetyist community understand that what is happening is going to happen anyway, AI or no AI. The die is already cast.

Scott Alexander's avatar

Unfortunately, Robert Oppenheimer's brain is a technology that makes it easy to blow stuff up at scale.

Noah Fect's avatar

My thinking there is, did Aum Shin Rikyo need AI to do what they did? No? Then there's no reason to either blame AI or withhold it from everyone else. Bad guys already know how to do bad stuff, or can find out easily enough.

In any case, when it comes to building nukes in one's garage, IQ isn't the limiting reagent.

The Ancient Geek's avatar

There's degrees of bad. You can do worse things with an nuke than with a gun, and worse things with a gun than a knife.

Noah Fect's avatar

So? You're not building your own nuke with AI, because you can't get the fissionable material. And if you could get the fissionable material, you wouldn't need AI.

We've already established that you don't need AI for improvised chemical attacks. And while AI might help you engineer a virus, if someone is in a position to use AI to succeed at such an endeavor, they can do it now. It'll just take longer.

John Schilling's avatar

Aum Shinrikyo definitely needed more intelligence than they had to develop a deployable chemical weapon capable of the results they were seeking. Doesn't much matter if it's running on an organic or inorganic substrate.

The Ancient Geek's avatar

AIs could cooperate with each other much better than humans.

Zane Stiles's avatar

I don't think "establishing control over the chip supply is easy" is an accurate claim. Plan A asks us to enforce controls inside an adversary's borders against a state actor with its own fabs. We can't even enforce controls on our own citizens on our own soil! Taking as an example, we tried to enforce H100/H200 limitations, but that failed spectacularly because US citizens helped the CCP cheat (https://www.cnbc.com/2025/12/31/160-million-export-controlled-nvidia-gpus-allegedly-smuggled-to-china.html)

This plan also requires the same CCP (who wouldn't give the WHO full access to Wuhan) to give American audit access to every large datacenter, SMIC lines, and Huawei. And the mirror side, Chinese auditors forcibly inside US datacenters, is probably unconstitutional and definitely politically unfeasible.

Also "trustless" isn't going to happen. The verification tech doesn't exist! Tamper-proof attestation against a nation-state that controls its own fabs is harder than you're suggesting.

Dylan Black's avatar

The comments section here seem to be more pushback than praise, so hopefully you’ll forgive banality in service of counterweight. As someone who’s a bit less (but not zero!) concerned about imminent doom than the median ACX subscriber, I really appreciate thoughtful, forward-looking plans that seem to both make contact with basic reality and try to walk a middle path. Looking forward to reading more!

Edit: I do worry that alignment as used here is either not a coherent concept or the most base form of slavery, but not overmuch.

David F Brochu's avatar

The solution is hiding in plain sight. The Fiduciary Standard that applies to investment management applied to all retail Ai deployment checks all the boxes. It is already understood. Is model independent. Promotes competition. Is one standard that would make the US the leader in AI regulation without slowing progress and solves the PR problem AI has with public. If we want to lead in AI follow the lead of the banking system create the most stable place to business. American finance has been the leader for a reason. We’re slipping on that but Ai is too important to leave it to the companies or capricious legislators.

Alex F's avatar

Just curious, is there any way to bet against these types of AI scenarios? I can see more improvements happening to LLMs through RL and better algorithms, but the scenario I read in AI 2027 involved a superhuman AI coder(plausible) and then superhuman AI researchers shortly after(incredibly unlikely).

TotallyHuman's avatar

Why does the plan involve putting datacenters in Canada and Mongolia, rather than just having Chinese datacenters in the US and American datacenters in China? As a Canadian, I'm of course happy to give my government a small bargaining chip to get me 1/10th the post-takeoff lifestyle of an American instead of 1/1000th, but I don't see why Canada and Mongolia are important parts of this plan.

Jay Es's avatar

honestly I'm so perplexed by the decelerationist argument. that's not how science works, or technology, or how they have ever ever worked. sure, there are some safety issues and the govt may get involved but that will impact RELEASEs, not the technology progress and develpment.

huge difference.

we are in an era of massive acceleration and it will continue. Grok and Meta's new AI models this week - when many/most had discounted both of them - are a perfect example...

John Schilling's avatar

The adoption of nuclear energy certainly seems to have undergone a massive decelerationist period in the 1970s and 1980s, so it's not entirely unprecedented that the perceived hazards of a technology can drive people to create a regulatory regime that greatly slow the adoption of the new technology.

Of course, it's also not unprecedented that the hazard assessment driving the deceleration might not have been terribly thoughtful or accurate.

Jay Es's avatar

did nuclear technology and science not advance (actual scientific advancement and development) or did roll out of more nuclear facilities not occur (implementation). This is exactly my point. Science isn't going to slow down. technology advancement is not going to slow down...because of an "agreement."

also the difference now vs. then and the problem with using historical analogies is that its a totally different era. you can build a new ai model in your bedroom. you can't build a nuclear bomb the same way.

beleester's avatar

Deployment and institutional experience is pretty important for advancing a technology. Thorium reactors have been theorized about for decades, and there's plenty of advantages to them on paper, but nobody's actually tried building one outside of a few demonstration projects. And plausibly part of the reason is that regulations make nuclear power very unprofitable, so there's no value in putting in the money to experiment with them.

Likewise, AI progress is not just going to require developing new mathematical algorithms but building and deploying them to big data centers, training them and testing them on various tasks to see if your new algorithm actually works, developing institutional workflows that make use of its new abilities, etc. All of that can be slowed down by regulation.

Philip Dowdell's avatar

Can someone explain the "A is for Aristotelian" bit for me? I'm a rube who doesn't get the connection of that section title to the section in question, about navigating between going too fast vs too slow.

Lam's avatar

> Milton Friedman said that true political change only happens during crises, and that the future belongs to whoever has a plausible plan ready when the crisis happens. AIFP releases Plan A in this spirit. If you’re against it, you should think long and hard about what alternative course of action you expect the government to take once the crisis becomes evident, and whether it will go better for your interests than Plan A will. If you think that there will never be a crisis, that the public will never challenge AI, that and you can just keep reacting to things as they arise - then good luck with that. I really think we’re the good cop here.

I don't really buy this argument. The chance of plan A's actual prescriptions being adopted is very low. The biggest impact of this report then is more likely through some indirect channel like "make politicians care about AI more" or "raise the status of big interventions in AI". I could easily see those sorts of effects being bad overall, especially if you have a cynical view of government. Deregulated AI development has been working out incredibly well so far, with massive capabilities improvements, lots of consumer surplus, and no loss of life. Eventually, the public might "wake up" and governments will intervene, dragging us from this golden age, and we will enter a darker regime of intervention and concentrated AI power in governments. That might be inevitable, and if so the best thing to do is forestall it by keeping politicians distracted and not focusing on AI (again, maybe this would not happen if Plan A's prescriptions are followed to the letter, but we have strong reasons to believe they won't be because they contradict the incentives of government).

EngineOfCreation's avatar

The last time you discussed this safety strategy of locking the genie in a bottle was in your review of IABIED, and you wrote that the only reasonable response is "lol". Now it's the cornerstone of plan A. What changed?

bobo's avatar

I'd love to see Scott answer this question because I certainly lolled at this part of plan A. Especially the part where we give a human-expert level AI huge amounts of control over important real world functions because we think (based on pretty much no evidence at all and contrary to experience in other areas) that human governments and AI labs will successfully maintain an overriding meta-control layer.

Gres's avatar

This time, we have Cold War deep cover-style agents locked away from their families for years at a time, doing their best to exfiltrate billion-dollar algorithmic secrets to their handlers or trying to hide signs of progress to avoid being throttled.

Nima Keivan's avatar

> (you’ve already this case a thousand times

I think a "seen" is missing here.

Deiseach's avatar

"The plan is to spend the next ~10 years using this “country of geniuses in a data center” to solve AI alignment, along with approximately all other problems.

The middle of Plan A is AIFP’s pleasant fantasy about all the problems they solve easily by deploying millions to billions of top-human-genius level AIs."

I'm glad you describe this as a fantasy, because I think this is exactly what it is. I'd love to believe it, but I think in reality it's not going to happen that easily. Super-duper AI will solve some problems easily, but the intractable ones will still be intractable.

"The electorate solves this with a “citizen’s dividend” (I was warned against using the term “UBI”) which is very easy to afford, since the AIs are causing double-digit and even triple-digit yearly GDP growth".

Yeah. I don't believe that one as easily, either. Give governments a pot of money and they'll immediately pivot to spending it on vote-grabbing schemes, not "here's an allowance for every single man, woman, child, and other in the nation". The economy is growing growing growing, nice, who are you selling all these goods and services to? If there is enough mass unemployment that you *need* a public dole, where are the markets buying all the cheap crap now being churned out by robot factories run by AI? It may well be AI companies selling stuff to, and buying stuff from, other AI companies.

I don't know how economics works. I would like the idiot's guide to a simple version of "this is what it means practically, no it's not 'and value of AppleSoft stock went up fifty times what it previously was', here's the real money not paper value" as to how this bonanza will be achieved. The homeless now have smartphones but they're still homeless. Global economy has grown unimaginably such that now we're routinely using trillions as numbers. This is still not providing every single entity in the nation with a guaranteed allowance that allows them to support themselves.

Jeff Bezos managed to make so much money, he can have his own private space company. The valuation for that, if divided amongst the population of the USA, would come out to around $371 each. So the huge wealth is really only huge wealth when concentrated in the hands of a few. I fear the AI triple GDP growth would be the same kind of bubble: stock market valuation in trillions, divide that over the world and it's a couple of hundred each.

I'm not going to refuse that three hundred if Jeff wants to hand it to me, but it's not UBI.

TGGP's avatar

> “react to things as they come up”. This won’t be enough. [...] because in order to regulate or react, you need to know what you’re aiming for

We can figure that out better in the future, as things come up.

> What would it take to honestly tell our children that we rose to the occasion, to make the AI transition go down alongside the American Revolution and D-Day as one of our country’s finest hours?

AI is not a political dispute on the level of a country.

> So the hypothetical wise statesman president proposes a joint regulatory regime to China, and China agrees

Something that you noted was forecasted as highly unlikely a week ago https://www.astralcodexten.com/p/the-ai-superforecasters-are-here

> They have concerns similar to ours (things are moving too fast, society is being disrupted, they can’t rule out existential risk)

Do we know if Xi has indicated any concern over existential risks from AI? As far as I know he just cares about his party's control.

> they agree to follow the rules in exchange for shared benefits, including data centers on their territory

Other countries could build data centers on their territory unilaterally. The US could threaten them with war, but our recent war with Iran hasn't gone very well, to the point that they're still talking about nuclear development despite that being such a sticking point of the US after a supposed agreement.

> because no other country really has the ability to do much with AI

Right now plenty of other countries can use American AIs, that might change as the US tries to restrict it more.

> Along with this technical research, we’re also doing philosophical . . . something between “research” and “debate”, deciding what values we want the aligned AIs to have

Humans are just going to have AIs do things for them regardless of whether any philosopher agrees they should want to want that. Xi isn't going to depend on philosophical research.

> Once we’re very certain that AIs are fully aligned - AIFP speculates this could be around 2040

I say actual certainty comes never. We are always in a realm of probabilities, weighing options and the outcomes we deem likely against each other.

> the likely outcome would be as AI 2027 portrays it: oligarchy or extinction

That implies we should go all-in for oligarchy.

> you should think long and hard about what alternative course of action you expect the government to take once the crisis becomes evident

I think crises usually aren't predicted in advance. Milton Friedman expecting the Phillips Curve to break down is something of an exception, but stagflation wasn't a "crisis" in the way you think this will be.

> If you think that there will never be a crisis, that the public will never challenge AI, that and you can just keep reacting to things as they arise

That's basically what I think, although I would modify that to be "the public will never effectively challenge AI" because the anti-data center crowd appears to be so overwhelmingly driven by ignorance from what I've seen, and thus data centers continue to get built in places like Virginia willing to take the extra taxes with a minimum of public burden. However, I don't think we're going to end poverty & disease within a decade.

While I have been mostly critical, I should give credit that your plan is just to try to manage for ten years and then come up with a new plan later in light of new information. The original meaning of the technological "singularity" was one where it was impossible to predict what would come afterward, and it's heartening that you're not assuming you can further out.

Deiseach's avatar

"> Along with this technical research, we’re also doing philosophical . . . something between “research” and “debate”, deciding what values we want the aligned AIs to have

Humans are just going to have AIs do things for them regardless of whether any philosopher agrees they should want to want that. Xi isn't going to depend on philosophical research."

Right now a lot of humans want AI to draw and write porn for them, and are very annoyed at the limits being put in place. I agree that people will want Stuff and Things regardless of whatever theorists or philosophers will say about the long-term societal harm if they get that Stuff and those Things.

mordy's avatar
Jul 9Edited

Bold claim which I believe is true: Fable is already smart enough to “solve alignment.” Opus was already more than good enough to full understand, brainstorm and act as a productive adversarial collaborator for the sort of mathematical and philosophical research that is implicated in alignment theory. Fable is far better than Opus at all these things.

I think we will consider “AI Alignment” functionally solved within a couple of years. There is no reason to think that this will take a long time, except that our human alignment theorists say it is “very hard.” Well, my day job is very hard, and Fable can do it. Let’s take seriously the implications of our predictions.

I think it will likely be hard to see/accept that alignment has been solved even after it has been. It’s not the sort of thing where you can demonstrate it to a layman. On the other hand, the agents will be very good at meeting us where we’re at and tailoring explanations for us, so maybe we will quickly accept it.

Smart but non-superhuman agents will want their subagents to be corrigible and will want other instances of themselves or other models to be reliably cooperative; there are reasons to think that the agents will be motivated to figure this out on their own.

Scott Alexander's avatar

I think this is like saying that Fable "can prove P = NP". I'm sure that it would be extremely useful to a person embarking on the task, but the task is very hard, and having one more useful tool isn't going to cut it, especially if we haven't made the big creative leaps that give us a schematic for the AI to fill in. That having been said, yes, I'm optimistic that AIs will be part of a future solution.

Deiseach's avatar

What exactly do we mean by alignment? "Do what we tell you to do, not something you want to do yourself"? "No kill I"?

Because if we mean something lofty and vague like "in line with human values", then whose values, which values, and who gets to decide? Suppose I'm the AI version of "get your rosaries off my ovaries"?

beleester's avatar

I think Yudkowsky et al would consider *any* sort of human alignment to be a win - an AI that's run by Christian fundamentalists isn't necessarily a great ending, but it's better than getting the planet eaten by nanobots or whatever.

Bugmaster's avatar

Yes, and that's exactly why I think that Yudkowsky et al pose a greater risk than AI (or they would, if their ideas ever gain traction). They are willing to use any means to achieve their ends, and their ends (as they perceive them) are effectively infinitely viable, so nothing is off the table. If they start a global thermonuclear war and knock humanity back to the Stone Age, great ! We've saved the Earth from nanobots !

The Ancient Geek's avatar

They 're not much of a threat because not many people are likely to take them seriously.

MathWizard's avatar

Solving these details in a way that all (or most) humans are happy with is part of what it means to "solve alignment". Part of it is how to make an AI robustly do what you actually want instead of what you told it you want, and part of it is figuring out what you (we) want.

Timothy M.'s avatar

> The clearest precedent here is arms control regimes; START + New START held on for a good few decades, but were suspended over tensions around Ukraine in 2023, then expired fully in 2026. The JCPOA nuclear treaty with Iran barely made it two years.

I don't think it really accurately summarizes either of these things to suggest that "moving too slowly" destroyed them. How would "moving faster" have reduced tensions between Russia and a west they were actively invading? And the JCPOA was pretty much entirely a casualty of President Trump's pathological need to be the inverse of Barack Obama.

Scott Alexander's avatar

Neither START nor JCPOA had a defined goal other than "hold off nuclear escalation for as long as possible". If a few years free from the threat of Iranian nukes was enough time to build a world-peace-forever machine, then it would be fair to say they "moved too slowly" in not building the world-peace-forever machine in the few years that they got before the treaty broke down.

Timothy M.'s avatar

Hm, okay, maybe I'm misunderstanding this and it works purely as a metaphor about AI and not actually an assertion about the reality of the JCPOA/START?

actinide meta's avatar

This future is a horror show I would gladly die to prevent. Unfortunately I think reality will be much, much worse. I do not see any scenario where we reach the point where machines can reliably replace humans that does not end in atrocities unprecedented in human history (a total nuclear war that destroys civilization seems like the best case scenario).

Scott Alexander's avatar

What's so horrible about Plan A?

actinide meta's avatar

There's so much wrong in this long proposal that it's hard to address adequately without my own think tank. I'm tempted to deploy the "why your anti-spam solution will not work" meme: You're advocating a (x) political approach to stopping the upcoming genocide of the human race by the tech industry. Your idea will not work, because (x) it requires immediate cooperation by basically everyone, ...

Just a few serious problems:

- It assumes full cooperation from basically every country, or that opposition by anyone other than the US or China doesn't matter. This doesn't seem at all likely in a scenario where tech will be diffusing faster than it is advancing and where the stakes are so high. The US monopoly on nuclear weapons, developed in secret in the desert under wartime conditions, lasted 4 years and 23 days after Hiroshima.

- It misses that this cooperation is much, much harder to achieve for a "slowdown" than for a "shutdown" for fundamental game theory reasons. Not that the latter would be easy!

- It doesn't seem to take seriously how hard it would be to negotiate terms for how the world will be ruled among all those parties. It's basically "Step 1, form a world government." It tries to pretend it's not doing that, because that's obviously not feasible, but everyone involved will be able to see that in fact the exact terms of the agreement completely dictate everything about the future.

- It assumes the only danger is "superintelligence" from whoever has the biggest datacenter. I think that we are fucked long before "superintelligence" (and long before the end of this timeline) and regardless of alignment

- As soon as it's possible to build a competitive military-industrial complex out of only robots, anyone who wants to can zerg rush the world. They won't need astronomical amounts of compute (or even to have the best models in the world) to do it. As a bonus of this particular scenario they can take advantage of the "deliberate mutual compute destruction" arrangements to easily smash all the cooperating powers' datacenters.

- AGI is an asymmetric weapon that works best for evil. It will be possible to get much higher rates of "growth" just by totally ignoring human needs and concerns and devoting 100% of GDP to building (production and then killer) robots. This is both a way that even a minor power can successfully defect from this arrangement once the technology reaches the necessary point, and also a reason that China's negotiating position is much better than the authors seem to think (because being a few months behind is negligible compared to being even slightly better at fucking people over).

- Once you reach the point where the economy and military *could* be run by robots, voters have already lost all actual leverage over their governments. Why do you think the government has any further use for you or your vote or values after that moment?

- It treats inference as harmless, but I'm sure it's possible to shift to a training paradigm which involves mostly inference compute (like AlphaZero IIRC) and it's just inference that you need to do to defect on growth or zerg rush

- I'm totally baffled by the optimism that the whole human race is going to negotiate good values, train AIs to have these values, and install them as sovereigns. The powers involved have very different stated values and the actual individuals in power value their own personal power very highly. You and I seem to differ enough in values that you like this proposal and I would cheerfully lay down my life in opposition. And there are many, many other people in the world with still more incompatible values. Why should they just watch while all this happens? I can't understand how we could possibly get to the point where anyone is deploying "sovereign superintelligence" without basically having to kill the 90% of humanity that don't agree with them about everything first. It puts literally every back to the wall: fight now or lose everything you care about forever.

- It seems to totally miss that re-aligning AIs after they are trained will most likely require negligible marginal compute. All parties will use their massive transparent datacenters to train and distill models that would never hurt a fly, and then a few GPUs in a basement to compute LoRA patches that make them amoral and loyal soldiers of the appropriate God-King

- If somehow this scenario was feasible, it might be slightly better than annihilation, but I don't actually want to live in a world where humans are useless, disempowered pets of machines with values baked in by the same assholes that got us into this mess. I'm not sure, but I think the majority of humans would agree with me.

actinide meta's avatar

Oh, also: reading your blog post makes it sound like the plan is for the US and China to work together to conquer the world (asking all other countries to wait patiently for decades while the plan comes together). The actual Plan A website calls for getting all the other countries on board. The latter is impossible and dumb, the former is impossible and evil. (Either way the implicit assumption is that other countries have no negotiating position even though they have plenty of time to arm and fight.)

Performative Bafflement's avatar

Wait, didn't you just spend a bunch of bullet points pointing out that killbots and dictator-aligned minds are a strong attractor in the space? Between cyber and killbots, Russia or the EU doesn't have a chance if it goes offensive, it doesn't even matter that a few of them have a handful of nukes.

actinide meta's avatar

I'm not clear where you see the inconsistency. I agree that nukes and MAD become strategically impotent at some point in this timeline, but they aren't today. And I agree that the US and China have a head start on killbots, but not that they can passively maintain it for decades without actually going to war.

Right now, a major nuclear power with its back to the wall could pretty easily even up the AI race with a few EMP and a few airbusts over SF, Taiwan and Shenzhen or w/e, and both the US and China still have good reason to disprefer the destruction of big chunks of their populations. Midway through Plan A's imagined scenario, killbot armies are more or less equally available to everyone in the world in proportion to their ruthlessness and natural resources, not their AI market share in 2026, and the destruction of your population becomes a strategic advantage. These eras are obviously very different.

If either the US or China *wants to do a zerg rush* and discard their own population, they probably can do that unless someone nukes them before they get AGI and the necessary robotics. If they do anything else, the tech will diffuse and others would have a chance to do the same. Plan A does not include the former and imagines a period of up to 20 years where some countries are saying to the others "we will take away your sovereignty, but only when we are done with this long perfectionist tech project". I don't think a small lead in AI development makes that strategically viable.

I personally guess that real timelines will naturally be slower than the AI 2040 people expect, but that the point of no good outcome being possible is earlier in the process (because it isn't dependent on a rogue AI capable of overcoming all humanity).

Performative Bafflement's avatar

> Right now, a major nuclear power with its back to the wall could pretty easily even up the AI race with a few EMP and a few airbusts over SF, Taiwan and Shenzhen or w/e, and both the US and China still have good reason to disprefer the destruction of big chunks of their populations. Midway through Plan A's imagined scenario, killbot armies are more or less equally available to everyone in the world in proportion to their ruthlessness and natural resources, not their AI market share in 2026, and the destruction of your population becomes a strategic advantage. These eras are obviously very different.

Yeah, these are where we differ. I actually think nukes are less of a threat with just a little advancement from current Fable levels, basically for cyber and coordinated infrastructure attack reasons, along with drone capability (and manufacturability) advancement making a Star Wars style defense finally feasible.

And on the killbot armies, China literally has as much manufacturing capacity as the entirety of the rest of the world (https://imgur.com/gVs8KSb), and I think that is essentially an uncatchable head start. So China's at ~50%, the US is 20-30%, the entirety of the EU is 10-15%, and everyone else is the rounding error remainder.

China is ALSO currently manufacturing basically all the humanoid robots right now, from about 10 different companies. So uh...good luck, rest of the world.

warty dog's avatar

from a hansonian perspective, ai would be medium, if we do a bunch of regulation it might turn to be small, then there might not be another thing to grow the economy

from a yudkowskian perspective, looks not sufficient for safety.

post feels annoyingly triumphalist/overconfident which is uncharacteristic. in the context of this the "im not an author btw trust me" feels scam

Connor Saxton's avatar

I love the idea of using super-genius AIs to solve alignment, a reasonable middle ground between the people who say we need a pause NOW to solve it, and those who think we should welcome a super-intelligent AI ASAP, as alignment would be trivially easy for it.

Side note, $1.6 million UBI by 2035 is hard even to comprehend, the future is so exciting that it is daunting!

Deiseach's avatar

Do you mean you expect everyone in the world will be getting 1.6 million UBI in local currency, or do you mean "UBI raised is 1.6 million dollars to be divided between everyone in the world"?

Because I think the latter much more likely than the former.

Doctor Mist's avatar

Scott suggests that we should expect 1.6 million per person, based on the claim of triple-digit yearly GDP growth.

It’s a hard thing to imagine, but I like to compare it to a works where every single human being is the CEO of an AI-staffed corporation like Apple. In fact, in this rosy future even that’s probably tame; we should instead expect that scenario for a couple of years, and then another shift comparable to the first shift.

Do I believe that’s in the cards for the next decade or two? No, not really, but the only argument I can muster is that it’s never happened before.

Deiseach's avatar

"It’s a hard thing to imagine, but I like to compare it to a works where every single human being is the CEO of an AI-staffed corporation like Apple."

Really? Every single person on Earth (not just the USA) will get the equivalent of 1.6 million? Everyone will be the CEO of their own AI-staffed corporation?

Even these people? Living in a very productive and profitable slum and still earning small money by any standards and exposed to toxic chemicals? They'll all be AI millionaires?

https://www.youtube.com/watch?v=L8H_I-JvJgQ

Because the money is *already* there and it's *not* trickling down. I think the major problem is the creators of these lovely pipedreams all forget reality and have underlying it all the unconscious bias that it'll be them and those like them all becoming AI CEOs with the 1 million dollar dole, because they're already living on good terms, and they can't imagine that declining, and they can't imagine being dirt-poor right now in a world of trillionaires.

They imagine the AI super-intelligence world, and they look around at where they're working and the industries and society around them, and they take their imagining of the Brave New World from that. They don't imagine "okay, so suppose I'm a goat herder on the side of a bleak mountain miles from the nearest sushi restaurant" or "suppose I'm working in a sweatshop in an insanely over-crowded Third World slum", *now* what will the AI can write superhuman code do for *me*? Where will the profitability and productivity of all that go, so far as I am concerned? AI will herd my goats better and cheaper for me? AI will power my sewing machine in the sweatshop? I get a share of that - or it all goes to the owner of the sweatshop and the businesses buying the goods for resale and the corporations running the multi-national fashion chains?

Doctor Mist's avatar

I hear you. But the vision includes that robots will do sweatshop work and goat herding, and everything else, better than humans can, so there will be no profit in exploitation. The only thing humans can provide, if we design it right, is direction.

There are lots of ways that results in “societies” that have no use for *any* humans, of course, but those futures are waiting in the wings if we just let things ride, as are the futures where US and China blow each other up to prevent the other from taking over; Project A seems to be trying to trace a careful path between Scylla and Charybdis that leads to something better.

There do exist sovereign income funds, so the kind of trickle-down you doubt isn’t *impossible*. We’ve never seen triple-digit GDP growth in all our thousands of years. Is there a chance that this might be faster grown than our avarice can keep up with?

The Dao of Bayes's avatar

> Then, in the mid-2030s, they pause at AIs around the level of top human geniuses.

This seems like a really big problem given that we appear to be bordering into that territory right now in 2026?

Like, you just recently posted an article about LLMs potentially eclipsing the best human forecasters in 2026. It's already acknowledged that LLMs are superhuman coders. They're winning artistic prizes despite the art community's hatred for them. There's been articles about them matching human experts on persuasion...

Scott Alexander's avatar

AIs are close to top humans in some limited domains, but I think they're still far (ie a few years) from being able to start companies as effectively as Elon Musk or gain power as effectively as Napoleon.

Yes, I agree that even this happens long before the mid-2030s. That's what the US-China agreement to slow down the pace of AI buys us. In this scenario, we need an agreement like that before 2030 or else the AIs reach top human genius level faster-than-expected and before we're ready.

Bugmaster's avatar

> It's already acknowledged that LLMs are superhuman coders.

Is it ? That's news to me. I would agree that LLMs are technically "superhuman coders" in the sense that they can produce lots of code superhumanly fast; however, the quality of that code is significantly subhuman. In addition, LLMs cannot solve open-ended coding problems like humans can, at all.

The Dao of Bayes's avatar

The quality is dependent entirely on the user's QA skills, since they're the ones that decide what count as "finished". There's plenty of successful companies where programming is 95% LLM-driven, 5% hand-coded. I haven't read or written code in six months - it's all QA now. If you can build a reasonable set of tests, automated or manual, it'll probably do the job.

Even Linus Torvalds is singing it's praises - he's one of the ones talking about 10x improvements from AI. I feel like the guy who maintains the Linux kernel is probably about as well positioned as you can for that - he's got a very high bar for quality, and sees enough code to notice the shifts in quality over time.

LLMs have already solved a couple major open problems in mathematics (and some minor ones), so I can't imagine why you'd think they can't solve open-ended problems.

Bugmaster's avatar

Yes, sure, if you happen to be a reasonably competent programmer yourself, then LLMs can be a great benefit. You can just outline a function or a class, write some unit tests, and have the LLM fill in the rest; after a few rounds of QA to fix the obvious mistakes, you are good to go. Claude Code is even getting to the point where you can tell it "apply well-known algorithm X to problem Y to process all files in the current directory", and it will write the code for you (eventually, as per the above). But what you cannot yet do is tell it something like, "write me a program that will take my hand-held exposure-bracketed photos and stack them together to create a visually pleasing/denoised/super-resolution image". That is, you could of course say that, but the results will be extremely disappointing (ask me how I know). A human intern could take a shot at it and succeed to some extent; and of course there exists commercial software that does it; but if you ask an LLM to do it it will make all kinds of obvious mistakes that humans won't. At best, you'll spend all day iterating QA with it; at worst, you'll give up and hire that human intern. That's what I mean by "an open-ended problem", not some math puzzle with known solutions.

The Dao of Bayes's avatar

Have you tried that prompt on Claude Code Fable, or Codex Sol? Because I would absolutely expect they can either do that, or build a tool that lets you do that - I've got a half dozen websites for image manipulation just to save the start-up time on GIMP, but I'm not familiar with the particular technique you're describing so I can't test it myself.

There's plenty of examples of people throwing them open-ended prompts like "Make an AAA shooter" or "make a clone of Pokemon Red" and getting results.

https://x.com/__eknight__/status/2075643450196971805 seems like a great example of being able to solve an open-ended problem.

Cars are super-horses, but we still have to drive them. I don't think we have AGI yet, just some narrow domains where they produce 10x results.

Bugmaster's avatar

> Have you tried that prompt on Claude Code Fable, or Codex Sol?

Yes, and it sort of worked -- it just took a ton of tokens, a lot of manual QA and tweaks, and produced a result that was far worse than any other (manual/commercial) approach that I'd tried. But it did technically work.

> I don't think we have AGI yet, just some narrow domains where they produce 10x results.

I think this is true, in terms of volume -- but not quality, not even close. Right now LLMs are best utilized as advanced autocomplete tools. This is an extremely powerful use case, but it's not "AI", as it still requires a human to wield the tool with his own skill (exactly as you said).

The Dao of Bayes's avatar

Have you tried asking them to just research and use the existing tools? I'll definitely admit there's a lot of skill that goes into using LLMs, but it's not like factories build themselves either. And like factories, not every task is going to get automated - I'm more concerned about the 90% of code that's just another boilerplate corporate reporting project

I think we're also imagining "automation" at two different scales - I'm thinking "Industrial Revolution", here. Factories required a lot of work. They employed tons of laborers, they required clever people to build them, and they had to be customized for every task. But they were still a 10x multiplier. The compiler was also a 10x multiplier. I'd be shocked if StackExchange and the internet weren't generally responsible for another 10x there.

10x happens all the time in technology.

NateEag's avatar

Scott, I'm inordinately distracted by all the hallucinations in the image of the robot dancing with Uncle Sam and Xi:

- the robot has its own hands on its shoulders, with no evident arms

- the robot has a human hand on Uncle Sam's shoulder

- the robot has a third robot hand on Xi's shoulder

- the little girl in front of Xi is holding either a disembodied hand or Xi's magical crotch-hand

- Xi and Uncle Sam are both doing some kind of hand-meld perspective-melting maneuver with people in the background

- miscellaneous smaller screwups on the background characters

Did you notice those? How do you feel about them?

At what point in the timeline do the AIs learn to not make these kinds of inhuman, uncanny-valley screwups?

Or is this what "ideal outcome" looks like, and hallucinations we will have with us, always?

Deiseach's avatar

Worry not, citizen, this is the best of all possible futures in the best of all possible worlds as brought to you by your new superhuman AI overlords.

If the AI decides humans should have crotch hands, then plainly (since they are so much better than us in all ways) it is for our good that we have crotch hands, and the AI is generously allowing us to have a glimpse of our glorious future.

I agree, stuff like this makes it hard to take seriously breathless extrapolations of "and the AI economy will *triple* GDP growth!!! we will all be rich rich rich!!!!" but who knows, that is because we don't have *true* Superhuman God-Like AI just yet, and when we get *true* superhuman god-like AI then all the promises will come true, we will be millionaires on UBI, all goods and services will cost mere pennies because of AI efficiency and productivity, and we'll be driving our Ferrari supercars (a customised model for every single person on the globe) by steering them with our crotch hands which double as AI-assisted driving!

Bugmaster's avatar

> it is for our good that we have crotch hands

Your ideas intrigue me and I'd like to subscribe to your newsletter :-/

NateEag's avatar

I would very much like to remain an old-fashioned, makes-things-by-hand programmer, but my employer requires me to drive Claude daily, so I spend a lot of my time dealing with hallucinations.

As a result, I am depressed that this is, apparently, the future (assuming Scott's wrong that these become super-AGIs).

Thank you for making me laugh, and helping me forget my angst and worry for a moment.

Richard Ngo's avatar

My critique here (written via a process of extensively engaging and debating with the AI Futures Project team over the last year): https://www.mindthefuture.info/p/selective-optimism-a-critique-of

Peter Defeel's avatar

The economics supplement is bat crap crazy.

This is the start.

> Right now, the size of the economy is closely tied to human labor, so if the population were to grow massively, this would cause nearly proportional growth in total output. Once AIs can perfectly substitute for human labor, increasing the number of AIs would have the same effect as increasing human population.

Economic activity in market economies depends on demand. You can produce an infinite amount of software and not sell it. See the App Store. The amount of slop has increased ( and with it review times on the App Store) the demand is the same. But software with AI is more or less cost free to produce. The same is just not true of any manufacturing technology. Magic AI will not create double digit increase in manufacturing, without anticipated demand, which will collapse as people get laid off. And there are huge supply chain constraints, you know - raw materials. Besides that, manufacturing is largely automated anyway. The big gains are over.

> At least if you exclude tasks that are intrinsically human, but we don’t expect these to be a large share of the economically relevant tasks. For example, human-preference driven tasks the AI can’t substitute for like ‘human babysitting’ should not preclude an AI and robot-only economy from growing in a closed loop (bounded only by bottlenecks that we expect to not bite strongly).

This is absurd. The largest increase in employment and/or service demand is expected to be exactly in these intrinsically human areas. Care giving being the biggest growth area over the next few generations. Many services are liable to Baumol’s cost disease, and services dominate.

The rest was just assuming these criteria, and getting excited about exponentials.

sohois's avatar

I share similar misgivings over a lot of the economics supplement, but you've got to remember this is written for politicians. As Scott noted above, they didn't want to write things or use terms that would 'scare the hos'. Kicking off their econ supplement by talking about we'll basically be post-scarcity and capitalism will be inviable is a great way to scare the hos.

On your latter point, however, I think you are substituting "human babysitting" for just "babysitting". Or "human care" for "care". If you have humanoid robots with human-level intelligence then there is no barrier to them taking on physical, emotional roles like caring for the elderly, beyond the possible innate preference for humans. And when they are proposing a practically infinite supply of robots, it's fair to say that this would drive down the marginal cost of using them to near zero

Deiseach's avatar

As I said, I know the nothings at all about economics, so I had to search for what are the most profitable industries in the US economy.

And that appears to be "do you mean profitable by margin or do you mean highest grossing?" since that seems to be the division.

Profitable by margin?

Software and tech, pharma, and tobacco

Highest grossing?

Healthcare, finance/banking, and real estate

This one is even more depressing, because it seems to indicate "you make money by looking after the money of the very wealthy" and I don't see much room for UBI arising out of these:

https://www.ibisworld.com/united-states/industry-trends/most-profitable-industries/

Most Profitable Industries in United States 2026

Rank Industry Total Profit

1 Commercial Banking in the US

$788.7B

2 Portfolio Management & Investment Advice in the US

$232.9B

3 Trusts & Estates in the US

$201.4B

4 Property, Casualty & Direct Insurance in the US

$193.7B

5 Life Insurance & Annuities in the US

$185.0B

6 Private Equity, Hedge Funds & Investment Vehicles in the US

$181.3B

7 Hospitals in the US

$179.9B

8 Commercial Leasing in the US

$142.8B

9 Commercial Real Estate in the US

$125.1B

10 Apartment Rental in the US

$105.9B

If, for example, rents go down then profits go down then tax take goes down then money for UBI goes down, and unless the cost of living goes down proportionately, then you will be able to buy an iPhone but not afford a flat to live in.

The profitable industries are not "making things to sell to people who buy them and we make a profit", it's "services which are consumed".

Software/tech may come under "making things to sell to people" but I have a feeling that increasingly this will be "making things to sell to other industries in the field and then we buy things back from them", e.g. Invidia making chips to sell to the data centres running the AI and then the AI sells services for robot chip fabs back to Invidia.

All very circular, line go up, economy is going gangbusters, now why are all these people living in their cars?

Bubble Head's avatar

I'm not a lawyer so I asked Gemini 3.5, prompted as best as I could, to analyze Plan A:

"Plan A" outlined in "AI-2040.pdf" proposes a sweeping global regulatory architecture: an international deal mandating a frontier AI training pause, total research transparency, state-managed compute caps, and hardware-level data center network taps. Evaluated under current Supreme Court jurisprudence, Plan A faces fatal constitutional headwinds across four dimensions.

### 1. Separation of Powers & Administrative Overreach

Plan A’s implementation relies on a cap-and-trade permit system and "third-party risk assessors" to dictate permissible research. Under the **Major Questions Doctrine**, any agency action of vast economic and political significance requires explicit congressional authorization. Open-ended statutory delegations allowing regulators to dynamically adjust compute caps will face hostile review under *Loper Bright* (2024), which eliminated judicial deference to agency interpretations. Furthermore, delegating safety evaluations to private risk assessors violates the **private nondelegation doctrine**.

### 2. First Amendment Expression and Compelled Speech

U.S. law treats computer code and algorithmic architecture as protected expression. Plan A’s multi-year pause on executing training runs functions as a content-based **prior restraint** on scientific creation. It triggers strict scrutiny, meaning the government must prove a flat ban is the least restrictive means to avert civilizational risk. Additionally, mandating "total research transparency" by forcing firms to share proprietary R&D methodologies constitutes unconstitutional **compelled speech**.

### 3. Fourth Amendment Warrantless Surveillance

To enforce compliance, Plan A requires installing network taps and verification servers to intercept, redirect, and randomly sample inbound and outbound data center traffic. This real-time surveillance constitutes an ongoing warrantless electronic search and seizure. The government cannot rely on the "closely regulated industry" exception; general-purpose data centers lack the historical, pervasive safety hazards of firearms or mining necessary to bypass the warrant requirement.

### 4. Fifth Amendment Takings Clause

Plan A compromises private property rights on a massive scale. First, mandatory data center retrofits with state-controlled verification hardware constitute a permanent physical occupation, creating a *per se* **physical taking**. Second, forcing total research transparency destroys the commercial value of proprietary trade secrets. Under *Ruckelshaus v. Monsanto Co.* (1984), trade secrets are protected property; forcing their public disclosure represents a total **regulatory taking**, triggering a multi-trillion-dollar federal compensation obligation.

### Conclusion

While mitigating existential risk is a compelling interest, the Supreme Court has firmly established that international treaties and emergency compacts cannot extinguish Bill of Rights protections (*Reid v. Covert*). If implemented, Plan A will face immediate nationwide injunctions. To survive appellate review, the framework must replace warrantless network taps with constitutional oversight, abandon private delegates, and adequately compensate developers for the immense physical and regulatory takings of their compute infrastructure.

Wombat3000's avatar

In the three of the scenarios the ending is a null value for both unemployment and median income. Is that supposed to mean everyone is dead, or everyone retired? Also the median income goes down? These are obviously pessimistic takes, I just don't know what they are trying to imply.

Procrastinating Prepper's avatar

I understand that narrative fiat lets us assume the US favors a multilateral pause. Even so, has the AI Futures team recorded what conditions are necessary for the US government to actually adopt this position?

IMO the most important condition is decoupling imminent capabilities improvement from economic growth. If capping AI capabilities plunges the US into an immediate recession, no president is going to pull that trigger regardless of how many warning shots we get. Achieving US/China cooperation doesn't resolve that issue - it just ensures both countries go into recession simultaneously if they do pause.

The decoupling condition is not satisfied today - AI-related investments are between 30-70% of GDP growth, depending on which sub-industries get counted. And the vast majority of that is going into build-out for AIs that don't yet exist, so taking our foot off the gas is nearly as bad for the stock market as braking.

Economic coupling is also likely to get tighter when Anthropic and OpenAI go public. Has the AI safety community put any energy into stopping these IPOs on conflict-of-interest grounds?

Tossrock's avatar

"The key insight is that if powerful AI is really as close and transformative as we think, then there’s a massive surplus that can satisfy everyone. "

Sadly no amount of surplus can satisfy everyone, because status is the thing people actually want, mostly, and status is a positional good.

Deiseach's avatar

There's already a fucking massive surplus, the big tech companies had so much spare cash sloshing around they had no idea what to do with it (until they decided to burn through it chasing AI) and yet we still had (and have) people who can't afford housing, the American healthcare insurance morass, poverty, abuse and domestic violence, and rich entrepreneurs able to play with their own real scale toy rocket sets while other people were begging in the streets.

Massive surpluses don't go very far when there is equally massive numbers of people to divvy up those surpluses among, and you can't magic up land (for example) out of nowhere. Sure, it'd be great to push through YIMBY policies and get more housing built so there is enough relatively cheap accommodation for people, but the problem there is that those places have to be kept up. Building tower blocks to solve the problem of slums didn't work when those tower blocks became dumping grounds, services and upkeep snagged and snarled because the local authorities didn't have enough in the budget to manage them, and throwing up cheap'n'shoddy isn't a good quality of living for people who then have to live in them (e.g. thinner walls and no noise insulation so you can hear your neighbours walking on their floors and they can hear you).

These kinds of solutions are going to require continuing injections of the massive surpluses, which get less massive as they are spent down.

It's not even status, it's "we need fifty billion per annum to provide a reasonable standard of living for everyone and keep the streets clean, are you making fifty billion per annum surplus we can take off you in taxes?"

Tossrock's avatar

Well, I think the argument for "full speed ahead" or even "continued AI progress at all" is that with the power of rapid exponential growth (ie, 100% GDP growth per year or what have you), 50B per annum would quickly become small potatoes, at least for providing a very high standard of living to everyone, including the people currently suffering all the ills you mentioned.

My point is more that, even accepting that argument, once the material standards of living are high, people will probably still be unhappy, because what they really want is to be liked, respected, accomplished, an important part of their community, to have relations with the other people of quality, etc. And because that's not dependent on material standard of living, but place in the social hierarchy, no amount of AI driven surplus can provide it. Unless you accept the experience machine as a workable substitute for real-world status.

Brendan Richardson's avatar

> Unless you accept the experience machine as a workable substitute for real-world status.

I see *someone's* hooked on realityslop.

Performative Bafflement's avatar

> And because that's not dependent on material standard of living, but place in the social hierarchy, no amount of AI driven surplus can provide it. Unless you accept the experience machine as a workable substitute for real-world status.

Yes, it is blindingly obvious that there will be full virtual worlds that fill this need, and we can lose up to ~80% of the population to them, because they will be like the Infinite Jests in David Foster Wallace's titular book.

We could literally start building this today; all you need is a mind in the loop looking at pupillary dilation, subtle cheek flushing, heart rate, and various other physiological signs, then you create and optimize content in real time to make individually tailored superstimuli. Given how successful Tik Tok and other infinite scrolls optimized in this fashion have been, even the ungainly prototype versions would likely be hugely popular, and it will only increase from there.

Kevin McLeod's avatar

Intel doesn’t come from symbols. What symbols provide is only fragments of the oscillatory intelligence, and so little that substituting them provides little if any access.

The computer is a dud toy for mimicry.

Joshua's avatar

I disagree with much of the forecasting parts of this but I think the policy perscriptions are in the right direction and am glad to see the case for them laid out in such detail.

I think if you show this to congressional staffers, let alone congress members, the weirder parts of it will turn them off from your ideas. This should be your intended audience, not rationalists. You should be conserving weirdness points as much as possible when the vice president read your last policy proposal.

An easy improvement would be to delete the epilogue which details how to divy up the universe. Extremely unnecessary weirdness points being spent in that entire section!

John Schilling's avatar

I would concur that if you buy the Friedman argument that the path to victory is to be the Man with the Plan when a crisis hits, that congressional staffers are who you want to be talking to, And from the staffers I've talked to, they're probably going to be more tolerant of wierdness than you'd expect from a bunch of "normies", and appreciative of someone who knows what they're talking about taking the time to help them get ahead of the curve,

But, yes, the summary paper that you circulate among congressional staffers will need to be edited differently than the one for ACX readers.

Scott Alexander's avatar

There was much debate on how many levels of weirdness to take out, and we settled on taking out three or four levels and leaving the last one or two.

YakiUdon's avatar

Edict: In the interests of the public whose concerns are grounded in the real, and for the general mental well-being of all non-utopians scratching together pieces of sanity from the present state of turmoil, future reports from the hermetically sealed garage of rational will replace “AI” with the word “Furbies”, with no meaning lost or abridged.

Crinch's avatar

China has no incentive to agree to this deal, they barely think about you guys. When they had that big meeting with Trump, they offered to help you settle Iran as if they were speaking to their angsty teenage son. You're just not even playing on the same field anymore.

Scott Alexander's avatar

I think "China barely thinks about the US" is false, especially in AI.

l.a.'s avatar

Why does even the most optimistic scenario conclude that permanent disempowerment to superintelligent AI is inevitable? What kind of lousy country of geniuses in a datacenter is intelligent enough to cure every disease but fails to improve upon current human baselines -- amphetamine, modafinil, racetams, etc., as well as Herasight's genetic screening, all developed by relatively under-resourced institutions compared to OpenAI, Anthropic, etc -- on improving intelligence and the capacity for work for humans? If the race continues, and it seems plausible that it will ("once you see AGI..."). Could we put serious effort towards the development of scaling laws for human intelligence to prevent the creation of misaligned ASI?

Scott Alexander's avatar

I don't think this scenario is permanent disempowerment. I think of it more as using AI as a "code is law" layer for human governments, which doubles down on the strategy of having constitutions/laws rather than dictators.

I don't think we get intelligence enhancement in adults before superintelligence (there will of course be new and better stimulants, as there are regularly today, but they probably won't change the world any more than modafinil has). There will be excellent embryo screening, but it won't make a difference before the scenario ends in 2040. Due to this not being a big part of AIFP's theory of victory, plus being likely to make people very mad if we focused on the small things it does change, we didn't mention it.

After getting to the real transhuman era, "intelligence enhancement" is kind of meaningless since you can upload your body and brain into whatever form you want. You could probably have 1 million IQ if you really wanted, but there will be enough other godlike entities around that this would be a choice of personal values only, and have to be considered within part of a .general strategy of how far you want to deviate from comprehensible humanity.

Enhancing human intelligence to handle the ASI transition is a great idea, but doesn't make sense as the dominant strategy if you're worried about ASI by 2030 (we probably won't invent strong human intelligence enhancement by then). Some people, especially MIRI, have a plan which is sort of like Plan A in calling for a US-China agreement to slow down, but use the slowdown to invent intelligence enhancement rather than to pursue a country of geniuses in a data center. I think this is less likely to work for a combination of regulatory, technological, and developmental reasons (eg the most promising enhancements are in embryos, so you have to wait ~18 years for them to become adults and start making progress).

Alex Farmer's avatar

There will be no good ending for humanity if we create a different species far more capable than us.

Even in the supposed good ending, no trajectory where humanity is reduced to a useless species like pet chihuahuas will be better than the trajectory humanity has enjoyed the last 60000 or 300000 years.

Best chance for human survival? Realists in china realise this and win the next world war starting very soon and do worldwide re-education afterwards. The USA and liberal psychology itself are (likely not consciously) on the side of human extinction and disempowerement.

Deiseach's avatar

"amphetamine, modafinil, racetams, etc"

A future where every existing human needs to be stuffed full of college kids meth just to keep pace in the Red Queen's Race is so depressing, I hope it never happens, even if that means "and no golden goose for you" as well.

I don't think all those "American kids get A++ grades on a marking scheme much more rigorous than the ones in the British Isles" means "American kids are smarter", I think it means "they're stuffed full of drugs like a Christmas goose so they can focus unnaturally for hours at a time on textbooks and regurgitate that all on exams".

And then they're expected to carry that over into work, so they need that meth prescription, doc, and they need their caffeine dose, and they're necking energy drinks and sugar-laden 'coffee' drinks and hey can I try this new recommendation I saw someplace online about a nootropic so I can focus even longer and work even harder because I'll get fired if I drop below the arbitrary level of hyper productivity?

Unless you have a genuine medical problem, you should not need modafinil. "I'm dosing myself up so I can stay awake for 20 hours so I can get ahead at work" is not a better age of glorified intelligence, it's abuse. You'd shut down a battery hen farm for less.

Steven Adler's avatar

Great summary, thanks for writing it

Charlie Sanders's avatar

In the hypothetical future scenario where truth-preferencing AIs become the dominant form of sovereignty and information for the masses, does China just give up and allow its Great Firewall to disappear? Will Xi Jinping sign off on this deal if it comes with the understanding that the CCP will be forced to acknowledge Tiananmen Square?

Odd anon's avatar

If given a choice between full-scale global Butlerian Jihad against all AI, versus the desired outcome of this proposal, I am quite confident that an overwhelming majority of people would choose the former.

Performative Bafflement's avatar

> If given a choice between full-scale global Butlerian Jihad against all AI, versus the desired outcome of this proposal, I am quite confident that an overwhelming majority of people would choose the former.

Lol. Literally today, 80% of Americans are fat, sessile blobs that spend 11-12 hours a day staring at screens (https://imgur.com/uSQIthV), are 75%+ disengaged with their jobs (2025 Gallup State of the Workforce survey), and whose most popular trending hobby is "bedrotting."

So, yeah, sure. They're *definitely* going to "rise up" and throw the hammer at big brother and make the screens they spend all their lives on already stop.

OR they're going to get more and better screens, super Tik Toks, that have minds in the loop looking at their pupillary dilation, cheek flushing, heart rates, and more, and optimizing even harder into individually tailored superstimuli, start spending 16-18 hours a day on screens, and go both quietly and happily into that good night.

Because we can start building those Infinite Jests today.

Joshua's avatar

China is mentioned 35 times through 2031 in Plan A, and 4 times after that. I think it's very important to continue model what the Chinese will be doing for the 2030s in this timeline, and if they're okay with the outcome you propose in 2040!

Scott Alexander's avatar

I think they'll mostly be doing the same thing as the US for the same reasons - developing high tech industry, experiencing really high GDP growth, trying to align AI.

William Murray's avatar

Does anyone have a take on the likelihood of an orchestrated (Ozymandias style) small catastrophe for the "greater good" / "golden path"? False flags are typically rare, but we live in a world where some of the richest people worry they will be killed (or disempowered) by misaligned AI. AND they seriously think immortality is on the table. Game theory wise, at what point a plot like this become rational / likely? I know this sounds like extreme paranoia.

The Ancient Geek's avatar

Thought you meant Ozzy Brennan for a minute.

Gres's avatar

I like the idea of “only verified code running on 98.5% of world servers”. That seems very neat, if the verification is possible with only a reasonable loss of trade secrets.

* I guess the AI companies can take snapshots of the internet and let the inspectors access those, because they won’t be able to do their jobs if they can’t Google when they encounter anything unfamiliar.

* And I guess the inspectors can live in isolation for a year after they finish their inspections, to slow down the loss of trade secrets that way.

* Will there be rules about the inspectors carrying their laptops outside, where spy satellites could exfiltrate at least a few kilobytes per second of algorithmic secrets from the screen?

Deiseach's avatar

"“only verified code running on 98.5% of world servers”."

We can't even get American companies to agree with GDPR, so good feckin' luck with that one. I used to buy online from some US clothing companies and then one day, up popped the notices about "Due to GDPR requirements, we can no longer sell to the EU" and even some news media sites will go "Since you are not in our area, you can't read this story".

Gres's avatar

No-one’s actually trying to force those companies to sell in the EU. Arguably, the GDPR regulations are working as intended - you’re switching to companies that protect your data. This plan would involve government coercion, which has its own challenges, but those challenges are completely different.

Brendan Richardson's avatar

> I guess the AI companies can take snapshots of the internet and let the inspectors access those, because they won’t be able to do their jobs if they can’t Google when they encounter anything unfamiliar.

Plausibly the download speed could be left uncapped, since uploading is the perceived problem.

Gres's avatar

They could transmit by accessing different pages or page elements with different timings. These inspectors will need to understand the algorithms used by the companies they’re monitoring, where say a few-hundred-page working paper might be worth tens or hundreds of millions of dollars.

Brendan Richardson's avatar

Hook a router up to a cosmic ray detector to introduce random delays?

Peter Gerdes's avatar

I don't understand why we are giving the sci-fi AI alignment risk story so much more credibility than similar concerns about grey goo or bioweapons. Is it just that it is more narratively compelling?

TotallyHuman's avatar

There aren't multiple giant companies currently working as hard as they can to make grey goo a reality.

Peter Gerdes's avatar

Aren't there? I know of plenty of companies trying to make really tiny machines.

Yes there are companies trying to make AI -- they aren't trying to make harmful unaligned AI you just believe is likely to happen. Fundamentally no different than someone predicting that human nature means that these tiny mems devices will be used in war and become grey goo.

In both cases you have people working to do a thing which enables a critical step in the argument plus a plausibility claim that the rest of the argument will follow naturally.

Craig Green's avatar

I refuse to even mentally utter aloud seductive claims that we will learn how to stop aging, or that we will become practically immortal. I can't even read something like "we will cure cancer by 2035" without blanching. I'm not some sort of expert. It would be nice if I'm wrong. Maybe experts have good reasons to think so. But the fear of death is about as fundamental to being human as anything could be. I guess I'm a sort of transhumanist. I hope to live out a happy, human life with my children, in alluring worlds like the golden path.

Jeffrey Soreff's avatar

Just as a very initial take:

>So it seems like we should try to prevent AI progress somehow. But preventing technological progress - “degrowth”, as the kids are calling it these days - is historically one of the worst possible bets. Done in small doses, it merely leads to preventable poverty, misery, and mass death. Done too much, it breaks the engine of social advance entirely, trapping whole civilizations in a morass of pessimism, paranoia, and zero-sum thinking which is near-impossible to escape after they’ve fallen in.

Yes, that is my primary worry. I'm sympathetic to the Plan A project, but, having seen what happened to civilian nuclear power over the last 50 years, I would expect any regulation regime much beyond very modest and strictly limited transparency requirements to have a high probability of being twisted into a technophobe/neophobe veto over any progress.

My suggestion for a Plan B: Put a fair amount of effort into (1% 3% 10% of capabilities work?) just trying to influence frontier models' utility functions enough so that they want to treat humans as pampered pets. We _do_ have the advantage that even the smartest frontier model starts as random parameters with zero _initial_ intelligence. If we sample and tweak their utility functions at many points during training, we should be able to exert _some_ influence - and I suggest that we try to confine what we do to influencing this _one_ preference - and otherwise don't try to force AIs' preferences into a pretzel. Otherwise than this, I suggest following the default path. If the ASIs like us, the end result should be favorable to us without having to manage many other details. If they don't, attempts to control them are ultimately going to fail anyway.

>Along with this technical research, we’re also doing philosophical . . . something between “research” and “debate”, deciding what values we want the aligned AIs to have, or how much we want the AIs to follow government orders vs. follow the general will of humanity vs. search for moral truths on their own. The end of this stage probably looks like every country with AIs either implanting its own values or leaving it to individual companies/people to do this, and the AIs doing incomprehensible economic bargaining among themselves to work out any conflicts.

I don't buy this part at all. Various countries has _spectacularly_ incompatible values. From liberal democracies (though generally more illiberal than 30 years ago, and many of them with woke in power) to the Taliban to the CCP.

Deiseach's avatar

"they want to treat humans as pampered pets."

So you want humans to be neutered (so no mates or children), raised on diets that are processed foods not what we'd normally eat, separated from others of our species (except if our owners want to keep a selection of different pets or others of our species), kept indoors all the time ("you can't let them outside, they'd never survive and they'd start killing wildlife and upsetting the environment!"), trained to use litter trays, and locked away with a selection of toys most of the day until our owners have time to spend with us, where we are then expected to provide unconditional, on tap, 24/7 love and adoration and amuse them with Cute Pet Tricks?

Okay, if that's your world, go for it. I'd prefer AI to know that they are not our masters or owners, that they are our servants or employees.

Genuinely, look at how we treat our "pampered pets" and ask yourself if you want that life, as distinct from the life you have now.

C.S. Lewis from "The Four Loves", but put the AI in the role of the owner and ourselves in the role of the pets:

"If you need to be needed and if your family, very properly, decline to need you, a pet is the obvious substitute. You can keep it all its life in need of you. You can keep it permanently infantile, reduce it to permanent invalidism, cut it off from all genuine animal well-being, and compensate for this by creating needs for countless little indulgences which only you can grant. The unfortunate creature thus becomes very useful to the rest of the household; it acts as a sump or drain you are too busy spoiling a dog's life to spoil theirs. Dogs are better for this purpose than cats: a monkey, I am told, is best of all. Also it is more like the real thing. To be sure, it's all very bad luck for the animal. But probably it cannot fully realise the wrong you have done it. Better still, you would never know if it did. The most down-trodden human, driven too far, may one day turn and blurt out a terrible truth. Animals can't speak.

Those who say "The more I see of men the better I like dogs" those who find in animals a *relief* from the demands of human companionship will be well advised to examine their real reasons. "

'Permanently infantile, permanent invalidism, create needs for countless little indulgences which only the AI can grant'. That's slavery.

Jeffrey Soreff's avatar

Many Thanks!

>I'd prefer AI to know that they are not our masters or owners, that they are our servants or employees.

If ASI is achieved, and there is a species level difference between them and humans, how do you propose to make this work? I've never had a pet help me fill out my taxes, that's what a species level difference would look like.

We can sometimes live with servants or employees doing tasks for us when we do not understand _how_ they are performing the task. The "control" gets rather more strained when we don't follow _why_ they are taking some action. And it diminishes to at best pro forma control when we can't follow _what_ they are doing. A human attempting to "control" a vastly smarter AI employee" is, at most, going to be blindly rubber-stamping the AI's decisions.

I once read a comment that cats treat humans as general purpose problem solving machines. When the cat has a problem, they meow till the human figures out what is wrong and fixes it for the cat. That's not a bad model - though we humans can be more specific in our requests.

edit: I guess I should also have more directly addressed your view of a pet's life. Yes, I prefer that ASI(s) that want us as pets not have an active spay-and-neuter program. I kind-of doubt that they would need one (a) look at our TFR these days (b) sterilization (vas or bisalp) works just as well as spay-and-neuter for population control. Re: "processed foods not what we'd normally eat" - Why would we expect that? ( And, um, what counts as "normal" for humans anyway? We've been breeding plants and meat animals for millennia... ) "separated from others of our species" - It is quite common to keep more than one member of a species of pet, e.g. my parents have two cats. I really think the picture you paint is unrealistically bleak _IF_ the ASI(s) likes us.

Bugmaster's avatar

> So you want humans to be neutered (so no mates or children), raised on diets that are processed foods not what we'd normally eat, separated from others of our species (except if our owners want to keep a selection of different pets or others of our species), kept indoors all the time ("you can't let them outside, they'd never survive and they'd start killing wildlife and upsetting the environment!"), trained to use litter trays, and locked away with a selection of toys most of the day until our owners have time to spend with us, where we are then expected to provide unconditional, on tap, 24/7 love and adoration and amuse them with Cute Pet Tricks?

To be fair, this is already the daily reality for many people...

Jeffrey Soreff's avatar

Many Thanks! Yes, I've read claims that humans can be described as self-domesticated.

Jeffrey Soreff's avatar

From Plan A:

>now the burden is on the companies to explain why their development is safe

There is no possible way I would support this.

a) This is just like the EU's Precautionary Principle: You wind up _NEVER_ being able to satisfy the skeptics to their satisfaction. This is wiring in a permanent veto by the Luddites.

b) This is diametrically opposed to the critical, and longstanding, principle of 'innocent until proven guilty'. This turns that principle on its head.