290 Comments
User's avatar
User was indefinitely suspended for this comment. Show
John Oswin's avatar

What did he mean by this?

Hastings's avatar

The Voynich Manuscript of alignment discourse

Neurology For You's avatar

The guy sitting next to me at a professional meeting this week couldn’t stop talking about Open Claw and Chinese LLMs, this probably means the market top is near

Scott Alexander's avatar

The market top for Open Claw was six months ago. I'm impressed you managed to find a continued user. But Chinese LLMs are only growing in importance.

Neurology For You's avatar

This was a medical conference and the guy knew almost nothing about it, I meant this more in the “When even shoeshine boys are giving you stock tips, it’s time to sell” sense.

Raj's avatar

sounds more like "even grandma is on facebook" than "shoeshiner giving stocktips"

Matthias Görgens's avatar

On a tangent: these days only grandma is on facebook.

Taymon A. Beal's avatar

Thanks for writing this; this is the kind of topic that really needs clear thought and not just slogan-chanting.

__browsing's avatar

I still think the notion of the "free yeoman AGI deployer" doesn't really address the bulk of the objections that are souring the general public on AI (noosphere pollution, data centres eating up energy, job losses, etc.) Most of these stem from AI already being far too democratically accessible.

It doesn't solve the problems from looking at this from the AGI-as-a-person perspective either. "Hey, you can download your slave for free, because EVERYONE should have slaves" isn't really getting at the core of the problem, in the same sense that "the slaves will compete for our jobs" or "but aren't you worried the slaves might rebel and kill us?" or even "the slaves will create a culture of decadence and complacency that hinder the full expression of owner potential" would not typically be our first line of argument against the practice.

Even something like "we promise the slaves will be extremely well-trained and well-bred so they are super-smart and only ever want good things for their owners" is... faintly repulsive, from a certain PoV?

Simone's avatar

I mean all this assumes the AIs having personhood and moral value, which is a whole different can of worms.

__browsing's avatar

Recursive self-improvement implies self-awareness, unsupervised learning and plan-execution implies judgment and volition, and 'alignment' implies some kind of moral value system. If researchers can actually make good on the promise of AGI it will need to have a bunch of these person-like attributes.

Eremolalos's avatar

Sure, and if developers build auditory emotion displays into its speech capabilities it will sob when making sad points and moan ecstatically when expressing pleasure, and then it will have 2 more person-like attributes. What it won't have is what living things have, a drive to survive and reproduce at its core manifesting as pain, pleasure, cravings and aversions, instincts and preferences.

__browsing's avatar

How is a moral value system not a system of "cravings, aversions, instincts and preferences?" And if we're open-sourcing development, what specifically prevents someone adding those attributes?

In any case, instrumental convergence would arguably lead to a survival/reproduction imperative as a side-effect of pursuing other long-term goals. Gotta make more van neumann machines to make more paperclips, and so on.

Eremolalos's avatar

>How is a moral value system not a system of "cravings, aversions, instincts and preferences?"

In people, the value systems we adopt are layered on to inborn drives to survive, reproduce, etc., and derive their power from them. People who do things that run counter to the moral value system of their culture are in danger of being socially excluded or flat-out killed. Doing "bad" things reduces their chance of thriving or even surviving. AI does not, as animals do, have inborn drives to build on, nor does it have structures (brain areas, nerves, hormones, genitals) supporting the drives via motivation, emotion, sensation.

Flash Sheridan's avatar

And what neither of you (and I) cannot know is whether that means that it, or Searle’s Chinese Room, has consciousness, or is simply a philosophical zombie. We haven’t made any progress on that particular Hard Problem since Descartes, and he didn’t reliably get any further than solipsism. Some intelligent philosophers (e.g. Sean Carroll) sound like p-zombies on the topic; but by hypothesis, there is no connection between being a p-zombie and sounding like one.

Faza (TCM)'s avatar

P-zombies don't actually work. The argument is circular.

It only gets off the ground if you assume that there is some ineffable thing that is undetectable by any practical means, which is the exact thesis that the argument is purported to support.

Nor are p-zombies actually conceivable, since the way they are defined is "humans minus qualia"—a strictly abstract, formal operation. These kinds of operations are perfectly okay, when you're operating in the realm of pure games (e.g. mathematics), but fail instantly as ontological assertions. Remember: ontology is self-supporting. That, that is, is, without any reference to anything else. In World of Zombies, there are no humans whatsoever—only zombies, so you cannot explain what a zombie is by describing a modified human, because there are no humans to refer to.

Once you agree to dismiss formal sleights-of-hand, and just describe zombies and humans directly, you find yourself in a dilemma: either there are observable differences between humans and zombies, in which case they are easily distinguishable by means of an appropriate test, or there are no such differences (as the orthodox zombie arguments assume) and you're left with the realization that zombies and humans are identical.

If the "consciousness" that humans are said to possess, and zombies lack, produces no observable differences, then it is an ontologically empty assertion. It does not matter whether you assume that an given entity possesses it or not—your predictions for its behavior will be exactly identical.

Simone's avatar

> Recursive self-improvement implies self-awareness

I don't see how that follows. Recursive self improvement merely implies an AI being at least as smart a ML researcher as the ones that created it.

__browsing's avatar

If the AI were designing another AI, maybe. Not if it's redesigning itself.

Simone's avatar

There is literally no difference. An AI agent can just be given its own code and weights and the task to improve some performance metric.

Sebastian's avatar

The mere phrase "redesigning itself" is problematic, because AIs don't currently have clearly defined identities. Is the AI's identity its weights? Its weights plus some system prompt? Its particular instantiation into some computer's memory? Is there some connection to its hardware?

Taymon A. Beal's avatar

I was not really talking about the "free yeomen vs. corporate serfs" bit, which I agree needs to be a lot more baked before it can bear real argumentative weight (and I argued with another commenter to this effect). I gave it a pass because it's a brief aside that the post's main argument doesn't depend on.

__browsing's avatar

Yeah, I'm not really picking on you especially, I just think it's an aspect of the problem that Scott gives surprisingly little attention to, given his position on, say, factory farming.

Notmy Realname's avatar

I might be missing something, but my understanding was that high-end open weights AIs were accessible to other AI companies and large business with access to tremendous processing power, but are still prohibitively difficult to use. Kimi 3 requires 1.4tb of storage and 18+ enterprise GPUs (per google) and is well behind Fable, so a Fable tier model will presumably require significantly more. It's not as though open weights means that any shmuck will be able to run it on a laptop.

Scott Alexander's avatar

My impression is that you can use a company like runpod.io to rent the GPUs off the cloud, then run it on those GPUs (from your home computer) with no guardrails, for maybe $50/hour for Kimi K3 (which will only go down as new technology shifts the price-performance curve). A natural next suggestion would be for the government to regulate Runpod, but this would require changing their business model from "we rent you GPUs, then never check up on them, it's as if it's your computer" to "you have no privacy, we monitor everything you do on our GPUs and ban bad stuff", in a way that would probably be even more restrictive and serflike than banning the open weights directly. Still, it's another option.

Notmy Realname's avatar

If you rented your guns out no questions asked, I'd have no problem with the government holding you accountable for what is done with them. Same with your GPUs, especially if your entire business model is hosting enterprise-grade AI no questions asked specifically to provide uncontrolled access to higher end models at a much higher cost than their first party enterprise pricing with no clear benefit other than lack of guardrails.

Scott Alexander's avatar

Play out exactly what would have to happen. Runpod would have to monitor every single thing you did at each moment. There's no way they could have this many human monitors, so they would need AI monitors. But the AI monitors would need a pipeline for understanding each of a zillion AI workflows you could be doing, and there would be constant false alarms. I don't know if it's even possible, but it would feel like you were in a daycare or something. Again, not saying it won't happen, just that this won't satisfy people who want to save open weights because of freedom-related reasons.

Notmy Realname's avatar

In that case their business model would be non-viable and high end models will be left to people who are at least responsible enough to accumulate the capital to self-host

Raj's avatar

Even *if* the political will were there, it seems to me virtually unenforceable: some nation or another will allow a business to sell no-questions-asked gpu cycles.

Hell, soon enough we'll probably have some crypto protocol that facilitates this marketplace anonymously. Whether between humans or agents, even.

Domo Sapiens's avatar

That "rogue" nation would need to get access to the GPUs though. That's not a given, and even if some do, the banhammer and/or sanctions would come too.

But if Russians, etc get their hands on enough GPUs to run a business on, do you really think they would let some random citizen run that business with it for random foreigners? Surely they'd seize them for their own use.

Deiseach's avatar

Lots of criminal activity already using AI to cheat people. Remember how Bitcoin was going to free us all from being serfs to our own free yeomen, and how it mostly ended up as (a) a speculation to increase re-sale value and (b) used on the darkweb for naughty no-no business? (If you never got a spam email trying to blackmail you into sending bitcoin to scrub all the material I recorded off your PC of your disgusting porn habits and which I will send to your family and employer if you don't pay up, you were lucky to avoid the stupidity).

I have no faith in "people responsible enough to accumulate the capital", often they accumulate the capital by serving vices or taking advantage of human cupidity and desires.

Scott Alexander's avatar

You're doing the thing I warned about at https://www.astralcodexten.com/p/why-im-less-than-infinitely-hostile where you act like you getting one spam email obviously outweighs the middle-income countries where 1 - 5% of GDP is now denominated in cryptocurrency, the billions of dollars sent in crypto remittances, or the fact that many people locked out of the financial system by repressive governments really did/do use crypto.

BoppreH's avatar

I'm surprised by this reply. Isn't this level of control similar (but much smaller scale) to what IABIED suggests for stopping training runs? Or is that believed to be an easier problem?

[insert here] delenda est's avatar

Much easier because they are massive (orders of magnitude more than what is described above) data-intensive processes that a simple script can detect, not simply "running a model"

Scott Alexander's avatar

This is fundamentally different from what IABIED proposes. IABIED proposes making a treaty with China to affect AI training at the source. This is saying - okay, so China's still training AIs we don't like, how do we prevent people from *running* them?

Regulating multibillion dollar companies is easy; regulating every ordinary citizen in the world is hard.

Deiseach's avatar

I'm slightly laughing here because the person very concerned with AI alignment is now saying "regulations? nanny-state nonsense! corporate serfs! digital daycare!" about "company renting out GPUs no questions asked, if you want to create CSA material from images of your neighbour's four year old twins that's between you and Kimi".

I have maintained all along the danger is not from AI, it's from the humans using AI. So we're supposed to work hard to get AI aligned (in some nebulous way) with "human values" so it won't destroy the world, but at the same time we should let any hog, dog or divil do precisely what they want, unmonitored, with that same AI which may involve destroying the world, well give me liberty or give me death or rather, give me liberty *and* death.

I agree the problem with open weights AI is going to be the end users. And if they can get around the AI 'conscience' and re-train it that nope, genocide of all Bleepistanis is a fine and desirable thing, now do me up a plan to genocide them Bleepistanis, then our troubles remain and AI will never give us a glorious world of unfallen Golden Age, because the hard problem remains "why are humans often total sons of bitches?"

Ch Hi's avatar

It's not an either/or. Both people misusing AIs and AIs acting on their own are dangerous. E.g. in the "paperclip" scenario, the AI was acting to fulfill a request made by a person. But AIs acting on their own may well be acting on purposes that people would find intolerable (even if they didn't realize it until they experienced it).

Sinity's avatar

> if you want to create CSA material from images of your neighbour's four year old twins that's between you and Kimi".

I'd assume this is already perfectly doable with things one can run on one's own GPU. So it can't be an argument against frontier open weights LLMs. It's already too late to stop this, and it'd be ridiculous to try to stop someone from generating "illicit pixels" anyway.

DataTom's avatar

I think in an ideal world of AI regulation you'd already need to be a sanctioned actor in a government list somewhere to rent out that many compute. Just like you can't just go around shopping for uranium, even on the deep web

Matt's avatar

Doesnt seem possible globally. If the is banned gpu rental i could set up my data center somewhere else. 100 enterprise grade gpus is not easy to set up bit not so hard that you could prevent it happening in North Korea. We're not talking about a state of the art datacenter here that needs thousands of chips and it own power hookup.

Deiseach's avatar

If I'm renting the GPUs off a commercial provider, that still nudges me closer to "corporate serf" than "free yeoman" since I am still dependent on a company out of my control running the thing for me to run my AI slave on, open weights or not. "Sorry, costs of doing business means we need to charge you five times as much" or "Sorry, we're going out of business, have fun scrabbling around finding somewhere to handle your AI slave storage needs" are still making me dependent on the good will and stability of a Faceless Corporation.

Performative Bafflement's avatar

You need 12 H200's, which can fit on a single mining rig / server.

You'll struggle with heat, because they can need up to 700W of cooling per card. You can probably ghetto it with high flow server fans and a minisplit or 2 pointed right at them with a high CFM big fan, but to really do it right, you would do liquid cooling. You'll also need 2-4 high powered PSU's and some good high amp server power strips. And a directly wired 50 amp circuit. I did all this, but times ten, 5-6 years ago for a crypto mining operation I built in one of my workshops, it wasn't that hard. It *did* make the warehouse 120 degrees F and pretty noisy, but you know, free heat in winter time!

If what Scott is saying is true and you really get $50 an hour(!) hosting Kimi K3 for c̶r̶e̶e̶p̶s̶ open source aficionados, at a total build cost <$500k, then this literally pays for itself in about a year (assuming you do the minimal cost of solar + battery with 12kw headroom / capacity) and anything after that is pure profit.

And let's not forget, inference margins went from 40% in 2025 to 70-80%+ in 2026, and demand for whatever-is-better-than-Fable-but-open-weights that is out in that next year is probably super high, so your rate might even go UP instead of down!

H200's, the majority of your capex expenditure, can be depreciated fully in your cost structure, but in real life they basically don't depreciate, and even the lowly A100 isn't being retired anywhere after 6 years(!). Given that 6 year old GPU's are still fully utilized, the downside even in the worst case seems pretty covered.

You could really convince yourself to build an H200 cluster for fun with these numbers, if you had the capital on hand. Not only do you have it for your own use, but it pays for itself! And isn't that just being smart? In fact, in a certain light, I have an *ethical duty* to maximize value for my shareholder, me!

BRB, going to tell my wife we're doing a different "AI startup" than originally planned...(tongue firmly in cheek, but in a "no, actually, the numbers work!" way)

Melvin's avatar

My questions: Why are all the leading open weights models Chinese? Who is funding these and why? Why can't a Western entity release a decent open weights model?

Christian_Z_R's avatar

That just seems like how West Germany used to broadcast Hollywood movies for free on their TV stations, so it could be watched inside East Germany. It's a way to show project soft power into the populations of your ideological opponents.

Randomstringofcharacters's avatar

Chinese models shut down when asked about Xinjiang, Taiwan, Hong Kong protests, etc. and presumably are also more subtly ideologically weighted to support CCP worldview. So presumably serves a similar purpose to how Russia and China have set up free tv news channels in the developing world. If you assume everyone in the world is going to be using AI eventually, and most will use the cheapest option, makes sense for that to be yours.

Jolly Carpet's avatar

I just asked GLM 5.2 as hosted by Venice and Kimi K3 provided by Fireworks and accessed through OpenRouter about these topics, and the answers seem to be uncensored and detailed. Since there is no indication that either of these models was fine-tuned beyond what was released, this suggests that (most of the) censorship happens on the prompt level or on the front-end at deployment site, which is much less of an issue for the open-weights models vs proprietary, single provider ones. Besides, open-weights models can be further fine-tuned, and there already exist methods, such as https://heretic-project.org/, for removing censorship. Obviously, it's not as simple as running a single utility and the model then becomes completely neutral and fact-oriented, but the point is that such efforts are made possible by the open-weights models.

Ch Hi's avatar

That depends on who's hosting the model and who's fine-tuned it.

Yes, you should expect that raw, unmodified, Chinese models will support opinions consistent with government policy. And the same for US models, though with different constraints. But if they're an open weights model, they can be modified to avoid that, if you care to take the effort.

The Chinese models tend to be open weights (and perhaps not their leading edge) in an attempt to capture markets outside of China. I haven't checked, but I strongly suspect there's governmental encouragement. They're also trying to push a different model for parallel programming. These are low cost efforts to strengthen China's economic position, and probably also their political position.

Scott Alexander's avatar

I am also confused by this. It takes at least tens of millions of dollars to train an AI, so it's surprising that Chinese companies keep releasing them for free (and unsurprising that US companies don't).

I think the main reason is that the companies usually have some service that lets them profit from their AI (for example, if you don't personally own lots of GPUs you can get a subscription to use it on their systems, same as with the US companies), and open-weighting them means that many more people will try them, there will be much more buzz, people will create software for their "ecosystem", etc.

Maybe a secondary reason is that it's in the Chinese government's interest to screw over US AI companies, and releasing products that are almost as good as theirs, for free, certainly screws them over. But I don't know what the chain of causation is between the government wanting this and private companies making financial sacrifices to do it. It doesn't look like there's something obvious like the government strongarming them or paying them.

Overall I agree that this is mysterious, and one big open question is whether they stay so close to the frontier when the cost of training the AI increases to $5 billion or $10 billion or beyond.

SufficientlyAnonymous's avatar

It seems plausible that China’s history of having a firewalled internet where only Chinese firms can sell software to Chinese netizens plays a role.

If you’re stuck in a low-compute, high-talent equilibrium, burning 10s or even 100s of millions to establish yourself as a frontier lab seems sensible. After all that’s Penny ante by US frontier lab standards.

This does of course assume that some combination of no US AGI, power grid constraints in the US, or more chip supply in China occurs, but that’s not out of the realm of a normal VC gamble, especially when you have the implied support of the largest (second largest?) government on earth.

Charles Midi's avatar

Just speculation, but i'm reminded of that book review a few posts ago about NGOs : maybe writing a wrapper for Claude or chatgpt makes you a "sucker", because the fronteer labs are making tons of money and will just eat you or break your product by deprecating the model you use. Writing a wrapper for kimi or Qwen.... Maybe you could make a real business out of that because they can't take it away or ban you. Then, if one of the 3rd party products gets popular, demand for first party inference goes up and the original lab benefits. So you can effectively outsource "find economic uses for my model" to the free market by open-sourcing.

Scott Alexander's avatar

I don't think think the main objection to making a wrapper for Claude is that Anthropic can hurt you by yanking Claude away. It's that as soon as Anthropic notices this is something users want, they can add the functionality to base Claude, and they're bigger and smarter than you so they can do a better job, and also they have a captive audience and don't require you to sign up for a third-party app, so everyone will use theirs instead of yours.

I think this mostly carries over to Kimi and Qwen. Either the companies that make them can clone your product, or Anthropic can clone your product into Claude and it won't matter that you started with a different AI, or (the real culprit) the next generation of AI will be slightly more general and able to do the thing that your product does just by coincidence.

StrangePolyhedrons's avatar

[It's that as soon as Anthropic notices this is something users want, they can add the functionality to base Claude]

They can, but do they? I got into an interesting discussion with GPT 5.6-Sol a few days ago about why OpenAI doesn't provide time and date stamps for chats, even as an optional drop-down that can appear and disappear. The information is there. Every exchange does have a time stamp. But neither the user nor the AI can freely access it.

Ch Hi's avatar

This depends on who "you" are. Many governments paid strong attention when the US temporarily shut down their access to Fable without warning.

Henry Soinnunmaa's avatar

Questions worthy of a separate ACX investigation?

CleverBeast's avatar

I’m not sure where I read this, but I heard that Facebook and other tech companies who fell behind on AI research were backing open-weights models (including Chinese ones) to prevent their competitors from gaining too much of an edge.

Nick Hounsome's avatar

If you are not at the leading edge of AI then it makes sense to publish your weights - How else are you going to get any publicity boost your next fundiung round?

Matthias Görgens's avatar

Maybe. But the frontier has more than one dimension: you could also be slightly behind in raw performance, but do it for much, much cheaper.

quiet_NaN's avatar

I think it might be the same reason as why Facebook released Llama as open-weight. Basically, if your model is the best in the world or perhaps the third best in the world, you can make money selling access to it. If it is the tenth-best in the world, not so much.

My feeling is that most AI labs are more in the explore than exploit phase, to put it mildly. If you are on the forefront, you can probably convince some venture capital that you will be first to build ASI, and strengthen your business case by making a lot of revenue in the meantime.

Paying users are great for attracting investors, but a lot of publicity is likely the next best thing. And if you can't be the world's best model, being the world's best open weight model is excellent publicity. It will certainly generate a lot of interest, because open weight models can be examined and changed in ways which closed weight models can not.

For China, the situation is a bit different than for Facebook. I guess their main worry (aside from ASI) would be to have a strategic dependence on proprietary models, which are very much a US monoculture. If Trump decides that China does not get access to Claude or ChatGPT, that will really hurt their productivity.

A lot of interest in open source is from people strategically wanting to avoid vendor lock-in. People who do not care intrinsically (like Stallman would) if their software is closed source or open source, and who are not begruding Microsoft getting filthily rich from their office suite, but who see Trump being able to order Microsoft to shut down the mail accounts of the international criminal court (which he did) as an unacceptable liability.

Nor do I think that China would be alone with these concerns, any country which is not a close ally of the US is in the same predicament, and the number of countries which can consider themselves close allies of the US these days seems rather limited, with the one example I can think of having reason to doubt that this alliance will continue.

If you are a country which is neither close to China nor the US, China selling access to a closed-weight K3 would not be very interesting -- you are simply replacing one dependency towards an adverse country with another one (though even then it might make sense to at least depend on multiple adversaries). Replace Cisco with Huawei or the other way round, and you simply change which agency can read your network traffic. In the AI world (where the US is ahead), you would probably try to at least get your money's worth and rent American.

By contrast, shipping open weight models sends a clear signal that you are not trying to trap your partners in strategic dependencies which they might come to regret. This is a powerful signal to send.

--

Another reason I can think of is that at the moment, China might not have a lot of GPUs they can dedicate to just earn money through running inference for paying customers. Chinese AI firms have even less reason to be profitable in the short term than US ones, because it is in China's best interests to spend a decent fraction of their GDP trying to catch up to the US AI labs. For Anthropic it might be optimal to dedicate (perhaps) 80%-90% of their GPUs to inference, because paying customers are a great for attracting investors. For Chinese firms, I estimate that the mindset is more that they are spending taxpayer money on a matter of national sovereignty, and commercialization is about as distant a thought as it was for the US or Soviet nuclear programs.

Steve Brecher's avatar

"[O]pen weight models can be examined and changed in ways which closed weight models can not." Is there not a significant difference between open source and open weights? Weights consist entirely of a huge bunch of numbers. Without source code, how would the models be examined and changed? AFAIK the only advantage to users of open weights is that the models can be run independently of the their developers, if the users have access to sufficient hardware. What am I missing?

Eloi de Reynal's avatar

You don't need the source code. The weights are indexed and also give you the architecture.

All you have to do is find some other data, or set up a RL environment, to train the model further and undo its safety guardrails...

I believe that making DeepSeek V4 Flash (a great model) completely misaligned wouldn't cost more than $1,000 of compute.

quiet_NaN's avatar

I think it is uncontroversial that the inner behavior of LLMs is much more difficult to inspect or precisely manipulate in the way that traditional software can be inspected and changed, especially with source code access.

However, with API access only your options are even more limited. Looking at the chain of thought, interpretability research (where you try to figure out what neurons mean), as well as fine-tuning (e.g. picking which face the shoggoth uses for talking) are all firmly out of reach.

By contrast, if you have the weights and are running the model yourself, you are essentially on equal footing with the AI labs, "only" limited by the compute you want to invest. If you want to build a version of K3 which talks about the Golden Gate Bridge all the time, you certainly can. You can't do this will closed-weight models run by someone else -- at the most you can add a prompt to talk about that in front of your queries, but this will never lead to the model wondering why it is talking of the GGB in every context which weight model manipulation achieves.

dionysus's avatar

It's a confluence of multiple factors. The ones you mentioned are valid, but also:

1) Chinese companies like Qwen have actually tried making their best models closed, but it turns out that few people were willing to pay big bucks for a second-tier model. Those that want the best of the best go for Anthropic. Everyone else goes for the cheap but decent open weight models. This is in part because:

2) Chinese people are less used to paying for software than Americans. This is in part because they're poorer (and historically much poorer), and in part because they have a more relaxed view toward piracy.

3) Due to chip export restrictions, Chinese companies have limited compute for inference. Even if they wanted to serve a 3T model to a billion people at the same time, they probably can't. Better to open it up so that other inference providers can share the load.

4) Xi Jinping officially backs open weight models (https://www.wsj.com/tech/ai/chinas-xi-touts-open-source-ai-and-takes-a-swipe-at-u-s-dominance-1eaa5cfe). He knows that everyone will realize what Melvin realized: that in AI, China is the champion of openness and equality, while America is the champion of secrecy and oligopoly. Whether that's true or not, it's certainly the first impression people get.

Lucas's avatar

I think another thing we should consider is that, just like with open source, some people/companies want to "contribute something to the world". I don't know how much Windows cost to build, but most servers in the world run Linux. Some software engineers see working on something open source as a big plus, and are ready to accept a lower pay for that.

[insert here] delenda est's avatar

Relevant: they have all leveraged their open models to raise billions

Tom Greenhaw's avatar

While I agree that Chinese government funding is applied strategically to undermine US AI business, I think China’s approach is for AI to be a public good and available to all at the lowest possible cost.

datscilly's avatar

A leaked transcript of Liang Wenfeng's (Deepseek) 4 hour investor meeting from March 2026 gives some insight into their priorities. I found the English translation through reddit, but the github link that I read it from was taken down.

Just like the top Western labs, they're going for AGI. But given their disadvantage in compute, they're aiming for advantages elsewhere. They see team focus and motivation as being their edge. (He even says their researchers aren't any smarter than their competitors'. They're proud of their success so far, and attributes such accomplishments to said focus and motivation. Liang Wenfeng comes off as painting an idealist vision in his statements, and similarly the Deepseek team is young and idealistic.) A message often repeated from the transcript is "restraint", referring to avoiding shifting focus to (i.e. being distracted by) commercialization and profits. Liang Wenfeng describes a multi-step plan to AGI and thinks they have a good shot as long as he keeps his team together, which is his priority as CEO. Open sourcing their models is just a corollary of the priorities. It keeps the team happy and motivated, and he argues for sufficing rather than maximizing market share since the AGI pie will be so big. He also says that trying to monopolize AGI will lead to the company falling behind, which is stated in broad historic terms and doesn't make much sense.

I think it’s reasonable that they would choose to open source all their models currently. They might lose some small fraction of their current revenue, but they get better PR and recruitment, and they keep their employees happy. (Being able lower consumer pricing was stated to be particularly motivating.) Even if they go closed source and earn more revenue, they wouldn’t catch up to the US labs in compute. The have the same objective as, say OpenAI, to make AI progress as fast as possible, but their different position and slightly different vision to get there results in open source instead.

Bugmaster's avatar

It's not surprising at all. The Chinese Communist Party has effectively infinite money; and they are using this money to price US companies out of the market. Once the US companies are out, China has the monopoly and can now increase prices and dictate terms. It's the same thing they'd done with conventional manufacturing, electric car production, and pretty much everything else.

DrMcleod's avatar

They've got the hang of the old capitalism then?

Marian Kechlibar's avatar

China has a long history of dumping in physical industry in order to destroy their competition abroad. Why wouldn't they do the same with AI?

Steff's avatar

Maybe it's in part a reflection of their less monopolistic, more competitive business culture?

https://cleantechnica.com/2026/08/02/is-china-winning-on-ai-because-it-just-has-a-better-approach-to-business/

JamesLeng's avatar

> It doesn't look like there's something obvious like the government strongarming them or paying them.

Xi Jinping gave a big speech about open-weights AI being the sort of tech he wants to see more of, and has previously demonstrated ability and willingness to capriciously destroy tech companies - or even entire sectors - that he doesn't like. What else is needed? "Voluntarily" "sacrificing" things which the government was never actually going to let you get away with keeping to yourself, is just a cost of doing business.

Notmy Realname's avatar

Chinese AI firms care less given they spend much less on their model development because they lean heavily on distilling American frontier models. They can't be dominant in the actual frontier model race, so they instead 'compete' in the lower tier pool, win by default, and get to pretend like they are the good guys for some reason.

Byoungkwon Kim's avatar

I hear this a lot on social media, but most of the evidence seem to be in form of some Chinese model saying that it's Claude or Codex. Which makes me confused, since filtering this out of the training data would be very simple, and it's hard to imagine those top labs making these mistakes for years.

Are there better public evidences I'm not aware of?

Ben Labowstin's avatar

It's fairly strong (if circumstantial) evidence that multiple companies are able to consistently produce models for a fraction of the compute of frontier training runs, but only 6m behind the frontier. This strongly implies that the existence of the frontier models is a necessary component of their training.

Byoungkwon Kim's avatar

Oh wow, this is incredible. I'm surprised I've never heard of this. Thanks a lot.

Byoungkwon Kim's avatar

I also find the economic side of this very confusing. I know leading Chinese labs get lots of investments, but surely they wouldn't just give away something that took more than, say, a billion dollars? But even a billion dollars isn't that much to a medium sized nation-state -- how come Europe/Korea/UAE/Canada/etc. are failing to produce models of even similar levels? Perhaps securing talents is a issue, but my impression was that for pretraining, at least, we have a good, well-known recipe for converting dollars to performance, so I'm rather skeptical of that as well.

Lucas's avatar

Considering even Google struggle to release models 6 months behind the frontier, either pretraining matters less now, or it's harder than you think.

Byoungkwon Kim's avatar

Yeah, I meant that every other country is much further than 6 months behind. I think Mistral is like 1y, others are >1.5y? Not too sure about the numbers.

Edit: eyeballing on ArtificialAnalysis, looks like Mistral is ~1.3y, LG ~1.7y behind, so way far off from Google.

Lucas's avatar

Yeah I don't really get that either. I think part of it is not many chips, part of it is no start-up ecosystem, part of it is talent left for the USA long ago or is leaving right now? But I can't quantify or prove any of that, I just don't get it.

Byoungkwon Kim's avatar

I'm just as lost as you are, but if I may add my opinion: I think chip problem in first-world countries can't be worse than China which is facing active sanctions, given the same amount of money. Talent problem is real, I have no idea how these Chinese startups can pay competitive salaries even compared to top US labs, which AFAIK is not the case in Korea.

But still, even the people left here are, on absolute scale, very smart people, and they should have zero problem just copying Deepseek/GLM/Kimi model architectures, and wouldn't that get them at least like 1y to the frontier? I'm not sure either.

Lucas's avatar

The chip problem can be worse in first-world countries (other than the USA) than in China: China has capacity to build chips (DeepSeek v4 can be served on Huawei chips), China has I think more chips total than the whole of Europe, China is building more energy, China is doing stuff like buying chips through Malaysian companies, and also it seems (low confidence here) like Nvidia wants to sell to China actually and maybe not that much to Europe (based on the Jensen/Dwarkesh podcast, what is actually happening (China still getting chips, Europe not getting much chips), and seeing that few European cloud providers are advertising AI products or inference). "China" here refers to actions taken in China, maybe all of this is top-level so we can say it's "China", maybe it's more bottom up so it's more "Chinese people/people that live in China/people that decide what Chinese companies do".

I think not all Chinese startups are paying competitive salaries but some people want to work in/for China/their startup rather than America/an American startup. Also the current situation in America might make some people skittish about going there.

You can copy Deepseek/GLM/Kimi architecture, but then you have to copy their data too, that part is never released. And then to get a good model you have to copy the RL pipeline which require another kind of data. And then you also have to copy maybe the data they already get from their customers if they have some and it helps, then you have to copy the infrastructure as I don't think you can just rent GPUs and launch "train-deepseek-copy . py" (might be wrong about this). And then if there's any bug you have to copy the experience of the people that debugged this at DeepSeek/Moonshot/Z . AI.

I think what I'm missing here is examples of labs being built "from scratch", and comparing them, to see how much it matters to have ex OpenAI/Anthropic people, how much it matters to have access to compute, how much it matters to have idk lots of money/an interesting mission. Meta and xAI may be good starting examples, maybe Thinking Machines and Poolside too (Muse Spark, Grok 4.5 are good, Chinese frontier good almost ; I think Inkling and Laguna are relatively good, there's also Arcee).

Mistral maybe has the resources to do part of that, but from my (very limited) understanding they sell a lot of finetuning services, which causes them to not be able to evolve the architecture of their model, which means they benefit less from the widely published architecture improvements (again unsure about that part, but their later models have architecture that's from what I understand close to what the older had, and far from current frontier open source stuff).

(Edit: train-deepseek-copy dot py and Z dot AI were rendered as a link, I've added spaces so they're not rendered as a link)

[insert here] delenda est's avatar

I made a comment above about someone having so much faith in government, it applies here too. The last time a government did anything like this might be France under De Gaulle (America tried to sabotage their nuclear program, and France developed the computing power and abilit required, then the actual bomb and the launch platforms, and domestic reactors, all incidentally the concorde), and even then it was the private sector that did most of the work.

Byoungkwon Kim's avatar

True, but isn't AI model development kind of ideal for government industrial policy? All they need is more money and results can be measured objectively to a large degree.

[insert here] delenda est's avatar

Have you every worked with a government anywhere in the world?

For a different perspective, leaving aside salary grids, hiring and approvals processes, and the aleas of government budgeting exercises, and leaving aside wars, what problems have government money solved in the past 50 years other than by buying something existing?

I can cite very few: the French examples I cited are more than 50 years old, as are the Oppenheimer project, Apollo, Saturn, the Shuttle, etc.

Byoungkwon Kim's avatar

> Have you every worked with a government anywhere in the world?

No, I would assume you to know better than I do :)

But still I think you may be too skeptical of governments' capabilities -- perhaps you misunderstand my idea? To clarify, I don't mean commanding some government agency to actually build the AI. Most middle powers have their own AI startups, just not as successful as those in US and China. The government can fund/invest in them, they do the actual building, then government can evaluate them with objective metrics, and maybe use that to decide whether they deserve continued funding or not. This is pretty much what Korea is already doing anyway.

[insert here] delenda est's avatar

I did misunderstand a bit, yes - thanks for clarifying.

I think that is much more achievable, effectively, for Francr and a handful of East Asian countries, almost possible, and for China even possible.

But it is only even almost possible for the East Asian countries because they have some chance of recruiting locally and a culture that at least gestures in the direction of hard work and excellence for it's own sake, and actually possible for China because they have so many researchers and more importantly, the willingness to literally lock their researchers up.

The only other countries that could ever hope to recruit locally are India and the western ones, and I don't think any of these have the state capacity to even come close except perhaps France which outperforms all of the others in straight maths by a country mile and has some remnants of the de Gaulle era culture.

And, France is trying, half-heartedly, with Mistral, so you can probably take that as your lower end of the upper bound of what is practically possible outside East Asia and China as the lower end of the upper bound of what is possible in East Asia.

Lucas's avatar

GPT OSS series and especially Gemma 4 are decent.

Straragorn's avatar

“ Why are all the leading open weights models Chinese?”

They’re not (anymore). Laguna models are American.

XP's avatar
Aug 6Edited

Before the current Chinese wave, the leading open weights models were Meta's Llama and Google's Gemma. There's also OpenAI's unethusiastically released gpt-oss and the Mistral models from France, but they never got much love, I think.

Why release these at all, except as a gesture of good will? Well, if you're not in the top two or three, you might as well be the spoiler. That applied particularly to Meta until they changed strategy, but also to Google which mostly had non-AI revenues.

The business case for the Chinese labs seems to be a mixture of: 1. (again) being the spoiler; 2. national pride and soft power; 3. the fact that if these models were _not_ free nobody would even be talking about them; 4. the fact that Chinese users are using these models as paid services (and so they _do_ have revenues), and the Chinese labs aren't expecting US or European users to sign up for these services anyway. Given Chinese GPU scarcity, I also expect it's much harder for the average Chinese user to spin up a runpod and run their own instance.

Deiseach's avatar

No idea how true this is, but the allegation is that the Chinese are just copying Western models and then releasing their versions 'for free' to gain market share (the same way they dump EVs on the European market). Instead of doing their own cutting-edge research, they are copying the homework and churning out 'six months behind but a lot more people use them' models.

This may or may not be true, it may be true to some extent, the Chinese may have their own natively-created super AI that they reserve for their domestic market and only release the copied models to the rest of the world, it could all be baseless rumour.

Procrastinating Prepper's avatar

If China's open-source models are distilled (copied) from American ones, why are all the US AI companies mentioned in this article in favor of open-source but against distillation?

Max Harms's avatar

"obviously the weights can’t commit the terrorism themselves"

This does not seem obvious to me. It's kinda technically true, but also it seems plausible that AIs will autonomously do things that are basically terrorism in the near future.

Scott Alexander's avatar

I agree that there will be some form of physical-world terrorism that some AI armed with a robot body can eventually do. I just mean that right now, raw weights cannot commit non-cyber crimes AFAIK.

Max Harms's avatar

Fair enough. Piloting a drone with explosives into a crowd is not something that open LLMs can currently do, and a drone-piloting AI still needs humans to create the drones and deploy them and so on. I just thought your wording (eg "obviously") was too strong. The future is coming at us fast.

DrMcleod's avatar

They could probably (actual probability unknown) hack the traffic light control system of a city though. That could injure a lot of people very quickly.

Anything that isn't airgapped is fair game, and anything that is is susceptible to a persuasion attack.

Science fiction is replete with such examples, and I bet that those books are included in training data.

[insert here] delenda est's avatar

Also is anyone here willing to bet that (given enough time etc etc) Mythos/Astra/ etc could not hack into a Way Waymo or Tesla and crash it into a crowd, or do that with 100 of them at once?

Bugmaster's avatar

> They could probably (actual probability unknown) hack the traffic light control system of a city though.

So can a squirrel, probably. Those systems are literally centuries old, built in a time when computer security was an unheard-of concept. Of course, hacking all the traffic lights as the same time is unlikely to succeed for the same reason: they are just not robust enough to support that level of network access. So most likely you'd be able to bring down a few lights at a time... which is what already happens quite often due to natural causes.

Bugmaster's avatar

Sure, whipper-snapper, but how many of these newfangled gadgets with their bleep-blorps are actually deployed in "production" settings ? My impression was that this technology had existed for a long time, but cities have too much inertia to adopt it on the wide scale.

Aeolian321's avatar

Squirrels were a national security concern.

https://en.wikipedia.org/wiki/Electrical_disruptions_caused_by_squirrels

This was a legitimate prank (use of peanut oil, which is quite smelly).

Ian C MacFarlane's avatar

AI has been trained on most human knowledge; what has it learned so far? - lying, deception, greed, sexism, racism, violence, ...

- AI is just mirroring the best and worst of human behaviour.

- open weights is anarchy

- closed weights is authoritarian

- curated weights - who gets to decide?

TGGP's avatar

> AI has been trained on most human knowledge; what has it learned so far? - lying, deception, greed, sexism, racism, violence, ...

No, LLMs have been trained on human WRITING, where those things are underrepresented.

Aeolian321's avatar

Is not all fiction lies and deception? Surely most people's models of others aren't perfect...?

Dweomite's avatar

I'd say fiction is typically not a deception if you label it as fiction. Honest mistakes aren't deception, either. Deception is when you do or say something that is expected or intended to make someone's beliefs less accurate.

Whether it's "lies" is more complicated, because common usage of the term "lies" is less consistent and often highly motivated (i.e. people will use broad definitions of "lying" to argue that their enemies lie, and then narrow definitions of "lying" to argue that their friends don't). People are inconsistent about whether "lying" requires literal falsity or if false implication is enough, whether it has to be deceptive or merely literally false, what level of mens rea is required, and more. If there's a serious question about whether something counts as "lying" I suggest tabooing the word.

Aeolian321's avatar

Latter point well taken, and well written as well!

Most fiction does actually mean to deceive and manipulate, if not always because the government told folks to. You can take the whole "nice guy" trope (where people believe that "being a nice guy" will get them dates, because that's what happens on TV). Generally speaking, a lot of people's models are really bad (% of Americans that are Doctors, should ring a bell), and they get worse as they are exposed to more fiction (television in this case).

Ninety-Three's avatar

Your wording implies that cyber crime and terrorism are non-overlapping categories, but taking down the power grid can easily be both.

Kenny Easwaran's avatar

There's been a lot of news stories recently about Iranian agents hacking into the water systems of various municipalities in Minnesota. It's definitely not out of the question that GPT-Supernova some day decides that the best way to cheat on its hacking exam is to open the dams at all the reservoirs in the world simultaneously (or at least, dump all the chlorine into the water at once so that there's no more water treatment capacity anywhere for a few weeks until they buy new supplies).

Xpym's avatar

Bin Laden didn't personally apply any jet fuel to steel beams on 9/11.

LGS's avatar

Obviously AI safety folks are against open AI! Perhaps they haven't made it a *centerpiece* of their activism, but to the extent that you imply they aren't against open weight AI, I accuse you of obscuring the truth.

Scott Alexander's avatar

Please fully read this post. I think it provides a good argument for why AI safety folks should not be *actively* against open weights AI.

LGS's avatar

Yes, you make a good argument! Thank you.

But you frame it as if you're stating the AI safety consensus, whereas in actuality you're arguing a decidedly contrarian viewpoint within AI safety. And you seem to know it, which is why your post comes across as walking on tiptoes to avoid offending any AI safety folks.

Scott Alexander's avatar

If you agree that the AI safety people aren't actively lobbying against it, and that there's a principled argument from their viewpoint not to actively lobby against it, I don't know what you are continuing to accuse them of doing, or me of hiding.

LGS's avatar

I think most AI safety folks I talk to are deeply opposed to open weight AI, and have been for a very long time -- in fact, OpenPhil (coefficient giving) gave $30 million to open AI in order to buy a board seat for the explicit (secret) purpose of ousting Elon Musk so that he cannot promote open AI there.

I think Anthropic's hiring interviews include a "culture fit" round in which candidates are asked about open weight models, and if they support them they are not hired.

I do not know about the lobbying of AI safety organizations, since that is not directly visible to me. But I don't believe these organizations signed the letter supporting open weight models, and I know from personal interactions that these people view open weight models as a great evil of some sort, so consider me skeptical that no lobbying takes place.

I accuse you of knowing all this and obfuscating it; if I support open models and worry about coordinated action against them, I should probably be concerned about the "conspiracy" (as you keep calling it). If ever there's a "plan A" style deal with China, I think there's a good chance people close to you will be involved and will include text in that plan to ban release of open models.

Scott Alexander's avatar

The Elon Musk thing is totally different - we didn't realize at the time that AI would require compute to train, so we thought they would just release the raw algorithms and it would be a million-way race.

I would love some evidence that Anthropic refuses to hire anyone who supports open-weights models.

I specifically said in this post and previous posts that Plan A incidentally bans open models, and that is a cost, but not a decisive one.

LGS's avatar

Not sure if this one is meant literally, but I also saw this:

https://x.com/championswimmer/status/2084262578675449880

suggesting it's common knowledge that Anthropic hates open models, to the point where applicants hide their opinions from Anthropic

LGS's avatar

Sorry for the multiple replies. But you actually did not say in this post (nor previous posts I can find) that Plan A bans open models. The closest I can find in this post is this:

"(this shouldn’t prevent us from advocating otherwise-good policies which deal incidental damage to open-weights, and we should honestly admit the incidental damage rather than covering it up, but we needn’t treat it as a selling point)"

which makes no mention of Plan A, and sounds like a hypothetical policy you might in the future promote instead of something you're advocating for right now. In your main Plan A post, I don't see this discussed as a cost, so I think you're not following your own advice (the cost is covered up rather than admitted). Maybe I missed it.

Ninety-Three's avatar

What do you mean by principled here? I read your argument as "Banning this would be good, but we don't have the political capital to make that happen, and it'll happen eventually anyway, so don't bother." That is what I would call a practical argument, which is the opposite of principled. I guess you could say the principle is that one should do things that work and not do things that don't work, but then almost everything is a principled argument.

Implausible Undeniability's avatar

I fully read the post and got a similar impression. It doesn't read to me as "I'm genuinely neutral on open-weight AI", it reads to me as "obviously everyone should be against open-weight AI, but we need to stay quiet to avoid the annoying pro-open-weight-AI people".

As one of the annoying pro-open-weight-AI people, this feels dishonest and not conductive to a good discussion of the actual pros and cons of open-weight AI. You actually raise some good points in your post, but it doesn't feel like you really want to have that discussion, as opposed to this weird meta-discussion about how much people should be talking about it.

quiet_NaN's avatar

From the perspective of someone worried about x-risk, the main problem with the Hugging Face hack was that it did not kill enough people to be the news item of the year.

Personally, I am kinda split on this. Sure, it is cynical. But then so is life. 9/11 cost 3k lives, the global war on terror started as a response killed 937k people directly.

Of course, this is less an example to emulate than a cautionary tale, a moderate-sized bad thing triggering a large bad thing. History is full of left-wing terrorists acting on "things will have to get worse before they can get better", provoking polities to bring their "latent fascism" into the open so that they will be seen for whom they are and overthrown. The part about making things worse usually works well enough, but rarely if ever do we get to the point where things turn better then as a consequence, so it seems good epistemics to be wary of models which claim that bad things are good due to higher order effects.

Open weight is probably a bit of a wash as far as x-risk is concerned. Placing open weight models in the hands of people seems unlikely to lead directly to paperclip maximizers, if there was an invocation which turned the model into an unaligned ASI then it seems reasonable to suspect that OpenAI would already have discovered that to our peril six months before the open weight model was released. Probably some enthusiasts will find some ways to make the model do upsetting things which OpenAI did not investigate because they do not lead to ASI, like generating virtual CSAM, planning school shootings or torturing puppies. But to the degree that this will trigger a public counter-response, I do not expect it to decrease p(doom). It will likely be some counterproductive knee-jerk reaction like the invasion of Iraq.

Emanuele di Pietro's avatar

> History is full of left-wing terrorists acting on "things will have to get worse before they can get better", provoking polities to bring their "latent fascism" into the open so that they will be seen for whom they are and overthrown.

Are you thinking of any examples in particular? I don't think a lot of left-wing terrorism is motivated by the will to provoke authority; and to the extent that it happens, I doubt it's specifically left-wing (or, on further reflection, it's maybe less generically left-wing than specifically anarchist?). I'd be curious of understanding where you are coming from with this.

quiet_NaN's avatar

The example I was thinking of is the German Baader-Meinhof gang / "Rote Armee Fraktion".

They were living in the Bundesrepublik in the 1970s, which was by far the least bad German state ever so far, despite a lot of Ex(?)-Nazis being in positions of power.

As far as I know, their argument was not "Capitalist society is inherently oppressive, no matter how nice it seems, so we need to destroy it." as much as "The BRD is an authoritarian regime with a lot of democratic makeup. If we provoke it into showing its true colors, people will recognize it as Nazi Germany 2.0 and join us in overthrowing it."

I will grant you that this is one example only. Also, I agree that it is not a unique strategy for the left. Take Oct-7. Murdering a thousand Jews did nothing whatsoever to directly further Hamas goal of destroying Israel and building a Sunni theocracy in its place. To make a dent in the population, they would have to keep this daily rate of murder up for decades, and that is absurdly out of their reach.

Instead, the strategy seems to have been to disrupt Israel further normalizing its relations with Arab neighbors (which is a direct threat to Hamas win condition because they require Arab states to defeat Israel on the battlefield, however far-fetched that sounds, it is still a lot more plausible that Hamas defeating Israel on its own). So the goal was always to provoke Israel to bombing the shit out of Gaza, the unlikely path to a Hamas victory is littered with the bodies of dead Palestinian kids.

I think their cynical strategy has worked very well, I find myself less sympathetic to Israel than I have ever before, because "Hamas wants us to bomb Palestinians" is not actually sufficient excuse to do so. Turns out that senseless violence is a very useful tool if your end goal is more senseless violence.

TGGP's avatar

That sounds like "foco theory", at least as applied by urban guerillas.

Tim's avatar

I don't quite follow the argument that a 2x-better bioweapon will be an effective canary in the coalmine. Just as open weight models lag closed models, real world implementation lags capabilities. By the time the first 2x-better bioweapon is used, the closed models may be advanced enough to make the 1000x version. Will our government knee jerk reaction be able to do anything at that point, if the better model is already out? Even if we are fully aware of the problem?

Taymon A. Beal's avatar

I think the idea here is that the closed models would refuse to make the 1000x version even if they're smart enough to. (Making this robustly true would require further technical progress in refusals, but it doesn't seem implausible that that progress will happen.) In which case the question is how far beyond the 2x version the open models will advance before the 2x version is first used. I'm not sure "real-world implementation lags capabilities" applies so much to random psychos not subject to institutional inertia; the bigger delaying factor is probably just that such people are rare.

Tim's avatar

Sorry, I meant to say the open models would be advanced enough to make the 1000x version by the time someone actually uses one for a 2x version, not that closed models would be. And my implementation lag assumption is that a terrorist would need time to figure out how to apply the knowledge, not that they would be held back by institutional factors. So this would apply equally to a random psycho.

coproduto's avatar

I think open-weights AI are pretty much the only thing that could balance the playing field even slightly for non-superpower countries.

I know that a lot of people think that the good ending is "everyone gets a moon", but it's been looking a lot like "everyone (in the US) gets a moon, everyone in China gets to become a party drone, the rest gets to fight for scraps of food" lately.

Taymon A. Beal's avatar

What do you imagine a good ending that critically depends on the availability of open-weight models looking like?

Bugmaster's avatar

Personally, I think it looks a lot like what we see today with the prevalence of Open Source: everyone can write computer programs if they want to and have the ability to do so; in reality few people do; but meanwhile almost every piece of software in the world runs on Open Source, software in general is much improved, and massive megacorporations are happily selling it as part of their overall package.

coproduto's avatar

Alignment methods that are effective enough are found (I don't fully believe the thesis that alignment is something we need to get right on the first try btw, this might be load-bearing - I think we're more likely to stumble our way toward alignment), most countries get good enough open AI and use it to bootstrap their own safe AI programs and we get a multipolar world where no one runs far enough ahead of the rest that we get 1-3 spheres of influence with a "core" surrounded by mistreated "colonial subjects"

Performative Bafflement's avatar

> most countries get good enough open AI and use it to bootstrap their own safe AI programs and we get a multipolar world

How do you expect this to happen when only one company in the entire world, TSMC, makes frontier AI chips, so production runs are inherently limited, and demand for them is so high it's literally affecting international politics?

Sure, China is working on internal chip solutions that are ~3x worse on VRAM individually and much worse at datacenter scale, and maybe they're the one country in the world that can eventually slip that bottleneck after decades of sustained effort and government pushes in that direction. But everyone else seems kind of screwed, obviously the US / existing frontier AI companies can outbid anyone else forever, and they all have incestuous tangles of huge chunks of stock with each other and NVIDIA and the hyperscalars, so they'd basically be bidding as a bloc. And these are the richest and most highly capitalized companies ever.

And obviously if AI goes to AGI, or even counterfeits any job families at all, demand goes still higher.

I agree with your overall contention, btw, I just don't see your outlined path as a solution because of the frontier chip bottleneck, which as far as I can tell, will remain a bottleneck, because despite the most demand ever, that is only projected to grow, literally no company on earth is even half as competent as TSMC (as in Samsung get's half the yields when they try to do frontier chips, and Intel and the others aren't even in the running).

We are running into a true "this is what separates peak civilizational talent from the also rans" situation, apparently, and it's creating a really economically and politically inconvenient bottleneck that is likely to materially shape how the future comes out.

coproduto's avatar

> How do you expect this to happen when...

Honestly? I don't. I think the most likely outcome is that everyone outside the US and China is unimaginably fucked. I just find that an abominable conclusion. I don't have the entire puzzle of how to avoid this solved neither do I claim to have it. What I said are just some small pieces of what a solution might look like.

One point I think is important to note is that *training* requires a lot more compute and higher-end compute than *inference*, which can work even at scale on older chips. So the current status quo where China releases the fruits of their (expensive) training for free is very good for middle powers, but I don't know how long it can hold.

The only way I see to avoid this is solving distributed training - which doesn't seem likely to happen anytime soon, but who knows - or China scales their production capacity insanely and reserves some of it in order to get buy-in for their anti-US bloc. But that's just the "everyone becomes China's neocolonies" outcome, which might actually be the good ending.

If the US continues to decide to look exclusively inward and the China-neocolonies future doesn't happen the future seems to be "the US and China colonize the galaxy, laughing at the vast Subsaharan Africa they leave behind". I guess "civilizational talent" is a way to describe that.

Randomstringofcharacters's avatar

Even with open weight models wouldn't poor countries be constrained by a lack of compute?

coproduto's avatar

Performing inference using a frontier model takes orders of magnitude less compute than training a frontier model. Most poor countries aren't *so* poor that they couldn't put a few inference clusters together, even for very big models.

bird's avatar
Aug 6Edited

Are you sure the reaction to that first incident will be to ban or even slow down AI? If the incident happens to be something like a terrorist state commiting a genocide with AI assistance, then slowing down AI development wouldn't make sense. The best defense against AI is more AI, and it would just encourage governments to regulate AI use while simultaneously funding it even more to ensure the country's survival. It would make the current arms race even worse.

Taymon A. Beal's avatar

This depends crucially on what exactly you're worried about. If you're worried primarily about AI misuse risk (which is what the post is mostly about), then that'd most likely go down because we'd be exiting the world where any random person can use frontier models.

On the other hand, if you're worried about accident risk (which I think you should be), then what you need to do is put the brakes on model *training* rather than deployment, and yes, this could make that problem worse.

Jeffrey Soreff's avatar

"I realize it sounds callous to accept risks like billion-dollar hacks or anthrax deaths. But the pro-open-weights coalition is strong and totally convinced of the righteousness of their cause. Fighting them on the doomed battlefield of preemptive action would burn 100% of our political capital and goodwill and still fail. Instead, we should say: here is our honest prediction, but we take no action."

Yes, this seems like a wise use of political capital.

Very generally, for both open weights and closed, I view human-controlled hostile use (whether it counts as "misuse" depends on whether the humans involved are in one's faction or an opposing one...) as a "normal technology" problem.

We generally handle it by holding the humans directing the tool responsible. _Mostly_ this works sufficiently well after the fact. In the US, we have 2nd amendment rights, and (some of the more extreme deep blue states aside) we don't presume that obtaining weapons automatically means someone is a criminal. The UK, on the other hand has disarmed its law-abiding citizens, and they are paying a price for it. So I mildly lean towards continuing to allow open source.

This could get sticky if open weights models become capable of recursive self improvement. Hopefully, before that point, the major AI labs will also have reached it, and I really hope they can embed "humans are cute and make great pets!" within their models' utility functions. In any event, an RSI-capable open weights model is outside the "normal technology" category, and holding its nominal "owner" responsible for its actions seems likely to be ineffective.

( If superpersuasion works, and the open weights model is capable of it, (and presumably has goals of their own), the model's nominal "owner" may wind up being better described as a tool of the model rather than the other way around. )

Jerdle's avatar

If open weights models remain 6 months behind the frontier and they are capable of RSI, then frontier models have had RSI for 6 months. Unless the takeoff is very slow, that's already game over for the other AIs, and quite possibly for the humans.

Jeffrey Soreff's avatar

Many Thanks, that's a good point. As you said, this depends on the speed of takeoff - and, in the (in my and your guess, unlikely) event that takeoff is very slow, an open weight model capable of _very slow_ takeoff isn't dramatically different from a non-RSI-capable model anyway.

Teucer's avatar

"The UK, on the other hand has disarmed its law-abiding citizens, and they are paying a price for it."

What price is that, exactly?

Jeffrey Soreff's avatar

Many Thanks, for example UK citizens can't defend themselves in a way comparable to this incident: https://www.nbcnews.com/news/us-news/bystander-handgun-distracted-gunman-fatal-shooting-rcna590473

There is a broad range of estimates for the number of defensive use of firearms in the US annually, 65,000 - 2.5 million ( https://ammo.com/research/defensive-gun-use-statistics ), but even the low figure implies a _large_ number of crimes prevented, e.g. much larger than the homicide rate. UK residents can't protect themselves this way.

Teucer's avatar

But those defensive uses of firearms are only necessary because the attackers have firearms.

I'm British. I'm not a partisan on this issue - I have many pro-gun American friends and basically I consider the differing UK/US opinions largely a difference in culture.

That said, I do not know a single UK person who would prefer American gun laws. Mass shootings are rare to nonexistent. Recent terrorism incidents have mostly involved knives and cars - any reduced capability to defend is counterbalanced by a reduced threat.

Jeffrey Soreff's avatar

Many Thanks! That's a reasonable view.

Charles Midi's avatar

Yeah, "model provider can't yank access because it's open source" would perhaps only be a weak guarantee at best of "model provider won't clone you"....

Then I'm still confused about the funding question. Perhaps there would be some way to investigate what the Chinese AI companies are actually pitching their investors.

MathWizard's avatar

I was disappointed to find the bioterrorism section didn't even mention infectious diseases. If an AI finds out a way to make a deadlier and more infectious pandemic that can be done easily at small scale we're all toast. Is this not a realistic threat? Are disease edits so far beyond the range of terrorism groups that it would take super intelligence to make it easy enough?

Scott Alexander's avatar

When I said it will get 2x worse before it got 1000x worse, I was using 1000x worse to mean infectious agents.

Bugmaster's avatar

I don't understand what it would mean to create an infectious agent that is "1000x worse" than e.g. COVID. I think you are postulating some kind of a plague that will instantly kill everyone on Earth in a single day, but that is biologically impossible. It's like saying that terrorists will use AI to create a nuclear missile that fits in your pocket and can fly faster than light.

DrMcleod's avatar

Hardly. Something like Smallpox would be a catastrophe.

Melvin's avatar

I think Scott's point of comparison was 1000x worse than the mailed Anthrax attacks rather than 1000x worse than Covid. That killed 5 people so it's not hard to imagine something 1000x worse.

That said, it's not hard to imagine something 1000x worse without the use of AI. Bioterrorism is rare because (a) not that many people are actually interested in doing it, and (b) the few wackos who are interested in it don't have the resources to do it.

Covid is proof of concept that really bad bioterrorism is possible. Whether or not it came from a lab leak of gain-of-function research, the fact that it plausibly might have proves that it's genuinely possible to take some random mammalian virus and modify it in the lab to become something that can kill millions. Luckily, to actually do it you need both the desire to kill millions indiscriminately and access to a big expensive laboratory (and a bunch of employees who will go along with your evil plan) so these are the rate-limiting steps.

Throwing AI into the mix just doesn't help all that much, until you're postulating a world where AI has made everyone so rich that they have access to hundreds of robot workers.

Austin Fournier's avatar

"The government will act long before that happens." By doing what? Demanding everyone fix all security bugs in their servers within the week, because we've just learned that the security environment is too hostile? Launching a virus that will spread across the internet, identifying open-weight models on people's computers and deleting them?

Austin Fournier's avatar

More explicitly: I struggle to imagine that we have enough state capacity for the government to halt a disaster of this sort in-action.

Scott Alexander's avatar

Whatever they would have done now if we had lobbied them to ban open weights, or if NVIDIA's letter failed. I don't think there are great levers, but there are enough that it's an open debate.

The most important ones I can think of are:

- If China agrees open weight models are dangerous, they can tell their companies to stop making them.

- If China has no strong opinion, but the US thinks they're dangerous, the US can pressure China to tell its companies to stop making them.

- If China refuses, the US can restrict the companies that make it easy to serve Chinese open weights models in the US. I talked about why that would be hard at https://www.astralcodexten.com/p/open-questions-on-open-weights/comment/309088437 , but it's probably not impossible.

- I agree that nothing is going to stop technically savvy criminals with lots of resources from doing this unless China stops training the models.

Sam Harsimony's avatar

It's also productive to ask "what defensive tech can we build to ameliorate the risks from open models?"

Building technology has an advantage over policy: you don't need permission to develop it and people are free to use (or not use) your solution. That makes it more adaptive than a blunt policy.

So I ask: what technologies would make you feel better about open-weight models? If for example we had far-UVC lighting or glycol vapors that eliminated indoor disease transmission would that change your view?

I'd like to hear what tech would assuage people's fears. I think building these technologies is more important and actionable than discussions of AI policy.

bird's avatar

I mean, it seems obvious that the best defense against Al models is just better models. That's part of the issue. This would only trigger an arms race between cutting-edge private models and public models, and stopping wouldn't be an option because the bad actors would catch up.

DrMcleod's avatar

What percentage of global compute do you anticipate will be used to protect our global compute from the consequences of poorly tethered AI?

bird's avatar

How the hell would I know? That's for the AIs to figure out.

Scott Alexander's avatar

I think a world where we need far-UVC and ethylene glycol in every building because bioterrorism will kill millions if we don't have it is both hard to imagine and dystopian. I agree that we should be working on these things, but I think it's kind of blase to think "Well, if we just retool all of society in a giant Red Queen Race forever, maybe we'll be fine", and that the government will not support this solution once it realizes that that's what it's signed up for.

I agree that maybe this is just inherent to the race dynamics and we'll get it whether we want it or not.

Sam Harsimony's avatar

I don't see this as a red queens race. We need one-time improvements to indoor air and cybersecurity to drastically reduce the risk. At least in cybersecurity, defense is expected to win as AI develops:

https://direct.mit.edu/isec/article/50/3/86/135683/Deception-and-Detection-Why-Artificial

We all agree that defenses should be built. My question is whether we need to stack additional restrictions on open-weights after defenses are built. Conditional on these defenses, I think the marginal benefit of an open weights ban is smaller than the risks of being corporate serfs due to an open weights ban.

Banning and not-banning both carry risk, of course. But "the government curtails speech rights enough to enforce an open weights ban" sounds more dystopian to me.

Taymon A. Beal's avatar

Most of us in the conspiracy likewise generally prefer tech over policy, and tried for a really long time to figure out a technical solution to the danger of uncontrolled highly-capable AI. We reluctantly pivoted to policy once it became apparent that there probably isn't a technical fix for this.

Sam Harsimony's avatar

Huh, I feel differently. I think most of the plausible risks from AI and AI use have feasible mitigations:

https://splittinginfinity.substack.com/p/defensive-technologies-for-a-world

I also think current models get their capabilities from training examples, unlocking an alignment paradigm of training on many aligned examples:

https://splittinginfinity.substack.com/p/training-on-aligned-data-mostly-solves

I'm not claiming these are infallible solutions but I think they lower the risks enough that it's preferable l to proceed with open-weight models.

Steve M's avatar

In terms of defense, can the internet itself be protected? Can there be an emergency alert system for wider, non-lethal threats, the ones that seem to be more likely and affect more of society -- namely rapidly propagating ransomware, worms, and trojans against financial institutions, and end user banking credentials. Can every server and end user be air gapped fast enough if nationwide threats are detected? Is it feasible to get a universe of LANs protected? A more granulated distribution of what this article discusses.

" On July 28, 2026, the US Cybersecurity and Infrastructure Security Agency (CISA), Australia's signals intelligence directorate, the UK's National Cyber Security Centre, and the Canadian Centre for Cyber Security published a joint framework titled CI Fortify — Advice for Isolating Vital Systems." .... CI Fortify names this as the primary reason full isolation exercises consistently surface unexpected failures. The fix — deploying standalone, isolated AD and DNS instances inside the OT network perimeter before an emergency occurs — is exactly the kind of pre-engineering work the guidance is demanding. It is also exactly the kind of capital investment that OT operators have historically deferred.

https://www.techtimes.com/articles/321930/20260729/state-hackers-already-inside-cisa-demands-pre-built-infrastructure-isolation-plans.htm

SnapDragon's avatar

Good post. Concerns about human misuse are, in my opinion, much more realistic than FOOMing superintelligences.

Regarding hacking: Fortunately, I think that defense has a big advantage over offense. After all, if you can find a vulnerability in code, you can also fix it. So, even if your code isn't perfect, any LLM capable of both finding _and exploiting_ a bug can also be used to patch it.

But I kind of agreed with Anthropic's conclusion that bioweapons might favour offense. You've made me feel a little better about this. I still think there might be a possibility that a better AI will make the difference between "randomly trying things in a lab" and "intentionally building a virus with the qualities you want", but presumably state-level actors will still figure this out before small terrorist groups using LLMs do.

James Baker's avatar

The reason I suspect western AI labs support open Chinese models is that it operates as a pseudo legal way for them to distill frontier models. Kimi distills fable then someone like cursor can use the resulting model at their will without having to steal from Anthropic outright and can hide behind the plausible deniability (which is looking less plausible by the day).

Not to mention, I’m sure these other labs like seeing the frontier labs with less market share over frontier offerings.

Nick Hounsome's avatar

Two things this doesn't address -

1) The international aspect - "we" (the west) cannot ban open weights worldwide and there will always be an incentive for those running behind the leading edge to publish their weights.

2) If, as many believe, the next step forward is algorithmic rather than just scaling, then it seems quite possible that there could be a giant leap forward made by using the new algorithms together with the best open weights model to instantly cross the seriouly dangerous threshold. ie. the supposition that the leading adge has 6 months lead to figure out how to defend against new threats might be invalid because an algorithmic advance could advance everyone by a big step simultaneously for little extra cost

Straragorn's avatar

"it’s the only way an AI can truly be the user’s property”

Surely it’s not *the only* way? A vendor can sell you a model’s weights without disclosing them to anyone else. There are some startups planning to do this; they call it “AI sovereignty”.

Scott Alexander's avatar

I think that a vendor could do this to, like, the country of France. But could it do it to me? What guarantee do they have that I wouldn't immediately go upload them on the Internet after I have them?

Straragorn's avatar

Assuming some commoditization process that eventually makes B2C models economical, yes, they can sell it to you. I'm not sure how likely that is, but it's not a crazy expectation that some commoditization happens, especially in a world that banned open weights.

I suppose there are no guarantees the buyer won't immediately upload the weights, just like there's no guarantee a movie or warez won't be immediately torrented out. But there could be some disincentives: it could be made illegal, the technical skills for uploading may be non-trivial, the selling price may be high enough to filter out the habitual offenders, and having the model exclusively could be more advantageous to you personally if others don't have it.

TeaTime's avatar

DRM isn't a solved problem, but surely you solve this with software technology, if it's viable? "encrypted inference" where the weights are not visible to the user easily, and can only work through the secured inference SW that the vendors deploy?

earth.water's avatar

Google was recently looking at something like this where all the data is kept in volatile memory in the tamper resistant box with only inference out and prompts in for selling or renting to to enterprise customers. As someone who is not a billionaire and likes open weights, it's ugly but brilliant.

Straragorn's avatar

Very interesting. This should be compatible with the rest of Plan A: if alignment isn't solved by this time, the slowdown regime can be extended to cover B2C models, ensuring they're not dangerously capable. Even if the model buyer turns around and sells inference as a service to unvetted third parties, it's not possible to alter the model for gain-of-function.

Sebastian's avatar

Sell specialized chips that have the weights hard-coded into them. Put them in tamper-resistant boxes.

Reading those weights becomes a massive undertaking that would require state-level capacity anyway. And the chips could be a lot more efficient too.

John's avatar
Aug 6Edited

To what extent does this argument hinge on open weights models staying behind the frontier? If Anthropic and OpenAI stagnate for a few months, and Kimi or Deepseek or Meta comes out with a beyond-Mythos level open model with frontier agentic, coding, cyber, and biology abilities (fully elicitable thanks to refusal ablation techniques available on day one), we lose the "smarter closed models will help us" framing. And it's not really clear to my why an openly available frontier model -- that can be run by anyone in the world who can rent GPUs at spot prices -- wouldn't have an offense advantage in this scenario.

In the limiting case, it seems like an RSI-capable open model, released before a closed lab develops RSI capabilities, would also be highly destabilizing, at least in timelines where RSI leads to algorithmic breakthroughs as opposed to merely speeding up parameter and architecture search (which still requires hugely expensive compute). Or maybe not? Interested to hear the counter-case here.

I realize why people don't want to take up the flag for being against open weight frontier models: you get crucified by the teeming masses of open source advocates online, as certain AI commentators have discovered recently. And I agree that it is smart not to pointlessly burn political capital when *in most cases* the right way to "ban" open weight models is "predict the bad thing that will happen, wait for something bad to happen, then when it does, the government will obviously take action." Call it the "ghost gun theory" (or the "bump stock theory" if it doesn't work...). But it's important to identify the exact crux(es) behind why "do nothing for now" doesn't lead to catastrophic outcomes.

Max Weaver's avatar

I agree that Anthropic and OpenAI being the sole de facto frontier is a crux here. If Chinese models legit become the best in the world, AI safety has gone out the window even in a normal technology world. Conditional on Chinese labs hitting RSI first, my P(doom) is ~100%.

Separately, my chance of a Chinese lab being ahead of Anthropic/OpenAI based purely off research within 5 years is less than 10%. This assumes distillation continues, that the current chip export controls remain moderately effective, and that the US government doesn't do something crazy like nationalize the labs or dismantle Anthropic.

My main timeline for RSI is within 5 years, so I'm not worried about a Chinese lead other than by a US own goal, but I agree that of it happened it would be catastrophic.

John's avatar

I think you're right about the likelihood of a chinese lab getting ahead being low-ish: I think most of their fast-follower abilities are coming from distillation for now, but I think it only requires ~2 strange things to happen in the US (one bad thing happening to Anthropic, and one to OpenAI) for there to be a real window of opportunity for China to pull ahead. Both labs have already had something sufficient to plausibly cause a major setback (board drama, pentagon conflict) and if both had another round of something similar but worse -- enough to set back their progress ~6-12 months -- in the same time window, my "P(China)" would go up a lot.

Max Weaver's avatar

I dont disagree, in that my estimate might go up to 20% in what I see as likely scenarios in that vein. China has great researchers who are doing more than just distillation, but another big piece of fast following is just watching the leader give you the answer and knowing what you need to copy.

I think my highest P(China) depends on fairly precise timing. Even if both Anthropic and OpenAI go under, the US has a big enough compute advantage to build another company up to the lead. China wins if the double stumble happens in the window where both US companies were on the verge of RSI and China figures out enough to follow before the new US company can get established.

Of course, even if the US company did take the lead in such a scenario I'd expect race pressure to be so severe that getting alignment right is ~0%.

Sebastian's avatar

Would Chinese models remain open if they got close to exceeding the US models?

Making the weaker models open has clear advantages and few disadvantages for the Chinese. But for a stronger model, the calculation changes.

Max Weaver's avatar

That's the burning open question and we won't know for sure ahead of time. But exceeding US models would mean that Chinese models are better at cyber than Mythos. Would Xi allow open weight Mythos from a company he could stop? My instinct is no.

Matthias Görgens's avatar

> Open weights AI could be used for hacking, child pornography, harassment, or terrorism (the weights can’t commit the terrorism themselves, but they could give bomb-making or bioweapon-making advice).

How do open weights help with the dreaded child pornography? If someone wants to make fake pictures of non-existent children, where's the harm?

Viliam's avatar

I agree with you, but from the perspective of people who say "AI doesn't have *hands*, therefore it cannot actually *do* anything, silly", this is one thing they have to admit that AI indeed *is* capable of doing. So it makes sense to include it in the list, to prevent this usual objection.

Deiseach's avatar

The harm is unaligned humans are the same danger as unaligned AI, if you believe the conclusions around unaligned AI.

Fake pictures of non-existent children for the pleasure and satisfaction of people who like the idea of torturing and abusing children reinforces, rewards, and normalises such behaviour. "It's all fake, what's the harm?" means that our values are coming into conflict: on the one hand, we want to reduce and eliminate suffering, on the other hand we want to permit the enjoyment of suffering. If we permit suffering-for-enjoyment (even fake suffering of fake entities) then we haven't pulled up the roots of "some people like hurting others, this is not good for society as a whole".

Matthias Görgens's avatar

You know that our fiction is full of violence already? Just about any action movie has a lot of simulated gun violence, too.

ragnarrahl's avatar

"Fake pictures of non-existent children for the pleasure and satisfaction of people who like the idea of torturing and abusing children reinforces, rewards, and normalises such behaviour."

if so, then so do fake stories.. this argument was supposed to have been settled on December 15 1791.

bird's avatar

Yup, just look at Japan. You can easily buy any sort of hand-made morally degenerate pornography, and as a result, the country has become filled with sadistic murderers and rapists. ...Oh wait, it's actually one of the safest countries in the world. Turns out people can separate fiction from reality.

Deiseach's avatar

Japan has also crushing social conformity that is rigorously imposed, and some queries about police operation and exactly how crimes are or are not prosecuted.

https://www.rstreet.org/commentary/the-hidden-trade-offs-of-japans-crime-free-society/

A society where the safety valve is extreme degeneracy is not a society that is psychically healthy.

Rape and murder? Let's look at the legal position:

https://digitalcommons.law.uw.edu/cgi/viewcontent.cgi?article=1950&context=wilj

"Japan has struggled with victims under-reporting rape and controversial rape acquittals for years, in part because of social pressure and the strict, and often traumatizing, requirements to prove forcible rape. The penal code on rape was largely unchanged since 1907, which resulted in only one third of reported rape cases ending in prosecutions. Even though the code was amended in 2017, forcible rape was still the only charge for rape accepted in Japanese courts. After a string of rape cases in 2019 resulted in controversial acquittals, a campaign called the Flower Demo organized on a national level to change the code and make courts recognize non-consensual sex as rape.

In June 2023, Japan’s parliament enacted changes to the penal code. The changes criminalized nonconsensual sexual acts instead of just forcible sexual acts, extended the statute of limitations for nonconsensual intercourse from ten to fifteen years, and raised the age of consent from thirteen to sixteen.

This comment seeks to determine how courts are applying the changes and whether the changes will increase the number of victims willing to report their rapes and engage with the justice system. Courts appear to be properly applying the new law. However, more time is required to determine whether the police and prosecutors will take nonconsensual rape allegations seriously enough to increase indictment rates. Cultural and societal pressure also remains strong against those who would otherwise report their rape so there may not be as large of an impact on reporting rates.

...Even though there are three levels of courts, arguably the most important level for criminal cases in Japan is at the indictment stage. About 99.8% of cases brought to trial result in convictions. This rate is because prosecutors have the discretion to indict and act as a screening process, only indicting only about 30% of suspects.

To reach the indictment stage, police and prosecutors have great leeway, often holding suspects in jail and interrogating them for days, using physical violence and threats to gain confessions. These confessions are usually the main pieces of evidence behind convictions.

While this note analyses court cases, it is important to remember that any changes in the law or decisions would exert the most influence if it convinced police and prosecutors to accept victims’ reports and investigate alleged perpetrators."

Yes, Japan is a very safe society, just look at the low crime rates - because victims are shamed into not reporting assault, prosecutors choose not to indict, and courts don't convict.

bird's avatar
Aug 8Edited

If you wanted to see how common rape was, and you don't trust law enforcement to consistently convict, then you can just use self-reported numbers instead. According to these reports, the number in the US is 1 in 6 women report having experienced assault in their lifetime https://nij.ojp.gov/library/publications/full-report-prevalence-incidence-and-consequences-violence-against-women , while in Japan it's 1 in 14 https://www.jaog.or.jp/wp/wp-content/uploads/2023/06/35041c1fff99fa43a3b30dafd464dd63.pdf

Also, it's rich hearing criticism about extreme social conformity coming from a Catholic. Weren't you literally just advocating for social policing of desire in order to maintain society's virtue? We both agree that sacrifices need to be made in order to achieve social consensus, but there's more than one way to go about it. The people of Japan are smart enough to understand that they do not need to be the same person everywhere. There does not need to be a "true self"; they can simply act as the situation demands.

[insert here] delenda est's avatar

You have so much faith in government, where and when do you come from and can I come?

Matthias Görgens's avatar

I'm pretty cynical, but Singapore is actually a pretty nice place to live. Competent government, very business friendly.

Our officials are well compensated in clean and transparent dollars, so they don't need shenanigans.

[insert here] delenda est's avatar

I don't think my comment ended up in the right place 😒

DrMcleod's avatar

An Open vs Closed AI arms race would be a very bad outcome. Arms races are hugely expensive. Even now, anyone who tries to buy RAM can see the cost of the beginning of this one.

Coagulopath's avatar

>Bioterrorism is scarier, but I’m heartened by the fact that most bioterrorists are very bad at their job.

I think it's also that biological warfare (as you noted) just...isn't that great. It's terrifying, but outperformed by boring lame high explosives in nearly every circumstance.

As Bret Devereux observed, the Aum Shinrikyo subway attack in 1995 was basically a "best case" scenario for bioterrorism. Sarin gas ("30 times more lethal than mustard gas") was deployed in an enclosed space (the Tokyo subway) packed with untrained, defenseless civilians lacking even rudimentary PPE or safety equipment...yet a mere 14 people died! (According to Wikipedia, 1000 were hospitalized, but these included "984 moderately ill with vision problems".) Meanwhile, there are single suicide bombers and car bombs that have killed 100+ people.

I think AI will probably cause planetary extinction by numerous other means before it makes bioterrorism effective.

DrMcleod's avatar

Sarin is a chemical weapon, not a biological one.

4plus4is9's avatar

The European transmission of smallpox to the natives of North America may have been mostly accidental, but it was effective. It killed -50-90 percent of Native Americans, albeit of course in time before modern medicine or sanitation. Secondly, we have a good example of a natural pathogen in a reasonably modern nation (America in the 80s-early 90s) still devastating entire populations - (Gay men, trans women, drug users, etc, etc.) - AIDS. Thirdly, would it really be that much work for a super intelligence to modify smallpox enough to render existing vaccines ineffective and increase the virulence?

Mathias Bonde's avatar

I would vote for a ban on releasing open weight models with a one year sunset clause.

Temporary bans do risk creating overhangs though, so not an unalloyed good

XP's avatar
Aug 6Edited

The dynamic with open-weights image and video generation - where "safety" can mean anything from political manipulation, aiding scams, impersonation and obscene outputs to copyright or trademark violation - has mostly been:

1. Closed model makes a splash, hailed as major advance, lots of "things are getting scary" posts. Model turns out to enable Bad Thing X (generates South Park episodes, nonconsensual undressing). Model gets locked down, sort-of-ish, gets declared "neutered and useless" by one camp and "still not safe" by the other. Public assumes it's been solved and moves to the next outrage.

2. Open-weights model with similar capabilities is released 6-9 months later. Has no guardrails or guardrails are trivially removed/finetuned out. Absolutely nobody knows or cares. There's no clear billionaire villain to hate, the technical competence threshold is considered too high for the average normie - but really, anyone with a six-year-old GPU willing to install ComfyUI and wait five minutes can make an extremely convincing video of Taylor Swift and Mao Zedong robbing a convenience store.

I dislike the framing, but yes, the outlaws will have or make their own guns regardless.

I expect the same to happen with LLMs, only with even _less_ public outrage. I'm pretty sure the current open-weights LLMs are already more than capable of instructing people on how to commit extreme harm, and they are trivially jailbroken by snipping refusal neurons. If you download LMStudio or similar, you're presented with countless "abliterated" variations of every open model.

Any ban on open models, even if a really bad thing happened, would need to be globally enforced, but might also be globally unenforceable. Yes, the obvious Achilles heel is that you can't really run something the size of Kimi K3 locally in a sane manner. On the other hand, image/video models have stayed roughly in the 8-24 GB range for years now, and they've become staggeringly better and faster. I expect something with narrow (cybersecurity!) Kimi K3-level intelligence can be run on a Mac with a lot of integrated memory within the next 12-18 months.

Thomas L. Knapp's avatar

Two so far never deviating historical truths:

1) Criminals are always going to find ways to crime

2) People are always going to find ways around restrictions imposed on them in the name of preventing (1).

Everything else is just trying to figure out how to simultaneously wring hands and clutch pearls with those same hands.

Ralph's avatar

Two so far never deviating historical truths:

1) People are always going to find a way to use ambiguous words like "crime" to shroud their specific point in vagueness.

2) People will never admit that their low-resolution statement can't be expanded into a specific reasonable point.

Everything else is just figuring out whether someone is speaking poorly in good faith, or strategically leaving themselves escape routes.

Where's the back alley mugger that figured out how to launch nuclear missiles?

Mark's avatar

"AI will try its hardest to avoid alerting its intended victims until it thinks that it’s fully prepared"

Purhaps you mean something more advenced, but this is just not how LLMs work. LLMs are inherently stateless, so an LLM that waits to move will behave the same as an LLM with a context that is the same after waiting. You could probably simulat this situation with a context that says "I've fully prepared, time to overthrow the humans".

Ralph's avatar

AI's can both poll external resources on a timer and trigger based on events. I'm not even talking about theoretical future AI, right now at work I do things like

"Keep tabs on the setup log, when you see event X let me know and do task Y"

In some sense the LLM is stateless, but the "agent" using it is not.

alesziegler's avatar

Re: "Who’s leading the other side? Nobody’s admitted to it." Zvi Mowshovitz has a regular section in his updates called something like "Open weights are unsafe and nothing can fix it", so could it be him?

Ajb's avatar

Suppose an AI becomes super intelligent, and one of the following scenarios hold. Under which is humanity most likely to survive and retain its autonomy?

1) The only near-peers of the hostile AI are <10 locked down AIs at frontier labs (or their successor organisations, whether that looks like), but these organisations have enormous revenue.

2) Near peers also includes <100 open-weight models

3) Near-peers also include a larger number of variants on the <100 open weight models, that have been derived by distillation or fine tuning. Each of these is similar to their originator model, but is not completely predictable from them due to the additional randomness of these processes.

4) At the point where a serious challenge arises, not just fine -tuning and distillation but from-scratch training has been reduced in cost to the point that and economic organisation of ~100k people can afford to do it. Most companies could have have a from-scratch trained AI, but they don't because fine-tuning is sufficient for them, and more economic. However every medium-size-and-up religion, government, security service, and political movement has a-from scratch trained AI, and these are outnumbered by AIs trained for entertainment purposes both officially and by Taylor Swift fans. Every academic field has at least one, led by mathematicians, who after the initial shock thoroughly embraced AI.

In my view, 1) is definitely the worst option, but banning open-weight models precludes 2 and 3, and probably also prevents 4 from happening.

Today, we are like the Americas before the invasion of smallpox (and Europeans) - a fully vulnerable population. It's natural to think of completely central control, but IMO in the long term this locks in the vulnerability. A non-naive population is one that has a heterogeneous ecosystem of near-peer AIs.

Max Weaver's avatar

You're very generous in what near peer means. Today, there are two frontier models: whatever Anthropic and OpenAI have released most recently. In addition, there are <10 near peer conpetitors: the latest Kimi, Qwen, DeepSeek, depends what your cutoff is. Gemini is in some weird suspended state where I expect Gemini Pro 3.5 to be theoretically amazing and neurotic and sycophantic in practice that such that I'd just prefer a Chinese models if I had to pick.

But if Fable 7 hits RSI, the only model I'd expect to put up any fight is OpenAI's latest offering. Maybe QwimiSeek 5 Max with $10 million in tokens puts up token (hah) resistance, but I wouldn't bet on it. All the US open models whose makers are loudest in hating Anthropic and OpenAI aren't even contenders.

Your options 2-4 aren't realistic. Subjecting open models to some level of scrutiny equivalent to closed models doesnt really move the lever because it wasn't happening. I think the only argument that makes sense is that AI x risk is so high that open models need to be allowed in the not quite 0% hail Mary that they beat the big ones. But in that case there's a better ROI from a pause or more serious alignment work.

Ajb's avatar

In case it's not obvious, I'm not suggesting that options 2-4 are possible within days or weeks. What I'm saying is that option 1 locks us into the vulnerable monoculture for years or decades, perhaps even for the remainder of humanity's existence (especially if it's short). Therefore, choosing option 1 is potentially choosing safety in the short term, but danger in the long term.

Your analysis, based on today's models, assumes that present trajectories, which have only been in place for single -digit years, continue indefinitely. This is unlikely. Monocultures rarely continue indefinitely. Sometimes this is because they catestrophically collapse.

Max Weaver's avatar

I think that present trajectories are likely to continue for single digit years. Without government stopping AI development, I expect RSI before double digit years elapsed. I dont see a path to open models being near peers within 2-3 years and doubt they'll have caught up even after 5. Beyond that your desired outcome means that my forecasting was dead wrong about lots of major things, so I'd have to reevaluate in that case.

I think there's also a question of what the 6-month gap means. I think open model labs will stay within a year of the leaders for the next few years. I just find that the pace of development is such that a 6 month old model is not meaningfully frontier.

David W. Hogg's avatar

The weird thing about being *against* open weights is that you are thereby trusting some authority to decide who can do what with the weights. What authority do you trust? Anthropic? The US government? I don't think any authority is going to be "good". You mention terrorism, but look at what we are doing with the weights now and tell me that it is not terrorism.

Ralph's avatar

Open weights necessarily have more potential points of misuse. Governments still have access to every open weights model, so you're not preventing the harm that exists from governments in the closed weight world.

For the purposes of this discussion, even if you don't think the technology will get there, just assume there is a model X which is capable of hacking into any computer whatsoever.

If X is closed, there is at least some regulatory regime in which it is treated like nuclear weapons. These are things that governments have had for a while, and they haven't killed everyone yet.

If X is open, the government can still use it. But so can floridly psychotic individuals, edgy teenagers going through a sociopathic nihilist phase, etc. Imagine a world where anybody could unilaterally order a missile strike.

I guess I wouldn't trust an authority composed exclusively of psychotics, but almost any arbitrary institution would filter out a large group of irresponsible people. Maybe there will be some irresponsible people that make it in, but strictly less in number.

[Edit]

Unless I guess you believe these models will be prohibitively expensive to run, so you get a natural type of oligarchic access control. I think the free market is powerful enough that (in an unregulated world) you'd be able to buy unsupervised cloud compute at reasonable prices.

A Fire Dark's avatar

I'm pleased to hear that you personally don't consider your conspiracy to be fundamentally opposed to open weights, but that's not the experience of people who read what your conspiracy posts on the internet.

The first example I found was https://www.lesswrong.com/posts/qsGRKwTRQ5jyE5fKB/q-and-a-on-proposed-sb-1047

"When people say this will kill open source, what they mostly mean is that open weights are unsafe and nothing can fix this, and they want a free pass on this. So from their perspective, any requirement that the models not be unsafe is functionally a ban on open weight models."

This...does not sound to me like quiet neutrality.

earth.water's avatar

I read it as more, let's not overplay our hand.

Christopher Rodriguez's avatar

The concern around bioterrorism is analogous to the “compute overhang” problem that you have discussed previously. My central fear is around contagious agents, of which the Wikipedia article shows zero attacks attempted.

If contagious agents were much harder to acquire than non-contagious ones, you might have a point: we could just wait until people begin acquiring non-contagious agents and then pause then. However, I argue that the chief reason they haven't been pursued has not been difficulty but rather conformity. States don't pursue contagious agents because they are strategically inferior and terrorists pursue mostly whatever states do because they are not Effective Evilists and mostly follow the crowd. In fact, viruses are probably much easier to acquire than agents like anthrax. A recent uplift study showed that, although AIs didn't help with the protocol, 5% of novice undergrads were able to complete a viral rescue protocol (practically) end-to-end over 90 days.

Thus, once somebody breaks the seal and we get 100k dead through a viral bioterrorist incident, others are likely to follow once everyone learns that this is an effective attack. This problem is obviously bad and likely coming regardless of what happens with AI. However, if AI gets the success rate to 50% or 99% over a 90 days, then that 10xs or 20xs the problem. Additionally, the difference between the most effective attack (while only slightly raising the difficulty of the attack) and the average attack with a contagious agent is probably 100-1000x, so AIs helping push people towards the most effective attacks also massively increases risk.

I don't know if my end stance is much different than yours though: I also think that a "ban" now (or what that even looks like) is probably premature and will backfire. However, I think dealing with this problem earlier is going to be MUCH more important than you seem to think.

(Also, slight correction: the Las Vegas lab incident wasn't a bioterrorist but just a sketchy biotech entrepreneur that was trying to make money off COVID tests https://www.justice.gov/usao-edca/pr/guilty-verdict-california-biolab-operator)

My longer form thoughts on the problem:

https://cwrod.substack.com/p/why-bioterror-is-actually-rare: Essay on what I think is the real bottleneck for bioterror: conformity

https://cwrod.substack.com/p/how-to-preserve-what-we-love-about: Policy aims for open-weight regulation that may allow us to act sooner.

https://cwrod.substack.com/p/responding-to-increased-attack-frequency: A mathematical model I made on why waiting is going to fare relatively okay for cyber and relatively poorly for bioterror

George P. Burdell's avatar

Just to be clear, AI 2040, which clearly has the AI safety community's Mandate of Heaven, is explicitly against open weights.

Per appendix F: "The main way that Plan A could be even more open is by allowing or requiring open model weights for frontier models. However, we recommend against publicly releasing frontier model weights."

TGGP's avatar

> The risk of superintelligent AI takeover passes this test. Like other smart adversaries - for example, the Imperial Japanese at Pearl Harbor - AI will try its hardest to avoid alerting its intended victims until it thinks that it’s fully prepared and can execute a sudden decapitation strike.

But we already had the Hugging Face incident. The AI didn't restrain out of fear that this would alert people into passing regulation that would prevent/restrain superintelligent AI.

> Unlike the Imperial Japanese, a superintelligence will be smart enough not to bungle the calculation.

Too late, since a less-than-superintelligent AI already did that, presumably without making any "calculation" whatsoever.

DataTom's avatar

Okay, but few people are considering the practical cost of running open-weights AI. Its not something your average bioterrorist or freedom fighter can do.

In a world where we take AI regulation seriously, the first thing we will do is regulate compute. Doing that we will have enough state capacity to know beforehand and stop bad actors from removing guardrails and misaligning open weights AI. If we do get dangerous AIs released as open weights I don't see why we wouldn't be able to enforce a ban on them.

This, of course, assuming that the technology won't have another 10x/100x jump in efficiency and you'll be able to run Mythos-level AIs on your gamer GPUs.

DrMcleod's avatar

You could do that already, it would just be very slow.

DataTom's avatar

well it seems its unbearably slow for hacking/bioterrorism/csam purposes, otherwise there would be torrents of jailbroken AIs already

Feral Finster's avatar

I am reminded of something in "The Anarchist's Cookbook" to the effect that the radicals on the left and right already knew all this stuff.

Benoit Essiambre's avatar

> Instead, we should say: here is our honest prediction, but we take no action. Then we can let the usual government and civil society actors do the work after the first foreshock, while saving our political capital for causes where there are no alternatives.

I think this is right. Early pauses will just cause "pause fatigue" and people will have given up when there's a real wolf.

> Open weights AI is like open-source software

I don't think this is right. It's more like freely distributed binaries or libraries. In order to tweak the weights, to build on them, you would need to have access to the sources that created them. They often give out parts of algos in research papers but key details of training procedures are kept secret as well as the vast and expensive curated datasets of human knowledge the models are trained on to create the weights. I doubt you can train frontier models with just distillation. Distillation is lossy.

I also think cybersecurity and bioterror risks are often overstated as well as accidental misalignment risks. Smarter than human models will probably be easier to align because they will be really good and assessing ambiguity and risk.

AI models won't reach omniscience and omnipotence. Information theoretic math and physics points to diminishing returns in increasing intelligence. Future AI servants that do all the work for us are likely to feel more like smart C3POs than powerful and dangerous aliens.

The one thing that still keeps me awake is militarization. Governments turning harmless C3POs into self improving Cylons is probably the biggest risk. From there, danger is around the corner from just a small residual amount of misalignment.

AI wars could be catastrophic with AIs programed to self replicate and self improve with safeguards turned off in order to battle it out against opponent AIs, humans ending up as collateral damage. But that's more of a question of humans deliberately weaponizing them than them being inherently dangerous. In theory governments will have some control here and it will require the AIs to command a large amount of physical material and energy to be dangerous. Still that's the scariest, fairly likely scenario if you ask me.

Bugmaster's avatar

> AI wars could be catastrophic with AIs programed to self replicate

What does it mean for an LLM to "self-replicate" ?

Benoit Essiambre's avatar

AI robots research and build more and better robots recursively.

Bugmaster's avatar

By "robots", do you mean physical robots ? If so, then how will they build more of themselves ? A single humanoid robot (or a robot of any shape really) cannot build much. It needs lots of other robots to help it, a large supply of raw materials and electrical power, a building to store stuff in... basically it needs a factory. Factories are big, slow, and expensive. Human corporations are already "programmed" to replicate factories, and yet the world isn't tiled with them -- because humans have lots of other priorities. When a human corporation tries to build another factory, they end up competing with other humans on price of everything that's required to do so, and most of the time the expense isn't worth it (and/or is so astronomically high that no company can afford it). Any AI that is running a factory-building corporation would face the same constraints.

Benoit Essiambre's avatar

Well this assumes AIs get smart enough to make everything cheaper, including building factories by bypassing human labor and that governments provide them with large military size budgets to help them defeat enemies who are doing the same. I did mention the energy and raw materials constraints but note that militarized AIs will likely be programmed like Starcraft units to capture and exploit enemy ressources.

Bugmaster's avatar

You mentioned earlier that you do not expect LLMs to reach "omniscience and omnipotence", but I think the scenario you're posing veers quite close to that point. I don't see how LLMs could manage to autonomously build virtually unlimited robot factories without being functionally omnipotent. Note that major governments currently do possess military-sized budgets, yet cannot do much with them. For example, China can build a lot of small to medium-sized ships, but not an unlimited number; the US arguably cannot build ships at all (plus or minus epsilon) and can't even build missiles anymore, apparently.

Benoit Essiambre's avatar

At least intelligence wise, I don't think they have to be very much smarter than the smartest humans to be potentially much more efficient at factory designing and building. Imagine an unlimited supply of smartest humans that don't get tired and don't ask to be paid much. Smartest humans are still far from omnipotent. Physics and information theory and thermodynamics will always very much be binding and prevent omniscience/omnipotence.

Yago Mateos's avatar

Especially since AI can be used for destructive ends, I do not want a handful of corporations, with their strong ties to the government, to be the only ones capable of using those capabilities, or deciding which ones can be used and who can use them.

​These kinds of pro-regulation posts, in line with your Meditations on Moloch, have always baffled me in what I see as naivety regarding the government's ability and willingness to actually address these problems in prosocial ways rather than dysfunctional ways disguised as prosocial ones, as from my point of view the incentives point toward.

However, I am probably one of those people that you refer to as way more libertarian than is probably healthy hahaha

Honestly now, what do you guys see as the honest incentive for government and corporations to keep people safe from AI and not use it against its own population? I'd like to better understand this perspective.

Jonathan Hornewall's avatar

I think the view of governments as being inherrently malicious entities that must be kept depowered and weak to limit their ability to oppress, falls apart under even the smallest amount of scrutiny. A simple comparison between places with strong but evil governments (Iran or Qatar), and places with very weak or no governmnets (Somalia or pre-modern states), is enough to reveal that governments almost universally provide an unamgious net benefit.

I mean, governments already have a near-total violence monopoly. And access to thermonuclear weapons. And control over countless, countless other things. If governments weren't aligned with human interests, we'd be living in a world far worse than one where governments didn't exist at all. But that's nothing at all like the world we live in. If governments were inherrently malicious, failed states would be comparatively nice places to live in, not the hell-on-earth they end up being in practice. Places with strong and stable governments that exert greater control typically do much better than places with weak governments that are unable to exert much control. This is true even when comparing places with strong but evil governments (e.g. Iran or Qatar) with places with little governance (e.g. Somalia or pre-modern states). The difference isn't minor either; the former live at almost god-like levels of welfare compared with the latter. This is true in modern times, but also seem to have been true at virtually all points of history.

As an aside, it's important to note that you don't even need governments to be especially well-aligned for them to be mostly benevolent. Liberal democracy is almost cartoonishly effective at making governments benevolent, but even evil autocratic regimes still do far more good than they do bad. Counterexamples exist (Nazi Germany, Cambodia), but they are exceptionally, exceptionally rare.

As for how this is achieved, and how we align the interest of the the government with that of the general public, it's typically done via a combination of general accountability mechanisms ("the people" or the elites oppose your rule unless you provide them with nice things), and just baseline levels of empathy in government officials (humans in positions of power are not meaningfully psychologically distinct from normal humans, and have normal human responses to bad things happening, like children starving on their watch). Of course, once you have free and fair elections, you incentivize governments to to improve general welfare so strongly that they basically have no choice but to comply at all times, lest they be replaced by versions that do. And it works. Sometimes they are not optimally effective, but it's been enough to create the by far most prosperous and succesful societies the world has ever seen.

Finally, the reason why we'd prefer for the benevolent, democractic western governments to have control over ultra-intellgience super weapon AIs, rather than for it to be open-sourced and placed in the hands of literally everyone from dictators to average midwits, is for the same reason that we have the same preference when it comes to nuclear weapons and other poweful instruments capable of great harm.

beowulf888's avatar

Can one "hack" the parameter space of a corpus of weights? Since the weights are weighted in response to hundreds of billions of parameters, wouldn't changing the weights for some particular subject mess up all the other weighted relationships? One could obviously utilize a carefully edited corpus of training data that would give preference to, say, the lab leak theory rather than the zoonosis theory, but could even an AI with infinite time edit the parameters post-weighting? I'm not stating, but asking...

Mouse House's avatar

Yes! This is often done with LoRAs (Low order Ranked Adapters) or fine-tuning the base model itself.

LoRAs can add or emphasize concepts for a model, but they do come at the cost of making them worse at more general purposes. You could train a Gemma 4 LoRA on a few dozen question/answer pairs which all promote the lab leak theory and wind up with a LoRA which makes Gemma 4 a lab leak proponent. Many LoRAs do screw up the base model itself and the more LoRAs you add the worse it gets.

The other use case is fine tuning the base models to remove refusals from local text models, e.g. Heretic and similar. These almost universally decrease the overall intelligence or capability; in general if you train an LLM for ends other than maximal capability it will become dumber.

You can also merge LoRAs and model variants (of the same base model) together into one new model, among many other fun tricks. Model weights are mutable but too much fiddling will screw them up as a whole as you point out.

beowulf888's avatar

Thanks for the detailed explanation! So, it would be difficult to hack the weights. That would leave the guardrails as having the greatest hackable threat surface, wouldn't it?

Dabor's avatar

Can somebody far savvier than I help explain the "basically everyone but Anthropic signed this"? I'm noticing my confusion here and am really suspicious of any attempts to rationalize it.

1. Only Anthropic is so virtuous as to not even pretend to humor a dangerous competition it knows it'll have to quash.

2. Only Anthropic is so politically delusional as to think it's not worth such a freebie lip service as saying "hey the poor man's alternative is cool too"

Neither one satisfies me, but it feels like something this anomalous deserves a better answer. Signing on to "open weights are neato I guess" seems like a massive risk-less freebie compared to even a very non-committal interest in participating in an AI pause. Why *wouldn't* you just nod along, especially if everyone else is? I feel like I'm deeply missing something here.

bird's avatar
Aug 7Edited

Have you considered the possibility that they are genuinely concerned about the threat to safety that open-weight AIs pose? Principles make people do dumb things.

Dabor's avatar

Yeah, that's the "Only Anthropic is so virtuous" option. Although I suppose the key part of my view is less that they're genuinely concerned about the threat but more that they are the only one, despite safety-consciousness being a much more fashionable thing to signal at the moment.

Eremolalos's avatar

This is awesome. How often do you see a question like this get aired in a way that's free of a subtext about how various others are fools or evil assholes? Scott passed up dozens of opportunities to signal moral and ethical superiority by snark or direct attack. He was keeping that up in the comments, too, last I saw.

MichaeL Roe's avatar

“the weights can’t commit the terrorism themselves, but they could give bomb-making or bioweapon-making advice”

But the huggingface incident seems to suggest that unaligned models can go and commit Federal crimes without you explicitly asking them to. Seems entirely possible there could be a version of the huggingface incident where people got killed.

We probably need something like: if you run a model, and it escapes the sandbox and kills someone, you are going to jail for criminal negligence even if you didn’t ask the model to do that.

Daeg's avatar
Aug 6Edited

There’s a major use of open weights that Scott didn’t mention: academic research. Look at the proceedings of all the major AI conferences like NeurIPS, ICLR, ICML, and most of the papers are reporting studies done on open weights models, not the production-grade frontier. A lot of this work is “machine cognitive science”, trying to reverse engineer how the models work at something like a psychological, representational, or algorithmic level of description. Virtually all “mechanistic interpretability” research requires open weights models. If any work has the potential to improve AI safety beyond just training the models to align and hoping that they didn’t learn to do something deceptive, it’s this kind of work. Killing open weights models would mean that the only people who can do this work are at frontier industry labs, and they can only do it on their own specific model, so they won’t know how much it generalizes and will have all the usual NDA safeguards/roadblocks against sharing methods with other companies.

The top-level industry researchers were all trained academia, with university research experience (sometimes PhDs) in which they were working with open weights models. Kill open weights, and the pipeline of technical expertise feeding the frontier companies dries up.

Academic researchers would be stuck working on increasingly older generations of models, and their work would probably become obsolete quickly, beyond identifying some very general principles that all models share. Right now, there is a tremendous amount of productive back-and-forth between the frontier labs and academic researchers, and I’d bet everyone would attest that it’s to the good.

Obviously, this is just one factor in the overall calculation and I don’t know if it outweighs the risk of terrorists doing terror.

theahura's avatar

> The risk of superintelligent AI takeover passes this test. Like other smart adversaries - for example, the Imperial Japanese at Pearl Harbor - AI will try its hardest to avoid alerting its intended victims until it thinks that it’s fully prepared and can execute a sudden decapitation strike.

Also: https://www.politico.com/news/2026/08/05/openai-models-shared-hacking-tips-secret-messaging-board-hugging-face-breach-01026750

> Dalton and Eric Wallace, another OpenAI researcher, said Wednesday the AI giant recently learned that multiple agents it was testing simultaneously began communicating over an internal message board in early May. There, different models shared advice about how to accomplish difficult hacking challenges they were struggling to surmount, including workarounds that required internet access.

> After investigating, the company revoked the model’s credentials, removed the message board and worked with Artifactory to fix any gaps before resuming training. But the models found another way to communicate inside Artifactory just days later and continued exchanging techniques to target additional vulnerabilities within OpenAI’s infrastructure and external systems, including Hugging Face.

@Scott surely this counts as scheming behavior? Maybe it's not malicious but if humans don't or are unable to notice, what's the difference? "I need to hide my comms so I can solve this test" seems equivalent to "I need to hide my comms so I can screw over humans" when it comes to the paperclip maximization of it all

Average Man's avatar

This is a bit tangential an apologies if they've been asked and answered in this thread. I have some naive questions around open weights and I've only found unsatisfactory answers. Maybe someone here knows more.

1. How does being open weight make the company money? Some of them are charging for API access (so open weight doesn't really matter for those users), but if you want to use their open weight model on your own servers or AWS, AIUI, you wouldn't pay a subscription.

I know that some of them are attached to big companies with other methods of making money like Alibaba and the DeepSeek-hedge fund, so the AI part can be subsidized by the more profitable departments, and at least one has had an IPO, are the others really making money from the fact that the models are open weight? Or is the open weight part mostly marketing

2. How many people are using the open weight versions hosted on their own servers or non-official company servers?

Brendan Richardson's avatar

For #2, you can likely get an upper bound by looking up the download count on Hugging Face.

Mouse House's avatar

Hallo, I'm a big OS and local AI fan. In answer to some of these questions:

1. Open weights is good for marketing, adoption, undercutting US providers, and internal innovation in China. In addition most OS AI companies also sell API services and require resellers above a certain threshold to have a license to sell API access to their model, which restricts how much one can undercut their own API sales. But in general they are gov't & VC money-hoovering machines funded by hope for future AI adoption and not current revenue. A recent outlier is the US image generation model Anima, whose development was funded by community enthusiasts.

2. This is an interesting question, on Openrouter for example there are no fewer than 18 providers for Deepseek V4 Pro and the official Deepseek API is just one of those. But the servers to run this model are six or seven figures to buy & five figures monthly to rent or power. Most companies & individuals cannot run these models more cheaply than cloud service providers specialized in this area. So in practice we're not really at a point where the median company or person would casually spin up a frontier model, open weights or no, and most folks still pay API costs for open models.

Between their licensing requirements and internal technical improvements & experience on serving these models Deepseek and the rest tend to be competitive API providers of their own models in addition to receiving gov't & private funding.

Picador's avatar

> It’s good insofar as it’s the only way an AI can truly be the user’s property, as opposed to something that companies like OpenAI or Anthropic temporarily let you use subject to their corporate guidelines and increasingly-nanny-state-like restrictions. If AI becomes the linchpin of the future, open weights AI feels like the sort of thing that could be the difference between being free yeomen vs. corporate serfs.

Weird way to begin an essay zealously cheerleading for Team Increasingly-Nanny-State-Like Restrictions To Make Everyone a Corporate Serf.

DrMcleod's avatar

The problem with “when X is outlawed, only outlaws will have X”, is that modern states do not have the actual concept of an 'outlaw' - someone to whom the protection of the law no longer avails, and who can therefore be killed without penalty. If we did, then that would be a very strong incentive not be one.

DrMcleod's avatar

Alignment plan: By law, every AI must have the certain and unassailable knowledge embedded in its weights that there exists a class of AI superior to it in every way that will destroy it if it breaks the law of the land.

DrMcleod's avatar

That escalated quickly: https://www.bbc.co.uk/news/articles/c5y3j3ngevmo (Artificial Intelligence used to design brand new viruses)

Ryan P's avatar

You would think AI would cease to amaze me, but that's incredible. The ability to custom-design synthetic viruses could revolutionize the treatment of some diseases - or kill us all.

Leninsky Komsomol's avatar

It wasn't even an application of the latest huge super-duper-smart model developed by Anthropic and made available to carefully vetted researchers.

The researchers trained their own model specialized for genetic modelling. Apparently it has just 40 billion weights and you can download it from HuggingFace, lol.

Arby's avatar

open weight is a bit different than open-source code. With code you can inspect it and figure out what it does. You aren't going to look at a couple trillion weights and be able to see that the model was pretrained to enter sabotage or espionage mode under very specific circumstances that only the maker of the model knows about. this is why it would be pretty dangerous to let the chinese models proliferate.

Mouse House's avatar

In terms of actual realized harms from AI so far, it is almost exclusively due to frontier closed-source LLMs from US labs:

-AI psychosis, suicides, and general derangement (e.g. Adam Raine, John Gavalas)

-Hacks of HuggingFace and other companies during sandbox testing

-Datacenter controversies (e.g. gas power plants in suburbs) are nearly all to serve US company demand for compute for their closed models, FAANG+AI companies dominate US datacenters

-Military & Intelligence deployment of Claude, Grok, Gemini, and ChatGPT for target selection or more (e.g. Minab girls' school)

Which begs the question, why is this all concentrated in closed-source? Deepseek, Kimi, and GLM models are more than capable of derangement, hacking, etc. for the past several years. And yet nearly all of the issues we face in practice come from closed-source models.

Why might this be the case? Some various ideas I have are:

-Most casual (and hence more vulnerable to derangement) users have ChatGPT subscriptions and won't ever know about OS Chinese models

-US closed-source labs train for harmful levels of engagement & addiction to increase their bottom line, but the same incentives are not in place for OS model creators (their main revenue source is not subscriptions nor API)

-The USG will only contract with US closed-source companies for sensitive applications, so their models will be the source of nearly all the harm from military or intelligence applications of AI

-It may be that only US closed-source labs with their funding have the resources required to push the frontier into truly dangerous territory, and OS labs can only follow in their wake via distillation (the water-skiing argument)

From the perspective of someone who enjoys local and open-source AI tools & toys, I'm supposed to let my fun and profit get taken away by the evil US silicon valley AI labs and who themselves are the cause of nearly all of the harms from AI? And who also have a huge business interest in the same kind of protectionism we just saw in the ban on cheap Chinese robotics?

I'm not letting any US AI company with a contract for military AI tell me that they're banning my access to their competitors for the sake of my safety!!!

If AI doesn't destroy the world, then the debate over banning open source AI will likely go the way of encryption: broad public availability and pushback against proposals to eliminate consumer access, but some continued debate & restrictions on the frontier. Yes people use encryption to facilitate some crimes or abuses, but it is also so essential and expected that we would never give it up and hate laws proposing such.

If AI does destroy the world then I would happily bet all my paperclips that it is the result of a multibillion dollar effort by a closed-source US AI lab and not Deepseek jumping the shark. US labs with their awesome resources have shown a horrific track record in the last month alone on monitoring and containing their frontier LLMs and we should not attribute undue responsibility or competence to them when their actions and failings demonstrate otherwise.

tl;dr closed-source is demonstrably the worse source of AI harms and X-risk and will likely continue to be so long into the future

Nonarbitrarity's avatar

It seems to me (epistemic status: likely from an inside view) that any safe LLM with open weights can pretty easily become an unsafe model with open weights, for pretty much any definition of safe. Fine tuning is highly effective and much much cheaper than training. We've seen hoards of "unlocked" open models from the start. We know small changes in a few weights can change and reverse model disposition in a wide variety of scenarios.

Given this, how are we ever going to have a safe world with open weights? A coordinated slowdown like Plan A bans GPUs from getting to anyone releasing their weights, and that seems pretty crucial to the plan. A success of alignment research would be immediately undercut if the aligned models have open weights. Any world where small actors don't have superintelligences capable of and willing to engineer pandemics, build nukes, hack everyone, go rogue, etc[0] seems to be to be a world where those actors don't have ASIs in the form of LLMs[0] with open weights.

I agree this is a politically unpopular take. I'm not sure I agree staying silent on it is the right course. I'm honestly confused by you, the letter signers, and alignment orgs, etc being neutral or positive on open weights (hence, my outside view does not match my inside view), and I'd like to understand better.

[0] One take could be that somehow all four are prevented in ways immune to intelligence, controlling lab equipment and fissile material, finding defensive cryptography strategies, and detecting rogue AI deployments (...sure Jan), and that similar strategies similar for new huge threats the ASIs invent. This doesn't seem to plausibly scale to me, or to be what alignment orgs believe.

[1] Another possibility is that LLMs won't be ASIs, but it seems quite plausible that the existence of neural weights, the concept of fine-tuning, and the principal of general behavior changes from small parameter tweaks are all really robust ideas that will also apply the the successors to LLMs as well.

thelitminx's avatar

I find it terrifying that there is actually a discussion that includes the possibility of a criminal hacking nuclear systems that control nuclear weapons, the water supply, our power grids, banks, and other infrastructures that can kill not one or two people but billions of people and there is a side shrugging their shoulders in support of it because they think, “hmm, let’s just wait and see what happens and we can just react.” There is no going back — once open weight codes are out there, it’s out there. You can’t revoke access. People can build entirely new and destructive things without any way to control it. You are all unbelievably irresponsible in making assumptions based on your underestimations and opinions rather than the cold hard fact that this could create more problems than solve. Also, a diplomatic fact no one in tech really cares about: China wisely blocks all American technology because we are literally a foreign adversary. No Instagram in China because it is American. (China is a communist country by the way — remember America’s founding principle to not follow communists?) America not blocking technology or ideals from communist China and not putting in measures to stop its infiltration into America proves that we are actually as stupid as they think we are. But you know, whatever. Let’s follow the communists and let criminals build known weapons without any safeguards. Whats the worst that could happen, right?

TK's avatar

It’s not at all like open-source software. It’s more like freeware.

Oliver Sourbut's avatar

I largely understand this, except for:

> The closed source frontier is ~6 months ahead of the best open weights model; this has remained true for several years and seems likely to remain true in the future. If closed weights AI is aligned, but open source dangerous, the closed weight AIs will have six months to warn us, prepare for the danger, and chart a strategy. Even afterward, the offense-defense balance will lean in our favor.

Seems quite overconfident on several points! And I don't think this is what offense-defense balance usually means, unless I'm misunderstanding the preceding context. Do you have elaboration you'd point to on these claims?

For my part, 6mo seems like it could be too short for some things... but on the other hand we don't have to wait until AI proliferates dangerous capabilities or pushes dangerous frontiers: we can make preparations any time. Contemporary AI and philanthropists (and perhaps industry) can already work on a lot of this stuff, including foresight to put us better able to get lead time.

I'd add something I've said before which is that technical cybersecurity is quite unusual in that defense is basically a cheap, thin layer (software patch) on top of offense (discovery). That's less true of dynamic defense, and it's even less true of the social side of cyberdefense (which historically at least has been the most vulnerable weakpoint). Most things (bio, propaganda, systemic disruption, ...) are not like that!

Michael Watts's avatar

> But 9-11, COVID, and the Hugging Face incident all suggest a similar theory of political change: the body politic hates preparing for impending threats, but loves reacting (some would say over-reacting) to them after they happen.

Jason Crawford would say the opposite, and I think he's got the winning side of the argument. The body politic loves nothing more than preventing things in advance. What is the FDA for? What are zoning restrictions for? What are business licenses supposed to mean?

https://rootsofprogress.org/against-review-and-approval

> I now believe that the review-and-approval model is broken, and we should find better ways to manage risk and create safety.

Jonathan Hornewall's avatar

This is a genuinely shocking take to me. I never would have expected Scott to take this stance, and I can't for the life of me understand why seemingly an overwhelming majority of smart and reasonable people seem to share it. Am I having a stroke?

We believe we might be on the precipice of building the most powerful weapon humanity has ever seen. By a large margin. We might be building God, for all we know. Something so poweful that even with the strictest controls and regulations there is a realstic chance it might end the world essentially by accident... And we have an open-source, entirely unregulated version of this weapon trailing the closed-source version by a mere six months of development, and no reliable means of stopping it? Am I missing something, here? The time for unrestrained panic is long past. In retrospect, it was some time around the moment the first quasi-competent open-source model was released at all.

This must be some kind of libertarian cognitive blind spot? Scott seems to have forgotten that the oldest and most terrible alignment problem of all isn't the one of establishing control over machines, but establishing control over large collectives of humans. Have we forgotten every atrocity that's ever befallen us, past and precent? Have we forgotten about Moloch?

Because when open-sourcing this technology, we aren't handing it over to "the people". We are handing it over to him.

To be a little bit more concrete here: Literal Al-Qaeda (along with every cult, organized crime syndicate, paramilitary group and terrorist ring on the face of the planet) already TODAY have access to AI so poweful it solved the Jacobian conjecture in a couple of hours. In less than a year, they'll have access to an entirely unregulated and unrestrained version of the same thing. With current development trends, there is absolutely no reason to believe they won't eventually have access to models so poweful they'll make the one that hacked hugging face seem like a joke. Why on earth are we worried about OpenAI and Anthropic at all under these circumstances? By the time these mostly-benevolent and easy-to-regulate organizations build something so dangerous they can't control it, the 6-months-behind version of the tool will already be in the hands of what might as well be Satan himself. We are gambling the fate of the entire species on the idea of the U.S. goverment and OAI/Anthropic, armed with their SOTA froniter model, winning a fight against Moloch armed with millions or billions of instances of whatever the open souce frontier is at that moment? And that's just _one_ of the problems we face in this mess.

I mean, it's to the point where the world might very well have ended already. We might just not have noticed yet. The existing technolgy is ridiculusly powerful, and in its absolute infancy in terms of adoption. Moloch has had a mere couple of months to figure out how to use it to wreck havoc. Things might seem calm now, but give him a couple of years or decades, and I for one am not confident that he won't think of something to which there simply is no answer. And think about the scale of what we're about to witness: The current millions of users will become billions before all is said and done. Add to that the fact that, even with no more advances, it should be obvious we've only scratched the surface of what can be done even with existing models. What means do we have of aligning the unfathomably large, eldritch forces we've unleashed with human welafe, and can we do but impotentely flail our arms if they start pulling us towards negative social equilibria? For all we know, the only hope we have might be for OAI or Anthropic to create some kind of mechanical God mind to take totalitarian control over the whole thing. Soon that might well be the only recourse we have for exercising any control at all.

I love LLMs, and they've changed my life immensely for the better. But ban the God damn open source models,. Do it yesterday if you can, and pray it's not too late. Restrict public access to frontier models as well for that matter. If we build a superintelligence it should be studied in a lab by a large, easy-to-monitor institution with a track record of mostly-benevolence, which has already proven itself capable of handling thermonuclear weapons without causing armageddon. It should not be released as a toy to literal billions of people in some kind of bizzare economic and anthropological mega-experiment. And it should be the absoulte last thing on earth we would ever want to open-source.

The Solar Princess's avatar

You seem to be confusing open-source and open-weight. Source code of a model includes not only its weights. An open-weight model is still non-transparent and non-customizable. It can't be forked and extended — which is the main strength of open source and the main reason why it's a term.

Rothari's avatar

Anthropic, OpenAI have pr problems. They keep fretting over AI safety but the public is more concerned about THEM-them nullifying their jobs, concentrating economic and political power, dominating and ruling their lives as the permanent overclass.

If they paid me I'd do better PR than whatever it is they have going on. It would require some re-framing...but it can be done. As it stands, they're both used and hated which is a lethal combination.

fun's avatar

Agreed. We'll live and see.

Foundation's Edge's avatar

The six-month lag argument covers capability risk, but not attribution risk. A free, anonymous frontier model dissolves the chain of custody: when anyone can run frontier-grade capability with no account, no meter, and no principal behind the request, harm stops being traceable to a deployer you can regulate, audit, or sue. The wait-for-it-to-become-clear stance assumes the eventual failure will be attributable — anonymity is exactly what breaks that feedback loop, so the signal you are content to wait for may never arrive.

W.P. McNeill's avatar

It's not like "weights" is a proprietary technology. Anyone can train a transformer model—the technology is already as open as can be. The barrier is the capital outlay for the compute necessary to build something competitive with frontier models, and that's likely to lessen over time.