171 Comments
User's avatar
User's avatar
Comment deleted
Aug 3Edited
Comment deleted
Linch's avatar

this seems less useful than the numbers Scott's post gave.

Scott Alexander's avatar

Did you read the post, or is this just a comment about the graph on the top and the section around it?

Ted Sanders's avatar

I did in fact read the whole post, though I suppose this primarily reacts to the intro. Apologies if my comment was low quality.

Performative Bafflement's avatar

> performance scales with effort

If performance scales with effort, this also vastly privileges AI, because you can have your harnessed AI crunching along at many times human speed for days and dedicate a lot more brain / clocktime to any given problem.

Mark's avatar

Does this imply that current AI forecasting scores are overestimates, because they are making money based on greater effort rather than greater understanding? (And AI progress over time might reflect growing ability to stay on task rather than growing understanding)

Or maybe I'm missing something.

Ch Hi's avatar

I'm rather sure that performance is highly sensitive to effort, but I'm also rather sure that there's a lot of chaos in predictions, and probably some amount of actual randomness. This will affect where the ceiling is, but my guess is that the ceiling will never be reached, because the effort to increase capability is non-linear. My guess is that it's worse than O(n^2).

Taleuntum's avatar

From the post:

"I find these numbers more interesting than the worn-out debate over whether AIs or humans have reached the “max” and whether “perfect forecasting is impossible in principle”. It should be uncontroversial that the max is somewhere above current levels, even if only a fraction of a percent (for example, because even the best human forecaster sometimes has a bad day). And it should be uncontroversial that even the best possible forecaster can’t get 100% on everything (because you could ask questions about quantum effects that are unpredictable even in principle). Now that we’ve demonstrated that the answer will be some particular number, we can get to debating what number it is."

So you are in agreement with Scott.

Alex Zavoluk's avatar

> Instead, we go to absurd lengths to keep the outcome uncertain.

This is... an oversimplification, at least. Sports outcomes range from "mostly random noise" to "almost pure measure of ability" (like running or swimming or weightlifting) and even the latter get watched. Many of the things mentioned in this paragraph don't apply to individual sports, for example. Certainly there's no strong correlation between uncertainty of outcomes and watchability. If anything, people probably intuitively underestimate how predictable outcomes are based on past performance, because inventing narratives where the underdog wins is fun and makes for TV that lots of people will watch.

Without looking at the data, I would guess that absolute error is highly sport-dependent (although if the simple model includes something like Vegas betting odds, the relative error of top forecasters compared to that simple model would probably be similar).

David J Higgs's avatar

Most of the relevant factors ensure that you have roughly equally "skilled" (or strong or whatever is important) competitors. Team dynamics certainly increase uncertainty further, but there are still all kinds of rules, social structures, competitive dynamics, etc. that ensure weightlifters for example are going to be very similar in strength competing at prominent events like the Olympics.

Geopolitical events like the US abducting Maduro from Venezuala via a spec ops raid are not like that in general (though some can be like competitive elections)

Alex Zavoluk's avatar

You're making the argument that sports are relatively *more* predictable, which seems to be the opposite point of the line I quoted.

David J Higgs's avatar

No the similarity in competitors makes the outcome less predictable, like flipping a coin is unpredictable (you can't predict heads or tails and do better no matter how smart you are, and it's almost as hard to predict who wins the weightlifting competition). By the same logic, the dissimilarity along countless axes makes Geopolitical events potentially more predictable

Ch Hi's avatar

FWIW, and IIRC, heads is slightly more probable than tails. Of course, that may depend on the coin being flipped. As I recall, the study didn't specify that a range of coins was flipped, only the number of flips and the results. (It was less than 1 part in a thousand difference. IIRC it was done at Stanford in the 1970s or 80s.)

Simone's avatar

Geopolitics is like boxing if the reigning heavyweight champion could challenge a student club to a title match, the student club couldn't refuse, and the match was to the death.

Pete's avatar

As you say, running is an example of "almost pure measure of ability", however, betting on races (both human and animal e.g. horse) is a widespread established tradition, so clearly there is enough unpredictability to make it nontrivial and interesting.

Alex Zavoluk's avatar

People will seemingly gamble on anything. That doesn't mean that the outcome is not mostly predictable from past performance. (Although I did sort of gloss over a distinction, it's theoretically possible for performance in a sport to be mostly driven by ability, but outcomes still mostly be luck dependent, if all competitors have very similar ability. However, my understanding is that running and swimming events have some of the lowest outcome variance of any sport.)

Simone's avatar

What makes races more unpredictable is selection effects, athletes go through several rounds and only the best move forward, so at the top end there's just a selection of extreme outliers. And even then the uncertainty usually boils down to "who will win, the favourite or the next best guy".

Hafizh Afkar Makmur's avatar

I don't know why I haven't ever thought about handicap match between AI and humans, but I never thought the gap would be that "small". 600 ELO difference just seems so big, but what I know in other boardgame like shogi, best human can beat best amateur through a bigger handicap than that.

Luke's avatar

Handicaps increase in Elo value, sometimes by a lot, as players get stronger. 3 pawns or a minor piece is game deciding for grandmasters, but a moderate advantage for a casual player.

Intuitively, as players get stronger, there's less variance in their play, so small advantages are much more persistent. Said differently, an expert is less likely to blunder away their advantage than an amateur.

John M's avatar

The chess anchor makes this post significantly understate things in my opinion. Chess is based on a far simpler set of rules than forecasting events in the real world. The more simple a game is, the lower the level of intelligence it takes to saturate it.

It's easy to understand this when you consider something even simpler than chess like tic-tac-toe. I could probably play tic-tac-toe about as well as a superintelligent AI but that doesn't tell you much about the AI's performance in more complex scenarios.

Jerry's avatar

Or even checkers, in between chess and tictactoe

Scott Alexander's avatar

What you say makes sense, but chess actually ended up being the most ambitious of the three anchors!

Skornne's avatar

This is a very tempting analogy but I don't think it applies. Increasing the complexity does add to the number of elements under consideration (AI advantage), but in real life scenarios it also inevitably makes it noisier. As Scott brings up, team sports prediction (say, American football) is much more complex than either tic tac toe or chess, but the lack of determinism in real life means predictive ability will saturate anyways regardless of intelligence.

John M's avatar

Team sports is designed for unpredictability though.

Aeolian321's avatar

Paterno's "trial in public" about sexual abuse of minors was a prime example of how sports prediction can get bollixed up by "known issues off the field."

Simone's avatar

It's the other way around IMO, the more abstract a game, the better intelligence can be leveraged for it. Roulette is a very simple game but an ASI couldn't predict it.

elopotion's avatar

Chess might also overstate the case for AI outperforming humans. Because the game is played with complete information, the ability to calculate (chess word for doing move-by-move predictions) is really useful compared to many other domains. Humans ***suck*** at calculations compared to computers. The difference between a grandmaster chess player's calculation ability and Stockfish's calculation ability is similar to the difference between a layperson and a graphing calculator. This is leveraged to great effect, with chess computers calculating millions (generic large number) of moves before making a decision in game states where human players would only do dozens. When Kasparov beat Deep Blue in individual games, it was because he was orders of magnitude better at deciding which lines were worth looking at. Nowadays, chess programs are much better than Deep Blue at identifying fruitful lines of inquiry and they also have more processing power. I don't know if they're close to humans yet in terms of selecting good lines to calculate.

In messier domains, this type of specific calculation is much harder because the number of possible outcomes explodes so quickly that we have to develop heuristics. AI might end up better than humans here, but it will probably outperform humans by a narrower margin in a domain where it can't leverage its ability to calculate specific positions.

Peter's avatar

What I want to know is about weather forecasting. I know as of four years ago the National Weather Service was against automation, standards, or even quality control, and vehemently against AI for the same reason, meteorologists are oracles working in entrails according to NWSEO and no mere mortal, supervisor, or AI, can handle that. Curious how in the rest of the non-Luddite weather world, the AI has been stacking up.

DrMcleod's avatar

Presumably they use computer modelling though?

Peter's avatar

As an optional input data point only, NWS mets are free to discount it as much as they want for the products they put out based on their own personal interpretation of the animal spirits. Also while they get multiple computer model outputs (from NCEP via AWIPS) , the models are highly currated politically, not necessarily the most accurate or accepted ones.

But like I said, this is a post about AI. I'm curious how the non-Luddite weather organizations are doing and how the AI forecasts are holding up vs human forecasters. I assume larger organizations which don't depend on NWS forecasts are moving ahead with AI forecasting as they care about accuracy and public safety (as opposed to union jobs), I.e. DoD, FedEx, other nations, etc

Alastair Williams's avatar

The ECMWF, which by most measures is far ahead of the equivalent American services, has been doing a fair bit of work on LLM assisted forecasts.

They have an operational AI assisted forecast: https://www.ecmwf.int/en/about/media-centre/news/2025/ecmwfs-ai-forecasts-become-operational

A year ago this was:

- outperforming "state-of-the-art physics-based models for many measures, including tropical cyclone tracks, with gains of up to 20%."

- offering "increased speed and a reduction of approximately 1,000 times in energy use for making a forecast."

Peter's avatar

Ty on that link

Steve Sailer's avatar

Weather forecasting has gotten vastly better over my lifetime.

Nancy Lebovitz's avatar

My weather app doesn't just get the weather wrong, which isn't too surprising, but it frequently doesn't agree with itself. That is, the hourly predictions don't necessarily match the overview for the day.

Sniffnoy's avatar

> Either humans have already come close to some fundamental limit on the predictability of world events - in which case AIs will plateau at or slightly above the human level - or the trend line will continue until AIs are far beyond top humans.

I don't think it's very likely, but I think there's a logically possible third scenario here -- humans and AIs both continue to improve but at about the same rates. For instance, if AI's reliance on training data somehow caused it to plateau at human-level, but that level improved as humans improved. I don't think this is a likely scenario as this doesn't seem to really be how modern AIs work, but worth noting, I think.

Xpym's avatar

The best human will trivially never perform worse than the best publicly available AI, because there's no good reason for him to not use it, so it would be difficult to separate these.

Domo Sapiens's avatar

you can try and pay the best humans to work as control groups free of AI. Chosen well, that could offset their expected loss of value from not using AI.

David Manheim's avatar

The same logic would predict that AI-human hybrids would never perform worse than pure AI at chess, but in fact, that is the current state of things.

João Bosco de Lucena's avatar

Important to note that this is true only in the sense that the human could exclusively use the AI output. In chess, where AI far exceeds the best human, any human input would only serve to make the game worse.

Xpym's avatar

Have there been serious competitions to establish this? It seems to me that there's at least one way human judgment can still contribute - there are several engines of similar strength that sometimes disagree, and there likely are classes of situations where some tend to perform better than others, which could be analyzed and exploited.

João Bosco de Lucena's avatar

So centaur chess was a thing for a while, where humans + engines did beat engines alone

As far as I know this mostly just does not work anymore, and from my brief research centaur chess seems to just have dwindled (in the 2017 Infinity Chess Ultimate Challenge a human-engine centaur got third in a championship that had both engines and centaurs). No one has really tested this in the last 5 years, but I suspect that maybe if you had a grandmaster playing with an engine in a correspondence game, they could at most draw against an actual engine if they were forced to actually play some moves themselves and not just follow the engine line directly.

It makes sense too because engines have an almost 1000 point ELO advantage over Magnus Carlsen, it's like expecting Magnus to get insight from a club player, just seems very unlikely.

> And there likely are classes of situations where some tend to perform better than others, which could be analyzed and exploited.

maaaaybe, but I'd expect these lines to be hard enough to calculate (e.x. 5-10 move tactics in complex positions), and the resulting positions hard to evaluate, which means the human would be spending an inordinate amount of time to calculate and then possibly still make the wrong choice.

I sort of expect the same thing to happen with predictions, where right now a human aided by AI does much better than AI by itself, but won't always be the case.

Xpym's avatar

>which means the human would be spending an inordinate amount of time to calculate and then possibly still make the wrong choice.

No, I mean something like (simplified imaginary example) Stockfish tends to do better in highly imbalanced tactical positions while Leela has better planning in quiet ones, so depending on what's happening you're taking the first choice from one or the other.

Ch Hi's avatar

The most likely reason for that third scenario is if the AIs discovered an algorithm that was simple enough for people to use. This isn't unreasonable, as an analogy has happened with some of the proofs of Erdos problems.

Bugmaster's avatar

> AI first beat the human champion at chess in 1997

No, it didn't. That is to say, yes, a computer did beat a human at chess; but this computer wasn't an "AI" as the term is usually understood today, i.e. some kind of a transformer-based LLM (and in fact modern LLMs tend to perform comparatively poorly at chess). Rather, it was a relatively simple minimax algorithm, supplied with a massive (for its time) amount of CPU and RAM.

The implications of this are IMO much more interesting than the same old "LLMs are superintelligent AGIs" take that dominates the discussion today. Chess was touted as the pinnacle of human intelligence, but as it turned out, the game was simple enough so that a one-page algorithm could play it (and win). This echoes the earlier history of computing, when the task of tabulating numbers was thought to be too abstract to ever be performed by machines.

My takeaway from this is that "intelligence" is not some simple scalar number that can describe everything that human brains do. Rather, there are lots of things we do very inefficiently compared to relatively simple machines; and there are other things we doing that appears to be beyound their capability for now. It is unfortunate that LLMs have eaten up all the R&D resources in search of that "one quick fix", because personally I'd love to see more discoveries that could potentially reduce other problems to simple algorithmic scenarios -- like chess. Or maybe we would've seen a definitive proof that this is impossible (though I suppose this probably borders on disproving P=NP). But I guess we will never know, until the LLM bubble bursts or at least deflates enough to allow other research to proceed...

Lucas's avatar

AI is not necessarily understood today as a LLM, lots of stuff is still called AI and is regular machine learning. For example, stockfish is obviously AI, obviously machine learning, obviously way better than deep blue was and also more efficient.

What makes LLMs especially interesting is they can absolutely help find ways to reduce other problems to simple scenario.

The rethoric about the LLM bubble and the "one quick fix" seems as pervasive as it is wrong, LLM work, scaling works, models keep getting better and cheaper, and there's no "one quick fix" needed, most (all?) benchmarks thrown at LLMs gets climbed sooner or later.

Other research is absolutely allowed to proceed, for example Anki is experimenting with LSTM to help reduce review count/increase retention. Good old fashioned machine learning is still very useful in retail, especially online. But I'm not aware of any other approach that has results as good as LLMs in terms of "general intelligence", or even just for searching/crafting an answer and coding.

Bugmaster's avatar

> For example, stockfish is obviously AI, obviously machine learning, obviously way better than deep blue was and also more efficient.

I would agree with all of that except for the "obviously AI" part. That is, I'd call it "AI", but then I'm old, and I still remember when things like alpha-beta pruning and simulated annealing were called "AI". I would argue that today the term "AI" is virtually synonymous with LLMs; this is how most people use it even inside of academia, and this is how Scott uses it the overwhelming majority of the time.

> scaling works, models keep getting better and cheaper

Are they ? We are now at the point where companies are resorting to shredding (and scanning) out of print books in a desperate bit to get just a bit more training data. Once they are done scanning the last book, then what ?

> for example Anki is experimenting with LSTM to help reduce review count/increase retention.

I'm not sure what Anki does, but I'll take your word for it. LSTM is still a NN-based approach, but of course it's not an LLM, so you are correct there. That said:

> But I'm not aware of any other approach that has results as good as LLMs in terms of "general intelligence", or even just for searching/crafting an answer and coding.

As I'd argued before, the term "general intelligence" is so poorly defined as to be borderline motte-and-bailey; but I agree with you on that last point: LLMs are great at interpolating an answer out of the training corpus. If you use a conventional search engine to index a bunch of documents on rabbits, and a bunch of other documents on pocket watches, you could find information on either topic very quickly; but only an LLM could generate a document (based on the same training corpus) that describes a rabbit wearing a pocket watch. LLMs are so good at coding (comparatively speaking) because most coding tasks are of the rabbit-pocketwatch variety. But there are lots of other tasks, such as e.g. navigating a car down the street in real time, which LLMs are unable to solve at all.

Andrew's avatar

> for example Anki is experimenting with LSTM

That would be news to me! I'm one of the people working on FSRS, a hand-crafted spaced repetition algorithm that only has 34 parameters (well, FSRS-7 which is not in Anki yet has 34, FSRS-6 has 21) that has nothing to do with neural nets. I am working on a neural net with ~500k params, but unless you're one of like 5 people on the Anki Discord server who care about that, you wouldn't know about it. And it uses a RWKV architecture, not LSTM.

I'm actually very curious what made you say that Anki is experimenting with LSTM, because I can't think of any source that could be saying so.

> The rethoric about the LLM bubble and the "one quick fix" seems as pervasive as it is wrong

Just wanted to say that I agree with that.

Bugmaster's avatar

I actually know next to nothing about RWKV besides its existence (and lack of attention); does it really perform as well as a transformer ?

> The rethoric about the LLM bubble and the "one quick fix" seems as pervasive as it is wrong

Oh, I agree as well -- LLMs are not any kind of "one quick fix". But whenever most people discuss AI research, especially on ACX, it's usually in the context of the research being over; and all that's left to do is increase the number of parameters and the amount of training data -- until the LLM becomes ASI, which should happen any day now (though I suppose it could be doing its own research after that).

Andrew's avatar

> I actually know next to nothing about RWKV besides its existence (and lack of attention); does it really perform as well as a transformer ?

For LLMs? No, as far as I'm aware, but that could be because much more research has been put into improving the Transformer architecture. For spaced repetition stuff? Idk, I'm not the guy who made the original RWKV for the srs-benchmark repo, I'm working on making it more accurate and more efficient. I don't think the guy who made it ever said why he chose this architecture.

I feel like we're talking past each other, so let me be more explicit about my position on LLMs an AI:

1) If you are expecting something like "OpenAI, Anthropic, and all other companies that make LLMs go out of business, LLMs vanish and the entire field of machine learning rolls back to pre-Attention Is All You Need era", then no, I am absolutely NOT expecting that to happen.

If you are expecting something like "stock market prices of some companies drop", sure.

2) I do think the first ASI - AI that is better than any human at all cognitive tasks - will either be an LLM or will come from LLMs doing research.

Bugmaster's avatar

> f you are expecting something like "OpenAI, Anthropic, and all other companies that make LLMs go out of business...

No, that'd be silly. LLMs are genuinely useful tools, just like CMS applications are (were) genuinely useful tools. They will never vanish entirely (unless they are replaced by something radically new, and probably not even then). But I also don't think that LLMs are the ultimate peak of human achievement and the last discovery we will ever need to make.

> I do think the first ASI - AI that is better than any human at all cognitive tasks - will either be an LLM or will come from LLMs doing research.

The only way I could possibly agree with this statement is to rephrase "LLMs doing research" as "humans doing research using all kinds of tools, including LLMs". In this case I do expect ASI to be developed with the aid of LLMs (as well as other tools)... eventually. The laws of physics do not appear to prohibit this from happening, at least (setting aside for the moment any magical powers often attributed to ASI).

Andrew's avatar

Present-day LLMs can already contribute to their own development: https://x.com/OpenAI/status/2082577277246972300

And given that neither compute scaling nor algorithmic improvements have been exhausted + the fact that numbers on pretty much any benchmark you can think of continue to climb + the fact that LLMs are *also* really good at math now (https://vibemathed.com/), which obviously helps with coming up with new NN architectures, it's pretty clear that we'll have recursive self-improvement in the next 2-5 years.

Lucas's avatar

First, thank you for your work on FSRS!

Yeah it seems like the LSTM stuff is old news. I do check from time to time stuff like https://github.com/open-spaced-repetition/srs-benchmark, I think I remembered it as "official Anki stuff" rather than "people working around spaced repetition in general are experimenting with LSTM". There's this repo and also some discussion on the anki web forums ("Suggestion to Damien: Migrating to a NN (Neural Net)", especially the posts by Expertium.

Andrew's avatar

> especially the posts by Expertium

*star_wars_well_of_course_I_know_him_hes_me.png*

Out of all the algorithms in srs-benchmark, only a handful are actually used in real life. HLR is used by Duolingo, FSRS is used in Anki, Remnote and a few other obscure apps, and...I think that's it. Ebisu is used only by its own creator. I'm not aware of DASH and ACT-R algorithms making their way from academic papers into software. RWKV is used in an Anki fork which is only shared in the Anki Discord server and used by <10 people, I'm working on a more accurate AND more efficient version of it.

Oliver's avatar

I would be interested in seeing how well AI can predict the weather because there are supposed to be mathematical limits due to chaos theory. My guess is that they will be better than humans but still a distance from theoretical perfection.

Bugmaster's avatar

FWIW I've been using an app called "Windy" which aggregates several predictive weather models, including some ML-based ones. I suppose you could say they are "AI", although AFAIK they are not LLMs. Anyway, Windy's results are better than the general-purpose model that's used by e.g. weather.com, but not by much.

Bugmaster's avatar

I don't think I understand this part correctly:

> Maybe those statistical models are close to the best that it’s possible to do; the rest is what the mathematicians call aleatoric uncertainty - irreducible complexity downstream of chaotic systems that entirely resist modeling ... Consider a question like “Will AI take most human jobs by 2050?” ... a question like this one might be the opposite of a sports game, and have almost no aleatoric uncertainty...

How could it not ? Are you saying that the world economy, which depends on the actions of billions of people, not to mention chaotic environmental factors like the weather (which strongly affect many aspects of economic activity), is actually *simpler* than a (deliberately controlled) sports game ?

Kenny Easwaran's avatar

He is thinking that in the medium run of a couple decades, what matters for the economic transition will be the fundamentals. The chaos of individual human actions will wash out over time.

This is certainly plausible for some economic variables. But I don’t know how true this should actually be for something like employment.

Scott Alexander's avatar

Yes. To give an obvious example, will more or fewer people be wearing jackets in California six months from now compared to today? This requires predicting the actions of millions of people plus the weather, but the answer is obviously "more" because it will be winter.

I think in retrospect, we can say that any world that invented the car would have fewer horses fifty years later, or any world that invented the gun would have less feudal / more democratic governments. I don't think these things were obvious at the time, but I think they're pretty deterministic, and although you could imagine situations where they don't hold (maybe every government in the world bans cars and institutes a horse breeding program for purely sentimental reasons) they're more predictable than something relatively-unpredictable like a sports game.

MicaiahC's avatar

See also, gwerns comments about "impossibility proofs" of things from chaos theory or "rigorous bounds to intelligence" are often bait and switches.

https://www.lesswrong.com/posts/epgCXiv3Yy3qgcsys/you-can-t-predict-a-game-of-pinball?commentId=wjLFhiWWacByqyu6a

Excerpt:

As Von Neumann was commenting about weather 'forecasting', and was a standard trope in cybernetics, prediction and control are duals, and 'muh chaos' doesn't mean you can't predict chaotic systems [snip snide aside] - neural networks are good at predicting chaotic systems (https://journals.aps.org/prresearch/abstract/10.1103/PhysRevResearch.5.043252#fulltext), and if you can't, it just means that you need to control them upstream as well. This is why the field of "control of chaos"(https://en.wikipedia.org/wiki/Control_of_chaos) can exist. (Or, to put it more pragmatically: If you are worried about hurricanes causing damage, you "pour oil over troubled waters" in the right places, or you can try to predict them months out and disrupt the initial warm air formation; but it might be better to control them later by steering them, and if that turns out to be infeasible, control the damage by ensuring people aren't there to begin with by quietly tweaking tax & insurance policy decades in advance to avoid lavishly subsidizing coastal vacation homes, etc; there are many places in the system to exert control rather than go 'ah, you see, the weather was proven by Mandelbrot to be chaotic and unpredictable!' & throw your hands up.)

That's the beauty of unstable systems: what makes them hard to predict is also what makes them powerful to control. The sensitivity to infinitesimal detail is conserved: because they are so sensitive at the points in orbits where they transition between attractors, they must be extremely insensitive elsewhere. (If a butterfly flapping its wings can cause a hurricane, then that implies there must be another point where flapping wings could stop a hurricane...) So if you can observe where the crossings are in the phase-space, you can focus your control on avoiding going near the crossings. This nowhere requires infinitely precise measurements/predictions, even though it is true that you would if you wanted to stand by passively and try to predict it.

Bugmaster's avatar

This is a very astute observation, but I think it's also a bit tangential.

Imagine that I just rolled a d6, a 6-sided die in the shape of a cube. Its final position is subject to many chaotic interactions: air molecules rubbing against in flight, slightly elastic collisions with the surface of the table, etc. Let's say that the die came up a "3", and that you are a superforecaster. What is your prediction on what the result of the next roll will be ?

I would argue that knowing everything you do about the behaviour of chaotic systems doesn't help you one bit. Absent any other information, the answer is, "it could be any number from 1 to 6, with a 1/6 probability". Maybe you (being a superforecaster) could go out and collect some data on the most popular dice manufacturers and their manufacturing flaws, and eke out 0.01% somewhere, but that's the best you could do. There's no AI/ML/LLM/etc. system that could improve your chances.

The smart play would be to wait for me to throw the die a few more times (ideally, lots more times); but once I do that, the results are once again obvious, and it doesn't take a superforecaster to predict the next roll: all you have to do is count.

Thus the only way a superforecaster can shine is when there's *some* available data, but not a lot. Maybe I rolled the die 6 times and got {2, 6, 4, 3, 2, 1}. Does this help you predict the next result ? I don't think so; at least, not enough to outperform a regular forecaster by any significant margin.

And that was just a single die. Real-world events often have many more than just 6 possible outcomes. The best strategy for outperforming the market is to collect more data; specifically, more than anyone else can collect in time. This is feasible when you are predicting things like political events or financial outcomes (assuming you have insider knowledge that others lack); it is not feasible when you are trying to predict outcomes of hitherto unprecedented situations -- though I could be wrong.

Bugmaster's avatar

Just for fun, I asked our resident "superintelligent AI", i.e. Claude Fable about the die-rolling scenario. Sadly, it was not able to outperform myself on this topic:

https://claude.ai/share/7a910da9-3c5c-4936-9646-31a78b21292a

"And a sobering endnote: even after 40,000 rolls of the cheap die, your prediction accuracy on the next roll would rise from 16.7% to about 17.7%. You'd be "better than chance" in a way that's statistically defensible and practically almost worthless."

Yug Gnirob's avatar

With a single roll of 3, the prediction should be "3 again"; you know the die can hit that number, so the odds of it being weighted have shrunk to "weighted to roll 3's". Any impurities are more likely to bias the die toward 3 and away from 4.

In the second example, you've gotten more 2's than other results, and no 5's, meaning the die is perhaps biased toward the 2 side of things.

Bugmaster's avatar

If you are a space alien who knows nothing about dice, this is correct. However, if you use this strategy to bet on prediction markets, you would lose, because as humans we do know that plastic gaming dice manufactured today are very close to being fair. Thus the fact that e.g. a die landed on a "3" exactly 1/1 times tells you virtually nothing about its next roll.

Jeffrey Soreff's avatar

>they're more predictable than something relatively-unpredictable like a sports game.

Yes. More generally, I expect that there is probably a lot of variation from field to field on how close human superforecasters today come to the ideal, noise-limited, predictions. I suspect that AIs will outperform humans where there are _many_ heuristics, each of which improves the prediction slightly (not dominated by just a few), and AIs can grind through applying _all_ of them without losing patience or losing track of which had already been applied.

Domo Sapiens's avatar

Yes, that's always been the use case for ML/AI: Where you have messy, multivariate and noise-covered data. Whether thats a spectroscopic sample of biological fluids in a lab or a noisy news-environment in a messy human-social-political world, they seem to have some shared fundamental properties in this sense.

Jeffrey Soreff's avatar

Agreed, Many Thanks!

Bugmaster's avatar

I don't think the clothing analogy holds, since we have so much data collected on seasonal temperature changes that winter chills are virtually certain. Thus, there'd be no room for any superforecaster or AI or whomever to make money on the prediction market; everyone would just bet on "yes, more people will wear jackets".

As for the cars vs. horses or guns vs. feudalism, hindsight is 20/20. I would argue that when these tools were invented, it was not at all clear which way the balance would tip (e.g. in China and Japan feudal governments made regular use of guns). I am not at all convinced that even a superforecaster would be able to witness the first bullet (or the first steam engine) being cast, and then confidently say, "ha, it's obvious, no more horses or kings in 20 years". Granted, he would be able to observe the trends, e.g. market share of horses vs. cars falling year after year; and perhaps he would outbid the market on his predictions -- but again, it's not at all obvious to me that he'd be able to do so by any significant margin (absent some kind of a crystal ball or time machine or other magic).

Aeolian321's avatar

Modern feudalism will be very much improved by modern cars (aka electric cars).

Simone's avatar

The difference is that the games are artificially manufactured to be fair and thus uncertain, while reality has no such obligation to us and can absolutely present us with nigh inescapable scenarios.

meanderingmoose's avatar

Regarding the chess ceiling and footnote 11 - I think that’s a reasonable range from Fable, with 3-4 the most likely range. A 5-pawn advantage is about equivalent to the stronger player playing without a rook, and in that scenario the nature of the game changes significantly. It goes from a positional battle to one in which the disadvantaged player is forced to rely on tactics and avoid piece trades (as every trade represents a dangerous simplification). If we assume that chess grandmasters would still have the first ~10 (book) moves memorized that doesn’t leave room for the type of tactics required for the disadvantaged player to get back in the game, even with perfect play. Though it could be a more interesting game if the grandmaster wasn’t allowed to prep for the rookless board (in the spirit of chess960).

abilash suresh's avatar

These pawn-advantages also heavily depend on the time control, we only have reliable data from grandmasters competing against engines in bullet or blitz. Hikaru played against Leela in 5-0 with rook-odds and lost almost every game (https://www.youtube.com/watch?v=m7N4qC1znDc). Hikaru lost a couple of games to Leela with queen odds in bullet as well (https://www.youtube.com/watch?v=m7N4qC1znDc).

meanderingmoose's avatar

That’s a great point - I was thinking about classical, but it makes sense to me that in bullet the gap is / can be much larger (perfect play is the same in any time format, but humans are less accurate with less time).

Alvin Ånestrand's avatar

I'm excited about AI forecasting in domains with low aleatoric uncertainty that are so difficult that even superforecasters can't get close to optimal predictions. But while ultraforecaster AIs could have even better gains there, current AIs are probably further behind the human frontier in those domains.

David J Higgs's avatar

Why do you think they're behind the human frontier? "So difficult" implies a lower amount of human tractability and therefore training data, but then again there's also Moravec's Paradox to consider: (current frontier) AI could just as easily be *relatively* better at those harder sub-domains (despite still being subpar compared to humans or at least not noticeably better yet)

Alvin Ånestrand's avatar

I've done a lot of forecasting and have experimented with letting AIs do it for me, and often feel like they're really good at straightforward questions but lack the judgment for more complex issues, like trying to predict things several years into the future. (They fail to consider everything relevant to the predictions and end up with incoherent forecasts, if you don't point things out for them.)

They're really good at doing background research though.

David J Higgs's avatar

Thank you for the information, that makes sense to me. That sounds like it mostly matches my experience doing some light, recreational forecasting research w/ Claude on a few Manifold Markets and AI timelines/graphs type of questions. Much more tenacious and thorough than me in pulling from lots of different sources and noticing a bunch of details, but neglecting some larger and/or strategic considerations that become relevant over longer time periods.

Daniel's avatar

>"Probably all of this pales into insignificance compared to the gains we could get by switching from our current strategy of making decisions based on vibes and ballroom-related-bribery to listening to markets and forecasters at all"

Yes, but we could have done this years (really decades) ago even without AI. The reality is that the people in charge (up to and including the voters themselves) love vibes and ballroom-related bribery, and hate markets and forecasters.

Have you read the comments on any mainstream news article on prediction markets? People want heads on pikes. There would be riots if the government explicitly delegated decision-making power to a prediction market.

Scott Alexander's avatar

Yes, I agree, although I think it will turn out more palatable to delegate decisions to prediction-market caliber AIs than to either prediction markets themselves or equal-quality human forecasters, for reasons I mention at https://www.astralcodexten.com/i/202397135/living-in-the-world-where-bots-approximately-equal-top-humans

Max Marty's avatar

What’s to stop those “people in charge” from privately consulting the superforecaster AI behind the scenes and then publicly talking all about the vibes sans any mention of the forecast, removing the real provenance of their beliefs from public view?

Maybe some are already doing this and we just don’t know about it.

David J Higgs's avatar

Oh I'm sure that some of that is or at least will be going on, but rationalizing an optimal decision in vibes language =/= picking a decision to optimize for vibes. There have been no shortage of efforts to make immigration (or a million other things) sound palatable or fit the right vibes or whatever, but all of that pales before decisions actually based around xenophobia (or whatever drives so much anti-immigration sentiment).

Edit: and ofc decision makers competing for power/public opinion can always ask AIs to help them maximize the right vibes of their decisions that were also selected by AI to maximize fitness to vibes

Dan Schwarz's avatar

On the topic of headroom above teams of superforecasters (40 or 50 / 100 in this system, I think more like 40 as per footnote 9):

There's an old LW post on Metaculus data that didn't see much correlation between accuracy and time-horizon of questions: https://www.lesswrong.com/posts/MquvZCGWyYinsN49c/range-and-forecasting-accuracy

I think this is evidence that aleatory uncertainty is far above the community. If it were closer, I think we'd see a closer relationship between accuracy and time-horizon, at least beyond short horizons (where sometimes more time reduces variance), e.g. 2 year questions should be less accurate than 1 year questions.

Maybe a Samotsvety dataset would show this relationship. I don't think there are enough questions to ever get that evidence, but if that was the case, I would update that teams of superforecasters are closer to irreducible complexity of the world.

Anyway, taking a guess, and acknowledging I'm talking the FutureSearch book here, I predict that the best AI forecaster in mid-2028 will score 60, e.g. Anchor 2. I further predict this will be very impactful, and that accuracy will continue to improve beyond that point.

Hidden Agent's avatar

questions at different horizons are not randomly chosen though. Do you think that matters?

100YoS's avatar

In which I once again link to the Ringer story on popcorn prediction markets: https://www.theringer.com/2018/11/15/movies/box-office-futures-dodd-frank-mpaa-recession

abilash suresh's avatar

Minor addendum, knight odds at classical have been tested once. There was an 8 round classical match between GM Joel Benjamin (2473 ELO at the time) and Leela with knight odds, the match ended 4.5–3.5 in favor of Joel.

Side note, these grandmasters are not playing pure chess engines and are instead playing Leela bots specially designed to play odds. They're trained against a model of aggregated human play, so they'll deliberately play objectively worse moves that avoid trades and keep the position complicated.

David J Higgs's avatar

Interesting. I wonder if chess AI trained to adversarially beat humans come closer to approximating perfect SAI chess play than simple "best chess" AI, since SAI would also exploit psychological flaws in its opponents.

Pete's avatar

What's interesting to me is at which level of odds chess becomes "practically solvable" - i.e. from game theory we know that chess must be solvable, there must either exist a strategy that allows one of the players to win or force a draw no matter what the opponent does, but it's impractical for us to find it or even determine if it's always winnable or always drawable.

The same also applies for every option of odds-chess. So where on the scale of odds (from zero odds to one player has just the king) is the boundary where a strong human player can force a win even against a theoretically perfect play and even an infinitely smart opponent can't get a draw? As you show, knight odds aren't enough, and trivially "everything but king" odds are sufficient..

Veedrac's avatar

It's worth considering that this match was using a notably worse version of Leela Knight Odds, and that even today there's a lot of low-hanging fruit in this domain.

Melvin's avatar

> we’re not talking about how practical it is to create superforecaster AIs, we’re talking about what chess can teach us about the degree to which top humans approach optimal play

Can chess tell us anything about the general case here? It seems very game-dependent.

In tic-tac-toe, humans can easily reach optimal play.

On the other hand you could imagine some very complicated games where humans would be even further from optimal play than they are in chess.

David J Higgs's avatar

You don't have to imagine them, they exist: Go is an example (if I understand correctly), and I think many modern competitive eSports like Starcraft 2 or DotA/League of Legends are even more complex/difficult examples (solo or team based respectively).

Melvin's avatar

Yes I've heard that Go is in some sense more complex, but I've never played it myself so I don't want to speculate too hard.

You can imagine arbitrarily complex games though, you can imagine a game where the rulebook is forty thousand pages long and it maxes out human brainpower just trying to figure out what's a legal move, let alone an optimal one.

David J Higgs's avatar

This is true, and is I think similar to the reason why so many rat-adjacent folks have such strong intuitions about superintelligence: the real world is kind of like an arbitrarily complex game with arcane rules that have strained even genius humans trying to figure out what's a legal move let alone an optimal one!

Yug Gnirob's avatar

Go has extremely simple rules; every piece is the same, you can put one anywhere you want any turn (or nowhere), and if you completely surround a connected block of your opponent's pieces, however many that may be, the surrounded pieces are removed from the board. Most pieces wins, game ends when the players agree who will have the most pieces.

It's very hard to properly wrap one's head around the implications of that on a 19x19 board. You want to make your block as wide as possible, but quickly enough that it can't get surrounded. If you put two holes in your block, it can never be completely surrounded, but that takes time you could be spending widening your area.

Ch Hi's avatar

The real trick to imaging those games, though, is to increase the number of dimensions. Just imagine trying to even define the rules for an actual 4-d chess. Then imagine trying to play it. AIs could handle 6-d chess. (I can't say how well they could handle it, but they can prove theorems by mapping them onto higher dimensional spaces, and then reducing them to ordinary space. Theorems that no mathematician had been able to prove.)

Stephen Saperstein Frug's avatar

Rather than speculate about this, someone should ask the best superforcasters (both human & AI) for *their* percentages about whether scenario A or B is more likely

Tom Liptay's avatar

GJP Superforecaster, ex-Metaculus, and Futuresearch staff here.

First, I really love this post! I've pondered how to estimate the irreducible uncertainty limit for probably a decade. The short answer is that I think being precise is tricky because it depends on a lot of considerations (most mentioned in the piece or the comments already).

I think it is worth drawing an explicit distinction between absolute & relative accuracy.

A Brier score in isolation is what I would call an absolute accuracy score. However, as noted in the piece, it is generally not especially interesting because it is a function of the question difficulty. I think of the Baseline score on Metaculus as an absolute score.

What I find more interesting is the relative accuracy between forecasters, along with a confidence interval to tell me if the difference is significant. If we asked 1000 questions where the true probability is 1% then a perfect forecaster would forecast 1%, while a biased forecaster might forecast 2%. Even if we ask enough questions to have a tight confidence interval, the effect size will be quite small (Brier difference of 0.0001 in expectation), even though the biased forecaster is off by a factor of 2x in odds space on every question. So, relative accuracy differences are also a function of which questions are being asked.

Let's assume we are somehow able to keep the question difficulty fixed over time.

At Futuresearch, we've seen roughly straight line improvement in absolute Brier score over time. I expected it to start plateauing, but so far that has not happened. Despite being wrong so far, I still expect the absolute scores to start flattening as we approach a irreducible uncertainty limit.

Regarding relative scores (like Metaculus's Peer score), I expect that to plateau even sooner since Pro forecasters will increasingly adopt AI forecaster assistants. Since the Peer score measures the relative accuracy, it seems difficult to imagine a big gap in the head-to-head Peer score between Pros and AI ultraforecasters. The only way to get a big difference is if Pros do not use AI, or make big and wrong corrections to what the AI says.

In very rough terms, scenario A seems more likely if we're talking about relative/Peer scores. Scenario B seems more representative of absolute scores (although I'd expect the Pro line to bend upwards and not remain flat since they'll use AI too).

Peter Gerdes's avatar

I fear that vague claims like: what if AI were as much better than the smartest human at X as they are over the average human tend to lead people to make really bad generalizations and are a big part of what causes unjustified exaggeration of AI x-risk.

The problem is that it really matters how you measure those figures but even though you mention was to measure most people are just going to imagine being as to AI as they were to a great chess player when they were 5. But an AI doesn't have to be to you as you are to a child to have a crazy higher ELO score -- that score could mostly he about reliability.

Similarly, I fear that when people hear about AI doing much better than people at prediction they imagine it being able to know exactly how really complicated situations play out rather than just reducing the variance in important ways on the same range of situations. I mean nothing here allows an AI to get around the fact that lots of predictions are essentially NP or exp complete problems (or worse).

---

I don't necessarily mean this as a criticism of you since it's not that clear if there is anything else you should do without directly making this statement which you may not believe but I felt it was worth adding.

Nicholas Halden's avatar

I'm willing to bet money on something like, in the next three years, there won't be clear evidence of AI making market forecasts at a superhuman level.

I think the mechanism of superforecasting is tough. Intelligence isn't a scalar and you might well get a situation where the AI is incredible at certain things (coding, math, cybersecurity, protein folding) but terrible at others (qualitative forecasting).

David J Higgs's avatar

Sure, all the things you said are true and you *might* well get such a situation. But why should I believe we probably will get that situation instead of believing that the straight line will continue on the (non-linear) graph?

Bugmaster's avatar

Eh, it all depends on the framing. If LLMs are better than humans on average by 0.1%, does that count as "superhuman" ?

Nicholas Halden's avatar

No. The best LLMs have to be significantly better than the best humans.

Bugmaster's avatar

I would agree with you there, but I would also bet that if LLMs did manage to out-predict humans by 0.1% or even 1%, we'd be seeing a lot of hype about ASI-based superforecasting, the end of human knowledge, etc... :-(

Nicholas Halden's avatar

I’ll believe humans are worse at forecasting when the top hedge funds use ai agents for trade ideas the way SWEs currently use ai for coding.

Bugmaster's avatar

That is not exactly a fair comparison (though I agree with your overall point). From what I understand, LLMs are still worse at coding than humans; however, they are able to produce a lot more code a lot faster. So faced with a choice between a perfectly crafted program that's ready in a month, and some sloppy mess that is ready tomorrow, most people pick the 2nd option. I don't think the same tradeoffs apply to prediction markets (at least, not as strongly).

Byoungkwon Kim's avatar

Does anyone here have experience with the actual reasoning traces from SOTA prediction AIs? Do they just do normal forecaster stuff but faster and broader? Or are there some hints of superhuman capabilities?

David J Higgs's avatar

That seems like an interesting question to be asking, though more useful in a few months/a year-ish when we'd expect to see those hints for sure if the straight line on the graph wasn't bending

Byoungkwon Kim's avatar

Agreed. I just tried FutureSearch and the reasoning (at least its summary of it) was pretty mundane, so I'm guessing we're not at that stage yet.

Pycea's avatar

Could you point to the math that calculates the market movements? Trying to run it myself I get numbers that are close but sightly different. Maybe the result of the perfect score used for the calculation being 125?

Scott Alexander's avatar

No because that's one of the things I used AI for. What numbers do you get?

Pycea's avatar

I'd be interested in seeing the agent's reasoning. After looking a bit further, I actually think there may be a couple problems here. (The numbers I got from my first test are junk and should be ignored.)

First, the way FutureEval is calculated means that you can't necessarily just take a score of 30 compared to 20 to be the same level of improvement as 20 is to 10. (This is due to their ridge regression step.) But even if you make that assumption, the scores are scaled to have the same range as Peer Scores, which means you can't go backwards from there to raw probabilities without knowing the scaling factor.

Trying to reconstruct the numbers, it seems like it may have just taken the FutureEval as peer scores and gone from there: p_max_ai = .5*e^(score_diff/200). Though there's a factor of 2 in there whose origin is unclear, the 200 should be 100 if dealing with peer scores directly. It's also unclear where the perfect score of 100 came from, since that implies a given probability of only 67%.

It's possible there's good reason for these assumptions, maybe it decided to set the FutureEval scaling factor to be 2, but a proper test would involve getting raw Metaculus data which is a little trickier.

Eremolalos's avatar

Scott seems to think there's "some fundamental limit on the predictability of world events," but to believe it's likely that neither human superforecasters nor AI has hit it yet. But is there a fundamental limit, other an 100% accuracy? It seems possible that in the far future, some extraordinary technology would allow for a gigantic gain, even a gain to 100%, in forecasting accuracy about such things. For instance, being able to travel forward in time would do it. Or maybe it will be possible to access and observe another version of our universe that at the point where the event we want to predict occurs is identical to ours, because it hasn't branched off in another direction yet. Or maybe it will be possible to build and watch a perfect replica of the thing we want to predict. Or . . .

Ch Hi's avatar

To hit 100% you'd need to overthrow quantum mechanics.

While I agree that quantum mechanics probably isn't correct, since it and relativity conflict in their predictions, I suspect that whatever extends or replaces it will keep the uncertainty.

Ryan P's avatar

I think Scott is assuming we don't completely overturn known physics, and won't have the compute power to build functionally identical worlds.

Where things might get interesting is if prediction markets start becoming self-fulfilling prophecies: for instance, the market predicts war between two countries, so both countries try to make a first strike to gain the advantage. The market predicts a peace deal being made, so both sides walk into the room with the expectation that they will make concessions (and the other side will too). A candidate for office is predicted to win, so no one bothers to fund the opposition....

Mike B's avatar

Unlike heightball, fastball, where all the athletes line up side by side and whoever is fastest wins, is surprisingly compelling over many of its rule variants.

DrMcleod's avatar

Also, prettyball, in which the competitors line up side by side and a small committee decides which one is the prettiest has an enduring appeal, despite being politically out of favour in the current year.

Melvin's avatar

Not to be confused with the competition to determine which dog looks most like a dog https://www.youtube.com/watch?v=Ch0RwCtNV1o

Xpym's avatar

"And the napkin math was increasingly reliant on AI inputs. There’s a slight risk this whole post is downstream of AI psychosis; I asked some smart people to proofread it, but probably they just asked their own AIs."

O brave new world etc etc

Lucas's avatar

Are there any domain where AI has plateau'd at human level because of the reliance on training data? For code it seems like reasoning models + lots of RLs unlocked superhuman coding ability, at least one shot superhuman coding. But I'm not sure where the data comes from for RL, I think some is handcrafted scenarios/problems by humans (what many SWEs at Meta have been doing lately, what a few companies sell too), but I don't know how far you can go without that

David Manheim's avatar

This has a lot to do with what I called the "Aleatory Baseline" ( https://x.com/davidmanheim/status/1349349979652034560 ) in a 2019 presentation at the Global Priorities Institute, (Since unfortunately removed - https://x.com/davidmanheim/status/1286345459636801537 .) This is much more focused on prediction skill over time, and whether accuracy falls off in longer time horizons, which I think is part of the critical question that isn't really addressed here. (As obviously distinct from the question here about how quickly forecasting skill increases.) I also discussed the fact that many prediction domains were unstable, that is, the predictions themselves will change behaviors in a way that invalidate them, or stable, in that the predictions are self-fulfilling if believed. This seems even more important with AI superforecasting.

I've discussed all of this since, but the paper I had planned on the issue was interrupted in 2020 by more urgent issues, and I never got back to it - it would be very interesting to do so now given that we have much more data.

DrMcleod's avatar

"ultraforecasting"? I think not. In the natural progression of superlatives, the next entry is "superduperforecasting".

Mikk14's avatar

Oh me oh my, I am finally in the position of making a shameless plug that is not even that OT. I actually wrote a paper on sport predictability: https://link.springer.com/article/10.1140/epjds/s13688-024-00448-3

The draft system does have the effect of increase unpredictability vs systems that do not have it (I know, there exist sports played outside the US, shocking), but it really depends on how and why a measure is introduced.

For instance, French rugby does have a salary cap, but it is still the most predictable sport league in the world. That is because the salary cap is not there to level the playing field, it is there to prevent clubs to bankrupt themselves.

Hoopdawg's avatar

I just fundamentally don't understand (and to the exent that I understand, I reject) the entire framing.

We predict future events by constructing models of how the world works. The people on prediction markets who outdo simple toy models on sports / box office performance do not do it by some arcane insights, they do it by utilizing better models with more information. Sports have advanced statistical models utilizing form, individual players' skill and performance, etc. Box office predictions use, e.g., advance ticket sales. They're actually impressively good at what they do, and the fact that it can only slightly improve the accuracy over simple models is a point towards these things being, well, I won't say unpredictable (everything is predictable once you build a 1:1 model of universe), but certainly difficult to predict in a "there's a hard limit to the approach" sense.

(Importantly, not all models are statistical. In fact, the best models we have aren't statistical. Contrast chess. Machine learning models of chess operate start with perfect information of how the "world" of chessboard works on an "atomic" level, all that's left to improve prediction/manipulation of the game outcome is more / better calculation. But also contrast engineering, weather prediction, for a various degree where the modeling goes deeper than bare statistics and into mechanical knowledge of how the world works.)

Can statistical models of [essentially all data humanity generates] have advantages over [humans essentially guessing]? Sure. How much of an advantage is an open question (I model LLMs as "average of what humans say", you ostensibly model them as "super-intelligence", this points to vastly different outcomes). But the thing is, it's fundamentally not the question we should be concerning ourselves about! If AI generates a better model of the world, great, we should be looking at what its model is, distilling and formalizing it, so it can augment human knowledge. Treating them (or "superforcasting" in general) as the end-point, a one-size-fits-all solution, black box that generates superhuman answers, is fundamentally self-defeating.

(Of course none of what I'm saying does not contradict the literal point you're making. Of course. I just think what your reasoning omits is... kind of vastly more important, you know?)

Scott Alexander's avatar

I disagree - I think we do things both through formal models and through informal incommunicable models.

Why am I not as good a forecaster as Nate Silver? Can't I just use his models? Sometimes I do (in the sense of checking his website to see who is likely to win elections). But sometimes Nate makes off-the-cuff forecasts which are better than my off-the-cuff forecasts, and either he can't teach me how to do that, or it would take years of being his apprentice for me to learn. And even beyond that, why is Nate better at making models than I am? Doesn't he have some meta-model of making models, and shouldn't he be able to communicate that so I can make models as well as he does?

I think this all bottoms out in things requiring both formal/communicable and informal/incommunicable skills, and forecasting involves many of the latter.

I don't think it makes sense to ask an AI forecaster for "its model" any more than it makes sense to ask Mozart for "his model" and expect that anyone can generate brilliant music on demand.

Amaury LORIN's avatar

This reads like wishful fiction to me. All of that requires the unstated caveat of "assuming AI doesn't make prediction markets moot before then", which is quite a large assumption!

YesNoMaybe's avatar

One thing I like about the current state of AI discourse is, that you can never tell what follows after "This reads like wishful fiction to me"!

In this case I expected the AI sceptic position, but ofc your point is just as fitting a follow up!

Scott Alexander's avatar

I think there will be AI that does very well on prediction markets before there's AI that makes prediction markets moot. I think the gap will be more like a year than a month, though I'm not confident, and a year is time we can work with.

Amaury LORIN's avatar

I'd appreciate elaboration on this point: What do you expect AI ultraforecasters to accomplish in the year between surpassing human superforecasters and making the economy irrelevant/destroying the world for human society?

It seems way too fast to turn that into e.g. better governmental policies; we don't even need ultraforecasters for that. It seems to me that the main obstacle to making prediction markets useful is integrating them into decision-making, rather than making them more accurate on extremely hard to predict events.

Ultraforecasting could be useful for individual actors with the potential to make good use of it which I imagine is either:

- Consultants like Samotsvety called upon to predict something, where by construction they are given the power to turn prediction into action. I'm not sure what to think of that case.

- Idealists like billionaires or OpenAI, personally interested in prediction due to having the power to take radical action if convinced by a prediction. I don't see how giving them better decision-making could possibly improve the world rather than just exacerbate concentration of power in their favor.

Nicolò Bagarin - 404_NOT_FOUND's avatar

- Plotting the best humans' forecasting accuracy as a line that remains flat with time is wrong, unless you expect humans not to improve with access to better tools, which would be absurd

- Metaculus questions on FutureEval are not mostly about geopolitics

- The FutureSearch chart assumes that in the current season the Pros' average score is identical to the ones of past seasons, but you just cannot make that comparison because the questions are different (and that information is not publicly available). Also, is "best unattributed bot" a mysterious, secret AI forecasting company, or just an individual hobbyist project that can apparently compete with well-founded companies?

DanielLC's avatar

How will we know which scenario we're in? Practically speaking, there's two possibilities:

1. The top superforcasters contain a mix of AI and humans.

2. The top superforcasters contain a mix of AI and AI pretending to be humans. And humans consulting AI and not giving it the credit it deserves.

Alastair Williams's avatar

I am continually disappointed not to see a single reference to either Hari Seldon or Psychohistory in these posts.

Daniel Reeves's avatar

I'm excited about this debate! I'm not sure how far apart we actually are at this point. Originally I characterized your position as AI being on track to significantly beat human superforecasters. I said I expected AI to get slightly rather than significantly superhuman. I like your conversion to percentage points of difference. Your 4-12pp range is non-crazy. I guess I'll stick to my guns and predict 0-4pp above top superforecasters over the next 2 years. So nonzero overlap. If AI hits 4 on the button then we were both right!

Regarding "No wonder the relative improvement number [for predicting box office revenue] comes out looking slightly anemic":

Our point at the time was that this data (number of movie screens, Google search volume) is just sitting there for anyone to see. Dump it into a dirt-simple regression and you have a prediction that's probably not far from optimal. But point taken that it's not clear there's a similar trick for things like geopolitics. Here's our defense from the second paragraph of section 5 of the paper:

~~

Although reasonable [to point out that movies may be easier than political outcomes, or that statistical models require regularity and consistency that may not apply in other domains], these doubts should be weighed against recent empirical evidence in political and policy analysis. [...] In predicting election winners, the Iowa Electronic Markets were outperformed both by statistically corrected polls and by a model based on single-issue voting preferences. [...] One study of expert political predictions found that statistical models outperformed not only individual experts, but also compared favorably with aggregate forecasts. Presumably, properly designed election markets would in time adjust to incorporate predictions from these alternatives. Thus markets in the long run may still regain their performance advantages as suggested by theory. Nevertheless, these findings are consistent with our claim that market and non-market forecasting techniques are often comparable.

~~

Steve Sailer's avatar

I got interested in Philip Tetlock's forecasting work around 2014 or so. I thought about trying out for his superforecasting competition, but then it became obvious that it would be a huge amount of work to constantly update forecasts for 52 weeks on boring subjects like the China-Philippines dispute over the Spratley Islands.

So, I could well see AI getting better than humans at Tetlock-style competitions where you can repeatedly update since AIs don't get bored.

But then it turned out in 2015 that the world-historical event of that year that has transformed European and, to a certain extent, American politics ever since -- German Chancellor Merkel's off-hand decision to let in a million marching Muslim men -- wasn't part of Tetlock's competition. Nobody, it appears, was thinking about it.

The really big events come out of the blue like that.

Will AI be better than humans at predicting the unpredictable?

Beats me.

Steve Sailer's avatar

A lot of these issues regarding forecasting were kicked around in depth in Finance Economics in the 1970s and 1980s following the 1970 publication of Eugene Fama's paper on the Efficient Market Hypothesis. Giant fortunes have been made by getting ever so slightly better at forecasting financial markets, so the intensity of thought devoted to that subject was significantly greater for decades than that devoted to predicting political and other non-financial events.

Scott Alexander's avatar

What was the result? Presumably people are still improving, but is there some way of measuring humanity's collective financial acumen over time?

Steve Sailer's avatar

Fifty years ago, a lot of people put money in mutual funds with an 8 percent fee ("load") for which you got the benefit of their stock picking geniuses.

Now, a lot more money goes into index funds that don't try to beat the market, like Vanguard. They just try to have really low expense ratios.

On the other hand, lots of pros invest in finding and profiting from ever tinier or more subtle market anomalies.

Of course, if everybody gave up trying to beat the market and just put their money in Vanguard, then the market wouldn't be efficient anymore, now would it?

Presumably, that won't happen.

Anyway, the point is that financial markets front-ran what prediction markets are now trying to do decades ago, so there is much to learn from them.

Frikgeek's avatar

I don't think Chess is a great example(because the #1 chess Engine is neurosymbolic, human-written heuristic alpha/beta pruning used for search combined with a small and efficient neural network used for evaluation) but if you do want to use it(or only focus on engines like Leela that are much more reliant on Machine Learning and use it for search along larger and more complicated evaluation) then Elo or even game result might not be the best measure.

So let's disregard that and talk about chess engines and top human play.

There are 2 main statistics used to evaluate human quality of play. Accuracy percentage - how many of a player's moves matched the top engine choice and when they deviated how much evaluation was lost due to the deviation. Top grandmasters regularly reach accuracy scores above 99% in classical time controls. ACPL(Average CentiPawn loss) measures how many "centipawns" a player is losing on average due to suboptimal play. 100 centipawns equals 1 pawn.

Top grandmasters regularly score under 10(meaning it would take 10 turns of play against an engine for them to lose the equivalent of a pawn, this loss can be entirely positional and doesn't need to actually result in losing a pawn of material). A top grandmaster on a good day, with plenty of rest, and playing near their peak might score 7 or lower.

When Gary Kasparov lost against Deep Blue in 1997(after beating it in 1996) it was organised as a serious match. Classical time control, 1 game per day, Kasparov had a full team to help with preparation. Modern top Grandmasters no longer play engines like this, even if the engine is a pawn or 2 pawns down. Therefore just using results makes it harder to gauge how far ahead current engines are since they're mostly playing "ordinary" GMs and not the absolute best(and the absolute best GMs can already beat many "normal" GMs a pawn down) and they're also not playing in conditions optimal for the GMs, like how a top GM would play another top GM in a serious tournament.

Looking at the accuracy scores and ACPL of top grandmasters I think they could still play even top modern engines to a draw if given a 2 pawn advantage in classical time controls. Including opening moves prepared in advance and endgames that can be purely theoretical the grandmaster would only have to play ~20 moves or fewer by themselves. That's just not enough for the engine to both overcome a 2 pawn handicap and then secure a winning advantage on top. Once the game has been "simplified"(most of the pieces and pawns have been traded) playing for a draw becomes rather simple.

In 2016 Hikaru Nakamura(one of the top Grandmasters) played Komodo Dragon NNUE(one of the top engines, using neural network evaluation) in a 4-game match. Every game had different odds. First game was pawn + move(Hikaru goes first and Komodo loses a pawn), 2nd game was pawn only(Komodo goes first), third game was exchange odds(Komodo goes first but loses a rook while Hikaru loses a knight. in standard chess scoring this is equal to 2 pawns), and the 4th game both had equal material but Hikaru started with both e and d pawns pushed and his kingside knight developed, essentially being given 4 moves(3 from odds, 1 from being white) before Komodo got to respond.

Hikaru drew the first 3 games and lost the 4th one. Komodo played a very closed structure which eventually neutralised Hikaru's starting move advantage, eventually turning it into a no-odds game where Hikaru had no chance. But that still means that a top GM managed to DRAW, not lose, against a top engine with odds lesser than 2 pawns.

Hikaru played Komodo again in 2020 and only drew 3 games out of 8 matches but these games were played with faster time controls which generally massively benefit engines compared to humans.

Not really related to the rest of my comment but also interesting.

For positions with 7 total pieces(this includes pawns) or less we do actually know what perfect play looks like, as those positions have been brute forced by computers and we're currently working on an 8 piece tablebase.

And while our current top chess engines can match the perfect solution in most endgames, that still leaves millions of positions where they can't.

As for why it would be so difficult to create a chess engine that can beat top human players down a queen it's because humans are not bad enough at chess. A good human player simply won't get pressured up a queen which means they can start trading pieces and then sacrifice their queen for a rook as soon as possible, then sacrifice their rook for a minor piece until they're in an endgame up a full piece. As long as some pawns remain on the board this endgame is essentially unwinnable for the side that's a piece down and even intermediate players could convert it against an engine with access to the previously mentioned perfect play tablebase.

A queen is just too big an advantage, you can sacrifice an exchange TWICE and still be up a full piece, this lets you simplify the game very quickly and the engine's choices are far too limited if it tries to stop you from making any trades or even exchange sacrifices.

Will Newsome's avatar

I'd wager it's more than two pawns advantage if the engine is trained a la LeelaKnightOdds to have high Contempt. Iirc LeelaKnightOdds mauled Hikaru in blitz, I don't think he'd win the classical match.

Mark Roulo's avatar

This was very interesting!

demost_'s avatar

I am somewhat confused by the 99% accuracy. That would mean that in games of 50 moves, the top grandmaster should make 0.5 mistakes on average (in the sense of not playing the top engine move). So in at least 50% of such games, they must literally play perfectly (= 100% aligned with the engine). Even more than 50% if the "mistakes" are clustered within games. But if they can play perfectly aligned games, why don't they draw against an engine?

Is it because the accuracy is somehow distorted? Like, many moves are concentrated in openings and end games, and grandmaster play perfectly there? But how can an opening even be accurate if a grandmaster can play hundreds of variants? Only one of them can be the best prediction, and all other should be "mistakes" in the sense that they are not the top choice of the engine.

zahmahkibo's avatar

IIUC it's a little fuzzier than that. something like your average loss in expected win% vs. the engine's moves, scaled 0 to 100. https://lichess.org/page/accuracy

demost_'s avatar

Ah, thanks, that explains it!

It's actually equivalent to the centipawn loss, just tranformed to a scale from 0 to 100. (Or actually, from -3 to 100, but whatever, there are only so many centipawns you can lose.)

Frikgeek's avatar

There are many middlegames that are inherently drawish so the engine rates 5 different moves as being 0.00. They all lead to a draw so they're all the top move.

When white isn't pushing for a win the position often develops into something like this and the engine evaluation ends up being extremely favorable. It's not uncommon for 40 move draws to end up with 99% accuracy and 3-5 ACPL for both players.

luke's avatar

"Hikaru played Komodo again in 2020 and only drew 3 games out of 8 matches but these games were played with faster time controls which generally massively benefit engines compared to humans"

I dont think this is true? These are the two examples that cause me to think differently

https://www.youtube.com/watch?v=BabwXdxdLB0

https://www.youtube.com/watch?v=9ObxTiA4wG8

luke's avatar

Also, just adding a link of Hikaru playing stockfish recently

https://www.youtube.com/watch?v=5JSHLEqDCYQ

Frikgeek's avatar

Stockfish level 8 on lichess is far from the best Stockfish can do. Its estimated Elo is only around 2600, so some super GMs can beat it.

luke's avatar

https://computerchess.org.uk/402.archive/cgi/engine_details.cgi?each_game=0&eng=Stockfish+8+64-bit&match_length=30&print=Details&utm_source=chatgpt.com

Now maybe im missing something here, I am not in expert in computer chess, but I think stockfish 8's elo in 2020 would have been 3400.

Frikgeek's avatar

What you're missing is the difference between Stockfish 8, which is a specific (outdated) version of Stockfish and level 8 on lichess, which is an implementation of the engine adjusted to present a specific level of challenge. Lichess simply never implemented a difficulty higher than level 8 because only IMs and GMs are able to beat it anyway.

It's like CPU opponents in video games, even at the highest difficulty they are designed to be beatable. Here the levels are just a difficulty slider, with 1 being the easiest and 8 being the hardest. They have nothing to do with which version of Stockfish is used.

Matthias Görgens's avatar

You seem to be generalising from American sports a bit too much?

Most of the sports the rest of the world enjoys don't have American socialism like wage controls built into them. Just look at the insane transfer prices in football.

fraudconcern's avatar

Someone in this audience has probably thought this through better than I. If so, help me out.

I recognize the top end of the distribution is where the excitement should be, but am concerned that the prediction market might be expected to fail specifically the top end.

My intuition is that market liquidity is low enough that predictors can earn more leveraging their ability outside the prediction market than they can inside the market. (Perhaps even hiding their ability to the broad world so as to maintain the arbitrage opportunity?)

If so, this would mean top end talent should be consistently drawn out of the market, leaving the remaining visible top end a skewed group.

I recognize this seems "conspiratorial" but it also seems reasonable if not logical, and if we are drawing strong conclusions about what the top end currently "visibly" looks like, we should keep in mind that a good portion of the top end might not be visible.

Trevor Vossberg's avatar

The obvious comparison here is Research Taste eg in the AI Futures model. Research Taste is what enables really strong takeoffs. Forecasting potential gives us an initial glimpse of if AI can predict things dramatically better than humans.

Sol Hando's avatar

I suspect the marginal value of having slightly better predictions than current superforecasters is almost zero.

While yes, they're currently employed by hedge funds and prediction markets a billion dollar industry, I suspect the value of their superforecasting is largely positional rather than an inherent good. In that being the only hedge fund with a team of superforecasters gives you a large advantage, but everyone having access to superforecasters only makes the market a bit more efficient on average.

Scott Alexander's avatar

Not sure what you mean. I agree it doesn't mean that every hedge fund will make beating-all-the-other-hedge-fund levels of money. I think that for the financial side, you've got to believe that there's some value in markets being efficient. For the nonfinancial side I think the benefits are clearer - we can know more about likely policy outcomes.

Sol Hando's avatar

I mean that I think the value of the slight improvement from current superforecasters AI may give us will not be large, and will become hard to see at all once AI superforecasters are essentially accessible to everyone.

My thought is that the ability to predict things slightly better than almost everyone else (via human superforecasters) is currently extremely valuable, because if you're the one guy or organization that has the best predictions, this can give you a large advantage in markets when trading against people who are making worse predictions. But once AI gets as good or better than the best, and more importantly is widely accessible, there will be no competitive advantage to superforecasting, since you'll be up against other people with AI that make equally useful predictions.

I.E. If I was the only guy on the planet with a magical oracle that told me the price of a stock in 60 seconds, I'd be a trillionaire in the next year or two. If everyone had the same oracle, then you couldn't make any money from it.

Markets will be slightly more efficient, but at least in the realm of prediction markets I don't think the value will be especially high. At least with efficient securities markets you can claim people getting an accurate price for their asset is a good in itself, but no one is holding "Will Trump say teeth in his next speech" long calls in their retirement account, or trading election-results-backed securities. It's as if someone promised me they could make sports betting more efficient, which I suppose is desirable in an abstract sense, but practically I don't see it.

And beyond that it's just my difference of opinion as to the value of increased accuracy in prediction markets is practically useful at all.

Nancy Lebovitz's avatar

A little off from your main point, but are there people who habitually use simple, sensible prediction methods?

Scott Alexander's avatar

This depends what you mean. A lot of medical algorithms are sort of descended from simple sensible prediction methods - things like "only run this test if this number is above X *or* the patient is complaining of at least two out of the three symptoms on list Y". Is that what you're asking?

Lucas's avatar

I don't agree that AI means LLMs all the time. TikTok recommendation algorithm are AI, they're not LLMs (as far as I'm aware) for one example.

Once they're done scanning the last book companies can do like deepseek and improve post training and get a huge jump in capabilities. Idk why companies scan books, probably because it's cost efficient? But I don't think it's out of desperation. To me it feels more like, it's probably cheap, it's harder than internet stuff so it may be an advantage vs other companies, and it would feel a bit silly to not do it.

Anki is spaced repetition software, spaced repetition is a way to review material so that you can retain more stuff in less time.

Agree that general intelligence is super poorly defined, but I think "interpolation an answer out of a corpus" is also super poorly defined, and the tasks that qualify as "rabbit pocketwatch" are also badly defined (in fact usually substituting general intelligence for someyhing else that's described as a kind of pattern matching quickly ends up with the same flaws as discussing general intelligence).

Agree that LLMs can't navigate a street in real time, but to me it feels like a very rabbit - pocketwatch task. But it seems like computer vision is hard and LLMs are "slow" so driving becomes hard. Is it a consequence of LLMs being fundamentally flawed, driving being fundamentally hard, or humans having a great visual system, or something else entirely? I have no idea.

Josh E's avatar

I think John Henry eventually loses this one in the prediction game, but it's some distance off. I was/am working on a Superforecasting bot myself, and all the data points at "practice, practice, practice" for me, and leave the code alone until your own Brier score is worth bragging about. I don't think the bar for usefulness is necessarily "Better than human" though. Even if I "Get Good" in the sense of market beating, there's a place where I might outperform AI, but not by enough to justify the time. If the money printer goes BRRRR in full auto, why would I spend my day crunching numbers for a nominal improvement in returns?

Rob's avatar

As someone who used to be very interested in NFL modeling, let me share a fact I found surprising. There is a reasonably-knowable cap on the predictive power that our models can have.

I am going to talk about model strengths in terms of game point differentials, because that is easily interpretable and has more signal than win-loss results. As a first pass, you might guess that the result of each game is due to home field advantage plus the difference in the strength of the teams plus some random contribution.

pointdifferential = HFA + strength.team1 - strength.team2 + random.game

It turns out that if you make a statistical model with these assumptions, you can show that roughly 75% of the variance in point differentials in NFL games is due to something other than the (static, game-invariant) team strengths or home field advantage. So no model that makes the assumption of static team strengths can even in principle explain more than ~25% of the variance. All real models are lower than that because they can only approximate the team strengths with limited data.

You would think that we just need to look for explanatory variables that might change between games. There are definitely some that matter, like player injuries. In an NFL context, adding those considerations has made very small improvements to the amount of variance that can be explained. A paper by Lopez in Annals of Applied Statistics in 2018 shows how much of the variance in Vegas lines is due to team strengths, changes in team strengths week-to-week, and game specific considerations. Based on the work in that paper plus my own arithmetic, we can say that with all the knowledge that Vegas has about player injuries, team matchups, "momentum", weather conditions, etc. they capture an additional 3% of the variance in point differentials. So assuming that Vegas has captured a significant portion of the stuff that varies between games, I think we can be confident in saying that no model is ever going to explain meaningfully more than 28% of variance in point differentials. (Vegas is at about 21% because they cannot perfectly capture the static team strengths with limited data, especially early in the season.) Such a model would be expected to pick the correct winner 68% of the time vs. the measured 66% for the Vegas line.

So I think 28% of point differential variance explained is a reasonable ceiling and the math says that the win prediction ceiling of 68% is very firm.

You can do the same kind of calculation for other sports.

DamienLSS's avatar

As a recreational bettor, this is really interesting and good insight. Thanks.

Would Scott's rejoinder (personally I'm less sanguine than he about AI capabilities) perhaps be that sufficiently advanced AI could identify additional variables to include in the equation? Weather, perhaps - maybe it could adjust more easily for which "team strength" is more affected by the projected rain/snow/sun. Or team propensity for nightlife affecting performance. Or perhaps it could finally quantify the elusive "clutch" trait, something related to mental lapses. Or brute force analyze matchup by matchup to identify likely performance weaknesses. Field conditions. I don't know, I'm grasping, but I assume that has to be the play for Scott's ASI to possibly exceed the predictability cap you're proposing.

Melvin's avatar

Any variable that you can think of has probably already been tried and incorporated into models by professional gambling syndicates.

I assume that understanding which team plays better in bad weather would be useful, but of course there's very limited data available on this sort of thing anyway (how many times has this team with its current player lineup played in the rain?) so it's not going to add a huge amount to your model. It makes sense to wrap this up into the extra couple of percent that you might be able to squeeze out.

It makes sense that there ought to be a ceiling in predictability based on publicly available data, and that we're probably already pretty close to the ceiling in any area where there's a lot of data and a lot of ability to make money, like horse racing.

Rob's avatar

I think the best explanation of the situation with sports is that there is an inherent variability in each performance from an athlete that cannot be explained by any summary statistic you would know before the game. I can't prove the shape of it from first principles, but we should expect a ceiling effect where more model sophistication eventually stops yielding better predictions.

Imagine someone trying to predict the weather by making a more and more advanced almanac. First you would just find the average temperature and precipitation by month, which would be a big step up from knowing nothing. Then you could try to incorporate more recent seasonal tendencies, solar cycles, and other variables. But eventually you would hit a wall in how much variance is explainable with that kind of analysis. The weather a few days from now is in principle predictable to a high degree of accuracy from physical laws, but not with the kinds of data that an almanac uses. To do better, you would need to gather lots of data about your region and the surrounding ones for the days immediately before your prediction. Predicting sports outcomes with summary statistics and publicly available information is equivalent to the almanac. Maybe an AI could do significantly better than us if it had access to tons of fine-grained data of some kind about all the athletes, but it doesn't.

When Scott compared prediction markets to "models" he referred to a model based on a statistically very weak indicator of team strength. A different, arguably simpler model I built does much better, hugely so at the beginning of the season. The gap in the remaining error like Scott graphed is tiny between that model and the best models that professionals have ever made with money on the line. For your question, if a thousand-fold increase in effort and sophistication only modestly improves the models' predictions, I am inclined to believe there is no more signal to strangle out of the data sources. And subjectively, I would expect factors like field conditions to be much less important than injuries, midseason starter changes, etc. which Vegas could only successfully use to explain 3% of the variance. If there is some modelling advance that will go past my proposed 28% ceiling, either A. there is an undiscovered variable/mechanism much more important than things like injuries which are already being captured by the best models, or B. all of the best modelers have somehow failed to capture the predictive power of the variables they are looking at.

Ohad Avnery's avatar

Box office receipts also seem optimized against prediction in the limit. Studios want to make money, and they'll run the best AI forecasters internally to decide what to produce. This will reduce the number of films that are "obviously" going to fail. Since the space will be more competitive, I believe it will also reduce the number of films which obviously succeed.

Actually, a lot of reality might behave the same way, including politics and elections.

Russell Hogg's avatar

Once the AI's start outpeforming the super forecasters on a regular basis (even if only by a bit) do they just give up on the field? Forecasting is a lot of work after all.

Dweomite's avatar

> A three pawn handicap in chess is about 3x the difference between the world champion and a average grandmaster.

Are you assuming that a 3-pawn handicap is "3x as big" as a 1-pawn handicap? I doubt it's linear like that.

That is, if Alice can just barely give a 1-pawn handicap to Bob, and Bob can barely give a 1-pawn handicap to Carol, and Carol can barely give a 1-pawn handicap to Dave, I doubt that implies that Alice's advantage over Dave is equal to 3 pawns. I'd guess it's less than 3.

AI says (agreeing with me) that the appropriate handicap between the best and worst players will generally be smaller than the sum of the individual handicaps in the chain, but also notes that the Elo equivalent of a given piece advantage is drastically different depending on the players' skill (a 1-pawn handicap is 70-100 Elo at lower levels but 200-250 Elo for grandmasters), and that you should remove a knight instead of literally removing 3 pawns, because missing 3 pawns is virtually unplayable because your pawn wall will collapse.

This also means that being "slower to achieve the third pawn than the first or the second" is even less evidence of a plateau than you'd otherwise think.

Dylan Richardson's avatar

I suspect that the major merit of AI forecasting bots is that they are machines; not some greater degree of insight or creativity. Just thinking longer. For geopolitical type questions, it seems like there is near infinite potential thinking to be done. Human superforcastors just prioritize what they can. But AIs could find 100s+ of distinct base rates to factor in, and make just as many updates on new factors.

What are some other p(Taiwan invasion 2035) base rates?

- rate of historical revanchist conflicts

- rate of expansion of Chinese control over regional waters

- the above, but while under control of different military leaders and CCP authorities

- survey data of expected invasion dates and rate of shrinking/growing horizons over time

- surveys of Western experts

- surveys/indications from, CCP members

- the current Metaculus, kalshi and poly market predictions

Ongoing updates?

- success/failure of Russia's war

- extent of US response to Russians war, relative to expectations

- Agents could even do original "work", like sentiment analysis of CCP speeches with Taiwan mentions over time.

- presence/absense of anti-war or preparation skeptical voices in CCP over time

- rate of military expenditure

- rate of military expenditure relative to economic growth

The super smart folk at Samotsvety do some of this anyways, but there's simply a limit to how many such tasks they can take on.

James's avatar

I would've gone with super-duper-forecasting

Warren Hatch's avatar

A note on terms, since it bears on the argument here. “Superforecaster” is a verifiable accreditation, not a compliment one hands out freely. Lowercased, it refers to the top 2% of forecasters in the Good Judgment Project (GJP), identified based on a documented track record across multiple tournament seasons. Capitalized, it’s Good Judgment’s registered trademark and refers to Good Judgment’s professional Superforecasters, many of whom came from the GJP. Applied loosely to anyone with a good run and little if any record-keeping to speak of, it stops meaning anything. Don't use it that way.

That misapplication shows up in your scale. “Professional human superforecasters: 35” sits in a list you describe as approximate and partly napkin-math, and all your three anchors relate to that number. But the data has nothing to do with Good Judgment’s professional Superforecasters. All human data points behind the scale refer to Metaculus and Samotsvety’s own track-record page, and you note that the latter reads as an advertisement.

There is only one public benchmarking competition that includes Superforecasters, ForecastBench, and those forecasts are now over two years old. We’ve written up our other reservations (https://goodjudgment.substack.com/p/the-hype-and-the-evidence-about-ai), along with an invitation to anyone who would like to set up a proper competition. Conversations are underway, so I hope the next time there’s a headline about AI and Superforecasting, it will be current and include actual Superforecasters.

None of this is anti-AI. There’s a productive division of labor to be had on real-world issues that can contribute to better decisions and the public discourse.

Chastity's avatar

> (some of my draft readers argued that chess is an almost optimal domain for AI - it can self-play millions of times for training data - and forecasting is much harder. I agree, but I don’t think that matters here - we’re not talking about how practical it is to create superforecaster AIs, we’re talking about what chess can teach us about the degree to which top humans approach optimal play.)

I mean, the main thing about chess is not that the AI can play itself, but that it's a game without secret information. There's absolutely a bunch of secret information in geopolitics. A better comparison would be poker, which is similarly mathematically solvable/vulnerable to self-play, but luck and hidden information play a major role. (Or the weather, where top AI models are now doing better than previous techniques, but still can't predict the future that far in advance.)

To borrow the terminology used by Tetlock in his works, chess is clocklike (governed by highly comprehensible mechanical rules), whereas geopolitics is cloudlike (governed by small random chance, e.g. Yeltsin's decision to make a certain random ex-KGB guy his successor).

draaglom's avatar

I'm one of the Metaculus Pros you are talking about in this article, and I am one of the only (if not only) person to have participated in every instance of FutureEval so far, so when comparing the lines here you're substantially talking abut my forecasting specifically.

It's not at all clear how you're imputing the relative score of Metaculus Pros vs Samosvety forecasters; in fact I don't see any information on the Samotsvety track record page that could even in principle be used for this purpose, given that the pool of both players and questions between INFER in 2020-2022 and Metaculus in 2024-2026 when FutureEval has been running has limited overlap, so it's not meaningful to e.g. take the Samotsvety vs INFER pro team scores in 2021 and apply that same margin to the Metaculus pro team in 2026 (I assume this is what you're doing, since that's the closest to a comparable?)

My best _subjective_ judgement is that there's not much of an edge at all between the current Metaculus pro team and Samotsvety; but this is a guess based on forecasting n =~ 5-10 questions where a Samotsvety person also did, rather than something rigorous.