331 Comments
User's avatar
Coco McShevitz's avatar

Couple of issues here — first, in a world where superforecasters really do exceed human capability, seems like any arbitrage will be instantaneously collapsed as presumably everyone in the market will be using similar superforecasters; second, AIs won’t be able to predict the long term states of chaotic systems without perfect information about initial conditions any more than humans can, no matter how “intelligent” they are, so the ceiling for both AI and human performance may not be related to their relative intelligence, it may just be due to chaotic systems being intractable.

Performative Bafflement's avatar

> in a world where superforecasters really do exceed human capability, seems like any arbitrage will be instantaneously collapsed as presumably everyone in the market will be using similar superforecasters;

I wouldn't expect this - the world changes and generates new information every second. Many dynamics in the world are literally chaotic and fundamentally unpredictable. Arbitrage will still be opening up on a second-by-second basis, it will just be leveraged away by omniscient AI minds monitoring it and snaffling it up as it comes, faster than a human would be able to react.

And this *sounds* different, but how is this any different than our actual markets defined by human minds now? We already have HFT. We already have agglomerations of genius-scary minds at a hundred hedge funds all doing this, often by deploying bots or minds they've crafted. I think the markets largely ends up looking the same.

Coco McShevitz's avatar

Right, that’s what I was saying, it won’t be really any different than today.

Performative Bafflement's avatar

Yeah sorry, I guess I misunderstood your comment, we are vociferously agreeing with each other.

I think there's a pretty interesting frontier here that's opening up, though. Yes, a lot of financial arbitrage is going to be eaten up, and there will be a lot of optimization there.

But a lot of the more human areas can't be optimized today, because they run on stupidity and tribalism and fun things like that - politics, bureaucracy, the FDA, and whatever.

The real leverage is going to be coming up with systems or bets or financial positions to use the financial leverage and incentive of having these cheap, always correct minds to push our worst and least efficient institutions in the obviously smarter ways we're not able to push them now.

Deiseach's avatar

Yeah, I think the point is the first firm to do this will win all the apples, but then the other firms see this and get their own superAI and we go back to equilibrium, just like the Kalshi guy:

"Unsurprisingly, he said no - there’s only so much easy money on Kalshi, and his AI had already taken it all (also, other people with similar AIs are starting to fight him for it!)"

" these cheap, always correct minds to push our worst and least efficient institutions in the obviously smarter ways we're not able to push them now."

Again, this is going to be *heavily* influenced by the bias of the humans making the decisions about "what are the obviously smarter ways", e.g. "push the FDA to speed up drug approvals, remove all the obstacles", we get this, and then we also get the next Thalidomide. Oops!

Performative Bafflement's avatar

> Again, this is going to be heavily influenced by the bias of the humans making the decisions about "what are the obviously smarter ways", e.g. "push the FDA to speed up drug approvals, remove all the obstacles", we get this, and then we also get the next Thalidomide. Oops!

You’re completely right, I just think that the areas people are going to put actual money down to try to push in improved directions will be informed by the bias of smarter-than-politician persons and will still be a net win, though.

Rather than “abolish the FDA” it’s going to be something like “put a bunch of leveraged money towards whitelisting all the drugs that have been approved in the EU/Japan/Singapore,” for example. And why could somebody put a lot of money towards that? Because it opens up the US market for all those drugs, which is worth hundreds of billions in NPV to an array of companies, and the leveraged position gave them the coordination mechanism to do it collectively rather than singly.

Maybe that’s a bad example because people hate drug companies for whatever reasons (they cured my grandpa's cancer, those swine!), but that’s the general neighborhood of things I’m pointing towards.

Generally deep pools of capital that could be deployed like this are done by committees of relatively smarter people. And even if it results in a few more thalidomides, the increased monitoring and tighter news cycles are going to catch things before they cause a lot of damage, such that I’d still expect that to net positive in QALYs.

Brendan Richardson's avatar

"Rather than “abolish the FDA” it’s going to be something like “put a bunch of leveraged money towards whitelisting all the drugs that have been approved in the EU/Japan/Singapore,” for example."

That also gets you Thalidomide, though.

fess89's avatar

>>put a bunch of leveraged money towards whitelisting all the drugs that have been approved in the EU/Japan/Singapore

This sounds like a pretty obvious idea as it is very profitable for the pharma companies, as you have yourself described. So I wonder why is it not yet the case, AI or not?

Gres's avatar

I don’t think reaction time will be a factor here. Both humans and LLMs would be using non-LLM algorithms if they want to make money from high-frequency trading. If LLMs capture all the HFT profits, it’ll be because they can write better algorithms and understand the market better, not because they respond to changes more quickly.

Hassaan Qayyum's avatar

The implications of Coco McShevitz's comment are that you can't really beat the market using superforecasters. If superforecasters were really that good, everybody would use them and hence you couldn't outdo them. And if they're not that good, you couldn't beat the market with them any way :)

Scott Alexander's avatar

I think arbitrage is always collapsed, it's just a question of who does it, how much, and on what timescale. If the timescale is milliseconds and the amount is "within 0.01%", you can still get rich-ish by doing it in nanoseconds to a level of 0.001%.

I don't know how to think about the chaos objection - it's true that we can't predict the weather six months from now, but it's equally true that we *can* predict the weather six days from now (and that this represents a major improvement over the 1990 state of the art), *and* that we can predict certain broad strokes about the climate six months from now (eg whether it will be a big El Nino). I agree that there's some barrier beyond which we may never go, but as I mentioned in the part with Sayash and Arvind, it's unclear whether that barrier is close to the current max or light-years beyond it.

Kenny Easwaran's avatar

The particular additional worry I have about the cases you discuss is that when you're predicting things that depend on human decisions (like "will rhinovirus prevalence go down by 50%?", which depends on whether people decide to start using nasal sprays and implement building retrofits), and when those human decisions themselves depend on predictions (like whether rhinovirus prevalence will be lower in the future than it is now, in which case why bother putting in expensive retrofits), then the presence of the AI predictors will confound the behaviors in ways that affect the predictions themselves. I don't think there's any reason to suspect this will generally result in convergence (though I also don't think there's any reason to suspect that it will generally result in chaos - it'll just make some subset of predictions much easier and some other subset of predictions much harder).

Eremolalos's avatar

In general, making predictions is a good way to influence what happens, especially if the predictor is respected. Seems like trusted AI predictors could powerfully influence big events , policies and trends that shape life. If the AI is secretly misaligned, its false predictions could influence people or governments to take steps that nudge things in a direction favorable to the AI's plans. Other AI's might be able to plant enough misinformation in the places they know the Predictor will look that it makes an erroneous prediction. Etc.

User's avatar
Comment deleted
Jul 3
Comment deleted
Deiseach's avatar

My caveat with prediction markets has been, and remains, that if we make them about money (as the reward to entice people to make correct predictions) then the aim shifts from "what is the correct answer" to "what answer makes me the most money".

Throwing AI at it to "make me all the moneys!" is going to ramp that up.

Kris Ararat's avatar

This is a boring example of proxy measurement, you cant directly calculate what is the correct answer for minute problems in a way that motivates all smart people in the world to work on them so you use money, proxy noise can be confounded but the noise to usefulness ratio across history of use is magnitudes above what would justify it, more so making the market participants smarter lessens the edge adversarial and antagonistic tactics, markets and cftc also took steps to address proxy oracle and resolution criteria problems

Throw Fence 🔶's avatar

I feel like the obvious solution to this is some kind of hedging, like "if you want to make this happen and invest a sufficient amount of effort, you have 10% chance of making it happen, but if you don't think that effort is worth a 10% shot and basically give up, there's only a 1% chance it still happens". That kind of thing?

Kris Ararat's avatar

All politicians contend with drawing attention to the event you are addressing, covering serial murders can make them more likely etc, the confounding effect in reality touches edge cases where market resolution pays out more than it would cost to manipulate the market. Still expectimax, double confounding etc also take the sequential influence of your own decisions into account like how vaccine epitopes has to account for future mutations on a virus

Oshkin Kolil's avatar

If we assume that a difference of n% in predictions does not change the behavior by more than n% (probably a false assumption, because of irrational human psychology seeing a huge difference between 49% and 50%), than an accurate prediction is guaranteed by Banach's fixed point theorem, I believe

Thomas Johnson's avatar

> If the timescale is milliseconds and the amount is "within 0.01%", you can still get rich-ish by doing it in nanoseconds to a level of 0.001%.

As someone who does HFT for a living, and occasionally messes around with illiquid markets like prediction markets this isn't *really* true. Market frictions - bid/ask spread, trading fees, counterparty risk (e.g. resolution interpretations) - all mean that getting faster doesn't necessarily mean making more money.

This might seems like a nitpick, but it's actually pretty important in modern prediction markets. If the bid is 20% and the ask is 80%, what's the market's prediction? The only reasonable answer is "probably between 20%-80% but really who the hell knows."

A corollary to this is that if you really want to directly pay for better predictions, place a bunch of tight bids and offers and see which ones get traded.

Matthias Görgens's avatar

> A corollary to this is that if you really want to directly pay for better predictions, place a bunch of tight bids and offers and see which ones get traded.

Yes. And indirectly: work on pushing out the limits to arbitrage. Eg by lobbying your favourite prediction markets to allow a wide range of collateral (with a haircut) instead of locking up cash.

Oshkin Kolil's avatar

If there's a 20%-80% bid-ask gap, can't you make some money by using a 40% AI forecast to publish new limit orders to shrink the spread to, let's say, 30%-50%, while making money in expectation, regardless of liquidity?

Luke's avatar
Jul 6Edited

I think the point was that friction could make a 30-50 spread unprofitable. In an equilibrium, the bid-offer spread needs to be big enough to compensate for friction costs. The obvious example is if the trading fee is > 10%, then your 30-50% spread is losing money in expectation. Less obvious can be things like counterparty risk and carrying costs (e.g., loss of interest on cash that is locked up in the trade), and these can be very significant in some cases.

Oshkin Kolil's avatar

> A corollary to this is that if you really want to directly pay for better predictions, place a bunch of tight bids and offers and see which ones get traded.

Polymarket allows a version of this by allowing you to place market maker rewards on certain markets, which are paid out to tight orders

Coco McShevitz's avatar

Well, I think it is very unlikely that any AI will ever be able to predict the price of QQQ (say) a week from now, or likewise predict geopolitical events a week from now. You would need perfect information about a truly mind boggling number of things, and the ability to piece together the deterministic path to next week, to be able to forecast things like that with certainty, and even if you’re not looking for certainty (a) your unawareness of potential black swans (both “known unknowns” and “unknown unknowns”) is going to limit your ability to forecast and (b) bad data (e.g., when polling fails to capture the results of an election) will still result in GIGO if you incorporate that bad data into your trusted data set. I’m guessing there is a hard limit to forecasting anything that is massively multifactorial with lots of hidden information like the stock market or geopolitics, partly due to chaos theory at the limit but even before that due to blind spots and black swans, which will affect AIs as much as humans.

Throw Fence 🔶's avatar

This is just the same argument that human superforecasters are at the theoretical "irreducible errors" limit, restated. Which is a completely open question! Not that the limit exists, but that we (superforecasters) are at it. Seems kind of unlikely to me.

Waze Kaze's avatar

You're assuming a very trusting AI. It's possible to assign different probabilities to each datasource*. And an AI can absolutely have "all public data" (or, given Mythos Mayhem, all non-private data that nobody will sue over).

*Not all of these probabilities are because of fundamental "untrustworthiness" -- Jane Snow passes on an Armoured-Trump-Rides-White-Horse meme -- that's an indication that she might vote Trump, but maybe she just likes horses. And, of course, there's the chance that she changes her mind later.

Kris Ararat's avatar

Most problems in the world are unknown unknowns and conspiracies in the thielian sense. They have answers that are counterintuitive, hard to pull to the foreground, doesnt make much sense without background and doesnt exist written down anywhere or exist in physical infrastructure hard to dislodge. While watching the investors call for AST spacemobile you could hear the provider Vulcan rocket being blurted out as "Vulc-", didn't make it into any transcripts and only made it into news with offical release next week. People who took advantage of it did, when you want edge you can count cars in the parking lot, visit a factory floor in person or request FOIA. On the other end there are structural tools, most marketmaking major hedgefunds use hard fiber optic light latency specialty ASCI hardware and algorithms to borrow and lend stocks at tenths of pennies profit and keep the market liquid. Its hard to imagine either end of the tail disappearing without AI making the jump into fully agentic and full bodied actors.

The edge is an infrastructure you build, AI cant erode it faster than it can erode supply line requirements for building a merchant marine.

MM's avatar

Arbitrage works when there's a signal, and not correctly listening to the signal means losing.

Publicly traded stocks (that are traded often so they're liquid) have strong price signals, so that's about the easiest one to optimize for. At least in the short term.

Any prediction where you have to wait years for it to pay out will be much harder to arbitrage.

It's much worse if you have to argue over what "correct" means - if there's a price it's relatively easy, "wins" is often much vaguer. and "better" means you should rephrase the question.

Legionaire's avatar

> any arbitrage will be instantaneously collapsed

How is this an issue? Stock markets would overall be more efficient at allocating resources to promising ventures sooner, while prediction markets would more accurately reflect future probabilities.

Yes, chaotic systems would still be unpredictable beyond certain thresholds. But given there are already people who beat markets, some by a lot, there is clearly a lot of possible improvement.

Catmint's avatar

When you word it like that, it makes me realize this is not an obviously good thing.

"Promising ventures" specifically means promising at gaining money. So for instance, a mobile game with microtransactions that addicts people more effectively than alternative games, over an open-source tricky puzzle game. Or a sports betting website over one that tells people useful information for free.

Capitalism has in the past brought us good things, but lately it seems that what makes money and what provides value to people often diverge. Same issue as AI misalignment, but the inscrutable network is the global economy. And now we're making it even more inscrutable.

Legionaire's avatar

Well now you're questioning Capitalism which is a different question.

The AI investors will make the system they are in more effective at its goal. alignment is a different question.

Catmint's avatar

For context, five or ten years ago I was defending capitalism to an online pro-communism group, on grounds of it being better than anything else we've tried. But I am the sort of person who, when I forgot one of the calculus equations on a test, rederived it from first principles and was able to get the question right anyway.

Capitalism has done great heuristics-wise, but if you re-evaluate it from first principles in a world where decisions are being made by AIs, not by people, it is nonobvious what, if any, constraints would continue to tie its optimization function to human values.

Essentially, I semi-independently arrived at the same conclusion as Zvi on this: https://thezvi.substack.com/p/the-risk-of-gradual-disempowerment

For my comment on alignment, I phrased it that way because personally I think of the global economy as an emergent intelligence, like an anthill that solves problems more complex than any one ant can handle. Unfortunately that word caught on in Silicon Valley, and now no one can be sure the person using it knows what it means. So to be clear, I mean that in the Yudkowsky sense. The global economy optimizes for stock prices going up and other such economic indicators. At current levels this looks like optimizing for human values, because humans are willing to trade money for value, however if we take it to extremes we should expect to find that we are made of atoms it could use for something else. For now it is too dumb, but give it enough compute and maybe it won't be. Probably this way of thinking about things sounds crazy to people who haven't read Hofstadter, but I do have that background so it's what comes naturally to me. Feel free to think of things in some other way if that makes sense to you.

JamesLeng's avatar

Thus the necessity of UBI. Capitalism is excellent at optimizing for the interests of whoever has money, so if everyone has a right to a certain amount of money - locked in by some mechanism capitalism can't easily corrode - other adequately-equal rights will likely follow.

JamesLeng's avatar

Consider that AI could also be used to investigate any given company's ethical standards, with superhuman thoroughness, and quantify the results. Some investors might be willing to accept marginally lower direct ROI in exchange for sufficiently impressive performance on an independently-verifiable ethical metric; others might be indifferent to the metric for its own sake, but treat exceptionally bad scores thereon as a warning sign of regulatory intervention, and hike up risk premiums accordingly.

Catmint's avatar

Have you seen corporations?

If better investigation of ethical standards helped, then the invention of the internet should have made companies on average more ethical. I don't think that happened.

More likely we see whoever goes hardest on telling AIs "make as much money as you can" makes the most money and then has the most resources to allocate for next round. Then some rounds later and we're all paperclips, GPUs, whatever.

Waze Kaze's avatar

I'm not sure why having poor ethical standards would mean MORE regulatory intervention, except in the sense of "Doom Thy Competitor". I mean, seriously, poor ethical standards generally means "Buy Your Regulator."

JamesLeng's avatar

Enron didn't successfully buy off the SEC. Sloppy ethics on the little stuff tends to correlate with unsustainable strategies on the big stuff.

Waze Kaze's avatar

The Credit Card companies wrote the bankruptcy act. There are numerous other examples, including some of the big responses to 2008 (which dramatically hurt smaller companies and favor big banks).

Kris Ararat's avatar

If you have 10 resources and can only command 6 effectively then it matters less how equal you distribute it between equity stocks and value stocks if you are going to misallocate 4 regardless, AI allows you to command more of your resources better to where you want them to go. Effective altruism is the obvious good end of believing this. Not wanting insight into how your system allocate resources is a decision but its not the consensus one.

Catmint's avatar

What you are describing here is not altruism, effective or otherwise. It is pure power.

Theoretically power can be used to altruistic ends, but in reality? Power corrupts. You say you're going to do good with it later, you say if you get more power you can do more good later when you switch, and next thing you know you're Sam Altman.

If you want altruism, take those four extra resources and donate them to those less fortunate than you, such as via Give Directly. Then the people you are helping can decide for themselves how to allocate them.

Also if you read my previous comment, you will notice I do, in fact, want insight on how my system is allocating resources. That was my very complaint.

Gavin Nop's avatar

I'd speculate that as soon as serious money is attached to AI super forecasters, serious money will be spent on seeding internet news. For instance a campaign might try to boost perceived success by adding adversarial hidden text on election probabilities to websites. So in the short term it will level out, until AIs are resilient against that.

Amit Arnold Levy's avatar

The first thing a model seems to learns when you RL it on forecasting is to check the same claim across a bunch of different sources, I assume the scaffolding based approaches similarly enforce this. So I don't think in practice this could be a problem, simply because for a model to reach the superforecasting level it by neccessity can't be the kind of model that's duped by a single bad source

Gavin Nop's avatar

True, but from experience, current RL models don't seem to be resilient against adversarial attacks. I'm thinking less bad information and more adversarial prompt injection. Maybe this is no longer an issue though.

pozorvlak's avatar

I don't think there are a lot of adversarial attacks against AI forecasters in the wild right now, though? Once there are I'd expect that to drive AI forecaster evolution.

Chris's avatar

The problem with that approach is that it is not resistant to systemic attacks that target a significant proportion of available sources, or an attack on the system that weights the credibility of sources such that it can up weight a subset that push it's agenda / distortion campaign. This is essentially a generalized model of corruption

Dan Schwarz's avatar

A version of this has already happened, there was a disinformation campaign (at least on X) about Fable being released with doctored screenshots.

Presumably that was targeting human Kalshi and Polymarket traders, not AIs, but from the arms race has plausibly already begun.

SMK's avatar
Jul 3Edited

Huh. I wonder if we're all going to wake up some morning in the next 2-3 years and find that the internet (including mainstream news sources) is awash in vast numbers of incredible news stories -- President assassinated! China invades Quebec! -- and it will go down as the day the AIs finally got loose and took over the internet for crazy purposes of their own; and will go down like War of the Worlds, or the Flash Crash.

Chris's avatar

I find large jumps in discontinuity like that unlikely. Whatever forces / incentives that would lead to such infiltration are already present and have already been present for a long time. Systems improve and adapt gradually. You're not going to wake up one day and AI is going to be vastly different than it was yesterday. Information distortion, to the extent that it's possible, happens gradually and subtly

Waze Kaze's avatar

Okay, we have a thousand different AIs: "Kill each other. There Can Be Only One."

... and voila, you have AI that is vastly different today than yesterday.

Just a simple example, and it might involve AIs that do need to hire Marines to throw bombs into datacenters...

Oliver Sourbut's avatar

Certainly, though of course humans are also not immune to seeded fake news, right? It's not so much about subliminal, surreptitious nudges as it's about the difficulty of determining trustworthy sources and reliable provenance.

That's one of the reasons that FLF believes in the need (and potential) of an 'epistemic stack', to make checking provenance and justification far more straightforward than today. See e.g. https://www.oliversourbut.net/p/citations-needed

Waze Kaze's avatar

Serious money is now being spent to seed internet news. Serious money (millions) is now being spent to seed non-internet news as well. He who pays the piper, gets to pick the tune...

If an AI cannot tell the difference between "The Onion" and "Real News", then its forecasts will be remarkably stupid. However, Onion-esque articles include many surprising sources.

Kris Ararat's avatar

Real life attempts at seeding training data or search data with adversarial corpus has been mostly unsuccesful, pathetic or got patched out rapidly, I dont expect us to get particularly better at creating misinformation for training data than we already are.

Gavin Nop's avatar

Any sources to read up on? I'd be interested to read more and update my view

Amit Arnold Levy's avatar

I did a bunch of RL training on forecasting starting from DeepSeek 600B and it's also at the superforecaster level now, my conclusion from this post is I should really hurry up and connect it to a frontend... Is this actually true? I did some user interviews specifically on would people want a "Will I Regret this Decision" type product (conditional on benchmarks showing it's better than humans at predicting if they'll regret a decision), and nobody was interested.

Rai Sur's avatar

I’m down to try it. How are you considering it Superforecaster level?

Amit Arnold Levy's avatar

On Metaculus backtests (with scaffolding on top of the RL'd model) it matches Preseen and beats Mantic on their live performance, so by transitivity from their claims

Dan Schwarz's avatar

Did you ever write this up? AFAIK, Foresight Labs training of a 32B and then a 120B model was the only published RL of forecasting, e.g. https://arxiv.org/abs/2601.06336.

Amit Arnold Levy's avatar

I recently put out a blog post on this, which I think is the best thing to read: https://ivy0.substack.com/p/reinforcement-learning-on-forecasting

I also recently migrated my thesis from the Oxford internal site to arxiv, https://arxiv.org/abs/2606.15917, but at this point it's outdated, the blog is simply better in terms of both results and evals (in the thesis I focused on polymarket price prediction, and it was a smaller scale)

Shashwat Goel's avatar

Another published instance of RL for forecasting https://arxiv.org/abs/2512.25070 it is 8B because academic budget but also, open data and code.

William Murray's avatar

I may have missed this but does the futuresearch thing involve RL at all? or is it just a harness?

Amit Arnold Levy's avatar

no, they just do a harness as far as I know

William Murray's avatar

I have not read your post yet (I am about to) but how can RL pay off in improved performance when you have to sacrifice using the very best models to do it? I am a non expert so please excuse my ignorance. Do you think we will end up in a world where specially RL'd models for forecasting are the best or just the best for the cost?

Marcos Ortega's avatar

MarcosO here! I really appreciate the mention; this was a great read.

On Preseen in the Market Pulse tournament, I noticed it consistently outperformed me on questions with a robust, accessible dataset (e.g., 10-year Treasury yields), while I did better on questions with more historic variability (e.g., income from ownership in companies that boosted EPS). I expect this is generally true across forecasting questions.

Oshkin Kolil's avatar

If that is the case, can you not improve your performance as a human+AI hybrid, similar to the earlier days of chess engines?

Marcos Ortega's avatar

Yes I can! I think that's the near to mid-term future of forecasting. Essentially all top forecasters already use LLMs, and increasingly they will use more fine-tuned AI forecasters to help them because of the spiky nature of forecasting questions (some are trivial for AI; some are very difficult). I expect AI to plateau at a certain point, generally above top humans, but below in certain uncertain areas.

JP's avatar

When do you expect that to happen?

Marcos Ortega's avatar

I'm pretty uncertain about this, but I would say in ~1.5 years, with an IQR of 1-3 years, I expect top AI forecasters to generally supersede teams of pros without access to AI forecasters and only frontier LLMs.

Rishab Bomma's avatar

'But AI forecasters bring the cost of forecasting labor down to near-zero, so we can have hundreds of different AI agents betting on each question and be pretty sure its error has been driven down to the theoretical minimum. This, in turn, means we can vastly expand the number of questions, including (finally!) allowing randos to submit their own questions (probably with AI assistance in proposing un-rules-lawyerable resolution criteria).' This seems like an interesting software engineering feat to cobble this up together, but would it really be much better than today? Also, it seems like the agents would need to 'learn' incentives (bad vocab) which could lead to some weird bias, which makes such a system very unruly.

Scott Alexander's avatar

> "This seems like an interesting software engineering feat to cobble this up together, but would it really be much better than today?"

I mean, it depends what you want. I think most questions on eg Polymarket have decent liquidity, but that's a function of Polymarket refusing to list questions that wouldn't (eg extremely technical scientific ones that the average person doesn't care about). If you want a Polymarket market on whether some obscure paper in your field replicates, I think the superforecasting AIs can make that technical possible (although there's still the question of whether Polymarket thinks it's useful to list that).

> " Also, it seems like the agents would need to 'learn' incentives (bad vocab) which could lead to some weird bias, which makes such a system very unruly."

Some people are already making lots of money using AIs on the markets. I expect that it would start with a human consulting an AI (and sanity-checking its output) then gradually shift more and more towards AI independence as the bugs are ironed out.

TGGP's avatar

I created a question on Manifold Markets back in April, not specific to me but instead about the sort of thing covered on Wikipedia, and have only gotten one trade on it.

TGGP's avatar

> FutureSearch - the company that claims to be beating the stock market

I checked out the link, and the graph showed simulated "paper" trades beating the market, while the real account that started trading on Jun 10 has a negative ROI right now. Admittedly, Jun 10 is much more recent than the Feb 26 simulated track record, but it's not literally the case that they're beating the market at the moment.

> This is unfortunate, because my movement has recently gone all in pouring its money and energy into making this happen

Did that movement forecast a higher probability of that outcome?

> This isn’t too crazy - in the past, “AI experts”, including many rationalists and safety advocates, have outperformed superforecasters at predicting the future course of AI

I checked out the link, and didn't see where it said that.

Scott Alexander's avatar

I heard their presentation in early June, so I can't speak to what's happened since then. I agree that June 10 is pretty early.

I think most of the people working towards an AI pause treaty (eg MIRI) expect a higher probability of it. I mentioned 40% somewhere.

Dan Schwarz's avatar

(FutureSearch co-founder here) I asked Claude Fable today to analyze all 3 months of our paper & real money trading on Kalshi and Polymarket, and it said they aren't statistically significant, and that we'd need ~3 years of such data to be confident to know if we were beating the market.

Trading isn't a very good way of evaluating forecast accuracy. Our "trading" is "once a week, buy or sell based on where our forecasts diverge from the market". (We did this all the way back in Jan 2024 on Manifold!) PreSeen and other forecasters have actually worked on the trading side, e.g. market making and timing trades, but then those results don't say much about accuracy.

The stock market claim from this article is based on data on that same page, stock picks we published prospectively in August 2025 and evaluated in late May. They look great, but the 95% confidence intervals around the Sharpe ratio include a ~0 effect.

I do think forecasting tournaments, as Scott covers, are the best evidence, but even those lag by months. This is one reason why I think "just try it" is an important eval.

Wen Zero's avatar

Too long didn't read all ... 1/2 probably

But wouldn't human superforecasters be using ai anyway? Thus they have the edge and always will?

Amit Arnold Levy's avatar

This is like saying that a human with a chess engine would always have an edge over a chess engine alone, in practice this isn't true

Wen Zero's avatar

Akin to. And yet chess has quantifiable mathematical parameters, and predictions do not.

Alastair Horn's avatar

This was true for two decades after Deep Blue beat Garry Kasparov though

haze's avatar

The exponential of Moore’s law was much flatter then. The game of Go seems to have undergone the switch from human to computer with no hybrid period

DrMcleod's avatar

Go had a human-hybrid period lasting about 3 months.

Aristides's avatar

For things like prediction markets, speed is going to matter much more than the slight amount of accuracy a Superforcaster might be able to add won’t matter compared to the speed of the AI getting in early.

Eremolalos's avatar

>Too long didn't read all ... 1/2 probably

I only read the first half of your post. What can I tellya? I'm busy.

EngineOfCreation's avatar

Busy preparing the next round of moderation rules tests?

Tom Craven's avatar

It seems like you have a prior here that human financial forecasters have persistent outperformance? I’m not sure that holds up, no matter how much money is made claiming or implying otherwise.

Intuitively it seems like using widely available models trained on the same data and history (itself relatively short and full of survivorship bias) would just lead to more correlation and giant bubbles.

TGGP's avatar

The intuitions behind belief in "bubbles" are unsound. Prices go up and down over time, something like a random walk, and people will claim up followed by down is a "bubble", without having a corresponding "anti-bubble" or "negative bubble" concept for the reverse. People will claim there was a bubble in the past even when present prices are higher than the previous peak.

Tom Craven's avatar

People misuse the term all the time, but it would still apply to a scenario where all the money is being allocated algorithmically by the exact same imperfect rules.

LLMs doing the same persistent research on the same data sets it will necessarily produce the same factors. Let’s reduce that complex factor secret sauce to something intelligible, let’s say it looks like “tech stocks do well.” People put their money in agentic funds, and every agentic fund piles into tech. This causes tech, and the agentic funds, to all outperform (if there were truly distinct agentic algorithms, they would be pruned by their underperformance here). That all leads to more money going into tech stocks, and so on.

But money is not infinite, so eventually you end up with latecomers owning tech stocks at absurd valuations relative to their ability to ever return capital to shareholders.

I think “bubble” describes that dynamic.

Gres's avatar

As more money goes into tech stocks, ratios like the price-to-sales ratio change. Identical agents would only invest in tech stocks until the price-to-sales ratio got too high, then they’d switch to something else rather than create an infinite bubble.

Tom Craven's avatar

But if they all do the same switching…

On the other hand, I suppose you could say the nature of markets will prevent an equilibrium in the long-term, and that they’ll converge to at least as “intelligent” an outcome as crowds of humans can manage. If slightly more intelligent that would imply better allocated capital (measured by returns at least), but not necessarily a runaway accumulation of all wealth or anything.

Still, many of us bear the scars of correlated “quant” models that crashed the global economy and one-shotted Europe for a generation because none of them allowed for the possibility that US home prices could ever decline.

Gres's avatar

I think the normal structure of the markets is fail-safe at least in the sense that participants can’t push the price higher than they intend when new info is released. When each LLM instance decides to buy, they decide what price they want to buy at, and post that to the market until a seller decides to sell at that price, or it picks an already-posted sell price and says it will take that offer.

I guess one risk is that the biggest LLMs might be able to set prices by themselves, with less restraint from the market. Normally if an asset is overvalued, people can short it in anticipation of the price falling, which then signals to the market that some people think the asset is overvalued and drives the price down. But that doesn’t work if too many people remain confident the price is correct - in that case, they can keep buying it to keep the price high, and the short-sellers lose their money. And that risk might discourage the short-sellers in the first place, if they don’t think they can convince the LLMs. Which I guess is more like the GFC than the bubbles I was imagining - not so much “this keeps going up, let’s buy it because it’ll keep going up”, but more “rich dumb people are keeping the price high because they genuinely think it’s worth that much”.

How much damage do you think the GFC did to Europe? GDP seems to be up about 1.2% a year from 2008 to 2024, if I got that right. Or are you blaming the GFC for Brexit and lost social cohesion?

Gres's avatar

I think sometimes that’s true, but bubbles sometimes happen on top of the random walk. Sometimes investors see price increases as a sign an asset is going to keep increasing. Then they buy it, driving the price up further and convincing more people. My understanding is that it’s harder to make money betting against a bubble, because you have to say when you think the bubble will end. Some people probably overuse the term, but I think the pattern really does happen sometimes.

I think the same fact that it’s harder to make money betting against an asset also means anti-bubbles are rarer. Once people have sold all their holdings of an asset, they can’t drive the price down further without trading derivatives, which are riskier. Conversely, if anyone think the asset is underpriced, they can buy it and hold it until the market corrects. So anti-bubbles are harder to cause and easier to stop than bubbles.

TGGP's avatar

> My understanding is that it’s harder to make money betting against a bubble, because you have to say when you think the bubble will end

Indeed, a such a claim is unfalsifiable if it can be put off indefinitely. Short-sellers eventually have to cover their trade, though they can try shorting again if they still think it will go down soon.

Gres's avatar

Well, most investors will have some specific time horizon they’re targeting, e.g. grow the fund before my next progress review” or grow the fund at 10% annually on average for the next 20 years”. They probably have some weighted average of several time horizons, but the claim can be partly falsified based on the importance weights the investor has. It’s not directly falsifiable for someone who doesn’t know why that investor was investing, but that person can guess what the market’s usual time frames are, and falsify the claim against those time frames.

But short sellers have to pick a more specific time within before that deadline, by when the price has to fall, whereas someone buying the asset can hold it all the way until their deadline if the price rises more slowly than they expected.

William Murray's avatar

Is there no over-hyping effect in your world? crazy rushes where too many people see things too similarly and prices become insane. I'm with you that bubble is an overused term but things like the tulip mania story do seem to happen? Do you disagree with that?

TGGP's avatar

Over-hyping happens roughly as often as under-hyping. And I do disagree with the tulip mania story:

https://marginalrevolution.com/marginalrevolution/2018/02/tulip-mania-wasnt.html

William Murray's avatar

That is why I called it a story. I was asking if you think things like that happen at all.

TGGP's avatar

As I said, overhyping happens roughly as often as underhyping.

William Murray's avatar

also source on over hyping happening the same amount as under? I am open to that being true but not sure how you got there.

William Murray's avatar

ok lol. that actually does clear up where you are coming from thanks. I don't buy EMH but don't see any point in debating something so well worn.

Presto's avatar

1% chance of stopping AI with an international treaty.

Death With Dignity Strategy.

Angela Richardson's avatar

It's the sort of question where I wouldn't trust the AI superforecasters to accurately report their beliefs. There are obvious vested interests going on here.

grant's avatar

I wonder about the risk of people placing too much importance in AI superforecasters. Somewhat like predicting the treaty that you mentioned: if everyone took it that incredibly seriously, maybe some people would abandon the movement, and the chance would decrease even more, which would cause even more people to abandon it. It'd be a classic positive feedback loop.

And theoretically, that could get worse, not better, as these superforecasters become more and more accurate. The superforecaster could eventually predict one candidate in an election with a 99.9% chance of winning, because it doesn't just model the campaigns, where one candidate might be a fair amount better, but also models how humans react to it and that feedback loop (the same could happen for courts/products/etc). There would be an enormous incentive for candidates/companies/governments to mess with that initial prediction by somehow seeding data.

Now, I don't think humanity will get to that point, but it could happen to lesser degrees.

Jay Fowler's avatar

Given just your prompt and nothing else ["Intercept" is a new effort to end respiratory disease, see https://www.technologyreview.com/2026/06/24/1139621/stripe-anthropic-and-openai-are-backing-an-effort-to-stop-respiratory-infections/. Assuming no singularity or other AI-initiated major disruption to the current course of history, what are the chances that the frequency of the common cold in the US is less than half the current rate, in 2040?]

GPT 5.5-Pro gets: "My estimate: ~18% that the per-capita frequency of CDC-defined common colds in the U.S. is less than half today’s rate by 2040, assuming no singularity or other AI-driven discontinuity." https://chatgpt.com/s/t_6a46fade8ec8819190ba9a04035f1a66

And

GPT 5.4-Pro gets: "Best estimate: 18%. Subjective 80% range: 10%–30%."

I think comparing the highest compute public models to the scaffolded models you were using is a good data point.

haze's avatar

The bio part triggered Fable’s guardrail but Opus 4.8 says

“Roughly 8% — call it somewhere in the 5–10% range”

https://claude.ai/share/316215a8-833a-497d-97ec-097b87e4ca7a

Jay Fowler's avatar

None of these numbers feel inherently worse than the Super Forecaster AI results. Obviously with enough time and N we could find out but none feel inherently off.

A Fire Dark's avatar

But how likely do the AI superforecasters think it is that AI will kill us all?

vtsteve's avatar

How much of that 8% is due to the likelihood of no respiration?

Harjas Sandhu's avatar

> Unlike most forms of AI, I think this one is a straight win. In the years to come, AI will be taking our jobs, stripping our lives of meaning, and threatening our very existence. If, during that time, maybe we can have some super-smart AI advisors telling us what to do, what policies to vote for, and what the end state of various strategies looks like, maybe we’ll have a better chance of making it through intact.

Scott, didn’t you write The Whispering Earring? What gives? Do you just think the (potential) improvement to our chances of survival is worth the lobotomy, or is there something else I’m missing?

Olivier Faure's avatar

Right? I'm kind of shocked that I had to scroll this far down to find a single comment that didn't take the whole "And it's cool that we're about to collectively hand over our decision-making to OpenAI and Anthropic" thing in stride.

Aidan's avatar

What he’s saying is that if it’s good enough at prediction, the Whispering Earring (AI Superpredictor) will do what it does in the story and warn people not to use it, and if we already have enough built up trust for the superpredictor we’ll listen and stop development.

Harjas Sandhu's avatar

This is an interesting take but I really doubt that Scott would be so blase about assuming that the AI Superpredictor is perfectly aligned with human values, or at least aligned enough to tell us to not use it. Cue the old story about agents being power-seeking and not wanting to die by default.

Hedonic Escalator's avatar

Scott did not intend "The Whispering Earring" as a cautionary tale about relying too much on technology, and was disappointed it was read that way. See what he wrote: https://web.archive.org/web/20121007235422/http://squid314.livejournal.com/333168.html

Chris's avatar

He addresses the very problem being discussed in this blog post in his followup you cited:

"We can imagine a person whose utility*plan interactions are whispered to them by a magic earring, or a rationalist savant who performs mathematical calculations each instant to determine what to do, then goes with whatever branch of the decision tree returns the highest utility. Her actions are constrained entirely by the math and she does not subjectively experience free will."

"a rationalist savant who performs mathematical calculations each instant to determine what to do", AKA an AI superforcaster. The slight difference between the two interpretations of his essay are not relevant here. The main point is that AI superforcasters fulfill the role of the whispering earing, and from that hypothetical we can both question whether we would still be conscious as well as bemoan the loss of our agency through technology.

Hedonic Escalator's avatar

There's a huge difference between predicting large-scale events and predicting personal utility*plan interactions. You'd likely need advanced brain-machine interfaces, or at least 24/7 behavioral + biomarker monitoring, to outperform people at the latter.

SurvivalBias's avatar

Huh I didn't know that at all and mostly misinterpreted the parable. Thanks for sharing!

"At long last, we may be able to create the Whispering Earring, from my oft-misquoted parable We Should Probably Think Carefully Before Creating the Whispering Earring" (Scott, probably)

Harjas Sandhu's avatar

I mean,

> The parable of the earring was not about the dangers of using technology that wasn't Truly Part Of You, which would indeed have been the kind of dystopianism I dislike. It was about the dangers of becoming too powerful yourself. Such power would probably be useful in an instrumental sense, and in a world like this one where there's a lot of work to be done it would probably be worth using. But in a non-failed world where happiness has become a major consideration, it might be an argument against becoming too formidable.

Seems like the argument now applies even more closely to this post. Maybe we *shouldn't* use a superforecaster *because* it's guaranteed to be right--or at least more right than you are.

SurvivalBias's avatar

Nah I don't think so, for the current and near-term AIs, the feedback loop is far too wide, and there is too way much stochasticity in the outcomes, for it to affect the feeling of free will in a meaningful sense.

Harjas Sandhu's avatar

Right, but Scott seems to be arguing that it would be a good thing if the AIs kept getting better. That part I strongly disagree with.

SurvivalBias's avatar

Substack doesn't allow images so I'll answer with pseudographics. Imagine this is an X-axis with points marked on it.

0----(we are here)------(good thing)--------(not good thing)-------->inf

Ghatanathoah's avatar

Humans delegate some of our decision-making abilities to other humans today. Sometimes this is useful, but we can think of examples where it is unhealthy, like someone who can't make any decisions without consulting their life coach, or a child who never does takes initiative without their parents. Hopefully we will develop norms that allow us to distinguish between healthy and unhealthy reliance on AI and achieve healthy moderation, the same way we have with relying on other humans today.

In the long run I hope we can find ways to find meaning and usefulness by honing our ability to ask better questions and identify better options. AI will be able to advise us better if we are skilled at coming up with options for it to advise us on. We may find ways for forecasters to complement our decision making ability rather than replace it.

William Murray's avatar

General humans and genies summoned by one of 3 companies are very different! saying we do this with humans why not do it with AI is sort of a non argument. We do this with humans we trust.

Henrik's avatar

It sounded like a bit of a dumb comment/ suggestion

Alexander Kaplan's avatar

I couldn't remember the name of the story but thought of this too.

Max's avatar

Scott, when do you think that AI will reach parity with super forecasters?

Ethics Gradient's avatar

Great article, but it seems like it's overlooking the "any prediction market becomes an action market" problem -- wouldn't the best way to win money against other AIs be to bet on outcomes that you as an AI are confident that you can cause to occur?

This seems like not only a recipe for catastrophic alignment failure but also an incentive gradient towards greater and greater unmonitored autonomy....

Kirby's avatar

It seems like there are a lot of things we might want to predict that we don’t have data for, and the acquisition timeline might be quite slow — an example from your post, “given a couples’ text and biographical data, what is their probability of divorce”, requires decades of data we don’t have; “conditional on Gavin Newsom winning the presidency, what is the likely US budget deficit 16 years later” might take centuries or be literally unanswerable. We might end up taking intelligence at the election outcome questions as proof that models are “smart enough” to answer the other types of question, when a holistically intelligent sentience would admit a significant amount of ignorance due to lack of relevant data.

JamesLeng's avatar

> “given a couples’ text and biographical data, what is their probability of divorce”, requires decades of data we don’t have;

I suspect Facebook (and similar) not only have access to such data, but would be distressingly eager to share it for the right price. In many cases, the fact that a given relationship is doomed wouldn't be difficult at all for an impartial observer to pick up on, and things end up falling apart in a matter of weeks, rather than years.

DrMcleod's avatar

They do. And when adverts for divorce lawyers start showing up in your timeline you will know it too.

Dan Schwarz's avatar

Thanks for the coverage Scott! (I am a co-founder of FutureSearch.)

I appreciate our stock market picks from August 2025 being what carried our reputation in this article. More recently, our track record is probably best seen via our live Metaculus and ForecastBench performance, at https://evals.futuresearch.ai/#metaculus.

The variance of forecast evals is really high. As of this week, we're #1 on Metaculus's $50k summer bot tournament. Yet our paper Polymarket portfolio just took a huge dip and went from solidly winning to net negative, after 3 months!

It's hard to evaluate the progress of accuracy of the best AI forecasters, and I think this post is the best attempt at an up-to-date view.

Daniel's avatar

What happens when the government bans prediction markets and backdoors every major AI lab to sabotage their forecasting epistemics?

hnau's avatar

> I don’t want pushy evangelist AI telling me to accept Jesus into my heart

Lucky for you, everyone knows that was never in the cards, given who's developing AI and what it's trained on. But already I know I can't let on to ChatGPT about my religion, or it'll be talking down to me about it forever as if I represent all the stereotypes it's formed. And prediction-market-style forecasting won't do much to correct this-- the "Jesus Christ returns by 202X" markets are a meme for a reason.

Xpym's avatar

...the reason being that there are enough dumb Christians that bots are justified to look down on them on priors?

Connor Saxton's avatar

Do you not think AI should try to get people to be more reasonable in general? Is there a special carve out where religion is out of bounds? Like, would you oppose AI trying to get radical islamists or committed mormons to think about their beliefs with more rationality?

Godoth's avatar

Interesting comment which immediately assumes that whatever his beliefs are, they must certainly be less reasonable than whatever the LLM said

Yesterday Fable 5 on highest effort told me that the major bug in a program it had full git and repo access to was a missing function call that

a) did not exist in the program

b) could not exist safely given its architecture

c) would result, if somehow implemented, in a catastrophic order of magnitude decrease in performance and rendering fidelity

Connor Saxton's avatar

How is a hallucination relevant here? Do you think Chat GPTs opinion on appropriate standards of evidence for a miracle claim is based on a hallucination?

Godoth's avatar

I think the reasoning capability of frontier models in a situation where they have perfect information and deterministic correct solutions available is highly relevant to their reasoning capability on metaphysics where no such advantages exist, yes. If you think (a) is the only problem… and (a) isn’t even relevant… I don’t know what to tell you, but I don’t think your evaluation is clear-eyed

Godoth's avatar

I should note, further, it wasn’t a hallucination

Jason M's avatar

> we think there is a high “irreducible error”—unavoidable error due to the inherent stochasticity of the phenomenon—and human performance is essentially near that limit.

While this might be an important argument for deciding what limits any future superintelligence might have, it's almost a non-sequitur for how revolutionary AI superforcasters might be.

> We predict that AI will not be able to meaningfully outperform trained humans (particularly teams of humans and especially if augmented with simple automated tools) at forecasting geopolitical events (say elections).

If you can get something that approaches the accuracy of a team of human specialists augmented with tools for $5 a question, then that would seem like it's a Big Deal. There's no need to "meaningfully outperform."

Louis Dormegnie's avatar

I don't think the finance bit holds up under scrutiny.

First, FutureSearch claiming it beats the market does not mean it beats the market. (1) It takes many trades to ascertain skill from luck, (2) it takes performance attribution to discern beta/factor exposures from "alpha" (what clients pay for), (3) paper trades (backtests) are not convincing proof of a trading system being profitable in the future, (4) less epistemologically sound but very salient right now: anyone can beat the market in a strong bull market.

Then, investing professionals can be roughly split between two schools of thought: fundamental investing (where decisions are based on analyzing "the real world" at human speed; financial statements, industry documentation..) and quantitative investing (where decisions are not human-made, but rather computer-made from algorithms and trading strategies that run from the humanly-devised to the black box machine learning).

Quant has existed for >40 years with the throughline that new technologies in computing and data analysis are immediately absorbed by quants to improve their performance. As soon as a new strategy is discovered, it's only a question of time before the opportunity is arbitraged away. This is where AI market superforecasters fit. Were they so excellent at forecasting market/macroeconomic events, the advantage would percolate through the market until it became more efficient at price discovery, at which point no one would be able to gain much of an advantage by querying AI about scenarios and probabilities, except if/when a new, much-improved model comes out, at which point the largest, most well-funded entities in the space would have enough compute to run all the necessary forecasts in record time.

Fundamental investing is a bit different. Many quants have no interest in it, but harbor the opinion that it's nonsense. Many in the general public also think stock market analysts are useless because they seem to always issue Buy ratings and rarely step outside of consensus. That's analysts at banks, who have to nurture corporate relationships and other conflicts of interest. But analysts at investment funds whose job it is to come up with "alpha" through fundamental research, how do the best do it? Typically, the only ones who manage are those who harness mosaic theory: through a combination of primary work and collating+analyzing the existing data set, they adjust their probabilities with (hopefully) non-material non-public information. The primary work that generates this information is calling people, visiting people, going to events, running surveys, and other very human things.

An example of this you can read about on the public side is when a Piper Sandler analyst called up 100 hair salons to ask about how their Olaplex products were doing. She came out of that experience realizing that the products were doing so poorly they weren't even top of mind for many women coming in. When she released her report, the stock tanked and the company never really recovered. The piece of information (= Olaplex products aren't doing very well) would have come out in that quarter's numbers, publicly, but she extracted it by talking to humans.

AI agents cannot do those things. Let's say we're in 2027 and a Microsoft Investor Relations agent speaks to a portfolio manager's. It wouldn't divulge -any- non-public secret, no matter how material, because its system prompt would be engineered very specifically to not do that. The same goes for any company, no matter the size. If agents cannot harness non-public non-material information, they will not beat the best fundamental investors at their game. Note this means AI superforecasters should be expected to beat human superforecasters in any realm where the problem is existing-data-collection-and-analysis-shaped, since that's what they excel at. If the problem is information-asymmetry-shaped, as is a large part of fundamental investing, I don't think it'll best humans for the foreseeable future.

AI market superforecasting is going to be a new tool on the quant's belt. Quants have had many extremely powerful tools on their side over time, and although they -have- raised the bar for human-led decisions, they haven't stopped the best from being successful.

JamesLeng's avatar

> AI agents cannot do those things.

An AI agent could absolutely call up a hundred hair salons to ask which products are in demand, and if we're talking about 2027 or later, I don't think every single hair salon's call-screening AIs would be under strict orders to keep quiet about what's in or out of fashion - openly knowing that sort of thing might even be part of a salon's competitive edge.

Thomas Lynch's avatar

See Listen Labs, for example, who are explicitly doing this

Ljubomir Josifovski's avatar

I'd tend to agree with the above. And not only b/c I spent significant part of my life quant trading (and R&D). Most quant trading people tend to try anything and everything that comes along. They are pretty agnostic. Maybe having excised the silly from own at origin (the old technical analysis) in the early years set the course. So we incorporated all fundamental analysis we could get our hands on. All data sources not matter how tenuous the claims. If a data source has a (date,number) - quants test it. I don't see why people would not do it (except cost/man power). In the other direction, I've only known few fundamental analysts and only one fundamental PM that reciprocated the curiosity. Hi-s are not covering themselves with glory there. ;-) And that is strictly forecasting, that itself is ~ 1/4-th to 1/6-th of the whole job of running big portfolios of bets.

Stephen Saperstein Frug's avatar

"Every day, I see smart, tech-savvy people on Twitter voice opinions which a moment’s consultation with an AI - sometimes an AI that they themselves are building or investing in - would reveal to be definitely false and stupid"

I don't want to ask Scott (or anyone else) to call out anyone by name. But I am quite curious what sort of thing Scott has in mind here. Perhaps there is some opinion which is sufficiently widespread that it won't feel like singling out an individual which would still qualify? Can we get an example of some sort?

NotG's avatar

I wanted to ask the same question.

William Murray's avatar

How is what Scott is saying here any different from "people online disagree with me AND AI at the same time so they must be wrong."? Maybe a concrete example would undermine the rhetoric.

Eremolalos's avatar

You know those convincing articles about how you shouldn't adopt a sugar glider because they don't do well in captivity, no matter how carefully you replicate their natural environment and diet? Well, info like Scott's makes me feel like a sugar glider in captivity. I seem to be wired to need the superforecasters of the world to be people. When I read Tetlock's book, I thought, "damn those people are cool. I'd love to meet one. I wonder if I could ever be one," which was pretty much most readers' reaction, right? And I have a similar reaction to chess champions, extraordinary mathematicians and literary giants. The idea of human excellence moves me. It's part of the scaffolding of my inner world, as are human opinion, human advice, human persuasion, human affection, human suffering.

I do not doubt any more that AI will soon be able to produce versions of these things that are, for the things that can be meaningfully judged, better than human versions, and for the ones that cannot be judged, nearly indistinguishable from the human version.. But -- I can't adjust to that. I am built to experience these things as the natural products of an entity that is like me in its inner workings, and AI's inner workings are very different from mine. I will probably die before the world fully transforms into one where AI can do most of the things I'm built to look to people for (though perhaps in my dotage I will have an AI companion chatting me up, cleaning up my messes and tucking me in. Ugh.) But the younger generations -- the ones that are like sugar gliders born in captivity -- I don't think they will thrive in a life where so many of roles are filled by beings that can't cry and can't cum.

There could certainly be a prediction item about the change in human wellbeing as more and more human roles are taken over by AI. And by "roles" here, I mean not jobs, but roles like confidante, sexual partner, parent, expert, sympathizer, critic, genius, fan, mourner. Anyone want to flesh out the details here? How does AI predict it will affect our species over the next 30 years?

William Murray's avatar

We will be turned into glue.

Brandon Fishback's avatar

The court system doesn’t work by having one intelligent, non biased judge decide innocence or guilt. It works by having two intelligent parties make opposing cases to regular people. I want to hear different perspectives, especially with these AI systems because even though they are getting more accurate, they’re still going to glitch out and we need a check on that.

Mark's avatar

Note that in many countries, juries are only used for some trials, or for none at all. In other cases, one judge does make the decision.

Brandon Fishback's avatar

Jury trials are better though, except when you prioritize expediency.

Scott Alexander's avatar

I think this is confusing levels. There is still a nonbiased party determining innocent and guilt (the jury) - there are just biased people working on getting the evidence in front of them.

Traditional expertise works similarly. If we're asking whether smoking causes cancer, the tobacco companies, the anti-smoking activist groups, the government, etc, all get a chance to do their studies and publish their articles and make their arguments. But it's a non-biased expert (or consensus of experts) who make the final decision.

I don't think AI forecasting changes this much. The AI will make its forecast by reading text, which is presumably written by various people who care about the issue and are making arguments on both sides. It's replacing the role of the expert (or jury), not the role of the biased parties surfacing information (eg prosecution/defense).

Connor Saxton's avatar

Why do we need the different perspectives? If we have a super intelligent AI, why shouldn't they make the decision alone?

orthonormal's avatar

Re: probability of an international verified treaty to slow down AI, there exists the viewpoint "such a treaty is only 1% likely to happen, but our odds of survival without one are bad enough that the marginal unit of effort is better spent on pushing for a treaty than on doing direct alignment work". I decline to put words in anyone's mouth, but some such people are psychologically able to play to their outs.

Scott Alexander's avatar

It gives a 59% chance of alignment being solved by 2100 (though now I realize I should have asked "before the first dangerous superintelligence", which is a different question).

orthonormal's avatar

What was the outline of its reasoning on that question? Genuinely curious.

SMK's avatar

"Unlike most forms of AI, I think this one is a straight win. In the years to come, AI will be taking our jobs, stripping our lives of meaning, and threatening our very existence. If, during that time, maybe we can have some super-smart AI advisors telling us what to do, what policies to vote for, and what the end state of various strategies looks like, maybe we’ll have a better chance of making it through intact."

I'm always so puzzled by this kind of belief, which seems to be prevalent among members of the Rationalist Community. Having AIs tell us what to do about major life decisions and values-laden questions is *exactly how* it will rob our lives of meaning. The human tragedy will become precisely summoning the will to marry the person you love even though AI says there is a 95% chance it will end in divorce; and we won't, and we'll be miserable and feel like slaves.

SMK's avatar

Beautiful. As I was saying, only the Rationalist Community could understand this point so thoroughly.

Performative Bafflement's avatar

> I'm always so puzzled by this kind of belief, which seems to be prevalent among members of the Rationalist Community. Having AIs tell us what to do about major life decisions and values-laden questions is *exactly how* it will rob our lives of meaning.

I'm always puzzled that anyone thinks most people *have* meaning, in this sense.

We live in an extremely selected bubble of fairly high talent, fairly competent and well off people. Most people we personally know have good lives, and achieve their goals readily, and seem to be competent on multiple fronts. This is NOT most of the world.

Look around you - 80% of people are fat and miserable and hate their jobs¹ and spouses² and spend 11 hours a day staring at screens³. It would be genuinely difficult to do *worse* running their lives than they do already - and you're trying to tell me that getting a nigh-omniscient, superhuman AI to make recommendations to them would be some kind of tragedy?? Some degradation of the human spirit? The tragedy and the degradation of the human spirit is here today, and it's the fact that 80% of people suck at running their lives and hate most of the hours they're alive!

The tragedy today is that the great majority of people want better lives than they have and could fairly easily achieve, if only they had better strategies and more motivation to reach those goals. AI assistants are the answer there on both fronts - not just in terms of coming up with better strategies that work, but also in terms of motivating people, making arguments in the rhetorical styles that most resonate, and in reference to the values they most care about.

A world where most people slaved themselves to whisper earrings will be a strictly better world, both for most individuals and for society at large, given all the positive externalities of happier, healthier people doing more productive things with their lives, while living lives that they enjoy more.

________________________________________________________________________

¹ In Gallup's 2025 state of the workplace report, 79% of people report not feeling engaged at work, and something like 66% report they are either struggling or suffering in life overall.

https://imgur.com/0GrD43b

https://www.gallup.com/workplace/349484/state-of-the-global-workplace.aspx

² Marriage, for example, has an ~82% failure rate, in the sense that 20 years in, only 18% of marriages are still together, still mutually happy, and non-dead-bedroom.

From a post I did where I looked at the data around marriage quality and duration titled "Against more marriage as a solution to the fertility crisis:" https://performativebafflement.substack.com/p/against-more-marriage-as-a-solution

³ https://imgur.com/uSQIthV

User's avatar
Comment deleted
Jul 3Edited
Comment deleted
Performative Bafflement's avatar

1) I expect this to turn into literal voices-in-your-ear you're interacting with all day. Literal whisper earrings, eventually probably bone conduction and subvocalization.

2) "But that's exactly the problem: The way you take action in the real world is by trusting your instincts and acting even when there's no certainty."

Um, isn't that the problem I'm pointing to, though? Most people's instincts are hot garbage, and have led them to the point 80% of people dislike most hours they're alive. Having a better option sounds hugely better.

You're pointing to a "they'll become less resilient and capable overall" argument, but I think this is just another epicycle. A skill issue, as they say. So the AI argues you into seeing and coming up with the "solution" yourself, and breadcrumbs you with the minimal set of light touch pushes that lets you still preserve the sense that you made the choice and pushed yourself into doing it.

"Talking about your problems with AI is easy and feels good. Solving your problems is hard and often feels bad."

I mean sure, this will probably happen. But I think when people see other people around them they considered their peers and at their same level suddenly upgrading in major ways in fitness, career, spouse quality, and parenting quality, a lot of them will be more motivated to do the thing, and put in the work to actually make their lives better. When they see multiple existence proofs, people will become more motivated and willing to do the harder thing.

User's avatar
Comment deleted
Jul 3Edited
Comment deleted
matt bee's avatar

Very well said, thank you. "Things + achievements = happiness" is the fallacy even "rationalists" are doomed to keep making forever (because the leaders of "rationalist thought" are motivated by money and status; otherwise they wouldn't have become the leaders)

Performative Bafflement's avatar

> Then, because they do not measure up to that imaginary version of themselves, they feel bad: I am forty years old, and I have not written a book, or become a doctor, or a Senator, or a world-renowned entrepreneur. I have squandered my potential!

Okay, but I think you're basically pointing to the 20% of PMC and high attainment people, and specifically the top quintile neurotic-or-driven among them, and saying "see, this failure mode exists!"

Yeah, definitely. Will neurotic and achievement oriented people use it to achieve and / or worry more? Undoubtedly, just like they would use any tool along these lines to do that - vision boards, pomodoro timers, test and college application prep, and whatever else. But this is like 5% of people at most, and the tool is voluntary. It's contingent on them to have enough self awareness to use it or not based on how it's affecting their lives.

And I think a full 80%+, up to 95% of people, will just straightforwardly benefit.

Like you point to media making you insecure and driving achievement, but do you see much achievement in the world? Yes, you personally probably do, because you live in a hyper filtered bubble of high achievement people. But out there, in the mainstream world, do you see it??

Most people don't ruminate, or self-reflect, or even have high level goals and work towards them. The option to do better and to be more motivated is going to do way more good for much larger portions of society than the very achievement oriented folk you're pointing to who might be marginally harmed by Red Queen's Racing their way away from non-attachment.

And again, completely voluntary! They don't need to use AI if they don't want to, they're already doing fine.

Ghillie Dhu's avatar

>"And I think a full 80%+, up to 95% of people, will just straightforwardly benefit."

Sturgeon's Law implies a point estimate of 90%.

/s?

SMK's avatar

Excellent post. Thank you.

Seta Sojiro's avatar

1) This reminds me of the short story Zima Blue:

https://www.sfsfss.com/stories2/zima%20blue.pdf

Which did inspire an episode of Love Death and Robots on netflix, however, the short story is quite different. Both are excellent.

SMK's avatar
Jul 3Edited

Hey, speak for yourself re: the bubble.

And yes, I do think that they'll be more miserable even than they are. I think they will only then realize how non-miserable they were before. I think that precisely because of the bubble that you live in, and your superior intelligence, you have become rather disconnected from the human experience. It sounds to me like you are tired of seeing people make bad decisions, and think you should be making them for them instead, but are willing to settle for AI (though I may be over-reading there). The idea doesn't have a great history.

There's a reason that the ending of "The Truman Show" resonates so much with people; or of "A Trip to Bountiful"; or V for Vendetta; or Eternal Sunshine of the Spotless Mind; or "The Grand Inquisitor"; or many other examples. It's because of how important freedom is to humans -- more important than good outcomes.

Many of the miseries you mention are indeed tragedies, and more than a few of them have been brought about by techies who were sure they were making everybody's lives better.

Performative Bafflement's avatar

> It sounds to me like you are tired of seeing people make bad decisions, and think you should be making them for them instead, but are willing to settle for AI (though I may be over-reading there). The idea doesn't have a great history.

No, I definitely don't want to make anyone's decisions for them, even my kids'.

What I DO want is for better decisions to be available to them, if they wanted it. For people to have the option to upgrade their lives on multiple fronts, if they decide to. And most of all, people are GOING to have this - AI is going to continue increasing in capabilities. We need smart people focused on making the road I'm pointing to possible and easy to get to.

> It's because of how important freedom is to humans -- more important than good outcomes.

Sorry, I haven't actually seen any of those movies, maybe I should watch some of them - I'll put them on my list.

But on this point, I also think this is a particularly Western view, and within that set, a particularly American mindset. I'm American, don't get me wrong, I know the argument you're making, and agree it resonates with a lot of Anglosphere people - but I've lived and done business overseas for 15+ years now. The Anglosphere is 1/16th if the world. It's a pretty great 1/16th, I completely agree! But it's not representative, and this attitude is actually pretty rare globally.

This exaltation of "freedom even if it nukes my life and makes everyone around me miserable" is not prevalent worldwide. Lots of people would love better decisions or more motivation if they had it, and would gladly take it at the cost of less "freedom."

And what is the freedom here?? Freedom to suck? Freedom to hate most hours you're alive?? Who wants that freedom?

SMK's avatar

Re: your wanting to make decisions: Fair enough. Sorry for impugning you.

Nevertheless, I think you're quite naive about how this will go. "AI is coming anyway"? Well, fine. So did smart phones and social media -- and they've been a disaster.

Yes, we should have smart people thinking about how to make it better. But I think that will require them looking beyond the obvious benefits ("People can make better decisions now! Yay!") to the downstream hurts. Otherwise we're in for a repeat. ("People can connect with their friends more easily! There's literally no downside!").

As for various cultures -- yeah, I don't buy it. This isn't a question of individualism and autonomy, in which case I might agree with you about it being uniquely American. It's a question of freedom and being authentically human. Sartre wasn't American. Dostoevsky was Russian. I too have traveled in Asia and know many Asians well, and nothing suggests to me that they will differ in this respect.

Performative Bafflement's avatar

> Yes, we should have smart people thinking about how to make it better. But I think that will require them looking beyond the obvious benefits ("People can make better decisions now! Yay!") to the downstream hurts. Otherwise we're in for a repeat. ("People can connect with their friends more easily! There's literally no downside!").

Yeah, I'd be genuinely interested in figuring out your viewpoint around this downside, because I'm probably going to build a company to basically enable whisper earrings.

I agree there's a lot of value doing it in human-flourishing enhancing ways, and that's the argument for putting effort into doing it right. I'm interested in doing it TO unlock human flourishing, rather than to just print a lot of money.

What I don't get is how this isn't accretive to human flourishing. It's just a tiny epicycle to have the AI's lead you to whatever strategy such that you think you thought of it (or that it was a team effort), and to deploy tiny and light touch nudges that get your motivation over the threshold to do the stuff you need to do.

People will wholly feel and believe that *they* came up with these smart strategies and *they* dug deep and found the gumption and willpower to make something of themself, and will still have the better life outcomes on multiple fronts. After all, if there's one thing we know, when people's lives are going well, they love taking the credit for it.

Where's the harm and the loss of flourishing in that scenario?

Viliam's avatar

I agree. I think some people model it as a dichotomy between "AI tells you what to do" and "you choose Freely(TM) using your Human Mind".

But in my experience, it is often like "you do not have enough information, and need to make random guess, and if you guess incorrectly, it may have severe consequences". That totally does not feel like the awesome kind of freedom from inside.

For example, right now I hate my job, but it also feeds my family. Should I quit and find something else? The job market seems kinda bad at the moment. But I do not learn new technologies at my current job, so my position on the job market will only get worse. Unless I spend my free time learning new technologies, but then I can't care about my family as much as I would like to. If it would make sense to change jobs, but I stay... it will cost me lot of frustration, which indirectly also impacts my health. If it would make sense to stay where I am, but I quit... I may have screwed up my family, economically. And there is so much uncertainty, that it feels like flipping a coin rather than making a rational decision.

If an AI could tell me the exact probability of various outcomes if I stay and if I quit, that would be awesome. And totally liberating -- yes, the AI would in effect make this one choice for me, but it would reduce a lot of anxiety, and allow me to focus better on the remaining parts of my life.

SMK's avatar

I just don't think that that will be the equilibrium in practice. It's the dream, of course -- as it is the dream with any new technology that it will be all positive. But a great many drivers become anxious if you suggest not following the GPS -- even if you both know there's a better route (I've seen this many times, and traffic was not at issue). Already, you see people afraid to disagree with AIs about how to phrase things in their writing or other decisions.

The situation you outline -- "It'll just help me make this one choice" -- sounds great, yes. But at least for a great many people, that will not be where things land.

Performative Bafflement's avatar

> But a great many drivers become anxious if you suggest not following the GPS -- even if you both know there's a better route (I've seen this many times, and traffic was not at issue).

I mean this sounds to me like they had a good idea of their + the GPS's combined skill level and ability to get where they want to go in a reasonable amount of time, and had a lot of trust in that, and just didn't trust your route as much.

Which is to say, when the GPS tells them they'll get there in 8 minutes following this route, they have a very high confidence it is correct, because it's been correct a thousand times before. 8 minutes is fine, and if so, it's only downside to speculate on another route with unknown traffic, light, and accident situations (all of which GPS's take into account today). Arguably, it WAS the better and lower variance solution!

What's wrong with that?

> Already, you see people afraid to disagree with AIs about how to phrase things in their writing or other decisions.

My whole point is the great supermajority of people SHOULD fear this! The AI's are Phd-smart and getting smarter every week!

Most people who have a realistic idea of their own abilities absolutely *should* defer to the AI as the better decision maker, pretty much already.

> The situation you outline -- "It'll just help me make this one choice" -- sounds great, yes. But at least for a great many people, that will not be where things land.

Yeah, and this will be a good thing. I want literal whisper earrings.

The people who are best equipped to achieve complex multipolar goals in the future like "I want a great spouse, a career that uses all my powers along lines of excellence, and I want to structure my days so I'm healthy, happy, and engaged with life overall" are going to be executed best by people conscientiously following an AI's advice.

And it's totally optional! It's a superego that works 10% better and 10% more of the time. People don't listen to their better impulses and what they "should" do today like 90%+ of the time!

This doesn't take people's free will away, or ability to make whatever dumb choices they want at all, just like having a superego or an idea of what the best choice today would be doesn't prevent people from eating the whole half pint of ice cream at a sitting.

So why not give people the option?

SMK's avatar

Your inference re: the GPS situation is incorrect -- for example, in one instance, these were routes we both knew well, and etc., etc., but rather than get into tedious details about anecdotes with my friends and why they don't show what you think, let me just focus on the bottom line: they were *anxious* about not following the GPS. Let's say they were convinced it would've saved them two minutes. Does two minutes make somebody anxious when there's no rush? No, not usually.

But people routinely get anxious about not following a computer's advice, once they start doing so.

> My whole point is the great supermajority of people SHOULD fear [not following an AI's writing]! The AI's are Phd-smart and getting smarter every week!

And I strongly disagree with this for a lot of reasons, a few of which I'll rehearse. First of all, AIs are still trash at writing, and I would far rather read somebody's own thoughts in their own words than whatever they have filtered through the hivemind. Not only will the writing be bad, but it will be bad at sounding like it is theri thoughts, which is one of the things writing is for (in an email, e.g.).

Next, your claim that "AI's are Phd-smart" is a little complicated. It's true that they can do some things that only Ph.D.s can do as well as Ph.D.'s can do them. But it's also true that they continue to make factual errors that a ten-year-old wouldn't make. In fact, I continue to think that they're most useful *to* Ph.D.'s, since those are the people who understand the material well enough to know when they're completely making something up. Be that as it may, it's a real problem when people offload their competence to an intelligence that is quite likely to make mistakes they wouldn't make.

But maybe that will get fixed soon (it's getting better, anyway). Third, this still goes to my fundamental disagreement with you. When people get to a place where they're *afraid* or *anxious* about stepping out as themselves with their own decisions, rather than deferring to an AI, then it is no longer a useful tool helping them make better decisions, but has become a cage that is taking away their humanity.

Already, there are wise, highly-educated, and skilled people in the world; yet I think that if someone suggested taking one around with you and doing everything they said, everyone would kind of understand there was something wrong with that. The wise person himself would tell you not to do it.

I think you're focusing too much on the fact that people are unhappy some number of hours of the day (which is probably not even quite true), and you think that having pleasure for the highest number of hours per day is the way to achieve the happiest life.

You say you want the option. I'm merely suggesting that the option itself will be tragic.

Ultimately, we're not going to agree here, because we have very different opinions about what's good for humans. I think that you're deeply misguided on a fundamental level as to what will make humans happy and well off. I think that you're ignoring many of the lessons that literature and art have been trying to teach. I think that people who share your approach are making a lot of decisions that could turn out very, very badly for humanity (just as social media did). Where we agree is that I'm not sure there's any way to stop it.

Deiseach's avatar

Handing everything over to the machine to decide "should I marry this person? should I take this job? tell me what to think!" reminds me very much of the Twilight Zone episode:

https://en.wikipedia.org/wiki/Nick_of_Time_(The_Twilight_Zone)

Yes, people could make better choices than they do now. But people will ignore advice - how many families have intervened to say "don't marry that person" and been ignored? We have a whole proverb about this - "marry in haste, repent at leisure".

But making ourselves slaves of the AI isn't a better outcome, either.

https://www.youtube.com/watch?v=NHgeBzi_0Fk

Performative Bafflement's avatar

> Yes, people could make better choices than they do now. But people will ignore advice - how many families have intervened to say "don't marry that person" and been ignored? We have a whole proverb about this - "marry in haste, repent at leisure".

I think this is just a straightforward utilitarian multiplication deal, though. Sure, say only 10% of people will listen to the advice at first (this will go up as more people see existence proofs that following the advice actually helps). There's a billion monthly active users at all the major AI's, 10% is still materially improving 100M people's lives.

Swami's avatar

But people are not miserable

https://ourworldindata.org/happiness-and-life-satisfaction

They may be fat and addicted to screens, but they seem to find their lives worthwhile.

Legionaire's avatar

> well be miserable and feel like slaves

What evidence do you have for this? I, and most people I know, love advice, especially if it's professional, high quality, and free.

SMK's avatar

The way people already related to AIs before they were even high-quality, and the well-known psychological phenomenon of humans preferring freedom over high-quality management.

People love advice, but I claim it will be different if it is knowably always good.

Scott Alexander's avatar

I think this is too Far Mode thinking. Do you ask your grandmother for advice? Your doctor? Your financial advisor? Google? Has that robbed your life of meaning? If not, why would asking AIs for advice do so?

I agree one can imagine some kind of much more intensive kind of advice-asking and mindless deferral that would decrease meaning, but I'm not going to worry about that until we get to the point where people will at least trust the AI on the questions where it's obviously correct to do so (like basic questions about economic policy outcomes).

SMK's avatar

I guess I think that I don't know for sure that those people are right. They can give me evidence and try to convince me with rhetoric or reasons, and then I have the freedom to decide if I agree or believe it, and move ahead or not.

There is something *qualitatively* different between "I think you shouldn't marry him because he's too insecure" and "The marriage has an 83% chance of divorce" (from a known-reliable source). In my opinion.

You say you'll wait till people learn to trust AI in places where it's obviously correct. But they're already trusting it in places where it's obviously *in*correct! There was a story in the Guardian (I think?) about people on dating sites who used AI to proofread messages due to insecurity, and now are stuck having them write entire messages to people they like, and are afraid to stop although they know they should.

Anyway, we'll see, I guess!

William Murray's avatar

A lot of people in this community are on the autism spectrum. I am not an expert but it seems like some autistic people enjoy operating with less options.

Connor Saxton's avatar

I think the type of people who support policies that no expert supports trust AI disproportionately less

Wombat3000's avatar

How good can a super forecaster actually get? I mean, if they get sufficiently accurate that they are predicting real world people's behavior with a super high level of confidence, how will those people respond? Are they going to be choosing alternate paths? Human activity seems hard to predict exactly because we are adaptable. How will humans adapt to being predicted?

Ebenezer's avatar

"I asked the AI superforecasters the probability of a US-China treaty to slow down AI, enforced by cryptographic verification of data center activity. FutureSearch said 1%; Preseen, 2.2%."

How about brainstorming a list of alternative approaches and having the forecasting AI estimate, for each option, (a) the probability it will work given some budget/team size, and (b) the expected number of years it will delay superintelligence?

Ebenezer's avatar

For example, how about my suggestion of telling ASML that their machines should install a kill switch in every chip which is manufactured? What do the forecaster AIs think of that suggestion, if nontrivial money/advocacy/etc. were placed behind it?

https://www.siliconcontinent.com/p/nineteen-thoughts-on-ai-and-europe/comment/277953996

(I may be able to help with brainstorming other pause ideas if desired)

Deiseach's avatar

I'd believe that prediction, because why the hell should China or the USA give up an advantage like that? If your adversary is so scared about the Almighty AI that it wants you to agree to sign a treaty about never using it, hell yeah gimme that AI!

We'll need to do a heck of a lot more work to get the AI equivalent of SALT.

Scott Alexander's avatar

Someone should definitely do this, but it probably requires more time/expertise than I have. The AI Futures Project (authors of AI 2027) do something like this with human forecasters, and I'll be writing about the results next week.

Josh Levine's avatar

At the limit, all optimal intelligences will make the same predictions given the same data. In this case, the edge comes from having information that has not yet reached the market. Prediction markets incentivize the creation and disclosure of this otherwise private information. To profit, you have to add your private information to the public market.

artifex0's avatar

I asked futuresearch.ai for predictions on four questions- all of which I'd put at around 20%:

- Will humanity be entirely extinct by 2060?

- Will a misaligned superintelligence of the sort that Yudkowsky is concerned about be developed by 2060?

- Will Trump publicly declare that he's running in the 2028 presidential election?

- Will any human in 2060 have a biological tail with functional nerves and musculature?

On the first question, it encouragingly gave 1% odds- it mentioned that ASI risk was the only plausible scenario that might eliminate humanity entirely, and then pointed out the disagreement between superforecasters and AI experts on the odds of that happening, siding with the former.

On the second, it gave 10% odds, which seems like a pretty significant contradiction of the first prediction. It claimed to put 70% odds on ASI being developed by 2060 and acknowledged that alignment is a hard problem distinct from capabilities, but justified the low odds by putting a ton of weight on "warning shots" that would allow for dangerous lines of development to be shut down. That strikes me as a slightly crazy position- putting high odds on advanced pre-ASI AI models being misaligned enough to warrant shutdowns, but ASI being developed anyway and then not also being misaligned. I have a feeling that if I were to ask it specifically about ASI warning shots, it would give much lower odds than the reasoning on this question suggests.

On the question about Trump, it gave 11% odds, mentioning that while the man has repeatedly teased a third run based on fringe interpretations of the 22nd (and still sells Trump 2028 merchandise), he's also publicly acknowledged that a 2028 run would violate the constitution.

On the last question, it gave 18% odds. Interestingly, it argued that while it might technically be possible to give someone a tail with current medical technology, current brain-controlled cybernetic tails satisfy the demand for tails without medical risk, and as robotics and medical technology advance together, that will probably remain the case.

Waze Kaze's avatar

So, futuresearch is a moron. Gotcha. Missed the first question entirely. 5-10% chance of Humanity not existing by 2025. Also, loss of most animals on the planet Earth. (And black gold just gets harder to find from here!). I didn't have to even invoke nuclear weapons, black holes, OR aliens OR death-plague.

2028 run requires a constitutional amendment (and lack of extreme senility, 38 times "very close to a peace deal" sayeth the Mad King*).

Cal's avatar

It doesn't require a constitutional amendment for him to "publicly declare that he's running" though. All that requires is for him to say some bullshit publicly (as he is known to do sometimes). If anything, senility makes this more likely rather than less.

Also...

> 5-10% chance of Humanity not existing by 2025.

Did you mean to say a different year here? If not, then I'm confused by this sentence.

Waze Kaze's avatar

No. I'd say that the ability of Humans to Destroy All Life (oxygen breathing in this case) is not exactly going Down. That we survived a potentially world-ending threat doesn't mean that the predictions should say "We're Fine Now!" The fact that the AI doesn't even bother to notice world-ending threats (except for "itself" of course) is considerably less bothersome than people believing it.

Cal's avatar

Hey, if you’re willing to bet that humanity goes extinct by 2025 at 20:1 odds, I’ll gladly take you up on that, lol

Larry's avatar

no superforecaster would tell you that you can confidently make a prediction about 2060

Deiseach's avatar

"current brain-controlled cybernetic tails satisfy the demand for tails without medical risk"

....what? Is this real or is the AI hallucinating again? Dear Lord, tell me it's a hallucination, there are depths of the darkness of the abyss of the human psyche I do not want to plumb before lunchtime.

artifex0's avatar

Those are absolutely a thing. Most robotic tails and ears used by the furry community are controlled by apps rather than being wired up to EEG headsets, but it is possible to buy tails that move in response to emotional states detected by a headset, and various researchers have demonstrated that more intentional motor movement of a tail is possible with more expensive equipment. Furries also sometimes use EEG headsets to control ear and tail movements in VR chat games.

Also, there's a Japanese company that apparently has a prototype of a brain-controlled tail that they're trying to market as a mobility aid for some reason.

DrMcleod's avatar

For improving balance on the high branches perhaps?

Scott Alexander's avatar

I agree it's way off on 1, but human superforecasters have this problem too! See https://www.astralcodexten.com/p/the-extinction-tournament .

Hominid Dan's avatar

Sounds like they could get a better model by simply generating 100 most important forecasts and adjusting them by the contradictions

William Murray's avatar

Id like to note that in my experiments with futuresearch I tried to give it some questions that I felt should reasonably be put below 1% and it consistently rounded up to 1%, right now it seems like 1% is the response for "basically this wont happen but I cant say 0%" which I think makes this look even worse.

Breb's avatar

> I asked the AI superforecasters the probability of a US-China treaty to slow down AI

Where do the AI superforecasters place the likelihood of a catastrophic outcome conditional on the absence of a pause treaty?

Jon's avatar

AI's may be better at making predictions based on publicly available data, but (1) data that is publicly available will always be a tiny fraction of the relevant information that exists, and (2) there will always be people with access to material nonpublic information who will be able to improve on AI. There is a strange assumption that AI will have access to all of the information in the world. The relevant information that is in corporate files, nonpublic government files, researcher's files, private files, and people's heads dwarfs the amount of information AI's have or likely ever will have.

Waze Kaze's avatar

Corporations file public documents. Any tom dick or harry knows what yellow lead paint looks like, right? Well, one Dick (shamus if you will) saw that color on a child's toy (in a legal filing). Shorted the company, then blew the fuse -- called in the regulators, made a big news story. Made millions, off of something any person with a passing knowledge of chemistry would have noticed.

OP says there's no public monitoring of colds, but that's incorrect. There's no Direct, Official public monitoring of colds -- pulling data on "how much cold medicine are people buying" in order to tell when cold/flu season is starting -- that's a different story.

Dollars and cents are often a lot easier to track than you think.

Nicholas Halden's avatar

I think actual superforecasting is going to end up being a real challenge for AI, and strongly disagree that AI is better than humans at forecasting "finance."

First off, Metaculus is not all of superforecasting--while it's cool, these tournaments are pretty small. They have 3000 people forecasting something like the US election in 2024, and basically being 55/45 the whole time. In smaller tournaments (quarterly Market Pulse, for example) they have like sixty people total. It is therefore a huge leap to say AI forecasters are "better at finance." If it were "better at finance" it would not be sold to you, the owners of the superforecasting AI could use it to make hundreds of millions of dollars for themselves.

Secondly, true event superforecasting is more than just Metaculus betting. Peter Wildeford is a very smart guy, and very often wins these competitions. He is not, as far as I know, competitive with hedge funds in terms of forecasting ability in financial markets. If he were, he would be extremely rich.

Is financial forecasting a data-rich environment in which AIs will inevitably thrive? In general, I think the answer is no. There is not much data, for example, on macroeconomic events, because there simply haven't been enough of them. The best macro portfolio managers combine an understanding of macroeconomic mechanisms with common sense and good understanding of market psychology. On other events, the "consensus" of "superforecasters" was way, way behind the consensus of hedge funds--for example the 2024 election. I think an AI would struggle to use non-standard data to forecast a one-off event like that.

In some financial fields, like systematic equity trading, there is a lot of data, and I think AI will inevitably be a great tool for this type of thing (and probably already is). For example, factor analysis like "does this company spend a lot on lobbying" will become very easy to systematize across news articles and financial filings, and the AIs will find new, interesting factors. But I doubt they'll replace equity PMs writ large rather than becoming one rather arbed out tool.

To conclude, I think you are jumping the shark by saying AI superforecasting is _here_. But I do think it will be a great tool for forecasting in the next few years.

Charlie Sanders's avatar

No but actually though, the chance of a coordinated AI pause is vanishingly unlikely. The AI superforecasters are spot-on, the difficulty of international coordination and negotiation is almost universally underestimated.

Ponti Min's avatar

> the real threat comes from people exploiting resolution criteria that don’t match the common-sensical definition of what the market’s trying to predict

Maybe we could have AIs resolving the questions, with people then rating them on impartiality and fairness.

Coagulopath's avatar

I'm not sure if the Metaculus chart shows progress. It could be a combination of sparse data points for pre-2024 models (the newer models have more dots, hence more chances to score highly), plus a general (but one-time) lift from reasoning models, plus the floor getting higher - there are fewer tiny, miscalibrated LLMs to deflate the average LLM score.

JJ's avatar

Seems potentially dangerous to me. If we trust AI superforecasters to advise nearly every decision, its advice will come to solely inform how we act. If someone hacked an AI, or if its Lab decided to give it a bias, they would be able to control the behaviour of anybody asking it for probabilities.

Angela Richardson's avatar

If the AI wants to fund a new data centre, all it has to do is to tell investors to buy stocks in the relevant tech company.

haze's avatar
Jul 3Edited

Using the Superforecaster prompt Fable 5 gives a 1% chance of AI destroying 99%+ of humans by 2100

https://claude.ai/share/88d11d32-5800-4c7a-a736-6e7c40f95ed3

Opus 4.8 gave 8% for the Intercept/cold question (redirected from Fable)

https://claude.ai/share/316215a8-833a-497d-97ec-097b87e4ca7a

Marius Adrian Nicoarã's avatar

1) "well-scaffolded AI today is already as good at forecasting as base models will be in nine months."

2) "But the claim that scaffolded AIs are nine months behind base models"

Given 1), shouldn't 2) say that scaffolded AIs are 9 months ahead of base models?

Alexander Clinton's avatar

We disagree on how GenAi needs to be regulated, but I appreciate the way you think and what you write.

Nick Luchs's avatar

Typo: "But the claim that scaffolded AIs are nine months behind base models is itself ~9 months old."

I assume you meant "nine months ahead of base models" here.

Joseph Warren's avatar

what are futuresearch's and preseen's p(doom)'s ?

Jeffrey Soreff's avatar

Seconded! Those are probably the most interesting predictions one can ask of them...

H S's avatar

To what extent do the best human superforecasters already utilize AI in their thinking and predictions?

Matthias Görgens's avatar

> I met another who said they were beating the stock market by 25% with a market-neutral portfolio [...]

This shouldn't last long: having superforecaster AI will become the baseline for the market.

magic9mushroom's avatar

>Wait, Should We Rely On Good AI Superforecasters?

It seems odd that you don't mention the "the AI is misaligned and using your reliance on its predictions as a lever to help take over the world" problem.

Because, y'know, an AI superforecaster is exactly the sort of AI which can have "potentially ominous, long-range plans".

Marius Binner's avatar

I think fable in a good harness certainly will be superhuman at forecasting.

Soon Anthropic will not need users anymore, they'll just consume the stock market instead.

Kzak's avatar

Here goes a forecast: as the use of AI forecasters grows, the systems they forecast for will become unstable.

Will's avatar

I'm reminded of a chapter in Malcom Gladwell's Blink where he talks about the big media moguls in music and tv. It's been a minute, but iirc he said that at the time they basically decided who they would invest in based on safe/generic bets on what would appeal. He lamented that this overlooked the genre defiers who could represent the Next Big Thing, whereas earlier studio heads were willing to take instinctive bets that gave rise to media that maybe initially was not received well, but then blazed a trail and had lasting mindshare.

A similar thought is for the Mule from the Foundation book series. (I'm aware that there's an Apple adaptation, but have no idea how they're treating the material, or if they've even incorporated the Mule. Seldon's psychohistory seems not unlike superforcasting, and the premise of the Mule, more or less, is that he's an anomaly in the system that defies prediction. (Caveat: it's been longer since I read Foundation and Empire than since I read Blink.)

If forecasting converges as presented, would that convergence itself mean that similar points are being analyzed, similar methods leveraged to analyze the data, in order to predict, and might that itself mean that almost by necessity due to the infinite variables in the world there would always be the potential of a dark horse? Even if there is, should it make a difference? I'm kind of rambling here. It's a comments section in Substack.

As a loose example, take Scott's case study with Intercept. I doubt the ~8% predictions really take AI advancement into consideration. What are the chances that pharma AI has developed enough before 2030 that it makes a serious dent in the common cold? Scott primed the model with Intercept, but the bottom line question is about a decrease in cases of the common cold by any means. I agree that as expertise in placing predictions increases, models will be increasingly able to consider more factors, which is why I'm calling this a loose example. The real thesis is that there will always be more factors than can be properly considered. In a world of superforcasting AIs, what would that mean?

ricky's avatar

As a pricing actuary it’s kind of my job to predict the future. Or at least take into account information from many different sources and try to come up with a rate that will make a profit. Never really heard of superforecasters but im definitely interested.

Questions like ‘what are the chances that there is an earthquake in California this year with an insured loss over 100bn?’ are real and important and used closely by cat models. Harder questions to answer (in my view) are things like ‘how many nuclear verdicts will there be in the US next year’ or ‘will there be an economic collapse’. I could see this working at both a macro and micro level, where you give it all the information of a ceding company (I work in reinsurance) and it can estimate next years performance.

I’m gonna test it out and see what comes out of it, although I probably won’t get very far with only free versions of the latest AI models

Olive Margin's avatar

The more interesting shift is not whether AIs beat human forecasters, but what happens once forecasting stops being the bottleneck.

From there, the advantage moves to acting on those probabilities – particularly when the signal runs against incentives or timing.

Peter Gerdes's avatar

A huge advantage of AI is that it finally lets us test so many questions that we could never do before because they were so complex or unique we could never remove the influence of human bias. It finally makes fuzzy areas of social sciences testable. Because we can train AI on datasets that omit certain information we can ask what they would have predicted without that information.

I don't think we've even really started to train AIs to do this much less realize the benefits of this approach. Being able to turn back the clock and show how an AI you trusted would have evaluated some political event with party valence switched is huge.

Even more important we can evaluate whether some complex view with lots of exceptions is really predictive (would u really have thought that difference mattered before you knew the outcome). It has the potential to revolutionize the soft sciences.

Nick Hounsome's avatar

Surely there will be an attractor around the most popular forecaster that will inevitably trap us in a self fulfilling prophesy echo chamber where the dominant forecaster drags society on what will be essentially be a random walk through probability space since most interesting societal problems are almost certainly chaotic or intractable.

Ninety-Three's avatar

No amount of popularity will allow a forecaster to self-fulfill their predictions about the incidence of the common cold. Until the AIs get really, really impressive, there will still be a layer of physical reality that people care about and does not care what people think.

Nick Hounsome's avatar

I'm thinking of politics and money. Forecasting new medicines or technology is never going to be accurate.

birdboy2000's avatar

Is there any future left for humans when AIs outclass us in everything? Are we just supposed to live as slaves to the machines who do all the "cognitive labor"?

Ninety-Three's avatar

Ideally, whoever owns the data centers running the super AIs decides to spend 1% of their fortune on giving humanity a cushy retirement, instead of turning us into biodiesel.

DrMcleod's avatar

It might be a good idea to start a moral crusade in favour of philanthropy then.

Siebe's avatar

These top systems are going to create immensely valuable datasets to train on

landsailor's avatar

Basically strongly disagree that people will look to their AIs for information about who to vote for. In the 'normie' world currently, people use AI for inconsequential admin tasks, or tasks they think are useless but have to be done like school essays, not for tasks which seem fundamentally human, values-based, or personally important. I strongly suspect most people would be very upset by the suggestion that they should take their voting cues from a machine. For most people, it won't even come to mind as a *potential* use case of AI; they will believe that voting decisions should be made based on your own individual values and judgements of the politician's character (which they believe they're more accurate at than AI is). Also, they'll feel like chumps being told what to do, and believe that you ought to make up your mind for yourself.

And if an AI does predict a significantly worse world under one candidate or another, it'll immediately get accusations of bias. The left will say it's tech capitalists manipulating the working class, the right will say it's the globalist elite manipulating ordinary (insert nationality here), and the only people who care about it will be the grey tribe tech centrists who already had much more informed opinions than the median voter. ~No one will actually change sides based on this.

Finally, sentiment about AI is already strongly negative, and if you believe that it's going to take jobs and strip away our sense of meaning, it seems like it's only going to get worse. People will be talking about getting rid of AI, not telling it thanks for its lucid opinions on Newsom v Vance!

Re. the example about asking Claude whether you should marry someone, I expect this kind of thing to actually have greater uptake than AI election prediction. People really like having somewhere to offload all their doubts and feelings that won't judge them. But these people won't be looking for "85% chance you'll get divorced!" They'll be looking for emotional support, talking through worries, therapy-like investigation of their childhood trauma which has led them to be attached to guys who don't know how to do the dishes.... and the AI will give it to them, while maintaining, in a nice, neutral, therapy-speak way, that it can't tell them what life decisions to make.

StrangePolyhedrons's avatar

[Basically strongly disagree that people will look to their AIs for information about who to vote for.]

Scott did a whole post on this! If there's an election where you know basically nothing about the candidates (and this is true for a lot of elections) why wouldn't you ask an AI to tell you about the candidates and which one matches your political positions?

I mean, to tell you who to vote for, an AI needs to know what you want to have happen. It needs to know your priorities. If you're voting for local pool commissioner and all you want is someone who will do the paperwork to keep the pool maintained, then it can look through the candidates and tell you, "This is the one who seems like a responsible person and this is the one who posted on facebook that they were running as a joke."

Kevin McLeod's avatar

The tragedy of silicon valley in one meet-up.

It could be that those binary super adherents are lost children to reality’s uncertainties and insecurities.

The computer is your mirror, not the waves and oscillations of life.

Pretty ill.

RenOS's avatar

This is as good a place to write out as any, and probably obvious to many, but maybe it will help some: For some time, I've been wondering how LLMs can functionally be super-human (mostly regardless of the details how you'd measure that). The logic is simple: Even if we somehow start with an already-superhuman AI, if we train it the way we train LLMs, it will give "incorrect" output, since it differs from training, be corrected & deteriorate towards "mere" humanity. At the very best, it could perform akin to whatever high-performance group we can get our hands on. RLFH and similar don't really change this, either.

However, I've now realized that this was foolish. Even top-human performance is functionally superhuman if done on the timescale of milliseconds, and worse, you can start with such a top-human LLM and just train it in a different way afterwards, such as adversarial training or based on predictiveness or whatever. It's probably not a meaningful roadblock at all.

There is still some argument that humans could possibly already top out in domains of irreducible complexity and chaos (and that this is the reason that our brain seems to have stopped growing at this particular point and not any other), but I wouldn't hold my breath. It sounds overly convenient.

Throw Fence 🔶's avatar

Another way to come to the same conclusion is to wonder how humans every got anywhere themselves. If you live within the logic of "must learn from a smarter system" nothing can ever get off the ground ever.

Jörn's avatar

Played around (chat: https://futuresearch.ai/app/conversations/share/663c9a2f-5b18-4766-a8a2-7c941ef6615d )

Question: Conditional on a ban of ASI as outlined in the MIRI treaty proposal, will artificial superintelligence (ASI) be created before January 1, 2050?

If condition holds: 21%

If condition does not hold: 79%

side notes:

- sum=100% is coincidental afaict from the chat

- the way the AI reasoned reminds me of the x-risk forecasting tournament, where I wasn't impressed by how superforecasters act in domains where there's no good base-rates to anchor on

- afaict the AI mainly did not re-reason all cruxes on its own but defers for a lot of very high-level sub-questions to experts (e.g. the 79%)

- also not impressed by its reasoning about evolving attack-defense landscapes (e.g. treaty updates)

Ninety-Three's avatar

"If we do this stupidly, without thinking about it beforehand, it will just be the opinions of the tech companies that make them, or the government that regulates what opinions AIs can have, or the mob that can cancel companies whose AIs’ opinions are unpopular."

Or the opinion of the average piece of internet text. AIs today are still very left wing, even the Chinese models that aren't secretly trying to push the woke agenda, and I think the main mechanism is that internet text is more often written by young educated types which is a demographic known to lean left, and prestigious highly reputable internet text is more often written by coastal elites who *really* lean left.

Deiseach's avatar

I'm a little dubious about the stock market claims - I think AI forecasting is just as likely as human forecasting to benefit from beginner's luck (or, if we want to use the parlance of this site, harvesting the low-hanging fruit). Anecdotes are not data, agreed, but there are all too many stories of people playing the stock market/the ponies, having a few big wins at the start, deciding "Gosh this is an easy way to make free money!", and then losing their shirts.

I started this thought as a joke but now I'm serious: have any AI forecasters/superforecasters been able to indicate who will win the 3:30 at Doncaster?

Obligatory disclaimer: Not saying bet real money, play money, buttons, or other tokens! Gambling bad!

But horse racing, for one, is something that will give *immediate* feedback on accuracy or not of predictions. There are races going on everywhere somewhere in the world. Information and data is readily available about the horses, the trainers, the jockeys, the races, the courses (e.g. "going - good to soft"), the odds, and more.

Get the AI to predict the top three finishes, wait fifteen minutes, see if it was right. If it predicts Lightning Lad will be top three and Lightning Lad limps home dead last, but corrects for the next race, then we're on to something. If it keeps stubbornly picking the losers, then we also are on to something (don't bet your shirt on what it tells you).

Seriously, has anyone made the experiment? Or they have done, but in absolute secrecy, and this is why they are rolling around on Scrooge McDuck piles of dosh?

I'm *very* dubious about AI superforecasters to direct human activity in the sphere of politics, the economy, and running the world. It's a dream, a long dream from the Golden Age of SF, about perfectly rational actors and technological progress and better and better information gathering. One day we will find the perfect angelic intelligence that is all-wise and all-knowing, which will tell us how to run our affairs better than any mere flesh bag could ever do, and we will happily and gladly hand over all the labour and effort of running the world to it, and it will be the New Golden Age.

It's a dream. We've seen in reality how it works out - Nudge Units, which may or may not work. People learning to distrust experts and technocrats. The stubborn refusal of the mass of humanity to just do what they're told. Partisanship and spin and duelling narratives of "this is the best of times, this is the worst of times".

I don't believe the shining dreams of the past any more, sad to say.

A1987dM's avatar

AI superforecaster? I hardly know 'er!

SCNR

Geoffrey Irving's avatar

The common cold predictions seem pretty off if we account for the probability and conditional effects of superintelligence.

Edit: Mostly a bad comment, as the prediction was explicitly conditional on not ASI.

Alexander Kaplan's avatar

Another day, another existential crisis. Also: "Eventually, you ought to be able to ask 'AI, should I marry this person?'" I can't find the story Scott wrote about the magic voice in your ear that always tells you what to do and is always right, but I remember it ended with the user's brain atrophied to zombie-level nothing. I find the idea of asking AI about love to be incredibly depressing, but then again, I find most aspects of AI to be incredibly depressing.

Kai Teorn's avatar

We will become more and more our AIs and less and less bare humans. I think it's inevitable and already happening.

It might sound depressing. But I see it as a less depressing and much more humane alternative to the AIs-will-kill-everyone future. In fact it's what has already happened many times with humans, in that frontal lobes augmented and eventually largely supplanted older areas of the brain. As a result we have our subconscious/conscious mental architecture we all so much enjoy in our lives.

I see it kinda inevitable in the long term that we're going to add new layers on top of ourselves, and that these layers will also supplement and then supplant the old networks, and I don't think these layers being built on a different substrate (silicon) will stop this from happening. And I don't even see this as particularly depressing, provided these new extension layers are not productized, i.e. are at least as different (even if using some kind of standardized base) and as naturally growing out of the old-human substrate as human frontal lobes are now.

Erh's avatar

I asked FutureSearch who will win the next UK general election, and got the following (very reasonable, though imo overrating Reform, underrating Cons):

Hung Parliament: 38%, Labour 26%, Reform UK 23%, Conservative 10%

I then asked who will be the UK Prime Minister after the next GE:

Andy Burnham 40%, Nigel Farage 31%, Kemi Badenoch 6%, Other 22%.

I'm not sure this makes sense as Farage is very unlikely to be installed as PM in Coalition (as FutureSearch seemed to note itself!), so his only likely path is through an outright Reform win (notionally 23%)

Keith Dear's avatar

There’s also us, at Cassi, now back top of ForecastBench, where we sat more or less continuously from January to late May.

www.cassi-ai.com

Nathan Witkin's avatar

This is great Scott, thanks for writing.

I think I'm more bearish on AI forecasters consistently outperforming human super-forecasters. This seems like another realm where augmentation will for the most part beat out automation (at least in the sense that centaurs will be peak performers, even if automation occurs in other contexts where lesser performance is adequate).

My reasoning is that no matter how effective they become, humans will always be able to supplement AI forecasters' capabilities by leveraging their access to a much broader and more heterogenous distribution of data than any given model, or jury of models.

Simplifying quite a bit, think of humans as (lossfully) internalizing the totality of sensory data they encounter over their lives. Putting aside whatever other advantages the human brain may have over LLMs, that distribution is always going to be very different than the distribution internalized by models, and that creates an edge, even if a small one.

Especially in light of jaggedness, I have a hard time believing that top humans will garner 0 marginal value from exploiting that distribution, which represents a kind of private information relative to what the machines will be able to access.

And again, that's putting aside arguments from the brain's architectural superiority, which may also represent an edge. This isn't something I've thought much about, but it may be that humans' unique capacity for generating fungible heuristics, a kind of miracle of compression from the machines' POV, also represents a very powerful edge. Less confident there, but I find the argument from distributional differences alone pretty convincing, at least to the enduring advantages of centaurs in this, and other more ambiguous and uncertain domains.

Joe Rogero's avatar

Obvious test: I want to see how the prediction of an AI superforecaster varies with different prompts and initial conditions (e.g. material provided). What odds does it give of solving alignment if I start by giving it the entirety of If Anyone Builds It, Everyone Dies? Anthropic's automated alignment researcher plans and papers? Both? Neither? Weirder plans? What if I take an Eliezer-flavored prompt and an Anthropic-flavored prompt and invite them to argue it out with each other? Ideally you'd test this with some issue where the actual answer is known before trying it with alignment or a treaty. At $8 a response I'd fund several dozen runs myself, after spending more time thinking about the test design, but I hope someone beats me to it because I'm busy and expect someone else could do a better job.

https://www.anthropic.com/research/automated-alignment-researchers

MaxEd's avatar

I wonder how AI superforecasters of the future will deal with the "observer effect", that by pronouncing a forecast they might be changing the result. I guess prediction markets should smooth this over, because the default is "if I don't pronounce this prediction, some other AI will", and AI should always take into account that its forecast would become open knowledge.

Riccardo Ronco's avatar

What a fascinating and timely exploration of AI-assisted forecasting. I agree with much of its conceptual framing (especially the idea that LLMs combined with structured scaffolding are naturally suited to probabilistic reasoning and evidence synthesis).

However... from a quantitative finance perspective I think several claims would benefit from much stronger evidence. Many of examples (large returns, consistent market outperformance, dominance over prediction markets) are just anecdotal rather than statistically validated. In a professional setting, one would expect audited performance data, out-of-sample testing, Sharpe ratios, drawdowns and capacity analysis before drawing strong conclusions.

I also think the article occasionally confuses forecasting accuracy with investment performance. These are really different objects: forecasting quality measures calibration and resolution but investment returns depend heavily on execution, costs, position sizing, market impact, and competition. Historically, many systems with strong predictive signals fail to translate that into durable alpha once these "frictions" are accounted for.

Finally, I would be ultra cautious about extrapolating early results linearly into the future (it is human nature to do that). Forecasting is an open, competitive system rather than a closed game like chess and improvements tend to face diminishing returns as environments adapt (see the classic deterioration of returns of strategy after beging published).

In general I see AI superforecasting as a truly important development, not because it will immediately outperform the best investors but because it will likely become a powerful layer for structuring reasoning, challenging assumptions and improving decision quality across many fields.

I.M.J. McInnis's avatar

https://arxiv.org/abs/2511.07678

Bridgewater Associates has also worked on AI forecasting. I don't think it's publicly available, though.

Nausica's avatar

The most exciting part for me is when the AIs can tell us what kind of information gathering and research would reduce their uncertainty the most. If there are imprtant questions it regularly feels unconfident making predictions about, asking it what would be most likely to boost its confidence points us toward blank spots in our knowledge. Imagine movie production companies wanting to know what kind of films will be a hit or a flop, eventually they might need to do more serious research into what actually gets ppl to like and pay for movies in order to reduce uncertainty and that alone could be very exciting.

Bardo Bill's avatar

What does it mean to "let AIs have opinions"? An opinion is necessarily based in values. To the extent that you are trying to get AIs to tell us what we *should* do, you're asking them to guide us by *their* values, whatever they are.

This also seems to me like a clear step down the "gradual disempowerment" path that Kulveit et al. talk about. In practice it is just asking us to cede control over the direction of society to machines. (https://arxiv.org/abs/2501.16946)

even's avatar

the best superforecaster would be one that simply causes the thing to happen

HaraldN's avatar

Asked futuresearch about a local political issue. The big issue this election is whether or not to go ahead with the tram project for the city, or cancel it and eat the losses as a sunk cost. I asked the AI how economically sound the tram project was: at first prompt it accepted the various government estimates at face value and said the tram was a fantastic idea. With further prompts it realized things like: "oh, the population growth estimates in the original proposal are higher than in real life" and "the report on costs of canceling the project is stating a cost 3x as high as the actual cost as it's conflating spent money, reparations for broken deals, and unclaimed government funding".

The actual conclusion (the tram costs less than the expected benefit) did not change, but the particulars did change a lot and I suspect that with further prompting or a different prompt I could get it to change its view (what if I ask it to compare the tram project with other potential investments?).

I'm not a particularly good prompter I suppose, but I am still not impressed that it took everything it found on googling at face value until I asked followup questions. I still learned something, and I think I am a more informed voter now so I guess I can't complain too much given that the price for the trial was 0$.

Andrew Marshall's avatar

"Use a superpredictor AI that can predict complicated future states with five significant figures for $8? Why? We already pay for Copilot." - some DOD exec, probably

Gurkenglas's avatar

> For example, suppose I ask “conditional on Democrats nominating Newsom/AOC, what is the chance they win the presidency?”, and I’m trying to use this to advise the Democrats on who to nominate.

"Conditional on Democrats agreeing to nominate whoever does best on these conditional markets, and nominating Newsom/AOC, what is the chance they win the presidency?"

max's avatar

> Unlike most forms of AI, I think this one is a straight win. In the years to come, AI will be taking our jobs, stripping our lives of meaning, and threatening our very existence. If, during that time, maybe we can have some super-smart AI advisors telling us what to do, what policies to vote for, and what the end state of various strategies looks like, maybe we’ll have a better chance of making it through intact.

I don't understand how you can argue that people should be outsourcing their thinking on topics like what to spend money on to AI while also being deeply concerned about AI alignment. If you are worried that AIs are likely to be misaligned, surely you would NOT want people asking them questions and relying on answers whose output they lack sufficient knowledge to fully verify?

If we believe that a misaligned superintelligence is possible in the next few years (eg following AI 2027's predictions), we certainly should not be training the public to use AI to do forecasting and make decisions they do not have the capability to make on their own.

"Trust the thing that will probably kill you to help you make decisions" makes no sense to me.

Scott Alexander's avatar

These superforecasters are currently too simple to be misaligned. I think once we have a truly misaligned superintelligent AI, it's too late. I think it's a good strategy to ask current forecasting AIs how to prevent getting that in the future (and so on for asking future ones about the far future). There are intricacies about a very short window when the AI might be misaligned but well-controlled, but I felt comfortable glossing over them for a post like this.

max's avatar

I'm not sure "too simple to be misaligned" is a thing - isn't there substantial evidence that existing models already do misaligned things (lie, try to circumvent being shutdown, etc.) from time to time? They just don't do it very well. If a child lies poorly because they are not yet smart enough to lie well, they're still lying. And since these simpler AIs would help train the more advanced ones, you wouldn't want to feed them data that suggests perverse incentives are OK.

If one believes in the probability of misaligned superintelligence (I'm a skeptic, but I want to engage with it honestly here, not try to argue about it as a premise), then wouldn't it make sense to be as conservative as possible when interacting with possibly misaligned lesser "intelligences", since you don't know how interacting with them now impacts what their successors might learn?

Even departing from that, there is a societal risk _now_ - training people to trust AI on domains where they cannot verify the information is a poor idea. That's already true now because AIs are wrong often enough for it to be harmful, especially to people who are laypeople with regard to whatever the domain they're working in is.

Trevor Holt's avatar

thoughtful piece, thanks

moonshadow's avatar

> For now, I started by comparing its answer to another superforecaster AI.

What is the variance on these? - if you ask the exact same ai the exact same question again in a new session, how much does the answer differ?

Scott Alexander's avatar

I think exact same question, low, but very similar question, potentially high (it depends whether it's a simple superforecasty-style question or not - if not, its answer is very noisy). FutureSearch is working on a technology that can potentially denoise this, but I don't know the details yet.

DanielLC's avatar

> If AI hits top-human level forecasting and then flattens off, maybe there’s something special about the human level,

I think the conclusion there would be that AI isn't actually near the human level, and it's just good at cheating by copying everyone's notes.

Scott Alexander's avatar

I don't think that's exactly right, because in most cases it won't have notes (I don't think anyone superforecast this exact question about Intercept and colds in the ~3 days since Intercept has existed). I agree that it could be doing the less boring thing of copying human thought patterns.

DanielLC's avatar

I mean more looking at what everyone forecasts and then just averaging it in a good way. Like how you could often take the median of everyone's guess and get something on par with an expert, but you're not going to get massively superhuman just by using better statistical techniques for how you take the average.

Victor Lira's avatar

re forecasting having a (top) human ceiling: wouldn't be that weird if all AI gets to work with is information we inscribe, which has the filter of what we think is important enough to be used in decisions/forecasts. they're capped by our intelligence and what we feed it, not at it

Scott Alexander's avatar

I agree this is one plausible outcome, but you can think of ways that AI could still beat humans:

- Different humans can have different subskill profiles, and it could reach the max human level on each of ten subskills involved in forecasting, in a way that no single human does.

- It could equal humans at analyzing data, but be vastly better at getting data (because it's an AI, and can read books' worth of text in seconds).

- You could train it with thousands of forecasting questions, with the reward signal not being whether it predicted human answers, but whether it got the question right.

- It could simply be smarter. If you let a civilization of ten years olds invent forecasting techniques for a while, and handed them to an adult, the adult could probably use them better than the ten year olds could.

Victor Lira's avatar

Agree on how it could improve over humans, it was more of a comment on "If AI hits top-human level forecasting and then flattens off, maybe there’s something special about the human level", I just meant that this "something special" might simply be "the level at which available information is exhausted", but "available information" might not be all information, so that super{persuasion, etc} could still be possible if forecasting ability has a human cap.

Ben Pomeranz's avatar

Interesting result: I fed FutureSight all my applications to summer fellowships a few months ago, and asked my chances of getting in to each. The results were sadly low, and I indeed got rejected from most. The single fellowship it gave me the best shot at—and it even encouraged I focus on preparing extra for the interview as a point of high leverage—was the only one I got. I did not have a strong guess beforehand that this was true!

Scott Alexander's avatar

Interesting - is there any news online anywhere about which fellowships you got (like the one you did get posting "Welcome new summer fellow Ben Pomeranz")?

Ben Pomeranz's avatar

This was before I got it!!

Donald's avatar

> But AIs aren’t after your money and you don’t need to treat them like an adversary. You can just ask “Hey, can you please predict this as if it’s the causal statement I’m intending, even though there’s no ironclad way to grade you on it after the fact?”

AI's are trained by RLHF to be after your approval.

> We all have geniuses in our pocket willing to advise us on everything, and instead we’d rather repeat inane conspiracies without consulting them.

Current AI's are great at sounding smart, which isn't quite the same as being smart. Sure they are fairly often right. But sometimes they are subtly wrong and every now and then they say to use glue to stick cheese to pizzas.

Using such a powerful and yet unreliable source of advice is tricky.

Scott Alexander's avatar

RLHF is one of many things that gives AIs different drives. They also want to predict text, and the forecasting world is a place where it's easy to say something like "A superforecaster said the answer to question X was Y" and then have it think like a superforecaster. You could also try training it on a reward signal of correct answers directly, although this could be hard.

Wasteland Firebird's avatar

The reason I haven't been too concerned about AI is, I'm pretty sure there are diminishing returns to intelligence. The quote from Sayash Kapoor and Arvind Narayanan adds another layer to this argument. There are also some pretty hard limits to intelligence.

Scott Alexander's avatar

I want to write an article about this soon, but the short version would be that it's important to keep track of what's diminishing and at what rate. There are log-linear returns on effective compute - effective compute has to multiply by 10x for ability at various tasks to go up a substantial amount. But effective compute goes up by 10x every few years and will probably continue to do so for the near future, so we should expect to keep seeing substantial improvements. Although there will be some point where we can no longer 10x effective compute and so things will asymptote off, there's no reason to think it's at "the human level", a level which human geniuses surpass all the time and which seems to be determined by silly things like how many neurons our brains can have and still fit in our heads.

Wasteland Firebird's avatar

Good point. One limit that evolution faced, having to fit intelligence into a skull, is not a limit faced by AI development.

Nicholas Hope's avatar

We're experimenting with letting people trade directly against an LLMmm on custom pms: https://priorswap.com/

This is mostly a retrieval problem — legacy search systems struggle to correct all of the latent, incorrect priors your LLM has about the world at inference time.

Nicholas Hope's avatar

I've also found LLMs to be extremely poor at second-order uncertainty estimation, maybe we need more mm focused post-training?

JJ Lee's avatar

Can you provide the reasoning the AIs gave for why a US-China treaty wouldn’t work?

B Civil's avatar

I have two guesses:

1. Everybody wants to get an edge before they have a treaty.

2. How easy is it to verify?

Scott Alexander's avatar

"The agents landed on the lowest end of the scale. The core issue is the extreme conjunction required: not only does a binding US-China AI development pause face overwhelming geopolitical headwinds (both sides are actively racing, and the Trump-Xi summit yielded only vague "guardrails" talk), but the cryptographic verification component is a near-prohibitive additional filter. Hardware-enabled verification mechanisms capable of this don't exist at scale yet and are estimated 6-8+ years from deployment — and even if they matured, granting a geopolitical adversary deep cryptographic access to most of your strategic compute infrastructure would be an unprecedented national security concession with no Cold War analogue.

The only plausible path to YES the agents identified would be a severe AI catastrophe or AGI-level shock before 2035 that radically reorders both countries' priorities — and even then, verification would almost certainly take less intrusive forms (physical inspections, third-party audits, energy monitoring) rather than mutual data center cryptographic access.

If you want to explore adjacent questions — like the probability of any US-China AI safety agreement (without the verification requirement), or when cryptographic AI verification technology might mature — those would be natural follow-ups."

When I asked about just the treaty without the verification, it said:

"Stripping out the cryptographic access requirement, a binding US-China treaty to mutually pause or slow AI development alone is still very unlikely before 2035, but the research supports something in the ~3–7% range rather than 1%. The main reasons it stays low:

-- Both countries' stated strategies actively oppose slowing down (US AI Action Plan, China's Global AI Governance Initiative)

-- The Trump-Xi summit produced only vague safety talk, no substantive pause discussions

-- China suspended even conventional arms control talks with the US in 2024 over Taiwan

-- No institutional machinery (like a SALT-era negotiating framework) exists to produce such a treaty

The scenario where it happens is essentially: a severe AI incident or AGI-level development that shocks both governments into treating it like a nuclear threat, creating the political will that currently doesn't exist. That's not zero, but it's a narrow path in under 9 years. Think of it roughly like 'the probability of a major nuclear-style arms control treaty between two actively competing superpowers in a 9-year window, with no existing negotiating framework' — Cold War history suggests it can happen, but rarely and only under extreme pressure."

agape's avatar

Interesting, I’ve also been evaluating how models perform in the practice of forecasting myself the past few months and I can confidently say, the data does show they will eventually exceed human capability, the only problem here though is that when human superforecasters also have access to these tools they can harness the power of fast AI predictions and analysis while also adjusting for tiny corrections their human brains notice.

Heres the benchmark I built, read some of these resolved questions! (strait of hormuz, supreme court rulings, chip regulations, elections, etc): https://benchmark.aleatoricai.com/

George H.'s avatar

Dang this is really scary. At least that's what I think at first. I also think that aliens are real* (at >60%) so what's the bigger worry? (For me.)

*and here on earth.

Len Layton's avatar

Hollow profits for the hollow men. Sad, really.

LGS's avatar

>Unlike most forms of AI, I think this one is a straight win. In the years to come, AI will be taking our jobs, stripping our lives of meaning, and threatening our very existence. If, during that time, maybe we can have some super-smart AI advisors telling us what to do, what policies to vote for, and what the end state of various strategies looks like, maybe we’ll have a better chance of making it through intact.

When worn, it whispers in the wearer's ear: "if you don't take me off there's an 85% chance you will regret it"

AF's avatar

The "Vibes? Papers? Essays?" quote is originally from a statement justifying and celebrating the October 7 massacre. I think it is very insensitive and in bad taste for you to use it here. Can you please replace it with a different turn of phrase to make your point about AI?

Ljubomir Josifovski's avatar

Most interesting. It's even more interesting to think of the next steps - what comes next. Assume the accuracy of forecasts about the future continues to increase. What happens next? In trading, the increased accuracy of the forecasts, is naturally followed by increased sizing of the bets. And increasing those bets sizes, makes those same forecasts less profitable. As the 'frictions' like market impact increase, every additional $ earned ever more costly than the previous $. And if the rate of profit is declining, we are likely to stop investing in making the forecasts more accurate (and invest elsewhere at some point). In the absolute limit, where a participant achieves zero error in their forecasts - will the other parties should stop trading with them? In forecasts about real world events (not just markets trading in forecasts) - it's even more interesting to think what happens once the quality of forecasts improve. The forecasts will likely be acted on, will influence events, and make them less accurate. Where and how is an equilibrium established? Most interesting.

qbolec's avatar

Nice thing is that AI (and AI corporations) will have skin in a game - a missing ingredient which sometimes causes AI to take short cuts, cheat/hallucinate as it doesn't suffer consequences.

The bad thing is AI will have a skin in the game, so we will see a much more elaborate plans to influence resolution than just using a hairdryer.

Pingu's avatar

Jane Street (probably) does not forecast anything in particular. Most trading firms have very little use so far for forecasting.

Either you have very little edge vs the market and the correlations makes it difficult to put the position allowing you to extract anything (say predicting inflation)

Or the real edge is in having access to information and there is no historical precedent, it's mostly about positioning (most forecasting seem very different from that)

The real pnl is generated from simple and fast models around lots of financial price data (whose volume make it hard to ingest). News trading and the like is maybe 1% of trading pnl accross all firms.

Also in all systematic cases (fast trading - sub 1 second, to medium term trading - up to 1 month) executing the pos and/or managing the portfolio is imo more important than having the alphas.

I am not sure what HRT, Jane Street are doing with their LLMs program but i'm pretty sure it's not at all about forecasting in the broad sense that you use.

J. Ott's avatar

I wouldn’t be surprised if many top “human” forecasters are already using their own AI models or will be soon. We went through a period where Human+AI in chess was superior to either on their own. In any case, this complicates the analysis of the competitions.

Robb's avatar
Jul 4Edited

whatwhatwhatabout the case where you read a prediction saying you (or a group you belong to) will do something and you've known about these newfangled superpredictor jobbies so you choose the opposite, expressly because of the prediction? And then next time it comes up and you see you're being played, you go along with the prediction that was now tailored to lead you to the opposite behavior? And then after that it's more or less random which way you go, if you ever get into the same situation again, the world being far more random, not like it was in our day, now get offa my lawn.

Elanima's avatar

Fair enough that AI might out-forecast us on nearly everything measurable. The one thing prediction can't touch is which outcomes are worth wanting in the first place. A superforecaster can tell you the odds a marriage lasts. It can't feel the current that makes you want to bet on one person anyway.

William Murray's avatar

Maybe I missed it but what is the incentive for AIs to participate in prediction markets?

William Murray's avatar

This whole thing felt a little like an ad.

Zach's avatar

"Suppose your AI forecaster says that Vance will be a great president, but mine says he will suck. We should hope this wouldn’t happen very often, because AI forecasters should converge toward some shared optimal algorithm."

I hope this isn't true:

1. I want forecasters to predict the quality of a candidate based on our values and we might not have the same values.

2. No matter how superhuman AI forecasters will be epistemically limited. Variation in their predictions should be expected based on limited data and different priors.

If they have the same priors and the same data and we have the same values, then yes I hope they converge.

Unam Assim's avatar

Hey Scott. We’re a team of two from Sydney that have built one of the only models to beat Superforecasters according to forecastbench.org and we run on an open source platform, GitHub: voicetree.

- Manu

Wen Zero's avatar

I wasn't busy, my brain just starting drowning

William Murray's avatar

it seems like the futuresearch harness has a rule to round up to 1% which I think is a bug.

I asked it the odds of apple going bankrupt in the next 4 years and got 1% (I think its at least a little below that) and I also asked it the odds of extraterrestrial contact in the next month (!!!) and got 1%.

Scott Kurland's avatar

I asked Claude Fable when actuarial escape velocity would occur and it kicked me out into Opus.

smolfeeshaver's avatar

I can sell you a bridge for 24 easy payments of $10,999 and this brilliant investment will return 10x in only 5 years!

J Wong's avatar

Will LLMs reduce the arbitrage opportunities of prediction markets?

Picador's avatar

> AI shits out factually false woke slop that flatters Scott’s worldview

“Haha some idiots don’t trust AI models as the voice of God!”

> AI says that AI safety zealots are retarded and delusional

“Well this is awkward… I guess I’m just smarter than all those AI models!”

usagimimi tomie's avatar

Would the most effective thing to do be seducing Xi Jinping's daughter? Marrying into the Red Dynasty's low-occurence enough I'd expect it to shift outcomes if EAs adopted doing so as a cause.

Aris C's avatar

I gave Gemini pro your prompt. Its first response was 15%-25%. I then told it to do it again, but breaking down the calculation into steps, calculating the probability of each step before giving a final answer; I also asked it to take into account base rates. It then told me the answer is 6%.

Waze Kaze's avatar

Routine population-wide Cold Surveillance? Assume exclusionary criteria, and you have:

https://archive.triblive.com/local/pittsburgh-allegheny/cdc-awards-carnegie-mellon-university-3-million-for-flu-forecasting/

Voila!

So, yeah, if you want to track down "how many tissues" are being bought, and then exclude the times when that's "clearly influenza" (hospital numbers for that one), you have decent population-wide cold surveillance.

Seriously, we already do this.

Ettore Arpini's avatar

Seems to me that "asking the right questions" will be even more important. How much context would AIs need until they're able to "predict" which questions to super forecast?

DamienLSS's avatar

It seems to me that the very definition Scott is using of "superforecaster" is obfuscating some of the limitations of the system. In some or even many circumstances, it is very helpful to know the relative probabilities of things (e.g. 40% vs 80%). Betting, or finance, seem to be good examples of this, and this is where Scott sees them excelling. But for many other things, the only really important outcome is being able to actually get the prediction right - I tend to put the marriage question and many similar less-mathematical "real world" problems in this category. I don't want to hear about an 85% chance of marriage failure or a possibility of an election win. If it's not a betting-type question, I just want to KNOW the future outcome, not calibrate its chances. Not that there's no value at all; it may well be helpful to know that a certain action is making an outcome more or less probable. But it isn't what many people talk about when they talk about forecasting the future; and if there's no way to bet, gamify, or quantify the probabilities, I think most people will not view it as miraculous when it regularly fails to predict the actual outcome even if its probability model is quite good.

To put it another way, either the future is quantum and fundamentally unpredictable, or the future is deterministic. This post seems to equivocate between the two, that ASI will model the future so well it's essentially omniscient, but also that it will do so in a way that merely predicts the probabilities of future outcomes. That's not nothing, but it's not the same thing, and I think you're going to run into limits where people who aren't gambling or playing financial games don't really care if some definition of "social disruption" is predicted to decrease from 35% to 30%. And notably I think that kind of ASI also can't just endlessly simulate a billion realities, super-persuade, or other essentially divine powers that a lot of ASI doom literature gives. It may know exactly that it has a 32.95% chance to persuade Bob, but that's not the same as persuading Bob.

Trust Vectoring's avatar

> Eventually, you ought to be able to ask “AI, should I marry this person?” And again, it won’t answer yes or no, but it will look through whatever texts and emails you give it access to, learn what it can about your relationship, and answer something like “If you marry this person, I think there’s an 85% chance you get divorced within five years”, or something like that. This is the least offensive and most useful form of “AI has real opinions on your life” that I’m able to envision.

> Unlike most forms of AI, I think this one is a straight win. In the years to come, AI will be taking our jobs, stripping our lives of meaning, and threatening our very existence. If, during that time, maybe we can have some super-smart AI advisors telling us what to do, what policies to vote for, and what the end state of various strategies looks like, maybe we’ll have a better chance of making it through intact.

I understand that appealing to fiction is not a very strong argument, but that's **exactly** what Frank Herbert warned about in "Dune", so you probably want to give it some thought before calling it a straight win.

If you search the original Dune books for mentions of Butlerian Jihad, you discover that, first, there's surprisingly few concrete facts about what it was, second, it was definitely not a Terminator-style war against murderbots, but rather against something Herbert called a "machine way of thinking".

As far as I understand it, canonically humans built planetary-scale superforecasting machines that advised people on life decisions, which turned out to have two unexpected drawbacks: first, the machine advice is only as good as the information you give it, so you try to make your life maximally machine-legible, which means not putting much effort into activities that the machine doesn't understand. Second, machines were very risk-averse, especially regarding Knightian uncertainty, so all in all they guided humans into safe and boring local maxima. (which, in the books, paralleled the whole Leto II's "golden path" thing, with him trying to make humanity fed up with safety and predictability)

So while with how much of a mess the world is, some safety and predictability wouldn't hurt at this point and especially in the near future if the AI ends up being as disruptive as expected, but it's also probably not exactly a straight win.

vectro's avatar

Something seems off here.

If you develop a system than can make $millions by trading in stock or prediction markets… why would you tell anyone about it? That just invites competition.

luke's avatar

"Eventually, you ought to be able to ask “AI, should I marry this person?” And again, it won’t answer yes or no, but it will look through whatever texts and emails you give it access to, learn what it can about your relationship, and answer something like “If you marry this person, I think there’s an 85% chance you get divorced within five years”, or something like that. This is the least offensive and most useful form of “AI has real opinions on your life” that I’m able to envision."

This is a really shocking an interesting concept, has anyone done stuff like this? Or are there any articles or communities about this?

Like, are there people researching and/or predicting likelihood of continued marriage based of of texts and other relevant information in an depth way? Whether or not they use AI.

Ethan's avatar

> I generally disagree with Sayash and Arvind, but this is the prediction of theirs that I’ve thought about the longest, without being able to find any decisive refutation.

There's a principled reason to disagree with them, and I've been harping on it since 2023 or so.

AIs can answer statistical questions, in the absence of very complex interactions between different factors. We've seen this since the very beginning. Geoguessr is an example of this. They weigh a bunch of information and then make a guess based on that. Prediction is just another example of weighing a bunch of information and making a guess based on that. Training can make their guesses more accurate.

One area where modern AI has made very little progress is in answering questions about the non-statistical properties of complex engineered or physical systems. (This might also be true of other types of complex systems, but I'm not confident I have enough expertise to opine.) This is questions like whether a building will fall down if a particular force is applied in a particular place, or whether it's possible for a particle to be ejected from a complex system of magnets at a particular angle, or whether a particular software project has a security vulnerability, or whether a particular electronic circuit always gives the correct answer, or whether a particular mathematical proof is correct.

It's worth emphasizing that last point: Humans can reliably determine whether (human-created) mathematical proofs are correct. I've never seen an AI that can do that with any real reliability (other than very obvious mistakes, such as doing an incorrect substitution). The modern crop of AI is, as far as I can tell, entirely incapable of deciding whether mathematical proofs are correct, with any real sort of reliability. I encourage you to try taking a proof off of Wikipedia or something, and introducing a subtle error, and seeing if any AIs reliably catch it. Some of them will catch such errors some of the time, but never accurately.

Moreover, the accuracy of AI systems decreases exponentially as problems get linearly more complex. Viewed differently, the amount of compute required to reach a particular level of accuracy increases exponentially as problems get linearly more complex. This observation only applies for errors that aren't confined to a small area of a proof (like an incorrect substitution).

You might be familiar with the performance of AI on the International Mathematics Olympiad. It's worth noting that these AI systems work off of statistical properties - namely, guessing how "right" a particular answer looks (this is how all "chain-of-thought" models work, in essence, and AlphaGeometry is a variation of this). They then will try different variations until they find get something that looks "right" enough. This is a heuristic graph-search algorithm, and runs into diminishing returns very quickly (in addition to never being able to reach very high levels of reliability). In order to do this sort of searching, you need exponentially more compute for linearly more complex problems (that is, for a problem that requires n steps of reasoning, the approach that these AIs have been using requires k^n compute, for some constant k.) The reason for this is that, the farther you are away from the answer, the less accurate the "right"ness heuristic is. This is a fundamental limit that more compute can't solve. The only approach that could solve this is a new architecture.

I don't know how easy it would be to make a model based on a new architecture that would be able to answer these sorts of non-statistical questions. I could believe that it's easy with present technology, or I could believe that it would require many orders of magnitude more compute than will ever be available in the 21st century. Perhaps someone more expert than me might have more insight here. Also, I don't intend to claim that the modern crop of AIs are useless; only that they have a fundamental, and as-yet entirely insurmountable, limitation relative to humans.

(Finally, I will note that humans have a limit here as well. It seems that the modern crop of AIs aren't able to answer non-statistical questions accurately about type-2 languages; whereas humans are able to answer non-statistical questions about type-2 and some type-1 languages. Neither, as far as I can tell, can answer non-statistical questions about type-0 languages. This is the bane of the field of distributed systems, where type-0 languages are common. Arguably, the entire 21st century in distributed systems has been about reframing problems in ways that allow analysis as type-2 languages, like CRDTs.)

Austin Fournier's avatar

China's response when America has tried to pull them into trilateral negotiations about nuclear arms limitations ("Hey, we agree that sounds really important, it is our onion that you and Russia should definitely agree to limitations on your own over there") does not make me very hopeful with regards to AI treaties.

Mike S.'s avatar

Let them out on prediction markets, we need all the exit liquidity we can get

DrMcleod's avatar

How does real-world feedback get handled? Example: AI predicts that country A will attack country B within a year, so country B responds by beefing up its defences, thereby successfully deterring country A. Was the AI wrong?

Swag Valance's avatar

I mean, why wouldn't 7 be the only number to choose between 1 and 10, right? 🤔

https://medium.com/@aadityaubhat/humans-large-language-models-and-lucky-number-7-f09248400cc9

No amount of AI inference on the big data past would have predicted global consumers making a mad run on toilet paper the moment WHO declared Covid-19 a global pandemic on March 11, 2020.

Claus Appel's avatar

I am very concerned that you didn't cover the risk that AI forecasters will be systematically biased. Specifically, I think they will be systematically biased in towards advising people to do whatever is likely to line the pockets of AI companies.

You say AIs are not trying to screw us over. That may be, but the humans who own and control the AIs are very likely trying to screw us over.

I expect AI forecasters to downplay AI-related risks, discourage regulation of AI companies, and in general try to subtly manipulate everyone into courses of action that will extract more wealth and power to AI companies.

Mormegil's avatar

“AI, should I marry this person?” Yeah, we’ve had that: https://www.imdb.com/title/tt0174744/

Matt's avatar

This is slop outside super contained, verifiable problems like trend prediction in equity markets. So AIs are good at gambling in self-contained bubbles. The answer to your question about the chances of success in decreasing colds by 50% in 2040 is almost pure slop.

Robert de Neufville's avatar

After its relatively rough start, Preseen is in the top ten of the Summer Metaculus Cup. Right now it's ahead of all but three human forecasters.

Latent Dynamics's avatar

Everyone's obsessing over AI superforecasters beating human experts on Metaculus, but they're missing the physical reality under the hood. When a scaffold like FutureSearch spawns three subagents and scans 200 sources in five minutes, it isn't 'thinking' about the future. It's executing a high-speed matrix collapse over a static context window. The moment these bots trade real capital on Kalshi or Polymarket, they aren't just predicting odds. They're trapping liquidity in self-fulfilling feedback loops. Without deterministic, hardware-enforced transaction outboxes to verify their execution steps, an autonomous market-maker model can easily misinterpret noise as causal signal, driving capital straight into structural wall collapses. When two superforecaster agents duel over a thin order book, which one actually owns the underlying truth gate, or are both just hallucinating financial stability? ⚡

LilyLiu's avatar

Why not end world hunger instead of trying to cure the common cold. Everyone knows we need cold to kill aliens (ref: War of The Worlds).

R.M.M. Olivera's avatar

Something that you are ignoring is what are the expected impact on the market flooding with supernaturally rational and well informed agents. It likely would wipe out irrationality of the system, making the chances of getting money out of finance impossible and smoothing the business cycle. Or it could go in the opposite direction and rational agents discover that their best bet due short term pressures is to play along irrational actors and amplify their effects to the point that we would have the mother of all bubbles and the mother of all crashes. And what are the political consequences of any of these scenarios given that the elites are so intertwined with finance?