Skip to content
All resources
Breakdown For Managing Partner, COO

Can AI Outforecast Your Team? The Case for Probability-Based Planning

8 min watch + read Published 4 Sep 2025 Video companion

Start here

The confident call versus the right call

Most professional services firms make their biggest planning calls on gut instinct, last quarter’s numbers, and whatever the senior team agrees in a Wednesday morning meeting. That works well enough when conditions are stable. When supply chains, regulation, and client budgets shift week to week, it leaves partners carrying a nagging doubt: are we making the right call, or just the confident one?

There is a new class of AI benchmarking that tests whether AI can outperform human judgement on real, unresolved questions — events where no one knows the answer yet and models cannot simply recall from training data. If your team makes forward-looking calls on anything from supplier risk to market conditions, the results are worth your attention.

Why the leaderboards mislead

Standard AI benchmarks are no longer enough

You have seen the leaderboards. GPT-5 scores 94% on FrontierMath. Claude Opus 4.8 tops coding evaluations. Models improve every six months and the charts go up and to the right. The problem is that these benchmarks are, by design, static. The questions do not change, the answers are already known, and because the questions exist publicly online, the models trained on internet data have almost certainly seen them before they sit the exam.

That creates three real weaknesses for anyone applying these scores to business decisions:

  • They go stale. Yesterday’s test questions do not reflect today’s operating context. A model that aces a 2022 knowledge quiz tells you little about reasoning over current pricing pressures, supply chain fragility, or regulatory change.
  • They leak. The exam questions end up in training data, so scores partly measure memorisation, not reasoning.
  • They are gameable. A model can be tuned to perform on benchmark tasks without building the underlying judgement that matters in the real world. A model optimised for a test is not the same as one that handles messy, moving data well.

Your firm cares less about a perfect score on a static quiz and more about getting a sensible call on a live question with incomplete information.

An honest signal

What prediction markets actually tell you

A prediction market is a platform where participants buy and sell tiny stakes in outcomes. If a contract for “Will interest rates fall this quarter?” trades at 65p in the pound, the market is saying there is roughly a 65% probability of that happening. Prices move in real time as new information arrives — the same mechanism as financial markets.

Platforms like Polymarket carry hundreds of active questions at any moment: election outcomes, central bank decisions, geopolitical events, commodity prices. Participants with money on the line aggregate information quickly and accurately. The collective belief of thousands of people, each incentivised to be correct, often produces a probability estimate more accurate than any single expert’s view.

The point is not that you should bet. It is that the probabilities are useful signals — the current best collective estimate of an outcome, updated continuously. You can use them as an input to your planning without placing a single wager.

Scored against reality

Live AI benchmarking: Profit Arena

Profit Arena is a public experimental platform that asks AI models to estimate the probability of real, unresolved world events — then scores their forecasts once those events actually happen. Researchers collect upcoming events from prediction market platforms, gather supporting information (news feeds, analyst reports, market data), and pass that pack to different AI models, each of which produces a probability estimate. When the event resolves, the models are scored against the actual outcome and against each other.

It is less “can this model memorise a textbook” and more “can this model weigh uncertainty sensibly when the answer is not yet known.”

The early results are instructive. Human markets — the collective wisdom of prediction market participants — score around 79.5% on the Brier accuracy metric used to evaluate probabilistic forecasts. Several AI models beat that baseline: GPT-5 and OpenAI’s o3 family score above it. Deepseek R1, despite strong performance on traditional benchmarks, performs significantly worse on this live test. A second ranking, based on simulated market return, has o3 mini outperforming GPT-5 despite GPT-5’s higher raw accuracy — because different models have different relationships with uncertainty. Some are more decisive, others more cautious, and those characteristics translate into different practical outcomes depending on the decision you are making.

The five-step method

Applying this to your own planning

This is not an academic exercise. The same approach — give an AI model good information, ask for a probability estimate on an unresolved question, then compare it to human judgement — can run inside your firm now. It starts at the Knowledge layer of the KWA framework: before you can run useful forecasts, your firm needs clean, accessible information an AI can actually work from.

  • Identify three live questions affecting your next 30 to 60 days. Commercial (will a major client renew?), operational (will a key supplier meet their next delivery?), or environmental (will proposed regulatory changes affect your practice area?). Make each specific enough to have a definable outcome.
  • Build an information brief for each. Gather what is publicly known — news, reports, market signals, your own operational data. The quality of the forecast depends heavily on the quality of the brief. This is the Knowledge layer doing its job.
  • Ask for a probability estimate with reasoning. Not a yes or no — a percentage likelihood and the key factors behind it. Instruct the model to search for current information rather than rely on what it already knows. One test with GPT-5 on a live political question saw the model run 57 searches across 21 sources over 11 minutes before producing its estimate. That is more thorough research than most planning meetings get.
  • Run your own team estimate independently. Before sharing the AI’s output, ask your team to assign probabilities to the same questions and record their individual answers.
  • Compare and discuss the differences. When AI and your team diverge, that gap is the most valuable output. It forces an explicit conversation about what information each side is using and what assumptions are embedded in each estimate.

Over time, this builds better collective judgement — which is exactly what the Amplify step of the 5 Steps for AI Leadership is designed to capture.

A worked example

Two estimates you can interrogate

To make it concrete: a recent test asked GPT-5 to independently research and estimate the probability of a specific candidate winning the 2028 US presidential election, instructed not to rely on prediction market prices and to gather fresh information instead. After its research, GPT-5 assigned a 38% probability to one candidate. The live Polymarket estimate at the same time was 27% for the same outcome.

An 11-percentage-point gap is meaningful — it reflects different information, different weighting, or different assumptions about how voters behave. Neither number is definitive; both are useful. You now have two independently derived estimates to compare, interrogate, and use to sharpen your own view. The same logic applies to a supplier risk question, a client renewal, or a regulatory outcome. The gap is not a problem to resolve. It is information about where your collective assumptions diverge from a model that has read everything available.

The villain worth naming

AI as a productivity toy, not a decision tool

The reason most firms do not plan this way is not that the tools are unavailable. It is that the dominant model of “using AI” in professional services still means asking ChatGPT to draft a paragraph or summarise a document. That is useful, but it operates entirely within the Workflow layer of KWA. It does not touch the firm’s decision-making capability.

Generic tools, vendor demos, and conference keynotes have trained leaders to think of AI as a productivity tool for junior tasks. That framing misses the larger opportunity: using AI to reason over uncertainty and improve the quality of forward-looking decisions made by senior people.

The practical gap is not whether AI can write faster. It is whether it can help your firm make clearer calls earlier, with less noise and less drama.

Caveats worth keeping

What probability-based forecasting is not

It is not a magic answer machine. A few things to hold onto:

  • Models can be wrong. An 80% probability still fails one time in five. AI estimates reflect available information and the model’s reasoning patterns; they do not account for information that does not yet exist or events with no historical parallel.
  • Avoid cherry-picking. A figure you like is not confirmation it is right; a figure you dislike is not confirmation it is wrong. Use the estimate as an input to your judgement, not a replacement for it.
  • Avoid anchoring. Run the independent team exercise first. The comparison is only useful if both sides started from their own reasoning.

The method is still relatively new — the platforms are early, the research is developing, the models are changing quickly. Treated as a craft to develop over six to twelve months rather than a finished product to deploy immediately, it is likely to improve your firm’s decision quality meaningfully.

The takeaway

Reasoning under uncertainty is the capability that counts

Traditional AI benchmarks measure what models know. Live forecasting benchmarks measure how well models reason under uncertainty — and the second capability is far more relevant to the decisions that actually shape a firm’s performance.

The entry point is low. Pick three questions affecting your next quarter, brief an AI model properly, get its probability estimate alongside your team’s, and compare them. That conversation — grounded in two independent estimates — is better planning than most firms are doing today. Make AI work so your team can deliver: not by automating a document, but by improving the quality of the calls that determine what gets delivered, and to whom.