First posted to SSRN in July 2026; earlier versions circulated as Grading the Graders and The Certainty Business. The September 2026 revision, retitled The Art of Never Being Wrong, adds an independent second-coder validation of the full grading ledger, with every disagreement adjudicated on the record.
Abstract
Every year, investors turn to year-ahead outlooks to learn what to watch and where markets are heading. This paper asks whether those forecasts identified, in actionable form, the events that actually moved markets. It grades 135 forecasts from fifteen leading institutions in finance and political risk across the 2024, 2025 and first-half 2026 cycles, with longer tests of bank index targets and continuously public risk rankings. The fifteen earned 13.5 points out of a possible 135, ten percent. More striking than the score is how much of the material avoided forecasting at all: the event that moved markets was named in 21 of the 135 cases, and the channel through which it would reach a portfolio in three.
The paper argues that the problem is the product, not simply the forecasters. Political risk was once priced as a recurring cost of doing business in emerging markets, while advanced economies were analyzed mainly through economics. It now often arrives as single-shot decisions in the largest economies, with little historical precedent to forecast from. Buyers, meanwhile, use outlooks to validate a view they already hold. That combination rewards narratives broad enough to survive almost any outcome. A more useful product would be if-then scenarios: what would have to change, and how that change reaches markets. The point is not to predict the event. It is to identify the signal that changes the odds.
KeywordsGeopolitical risk; market risk; political risk; forecast evaluation; forecast accuracy; expert judgment; superforecasting; scenario planning; scoring rules; year-ahead outlooks; Knightian uncertainty; ambiguity aversion.
JEL classificationC53; D81; D84; G17; G41; F51.
The question
Every January, political risk houses, global banks, asset managers, and multilaterals publish year-ahead outlooks that corporations and investors buy to guide capital allocation. This paper asks a simple question of that product: when the calls are graded against what actually happened, how do they score?
The design
Every forecaster is graded at the maximum depth its public record permits, in three tiers. Tier one is a full cross-section: fifteen institutions, spanning political risk and macro research houses, global banks and asset managers, one multilateral, and two prominent individual forecasters, graded across the 2024, 2025, and 2026 cycles. Tier two extends the banks on their flagship annual call, the published S&P 500 year-end target, across six cycles from 2021 through 2026, using dated press documentation. Tier three extends the two institutions whose complete ranked forecasts are continuously public, Eurasia Group’s Top Risks and the World Economic Forum’s Global Risks Report, across a full decade, 2016 through 2026.
The depth distribution is itself a finding. The industry that sells prediction mostly does not preserve its predictions in checkable form.
The rubric
A call earns credit only if it is specific enough to guide a real decision: it must name an actor, a mechanism, a direction, and a time window. Calls are scored against an answer key restricted to events passing a five-part materiality test, so forecasters are graded only on the events that actually moved markets.
Validation
The full grading ledger was independently recoded by a second coder, with every disagreement adjudicated on the record. Post-adjudication agreement is 87.1 percent (Cohen’s κ 0.76); 16 disagreements stand unresolved and are published as such. Excluding any single institution moves the overall credit rate between 6.25 and 11.25 percent. The adjudication log, including the 16 standing disagreements, is published in full: download the adjudication log (CSV).
What the record shows
- Across the 135-call cross-section, total credit is 13.5. One call earned full credit: a bank’s 2026 reserve-diversification call built on three years of already-visible central-bank gold buying.
- Fewer than four in ten calls named the relevant actor or theme. The event itself was named in 21 of the 135 cases, and the channel through which it would reach a portfolio in three.
- Of the 81 misses, 73 are omissions. Only 8 are wrong calls.
- By cycle: 2024, 2 of 45. 2025, 7.5 of 45. 2026 H1, 4 of 45.
- The street consensus on the S&P 500 missed direction in consecutive opposite years: a median target of 4,825 ahead of the 19 percent decline of 2022, and a median of 4,000 ahead of the 24 percent rally of 2023, with survey evidence that targets are revised intra-year to follow realized prices.
- Each January’s consensus is well approximated by a persistence projection of the prior year’s realized environment.
- The instrument transmitting each cycle’s largest market event never appears on any list.
- The genuinely strong calls in the record share one feature: each described a visible, already-formed present condition. None predicted a discontinuity.
Why the product fails
The pattern reflects structure, and competence cannot fix it. The process being forecast is nonstationary. The events are single-shot, so forecasters never receive calibration feedback. Incentives reward vagueness. And the payoff matrices are specified at the level of the state, while the decisions that move markets are made by leaders optimizing personal survival, a game that offers no base rates to extrapolate and no stable logic to solve.
What works instead
Credit accrues where a product describes conditions already visible in the present, and it evaporates at discontinuities. The practical conclusion for corporations and investors is to stop paying for point predictions and buy the things that hold up: scenario preparation, monitoring of visible conditions, and short forecast horizons. Conditions are readable even when events are not.
All grading tables are published with the paper for replication.