Beyond Amazing
The Strategy Toolkit

Risk

Reference Class Forecasting

A forecasting method that starts from recorded outcomes of comparable past projects rather than the plan in front of you. Locating your case in that distribution, adjusting only for defensible differences, corrects the optimism and strategic misrepresentation of the inside view.

Also known as Outside view, RCF, Comparison class forecasting. First set out by Daniel Kahneman and Amos Tversky (concept); Bent Flyvbjerg (operational method) in 1979; the primary source is cited in full below.

Format
Scoring model
Level
Corporate · Business unit
Best for
Assess risk · Plan execution
Decision stage
Decide · Plan · Review
Difficulty
Advanced
Time to apply
A few days with an existing dataset; weeks if the reference class must be assembled from scratch.

Plate · The model

Choose the referenceclassEstablish the outcomedistributionPosition the project initAdjust for reliablespecifics
The 4 steps of Reference Class Forecasting, worked in sequence.
I

The components

1

Choose the reference class

Define the set of completed projects the forecast will be anchored to. This is the method's most consequential judgement: the class must share the drivers of outcome with the project at hand, and its selection should be documented, justified and open to challenge.

Signals of strength
Broad enough to be statistically meaningful · Narrow enough to be genuinely comparable · Selection made before the answer is known · Class definition survives challenge by a sceptic

2

Establish the outcome distribution

Turn the class into a probability distribution of outcomes, usually expressed as ratios of actual to forecast cost, duration or demand. Percentiles, skew and tail behaviour carry the information; a single average hides the disasters that contingency exists for.

Signals of strength
Actuals measured against forecasts at a consistent baseline date · Distribution plotted, not just averaged · Skew and outliers examined rather than trimmed away · Data quality of the class honestly assessed

3

Position the project in it

Set the baseline forecast from the distribution at the percentile your risk appetite demands. Budgeting at the median accepts a roughly even chance of overrun; regulated UK practice often budgets nearer the 80th percentile. The choice is a governance decision made visible.

Signals of strength
Percentile chosen explicitly and signed off · Gap between inside estimate and class baseline confronted, not split · Implied contingency stated in money and time · Acceptance criteria set before advocates react to the number

4

Adjust for reliable specifics

Amend the baseline only for demonstrable, evidenced differences between this project and the class, such as a fixed-price contract transferring a risk the class bore. Unevidenced claims of being smarter, better managed or further advanced are precisely the optimism the method exists to remove.

Signals of strength
Each adjustment tied to objective evidence · Adjustments cut both ways, not only downward · Total adjustment small relative to the uplift · Reasons recorded for the post-project review

II

When it earns its keep

  • You are approving a major project whose budget and benefits were estimated by the people who want it to happen, and you need an independent check before committing.
  • Your organisation's projects show a consistent pattern of overruns and shortfalls, which is evidence of bias rather than bad luck, and case-by-case scrutiny has not fixed it.
  • Usable outcome data exists for a class of similar projects, whether your own portfolio history, published datasets or sector benchmarks.
  • You must set contingency or an optimism-bias uplift defensibly, for instance under HM Treasury Green Book appraisal rules.

And when it doesn't

  • The project is genuinely unprecedented and no honest reference class exists. Forcing a poor comparison class produces confident numbers with no foundation; use scenario planning and staged commitment instead.
  • The forecast question is about which option to pursue rather than what an option will cost. RCF disciplines an estimate; it does not generate or compare alternatives.
  • The stakes are too small to justify assembling a distribution. A premortem or a simple contingency rule gives most of the debiasing benefit at a fraction of the effort.
  • Your only motive is to legitimise a number already chosen. Selecting a reference class to hit a target answer is the inside view wearing a lab coat.
III

How to run it

Before starting, gather the inputs the analysis depends on:

  • A precise definition of what is being forecast: final cost, schedule, demand or benefit, and the point of estimation it is measured from.
  • Outcome data for a reference class of completed comparable projects, ideally dozens, with actuals against their original forecasts.
  • The project's own base estimate, so the outside view has something to correct.
  • A stated risk appetite, because choosing which percentile of the distribution to budget at is a governance decision, not a statistical one.
  1. 1

    Commit to the outside view

    Set aside the project's internal logic before it anchors you. Kahneman and Tversky's finding is that inside views fail predictably because they treat the case as unique; the corrective is to begin with how projects of this kind actually turn out.

  2. 2

    Choose the reference class

    Assemble completed projects similar enough to be informative and numerous enough to be statistical. Flyvbjerg's rule of thumb is that the class must be broad enough for meaning and narrow enough for comparability, and the choice deserves to be documented and challenged.

  3. 3

    Establish the outcome distribution

    Plot the actual outcomes of the class against their forecasts, typically as cost-overrun or benefit-shortfall ratios. The shape matters as much as the average; construction classes are usually right-skewed, with a fat tail of disasters.

  4. 4

    Position the project in the distribution

    Take the class median or the percentile matching your risk appetite as the baseline prediction. The UK Green Book approach applies uplifts so that, for example, a budget can be set at the 80th percentile of the class rather than at the estimator's number.

  5. 5

    Adjust only for reliable specifics

    Move off the baseline solely where there is objective evidence this project differs from the class, and expect such evidence to be rarer than the team believes. Every claimed difference is also a claim the class was wrongly chosen; treat it with the same scepticism.

  6. 6

    Record the outturn and feed the class

    When the project completes, add its actuals to the dataset. RCF improves with every honest data point, and the review loop is what turns a one-off correction into an institutional capability.

IV

Reading the result

A debiased forecast: the chosen reference class and its outcome distribution, the project's baseline position at a stated percentile, the evidenced adjustments, and the resulting budget, schedule or demand figure with its implied contingency.

  • The gap between the team's estimate and the class baseline is the finding. A large gap does not prove this project will overrun; it proves that projects like it, forecast by people like these, usually did.
  • Read the percentile as a risk decision. P50 versus P80 is a statement about how much overrun the organisation is prepared to absorb, and it belongs to the board, not the analyst.
  • Small print matters at review: if the outturn lands far outside the class distribution, question the class definition before celebrating or despairing.
V

A worked example

A conference-centre expansion is retested against the outside view

The board of a city-centre conference venue in the North East is weighing a £24m extension adding a second exhibition hall. The development team forecasts completion in 30 months and a 15% uplift in event revenue in the first full year. The finance committee, aware that the team's last two capital projects both overran, commissions a reference class forecast before the funding decision.

Choose the reference class
The analyst assembles 41 completed UK and northern European venue projects: conference-centre extensions, exhibition halls and comparable civic buildings over £10m, completed in the last 20 years, with published or FOI-obtained actuals. Pure new-builds and stadium projects are excluded as structurally different. The team objects that none used their contractor; the objection is noted as exactly the uniqueness claim RCF exists to test.
Establish the outcome distribution
Against final business-case budgets, the class shows a median cost overrun of 21%, an 80th percentile of 46%, and a right skew driven by three projects that hit ground-condition and procurement failures. First-year demand outcomes are worse: the median venue achieved only 78% of its forecast first-year uplift.
Position the project in it
The committee sets cost at P80, implying a budget of £35m rather than £24m, and tests revenue at the class median, cutting the assumed first-year uplift from 15% to about 12%. On those numbers the scheme still clears the board's hurdle rate, but with roughly a third of the headroom the original case claimed.
Adjust for reliable specifics
Two adjustments survive scrutiny. The shell-and-core works are under a signed fixed-price contract, transferring a risk most of the class carried, worth a modest reduction from P80. Against that, the venue must trade through construction, which several class members did not, adding a disruption allowance. Net effect: budget settles at £33m. The team's claimed 'experienced delivery partner' adjustment is rejected as unevidenced.

The read. The outside view moved the budget from £24m to £33m and trimmed the revenue case, converting the decision from an easy yes into a marginal one that the board approved with staged funding gates and a scope-reduction option held in reserve. The honest caveat is that the class blends extensions with civic new-builds because pure comparators were scarce, so the P80 figure is an informed judgement rather than a law of nature. Two years on, the outturn will be added to the dataset either way, which is the part of the method most organisations skip.

VI

Pitfalls

  • Choosing the reference class after seeing what different classes imply. Class selection is the method's soft underbelly and must be fixed, documented and challenged before the distribution is computed.
  • Averaging away the tail. Skewed classes mean the mean flatters the risk; contingency exists for the 80th percentile, not the 50th.
  • Allowing generous 'specific adjustments' to claw the forecast back to the inside estimate. If adjustments roughly cancel the uplift, the exercise has been captured.
  • Applying uplifts mechanically forever. Uniform uplifts dull the incentive to estimate well and can become a self-fulfilling budget to be spent; the Green Book regime expects uplifts to shrink as risk management improves.
  • Forgetting the review loop. An organisation that never feeds outturns back into its dataset is buying the debiasing once and letting it decay.
VII

What the critics say

The choice of reference class is subjective and consequential: the class must be broad enough to be statistically meaningful yet narrow enough to be comparable, different defensible choices yield materially different forecasts, and equal weighting of dissimilar class members can bias the estimate. Recent reviews identify class construction and data availability as the method's central unresolved problems.

'Reference class forecasting: promises, problems, and a research agenda moving forward' (2025) Production Planning and Control.

The empirical foundation has been challenged from the inside view. Eliasson and Fosgerau show that observed cost overruns can arise from selection among noisy estimates rather than biased forecasting, in which case blanket uplifts misdiagnose the problem, while Love and Ahiaga-Dagbui argue cost growth is largely explained by scope change and estimation error, so outside-view uplifts treat a symptom and discard genuine project-specific knowledge.

Eliasson, J. and Fosgerau, M. (2013) 'Cost overruns and demand shortfalls: deception or selection?', Transportation Research Part B, 57, 105-113; Love, P. E. D. and Ahiaga-Dagbui, D. D. (2018) 'Debunking fake news in a post-truth era', Transportation Research Part A, 113, 357-368.

Institutionalised uplifts create their own incentives: teams may treat the uplifted figure as the real budget and spend to it, and promoters learn to game the class or the adjustments, which is why Flyvbjerg pairs the method with accountability measures rather than presenting it as sufficient on its own.

VIII

Sources and further reading

  • Kahneman, D. and Tversky, A. (1979) 'Intuitive prediction: biases and corrective procedures', TIMS Studies in Management Science, 12, 313-327. ↗
  • Lovallo, D. and Kahneman, D. (2003) 'Delusions of Success: How Optimism Undermines Executives' Decisions', Harvard Business Review, 81(7), July 2003. ↗
  • Flyvbjerg, B. (2006) 'From Nobel Prize to Project Management: Getting Risks Right', Project Management Journal, 37(3), 5-15. ↗
  • HM Treasury, Green Book supplementary guidance: optimism bias. London: HM Treasury. ↗

Pairs well with Pre-mortem·Cost-Benefit Analysis·Scenario Planning·Risk Matrix·Decision Trees·compare side by side

Near neighbours (computed from shared tags)·AI Risk Management Framework (NIST AI RMF)·Break-even Analysis·Failure Mode and Effects Analysis