Site Overlay

Predictive Reference Class Retrieval

Making Reference Class Forecasting Context-Specific

Estimated reading time: 4 minutes

Reference Class Forecasting (RCF) takes an outside view: rather than relying only on assumptions about the case at hand, it asks what happened in a relevant class of comparable historical cases.

But this raises a fundamental question: What makes historical cases comparable enough to form the relevant reference class?

Analysts often define comparability in advance, using characteristics they expect to matter. Yet cases that look similar do not necessarily provide the most useful evidence for predicting outcomes.

At Asset Mechanics, we address this challenge using Predictive Reference Class Retrieval (Predictive RCR).

From similarity to predictive relevance

Predictive RCR dynamically constructs a context-specific reference class for each new case. Rather than prescribing in advance which characteristics determine comparability, it identifies the available features and historical cases that provide the strongest predictive evidence for that target.

Because predictive relevance can vary from one case to another, both the selected features and the composition of the reference class can change with the target.

For an infrastructure project, relevant features may include project, location and technical characteristics. For land prices or interest rates, they may involve different combinations of economic and market conditions.

Conceptually: Target case → Target-specific predictive features → Context-specific reference class → Forecast

The objective is not to find cases that look most similar according to a predetermined definition, but to identify the historical evidence that is most informative for the forecast at hand.

Transparent empirical evidence

Predictive RCR makes the empirical basis of a forecast transparent by showing which historical observations are most relevant and which features have predictive value in different situations.

For infrastructure costs, the retrieved reference class provides a top-down empirical forecast of project cost and risk, based on observed outcomes from relevant historical projects. For economic risks, historical periods relevant to current conditions can provide an empirical basis for estimating value and downside risk.

This makes the evidence behind the forecast easier to inspect and challenge, which is particularly valuable when estimation errors can have material consequences.

Validation matters

Allowing an algorithm to search for predictive features and reference cases also creates a risk of finding patterns that do not generalise. Out-of-sample validation is therefore central to Predictive RCR.

We evaluate feature selection, reference-class retrieval and forecasting using only information that would genuinely have been available at the time of the forecast. We then test performance on unseen outcomes, assessing forecast accuracy and, where relevant, the calibration and coverage of predicted risk.

The key question is simple: does the dynamically constructed reference class improve our ability to forecast unseen outcomes?

Similarity alone does not establish whether historical cases form an appropriate reference class for forecasting. Predictive relevance needs to hold out of sample. This reflects our broader approach to Trustworthy AI, where empirical grounding, transparency and validation matter as much as predictive power.

Building on Prior Research

The question of how to predict an individual case from relevant historical evidence has a long history in statistics and decision science. John Venn recognized the reference-class problem — how to determine the appropriate class for an individual case — and Hans Reichenbach later formalized it. The outside-view approach developed by Daniel Kahneman and Amos Tversky brought this reasoning explicitly into forecasting and decision-making.

Subsequent research in Reference Class Forecasting (RCF) and similarity-based forecasting developed practical ways to use comparable historical cases and their observed outcomes as an empirical basis for prediction. RCF typically defines a relevant class of comparable cases, while similarity-based approaches select or weight observations according to measured similarity.

More recently, LLM/RAG systems have shown how systems can dynamically retrieve and rank primarily textual information according to its semantic or contextual relevance to a query.

Across these approaches, however, the central challenge remains: what makes an observation sufficiently relevant to improve the prediction, quantify the risk around it, and identify the underlying drivers?

Predictive RCR addresses this as a predictive relevance problem. A historical case may be descriptively, statistically or semantically similar without being particularly informative about the outcome of a new case. Rather than relying on predefined comparability, generic similarity or semantic relevance, Predictive RCR lets predictive performance determine which available features and historical cases are most informative for each target.

Summary

RCFLLM / RAG retrievalPredictive RCR
ObjectiveOutside-view forecastingInformation retrievalContext-specific reference-class forecasting
RelevanceCase comparabilitySemantic / contextualPredictive
Typical dataHistorical cases and outcomesPrimarily textData agnostic
SelectionTypically based on predefined domain characteristicsQuery-specificTarget-specific features and cases
EvaluationNot inherentRetrieval relevanceOut-of-sample predictive performance

Predictive RCR is algorithm- and data-agnostic: different machine-learning methods and data types can determine predictive relevance. The defining principle is not the algorithm or data type, but whether the evidence helps predict observed outcomes. Out-of-sample testing then evaluates the resulting retrieval and forecasts.

Home » Insights » Predictive Reference Class Retrieval

Risk and Data scientist at Asset Mechanics | https://assetmechanics.org/

Risk and Data scientist at Asset Mechanics

Risk and Data scientist at Asset Mechanics R&D | https://assetmechanics.org/

Risk and Data scientist at Asset Mechanics R&D