Better Reference Class Retrieval Through Predictive Relevance
Estimated reading time: 6 minutes
Reference Class Forecasting (RCF) takes an outside view: rather than relying only on assumptions about the case at hand, it asks what happened in a relevant class of comparable historical cases.
But this raises a fundamental question: What makes historical cases comparable enough to form the relevant reference class?
Analysts often define comparability in advance, using characteristics they expect to matter. Yet cases that look similar do not necessarily provide the most useful evidence for predicting outcomes.
At Asset Mechanics, we address this challenge using Predictive Reference Class Retrieval (Predictive RCR), an algorithm- and data-type-agnostic architectural framework for combining predictive learning with reference-class forecasting. It is designed so that the historical evidence used for a forecast is predictively relevant, statistically supported, validated out of sample, and retained as a transparent basis for a case-specific forecast and conditional outcome distribution.
From similarity to predictive relevance
Predictive RCR dynamically constructs a context-specific reference class for each new case. Rather than prescribing in advance which characteristics determine comparability, it selects the available features and historical cases with the strongest demonstrated predictive relevance for that target. In addition, the resulting empirical reference classes provide auditable historical evidence for the forecast and for the quantile estimates that characterize the conditional outcome distribution.
Because predictive relevance can vary from one case to another, both the selected features and the composition of the reference class can change with the target.
Conceptually: Target case → Target-specific predictive features → Context-specific reference classes → Conditional outcome distribution → Forecast + risk measures
Transparent empirical evidence
Unlike standard machine-learning and LLM-based approaches, Predictive RCR retains the historical evidence supporting the forecast and conditional outcome distribution, making the empirical basis for prediction and risk directly inspectable.
For infrastructure costs, the retrieved reference class provides a top-down empirical estimate of project cost and risk, grounded in observed outcomes from relevant historical projects. For economic risks, historical periods relevant to current conditions support estimates of expected outcomes and downside risk.
This makes the evidence behind the estimate easier to inspect, interpret and challenge, and makes it possible to trace the result back to the relevant historical reference classes and supporting evidence. This is particularly useful for audit and assurance.
Validation matters
Because Predictive RCR adaptively selects features and reference cases, the full selection process must be validated out of sample. We therefore evaluate feature selection, reference-class retrieval and forecasting using only information that would genuinely have been available at the time of the forecast. Performance is assessed on unseen outcomes, including forecast accuracy and, where relevant, the calibration and coverage of predicted risk.[10]
The resulting reference class must also provide sufficient historical support for reliable estimation of both the forecast and its associated risk. Where individual reference classes are small, appropriately designed ensemble methodologies can improve statistical robustness and reduce sensitivity to any single model or selection. Similar principles are used in analog ensemble weather and hydrological forecasting, as well as in tree-based ensemble methods such as random forests and gradient boosting. [6]
When the data do not support an informative context-specific reference class, the method falls back to the broader historical dataset rather than forcing a weak local match. If a target lies outside the range of historical experience, this lack of support is detected and flagged as an extrapolation or out-of-distribution case so that it can be treated separately.[11] Material regime changes require similar caution, since historical relationships may no longer remain predictive.
Together with the transparency provided by the underlying reference cases, these safeguards support the principles of Trustworthy AI: predictions should be empirically validated, interpretable and explicit about the limits of the available evidence.[12]
Building on prior research
The question of how to predict an individual case from relevant historical evidence has a long history in statistics and decision science. John Venn recognized the reference-class problem — how to determine the appropriate class for an individual case — and Hans Reichenbach later formalized it.[1,2] The outside-view approach developed by Daniel Kahneman and Amos Tversky brought this reasoning explicitly into forecasting and decision-making.[3]
Subsequent research in Reference Class Forecasting (RCF), analog forecasting and machine learning developed practical ways to use comparable historical cases and their observed outcomes as an empirical basis for prediction. RCF typically defines a reference class of comparable historical cases using predetermined comparability criteria,[4,5] while analog ensemble methods combine multiple historical analogues to produce probabilistic forecasts.[6] Related ideas appear in Meucci’s Flexible Probabilities framework, which assigns different weights to historical observations based on conditioning information to construct context-dependent empirical distributions.[7] Machine-learning methods instead learn predictive relationships from data, including which features or similarity measures are most informative for prediction.[8]
More recently, LLM/RAG systems have shown how systems can dynamically retrieve and rank primarily textual information according to its semantic or contextual relevance to a query.[9]
Across these approaches, however, the central challenge remains: what makes an observation sufficiently relevant to improve the prediction, quantify the risk around it, and identify the underlying drivers?[5]
Predictive RCR addresses this as a predictive relevance problem. A historical case may be descriptively, statistically or semantically similar without being particularly useful for predicting or explaining the outcome of a new case. Rather than assuming such similarity, Predictive RCR lets predictive performance determine which available features and historical cases are most informative for each target. Both the drivers and historical evidence underlying the forecast or risk assessment are made explicit.
Summary
| RCF | Predictive ML | LLM / RAG | Predictive RCR | |
|---|---|---|---|---|
| What is learned | Predefined similarity criteria | Predictive relationship | Semantic / contextual relevance | Out-of-sample predictive relevance |
| Basis for prediction | Outcomes of comparable cases | Learned model | Retrieved contextual information | Outcomes of target-specific predictive reference class |
| Typical data | Historical cases and outcomes | Various data types | Primarily text | Various data types |
| Historical input cases retained as evidence | Yes | No | No | Yes |
| Target-specific feature weights | Usually predefined / implicit | Model-derived | No | Yes |
| Risk | Empirical | Model-derived | No | Empirical |
| Evaluation | Reference-class calibration | Out-of-sample predictive performance | Retrieval / generation quality | End-to-end out-of-sample predictive relevance |
The defining principle is that the evidence retained for a forecast must demonstrate out-of-sample predictive relevance.
References
- Venn, J. (1866). The Logic of Chance: An Essay on the Foundations and Province of the Theory of Probability. Macmillan.
- Reichenbach, H. (1949). The Theory of Probability: An Inquiry into the Logical and Mathematical Foundations of the Calculus of Probability. University of California Press.
- Kahneman, D., & Tversky, A. (1979). Intuitive prediction: Biases and corrective procedures. TIMS Studies in Management Science, 12, 313–327.
- Flyvbjerg, B. (2006). From Nobel Prize to project management: Getting risks right. Project Management Journal, 37(3), 5–15. DOI: 10.1177/875697280603700302.
- Cantarelli, C. C., Davis, K., Pinto, J. K., & Turner, N. (2026). Reference class forecasting: Promises, problems, and a research agenda moving forward. Production Planning & Control, 37(7), 691–709. DOI: 10.1080/09537287.2025.2578708.
- Delle Monache, L., Eckel, F. A., Rife, D. L., Nagarajan, B., & Searight, K. (2013). Probabilistic Weather Prediction with an Analog Ensemble. Monthly Weather Review, 141(10), 3498–3516. DOI: 10.1175/MWR-D-12-00281.1.
- Meucci, A. (2010). “Historical Scenarios with Fully Flexible Probabilities.” GARP Risk Professional, December, pp. 47–51.
- Weinberger, K. Q., & Saul, L. K. (2009). Distance metric learning for large margin nearest neighbour classification. Journal of Machine Learning Research, 10, 207–244.
- Lewis, P., Perez, E., Piktus, A., et al. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. Advances in Neural Information Processing Systems, 33.
- Cawley, G. C., & Talbot, N. L. C. (2010). On over-fitting in model selection and subsequent selection bias in performance evaluation. Journal of Machine Learning Research, 11, 2079–2107.
- Yang, J., Zhou, K., Li, Y., & Liu, Z. (2024). Generalized out-of-distribution detection: A survey. International Journal of Computer Vision, 132(12), 5635–5662. DOI: 10.1007/s11263-024-02117-4.
- National Institute of Standards and Technology. (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1. DOI: 10.6028/NIST.AI.100-1.