• HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
Thursday, October 1, 2026
BIOENGINEER.ORG
No Result
View All Result
  • Login
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
No Result
View All Result
Bioengineer.org
No Result
View All Result
Home NEWS Science News Technology

New R package puts an end to data snooping in forecast comparisons

Bioengineer by Bioengineer
October 1, 2026
in Technology
Reading Time: 5 mins read
0
New R package puts an end to data snooping in forecast comparisons
Share on FacebookShare on TwitterShare on LinkedinShare on RedditShare on Telegram

Every forecaster faces a seductive trap. When dozens of competing models are lined up against a benchmark, at least one of them will almost always look brilliant by pure chance — and if you cherry-pick that lucky winner, you have fallen victim to what econometricians call data snooping bias. The problem, formalized by Halbert White in his landmark 2000 Reality Check paper, is that running many pairwise comparisons inflates the probability of falsely declaring a winner far beyond the nominal significance level. A new open-source software package called RCtest, described in the journal SoftwareX by Joanna Jędrzejewska and Krzysztof Drachal of the University of Warsaw, now gathers the entire arsenal of Reality Check methodology into a single, coherent R workflow, promising to make rigorous multi-model forecast evaluation accessible to anyone who can install a package from CRAN.

The core insight behind White’s Reality Check is deceptively simple. Instead of testing each competing model separately, the Reality Check tests a joint null hypothesis: that no competing model has lower expected loss than the chosen benchmark. The test statistic is the maximum across models of the sample mean loss differential, and because its distribution under the null is analytically intractable, p-values are obtained through bootstrap resampling. RCtest implements this using the Moving Block Bootstrap of Hans Künsch, which resamples overlapping blocks of consecutive observations to preserve the short-run autocorrelation structure of time series data — a prerequisite for valid inference when forecasts are serially dependent, as they almost always are in economics and finance.

White’s original test, however, is known to be conservative: it can fail to detect genuinely superior models, particularly in small samples. Peter Hansen’s 2005 Superior Predictive Ability test addresses this by studentizing each model’s mean loss differential with its own heteroskedasticity-and-autocorrelation-consistent standard deviation, sharpening the test’s power. RCtest reports two p-values for the SPA test: a consistent p-value, obtained by recentering the bootstrap statistics at each model’s sample mean loss differential, and a conservative p-value computed under the so-called Least Favourable Configuration, which is equivalent to White’s original bootstrap. The consistent p-value is the recommended output and is always no greater than its conservative counterpart — a property the package’s automated test suite explicitly verifies.

Beyond point forecasts, the package reaches into the harder territory of density forecast evaluation. The Kullback-Leibler Information Criterion test of Corradi and Swanson compares predictive distributions by evaluating negative log-likelihood scores at each realized outcome, under a Gaussian predictive density assumption; lower expected negative log score corresponds to lower Kullback-Leibler divergence from the true data-generating process. A companion ZP test examines tail probability accuracy by scoring whether realized outcomes fall below a user-specified threshold, typically the fifth percentile, against the probabilities implied by each model’s predictive distribution. Both use SPA-style studentization and Moving Block Bootstrap inference. The package also computes the Continuous Ranked Probability Score, a strictly proper scoring rule that jointly rewards calibration and sharpness, and applies a CDF-based Reality Check comparison to series of CRPS loss differences.

Conditional superiority is another dimension the package captures. The Conditional Predictive Ability test of Giacomini and White, extended to the multiple-model setting with Hansen-type studentization, weights each loss differential series by a conditioning instrument observable at the time the forecast is made. Recommended instruments include the absolute value of realizations as a volatility proxy, lagged realizations, or any external economic indicator. A constant instrument reduces the test to the unconditional Reality Check, which the authors suggest as a built-in sanity check. This matters because a model may dominate only in turbulent periods — during the 2020 pandemic shock, say, or the 2021–2022 commodity supercycle — and such state-dependent superiority can escape entirely from unconditional tests.

To demonstrate the machinery, the authors apply RCtest to a built-in dataset of 165 monthly observations of World Bank base metals price index forecasts from March 2011 to November 2024, spanning fourteen competing models including Bayesian Dynamic Mixture Models, Dynamic Model Averaging, Bayesian LASSO and RIDGE regressions, time-varying parameter regression, and a simple AR(1) benchmark. Using the historical average as the realized series and AR(1) as the benchmark, the joint tests delivered a strikingly consistent verdict: White’s Reality Check, the consistent SPA test, and the CPA test all rejected the null across MSE, MAE, and MASE loss functions, with p-values below 0.01 in most cases, while the conservative SPA variant failed to reject — exactly the conservatism it is designed to exhibit. The CRPS-based CDF comparison also rejected, whereas the KLIC and ZP distributional tests favored the benchmark.

The package’s credibility rests on an unusually thorough validation program. Monte Carlo experiments with 500 replications per configuration assessed finite-sample behavior across sample lengths of 100 to 250 observations, autoregressive dependence coefficients from 0 to 0.6, and varying cross-model correlation and heteroskedasticity. Under the baseline null, rejection rates hovered near nominal size — 7.2 percent for WRC, 8.4 percent for SPA, 7.6 percent for CPA — while under the baseline alternatives the tests detected inferior benchmarks with power reaching 84.4 percent for WRC and 85.8 percent for SPA, and effectively 100 percent for the KLIC test. Selected routines were also checked against independent references: Diebold-Mariano results matched the forecast package’s dm.test() decisions for all thirteen model comparisons, CRPS values converged to the closed-form Gaussian result from scoringRules, and the Kupiec likelihood ratio statistic agreed with ExactVaRTest to numerical precision.

Practical usability clearly drove the design. A single high-level function, run_comprehensive_erc_analysis(), orchestrates the entire pipeline from raw forecast matrices through nine test variants across three loss functions, returning structured hypothesis-test objects compatible with base R’s printing conventions. Companion functions flatten all results to spreadsheet-ready data frames, generate automated Markdown narrative reports, and produce publication-quality ggplot2 visualizations — including cumulative loss difference plots that directly label the best and worst performers, and scatter plots linking forecast accuracy to portfolio risk-contribution weights. Runtime benchmarks on an Apple M2 machine show the full WRC-SPA-CPA battery completing in roughly a tenth of a second for thirteen models with 999 bootstrap replications, and under a second even with thirty models and nearly five thousand replications. The code is released under the GPL-3 license, with automated testing across Ubuntu, macOS, and Windows, and an archived Zenodo release for reproducibility.

The significance of this release extends well beyond econometrics. The authors point to applications in financial risk backtesting, where the package’s Kupiec unconditional coverage test verifies whether Value-at-Risk forecasts breach at their nominal rate; in epidemiological forecast comparison, where model tournaments are now routine; and in climate model selection. In the illustrative metals application, no model violated the 5 percent VaR coverage null, while the per-model tail probability scores ranked the Bayesian models most favorably. Perhaps most tellingly, the cumulative loss plots revealed that no single model dominated uniformly: several Bayesian models built an early advantage that eroded during the high-volatility years of 2020 to 2022. That kind of honest, jointly tested nuance — rather than a cherry-picked champion — is precisely what the Reality Check tradition was invented to deliver, and what RCtest now makes routine for the R community.

Subject of Research: An R package for comprehensive multi-model forecast evaluation using Reality Check and predictive density tests

Article Title: RCtest: An R Package for comprehensive forecast evaluation via reality check and predictive density tests

Article References: Jędrzejewska, J., & Drachal, K. (2026). RCtest: An R Package for comprehensive forecast evaluation via reality check and predictive density tests. SoftwareX, 36, Article 103078. https://doi.org/10.1016/j.softx.2026.103078

Image Credits: AI Generated

DOI: 10.1016/j.softx.2026.103078

Keywords: forecast evaluation, data snooping, Reality Check test, R package, econometrics, bootstrap methods, predictive density, CRPS, Value-at-Risk, SPA test, Diebold-Mariano test, open-source software

Cite Scienmag News

APA
MLA
Chicago

Denise Maddox. (October 1, 2026). New R package puts an end to data snooping in forecast comparisons. Scienmag. https://scienmag.com/new-r-package-puts-an-end-to-data-snooping-in-forecast-comparisons/

Denise Maddox. “New R package puts an end to data snooping in forecast comparisons.” Scienmag, 1 October 2026, https://scienmag.com/new-r-package-puts-an-end-to-data-snooping-in-forecast-comparisons/. Accessed 1 October 2026.

Denise Maddox. “New R package puts an end to data snooping in forecast comparisons.” Scienmag. October 1, 2026. https://scienmag.com/new-r-package-puts-an-end-to-data-snooping-in-forecast-comparisons/

Copy citation
Download RIS

Tags: bootstrap methodsCRPSdata snoopingdata snooping correctionDiebold-Mariano testeconometricseconometrics model testingforecast comparison biasforecast evaluationjoint null hypothesis testing in forecastingmodel performance comparisonmulti-model forecast validationopen-source forecast assessment toolsopen-source softwarepredictive densitypreventing cherry-picking in forecastsR packageR package for forecast evaluationReality Check methodologyReality Check testsoftware for robust forecast evaluationSPA teststatistical significance in model selectionValue-at-Risk

Share12Tweet7Share2ShareShareShare1

Related Posts

Signing in Mid-Air: VR Headsets Learn to Verify Your Signature in 3D

Signing in Mid-Air: VR Headsets Learn to Verify Your Signature in 3D

October 1, 2026
Machine Learning Meets Epidemiology to Catch Disease Outbreaks in Real Time

Machine Learning Meets Epidemiology to Catch Disease Outbreaks in Real Time

October 1, 2026

School Lunches Show Modest Diet Gains, but Menu Cost May Shape Child Nutrition

October 1, 2026

Ground Limestone Makes Hemp-Starch Building Insulation Stronger and Water-Resistant

October 1, 2026

POPULAR NEWS

  • Lamp Soot From Sesame Oil Turns Into a Five-Minute Microwave Miracle for Supercapacitors

    29 shares
    Share 12 Tweet 7
  • Machine Learning Framework Distills Crohn’s Disease Risk From Half a Million Genetic Variants

    29 shares
    Share 12 Tweet 7
  • Signing in Mid-Air: VR Headsets Learn to Verify Your Signature in 3D

    29 shares
    Share 12 Tweet 7
  • Millets’ Chloroplast Genomes Reveal Codon Preferences Shaped by Natural Selection

    29 shares
    Share 12 Tweet 7

About

We bring you the latest biotechnology news from best research centers and universities around the world. Check our website.

Follow us

Recent News

Lamp Soot From Sesame Oil Turns Into a Five-Minute Microwave Miracle for Supercapacitors

Machine Learning Framework Distills Crohn’s Disease Risk From Half a Million Genetic Variants

Signing in Mid-Air: VR Headsets Learn to Verify Your Signature in 3D

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 85 other subscribers
  • Contact Us

Bioengineer.org © Copyright 2023 All Rights Reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Homepages
    • Home Page 1
    • Home Page 2
  • News
  • National
  • Business
  • Health
  • Lifestyle
  • Science

Bioengineer.org © Copyright 2023 All Rights Reserved.