Bitcoin has long been the stormiest sea in finance, a market where prices can swing by thousands of dollars in hours and where even the most sophisticated statistical models have historically struggled to keep their footing. Now, a team of researchers led by Ashwani Kharola of Graphic Era Deemed to be University, writing in the journal Discover Artificial Intelligence, reports that a carefully engineered machine learning framework can predict Bitcoin’s short-term closing prices with remarkable precision, while also explaining exactly why it makes each prediction. The work tackles one of the most persistent criticisms of artificial intelligence in finance: that its most powerful models are inscrutable black boxes whose forecasts cannot be trusted or audited.
The centerpiece of the study is a stacked ensemble architecture, a technique in which several different machine learning models are trained to make predictions and a second-level model, called a meta-learner, learns how best to combine their outputs. The researchers assembled their ensembles from four complementary algorithms: AdaBoost, which sequentially focuses on observations that previous models got wrong; CatBoost and XGBoost, two gradient-boosting methods that build trees iteratively to correct residual errors; and Random Forest, which averages many decision trees trained on random subsets of data to reduce variance. Because these algorithms capture different aspects of nonlinear relationships, their errors tend to be uncorrelated, giving the meta-learner genuinely diverse information to work with.
The winning configuration used AdaBoost, XGBoost, and Random Forest as base learners, with CatBoost acting as the meta-learner. On a held-out test set of 1,745 unseen observations, this architecture achieved a mean absolute error of 169.65 dollars, a root mean square error of 241.93 dollars, and a coefficient of determination of 0.975, meaning it explained roughly 97.5 percent of the variance in Bitcoin’s closing price. Crucially, it ranked first across every one of the five evaluation metrics the team employed, including the weighted mean absolute percentage error of 0.35 and a mean absolute scaled error of just 0.13 relative to a random walk baseline, the standard benchmark against which financial forecasters are judged.
The data behind these results came from a Kaggle repository of Bitcoin-to-US-dollar market observations recorded at 15-minute intervals, spanning roughly 94 days from late September to late December 2021. That window captured the full drama of the period: prices climbing from around 40,000 dollars to a peak near 67,000 dollars in early November before sliding back toward 45,000 to 50,000 dollars by December. Each record contained the standard open, high, low, close, and volume variables, with the closing price designated as the target. The researchers split the 8,721 observations chronologically, reserving the final 20 percent as a test set, a design choice that prevents any information from the future leaking into the training process, a subtle but fatal flaw in many published forecasting studies.
Before any model was trained, the team applied a battery of preprocessing and validation safeguards. Outliers were flagged using the interquartile range criterion, a distribution-independent method that does not assume the data follow a normal curve. Min-max scaling was fitted exclusively on training data and then applied unchanged to the test set. Hyperparameters were tuned with the Optuna framework, which uses a Tree-structured Parzen Estimator to search parameter spaces efficiently, and model stability was assessed through 20-fold expanding-window time-series cross-validation, in which each validation set consisted only of observations occurring after the training period. A fixed random seed of 42 was used throughout to guarantee reproducibility.
What distinguishes this study from the crowded field of Bitcoin prediction papers is its insistence on explainability. The researchers used SHAP, or SHapley Additive exPlanations, a technique borrowed from cooperative game theory that assigns each feature a contribution value for every prediction. The global SHAP analysis revealed that the intraperiod high price was by far the most influential predictor, with an importance score of 172.72, followed by the low price at 82.60 and the open price at 34.25. Trading volume, by contrast, contributed almost nothing, scoring just 1.06. In other words, the model’s forecasts are driven almost entirely by price-based variables, with volume adding negligible predictive signal.
Local interpretability came from LIME, or Local Interpretable Model-agnostic Explanations, which probes a model by perturbing individual observations and observing how predictions change. Applied to five representative test observations with 5,000 perturbed samples each, LIME consistently identified the immediately preceding closing price, the first lag of the target variable, as the strongest contributor to each prediction, followed by the one-step lags of the low and high prices. Longer lags and volume-related features contributed smaller and more variable amounts. The researchers caution that because the open, high, low, and close variables are strongly correlated, SHAP and LIME values should be read as model-specific predictive contributions rather than independent causal effects, an honest caveat that reflects the team’s methodological rigor.
The comparison against traditional methods was stark. Classical statistical baselines, including ARIMA, SARIMA, the historical mean model, and the random walk method, all produced substantially higher errors, with the historical mean model performing worst of all, posting a mean absolute error of 8,737.18 dollars and a coefficient of determination near zero. A hybrid configuration pairing an Inverted Transformer with Random Forest also fared poorly, achieving an R-squared of only 0.253, a reminder that attention-based deep learning is not automatically superior for every forecasting task. A Diebold-Mariano test comparing the two best ensembles yielded a statistic of minus 8.899 with a p-value below 0.001, providing formal statistical evidence that the CatBoost-meta-learner architecture genuinely outperforms its closest rival rather than winning by chance.
Robustness checks reinforced the result. Across five different random seeds, the winning model’s average mean absolute error was 170.68 dollars with a standard deviation of only 11.82, and its R-squared averaged 0.980 with a standard deviation of just 0.002, indicating that performance does not hinge on a lucky initialization. The authors suggest that ensemble performance depends not only on which algorithms are included but critically on which one is chosen as the meta-learner, since swapping CatBoost into the combining role transformed results that were otherwise mediocre. They propose extending the framework with Transformer variants, Temporal Fusion Transformers, and hybrid CNN-LSTM models, and validating it across bull, bear, and high-volatility regimes. For now, the study stands as a demonstration that in the wild west of cryptocurrency markets, accuracy and transparency can finally be engineered into the same forecasting machine.
Subject of Research: Explainable stacked ensemble machine learning for short-term Bitcoin closing price forecasting
Article Title: An explainable SHAP–LIME integrated stacked ensemble framework for bitcoin closing price forecasting
Article References: Kharola, A., Singh, H., Kumar, K., Kumar, N., Kumar, A., John, V., Gill, R., Singh, S., & Mengist, Y. (2026). An explainable SHAP–LIME integrated stacked ensemble framework for bitcoin closing price forecasting. Discover Artificial Intelligence, 6(1), Article 1346. https://doi.org/10.1007/s44163-026-02377-8
Image Credits: AI Generated
DOI: 10.1007/s44163-026-02377-8
Keywords: Bitcoin, price forecasting, stacked ensemble learning, SHAP, LIME, explainable AI, XGBoost, CatBoost, AdaBoost, Random Forest, time-series cross-validation, cryptocurrency markets
News Source: Denise Maddox. (October 11, 2026). Explainable AI ensemble cracks Bitcoin price forecasting with 97.5% accuracy. Scienmag.



