Live-streaming e-commerce has become one of the most explosive forces in modern retail, turning hours-long video broadcasts into flash sales that can move thousands of units of a single product in minutes. But behind every viral shopping stream lies a logistical nightmare: how much inventory should a company position before the cameras even switch on? A new study published in Knowledge and Information Systems offers one of the most complete answers yet, presenting an integrated artificial intelligence framework that predicts demand for individual live-streaming sessions and then uses reinforcement learning to decide exactly how much stock to allocate to each event. The research, led by Thi-Linh Ho of Ton Duc Thang University in Vietnam together with colleagues across Vietnam and Australia, addresses a problem that traditional supply chain planning has consistently failed to solve.
The core difficulty is that demand in live-streaming commerce is not a smooth, continuous curve that historical averages can capture. Each broadcast is a discrete, high-stakes event shaped by idiosyncratic factors: the charisma and reputation of the host, the specific assortment of products on offer, the timing of the session, and the unpredictable dynamics of viewer engagement. A misjudgment is punishing in both directions. Understock a session and the company loses sales in the very moment consumer excitement peaks, because viewers who cannot buy simply leave. Overstock it and capital sits idle in warehouses, tying up resources that a fast-moving retail operation cannot spare. Conventional forecasting methods, which typically extrapolate from past sales trends, struggle because the session itself, not the underlying trend, is the dominant driver of demand.
The researchers’ framework operates in two coupled stages. The first is session-based demand forecasting. Rather than treating each product’s sales history in isolation, the model blends historical transaction data with engagement-proxy features, signals that stand in for the human and contextual elements of a live broadcast, such as expected viewer activity and session characteristics. The goal is to predict total demand for each product within a given live-streaming event before that event takes place, giving the supply chain a forward-looking estimate it can act on during the crucial pre-session planning window.
To build and test this forecasting engine, the team constructed a synthetic dataset of 500 live-streaming sessions derived from two well-known public sources: the User Behavior Data from Taobao, which captures real shopping activity from one of the world’s largest e-commerce platforms, and the Twitch Live-Streaming Interactions Sample Dataset, which documents how audiences behave during live video streams. By combining these sources through probabilistic models, the researchers created a multi-dimensional simulation of a live-streaming e-commerce ecosystem that is statistically grounded in real-world behavior while remaining controllable for experimentation. Notably, the pipeline explicitly simulates the zero-demand phenomenon, the common but frequently overlooked situation in which a product sells nothing at all during a session, a challenge that public datasets rarely represent and that can badly distort models trained without it.
The forecasting results highlight a finding that echoes a broader trend in applied machine learning: on tabular data, carefully tuned gradient boosting often beats deep learning. The team compared multiple model families and found that an arithmetic-mean ensemble of ten independently trained Gradient Boosting Machines delivered the highest accuracy, achieving a mean R-squared of 0.9364 across 20 independent runs, along with the lowest standard deviation of any approach tested. That combination of high accuracy and low variance matters enormously in an operational setting, where a forecasting engine that performs brilliantly one week and poorly the next is nearly as damaging as one that is consistently mediocre. The ensemble’s stability suggests it can be trusted to feed downstream decisions reliably.
The second stage of the framework is where the work becomes genuinely distinctive. Accurate forecasts alone do not move boxes; someone or something must decide how much inventory to pull from a central supply and position for each session. The researchers automated this orchestration with a tabular Q-learning agent, a classic reinforcement learning algorithm that learns optimal ordering policies directly from sequences of demand. The agent operates in an environment where every decision carries a cost: holding inventory incurs carrying costs, running out triggers stockout penalties, and placing orders costs money. By repeatedly simulating allocation decisions and observing the resulting rewards, the agent converges on a policy that dynamically balances these competing pressures without any hand-tuned rules.
The training dynamics of the Q-learning agent are instructive. The agent converged within approximately 900 episodes, a remarkably small number by reinforcement learning standards, and achieved a 93.3 percent fill rate, meaning it successfully met demand in nearly every simulated scenario while keeping the trade-off between holding, stockout, and ordering costs in check. This efficiency stems partly from the tabular formulation, which is well suited to the discrete, session-by-session structure of the problem. Each live-streaming event becomes a decision point, and the agent’s accumulated experience across sessions teaches it when to stock aggressively, when to hold back, and how to hedge against the uncertainty that even the best forecast cannot eliminate.
The methodological contribution, the authors argue, is the integration itself: a session-aware framework that couples ensemble forecasting with reinforcement learning-driven orchestration in a supply chain context designed specifically for live-streaming commerce. Most prior work has treated demand prediction and inventory optimization as separate problems, solved by separate teams with separate tools. By binding them together, the framework ensures that the uncertainty captured by the forecaster is directly consumed by the orchestrator, so that allocation decisions reflect not just a point estimate of demand but the operational realities of cost and availability. The result is a pipeline that can, in principle, run automatically: data flows in, forecasts are generated, and stock is positioned, all before a host says a single word on camera.
The practical implications extend well beyond the platforms where live-stream selling originated. What began as a Chinese e-commerce phenomenon has spread rapidly across Southeast Asia, and Western retailers are increasingly experimenting with shoppable video formats on social platforms. Every one of these operations faces the same pre-session allocation dilemma, and the framework’s reliance on publicly available behavioral data means the approach is reproducible without proprietary access. The explicit modeling of zero-demand sessions is particularly valuable for smaller sellers, whose product catalogs often include items that simply fail to sell in a given broadcast, and whose forecasting systems are frequently blindsided by exactly those cases.
There are, of course, caveats. The forecasting engine was validated on a synthetic dataset, however realistically constructed, and real-world deployment would require adaptation to platform-specific behaviors, promotional dynamics, and shifting audience tastes. The Q-learning agent’s tabular structure, while efficient, may need scaling strategies as the number of products and sessions grows. Yet the study demonstrates that the pieces of an automated, intelligent inventory system for live-streaming commerce can be assembled from proven, well-understood machine learning components, and that doing so yields measurable gains in both predictive accuracy and operational performance. As live-streaming continues to compress the distance between marketing and fulfillment, research like this points toward supply chains that are no longer reactive afterthoughts but active participants in the show itself, quietly ensuring that when the audience rushes to buy, the products are already waiting.
Subject of Research: Machine learning-based session-level demand forecasting and reinforcement learning inventory allocation for live-streaming e-commerce supply chains
Article Title: Discovering session-level demand patterns for reinforcement learning-based dynamic inventory allocation in live-streaming e-commerce
Article References: Ho, T.-L., Pham, H., Hoang, Q. D., Binh, A. D. T., Lam, T.-P., Ngo, D.-T., & Nguyen, T. T. (2026). Discovering session-level demand patterns for reinforcement learning-based dynamic inventory allocation in live-streaming e-commerce. Knowledge and Information Systems, 68(1), Article 278. https://doi.org/10.1007/s10115-026-02895-y
Image Credits: AI Generated
DOI: 10.1007/s10115-026-02895-y
Keywords: live-streaming e-commerce, demand forecasting, reinforcement learning, Q-learning, gradient boosting, inventory optimization, supply chain management, session-based prediction, machine learning, Taobao, Twitch, ensemble learning
News Source: Denise Maddox. (October 7, 2026). AI Learns to Stock the Shelves Before the Livestream Begins. Scienmag.



