Bronchopulmonary dysplasia, or BPD, remains one of the most important chronic complications affecting babies born extremely preterm. The condition develops in immature lungs that are exposed to prolonged oxygen therapy, mechanical ventilation, inflammation and other stresses during the neonatal period. Clinicians often need to estimate which infants are most likely to develop BPD long before the diagnosis can be formally confirmed, because that forecast can influence respiratory support, nutritional planning, monitoring and discussions with families. A new article in Pediatric Research asks whether the growing volume and precision of clinical data can genuinely improve that prediction. In “Does more granular data equal better bronchopulmonary dysplasia prediction?”, E.M. Taylor, D.G. Tingay and S.B. Axford examine a question at the centre of modern neonatal medicine: whether more detailed information automatically produces more reliable medical insight.
The appeal of granular data is easy to understand. A conventional clinical record may describe an infant’s respiratory status using broad categories, such as whether the baby is receiving oxygen or mechanical ventilation on a particular day. More granular systems can capture oxygen concentration, airway pressure, respiratory rate, oxygen saturation and heart rate continuously or at very short intervals. They may also record the timing of changes in ventilator settings, episodes of desaturation, blood-gas measurements, medication exposure, weight changes and patterns of respiratory instability. In theory, these high-resolution data streams could reveal subtle differences between infants whose lungs are recovering and those who are moving toward chronic respiratory disease. Advanced statistical models and machine-learning systems can then search for combinations of variables that might be difficult for clinicians to recognise in real time.
Yet the relationship between data detail and predictive accuracy is not straightforward. A larger dataset may contain more biological information, but it may also contain more noise. Neonatal monitoring systems generate frequent measurements that can be affected by motion, poor sensor contact, calibration problems or clinical interruptions. A transient change in oxygen saturation may represent a meaningful respiratory event, or it may reflect an artefact. Ventilator data can be equally complex: a recorded pressure setting does not always indicate the pressure actually reaching the infant’s lungs, and the same setting may have different physiological effects depending on lung compliance, airway resistance and the infant’s breathing effort. Prediction models must therefore distinguish clinically meaningful signals from the enormous background of imperfect observations.
BPD itself also complicates the task. It is not a single, uniform disease with one clearly defined biological pathway. The condition can emerge through different combinations of immaturity, inflammation, infection, oxygen toxicity, fluid exposure and mechanical injury. Two infants may meet the same clinical definition while having very different lung structure and future respiratory needs. Definitions of BPD have also evolved, with different approaches incorporating the level of respiratory support, oxygen requirement and the timing of assessment. If the outcome being predicted is not consistently defined, even an exceptionally sophisticated algorithm may appear inaccurate simply because it is learning to predict a moving target.
The timing of prediction is another critical issue. A model that forecasts BPD shortly before the diagnostic assessment may perform well because it has access to information that already reflects established lung disease. That may be statistically useful but clinically less valuable than a model capable of identifying risk earlier, when treatment and surveillance decisions could still be changed. Researchers must also prevent “data leakage,” a problem that occurs when information recorded after the relevant prediction point accidentally enters the model. For example, later respiratory support or outcomes may be embedded in a dataset in ways that allow an algorithm to anticipate the diagnosis without producing a genuinely early forecast. Granular data can increase this risk because the record contains a detailed timeline of events.
The article’s central question therefore reaches beyond the technical performance of any single prediction tool. It asks what kind of information should be considered valuable in neonatal care. A model may achieve a higher numerical accuracy by using hundreds or thousands of measurements, but that does not necessarily mean it will improve decisions at the bedside. Clinicians need predictions that are interpretable, timely and robust across hospitals, equipment platforms and patient populations. If an algorithm identifies a baby as high risk, medical teams need to understand which features drove that assessment and whether those features represent modifiable factors, unavoidable consequences of extreme prematurity or merely statistical associations. Without that context, a prediction may be precise in a mathematical sense while remaining difficult to act upon.
Data quality and missingness are particularly important in neonatal intensive care. Measurements are not collected evenly from every infant. A baby who is unstable may undergo more blood tests and receive more frequent monitoring than a baby who is improving. Some variables may be absent because they were not clinically necessary, while others may be missing because of equipment limitations or transfer between units. These patterns are not random; they can themselves reflect illness severity, staffing, local practice and resource availability. If a model interprets missing data as if it were neutral, it may produce biased predictions. A system developed in one neonatal unit may also struggle elsewhere, where clinicians use different ventilation strategies, monitoring devices or documentation practices.
There is also a practical cost to collecting, storing and processing high-resolution information. Continuous physiological data require technical infrastructure, secure data management and carefully designed systems capable of translating streams of numbers into clinically meaningful summaries. Neonatal teams already work in environments saturated with alarms and digital records. Adding more alerts or complex risk scores could increase cognitive load rather than improve care. For a prediction model to be useful, its output must fit into clinical workflows, communicate uncertainty and avoid encouraging unnecessary interventions. A risk estimate should support professional judgment, not replace it. In this setting, the most effective model may not be the one that uses the greatest number of variables, but the one that extracts a small set of dependable signals and presents them at the right moment.
The discussion by Taylor, Tingay and Axford arrives as neonatal researchers increasingly explore artificial intelligence, continuous monitoring and large-scale clinical databases. The promise is substantial: better prediction could help identify infants who need closer follow-up, guide research into prevention and improve the design of clinical trials. But the paper’s title highlights a necessary caution. More granular data may reveal important physiology, yet detail alone does not guarantee truth, fairness or clinical usefulness. The value of a dataset depends on how accurately it measures biology, how consistently outcomes are defined, how transparently models are evaluated and whether predictions work beyond the environment in which they were created. For BPD, the next advance may come not from collecting every possible data point, but from combining technically sound measurement with careful clinical reasoning. In neonatal medicine, better prediction will ultimately be judged not by the size of the database, but by whether it helps vulnerable infants receive safer and more timely care.
Subject of Research: Bronchopulmonary dysplasia prediction in preterm infants
Article Title: Does more granular data equal better bronchopulmonary dysplasia prediction?
Article References: Taylor, E.M., Tingay, D.G. & Axford, S.B. “Does more granular data equal better bronchopulmonary dysplasia prediction?” Pediatric Research (2026). https://doi.org/10.1038/s41390-026-05374-w
Image Credits: AI Generated
DOI: https://doi.org/10.1038/s41390-026-05374-w
Tags: bronchopulmonary dysplasia risk predictioncontinuous physiological data in neonatologyearly detection of bronchopulmonary dysplasiagranular clinical data in neonatal careimpact of detailed clinical data on neonatal outcomesneonatal clinical data granularity and accuracyneonatal data-driven decision-makingneonatal intensive care unit data analysisneonatal respiratory monitoringneonatal respiratory monitoring technologypredictive modeling for BPDpreterm infant respiratory support


