For millions of people, the hardest part of hearing is not hearing at all. It is following a conversation in a crowded restaurant, a bustling train station, or a noisy classroom. Difficulty understanding speech in noise is one of the most common and disabling auditory complaints across the human lifespan, yet it remains stubbornly hard to predict with standard clinical tests. Now, a systematic review and meta-analysis published in the Annals of Biomedical Engineering has taken stock of how well machine learning and related computational approaches can forecast speech-in-noise performance, and the verdict is both encouraging and cautionary: the models work, but not well enough yet for the clinic.
The stakes are enormous. According to the World Health Organization, nearly 2.5 billion people, roughly one in four worldwide, are projected to have some degree of hearing loss by 2050, with an estimated 700 million requiring rehabilitation services. But the problem extends far beyond those with diagnosed hearing loss. Many people whose audiograms look perfectly normal still report substantial difficulty following conversations in noisy environments. This gap between what the audiogram says and what the listener experiences has long frustrated clinicians, because hearing thresholds alone provide an incomplete and often misleading picture of real-world functional hearing.
Scientists have proposed several mechanisms to explain these hidden deficits. Emerging evidence points to conditions such as cochlear synaptopathy, a loss of connections between sensory cells and auditory nerve fibers, and impairments in extended high-frequency hearing, both of which can degrade the neural coding of sound without shifting standard threshold measures. The result is a central challenge for hearing healthcare: current clinical tools do not accurately predict how well an individual will understand speech in the acoustically messy conditions of everyday life. Conventional linear statistical models built on audiometric and demographic variables capture some group-level trends, but they miss the nonlinear interplay among auditory, acoustic, and cognitive factors, such as attention and working memory, that shape each listener’s performance.
Enter machine learning. The new review, conducted by Prabhani Athukorala, Anu Nair, and Srikanta K. Mishra of The University of Texas at Austin and the University of South Dakota, followed PRISMA 2020 guidelines and prospectively registered its protocol. The team searched PubMed, Scopus, Web of Science, IEEE Xplore, and Google Scholar, screening 6,468 records after duplicate removal. They included studies using traditional machine learning models such as support vector machines and random forests, neural network architectures ranging from early artificial networks to deep and convolutional models, automatic speech recognition systems, and mechanistic auditory models that simulate the physiology of hearing itself.
The quantitative synthesis focused on the model-metric combinations with enough comparable studies to pool: support vector machine accuracy, support vector machine area under the curve, random forest accuracy, and random forest recall. Using random-effects meta-analysis with the DerSimonian-Laird method, the researchers found that support vector machines pooled at an accuracy of 0.71 and an area under the curve of 0.84, the latter indicating good discrimination across the included classification tasks. Random forests fared better on the available evidence, pooling at an accuracy of 0.86 and a recall of 0.81. Individual study estimates varied widely, however, with accuracies ranging from 0.53 to 0.96 depending on the task and dataset.
Those wide ranges point to the review’s most important caveat: heterogeneity. Between-study variability ranged from moderate to very high across analyses, reflecting genuine differences in participant populations, prediction targets, feature sets, outcome definitions, and validation procedures. Because each pooled analysis included only three or four studies, the authors urge cautious interpretation of the heterogeneity statistics and pooled estimates. Bootstrap resampling with 10,000 study-level draws produced estimates broadly consistent with the primary analyses, and sensitivity analyses, including exclusion of a large dataset of 12,697 subjects, shifted pooled values only modestly. But the evidence base remains too small and too varied to declare any single algorithm the winner.
Indeed, the review explicitly rejects the idea that one model class is universally superior. Random forests may excel at modeling nonlinear interactions among correlated clinical predictors, while support vector machine performance appears more sensitive to feature scaling and kernel specification. Neural networks showed real promise for problems involving complex spectrotemporal representations, with deep architectures outperforming conventional objective measures in speech intelligibility estimation. Yet greater complexity did not consistently help: in several studies, deep neural networks and multilayer perceptrons underperformed simpler algorithms when datasets were small relative to model capacity. The practical lesson is that neural networks suit high-dimensional acoustic or signal-level data, whereas traditional machine learning remains highly effective for smaller, structured clinical datasets.
Two other model families offered complementary perspectives. Automatic speech recognition systems, including hybrid deep neural network and hidden Markov model architectures, reproduced key patterns of human performance, such as declining recognition scores as signal-to-noise ratios worsen, and incorporating listener-specific auditory parameters improved predictions for hearing-impaired profiles. These systems also proved useful for automating the design of speech-in-noise tests themselves. Mechanistic auditory models, represented by a single study predicting consonant recognition, embedded cochlear filtering, compression, and hearing loss simulations directly into the prediction pipeline, demonstrating that reduced audibility and suprathreshold distortions jointly degrade speech perception. Such physiologically grounded models offer interpretability that purely data-driven approaches lack, and the authors suggest that future progress may lie in hybrid frameworks combining both.
Perhaps the most clinically intriguing finding concerns what the models use as inputs. Several studies showed that speech-in-noise outcomes can be predicted more effectively when models incorporate information beyond conventional audiometric thresholds. Extended high-frequency thresholds, acoustic features, behavioral measures, and even cortical speech-evoked neural responses all contributed predictive information in different analyses. Extended high-frequency hearing, in particular, was implicated in speech-in-noise variability among listeners with clinically normal audiograms, hinting that deficits invisible to the standard audiogram may underlie real-world listening difficulty. This multidimensional picture supports the hope that machine learning could eventually identify people whose communication struggles are disproportionate to their measured hearing loss.
The authors are careful, however, to distinguish proof of concept from clinical utility. Only one included study reported independent external validation on an additional dataset, and most relied on retrospective data with internal validation alone. Metrics such as accuracy, area under the curve, recall, and correlation are not directly comparable across tasks, and the small number of studies precluded publication-bias assessment and subgroup analyses. The path forward, the review concludes, requires larger, standardized, externally validated, and prospectively evaluated studies spanning multiple clinics, populations, languages, and listening environments. Until then, the models remain promising prototypes rather than diagnostic tools. Still, the trajectory is clear: as datasets grow and validation standards mature, data-driven prediction could transform how clinicians detect hidden hearing deficits, personalize hearing aids and cochlear implants, and design smarter listening technology for a world that is only getting noisier.
Subject of Research: Machine learning prediction of speech-in-noise perception in listeners with normal and impaired hearing
Article Title: Data-Driven Prediction of Speech-in-Noise Performance: A Systematic Review and Meta-analysis of Machine Learning Models
Article References: Athukorala, P., Nair, A., & Mishra, S. K. (2026). Data-Driven Prediction of Speech-in-Noise Performance: A Systematic Review and Meta-analysis of Machine Learning Models. Annals of Biomedical Engineering. https://doi.org/10.1007/s10439-026-04407-z
Image Credits: AI Generated
DOI: 10.1007/s10439-026-04407-z
Keywords: machine learning, speech-in-noise, hearing loss, support vector machine, random forest, neural networks, automatic speech recognition, auditory models, meta-analysis, audiometry, cochlear synaptopathy, hearing rehabilitation
News Source: Ophelia Keating. (October 11, 2026). Machines Learn to Predict Who Struggles to Hear in a Noisy World. Scienmag.



