When a patient arrives in the emergency room with sudden, severe vertigo, doctors face one of medicine’s most nerve-wracking dilemmas: is this a harmless inflammation of the inner ear, or the first sign of a stroke in the brainstem and cerebellum? A new study published in the Journal of Neurology suggests that machine learning algorithms can make that call with accuracy matching, and in some settings exceeding, that of specialist neurologists. The research, led by Chao Wang and Miriam Welgampola of the University of Sydney and Royal Prince Alfred Hospital together with data scientists at the University of Technology Sydney, trained artificial intelligence models on the clinical details of 294 patients and showed that the algorithms could reliably separate posterior circulation stroke from acute unilateral vestibulopathy, the two conditions that account for the vast majority of acute vestibular syndrome cases.
The stakes in this diagnostic puzzle could hardly be higher. Acute vestibular syndrome, defined as new-onset, severe and persistent vertigo or imbalance, makes up roughly 10 to 20 percent of all acute dizziness presentations to emergency departments. Posterior circulation stroke is life-threatening and demands urgent reperfusion therapy, while acute unilateral vestibulopathy, commonly known as vestibular neuritis, is a self-limiting inflammation of the vestibular nerve that resolves on its own. Yet these two conditions are among the most frequently misdiagnosed causes of dizziness, even by specialists. Focal neurological deficits, the classic red flags for stroke, are absent in as many as two-thirds of posterior circulation strokes, and early MRI can be falsely reassuring, missing about 15 percent of strokes when performed within the first 48 hours of symptom onset.
The traditional bedside answer to this problem is the HINTS examination, a three-step assessment comprising the head impulse test, nystagmus analysis and the test of skew. When performed by an experienced neuro-otologist, HINTS can exclude stroke with impressive accuracy, often outperforming early MRI. The catch is that frontline emergency clinicians, who are the ones actually evaluating these patients at 3 a.m., are rarely experts in the technique. Studies show they frequently avoid the examination, perform it incorrectly or misinterpret the results, particularly the subtle head impulse component. This gap between what is possible in expert hands and what happens in a busy emergency department is precisely what the Sydney team set out to close with algorithms rather than additional years of subspecialty training.
The researchers recruited 294 consecutive patients presenting to the emergency department of Royal Prince Alfred Hospital between March 2018 and March 2023, all reviewed within seven days of symptom onset by a clinician with neuro-otology expertise. Of these, 163 had acute unilateral vestibulopathy and 131 had imaging-confirmed posterior circulation stroke. Each patient underwent a standardised history covering demographics, vertigo characteristics, aural symptoms, focal neurological complaints and cardiovascular risk factors, a detailed bedside examination, and a battery of modern vestibular function tests: video-nystagmography to quantify spontaneous nystagmus with visual fixation removed, the video head impulse test assessing all six semicircular canals, cervical and ocular vestibular-evoked myogenic potentials probing the saccule and utricle, and the subjective visual horizontal test of utricular function. In total, 203 variables fed the modelling pipeline.
Rather than building a single model that demanded every possible test, the team deliberately designed three tiers of data availability to mirror real-world clinical settings. Tier 1 simulated a fully equipped emergency department with neuro-otology support, using history, specialist examination, video-nystagmography, horizontal canal video head impulse testing and ocular vestibular-evoked myogenic potentials. Tier 2 represented a non-expert clinician with basic examination skills and access to the video head impulse test, deliberately excluding the bedside head impulse and test of skew that trip up inexperienced examiners. Tier 3 captured the most resource-limited scenario, relying on history and basic examination alone. The researchers trialled three gradient-boosted tree algorithms, XGBoost, LightGBM and CatBoost, using five-fold stratified cross-validation, and compared the results against HINTS applied by experts to the same evaluation sets.
The results were striking. The best Tier 1 model, built on CatBoost, identified stroke with 96.6 percent accuracy, with a 95 percent confidence interval of 93.3 to 99.9 percent. The Tier 2 model, requiring no specialist examination skills at all, matched HINTS almost exactly, achieving 94.6 percent accuracy, statistically indistinguishable from the expert-performed bedside algorithm. Even the minimalist Tier 3 model, working from history and basic examination alone, reached 88.8 percent accuracy, significantly outperforming a conventional logistic regression model trained on the same data. HINTS itself achieved 94.6 percent accuracy in this cohort, confirming that the machine learning approaches were operating at the ceiling of what expert clinical assessment can deliver.
Feature importance analysis revealed that the algorithms were, reassuringly, learning clinically sensible patterns. In the richest data tier, the bedside head impulse test was the single most influential variable, followed by the slow phase velocity of spontaneous nystagmus measured on video-nystagmography and the presence of focal neurological symptoms on history. In Tier 2, where the head impulse test was unavailable, focal neurological symptoms dominated, with video head impulse metrics effectively substituting for the bedside manoeuvre. In the history-only tier, age emerged as the most powerful predictor, reflecting its status as a recognised stroke risk factor. The dataset explained why focal symptoms carried such weight: 65.6 percent of stroke patients reported them, compared with just 3.0 percent of those with vestibulopathy.
Perhaps most intriguingly, the models handled some of the trickiest atypical cases better than HINTS. Two patients with vestibulopathy but no spontaneous nystagmus on video-oculography, and two with the misleading test-of-skew finding, were correctly classified by the tiered models in several instances where HINTS failed. Among the small handful of patients misclassified by every model were two atypical vestibulopathy cases and a pontine stroke with no focal signs, no nystagmus and an abnormal bedside head impulse test, a genuinely deceptive combination. The researchers also showed that lowering the probability threshold for flagging stroke could push sensitivity above 99 percent in the better-resourced tiers, an attractive trade-off given the far greater cost of missing a stroke than of over-investigating a benign vertigo case.
The study’s implications reach well beyond the algorithm itself. The video head impulse test can be performed at the bedside in about five minutes, is well tolerated even in acutely unwell patients, and yields an objective quantified result that is easier to interpret than the naked-eye equivalent, especially when vigorous spontaneous nystagmus obscures the examination. A Tier 2-style workflow, combining structured history, basic neurological examination and video head impulse testing, could therefore provide rapid risk stratification early in the diagnostic pathway, potentially reducing unnecessary neuroimaging and the complications of inappropriate reperfusion therapy. The authors even found that video-nystagmography could substitute for the head impulse test in their Tier 2 model with similar performance, opening another accessible route for non-specialist departments.
The researchers are careful to frame this as a proof-of-concept rather than a finished clinical tool. Their history and examination data were collected by experts, which may flatter the models’ real-world performance, and the HINTS comparison benefited from video-oculographic nystagmus assessment that most emergency departments lack. The cohort came from a single tertiary centre, vestibular migraine patients were excluded, and external validation on larger, multi-centre datasets remains essential before deployment. The team also notes that practical hurdles, from user interface design and offline functionality to patient consent, data privacy and clinician adherence, must be solved before any algorithm enters the emergency workflow. Still, the central message is hard to ignore: with nothing more than a structured history, a basic examination and, ideally, a five-minute video head impulse test, machine learning can now tell a dangerous stroke from a self-limiting ear problem with the accuracy of the world’s leading dizziness specialists.
Subject of Research: Machine learning differentiation of posterior circulation stroke from acute unilateral vestibulopathy in acute vestibular syndrome
Article Title: Separating stroke and acute unilateral vestibulopathy using history, examination and vestibular tests: a machine learning approach
Article References: Wang, C., Chaturvedi, K., Nham, B., Reid, N., Bradshaw, A. P., Rosengren, S. M., Black, D. A., Bein, K. J., Halmagyi, G. M., Braytee, A., Bharathy, G. K., Prasad, M., & Welgampola, M. S. (2026). Separating stroke and acute unilateral vestibulopathy using history, examination and vestibular tests: a machine learning approach. Journal of Neurology, 273(10), Article 573. https://doi.org/10.1007/s00415-026-13779-0
Image Credits: AI Generated
DOI: 10.1007/s00415-026-13779-0
Keywords: machine learning, stroke, acute vestibular syndrome, vestibular neuritis, HINTS, video head impulse test, emergency medicine, neurology, vertigo, diagnosis, XGBoost, CatBoost
News Source: Cassandra Pierce. (October 7, 2026). AI Spots the Stroke Hiding Behind Sudden Dizziness With Expert-Level Accuracy. Scienmag.



