• HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
Sunday, October 11, 2026
BIOENGINEER.ORG
No Result
View All Result
  • Login
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
No Result
View All Result
Bioengineer.org
No Result
View All Result
Home NEWS Science News Technology

AI That Listens and Reads: Multimodal Deep Learning Scores Spoken English With 97.5% Accuracy

by
October 11, 2026
in Technology
Reading Time: 4 mins read
0
AI That Listens and Reads: Multimodal Deep Learning Scores Spoken English With 97.5% Accuracy

AI That Listens and Reads: Multimodal Deep Learning Scores Spoken English With 97.5% Accuracy

Share on FacebookShare on TwitterShare on LinkedinShare on RedditShare on Telegram

Grading spoken English has long been one of the most stubborn bottlenecks in language education. Human raters are expensive, slow, and notoriously inconsistent, with scores that can vary depending on the assessor’s mood, background, or expectations. Now a study published in Discover Artificial Intelligence describes a multimodal deep learning framework that automates oral English fluency assessment by listening to speech and reading its transcript at the same time, achieving 97.53% accuracy in classifying learners’ proficiency levels.

The framework, developed by Ping Zhang of the Shanghai University of Political Science and Law, rests on a simple but powerful insight: fluency is not just a property of sound or of words, but of both together. Most existing automated systems analyze either the acoustic signal or the transcribed text in isolation, missing the interplay between how something is said and what is actually said. By fusing both streams of information, the new model captures pronunciation, rhythm, prosody, semantic coherence, and even emotional expression within a single end-to-end pipeline.

Technically, the system processes each speech recording in two parallel branches. The audio is normalized, filtered for background noise, and segmented using voice activity detection before being converted into a Log-Mel spectrogram, a time-frequency representation that maps frequencies onto a scale approximating human auditory perception. This representation preserves rich spectral detail, including pitch contours and pause durations, that coarser features like MFCCs tend to discard. Meanwhile, the corresponding transcript is tokenized and passed through BERT, the transformer-based language model, whose bidirectional self-attention produces contextual embeddings that encode the meaning of each word in relation to the whole sentence.

The heart of the framework is its fusion strategy. Rather than simply concatenating acoustic and semantic features and hoping for the best, the model applies a self-attention mechanism over the combined feature vector. This mechanism computes query, key, and value projections, scores the relevance of every feature to every other feature, and produces a weighted representation that emphasizes whichever pieces of information matter most for judging fluency. The authors chose self-attention over cross-attention because it captures dependencies between the already-merged modalities without adding computational complexity.

The fused representation is then fed into a multilayer perceptron with a softmax output layer that classifies each response into low, medium, or high fluency. Training relied on the Adam optimizer, ReLU activations in the hidden layers, and dropout regularization to prevent overfitting. Notably, the pipeline was built across two deep learning ecosystems: TensorFlow handled model training and optimization, while PyTorch implemented the transformer-based BERT embeddings, exploiting the strengths of both frameworks in one system.

The experiments drew on a Kaggle dataset of 8,520 English speech samples from 426 speakers, roughly balanced between male and female, and spanning American, British, Indian, Chinese, and other accents. The recordings total 28.1 hours, with utterances averaging 11.8 seconds, and are split into 5,964 training, 1,278 validation, and 1,278 testing samples across the three fluency classes. To harden the model against real-world messiness, the researchers applied audio augmentation techniques such as time stretching and white noise injection, and used ANOVA-based feature selection to retain only the most discriminative acoustic and semantic features.

The results are striking. The framework achieved 97.53% accuracy, 97.68% precision, 97.53% recall, and a 97.52% F1-score, outperforming single-modal baselines and well-known architectures including CNNs, LSTMs, BiLSTMs, SpeechTransformer, wav2vec 2.0, HuBERT, and Whisper. An ablation study underscores the value of the multimodal design: speech-only and text-only versions reached just 90.52% and 91.81% accuracy respectively, and removing the self-attention module, BERT, or the Log-Mel features each dragged performance down by several percentage points. The model also correlated strongly with human expert ratings, with Pearson and Spearman correlations of 0.941 and 0.932, and quadratic weighted kappa of 0.921.

The system is not lightweight. It carries 112.4 million parameters, occupies 427 MB, and demands 5.6 GB of GPU memory, though inference is fast at 21 milliseconds per sample. The authors are candid about the limitations: performance can degrade with poor-quality audio, heavy background noise, or erroneous speech recognition, and the framework was validated on a single dataset. Fairness across accents remains an open question, and no formal usability study with students and educators has yet been conducted, though the authors flag user-centered testing as a priority for future work.

Even so, the implications for education are considerable. English is the world’s lingua franca, and demand for scalable, objective speaking assessment far exceeds the supply of trained human raters. A framework that scores fluency, pronunciation, rhythm, and expressiveness consistently and instantly could be embedded in online courses, virtual classrooms, and language-learning apps, giving learners detailed feedback that no exam hall could provide at scale. The authors envision future deployment on cloud and edge platforms, validation on larger multilingual datasets, and adaptation to languages beyond English. If those steps succeed, the era of waiting weeks for a speaking test score may finally be drawing to a close.

Subject of Research: Automated oral English fluency assessment using multimodal deep learning with acoustic and semantic feature fusion

Article Title: A multimodal deep learning framework for automated oral English fluency assessment

Article References: Zhang, P. (2026). A multimodal deep learning framework for automated oral English fluency assessment. Discover Artificial Intelligence, 6(1), Article 1387. https://doi.org/10.1007/s44163-026-02325-6

Image Credits: AI Generated

DOI: 10.1007/s44163-026-02325-6

Keywords: multimodal deep learning, oral English fluency, automated assessment, Log-Mel spectrograms, BERT embeddings, self-attention, speech processing, natural language processing, educational technology, speech recognition, language testing, artificial intelligence

News Source: Blake Davidson. (October 11, 2026). AI That Listens and Reads: Multimodal Deep Learning Scores Spoken English With 97.5% Accuracy. Scienmag.

Tags: Artificial Intelligenceautomated assessmentBERT embeddingsEducational Technologylanguage testingLog-Mel spectrogramsmultimodal deep learningNatural Language Processingoral English fluencyself-attentionspeech processingspeech recognition
Share12Tweet7Share2ShareShareShare1

Related Posts

Lactate Marks on Histones Drive Wilms Tumor Growth Through a Self-Reinforcing Genetic Loop

Lactate Marks on Histones Drive Wilms Tumor Growth Through a Self-Reinforcing Genetic Loop

October 11, 2026
Rolling Between Critical Temperatures Keeps Nano-Oxides Fine in Nuclear Steel

Rolling Between Critical Temperatures Keeps Nano-Oxides Fine in Nuclear Steel

October 11, 2026

AI Network Spots Hidden Wind Turbine Blade Damage From Drone Photos

October 11, 2026

Dead Probiotic Cells Show Promise for Healing Infected Burn Wounds in Mice

October 11, 2026

POPULAR NEWS

  • Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

    Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

    29 shares
    Share 12 Tweet 7
  • Endurance Exercise Reshapes the Liver in Males and Females Through Distinct Molecular Routes

    29 shares
    Share 12 Tweet 7
  • Single Transcription Factor PU.1 Rapidly Converts Fibroblasts into Macrophage-Lineage Cells

    29 shares
    Share 12 Tweet 7
  • New Scale Measures How Ready Nurse Educators Really Are for the AI Era

    29 shares
    Share 12 Tweet 7

About

We bring you the latest biotechnology news from best research centers and universities around the world. Check our website.

Follow us

Recent News

Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

Endurance Exercise Reshapes the Liver in Males and Females Through Distinct Molecular Routes

Single Transcription Factor PU.1 Rapidly Converts Fibroblasts into Macrophage-Lineage Cells

Subscribe to Blog via Email

Success! An email was just sent to confirm your subscription. Please find the email now and click 'Confirm' to start subscribing.

Join 85 other subscribers
  • Contact Us

Bioengineer.org © Copyright 2023 All Rights Reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Homepages
    • Home Page 1
    • Home Page 2
  • News
  • National
  • Business
  • Health
  • Lifestyle
  • Science

Bioengineer.org © Copyright 2023 All Rights Reserved.