• HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
Friday, October 9, 2026
BIOENGINEER.ORG
No Result
View All Result
  • Login
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
No Result
View All Result
Bioengineer.org
No Result
View All Result
Home NEWS Science News Technology

Attention-Based AI Reads News Word by Word to Catch Multilingual Fake Stories

by
October 9, 2026
in Technology
Reading Time: 5 mins read
0
Attention-Based AI Reads News Word by Word to Catch Multilingual Fake Stories

Attention-Based AI Reads News Word by Word to Catch Multilingual Fake Stories

Share on FacebookShare on TwitterShare on LinkedinShare on RedditShare on Telegram

Fake news does not respect borders, and it certainly does not respect languages. While social media platforms have invested heavily in automated misinformation detection for English content, the vast majority of the world’s languages remain poorly served by these systems. A new study published in Discover Artificial Intelligence by Sushama Nandgaonkar and Sunil Mane of COEP Technological University in Pune tackles this gap head-on, presenting a Hierarchical Attention Network that detects false news across English, Hindi, and Marathi with remarkable accuracy, achieving 98.9 percent accuracy and a 98.8 percent F1-score on a combined multilingual test set.

The stakes are particularly high in India, which has the second-largest population of internet users worldwide. According to figures cited in the study, India had more than 375 million social media users in 2019, rising to 518 million in 2020, with projections suggesting the number could reach 1.5 billion by 2040. Affordable mobile data has driven this explosive growth, and platforms such as Facebook, Twitter, and WhatsApp now carry content in dozens of languages. Misinformation that circulates in regional languages can provoke local sentiments and influence health decisions, as demonstrated during the COVID-19 pandemic, when false claims about remedies and vaccines spread faster than official corrections could keep up.

The researchers distinguish carefully between related concepts that are often conflated. Disinformation refers to deliberately created and spread falsehoods, rumors are unverified pieces of information circulated with intent to mislead, and misinformation is false content shared without necessarily intending to deceive. The motivations behind such content range from spreading religious hatred to gaining political advantage or disseminating false health information. Because fact-checking websites themselves can carry biases, and because human verification cannot scale to billions of posts, automated detection systems have become an essential line of defense, and the study argues that these systems must work across the languages people actually use.

At the heart of the new work is a two-level attention architecture that mirrors how humans read documents. The model first processes each sentence word by word using a Bidirectional Long Short-Term Memory network, which reads text in both forward and backward directions to capture context from either side of every word. A custom attention layer then assigns dynamic weights to individual words, learning during training which terms matter most for distinguishing fake from genuine content. The weighted word representations are combined into sentence vectors, and a second Bi-LSTM layer with its own attention mechanism models relationships between sentences, producing a document-level representation that feeds into a final sigmoid classifier.

This hierarchical design is not merely an engineering convenience. Fake news articles often contain misleading lexical patterns, emotionally polarized expressions, and contextual inconsistencies distributed across multiple sentences rather than concentrated in any single phrase. By attending to both words and sentences, the model can capture local semantic cues and document-level structure simultaneously. The attention weights also improve interpretability, allowing researchers to see which parts of an article drove a classification decision, a significant advantage over opaque black-box approaches.

A crucial contribution of the study is linguistic rather than architectural. Because no publicly available fake news dataset existed for Marathi, the team built one from scratch, collecting 1,957 news articles from Marathi outlets including Loksatta, Lokmat, Maharashtra Times, and Zee News, of which 707 were labeled fake and 1,250 true. Labels were assigned based on source credibility, fact-checking reports, and consistency across multiple platforms, with manual review of every article. For Hindi, the researchers combined two existing datasets from Kaggle and GitHub, while English experiments used the widely adopted ISOT dataset. To address severe class imbalance in the low-resource languages, the team applied translation-based augmentation using the IndicTrans neural translation framework, generating additional training samples while keeping the test set untouched to prevent data leakage. The final augmented corpus contained 45,386 English, 10,481 Hindi, and 5,000 Marathi samples.

The choice of word embeddings proved important. FastText, developed by Facebook’s AI Research lab, represents each word as a bag of character n-grams, allowing it to capture morphological information and generate vectors for out-of-vocabulary words from their sub-word components. This property is especially valuable for morphologically rich languages like Hindi and Marathi, where word forms vary extensively. The researchers compared FastText against GloVe and random initialization under identical hyperparameters, finding that pretrained embeddings substantially improved performance, with FastText also converging fastest at 383.06 seconds of training time compared with 501.22 seconds for GloVe.

The experimental results were striking. The HAN model achieved 99.42 percent accuracy on English news, 99.04 percent on Hindi, and 80.90 percent on Marathi, outperforming traditional machine learning baselines including Logistic Regression, Linear Support Vector Machines, and Random Forest, as well as deep learning models such as CNN and standalone Bi-LSTM. Against multilingual transformers, the picture was more nuanced: mBERT reached 98.72 percent overall accuracy and DeBERTa-v3-base 98.66 percent, slightly below the proposed model’s 98.90 percent, but DeBERTa’s performance collapsed to 74.42 percent on Marathi, exposing the sensitivity of large pretrained transformers to data imbalance. Notably, the instruction-tuned LLaMA 3.1 8B model, evaluated through prompting rather than fine-tuning, managed only 67.98 percent accuracy in zero-shot settings and 84.76 percent with six examples in context, with 1,638 failed predictions out of 10,965 test instances.

Ablation experiments confirmed that both attention levels contribute meaningfully. A Bi-LSTM-only model achieved an F1-score of 97.43 percent, rising to 97.82 percent with word-level attention alone and 98.17 percent with sentence-level attention alone, while the complete hierarchical architecture reached 98.72 percent. Sentence-level attention proved more influential than word-level attention in isolation, suggesting that contextual dependencies across sentences play a decisive role in identifying deceptive content. Paired t-tests across multiple random seeds showed all improvements over baselines were statistically significant, with p-values below 0.01.

The study is candid about its limitations. Marathi’s error rate of 19.10 percent, compared with just 0.58 percent for English and 0.96 percent for Hindi, reflects the challenge of low-resource detection, and qualitative analysis showed that the worst failures involved fake articles written in a style so similar to legitimate news that the model assigned incorrect predictions with high confidence. The authors note that generalization to other Indic languages remains unexplored and that multimodal signals such as images and metadata are not yet incorporated. Future work will expand the dataset to additional Indian regional languages and integrate textual and visual information using transformer-based Indic language models. For now, the research demonstrates that carefully designed attention mechanisms, paired with sub-word embeddings and thoughtful data augmentation, can bring state-of-the-art misinformation detection to languages that have long been left behind.

Subject of Research: Multilingual fake news detection using hierarchical attention networks for English, Hindi, and Marathi news articles

Article Title: Hierarchical attention network for multilingual fake news detection

Article References: Nandgaonkar, S., & Mane, S. (2026). Hierarchical attention network for multilingual fake news detection. Discover Artificial Intelligence, 6(1), Article 1417. https://doi.org/10.1007/s44163-026-02425-3

Image Credits: AI Generated

DOI: 10.1007/s44163-026-02425-3

Keywords: fake news detection, hierarchical attention network, multilingual NLP, Bi-LSTM, FastText embeddings, Marathi language, Hindi language, misinformation, deep learning, low-resource languages, data augmentation, natural language processing

News Source: Denise Maddox. (October 9, 2026). Attention-Based AI Reads News Word by Word to Catch Multilingual Fake Stories. Scienmag.

Tags: Bi-LSTMdata augmentationdeep learningFake News DetectionFastText embeddingshierarchical attention networkHindi languagelow-resource languagesMarathi languagemisinformationmultilingual NLPNatural Language Processing
Share12Tweet7Share2ShareShareShare1

Related Posts

Quantum Noise Becomes a Friend: Physicists Turn Hardware Errors Into Generative AI Fuel

Quantum Noise Becomes a Friend: Physicists Turn Hardware Errors Into Generative AI Fuel

October 9, 2026
Sunken Ice Age World Found Beneath North Sea Wind Farm Cores

Sunken Ice Age World Found Beneath North Sea Wind Farm Cores

October 9, 2026

Wind Veer Reshapes Turbine Wakes and May Outperform Yaw Control, Wind Tunnel Study Finds

October 9, 2026

Yolk-Shell Silicon Anode Reinforced with Bimetallic MOF-Derived Carbon Boosts Battery Durability

October 9, 2026

POPULAR NEWS

  • Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

    Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

    29 shares
    Share 12 Tweet 7
  • Endurance Exercise Reshapes the Liver in Males and Females Through Distinct Molecular Routes

    29 shares
    Share 12 Tweet 7
  • Single Transcription Factor PU.1 Rapidly Converts Fibroblasts into Macrophage-Lineage Cells

    29 shares
    Share 12 Tweet 7
  • New Scale Measures How Ready Nurse Educators Really Are for the AI Era

    29 shares
    Share 12 Tweet 7

About

We bring you the latest biotechnology news from best research centers and universities around the world. Check our website.

Follow us

Recent News

Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

Endurance Exercise Reshapes the Liver in Males and Females Through Distinct Molecular Routes

Single Transcription Factor PU.1 Rapidly Converts Fibroblasts into Macrophage-Lineage Cells

Subscribe to Blog via Email

Success! An email was just sent to confirm your subscription. Please find the email now and click 'Confirm' to start subscribing.

Join 85 other subscribers
  • Contact Us

Bioengineer.org © Copyright 2023 All Rights Reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Homepages
    • Home Page 1
    • Home Page 2
  • News
  • National
  • Business
  • Health
  • Lifestyle
  • Science

Bioengineer.org © Copyright 2023 All Rights Reserved.