• HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
Wednesday, September 23, 2026
BIOENGINEER.ORG
No Result
View All Result
  • Login
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
No Result
View All Result
Bioengineer.org
No Result
View All Result
Home NEWS Science News Technology

Teaching AI to Think Before Flagging Hateful and Propagandistic Memes

Bioengineer by Bioengineer
September 23, 2026
in Technology
Reading Time: 5 mins read
0
Teaching AI to Think Before Flagging Hateful and Propagandistic Memes
Share on FacebookShare on TwitterShare on LinkedinShare on RedditShare on Telegram

Memes have become one of the most pervasive modes of communication on social media, blending images and text with humor, irony, and cultural references. While often harmless, this format can be exploited to spread hate speech, disinformation, and propaganda, and the very humor that makes memes shareable can trivialize toxic content and normalize hostile views. Detecting harmful memes automatically is notoriously difficult because their meaning rarely lives in the image or the caption alone; it emerges from the interaction between the two, often through implicit stereotypes or satirical framing that a naive classifier will miss entirely.

A new study published in Machine Learning with Applications addresses this challenge with a reasoning-centric training methodology for multimodal large language models, or MLLMs. Led by Mohamed Bayan Kmainasi of Qatar Computing Research Institute and colleagues including Mucahid Kutlu, Ali Ezzat Shahroor, Abul Hasnat, and Firoj Alam, the work demonstrates that reinforcement learning with chain-of-thought supervision can push explainable meme moderation to state-of-the-art performance on both English hateful memes and Arabic propagandistic memes. The research represents, according to the authors, the first systematic study of group relative policy optimization, known as GRPO, in multimodal reasoning under cross-lingual, fine-grained, and self-training settings.

The team’s central insight is that thinking-based MLLMs, which generate explicit intermediate reasoning steps before committing to an answer, are ideally suited to memes because their meaning depends on image-text interaction rather than unimodal cues. But whether the reasoning capabilities of such models, typically honed on mathematics and code generation, transfer to subjective, culturally situated tasks like meme moderation remained an open question. Standard supervised fine-tuning alone provides limited control over the balance between prediction correctness and rationale faithfulness, since cross-entropy loss treats all output tokens equally regardless of their functional role.

To resolve this, the researchers designed a three-stage training pipeline. The first stage is a supervised fine-tuning warm-up that aligns the model with gold labels, natural language explanations, and distilled reasoning traces produced by GPT-4.1. The second stage applies GRPO with a composite reward function that jointly optimizes classification correctness, output-format compliance, reference-based explanation similarity measured by METEOR, explanation length, and a novel thinking-length reward called Rthink. The third stage, self-training GRPO or ST-GRPO, extends the approach to unlabeled data using consensus-based pseudo-labels derived from the model’s own majority-vote predictions.

The composite reward is carefully weighted so that the two binary objectives, label correctness and format compliance, each receive 0.35 and together dominate the three auxiliary components, which contribute 0.30 combined. The thinking-length reward is particularly significant. Without it, the authors observed a consistent reward-hacking pattern: the model learned to compress or empty its reasoning traces while still collecting high reward, a length-shortening bias that was especially pronounced on the harder Arabic dataset. Rthink penalizes only reasoning traces shorter than a minimum threshold of 150 words, discouraging degenerate outputs without incentivizing verbosity.

The evaluation spans two distinct tasks and two languages. The English benchmark is the Facebook Hateful Memes dataset, containing roughly 11,000 memes where classification requires joint multimodal understanding. The Arabic benchmark, ArMeme, contains about 5,700 memes with four labels covering propaganda, not-propaganda, not-meme, and other. Because no fine-grained propaganda annotations existed for ArMeme, the team built a dual-annotator pipeline using GPT-4.1 and Llama-4-Scout as independent labelers of 23 propaganda techniques, consolidated by Gemini-3-Pro as an arbiter. Human validation on 584 memes showed the consolidated annotations aligned better with human reference labels than either single-model source.

The results are striking. On the Hateful Memes benchmark, the best supervised GRPO configuration with thinking-length regularization achieved 82.0 percent accuracy and 0.80 macro-F1, outperforming prior reported results and strong sequence-classification baselines such as Qwen3-VL-8B-Instruct and Gemma-3-12B-IT. On ArMeme, self-training GRPO reached 0.612 macro-F1, improving over previous work by 7.6 points and over the original ArMeme benchmark by 6.1 points. Notably, unimodal baselines lagged far behind: the best text-only model on ArMeme reached only 0.509 macro-F1, and image-only models averaged just 0.267, confirming that cross-modal reasoning captures signals that neither modality provides alone.

The self-training stage showed a telling asymmetry. On ArMeme, where unlabeled data was collected from the same social media sources as the labeled set, ST-GRPO improved macro-F1 by 1.5 points with gains concentrated in minority classes. On the English dataset, it slightly degraded performance, a result the authors attribute to distribution mismatch between the unlabeled pool and the benchmark, and to majority-vote bias amplification in binary classification, where small prediction biases can produce near-unanimous consensus that reinforces rather than corrects class skew. The finding suggests that consensus-based pseudo-labeling requires both distribution alignment and sufficient label-space diversity to provide reliable supervision.

Beyond raw accuracy, the model produces natural language explanations alongside its predictions, a property that sequence classifiers lack. Modality ablations showed the model genuinely depends on both inputs, with image removal hurting the English task most and OCR text removal crippling Arabic propaganda detection. An LLM-as-judge evaluation using GPT-4.1 and Gemini-2.5-Pro found the trained models substantially outperformed the zero-shot backbone on grounding, correctness, and usefulness for moderation, approaching the quality of human-verified reference explanations. The team has released all code, data extensions, prompting templates, and evaluation resources on GitHub.

The implications extend well beyond memes. The study shows that fine-grained supervision and distilled chain-of-thought rationales complement reinforcement-learning-based optimization, that multi-LLM annotation pipelines can scale fine-grained labeling to previously unlabeled domains, and that even a simple thresholded reasoning-length penalty can stabilize RL training against reward hacking in multimodal settings. For content moderation at scale, where subjective interpretation of culturally embedded imagery is the daily reality, the work offers a concrete template for building systems that not only flag harmful content but also articulate why, enabling meaningful human review rather than opaque automated verdicts.

Subject of Research: Reinforcement learning with chain-of-thought supervision for explainable detection of hateful and propagandistic memes

Article Title: Adapting reinforcement learning with chain-of-thought supervision for explainable detection of hateful and propagandistic memes

Article References: Kmainasi, M. B., Kutlu, M., Shahroor, A. E., Hasnat, A., & Alam, F. (2026). Adapting reinforcement learning with chain-of-thought supervision for explainable detection of hateful and propagandistic memes. Machine Learning with Applications, 26, Article 101003. https://doi.org/10.1016/j.mlwa.2026.101003

Image Credits: AI Generated

DOI: 10.1016/j.mlwa.2026.101003

Keywords: reinforcement learning, chain-of-thought, multimodal large language models, hateful memes, propaganda detection, content moderation, GRPO, self-training, explainable AI, Arabic memes, reward hacking, fine-grained annotation

Cite Scienmag News
APA MLA Chicago

Denise Maddox. (September 23, 2026). Teaching AI to Think Before Flagging Hateful and Propagandistic Memes. Scienmag. https://scienmag.com/teaching-ai-to-think-before-flagging-hateful-and-propagandistic-memes/

Denise Maddox. “Teaching AI to Think Before Flagging Hateful and Propagandistic Memes.” Scienmag, 23 September 2026, https://scienmag.com/teaching-ai-to-think-before-flagging-hateful-and-propagandistic-memes/. Accessed 23 September 2026.

Denise Maddox. “Teaching AI to Think Before Flagging Hateful and Propagandistic Memes.” Scienmag. September 23, 2026. https://scienmag.com/teaching-ai-to-think-before-flagging-hateful-and-propagandistic-memes/

Copy citation Download RIS

Tags: AI-driven social media content filteringArabic memeschain-of-thoughtchain-of-thought supervision in AI modelschallenges in automatic harmful content recognitioncontent moderationcross-lingual meme analysisexplainable AIexplainable AI for social media moderationfine-grained annotationgroup relative policy optimization in AIGRPOHate speech detection in memeshateful memesmultilingual meme moderation techniquesmultimodal large language modelsmultimodal large language models for content moderationpropaganda and disinformation detection in memespropaganda detectionreasoning-based AI training for harmful content identificationreinforcement learningreinforcement learning in hate speech detectionreward hackingself-training

Share12Tweet7Share2ShareShareShare1

Related Posts

One Nanomaterial, Two Jobs: MOF-Derived Cobalt Ferrite Hybrid Cleans Water and Boosts Solar Cells

One Nanomaterial, Two Jobs: MOF-Derived Cobalt Ferrite Hybrid Cleans Water and Boosts Solar Cells

September 23, 2026
Intraosseous Access in Newborns: New European Standard Aimed at Saving Lives

Intraosseous Access in Newborns: New European Standard Aimed at Saving Lives

September 23, 2026

New AI Method Spots Hidden Threats in Social Networks Using Static and Dynamic Clues

September 23, 2026

High-Strength Aluminum Alloy Matches Steel Strength at One-Third the Weight, New Joint Tests Show

September 23, 2026

POPULAR NEWS

  • Plant Arginine Mimic Canavanine Disrupts Cancer Cell Metabolism and Signals

    29 shares
    Share 12 Tweet 7
  • Bone Crystals and the Clock of Death: X-ray Study Tests a Forensic Dating Dream

    29 shares
    Share 12 Tweet 7
  • Hidden Fractal Geometry Explains the Strange Scaling Laws of Cities

    29 shares
    Share 12 Tweet 7
  • One Nanomaterial, Two Jobs: MOF-Derived Cobalt Ferrite Hybrid Cleans Water and Boosts Solar Cells

    29 shares
    Share 12 Tweet 7

About

We bring you the latest biotechnology news from best research centers and universities around the world. Check our website.

Follow us

Recent News

Plant Arginine Mimic Canavanine Disrupts Cancer Cell Metabolism and Signals

Bone Crystals and the Clock of Death: X-ray Study Tests a Forensic Dating Dream

Hidden Fractal Geometry Explains the Strange Scaling Laws of Cities

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 85 other subscribers
  • Contact Us

Bioengineer.org © Copyright 2023 All Rights Reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Homepages
    • Home Page 1
    • Home Page 2
  • News
  • National
  • Business
  • Health
  • Lifestyle
  • Science

Bioengineer.org © Copyright 2023 All Rights Reserved.