• HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
Sunday, October 4, 2026
BIOENGINEER.ORG
No Result
View All Result
  • Login
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
No Result
View All Result
Bioengineer.org
No Result
View All Result
Home NEWS Science News Technology

Fine-tuned AI models learn to name the target of online hate

Bioengineer by Bioengineer
October 4, 2026
in Technology
Reading Time: 5 mins read
0
Fine-tuned AI models learn to name the target of online hate
Share on FacebookShare on TwitterShare on LinkedinShare on RedditShare on Telegram

Hate speech on social media has long been treated by automated moderation systems as a single, undifferentiated problem: a post is either hateful or it is not. A new study argues that this binary framing is precisely why so many moderation tools fail the communities that need them most. Researchers Sanaa Kaddoura of Zayed University and Sumaia Al-Kohlani of United Arab Emirates University have systematically evaluated how well modern large language models can classify hateful tweets not just as hate, but by the specific group being targeted, spanning gender, race and ethnicity, immigration and xenophobia, religion, and general abuse. Their findings, published in Discover Artificial Intelligence, show that supervised fine-tuning transforms these models from unreliable label-generators into the strongest classifiers yet tested on this task.

The motivation for the work lies in a structural weakness of earlier approaches. Traditional hate speech detectors, built on support vector machines, logistic regression, convolutional and recurrent neural networks, and eventually transformer models such as BERT, were trained on datasets that collapsed all forms of hate into one category. That conflation forces models to rely on surface-level lexical cues rather than deeper contextual meaning, and the consequences are uneven: hate directed at minority or rarely referenced groups gets misclassified far more often than hate against well-represented targets. Existing multiclass datasets offered some improvement, but most were small, inconsistently annotated, or still failed to model the target of abuse explicitly, creating a persistent trade-off between scale, label granularity, and annotation quality.

To test what large language models could add, the team assembled a dataset of 12,957 English tweets, built by re-annotating the hate speech subset of the widely used TweetEval benchmark. Of these, 5,455 samples carry one of five target-specific hate labels and 7,502 are non-hateful. Three independent annotators labeled each instance, with a fourth resolving disagreements, achieving a moderate inter-annotator agreement of 0.6003 on Krippendorff’s alpha, a figure the authors note reflects the inherent subjectivity of judging hateful content. Because some original categories were too sparse to train on reliably, the researchers consolidated them: threats and violence, with only 106 instances, was merged into profanity and general abuse, while hate toward countries, with just 18 samples, was folded into racial and ethnic hate. The result is a six-class problem that mirrors the imbalance of real platforms, where non-hate content makes up nearly 58 percent of the data and religious hate speech barely 1.3 percent.

Five models with distinct architectures and training histories were put through their paces: DeepSeek-R1-Distill-Qwen-32B, Mistral-7B-Instruct, WizardLM-13B, Phi-4-mini, and OpenAI’s GPT-4o, alongside a fine-tuned BERT-base-uncased baseline. Each was evaluated three ways: zero-shot prompting, few-shot prompting with six labeled examples, and supervised fine-tuning. The open-source models were adapted using QLoRA, a parameter-efficient technique that freezes the pretrained weights and trains only small low-rank adapter matrices inserted into each transformer block, with the base model loaded in 4-bit precision to fit on a single NVIDIA A100 GPU. GPT-4o was fine-tuned through OpenAI’s API at a total training cost of 85.75 dollars.

The zero-shot results were sobering. Every model struggled, with macro F1 scores, which weight all classes equally and thus expose weakness on minority categories, remaining low across the board. GPT-4o led with a macro F1 of 0.4503, while WizardLM, which had performed well in an earlier benchmark, scored close to zero on this dataset. More troubling still, several models refused to stay within the label space: WizardLM produced 1,574 unclassified outputs, including refusals, reasoning artifacts, and requests for context, while Mistral managed only four. Few-shot prompting helped some models, improving Phi and Mistral, but actually degraded DeepSeek and GPT-4o, demonstrating that in-context examples alone do not guarantee better behavior on complex classification tasks.

Fine-tuning changed the picture dramatically. The fine-tuned DeepSeek model achieved the best multiclass performance of any system tested, with macro and weighted F1 scores of 59.91 and 76.71 percent respectively, and it eliminated unclassified outputs entirely. Mistral followed closely with a weighted F1 of 74.12 percent, while fine-tuned GPT-4o, despite its strong showing in prompting settings, reached only 61.54 percent. The replicated BERT baseline remained competitive at 70.89 percent weighted F1 but was clearly outperformed by the best LLM. Against a suite of replicated classical machine learning and deep learning baselines, including logistic regression, naive Bayes, CNN, and BiLSTM models, the fine-tuned DeepSeek improved weighted F1 by margins ranging from 2.59 to 29.5 percentage points. McNemar statistical tests confirmed the differences were significant against BERT, BiLSTM, and logistic regression.

Crucially, the team did not stop at same-dataset evaluation, a common weakness in prior work. They assessed the best DeepSeek model on a remapped subset of the external ETHOS benchmark, aligning its gender, religion, and race labels with the study’s taxonomy and yielding 789 evaluation instances. The model scored an F1 of 88.31 percent on this external test, with no unclassified outputs, evidence of genuine cross-dataset generalization. A control experiment sharpened the interpretation: when a DeepSeek model was fine-tuned directly on ETHOS, it performed well on ETHOS itself but dropped sharply when tested on the original multiclass dataset, suggesting that the multiclass-trained model had learned richer, more transferable linguistic patterns rather than dataset-specific shortcuts.

The gains came with caveats. Per-class analysis revealed that minority categories remained hard: racial and ethnic hate speech and religious hate speech, each under two percent of the data, achieved F1 scores of only 0.4211 and 0.5098 respectively, while immigration and xenophobic hate, the best-annotated class, reached 0.7624. Confusion matrices showed the dominant error was not mixing up hate categories but distinguishing hate from non-hate at all, particularly for implicit or context-dependent hostility that lacks overt slurs. Computational cost also matters: the DeepSeek model required 4.71 hours of training and 32.5 minutes to classify the 1,620-tweet test set, whereas logistic regression completed both in seconds. The authors frame this as a practical trade-off, recommending LLMs where fine-grained discrimination justifies the expense and simpler classifiers for high-throughput, resource-constrained settings.

The study also confronts the ethics of deployment. False positives can suppress counter-speech, satire, or educational discussion, and may disproportionately silence the very communities these systems claim to protect, while false negatives leave targeted groups exposed. The authors stress that model predictions should never serve as the sole basis for content removal or user sanctions, and they call for human oversight, transparent governance, and periodic auditing. Looking forward, they outline plans to extend the framework to multilingual settings, adopt multi-label formulations for posts targeting multiple groups, and incorporate community-centered evaluation. The fine-tuned model and dataset have been released publicly, inviting the research community to build on a benchmark that finally asks not just whether a post is hateful, but whom it harms.

Subject of Research: Fine-tuning large language models for target-specific multiclass hate speech detection on social media

Article Title: Fine-tuning large language models for binary and multiclass target-specific hate speech detection

Article References: Kaddoura, S., & Al-Kohlani, S. (2026). Fine-tuning large language models for binary and multiclass target-specific hate speech detection. Discover Artificial Intelligence, 6(1), Article 1296. https://doi.org/10.1007/s44163-026-02354-1

Image Credits: AI Generated

DOI: 10.1007/s44163-026-02354-1

Keywords: large language models, hate speech detection, fine-tuning, QLoRA, DeepSeek, BERT, multiclass classification, content moderation, natural language processing, social media, ETHOS dataset, class imbalance

Cite Scienmag News
APA MLA Chicago

Denise Maddox. (October 4, 2026). Fine-tuned AI models learn to name the target of online hate. Scienmag. https://scienmag.com/fine-tuned-ai-models-learn-to-name-the-target-of-online-hate/

Denise Maddox. “Fine-tuned AI models learn to name the target of online hate.” Scienmag, 4 October 2026, https://scienmag.com/fine-tuned-ai-models-learn-to-name-the-target-of-online-hate/. Accessed 4 October 2026.

Denise Maddox. “Fine-tuned AI models learn to name the target of online hate.” Scienmag. October 4, 2026. https://scienmag.com/fine-tuned-ai-models-learn-to-name-the-target-of-online-hate/

Copy citation Download RIS

Tags: AI bias in hate speech detectionBERTchallenges in automated hate speech moderationclass imbalancecontent moderationDeepSeekdifferentiating hate speech targetsETHOS datasetfine-tuningfine-tuning AI for hate group detectionHate speech classificationhate speech detectionimproving hate speech classifierslarge language modelsmachine learning for online hatemulticlass classificationnatural language processingQLoRAsocial mediasocial media moderationsupervised learning in hate speech detectiontargeted hate speech identificationtransformer models in hate speech analysis

Share12Tweet7Share2ShareShareShare1

Related Posts

Kitchen Blender Beats Ball Milling in Greener Route to Superstrong Conductive Nanocomposites

Kitchen Blender Beats Ball Milling in Greener Route to Superstrong Conductive Nanocomposites

October 4, 2026
New JMIR Cardio Section Seeks Research on Generative and Multimodal AI in Heart Care

New JMIR Cardio Section Seeks Research on Generative and Multimodal AI in Heart Care

October 4, 2026

Deep Mines Run Hot: Why 40°C Makes Coal Waste Concrete Stronger, Then Weaker

October 4, 2026

Plant Tannins Reshape Soil Minerals Into Super Fertilizers That Boost Crops

October 4, 2026

POPULAR NEWS

  • Scientists Find the Perfect Way to Brew a Rare Chinese Bud Tea

    29 shares
    Share 12 Tweet 7
  • Kitchen Blender Beats Ball Milling in Greener Route to Superstrong Conductive Nanocomposites

    29 shares
    Share 12 Tweet 7
  • Camouflaged Nanoparticles Could Rewrite the Rules of Cancer Vaccines

    29 shares
    Share 12 Tweet 7
  • New JMIR Cardio Section Seeks Research on Generative and Multimodal AI in Heart Care

    29 shares
    Share 12 Tweet 7

About

We bring you the latest biotechnology news from best research centers and universities around the world. Check our website.

Follow us

Recent News

Scientists Find the Perfect Way to Brew a Rare Chinese Bud Tea

Kitchen Blender Beats Ball Milling in Greener Route to Superstrong Conductive Nanocomposites

Camouflaged Nanoparticles Could Rewrite the Rules of Cancer Vaccines

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 85 other subscribers
  • Contact Us

Bioengineer.org © Copyright 2023 All Rights Reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Homepages
    • Home Page 1
    • Home Page 2
  • News
  • National
  • Business
  • Health
  • Lifestyle
  • Science

Bioengineer.org © Copyright 2023 All Rights Reserved.