• HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
Wednesday, October 7, 2026
BIOENGINEER.ORG
No Result
View All Result
  • Login
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
No Result
View All Result
Bioengineer.org
No Result
View All Result
Home NEWS Science News Biology

AI Learns to Name the Bully, the Victim and the Defender in Online Abuse

by
October 7, 2026
in Biology
Reading Time: 5 mins read
0
AI Learns to Name the Bully, the Victim and the Defender in Online Abuse

AI Learns to Name the Bully, the Victim and the Defender in Online Abuse

Share on FacebookShare on TwitterShare on LinkedinShare on RedditShare on Telegram

Cyberbullying has long been treated by automated systems as a blunt binary: a post is either harmful or it is not. But a new study published in Heliyon argues that this framing misses the most important question. When an online attack unfolds, who is doing what? Is the author of a message the aggressor, the target, or a bystander stepping in to defend someone? A research team led by Teoh Hwai Teng and Kasturi Dewi Varathan, working with Fabio Crestani, has now built machine learning models that attempt to answer exactly that, classifying participants in cyberbullying episodes into distinct roles and, in the process, setting a new performance benchmark on one of the field’s most respected datasets.

The significance of the role question goes back to foundational work in bullying research. In the 1990s, psychologists studying school bullying identified a cast of characters beyond the bully and the victim: assistants who support the aggressor, reinforcers who egg the behavior on, defenders who protect the target, and outsiders who simply look away. Cyberbullying inherits this social geometry, but with a twist. Because it unfolds across digital platforms, it is more pervasive, and the bystanders who witness it can respond in seconds, either with hostile retaliation against the bully or with supportive, victim-focused intervention. Prior research has shown that roughly half of university students involved in cyberbullying occupy dual roles, acting as both bully and victim, which makes the social dynamics even harder to untangle.

Most automated detection systems, however, have focused almost exclusively on the bully’s voice. The new study tackles the harder, less-studied problem of identifying victims and bystanders from text alone, a task constrained by the fact that privacy rules typically block researchers from accessing user metadata. The team worked with the AMiCA corpus, a collection of 113,694 English posts gathered from the question-and-answer platform Ask.fm between April and October 2013 by a Belgian research project. Ask.fm was popular among teenagers and frequently associated with cyberbullying, making it an ideal hunting ground. After cleaning, the dataset contained 112,247 posts annotated with roles: non-bully, harasser, victim, and bystander defender. The distribution was starkly imbalanced, with more than 106,000 non-bully posts but only 3,596 harasser posts, 1,354 victim posts, and 425 bystander defender posts.

The researchers pursued two parallel strategies. The first was conventional machine learning, in which features must be engineered by hand from the raw text. This was no small undertaking: the team built roughly 620,000 attributes spanning six feature families. Textual features included word- and character-level n-grams, part-of-speech counts, and named entity frequencies. Sentiment and emotion features drew on tools such as VADER, TextBlob, and the NRC emotion lexicon. Psycholinguistic features came from LIWC 2022 and Empath, while word embeddings ranged from static vectors like Word2Vec, GloVe, and FastText to contextual embeddings from BERT and its many derivatives. Notably, the team also crafted toxicity features using the Detoxify framework, extracting probabilities for labels such as toxic, obscene, threat, insult, and identity hate, an approach no prior role-identification study had taken.

The preprocessing pipeline itself reveals how messy real social media text can be. The researchers merged spaced-out letters, collapsed elongated characters like ‘youuuuuuu’ into ‘you’, preserved emojis with the emot package, resolved slang from online chat dictionaries, and used fuzzy string matching with a 90 percent similarity threshold to de-obfuscate profanity that users had deliberately disguised. Grammar and spelling were corrected with LanguageTool, and stopwords were deliberately retained because negations and pronouns carry crucial contextual signals in harassment scenarios. The class imbalance was handled by randomly downsampling the dominant non-bully class in the training data, while the test set was left untouched to preserve an unbiased evaluation.

The conventional models, multinomial logistic regression and a linear support vector classifier, were then trained with forward feature selection to find the best combinations. Textual features proved the strongest single group, achieving a macro F-measure of roughly 50 percent, while static word embeddings languished near 30 percent because they cannot capture the context-dependent tone that defines cyberbullying. Sentiment and emotion features performed poorly, suggesting that sarcasm and subtle aggression do not map neatly onto traditional sentiment dictionaries. The best logistic regression model, combining textual features, DistilBERT embeddings, psycholinguistic features, term lists, and toxicity scores, reached a macro F-measure of 54.58 percent on the hold-out test, beating the previous benchmark on the same dataset by 1.58 percentage points.

The second strategy, transfer learning, proved decisively more powerful. The team fine-tuned three compact pre-trained language models: DistilBERT, a distilled version of BERT that is 40 percent smaller and 60 percent faster; DistilRoBERTa; and Electra-small, which uses a replace-token-detection objective instead of masked language modeling. DistilBERT emerged as the clear winner. Fine-tuned for four epochs, it achieved a macro F-measure of 70.57 percent on the hold-out test, a roughly 10 percentage point improvement over the previous best result on the AMiCA corpus. At the level of individual roles, it scored an F-measure of 76.32 percent for harassers, 60.67 percent for bystander defenders, and 46.46 percent for victims, with recall improvements across every minority class compared to the conventional approach.

The victim role remains the hardest problem, and the error analysis explains why. Many misclassified victim posts were short, context-free utterances such as ‘yeah funny’ or ‘nice try goodbye honesty hour’, whose meaning depends entirely on the preceding conversation. Sarcasm also fooled the model: posts like ‘apparently you are beautiful’, actually harassment, were labeled as harmless because the literal words are benign. Long, emotionally mixed defender posts were similarly misclassified. Because the model reads each post in isolation, it cannot see the interaction history that would reveal who is responding to whom. The authors argue that incorporating conversational context and, eventually, large language models capable of explaining their predictions could address these failures.

The practical payoff is already visible. The team has released a public role-checker application built on the fine-tuned DistilBERT model, which classifies a submitted post as coming from a non-bully, harasser, victim, or bystander defender. The broader vision is intervention: platforms could use role-aware detection not merely to delete harmful content but to act against perpetrators, route support to victims, and encourage defenders, thereby interrupting the escalation dynamics that make cyberbullying so damaging. Because bystanders’ reactions critically shape whether an episode persists, knowing who is who in an interaction gives moderators leverage that binary detection never could.

The study also carries a methodological lesson for the field. Carefully engineered features still delivered a benchmark-beating conventional model, showing that domain knowledge retains value even in the era of transformers. But the efficiency gap is telling: logistic regression needed nearly nine minutes of cross-validation training with hand-crafted features, while DistilBERT required no feature engineering at all and processed four to five iterations per second during training. As the authors note, the findings may not generalize directly to other languages and cultures, where slang, sarcasm, and indirect communication differ. Still, by moving the field from asking whether cyberbullying occurs to asking how each participant is positioned within it, the work reframes automated moderation as a genuinely social problem, one where the victim’s plea and the defender’s intervention finally count as signals worth detecting.

Subject of Research: Automatic identification of participant roles in cyberbullying using machine learning and transfer learning

Article Title: Automatic classification of participants’ role in cyberbullying

Article References: Teng, T. H., Varathan, K. D., & Crestani, F. (2026). Automatic classification of participants’ role in cyberbullying. Heliyon, 12(15), Article e45491. https://doi.org/10.1016/j.heliyon.2026.e45491

Image Credits: AI Generated

DOI: 10.1016/j.heliyon.2026.e45491

Keywords: cyberbullying, natural language processing, machine learning, transfer learning, DistilBERT, text classification, bystanders, online harassment, feature engineering, Ask.fm, toxicity detection, class imbalance

News Source: Drew Townsend. (October 7, 2026). AI Learns to Name the Bully, the Victim and the Defender in Online Abuse. Scienmag.

Tags: Ask.fmbystandersclass imbalancecyberbullyingDistilBERTfeature engineeringMachine LearningNatural Language Processingonline harassmenttext classificationtoxicity detectionTransfer Learning
Share12Tweet7Share2ShareShareShare1

Related Posts

Immune Enzyme PTPN2 Emerges as Master Switch in Sepsis Inflammation

Immune Enzyme PTPN2 Emerges as Master Switch in Sepsis Inflammation

October 7, 2026
Oleic Acid Blend Selectively Suppresses Vaginal Pathogens While Sparing Key Beneficial Bacterium

Oleic Acid Blend Selectively Suppresses Vaginal Pathogens While Sparing Key Beneficial Bacterium

October 7, 2026

Hospital study finds bacteria carrying both last-resort resistance genes spreading in wards

October 7, 2026

Parasite Protein SporoAMA1 Emerges as Key Switch Linking Chronic Toxoplasma Infection to Transmission

October 7, 2026

POPULAR NEWS

  • Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

    Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

    29 shares
    Share 12 Tweet 7
  • Endurance Exercise Reshapes the Liver in Males and Females Through Distinct Molecular Routes

    29 shares
    Share 12 Tweet 7
  • Single Transcription Factor PU.1 Rapidly Converts Fibroblasts into Macrophage-Lineage Cells

    29 shares
    Share 12 Tweet 7
  • New Scale Measures How Ready Nurse Educators Really Are for the AI Era

    29 shares
    Share 12 Tweet 7

About

We bring you the latest biotechnology news from best research centers and universities around the world. Check our website.

Follow us

Recent News

Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

Endurance Exercise Reshapes the Liver in Males and Females Through Distinct Molecular Routes

Single Transcription Factor PU.1 Rapidly Converts Fibroblasts into Macrophage-Lineage Cells

Subscribe to Blog via Email

Success! An email was just sent to confirm your subscription. Please find the email now and click 'Confirm' to start subscribing.

Join 85 other subscribers
  • Contact Us

Bioengineer.org © Copyright 2023 All Rights Reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Homepages
    • Home Page 1
    • Home Page 2
  • News
  • National
  • Business
  • Health
  • Lifestyle
  • Science

Bioengineer.org © Copyright 2023 All Rights Reserved.