• HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
Thursday, October 1, 2026
BIOENGINEER.ORG
No Result
View All Result
  • Login
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
No Result
View All Result
Bioengineer.org
No Result
View All Result
Home NEWS Science News Technology

Self-Supervised Neighborhood Probing Shields Tabular AI From Poisoned Labels

Bioengineer by Bioengineer
October 1, 2026
in Technology
Reading Time: 5 mins read
0
Self-Supervised Neighborhood Probing Shields Tabular AI From Poisoned Labels
Share on FacebookShare on TwitterShare on LinkedinShare on RedditShare on Telegram

Tabular machine learning quietly runs the modern world. Models trained on rows and columns of data decide whether a credit card transaction is fraudulent, whether a patient’s health indicators suggest diabetes, and whether a network connection is an intrusion. But a new study from researchers at Soongsil University in Seoul, published in Applied Intelligence, highlights a disturbing weakness in these systems: an attacker does not need to touch a single input feature to sabotage them. By simply flipping the labels attached to training samples, an adversary can distort the decision boundaries that these high-stakes models learn, and the corruption can be nearly invisible to conventional validation checks.

The research team, led by Jinhyeok Jang and corresponding author Daeseon Choi, frames the problem with unusual clarity. Label flipping is a form of data poisoning in which the observed class of a training example is changed while the underlying features remain intact. Because the features themselves look perfectly normal, simple feature-level screening catches nothing. Worse, verifying whether a label is truly correct in domains like finance or healthcare often requires scarce expert knowledge, so corrupted labels can survive unnoticed all the way into production. As organizations increasingly rely on crowdsourced annotation, outsourced labeling, and third-party data integration to cut costs, the attack surface only grows.

What makes the new work particularly timely is its focus on how modern label-flipping attacks have evolved. Early attacks flipped labels at random, scattering noise across the dataset. Recent strategies are far more surgical. Decision-boundary attacks use surrogate models to identify low-margin samples sitting perilously close to the classifier’s decision frontier, while optimization-based methods such as ALFA and its ALFA-Tilt variant solve constrained optimization problems to select the flip set that maximally degrades the learner. The authors’ visualizations show that these targeted attacks produce localized poisoning patterns, with flipped samples clustering together in feature space rather than appearing as isolated outliers. That clustering defeats density-based and global outlier detectors, because the poisoned points masquerade as a legitimate local group.

The team’s answer is a defense built on a counterintuitive idea: learn what the data looks like without trusting any labels at all. Their method, called BYOL-Union, first trains a self-supervised encoder using BYOL, or Bootstrap Your Own Latent, a technique originally developed for images that learns representations by predicting augmented views of the same instance without any negative pairs or labels. Applied to tabular data, the encoder absorbs the intrinsic geometric structure of the feature space, untouched by whatever corruption may lurk in the observed labels. Only after this label-free representation learning is complete does the system consult the labels, and then solely to measure how much each sample disagrees with its neighbors.

The disagreement scoring is where the method earns its name. For each sample, the system finds its k nearest neighbors in the embedding space and counts how many of them carry a different observed label. A genuinely flipped sample tends to sit among neighbors whose true labels differ from its own forged one, producing a high disagreement score. Crucially, the method does not rely on a single fixed view of the neighborhood. Just as image-based self-supervised learning uses crops and flips to see one object from multiple angles, the framework generates a second view by adding Gaussian noise, with a standard deviation of 0.1, to continuous features only, leaving categorical features untouched to avoid inventing invalid categories. Each sample is then probed in both the original and the perturbed embedding spaces, and suspicious sets from the two views are merged by a union rule.

That union rule is deliberately recall-oriented. A sample flagged as suspicious in either view enters the final suspicious set, maximizing the chance of catching poisoned labels that are exposed in at least one neighborhood perspective, at the cost of some additional false positives. The authors also designed for realistic deployment, where the true poisoning ratio is unknown: instead of assuming an oracle budget, the practical version uses z-score thresholding at 1.0 with any-view exceedance, which their ablations show performs close to the oracle reference while remaining entirely oracle-free.

The evaluation is unusually broad. Six public tabular benchmarks spanning network intrusion detection, finance, and healthcare, including NSL-KDD, UNSW-NB15, Bank Marketing, Credit Card Fraud, BRFSS 2015 Diabetes Health Indicators, and the Diabetes 130-US Hospitals dataset, were poisoned at ratios of 10, 20, and 30 percent under random, decision-boundary, and optimization-based attacks. Four heterogeneous target models, spanning support vector machines, deep neural networks, FT-Transformers, and XGBoost, were then trained on sanitized data. The results show a consistent pattern: under decision-boundary flipping, the original feature-space kNN detector achieved a recall of only 0.601, while the BYOL-based and SCARF-based variants reached 0.773 and 0.792 respectively, and paired Wilcoxon tests confirmed that the SSL-based detectors significantly improved both recall and F1 over Curie, LS-SVM, and original-space kNN baselines.

The study is equally candid about limits. Under ALFA-Tilt, the strongest optimization-based attack tested in a small-scale setting, detection became markedly harder for every detector, and the robust-training defense FLORAL achieved the smallest downstream accuracy gap, though it cannot identify which specific samples are poisoned and is tied to particular model architectures. The authors stress that improved detection does not always translate directly into recovered test accuracy, because removing suspicious samples changes the training set itself and can thin out supervision for minority classes. Detection quality, they argue, should be judged on its own terms, with downstream recovery treated as a complementary, setting-dependent benefit rather than the sole criterion.

Perhaps the most striking result concerns adaptive adversaries. The team constructed defense-aware attacks in which the attacker knows the sanitization mechanism and selects flips that remain locally plausible in the detector’s own neighborhood space. Against a raw feature-space detector, such an adaptive attack cut recall from 0.668 to 0.435. Yet the SSL-based union detectors held firm: BYOL-Union maintained a recall of 0.760 even when the attacker targeted its own representation space, and SCARF-Union reached 0.800 under the corresponding adaptive attack. The layered evidence from learned representation spaces, it appears, is not fully dismantled by an adversary who optimizes against any single neighborhood view.

The broader lesson resonates beyond this one defense. As machine learning systems are deployed in domains where a single misclassified transaction or missed intrusion carries real consequences, the integrity of training labels deserves the same security attention as model architecture. The Soongsil team’s framework, which will see code released on request, offers a practical, model-agnostic preprocessing step: it flags suspicious samples before any downstream classifier is fit, works alongside tree ensembles and neural networks alike, and leaves room for future extensions toward label correction, sample reweighting, and robust retraining with uncertainty. In a field where attackers increasingly aim at the data rather than the model, defenses that learn to see the data on its own terms may prove essential.

Subject of Research: Label-flipping attack detection and data sanitization in tabular machine learning using multi-view self-supervised representations

Article Title: Multi-view self-supervised learning for label-flipping robustness in tabular data: a comparative study

Article References: Multi-view self-supervised learning for label-flipping robustness in tabular data: a comparative study. (n.d.). https://doi.org/10.1007/s10489-026-07452-2

Image Credits: AI Generated

DOI: 10.1007/s10489-026-07452-2

Keywords: label flipping, data poisoning, self-supervised learning, BYOL, tabular data, data sanitization, k-nearest neighbors, adversarial machine learning, SCARF, network intrusion detection, machine learning security, Applied Intelligence

Cite Scienmag News
APA MLA Chicago

Denise Maddox. (October 1, 2026). Self-Supervised Neighborhood Probing Shields Tabular AI From Poisoned Labels. Scienmag. https://scienmag.com/self-supervised-neighborhood-probing-shields-tabular-ai-from-poisoned-labels/

Denise Maddox. “Self-Supervised Neighborhood Probing Shields Tabular AI From Poisoned Labels.” Scienmag, 1 October 2026, https://scienmag.com/self-supervised-neighborhood-probing-shields-tabular-ai-from-poisoned-labels/. Accessed 1 October 2026.

Denise Maddox. “Self-Supervised Neighborhood Probing Shields Tabular AI From Poisoned Labels.” Scienmag. October 1, 2026. https://scienmag.com/self-supervised-neighborhood-probing-shields-tabular-ai-from-poisoned-labels/

Copy citation Download RIS

Tags: adversarial data poisoning in machine learningadversarial machine learningApplied IntelligenceBYOLdata poisoningdata sanitizationdata validation challenges in healthcaredefenses against label-based attacksfinancial fraud detection vulnerabilitieshigh-stakes decision-making securityimportance of label integrity in AIk-nearest neighborslabel flippinglabel flipping attack in supervised learningmachine learning securitynetwork intrusion detectionnoise and corruption in training datanovel techniques for poisoning resistancepoisoned label detection in tabular datarobustness of tabular AI modelsSCARFself-supervised learningSelf-supervised neighborhood probingtabular data

Share12Tweet7Share2ShareShareShare1

Related Posts

Gotu Kola Powers a New Solid Electrolyte for Magnesium Batteries

Gotu Kola Powers a New Solid Electrolyte for Magnesium Batteries

October 1, 2026
Mild detergent preserves bacteria for rapid sepsis susceptibility testing

Mild detergent preserves bacteria for rapid sepsis susceptibility testing

October 1, 2026

AI Learns One Person at a Time to Predict Daily Actions and Flag Danger Early

October 1, 2026

Graphs Meet Transformers: New AI Model Reads the Mood of Twitter

October 1, 2026

POPULAR NEWS

  • Mothers and Babies Show Heart Rhythm Synchrony With Surprising Time Lags

    29 shares
    Share 12 Tweet 7
  • New Uracil-Based Compound Targets SARS-CoV-2 Enzyme Nsp15 With Promising Antiviral Activity

    29 shares
    Share 12 Tweet 7
  • Gotu Kola Powers a New Solid Electrolyte for Magnesium Batteries

    29 shares
    Share 12 Tweet 7
  • Intensive Farming Disrupts Underground Fungal Networks That Sustain Wheat Yields

    29 shares
    Share 12 Tweet 7

About

We bring you the latest biotechnology news from best research centers and universities around the world. Check our website.

Follow us

Recent News

Mothers and Babies Show Heart Rhythm Synchrony With Surprising Time Lags

New Uracil-Based Compound Targets SARS-CoV-2 Enzyme Nsp15 With Promising Antiviral Activity

Gotu Kola Powers a New Solid Electrolyte for Magnesium Batteries

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 85 other subscribers
  • Contact Us

Bioengineer.org © Copyright 2023 All Rights Reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Homepages
    • Home Page 1
    • Home Page 2
  • News
  • National
  • Business
  • Health
  • Lifestyle
  • Science

Bioengineer.org © Copyright 2023 All Rights Reserved.