• HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
Wednesday, September 23, 2026
BIOENGINEER.ORG
No Result
View All Result
  • Login
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
No Result
View All Result
Bioengineer.org
No Result
View All Result
Home NEWS Science News Technology

New AI Method Learns From Rare Positives Hidden in Unlabeled Data

Bioengineer by Bioengineer
September 23, 2026
in Technology
Reading Time: 6 mins read
0
New AI Method Learns From Rare Positives Hidden in Unlabeled Data
Share on FacebookShare on TwitterShare on LinkedinShare on RedditShare on Telegram

Machine learning has transformed fields from medicine to finance, but one stubborn problem continues to undermine its reliability in the real world: most of the data we care about is either missing or mislabeled. In many practical settings, practitioners have only a handful of confirmed positive examples, surrounded by a vast ocean of unlabeled data that contains a hidden mixture of positives and negatives. This scenario, known as positive-unlabeled (PU) learning, has been studied intensively for two decades, yet most existing methods quietly assume that the classes are roughly balanced and that labeled examples are drawn randomly from the positive population. A new study published in Data Mining and Knowledge Discovery by Elias Zavitsanos and Georgios Paliouras of the Institute of Informatics and Telecommunications at NCSR Demokritos in Greece confronts those assumptions head-on and delivers a method that thrives precisely where others falter.

The core difficulty is easy to state but hard to solve. In classical binary classification, an algorithm sees both labeled positives and labeled negatives and learns a decision boundary between them. In PU learning, negative labels simply do not exist. Instead, the algorithm receives a small set of confirmed positives and a large unlabeled pool. Because the unlabeled pool is dominated by negatives but contaminated with positives, treating every unlabeled example as negative injects label noise into training. Traditional approaches either try to identify reliable negatives from the unlabeled set before training, or they incorporate assumptions about the class prior, the underlying proportion of positive examples in the data. Both strategies become brittle when the dataset is severely imbalanced, meaning positives may account for only one to a few percent of all examples.

Zavitsanos and Paliouras observed that in imbalanced PU data, the mathematics of existing risk estimators works against the practitioner. The estimators used by state-of-the-art methods, such as the unbiased PU estimator (uPU) and its non-negative successor (nnPU), contain a term weighted by the positive class prior. When that prior is tiny, the contribution of the few labeled positives nearly vanishes, and the training objective is dominated by the unlabeled data. The model therefore learns to be excellent at recognizing negatives, and mediocre at catching the rare positives that matter most. Matters worsen under the Probabilistic Gap assumption, a realistic refinement of labeling theory in which positive examples that resemble negatives are the least likely to have been labeled. The very positives a model most needs to learn from are the ones most likely to be hidden in the unlabeled pool.

The researchers’ answer is a new empirical risk estimator they call iFPU, for imbalanced focused PU learning. The key ingredient is focal loss, a function originally developed in computer vision for dense object detection, where positive targets are similarly swamped by an enormous number of easy negatives. Focal loss reshapes the standard cross-entropy objective by multiplying each example’s loss by a factor that shrinks as the model’s confidence in the correct answer grows. Easy examples, which the model already classifies correctly with high probability, contribute almost nothing to the gradient. Hard examples, sitting near the decision boundary, receive exponentially amplified weight, controlled by a focusing parameter gamma. The result is that training concentrates its capacity on exactly the boundary cases that define the Probabilistic Gap.

Incorporating focal loss into a PU risk estimator is technically delicate. The authors derive a non-negative risk estimator that combines three terms: the focal loss of the labeled positives treated as positive, an unbiased correction subtracting the focal loss the positives would incur if treated as negative, and the focal loss of the unlabeled pool treated as negative. A max operator clamps the correction at zero to prevent the notorious problem of negative empirical risk, which causes overfitting in neural networks trained with earlier unbiased estimators. When the correction term goes negative within a training mini-batch, the method performs a step of gradient ascent instead, deliberately nudging the model away from overfitting that batch. The whole procedure slots into standard stochastic optimization, requires no preprocessing, resampling, or manipulation of the data, and can be attached to essentially any classifier trained by cost minimization, from multilayer perceptrons to gradient-boosted trees.

The authors also supply a rigorous theoretical analysis. They prove that the population-level iFPU risk is identical to the focal classification risk under full supervision, for any labeling mechanism, by virtue of a mixture identity between the unlabeled distribution and the class-conditional densities. They further characterize the bias that appears when labeled positives are not selected completely at random, showing that this bias depends only on the labeling propensity and is orthogonal to the choice of loss function. Under the SCAR assumption, they establish an estimation error bound that vanishes at the standard statistical rate, proportional to the inverse square root of the number of labeled positives and unlabeled examples combined, using Rademacher complexity tools. Crucially, they show focal loss is Lipschitz continuous and bounded on a restricted score range, which makes these guarantees possible. Although focal loss itself introduces a deliberate bias by reweighting errors, it remains classification-calibrated, meaning the Bayes-optimal classifier under focal loss coincides with the Bayes-optimal classifier under zero-one loss.

Empirically, the method was tested on 14 publicly available benchmark datasets for imbalanced binary classification, spanning positive rates from roughly 1 to 14 percent. The experimental design was deliberately demanding: positive examples were progressively hidden in the unlabeled pool at rates of 25, 50, and 75 percent, under both the SCAR and the more realistic SAR labeling assumptions, generating 840 experimental runs in total. Notably, the authors avoided hyperparameter tuning altogether, using the recommended default gamma of 3, because tuning in PU settings is itself fraught with assumptions due to the absence of negatively labeled validation data. Under SCAR, iFPU outperformed the neural risk estimators uPU, nnPU, and i-NNPU, and when paired with an XGBoost classifier it matched or exceeded strong competitors including the two-step NNIF anomaly-detection method, PU Hellinger Decision Trees, the label-bias estimation method LBE, and the SAREM expectation-maximization framework, trailing only slightly behind the PU Hellinger Random Forest ensemble.

The picture shifts decisively in favor of iFPU under the more realistic SAR assumption, where positives resembling negatives are less likely to be labeled. Here the competing methods degrade noticeably, while iFPU maintains its performance, and in the hardest scenario, with only 25 percent of positives labeled, it ranks first overall, surpassing the PU Hellinger Random Forest by five percentage points in PR-AUC. Statistical tests confirmed significant differences among methods, and paired comparisons showed iFPU significantly outperforming all alternatives in the most challenging configuration. A sensitivity analysis demonstrated that the method remains robust even when the class prior is misspecified by factors of two or four in either direction, degrading meaningfully only when the prior is severely underestimated. On the three most imbalanced datasets, Cover, Poker, and Satellite, iFPU showed clearly higher mean and median PR-AUC than its strongest rival.

To demonstrate real-world value, the researchers applied their method to financial misstatement detection, a problem where PU data arise naturally. Auditors and regulators typically discover accounting misstatements years after reports are filed, and often only a fraction of misstatements have been identified when a model is trained. Using data on publicly traded US companies spanning 2000 to 2014, with 47,086 firm-year records described by 28 financial indices and derived accounting features, the team simulated realistic detection delays in which approximately 40 percent of positive training labels were missing at training time. Misstatements, whether deliberately concealed frauds or subtle errors resembling normal accounts, fit the Probabilistic Gap assumption perfectly. Models built on a TabTransformer architecture with gated MLP modules and equipped with the calibrated iFPU risk achieved R-precision scores roughly three times higher than prior baselines such as RUSBoost, and outperformed both earlier specialized models and the PU Hellinger Random Forest, setting a new state of the art in this application.

The significance of this work extends beyond any single benchmark. By combining a principled risk-estimation framework with a loss function engineered for imbalance, the authors show that PU learning can be made practical in exactly the conditions that dominate high-stakes applications, from disease gene identification to fraud detection, where positives are rare, partially labeled, and deceptively similar to negatives. The method’s plug-and-play compatibility with modern classifiers, its robustness to prior misspecification, and its theoretical grounding distinguish it from heuristic preprocessing pipelines. The authors point to extensions toward semi-supervised and multi-class settings, integration with pre-trained tabular foundation models such as TabPFN, and output calibration via temperature scaling as promising future directions. For now, iFPU offers practitioners a rare commodity in weakly supervised machine learning: a method whose assumptions match reality rather than convenience.

Subject of Research: Positive-unlabeled machine learning from highly imbalanced datasets

Article Title: Focused PU learning from imbalanced data

Article References: Focused PU learning from imbalanced data. (n.d.). https://doi.org/10.1007/s10618-026-01264-1

Image Credits: AI Generated

DOI: 10.1007/s10618-026-01264-1

Keywords: PU learning, imbalanced classification, weakly supervised learning, focal loss, risk estimation, machine learning, fraud detection, financial misstatement detection, XGBoost, SCAR assumption, SAR assumption, data mining

Cite Scienmag News
APA MLA Chicago

Denise Maddox. (September 23, 2026). New AI Method Learns From Rare Positives Hidden in Unlabeled Data. Scienmag. https://scienmag.com/new-ai-method-learns-from-rare-positives-hidden-in-unlabeled-data/

Denise Maddox. “New AI Method Learns From Rare Positives Hidden in Unlabeled Data.” Scienmag, 23 September 2026, https://scienmag.com/new-ai-method-learns-from-rare-positives-hidden-in-unlabeled-data/. Accessed 23 September 2026.

Denise Maddox. “New AI Method Learns From Rare Positives Hidden in Unlabeled Data.” Scienmag. September 23, 2026. https://scienmag.com/new-ai-method-learns-from-rare-positives-hidden-in-unlabeled-data/

Copy citation Download RIS

Tags: advancing reliability of AI modelsassumptions in positive-unlabeled learningchallenges of unlabeled dataclass imbalance in machine learningdata miningdata mining for hidden positive signalsfinancial misstatement detectionfocal lossfraud detectionhandling missing and mislabeled dataimbalanced classificationMachine learningmachine learning in medicine and financenew algorithms for unbalanced datasetsnovel methods for PU learningpositive-unlabeled learningPU learningrare positive example detectionrisk estimationSAR assumptionSCAR assumptionsemi-supervised learning in machine learningweakly supervised learningXGBoost

Share12Tweet7Share2ShareShareShare1

Related Posts

New Bidirectional Grover Search Slashes Quantum Database Iterations

New Bidirectional Grover Search Slashes Quantum Database Iterations

September 23, 2026
AI model predicts earthquake vulnerability of existing concrete buildings in milliseconds

AI model predicts earthquake vulnerability of existing concrete buildings in milliseconds

September 23, 2026

Ankle Nerve Signals Reveal Hidden Timing Errors in Unstable Ankles

September 23, 2026

How Shrinking Manganite Crystals to Nanoscale Rewrites the Rules of Magnetic Anisotropy

September 23, 2026

POPULAR NEWS

  • Chemical Fingerprints Reveal Where China’s Sauce-Flavor Baijiu Truly Comes From

    29 shares
    Share 12 Tweet 7
  • Liver Tumour Ablation Enters Mainstream Oncology With New Global Standards

    29 shares
    Share 12 Tweet 7
  • Rare Congenital Lung Anomaly Masquerades as a Lookalike Condition on CT Scans

    29 shares
    Share 12 Tweet 7
  • Massive Protein Interaction Map Reveals How Muscles Fall Silent to Insulin

    29 shares
    Share 12 Tweet 7

About

We bring you the latest biotechnology news from best research centers and universities around the world. Check our website.

Follow us

Recent News

Chemical Fingerprints Reveal Where China’s Sauce-Flavor Baijiu Truly Comes From

Liver Tumour Ablation Enters Mainstream Oncology With New Global Standards

Rare Congenital Lung Anomaly Masquerades as a Lookalike Condition on CT Scans

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 85 other subscribers
  • Contact Us

Bioengineer.org © Copyright 2023 All Rights Reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Homepages
    • Home Page 1
    • Home Page 2
  • News
  • National
  • Business
  • Health
  • Lifestyle
  • Science

Bioengineer.org © Copyright 2023 All Rights Reserved.