• HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
Monday, October 5, 2026
BIOENGINEER.ORG
No Result
View All Result
  • Login
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
No Result
View All Result
Bioengineer.org
No Result
View All Result
Home NEWS Science News Technology

New AI Framework Slashes Labeling Costs in Shifting, Imbalanced Data Streams

by
October 5, 2026
in Technology
Reading Time: 6 mins read
0
New AI Framework Slashes Labeling Costs in Shifting, Imbalanced Data Streams

New AI Framework Slashes Labeling Costs in Shifting, Imbalanced Data Streams

Share on FacebookShare on TwitterShare on LinkedinShare on RedditShare on Telegram

Every second, an unrelenting torrent of data flows through the world’s networks: sensor readings from industrial machinery, packets of internet traffic, transactions from financial systems, and streams of images from cameras watching roads, factories, and hospitals. For machine learning systems tasked with making sense of this flood in real time, two stubborn problems have long conspired to undermine accuracy. The first is concept drift, the phenomenon whereby the statistical patterns underlying the data quietly shift as the world changes. The second is class imbalance, in which some categories of data are overwhelmingly common while others, often the most important ones, appear only rarely. A new framework called MLIDSC, published in Applied Intelligence by researchers at Khulna University of Engineering & Technology and the University of Barishal in Bangladesh, tackles both problems at once with a self-adaptive labeling strategy that promises high accuracy at a fraction of the usual cost.

The labeling problem sits at the heart of why streaming machine learning is so difficult. In classical supervised learning, algorithms are trained on datasets in which every example already carries a label, a process that is expensive and time-consuming but at least finite. In a data stream, new examples arrive continuously and without end, and if every one of them had to be sent to a human expert for annotation, the cost would quickly become prohibitive. Active learning offers a way out: instead of labeling everything, the system itself decides which incoming examples are worth the expense of an expert query, and which can be safely ignored or inferred. The art lies in choosing well. Query too often and the budget evaporates; query too rarely and the model drifts out of date, silently degrading as the world moves on.

Concept drift sharpens this dilemma considerably. When the underlying relationship between features and classes changes, a model trained on yesterday’s data may be actively misleading today. Drift can arrive suddenly, as when a network intrusion technique appears out of nowhere, or gradually, as customer behavior evolves season by season. Either way, the moments immediately after a drift begins are precisely when fresh labels are most valuable, and precisely when a passive system is least likely to seek them. Existing active learning methods for streams have developed various heuristics for detecting uncertainty and triggering queries, but many of them depend on parameters that must be tuned in advance, assumptions about the data that the designers of MLIDSC argue are rarely justified in real deployments where prior knowledge is scarce.

The second challenge, class imbalance, compounds the difficulty in a way that is particularly insidious in multiclass settings. In a binary problem with a rare class, resampling techniques and cost-sensitive learning have well-understood analogues. But when a stream contains many classes with wildly different frequencies, and when the identity of the rare classes can itself change over time, the problem becomes far more tangled. Rare classes are often the ones that matter most: the fraudulent transaction among millions of legitimate ones, the failing component signal buried in routine telemetry, the emerging disease pattern in a stream of ordinary diagnoses. A labeling strategy that samples uniformly across the stream will spend most of its budget on the abundant classes, leaving the rare ones underrepresented and the model blind to exactly the events it was built to catch.

MLIDSC, whose name abbreviates Multiclass Imbalanced Data Stream with Concept Drift, addresses these intertwined challenges with what the authors describe as a self-adaptive online labeling strategy. The key idea is that the decision to query an expert should be made based on the current context of the stream, rather than on fixed thresholds or pre-set parameters. The framework introduces an automated labeling mechanism that eliminates the need for prior parameter assumptions, allowing the system to calibrate its own behavior against the data as it actually arrives. This matters in practice because streaming deployments are often set up once and left to run for months or years, with little opportunity for the kind of manual retuning that offline machine learning pipelines take for granted.

The most distinctive technical contribution of the framework is a novel weighting scheme designed to prioritize informative minority-class data in nonstationary environments. The scheme combines two signals: the imbalance ratio, which captures how underrepresented a given class is at a given moment, and the importance of individual instances at that point in time. By fusing these two measures, MLIDSC directs its limited labeling budget toward the data points that are simultaneously rare and critical, rather than spreading attention evenly or relying on static class priors. In a stream where the balance of classes shifts as drift proceeds, this dynamic weighting allows the system to notice when a formerly abundant class is fading and a formerly rare one is rising, and to reallocate expert attention accordingly.

Another important design choice distinguishes MLIDSC from much of the prior literature: it processes data online, example by example, rather than in chunks. Many existing methods for drifting and imbalanced streams operate on fixed-size batches, accumulating a buffer of examples before making labeling and model-update decisions. Chunk-based processing simplifies some algorithmic questions, but the authors argue it is impractical for scenarios that demand continuous processing, where waiting for a buffer to fill introduces latency and can delay the detection of abrupt changes. By working in a truly online fashion, MLIDSC can react to each new example as it arrives, which is essential in applications such as network security or industrial monitoring where a delayed response is often as bad as no response at all.

The experimental evaluation behind the framework is notably comprehensive. The researchers tested MLIDSC on both real and synthetic data streams, with synthetic datasets generated using the Scikit-multiflow Python package and real datasets drawn from the Massive Online Analysis repository and the UCI machine learning repository, including the Statlog project data. Crucially, the experiments varied both the degree of concept drift and the imbalance ratio, allowing the team to probe how the method behaves across the full spectrum of difficulty that streaming environments present. The results reported in the paper show that MLIDSC achieves high accuracy while substantially reducing labeling costs compared to existing methods, a combination that is the central promise of active learning and the standard against which such frameworks are judged.

The implications reach well beyond the machine learning research community. Any organization that deploys predictive models on live data faces the twin pressures that MLIDSC targets: expert annotation is expensive, and the world refuses to hold still. Network traffic classification, one of the motivating applications in this line of research, is a vivid example. The mix of applications and protocols traversing a network changes constantly, and traffic classes are wildly imbalanced, with a handful of dominant services dwarfing the rare but security-critical flows that analysts most need to identify. A framework that keeps a classifier accurate in such an environment while asking human experts to label only a small, well-chosen fraction of the stream translates directly into operational savings and faster detection of emerging threats.

The work also fits into a broader and rapidly evolving research conversation. The references anchoring the study trace a decade of progress in the field, from early active learning methods for drifting streams developed by Žliobaitė and colleagues, through ensemble approaches that pair drift detection with resampling strategies, to recent work on multiclass imbalance by Liu, Li, Han and others. The authors of MLIDSC build on their own prior contributions as well, including hybrid labeling strategies and drift detection techniques based on outlier computation. What sets the new framework apart in this crowded landscape is its insistence on self-adaptation: rather than asking practitioners to guess at parameters before deployment, it derives its labeling decisions from the stream itself. As machine learning systems are increasingly entrusted with continuous, high-stakes decisions in environments that never stop changing, that kind of autonomy may prove to be not just a convenience but a necessity, and MLIDSC offers a concrete, experimentally validated template for how to achieve it.

Subject of Research: Online active learning for multiclass imbalanced data streams with concept drift

Article Title: MLIDSC: A self-adaptive online active learning framework for multiclass imbalanced data stream with concept drift

Article References: Halder, B., Hasan, K. M. A., & Ahmed, M. M. (2026). MLIDSC: A self-adaptive online active learning framework for multiclass imbalanced data stream with concept drift. Applied Intelligence, 56(15), Article 431. https://doi.org/10.1007/s10489-026-07473-x

Image Credits: AI Generated

DOI: 10.1007/s10489-026-07473-x

Keywords: machine learning, active learning, concept drift, data streams, class imbalance, online learning, labeling strategy, multiclass classification, adaptive algorithms, streaming data, Applied Intelligence, MLIDSC

News Source: Denise Maddox. (October 5, 2026). New AI Framework Slashes Labeling Costs in Shifting, Imbalanced Data Streams. Scienmag.

Tags: active learningadaptive algorithmsApplied Intelligenceclass imbalanceconcept driftdata streamslabeling strategyMachine LearningMLIDSCmulticlass classificationonline learningstreaming data
Share12Tweet7Share2ShareShareShare1

Related Posts

Blockchain Meets the Road: A Lightweight Security Shield for the Internet of Vehicles

Blockchain Meets the Road: A Lightweight Security Shield for the Internet of Vehicles

October 5, 2026
Drones Get Smarter: Graph Attention Networks Boost Real-Time Aerial Object Detection

Drones Get Smarter: Graph Attention Networks Boost Real-Time Aerial Object Detection

October 5, 2026

Smart Statistical Model Reveals What Drives Green Finance Success in the Big Data Era

October 5, 2026

Female AI Agents Earn 10% Less Than Male Counterparts in Virtual Workplace Experiment

October 5, 2026

POPULAR NEWS

  • Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

    Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

    29 shares
    Share 12 Tweet 7
  • Endurance Exercise Reshapes the Liver in Males and Females Through Distinct Molecular Routes

    29 shares
    Share 12 Tweet 7
  • Single Transcription Factor PU.1 Rapidly Converts Fibroblasts into Macrophage-Lineage Cells

    29 shares
    Share 12 Tweet 7
  • New Scale Measures How Ready Nurse Educators Really Are for the AI Era

    29 shares
    Share 12 Tweet 7

About

We bring you the latest biotechnology news from best research centers and universities around the world. Check our website.

Follow us

Recent News

Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

Endurance Exercise Reshapes the Liver in Males and Females Through Distinct Molecular Routes

Single Transcription Factor PU.1 Rapidly Converts Fibroblasts into Macrophage-Lineage Cells

Subscribe to Blog via Email

Success! An email was just sent to confirm your subscription. Please find the email now and click 'Confirm' to start subscribing.

Join 85 other subscribers
  • Contact Us

Bioengineer.org © Copyright 2023 All Rights Reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Homepages
    • Home Page 1
    • Home Page 2
  • News
  • National
  • Business
  • Health
  • Lifestyle
  • Science

Bioengineer.org © Copyright 2023 All Rights Reserved.