• HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
Wednesday, October 7, 2026
BIOENGINEER.ORG
No Result
View All Result
  • Login
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
No Result
View All Result
Bioengineer.org
No Result
View All Result
Home NEWS Science News Technology

New AI Architecture Tames Extreme Data Imbalance in Multi-Task Recommendation

by
October 7, 2026
in Technology
Reading Time: 5 mins read
0
New AI Architecture Tames Extreme Data Imbalance in Multi-Task Recommendation

New AI Architecture Tames Extreme Data Imbalance in Multi-Task Recommendation

Share on FacebookShare on TwitterShare on LinkedinShare on RedditShare on Telegram

Recommendation systems face a deceptively simple question every time a user watches a video: did they like it, and did they share it? Behind that question lies one of the hardest practical problems in modern machine learning. Likes and shares are extraordinarily rare compared with ordinary views, and when a model is trained to predict several such signals at once, the abundant signal tends to drown out the scarce ones. A new study published in Applied Intelligence by Jinghao Xue of Keio University’s Graduate School of Science and Technology, together with Tianxiang Yang and Hideo Suzuki of the Department of Industrial and Systems Engineering, proposes a framework called H-PLE that confronts this imbalance head-on. The work, published on 8 September 2026 as volume 56, article 409 of the journal, demonstrates that careful architectural design can rescue the performance of sparse tasks without sacrificing the dense ones.

The starting point for the research is a well-known phenomenon called negative transfer. Multi-task learning rests on the appealing idea that a shared neural representation can serve several prediction tasks simultaneously, improving data efficiency because each task benefits from features learned for the others. Rich Caruana articulated this principle in 1997, and it has since become a cornerstone of industrial recommendation engines. But the bargain only holds when the tasks are roughly compatible in difficulty and data volume. When one task, such as predicting whether a user will share a video, has orders of magnitude fewer positive examples than another, such as predicting a click or a view, the shared representation is shaped almost entirely by the high-frequency task. The rare task then inherits features that are poorly suited to it, and attempts to improve it can actively degrade the common tasks.

Two influential architectures have dominated attempts to manage this interference. Multi-gate Mixture-of-Experts, introduced by Jiaqi Ma and colleagues at Google in 2018, maintains a pool of expert subnetworks and lets each task learn its own soft weighting over them through a gating network. Progressive Layered Extraction, or PLE, proposed by Hongyu Tang and colleagues in 2020, refined this idea by separating shared experts from task-specific experts and stacking the routing across multiple layers, so that task-specific knowledge is progressively extracted from the shared representation. Both approaches, however, rely on fully learned soft gating. As Xue and colleagues point out, under severe label imbalance the gating networks themselves become part of the problem: they are trained predominantly on gradients flowing from the high-frequency tasks, so the routing decisions drift toward serving those tasks, leaving the sparse ones with unstable expert selection and insufficiently expressive representations.

H-PLE, which stands for hierarchical progressive layered extraction, builds on PLE but adds three complementary design elements. The first is a hierarchical extraction structure that progressively separates shared and task-specific representations across stacked layers, giving the model a more disciplined pathway for carving out what is common from what is unique. The second, and arguably the most novel, is a raw-input-aware gating mechanism the authors call AGC. In conventional expert routing, the gate sees only learned intermediate representations, which are themselves contaminated by the imbalance problem. AGC instead injects a projection of the raw input features back into the gating decision as a residual candidate. This gives the router a stable, task-independent anchor: no matter how skewed the training gradients become, the gate can always consult an uncorrupted view of what the user and the content actually look like, which stabilizes routing precisely where supervision is sparsest.

The third element is a lightweight expert-level interaction module that introduces factorization-machine-style second-order feature interactions into the experts. Factorization machines, introduced by Steffen Rendle in 2010, model pairwise interactions between features through learned latent vectors, and they remain a powerful inductive bias for recommendation data, where the conjunction of two features, say a user attribute and a video genre, often carries more predictive signal than either feature alone. By embedding this bias at the expert level, H-PLE equips its subnetworks with an explicit mechanism for capturing cross-feature effects, which the authors find particularly valuable for the rare engagement tasks where every scrap of signal counts.

The evaluation is notable for its rigor. The researchers tested their framework on two real-world scenarios from the Tenrec benchmark, a large-scale multi-purpose recommendation dataset released by Tencent, using the QK-video and QB-video settings. They constructed binary prediction tasks for Like and Share, two engagement signals with highly imbalanced positive rates. Rather than relying on a single train-test split or a single random seed, the team ran multi-seed experiments with independent random splits, and they compared against a broad slate of modern baselines: MMoE, PLE, Cross-Stitch networks, PCGrad gradient surgery, and loss-rebalancing techniques including focal loss and class-balanced binary cross-entropy. Metrics went beyond the usual area under the ROC curve to include PR-AUC, which is far more informative under imbalance, GAUC for user-level ranking quality, calibration measures, and threshold-optimized F1 scores.

The results tell a story that is both encouraging and sobering. H-PLE-family variants consistently improved sparse-task ranking and minority-class retrieval over MMoE and PLE, the two architectures that dominate industrial practice. They also remained competitive with, or stronger than, the Cross-Stitch, PCGrad, and loss-rebalancing baselines. Perhaps the most striking finding concerns those rebalancing baselines themselves: focal loss and class-balanced BCE, techniques that are widely recommended for imbalanced classification, actually underperformed vanilla MMoE in the study’s most imbalanced setting, the QB-Share task. That result underscores just how difficult extreme sparsity is, and it serves as a caution against assuming that loss-level fixes can substitute for architectural ones when the imbalance is severe enough.

Equally important is the paper’s honesty about its own limits. The full H-PLE model was not uniformly best on every QB metric. The ablation studies, which systematically removed individual components, revealed that AGC, the raw-input-aware gating mechanism, is the most robust component under the most extreme sparsity, while the factorization-machine interaction module delivers benefits only when its interaction rank is carefully controlled. In other words, the gains come from specific, identifiable mechanisms rather than from a blanket advantage, and practitioners deploying the framework would need to tune the interaction rank rather than simply switch on every component. This kind of component-level attribution, made possible by the ablation design, is rarer in the recommendation literature than it should be, and it makes the paper’s positive claims considerably more trustworthy.

The practical implications reach well beyond video recommendation. Any production system that predicts multiple user behaviors, from e-commerce platforms jointly modeling clicks, add-to-carts, and purchases, to content platforms balancing watch time, comments, and subscriptions, confronts the same structural problem: dense signals dominate shared representations and starve sparse ones. The insight that gating networks are themselves vulnerable to imbalance-driven bias, and that anchoring them to raw inputs can counteract this, offers a general design principle that could be ported to other multi-task architectures. The finding that second-order interactions help only under careful rank control similarly generalizes as a reminder that inductive biases are tools with operating ranges, not free lunches.

The study also reflects a commendable degree of transparency about data and reproducibility. The authors state that the Tenrec dataset is available from Tencent’s public benchmark page under its own access conditions, which prevents redistribution by the authors, but that preprocessing scripts, experimental configuration files, and result summaries can be obtained from the corresponding author upon reasonable request, subject to the dataset’s terms of use. The work was supported by JST SPRING grant JPMJSP2123 and JSPS KAKENHI grant 25K17794, and the authors declare no competing interests. As recommendation systems increasingly mediate what billions of people see, ensuring that rare but meaningful signals, the shares and the passionate endorsements rather than the passive clicks, are not lost in the statistical noise becomes a matter of both engineering quality and platform health. H-PLE offers a carefully validated step in that direction, and its candid reporting of where the method falls short may prove as influential as the architecture itself.

Subject of Research: Multi-task learning for recommendation systems under extreme label imbalance

Article Title: H-PLE: hierarchical progressive layered extraction with raw-input-aware gating for multi-task recommendation under extreme label imbalance

Article References: Xue, J., Yang, T., & Suzuki, H. (2026). H-PLE: hierarchical progressive layered extraction with raw-input-aware gating for multi-task recommendation under extreme label imbalance. Applied Intelligence, 56(14), Article 409. https://doi.org/10.1007/s10489-026-07441-5

Image Credits: AI Generated

DOI: 10.1007/s10489-026-07441-5

Keywords: multi-task learning, recommendation systems, label imbalance, negative transfer, mixture-of-experts, progressive layered extraction, gating networks, factorization machines, Tenrec benchmark, PR-AUC, neural networks, Applied Intelligence

News Source: Blake Davidson. (October 7, 2026). New AI Architecture Tames Extreme Data Imbalance in Multi-Task Recommendation. Scienmag.

Tags: Applied Intelligencefactorization machinesgating networkslabel imbalancemixture-of-expertsMulti-task learningnegative transferNeural NetworksPR-AUCprogressive layered extractionrecommendation systemsTenrec benchmark
Share12Tweet7Share2ShareShareShare1

Related Posts

Fuzzy Math Tames Uncertainty in Multi-Modal Freight and Disaster Logistics

Fuzzy Math Tames Uncertainty in Multi-Modal Freight and Disaster Logistics

October 7, 2026
AI That Teaches Robots to See: New Deep Learning System Slashes Navigation Errors in Visual SLAM

AI That Teaches Robots to See: New Deep Learning System Slashes Navigation Errors in Visual SLAM

October 7, 2026

Explainable AI Model Predicts India’s Air Quality With Near-Perfect Accuracy

October 7, 2026

Bamboo Rings Meet Fiber Cement in Lightweight Sandwich Panels for Greener Walls

October 7, 2026

POPULAR NEWS

  • Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

    Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

    29 shares
    Share 12 Tweet 7
  • Endurance Exercise Reshapes the Liver in Males and Females Through Distinct Molecular Routes

    29 shares
    Share 12 Tweet 7
  • Single Transcription Factor PU.1 Rapidly Converts Fibroblasts into Macrophage-Lineage Cells

    29 shares
    Share 12 Tweet 7
  • New Scale Measures How Ready Nurse Educators Really Are for the AI Era

    29 shares
    Share 12 Tweet 7

About

We bring you the latest biotechnology news from best research centers and universities around the world. Check our website.

Follow us

Recent News

Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

Endurance Exercise Reshapes the Liver in Males and Females Through Distinct Molecular Routes

Single Transcription Factor PU.1 Rapidly Converts Fibroblasts into Macrophage-Lineage Cells

Subscribe to Blog via Email

Success! An email was just sent to confirm your subscription. Please find the email now and click 'Confirm' to start subscribing.

Join 85 other subscribers
  • Contact Us

Bioengineer.org © Copyright 2023 All Rights Reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Homepages
    • Home Page 1
    • Home Page 2
  • News
  • National
  • Business
  • Health
  • Lifestyle
  • Science

Bioengineer.org © Copyright 2023 All Rights Reserved.