• HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
Sunday, October 4, 2026
BIOENGINEER.ORG
No Result
View All Result
  • Login
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
No Result
View All Result
Bioengineer.org
No Result
View All Result
Home NEWS Science News Technology

New Graph-Based Algorithm Tames Streaming Features in High-Dimensional Data

Bioengineer by Bioengineer
October 4, 2026
in Technology
Reading Time: 5 mins read
0
New Graph-Based Algorithm Tames Streaming Features in High-Dimensional Data
Share on FacebookShare on TwitterShare on LinkedinShare on RedditShare on Telegram

Machine learning models are often drowning in data, but the problem is not always the sheer volume of samples. Increasingly, it is the flood of features—the individual measurable properties of each data point—that threatens to overwhelm algorithms. In domains ranging from gene expression profiling to spam filtering, features do not arrive all at once. They stream in sequentially, often in predefined groups, and waiting for the complete feature set before analysis begins is computationally prohibitive. A new study published in Applied Intelligence proposes a method designed precisely for this challenging scenario, and its results suggest a significant step forward in how machines can learn from data that never stops arriving.

The method, called FS-GSR—Feature Selection with Group Streaming via Graph-based Scoring and Sparse Reconstruction—was developed by Tianyuan Jia of The Hong Kong Polytechnic University. It tackles a core weakness of existing streaming feature selection techniques: most process features one at a time, ignoring the structural relationships and redundancies that exist both within and between groups of features. When features arrive in batches, as they do in multi-view learning systems or modular sensor deployments, sequential individual processing can disrupt group structures and degrade performance. FS-GSR instead operates at the group level, updating the selected feature subset in real time as each new batch arrives.

At the heart of the approach is a three-stage framework that functions like a computational funnel. The first stage performs a hybrid, graph-guided filtering of each incoming feature group. The algorithm constructs a similarity graph over the samples, connecting points that are close in the feature space, and computes a Laplacian Score for every feature. This score measures how smoothly a feature varies across locally similar samples—features that preserve the intrinsic geometric structure of the data receive low scores, while noisy, irregular features are penalized. In parallel, a supervised F-score evaluates each feature’s discriminative power by comparing between-class variance to within-class variance. The two measures are blended with equal weight, and only the top fraction of features by combined score survives this initial screening.

The second stage refines this candidate pool by preserving structure at a finer granularity. The algorithm performs a spectral decomposition of the graph Laplacian, extracting ten eigenvectors that form a low-dimensional embedding of the sample manifold. Each surviving feature is then evaluated by how well it can reconstruct this structural representation, using a closed-form projection solution that avoids costly iterative optimization. Features that best approximate the manifold’s skeleton are retained, ensuring that the selected subset captures not just individual relevance but the collective geometry of the data. Because the reconstruction coefficients have an explicit analytical solution, this stage adds minimal computational overhead.

The third and final stage is where FS-GSR departs most sharply from its predecessors. All candidate features accumulated from the first two stages are merged into a global pool, and a multi-task Lasso regression is applied across the entire set simultaneously, using one-hot encoded class labels as the response. The L2,1-norm penalty drives the coefficients of redundant features to exactly zero, yielding a compact, non-overlapping final feature set. The authors also provide theoretical backing for this stage, formulating it as a row-support recovery problem in high-dimensional multivariate regression. Under standard conditions on the covariance structure, they prove that the sparse reconstruction can recover the true set of informative features with high probability, provided the effective sample complexity exceeds a specific threshold.

The computational design is notable for its scalability. By confining expensive cubic operations to the fixed sample space rather than the growing feature dimension, the algorithm’s cost per streaming group scales linearly with the number of incoming features. Because the first two stages aggressively compress each group before global optimization, the feature pool fed into the multi-task Lasso remains small, preventing the exponential blow-up that plagues conventional structure-preserving methods. Scalability experiments on the Ovarian Cancer dataset confirmed near-linear growth in running time as feature dimension increased, and revealed a V-shaped trade-off in execution time as the number of streaming groups varied—very few large groups incur heavy intra-group matrix operations, while very many small groups accumulate loop overhead.

Empirically, the method was tested on five benchmark datasets spanning biomedical, image, and text domains, including Colon (62 samples, 2,000 features), Ovarian Cancer (253 samples, 15,154 features), MLL, COIL20, and PCMAC. These datasets deliberately covered contrasting correlation structures: the Ovarian Cancer data exhibited strong intra-group correlations of up to 0.88, while the PCMAC text dataset showed extremely sparse correlations of roughly 0.025. Against four established baselines—OGFS, Group-SAOLA, Fast-OSFS, and alpha-investing—FS-GSR consistently achieved the highest or near-highest classification accuracy and F1 scores under both KNN and SVM classifiers. On the Ovarian Cancer dataset, it reached accuracy and F1 values above 0.99.

Perhaps the most striking result concerned feature compression. In one experimental configuration, FS-GSR retained an average of just 8.12 features from the 15,154-dimensional Ovarian Cancer dataset, while simultaneously improving KNN accuracy from 0.9845 to 0.9960 and SVM accuracy from 0.9885 to 0.9980—a 63 percent compression relative to the two-stage variant alone. Ablation studies confirmed that each stage contributes distinct benefits: the first enables efficient local filtering, the second improves structural consistency, and the third enforces global sparsity and discriminative refinement. The authors note that the degree of compression adapts to each dataset’s intrinsic correlation structure, with strongly correlated data requiring more retained features to preserve associative signals.

Parameter sensitivity analyses reinforced the method’s practicality. Classification accuracy remained stable across broad plateaus for the retention ratios governing the first two stages, and the mixing weight balancing structural and discriminative scoring performed robustly between 0.2 and 0.8, peaking near 0.5. The embedding dimension saturated quickly, with accuracy plateauing once it exceeded five, justifying the default setting of ten. This robustness means practitioners can deploy the method without exhaustive grid searches, using a simple boundary-driven tuning strategy on validation data.

The implications extend well beyond the benchmark datasets. Streaming, group-wise feature arrival is characteristic of stepwise sensor deployment, progressive module activation in multi-view systems, and sequential extraction pipelines in bioinformatics and text mining. By combining graph-based manifold learning with sparse global reconstruction in an online framework, FS-GSR offers a template for learning systems that must remain both accurate and parsimonious as data evolves. The author identifies future extensions toward multi-label settings, incomplete data systems, and broader generalization, suggesting that the challenge of features that never stop flowing is one the field is only beginning to master.

Subject of Research: Group streaming feature selection for high-dimensional data using graph-based scoring and sparse reconstruction

Article Title: FS-GSR: Graph-based scoring and sparse reconstruction for group streaming feature selection

Article References: Jia, T. (2026). FS-GSR: Graph-based scoring and sparse reconstruction for group streaming feature selection. Applied Intelligence, 56(15), Article 445. https://doi.org/10.1007/s10489-026-07498-2

Image Credits: AI Generated

DOI: 10.1007/s10489-026-07498-2

Keywords: feature selection, streaming features, graph-based scoring, sparse reconstruction, multi-task Lasso, Laplacian score, machine learning, high-dimensional data, data mining, classification accuracy, dimensionality reduction, Applied Intelligence

Cite Scienmag News

APA
MLA
Chicago

Juliet Wilcox. (October 4, 2026). New Graph-Based Algorithm Tames Streaming Features in High-Dimensional Data. Scienmag. https://scienmag.com/new-graph-based-algorithm-tames-streaming-features-in-high-dimensional-data/

Juliet Wilcox. “New Graph-Based Algorithm Tames Streaming Features in High-Dimensional Data.” Scienmag, 4 October 2026, https://scienmag.com/new-graph-based-algorithm-tames-streaming-features-in-high-dimensional-data/. Accessed 4 October 2026.

Juliet Wilcox. “New Graph-Based Algorithm Tames Streaming Features in High-Dimensional Data.” Scienmag. October 4, 2026. https://scienmag.com/new-graph-based-algorithm-tames-streaming-features-in-high-dimensional-data/

Copy citation
Download RIS

Tags: Applied Intelligenceclassification accuracydata miningdimensionality reductionfeature redundancy reductionfeature selectionfeature selection in gene expression and spam filteringgraph-based feature selectiongraph-based scoringgroup streaming feature selectionhigh-dimensional datahigh-dimensional data algorithmsLaplacian scoreMachine learningmodular sensor data processingmulti-task Lassomulti-view learning systemsreal-time data streaming analysisscalable machine learning methodssparse reconstructionsparse reconstruction in machine learningstreaming featuresStreaming high-dimensional data analysisstructural relationship preservation

Share12Tweet7Share2ShareShareShare1

Related Posts

Two Weeks of Overeating Weakens the Gut Barrier and Ignites Liver Immunity in Healthy Men

Two Weeks of Overeating Weakens the Gut Barrier and Ignites Liver Immunity in Healthy Men

October 4, 2026
Perovskite Electrodes Could Supercharge the Next Generation of Lithium-Ion Batteries

Perovskite Electrodes Could Supercharge the Next Generation of Lithium-Ion Batteries

October 4, 2026

Hybrid CNN-ViT Model With Triple Loss Boosts Image Search Accuracy

October 4, 2026

Ancient Chinese Herb Supercharges Lab-Grown Bone Organoids to Heal Defects

October 4, 2026

POPULAR NEWS

  • Two Weeks of Overeating Weakens the Gut Barrier and Ignites Liver Immunity in Healthy Men

    Two Weeks of Overeating Weakens the Gut Barrier and Ignites Liver Immunity in Healthy Men

    29 shares
    Share 12 Tweet 7
  • How Learning Health Systems Could Transform Cardiovascular Care

    29 shares
    Share 12 Tweet 7
  • Chemical Tags on RNA May Steer a Sheep’s Dramatic Black-to-White Coat Change

    29 shares
    Share 12 Tweet 7
  • Perovskite Electrodes Could Supercharge the Next Generation of Lithium-Ion Batteries

    29 shares
    Share 12 Tweet 7

About

We bring you the latest biotechnology news from best research centers and universities around the world. Check our website.

Follow us

Recent News

Two Weeks of Overeating Weakens the Gut Barrier and Ignites Liver Immunity in Healthy Men

How Learning Health Systems Could Transform Cardiovascular Care

Chemical Tags on RNA May Steer a Sheep’s Dramatic Black-to-White Coat Change

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 85 other subscribers
  • Contact Us

Bioengineer.org © Copyright 2023 All Rights Reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Homepages
    • Home Page 1
    • Home Page 2
  • News
  • National
  • Business
  • Health
  • Lifestyle
  • Science

Bioengineer.org © Copyright 2023 All Rights Reserved.