• HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
Sunday, October 4, 2026
BIOENGINEER.ORG
No Result
View All Result
  • Login
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
No Result
View All Result
Bioengineer.org
No Result
View All Result
Home NEWS Science News Technology

New Graph-Based Algorithm Tames Streaming Features in High-Dimensional Data

Bioengineer by Bioengineer
October 4, 2026
in Technology
Reading Time: 5 mins read
0
New Graph-Based Algorithm Tames Streaming Features in High-Dimensional Data
Share on FacebookShare on TwitterShare on LinkedinShare on RedditShare on Telegram

Machine learning models are often drowning in data, but the problem is not always the sheer volume of samples. Increasingly, it is the flood of features—the individual measurable properties of each data point—that threatens to overwhelm algorithms. In domains ranging from gene expression profiling to spam filtering, features do not arrive all at once. They stream in sequentially, often in predefined groups, and waiting for the complete feature set before analysis begins is computationally prohibitive. A new study published in Applied Intelligence proposes a method designed precisely for this challenging scenario, and its results suggest a significant step forward in how machines can learn from data that never stops arriving.

The method, called FS-GSR—Feature Selection with Group Streaming via Graph-based Scoring and Sparse Reconstruction—was developed by Tianyuan Jia of The Hong Kong Polytechnic University. It tackles a core weakness of existing streaming feature selection techniques: most process features one at a time, ignoring the structural relationships and redundancies that exist both within and between groups of features. When features arrive in batches, as they do in multi-view learning systems or modular sensor deployments, sequential individual processing can disrupt group structures and degrade performance. FS-GSR instead operates at the group level, updating the selected feature subset in real time as each new batch arrives.

At the heart of the approach is a three-stage framework that functions like a computational funnel. The first stage performs a hybrid, graph-guided filtering of each incoming feature group. The algorithm constructs a similarity graph over the samples, connecting points that are close in the feature space, and computes a Laplacian Score for every feature. This score measures how smoothly a feature varies across locally similar samples—features that preserve the intrinsic geometric structure of the data receive low scores, while noisy, irregular features are penalized. In parallel, a supervised F-score evaluates each feature’s discriminative power by comparing between-class variance to within-class variance. The two measures are blended with equal weight, and only the top fraction of features by combined score survives this initial screening.

The second stage refines this candidate pool by preserving structure at a finer granularity. The algorithm performs a spectral decomposition of the graph Laplacian, extracting ten eigenvectors that form a low-dimensional embedding of the sample manifold. Each surviving feature is then evaluated by how well it can reconstruct this structural representation, using a closed-form projection solution that avoids costly iterative optimization. Features that best approximate the manifold’s skeleton are retained, ensuring that the selected subset captures not just individual relevance but the collective geometry of the data. Because the reconstruction coefficients have an explicit analytical solution, this stage adds minimal computational overhead.

The third and final stage is where FS-GSR departs most sharply from its predecessors. All candidate features accumulated from the first two stages are merged into a global pool, and a multi-task Lasso regression is applied across the entire set simultaneously, using one-hot encoded class labels as the response. The L2,1-norm penalty drives the coefficients of redundant features to exactly zero, yielding a compact, non-overlapping final feature set. The authors also provide theoretical backing for this stage, formulating it as a row-support recovery problem in high-dimensional multivariate regression. Under standard conditions on the covariance structure, they prove that the sparse reconstruction can recover the true set of informative features with high probability, provided the effective sample complexity exceeds a specific threshold.

The computational design is notable for its scalability. By confining expensive cubic operations to the fixed sample space rather than the growing feature dimension, the algorithm’s cost per streaming group scales linearly with the number of incoming features. Because the first two stages aggressively compress each group before global optimization, the feature pool fed into the multi-task Lasso remains small, preventing the exponential blow-up that plagues conventional structure-preserving methods. Scalability experiments on the Ovarian Cancer dataset confirmed near-linear growth in running time as feature dimension increased, and revealed a V-shaped trade-off in execution time as the number of streaming groups varied—very few large groups incur heavy intra-group matrix operations, while very many small groups accumulate loop overhead.

Empirically, the method was tested on five benchmark datasets spanning biomedical, image, and text domains, including Colon (62 samples, 2,000 features), Ovarian Cancer (253 samples, 15,154 features), MLL, COIL20, and PCMAC. These datasets deliberately covered contrasting correlation structures: the Ovarian Cancer data exhibited strong intra-group correlations of up to 0.88, while the PCMAC text dataset showed extremely sparse correlations of roughly 0.025. Against four established baselines—OGFS, Group-SAOLA, Fast-OSFS, and alpha-investing—FS-GSR consistently achieved the highest or near-highest classification accuracy and F1 scores under both KNN and SVM classifiers. On the Ovarian Cancer dataset, it reached accuracy and F1 values above 0.99.

Perhaps the most striking result concerned feature compression. In one experimental configuration, FS-GSR retained an average of just 8.12 features from the 15,154-dimensional Ovarian Cancer dataset, while simultaneously improving KNN accuracy from 0.9845 to 0.9960 and SVM accuracy from 0.9885 to 0.9980—a 63 percent compression relative to the two-stage variant alone. Ablation studies confirmed that each stage contributes distinct benefits: the first enables efficient local filtering, the second improves structural consistency, and the third enforces global sparsity and discriminative refinement. The authors note that the degree of compression adapts to each dataset’s intrinsic correlation structure, with strongly correlated data requiring more retained features to preserve associative signals.

Parameter sensitivity analyses reinforced the method’s practicality. Classification accuracy remained stable across broad plateaus for the retention ratios governing the first two stages, and the mixing weight balancing structural and discriminative scoring performed robustly between 0.2 and 0.8, peaking near 0.5. The embedding dimension saturated quickly, with accuracy plateauing once it exceeded five, justifying the default setting of ten. This robustness means practitioners can deploy the method without exhaustive grid searches, using a simple boundary-driven tuning strategy on validation data.

The implications extend well beyond the benchmark datasets. Streaming, group-wise feature arrival is characteristic of stepwise sensor deployment, progressive module activation in multi-view systems, and sequential extraction pipelines in bioinformatics and text mining. By combining graph-based manifold learning with sparse global reconstruction in an online framework, FS-GSR offers a template for learning systems that must remain both accurate and parsimonious as data evolves. The author identifies future extensions toward multi-label settings, incomplete data systems, and broader generalization, suggesting that the challenge of features that never stop flowing is one the field is only beginning to master.

Subject of Research: Group streaming feature selection for high-dimensional data using graph-based scoring and sparse reconstruction

Article Title: FS-GSR: Graph-based scoring and sparse reconstruction for group streaming feature selection

Article References: Jia, T. (2026). FS-GSR: Graph-based scoring and sparse reconstruction for group streaming feature selection. Applied Intelligence, 56(15), Article 445. https://doi.org/10.1007/s10489-026-07498-2

Image Credits: AI Generated

DOI: 10.1007/s10489-026-07498-2

Keywords: feature selection, streaming features, graph-based scoring, sparse reconstruction, multi-task Lasso, Laplacian score, machine learning, high-dimensional data, data mining, classification accuracy, dimensionality reduction, Applied Intelligence

Cite Scienmag News

APA
MLA
Chicago

Juliet Wilcox. (October 4, 2026). New Graph-Based Algorithm Tames Streaming Features in High-Dimensional Data. Scienmag. https://scienmag.com/new-graph-based-algorithm-tames-streaming-features-in-high-dimensional-data/

Juliet Wilcox. “New Graph-Based Algorithm Tames Streaming Features in High-Dimensional Data.” Scienmag, 4 October 2026, https://scienmag.com/new-graph-based-algorithm-tames-streaming-features-in-high-dimensional-data/. Accessed 4 October 2026.

Juliet Wilcox. “New Graph-Based Algorithm Tames Streaming Features in High-Dimensional Data.” Scienmag. October 4, 2026. https://scienmag.com/new-graph-based-algorithm-tames-streaming-features-in-high-dimensional-data/

Copy citation
Download RIS

Tags: Applied Intelligenceclassification accuracydata miningdimensionality reductionfeature redundancy reductionfeature selectionfeature selection in gene expression and spam filteringgraph-based feature selectiongraph-based scoringgroup streaming feature selectionhigh-dimensional datahigh-dimensional data algorithmsLaplacian scoreMachine learningmodular sensor data processingmulti-task Lassomulti-view learning systemsreal-time data streaming analysisscalable machine learning methodssparse reconstructionsparse reconstruction in machine learningstreaming featuresStreaming high-dimensional data analysisstructural relationship preservation

Share12Tweet7Share2ShareShareShare1

Related Posts

Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

October 4, 2026
Heavy Metal Mixtures in Blood Linked to Higher Obesity Risk in Children

Heavy Metal Mixtures in Blood Linked to Higher Obesity Risk in Children

October 4, 2026

Federated AI Learns to Spot IoT Cyberattacks Without Sharing Private Data

October 4, 2026

Eggshells and Ceramic Trash Turned Into Ultra-Strong, Low-Carbon Cement Alternative

October 4, 2026

POPULAR NEWS

  • Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

    Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

    29 shares
    Share 12 Tweet 7
  • Endurance Exercise Reshapes the Liver in Males and Females Through Distinct Molecular Routes

    29 shares
    Share 12 Tweet 7
  • Single Transcription Factor PU.1 Rapidly Converts Fibroblasts into Macrophage-Lineage Cells

    29 shares
    Share 12 Tweet 7
  • New Scale Measures How Ready Nurse Educators Really Are for the AI Era

    29 shares
    Share 12 Tweet 7

About

We bring you the latest biotechnology news from best research centers and universities around the world. Check our website.

Follow us

Recent News

Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

Endurance Exercise Reshapes the Liver in Males and Females Through Distinct Molecular Routes

Single Transcription Factor PU.1 Rapidly Converts Fibroblasts into Macrophage-Lineage Cells

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 85 other subscribers
  • Contact Us

Bioengineer.org © Copyright 2023 All Rights Reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Homepages
    • Home Page 1
    • Home Page 2
  • News
  • National
  • Business
  • Health
  • Lifestyle
  • Science

Bioengineer.org © Copyright 2023 All Rights Reserved.