• HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
Tuesday, September 22, 2026
BIOENGINEER.ORG
No Result
View All Result
  • Login
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
No Result
View All Result
Bioengineer.org
No Result
View All Result
Home NEWS Science News Biology

New Statistical Tool Speeds Up Gene Perturbation Screening in Single Cells

Bioengineer by Bioengineer
September 22, 2026
in Biology
Reading Time: 6 mins read
0
New Statistical Tool Speeds Up Gene Perturbation Screening in Single Cells
Share on FacebookShare on TwitterShare on LinkedinShare on RedditShare on Telegram

Single-cell CRISPR screens have transformed functional genomics by allowing researchers to knock out or modulate thousands of genes at once and then read out the consequences in individual cells. In the most popular implementation, known as Perturb-seq, a pooled library of guide RNAs is delivered into a population of cells, and single-cell RNA sequencing captures both each cell’s transcriptome and the identity of the guide RNA it received. A deceptively difficult computational step sits at the heart of every such experiment: deciding, from noisy sequencing counts, which guide RNA actually landed in which cell. A newly published method called Fishash promises to make that step faster, more statistically rigorous, and easier to apply at scale.

The method, described in the journal BMC Bioinformatics by Jack Kamm, Jake Yeung, and William F. Forrest of Genentech, reframes guide assignment as a problem in classical statistics. Rather than fitting an elaborate probabilistic model to the guide count matrix, the authors propose treating the matrix of guide RNA counts across cells as a contingency table. For every cell-guide pair, they construct a two-by-two table that contrasts the counts involving that cell and that guide with all other counts, and then apply Fisher’s Exact Test, a workhorse of categorical data analysis, to ask whether the cell and guide barcodes are statistically associated. A significant association signals that the guide is genuinely present in the cell rather than appearing by chance through sequencing noise.

The elegance of this formulation lies in what it provides automatically. Because Fisher’s Exact Test conditions on the margins of the contingency table, the resulting p-values implicitly normalize for both cell-specific and guide-specific size factors, the uneven sequencing depth that plagues single-cell data. Cells that received more total reads and guides that were more abundant in the library are handled without any explicit normalization step. The method also produces a p-value for every cell-guide pair, giving users a graded measure of confidence rather than a hard binary call, and enabling downstream analyses such as full precision-recall curves when benchmarking alternative cutoffs.

Two statistical subtleties demanded additional innovation. First, because the method performs one test for every cell-guide combination, large screens can involve tens of millions of hypothesis tests. Naive application of standard multiple testing corrections such as the Benjamini-Hochberg procedure would be inappropriate, because the tests are strongly correlated: a cell that tests positive for one guide is less likely to test positive for others, and counts involving the same guide or the same cell share information. The authors developed a multiple testing correction strategy that accounts for this correlation structure, preserving control of the false discovery rate while avoiding the excessive conservatism that would come from ignoring the dependencies.

The second subtlety is a form of confounding that can flip conclusions in exactly the wrong way. Ambient RNA contamination, sometimes called soup or background noise, introduces guide molecules into cells that never received them. If the contamination is not independent across cells and guides, a phenomenon analogous to Simpson’s paradox can arise: an association that appears significant in the aggregate data disappears, or even reverses, once the noise component is considered separately. Fishash addresses this by testing an adjusted odds ratio that replaces the margins involving the focal cell and guide with their noise-based counterparts. To estimate the unobserved noise counts, the method performs an iterative rank-1 Poisson matrix completion: entries that have already been confidently assigned are masked out, the remaining background structure is fit under a Poisson likelihood by alternating updates of guide and cell abundance parameters, and the imputed noise counts feed back into refined p-values. The procedure repeats until the assignments stabilize.

Recognizing that fair comparison of guide assignment methods has been hampered by a lack of realistic test data, the authors also introduced a simulation framework that generates synthetic guide counts under a detailed model of sequencing noise. The simulator draws on the contamination model popularized by the Cellbender software, incorporating latent cell-level factors such as ambient RNA fractions, capture efficiencies, and dropout rates, along with guide-level expression variation. Crucially, the simulations can vary the number of guides in the library and the multiplicity of infection, the average number of guides delivered per cell, allowing benchmarking across the parameter regimes that matter in practice. The code to reproduce all results is publicly available alongside the method itself.

In benchmarks on both simulated and real datasets, Fishash compared favorably with existing approaches in both accuracy and runtime. On simulations varying the number of guides, the method achieved the top median F1 score in the majority of settings, performing particularly well when the guide library was large or guide RNA expression was low. In simulations varying the multiplicity of infection, Fishash remained competitive at low multiplicity, where most cells carry at most one guide, though other methods such as crispat-NB and dcCLEANSER edged ahead at high multiplicity. The authors emphasize that Fishash was conservative in its error control: its observed precision stayed above its nominal lower bound across the sweep, whereas several competing Bayesian and frequentist methods failed to control the false discovery rate at least once across the scenarios tested.

The stress tests also revealed the method’s limits, which the authors report transparently. When extra overdispersion was injected into the noise counts by replacing the Poisson distribution with a Geometric distribution, one of the most extreme overdispersion regimes available, precision dropped for many methods and Fishash no longer controlled the false discovery rate at its nominal five percent level in every scenario. Even so, its F1 scores remained qualitatively similar to the original benchmarks, indicating a competitive balance between precision and recall even under unusually harsh noise. The authors note that the conditional low-rank Poisson structure assumed for background counts is common to many single-cell contamination models, including Cellbender, SoupX, DecontX, and scAR, so the simulation represents a deliberately challenging departure from the assumptions shared across the field.

Practical accessibility was a design priority. Fishash is distributed as an easy-to-use R package on GitHub, and because it relies on exact tests over contingency tables rather than iterative model fitting, it scales comfortably to screens with tens of thousands of cells and guides, where computationally expensive probabilistic approaches can become prohibitive. The software also outputs its test statistic for every entry of the cell-guide matrix, which allows users to visualize the distribution of scores, select alternative significance cutoffs informed by the data, and compute full precision-recall curves during method benchmarking. The work was carried out entirely within Genentech, with the authors crediting colleagues in the company’s AI Biology and Translation department for discussions that shaped the manuscript.

For the growing community running Perturb-seq experiments, the significance of the method is straightforward: guide assignment is the gatekeeper step that determines whether downstream perturbation effects are real or artifacts of misassignment, and errors made here propagate through every subsequent analysis. By combining a classical, well-understood statistical test with modern corrections for multiplicity, confounding, and background contamination, and by packaging the result in fast, open software, Fishash lowers both the computational and conceptual barriers to reliable pooled single-cell screening. As CRISPR screens continue to expand in scale and ambition, tools that trade heavy model fitting for careful, assumption-aware statistics are likely to play an increasingly central role in turning raw sequencing counts into trustworthy biological conclusions.

Subject of Research: A contingency table and Fisher’s Exact Test based method for assigning guide RNAs to cells in single-cell pooled CRISPR screens (Perturb-seq)

Article Title: Fishash: a contingency table approach to Perturb-seq guide assignment

Article References: Kamm, J., Yeung, J., & Forrest, W. F. (2026). Fishash: a contingency table approach to Perturb-seq guide assignment. BMC Bioinformatics. https://doi.org/10.1186/s12859-026-06631-9

Image Credits: AI Generated

DOI: 10.1186/s12859-026-06631-9

Keywords: Perturb-seq, CRISPR screen, single-cell RNA-seq, guide assignment, Fisher’s exact test, Simpson’s paradox, multiple testing correction, contingency table, false discovery rate, ambient RNA contamination, functional genomics, BMC Bioinformatics

Cite Scienmag News
APA MLA Chicago

Juliet Wilcox. (September 22, 2026). New Statistical Tool Speeds Up Gene Perturbation Screening in Single Cells. Scienmag. https://scienmag.com/new-statistical-tool-speeds-up-gene-perturbation-screening-in-single-cells/

Juliet Wilcox. “New Statistical Tool Speeds Up Gene Perturbation Screening in Single Cells.” Scienmag, 22 September 2026, https://scienmag.com/new-statistical-tool-speeds-up-gene-perturbation-screening-in-single-cells/. Accessed 22 September 2026.

Juliet Wilcox. “New Statistical Tool Speeds Up Gene Perturbation Screening in Single Cells.” Scienmag. September 22, 2026. https://scienmag.com/new-statistical-tool-speeds-up-gene-perturbation-screening-in-single-cells/

Copy citation Download RIS

Tags: ambient RNA contaminationanalysis of pooled guide RNA librariesbioinformatics tool developmentBMC Bioinformaticscomputational tools for gene editingcontingency tableCRISPR screenfalse discovery rateFisher’s exact testFisher’s Exact Test in genomicsfunctional genomicsgene perturbation screeningguide assignmentguide RNA assignmentmultiple testing correctionnoise reduction in single-cell sequencingPerturb-seqscalable gene perturbation analysisSimpson’s paradoxsingle-cell CRISPR screenssingle-cell RNA sequencing analysissingle-cell RNA-seqstatistical methods in functional genomics

Share12Tweet7Share2ShareShareShare1

Related Posts

Novel KRIT1 Gene Variant Linked to Severe Pediatric Familial Brain Vessel Malformations

Novel KRIT1 Gene Variant Linked to Severe Pediatric Familial Brain Vessel Malformations

September 22, 2026
How Plants Turn Stress Into Survival: Hormone, Kinase and Metabolite Networks Revealed

How Plants Turn Stress Into Survival: Hormone, Kinase and Metabolite Networks Revealed

September 22, 2026

READ INSTRUCTIONS

September 22, 2026

Induced Pluripotent Stem Cells Reshape Drug Discovery for Alzheimer’s, Diabetes and Cancer

September 22, 2026

POPULAR NEWS

  • Liposomal Paclitaxel With HER-2 Blockers Shows Strong Real-World Results in Advanced Breast Cancer

    29 shares
    Share 12 Tweet 7
  • Transplanted and Supported Hearts Keep a Hormonal Memory of Heart Failure, Study Finds

    29 shares
    Share 12 Tweet 7
  • How MoS2 Heterojunction Band Alignment Controls Sunlight-Driven Dye Breakdown

    29 shares
    Share 12 Tweet 7
  • Novel KRIT1 Gene Variant Linked to Severe Pediatric Familial Brain Vessel Malformations

    29 shares
    Share 12 Tweet 7

About

We bring you the latest biotechnology news from best research centers and universities around the world. Check our website.

Follow us

Recent News

Liposomal Paclitaxel With HER-2 Blockers Shows Strong Real-World Results in Advanced Breast Cancer

Transplanted and Supported Hearts Keep a Hormonal Memory of Heart Failure, Study Finds

How MoS2 Heterojunction Band Alignment Controls Sunlight-Driven Dye Breakdown

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 85 other subscribers
  • Contact Us

Bioengineer.org © Copyright 2023 All Rights Reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Homepages
    • Home Page 1
    • Home Page 2
  • News
  • National
  • Business
  • Health
  • Lifestyle
  • Science

Bioengineer.org © Copyright 2023 All Rights Reserved.