Single cells are transforming biology from a study of averages into a census of individual molecular lives. A new computational framework called MIRACLE is designed to make that census expandable, allowing researchers to integrate new single-cell multimodal datasets continuously rather than rebuilding an analysis from the beginning every time fresh data arrive. The work, reported by J. Zhou, J. Wang, S. Hu and colleagues in Nature Computational Science, addresses one of the fastest-growing challenges in modern genomics: how to combine measurements made with different technologies, in different laboratories and at different times, without losing the biological signal hidden inside each cell.
Single-cell experiments can measure several molecular layers at once, including gene activity, chromatin accessibility and surface proteins. These readouts describe different aspects of the same cell, but they do not naturally share the same scale, noise profile or biological vocabulary. A transcriptomic measurement records RNA abundance, for example, while an epigenomic assay can reveal which regions of DNA are accessible to regulatory machinery. Protein measurements add another layer of information, often closer to the cell’s functional state. Integrating these modalities can reveal cell types and transitions that remain invisible when each dataset is analyzed separately.
The problem becomes substantially harder when datasets are generated sequentially. Conventional integration pipelines often assume that all samples are available at once. Researchers may therefore need to rerun a complete analysis whenever a new experiment is added, a process that can be computationally expensive and can alter the interpretation of previously analyzed cells. More importantly, repeatedly training a model on new data can cause “catastrophic forgetting,” in which the system adapts to the latest dataset but loses information learned from earlier ones. MIRACLE is presented as a strategy for continual integration, enabling a growing collection of multimodal data to be updated over time.
At the center of such systems is a shared representation, often called a latent space. Instead of comparing raw measurements directly, a computational model converts data from different modalities into a lower-dimensional map in which cells with related biological states should occupy nearby positions. The challenge is to make that map meaningful across technologies while preserving genuine differences between cell populations. If integration is too weak, equivalent cells measured by different assays remain artificially separated. If it is too aggressive, the algorithm may erase real biological variation along with technical differences.
MIRACLE is designed to address this balance as new datasets enter the analysis. Its continual-learning framework can update an existing representation rather than treating every incoming experiment as an entirely new problem. In principle, this allows a reference atlas to grow as additional tissues, conditions, developmental stages or disease samples are collected. New cells can be compared with previously characterized populations, while the model can also accommodate states that were not represented in the original reference. This is particularly important for biology, where rare cell types may appear only after many experiments or under specific pathological conditions.
The framework also speaks to the practical reality of biomedical research. Single-cell datasets are frequently produced on different sequencing platforms, with distinct antibody panels, varying levels of sequencing depth and laboratory-specific protocols. These technical differences can create batch effects—patterns that reflect how a sample was measured rather than what the cells were doing biologically. A useful integration method must reduce those unwanted effects without flattening meaningful signals, such as disease-associated activation, cellular stress or developmental progression. MIRACLE’s purpose is to provide a computational route through that tension while retaining the multimodal structure of the data.
Continual integration could be especially valuable for large biological atlases and clinical research programs. A hospital or consortium may accumulate samples over months or years, with new patients, tissues and disease subtypes added as they become available. Instead of freezing an atlas at the moment of its original construction, researchers could update it as new evidence emerges. That creates the possibility of more responsive reference maps for identifying unusual cell states, comparing treatment responses and tracking how molecular programs change across individuals. The method does not eliminate the need for careful experimental design, but it could make the resulting datasets easier to reuse.
The technical significance of MIRACLE lies not simply in combining multiple data types, but in treating integration as an ongoing process. This distinction matters because biological knowledge is rarely assembled in one final batch. New technologies reveal new molecular layers, and improved experiments can challenge earlier classifications. A continually updated model must therefore preserve stable information, learn from incoming observations and remain capable of representing novelty. These requirements bring single-cell analysis closer to modern machine-learning systems that are expected to operate on data streams rather than on one fixed dataset.
For scientists, the larger promise is a more durable computational foundation for single-cell biology. Instead of generating isolated maps that are difficult to compare, laboratories could build connected, evolving representations of cells across tissues, conditions and experiments. Such maps may help researchers distinguish universal cellular programs from context-specific responses and identify relationships between molecular layers that would otherwise be missed. As multimodal technologies become more common, tools capable of integrating data continuously could become essential infrastructure for translating an ever-expanding flood of single-cell measurements into biological discovery.
Subject of Research: Continual integration of single-cell multimodal data using a machine-learning framework.
Article Title: Continual integration of single-cell multimodal data with MIRACLE
Article References: Zhou, J., Wang, J., Hu, S. et al. Continual integration of single-cell multimodal data with MIRACLE. Nature Computational Science (2026). https://doi.org/10.1038/s43588-026-01030-9
Image Credits: AI Generated
DOI: https://doi.org/10.1038/s43588-026-01030-9
Keywords: single-cell genomics, multimodal data integration, continual learning, machine learning, computational biology, latent representations, batch effects, cell atlases, transcriptomics, epigenomics
Tags: biological signal preservation in multi-omicscomputational framework for multi-omicscontinuous single-cell data analysiscross-laboratory single-cell genomicsMIRACLE single-cell data platformmulti-layered single-cell data analysismulti-modal data normalizationmulti-technology single-cell measurementsscalable single-cell data analysis toolssingle-cell data integration challengessingle-cell multimodal data integrationsingle-cell transcriptomics and epigenomics integration


