Environmental exposures rarely arrive one at a time. People encounter complex mixtures of pesticides, industrial chemicals, pharmaceuticals, plastics, combustion products and other contaminants through air, food, water, workplaces and consumer goods. Yet conventional exposure studies often examine only a limited number of compounds, leaving much of an individual’s chemical environment invisible. A new protocol published in Nature Protocols presents an expanded version of MetaboAnalyst 6.0 designed to make exposomics—the large-scale study of environmental exposures and their biological effects—more accessible, systematic and biologically informative. The workflow connects high-resolution liquid chromatography–tandem mass spectrometry, statistical analysis, dose–response modeling and genetic evidence in a single analytical framework.
The protocol addresses a central challenge in modern exposomics: converting enormous chemical datasets into reliable biological conclusions. Liquid chromatography separates molecules in a sample according to their chemical properties, while tandem mass spectrometry measures their mass-to-charge ratios and fragments them into characteristic product ions. These fragmentation patterns can act like molecular fingerprints, helping researchers determine which compounds are present in blood or other biological materials. However, identifying chemicals from LC–MS2 data is technically demanding. Signals may come from previously uncharacterized substances, isomers with nearly identical masses or metabolites produced by the body after exposure. MetaboAnalyst 6.0 adds improved support for interpreting these spectra and organizing the resulting chemical annotations.
The updated workflow is divided into four connected stages. Stage 1 focuses on LC–MS2 spectra processing and compound identification, beginning with the preparation and evaluation of raw spectral data. Researchers can process detected features, compare experimental fragmentation patterns with reference information and assign identification confidence according to the quality of the available evidence. This distinction is critical because a measured signal is not automatically a confirmed compound. Strong identification generally requires agreement between retention behavior, precursor mass and fragment ions, whereas weaker matches may only support a tentative structural class. By helping users manage these levels of certainty, the platform is intended to reduce overinterpretation while retaining useful information from unknown or partially characterized chemicals.
Stage 2 turns the identified or annotated features into an exposomics dataset suitable for statistical analysis. Data processing can include quality assessment, filtering of unreliable signals, normalization, transformation and evaluation of missing values. These steps help control technical variation caused by sample preparation, instrument drift or differences in ionization efficiency. Once processed, the data can be explored using multivariate statistical methods and visualization tools that reveal patterns among samples, exposure groups and chemical features. The protocol demonstrates this stage with data from a blood exposomics investigation involving electronic-waste exposure, an especially complex setting in which workers and nearby communities may encounter metals, flame retardants, solvents and combustion-related chemicals simultaneously.
Such datasets can contain thousands of molecular signals, many of which are correlated because they originate from the same source or biological pathway. Exploratory analysis helps researchers determine whether samples separate according to exposure status, identify chemicals that contribute most strongly to group differences and detect unusual observations that might distort downstream results. MetaboAnalyst’s statistical environment can support approaches such as principal component analysis, clustering and supervised group comparisons, while functional interpretation can connect chemical changes with metabolic pathways or biological processes. The goal is not simply to produce a longer list of altered molecules, but to build a coherent picture of how environmental mixtures are associated with changes in human physiology.
Stage 3 adds a dose–response perspective, moving beyond the question of whether an exposure group differs from a comparison group. In environmental health research, a key signal of biological relevance is whether molecular responses change systematically as exposure levels rise. The protocol describes modeling relationships between measured exposure concentrations and metabolic features, allowing investigators to identify compounds or pathways that increase, decrease or follow more complex patterns across an exposure gradient. Dose–response analysis can help distinguish potentially meaningful responses from random group variation, although it still requires careful interpretation. Confounding factors, nonlinear biology, exposure misclassification and differences in metabolism can all influence the observed relationship.
The fourth stage extends the analysis toward causal inference by incorporating known genetic associations. The protocol illustrates this approach through an investigation of the potential causal relationship between ʟ-isoleucine and type 2 diabetes. Genetic variants associated with lifelong differences in a molecular trait can sometimes be used as instrumental variables in Mendelian randomization analyses. Because genetic variation is established before disease develops and is generally less affected by later environmental or behavioral factors, it may provide evidence that complements conventional observational associations. However, genetic instruments must satisfy important assumptions, including a reliable association with the exposure of interest and limited influence through alternative biological pathways. The approach therefore strengthens causal assessment but does not eliminate the need for biological validation.
The integration of these capabilities reflects a broader shift in exposomics. Rather than treating exposure measurement, metabolite profiling, statistical analysis and causal biology as separate tasks, the new protocol presents them as stages of one reproducible research process. This is particularly important for studies of real-world chemical mixtures, where exposure patterns are difficult to simplify and where a single pollutant may interact with diet, medication, occupation, socioeconomic conditions and genetic background. A unified workflow can make it easier for researchers to document analytical decisions, compare findings across studies and move from raw spectral signals to hypotheses that can be tested experimentally or in larger populations.
The authors estimate that Stage 1 may require approximately two hours, depending on server load, while Stages 2 through 4 can be completed in roughly 90 minutes. These timelines are intended as practical guidance rather than guarantees, since data size, annotation complexity and statistical model selection can substantially affect the duration of an analysis. Even so, the protocol could lower the technical barrier for laboratories entering exposomics, particularly those that lack extensive bioinformatics infrastructure. As environmental chemical exposures continue to multiply and concerns about their health effects grow, tools that connect molecular detection with dose response and genetic evidence may help transform exposomics from a catalog of chemical signals into a more powerful system for understanding disease risk.
Subject of Research: Exposomics data analysis, LC–MS2 compound identification, dose–response modeling and genetic causal inference.
Article Title: Using MetaboAnalyst 6.0 for exposomics data analysis—from LC–MS2 spectra processing to dose–response modeling and causal inference.
Article References: Pang, Z., Lu, Y., Zhou, G. et al. Using MetaboAnalyst 6.0 for exposomics data analysis—from LC–MS2 spectra processing to dose–response modeling and causal inference. Nature Protocols (2026). https://doi.org/10.1038/s41596-026-01415-0
Image Credits: AI Generated
DOI: https://doi.org/10.1038/s41596-026-01415-0
Keywords: Exposomics, MetaboAnalyst 6.0, liquid chromatography–tandem mass spectrometry, LC–MS2, compound identification, electronic-waste exposure, dose–response analysis, metabolomics, Mendelian randomization, type 2 diabetes.
Tags: advanced analytical framework for exposomicsbiological effect of environmental chemicalschemical mixture analysisdose–response modelingenvironmental chemical identificationenvironmental exposure analysisexposomics workflowgenetic evidence in exposure studieshigh-resolution mass spectrometryLC–MS2 data processingmetabolomics and exposomics integrationstatistical analysis of chemical exposures



