PHILADELPHIA, PA — A new study is bringing scientists closer to an idea that has long seemed almost impossible: a quantitative map of smell. While color can be described using standardized coordinates for hue, saturation and brightness, odors have resisted similar treatment because they are often produced by complex mixtures rather than single substances. Researchers at the Monell Chemical Senses Center and collaborating institutions have now developed a machine-learning framework that can estimate how similar or different two odor mixtures will seem to people. The work, published online August 4 in the Proceedings of the National Academy of Sciences, could become an important step toward digital olfaction—the ability to record, compare and eventually reproduce smells using data.
The researchers’ goal was not simply to identify whether a substance smelled like roses, gasoline or fresh bread. Instead, they wanted to measure the perceptual distance between two odors. In their system, a score of 0 represents mixtures that people perceive as indistinguishable, while a score of 1 represents mixtures judged to be maximally different. This continuous scale gives researchers a way to compare smells mathematically, much as scientists compare colors or sounds. Such a tool could help determine whether two fragrances are genuinely close in perception, even when their chemical compositions are very different, or whether two chemically related mixtures create surprisingly different experiences.
The project was led in part by Joel Mainland, a scientist at the Monell Chemical Senses Center, who has spent years investigating how odor perception might be represented computationally. Previous efforts, including the 2015 IBM DREAM challenge, focused largely on predicting the smell of individual molecules from their chemical structures. Those studies helped establish links between molecular features and odor descriptors, but everyday smells rarely come from one molecule alone. Coffee, perfume, cooked food, cleaning products and human odors may contain dozens or even hundreds of volatile compounds. Understanding how those components combine is therefore essential if researchers hope to digitize real-world odors rather than laboratory chemicals in isolation.
To build a resource for the new challenge, Mainland and his colleagues standardized and combined six existing datasets drawn from three separate studies of odor similarity. The resulting collection included 168 unique individual molecules, 731 distinct mixtures and 507 measured pairs of mixtures. Participants were given information that allowed them to train predictive models, but the true similarity ratings for a hidden test set remained concealed. Over three months, 26 international teams competed in a DREAM challenge to predict how similar 46 unseen mixture pairs would appear to human observers. The contest ended in a four-way tie, reflecting the difficulty of the task and the comparable performance of several machine-learning strategies.
The scientists then created an ensemble model, a common technique in modern predictive analysis that combines the outputs of multiple models. Rather than relying on one algorithm, they averaged the predictions produced by the four winning teams and two other high-performing systems. Ensemble methods can improve reliability because different models may capture different patterns in the data: one may recognize relationships among odor descriptors, while another may respond more strongly to molecular or mixture-level features. After the challenge, the combined system was tested on an independent collection of 50 additional odor-mixture pairs, providing a separate check against overfitting the original competition data.
The final model performed with a median root mean squared error of 0.08. Root mean squared error measures the average size of prediction errors, while giving greater weight to larger mistakes; a value of 0 would indicate perfect predictions. The model also reached a Pearson correlation of 0.57 on the challenge test set, indicating a moderate positive relationship between its predictions and the measured human judgments. These results do not mean that machines can perfectly understand the subjective experience of smell. Rather, they show that computational models can capture enough of the structure of odor perception to make useful predictions about the relative similarity of complicated mixtures.
One of the study’s most striking findings was that the models benefited substantially from semantic language. Descriptions such as “fruity,” “sweet,” “woody” or “floral” were more useful than researchers might have expected, sometimes contributing more to prediction accuracy than strictly chemical information such as molecular weight or the presence of particular chemical groups. When the investigators removed semantic features from the models, performance became considerably worse. This suggests that words used by people to describe smells may act as a powerful bridge between chemistry and perception. A label such as “fruity” does not identify one molecule, but it can summarize a pattern of sensory effects that emerges from many compounds acting together.
The findings also challenge a common assumption in sensory science: that mixtures would be dramatically more difficult to predict than their individual components. According to Mainland, the results indicate that if researchers can estimate how the separate ingredients smell, they can make a reasonably good prediction of the overall mixture. The relationship is not necessarily a simple chemical average, because odor interactions can involve masking, enhancement and changes in perceived intensity. Yet the study suggests that the perceptual character of a mixture often remains connected to the sensory profiles of its ingredients. A subsequent DREAM challenge is examining this question more directly by asking teams to use people’s descriptions of individual components to predict how the complete mixture will be perceived.
A reliable odor-similarity metric could have applications far beyond fragrance research. Some diseases, including diabetes and liver failure, can produce characteristic volatile chemical signatures, and computational odor maps might eventually help researchers compare those signatures or support diagnostic technologies. Food scientists could use the models to quantify flavor and aroma profiles, while fragrance and consumer-product companies might replace some trial-and-error testing with mathematical design. A manufacturer could seek a new detergent scent that is perceptually close to an established product, or deliberately move in a different sensory direction while preserving selected characteristics. Digital olfaction could also support electronic noses, automated quality control and systems designed to transmit or reproduce scent information.
The researchers emphasize that the work is a foundation rather than a finished “Pantone wheel” for smell. Human odor perception varies across individuals and can be influenced by experience, context, concentration and biology. The current datasets are also limited compared with the enormous number of possible chemical mixtures encountered in the real world. Even so, the study establishes a benchmark against which future models can be tested and improved. Its combination of human sensory data, machine learning and semantic descriptions offers a practical route toward representing odors in a shared computational language. Mainland and colleagues are now preparing results from another olfaction challenge for publication, continuing an effort that could eventually make smell as measurable, searchable and digitally manipulable as images and sound.
Subject of Research: Data and statistical analysis of odor-mixture perception
Article Title: A Semantic-Based Community Model for High-Fidelity Tuning of Olfactory Mixture Distances
News Publication Date: August 20, 2026
Web References: https://www.pnas.org/doi/10.1073/pnas.2611057123; DOI: https://doi.org/10.1073/pnas.2611057123
References: Proceedings of the National Academy of Sciences, article published August 4, 2026; DOI: 10.1073/pnas.2611057123
Keywords: digital olfaction, machine learning, odor mixtures, smell perception, sensory science, computational modeling, fragrance science, human health, disease biomarkers, semantic odor descriptors
Tags: complex odor mixture analysiscomputational olfactory modelingdigital olfaction developmentdigital smell mappingmachine learning for olfactionmachine learning in scent researchodor mixture similarity scoringodor similarity measurementolfactory data analysisperceptual distance between odorsquantitative odor comparisonscent perception measurement


