Personal data are becoming the raw material of modern science. Every smartphone interaction, electronic health record, financial transaction and medical image can help researchers identify patterns that improve treatment, train artificial intelligence and guide public policy. Yet the same information can expose deeply personal details about the people who generated it. A single dataset may reveal a patient’s illness, a family’s financial difficulties or an individual’s identity, even when obvious identifiers such as names and addresses have been removed. At the University of Virginia, computer scientist Tianhao Wang is working on a way to make sensitive information useful without turning privacy into the price of scientific progress.
Wang, an assistant professor in UVA’s School of Engineering and Applied Science, has received a $677,866 National Science Foundation Faculty Early Career Development Program, or CAREER, Award to develop privacy-preserving methods for creating synthetic data. His project, titled “Advancing Differentially Private Data Synthesis: A Holistic Approach,” focuses on generating artificial datasets that retain the patterns researchers need while reducing the risk that information about real individuals can be recovered. The award recognizes early-career faculty members whose research and teaching show the potential to influence their fields and mentor future generations of scientists.
The central technology behind Wang’s work is differential privacy, a mathematical framework designed to control how much an analysis can reveal about any one person. In simplified terms, a differentially private algorithm is constructed so that the outcome of a computation changes only slightly whether a particular individual’s data are included or excluded. If an attacker examines the result, it should be difficult to determine whether a specific person contributed to the original dataset. This protection is typically expressed through a formal privacy budget, often represented by a parameter called epsilon. Smaller values generally provide stronger privacy, although they can also make it harder for an algorithm to preserve fine-grained information.
“ My research focuses on a simple but increasingly important question: How can we use data to advance science and AI without exposing people’s private information?” Wang said. “The goal is to make more data safely usable for research and innovation.” He joined UVA in 2022 after earning his doctorate from Purdue University and completing postdoctoral research at Carnegie Mellon University. His work addresses a growing problem in artificial intelligence: the systems that learn from data can sometimes memorize unusual or sensitive examples instead of learning only broad, reusable patterns.
Synthetic data offer one possible solution, but simply labeling a dataset “synthetic” does not automatically make it safe. Modern generative models can produce records, images, text and other information that look entirely artificial while still reproducing details from their training data. If a model has memorized a rare medical image or an unusual combination of personal characteristics, a determined user may be able to extract that information. Wang’s project applies differential privacy during the generation process, limiting the influence any single record can have on the model. The goal is to teach an AI system the statistical “big picture” without allowing it to copy the private details of the people represented in the data.
The potential applications are enormous. A privacy-protected model might generate artificial medical images that preserve features associated with a disease while avoiding the reproduction of an identifiable patient. It could create a synthetic table of health records that reflects relationships between age, symptoms, treatments and outcomes without revealing an actual person’s diagnosis. Similar techniques could support research involving financial transactions, personal photographs, mobility data, educational records or consumer behavior. Organizations could share carefully evaluated synthetic datasets with collaborators, allowing researchers to test software and statistical methods without directly distributing the original sensitive records.
The scientific difficulty lies in preserving usefulness and privacy at the same time. Adding stronger privacy protections can reduce the amount of information a model retains, particularly when the original data are complex. Simple numerical tables may be relatively easy to summarize, but medical images, clinical notes and multimodal datasets combine many forms of information whose relationships can be essential. A synthetic record that looks realistic in isolation may still fail to preserve the correlations needed for a medical study. Conversely, a dataset that reproduces too many rare patterns may become vulnerable to re-identification. Wang’s research seeks methods that navigate this balance rather than treating privacy and accuracy as unrelated goals.
His CAREER project will pursue three connected objectives: preserving important statistical relationships in sensitive datasets, improving the quality of synthetic images and multimodal data, and using public datasets and foundation models to strengthen generation performance. Foundation models are large AI systems trained on broad collections of data and later adapted for specific tasks. They can provide powerful general representations, but they may also carry privacy risks inherited from their training material. Wang’s work will examine how such models can be incorporated into privacy-preserving pipelines while maintaining formal guarantees. The broader ambition is to discover principles that apply across different data types instead of requiring a completely separate technique for every domain.
Determining whether synthetic data are genuinely useful will require more than computer science benchmarks. A medical researcher may care about whether a synthetic dataset preserves clinically meaningful relationships between symptoms, treatments and outcomes. A public-policy analyst may need accurate population-level trends, while a financial researcher may focus on rare but consequential events. These standards cannot be defined by an algorithm alone. Wang hopes to collaborate with researchers at UVA’s School of Medicine and with specialists in other fields to evaluate synthetic datasets according to the scientific questions they are meant to answer. Such partnerships could reveal when a dataset is sufficiently realistic for research and when privacy safeguards have removed information that cannot be replaced.
Sandhya Dwarkadas, Walter N. Munster Professor and chair of UVA’s Department of Computer Science, said Wang’s research could help unlock data that organizations currently hesitate to share. The project will also train graduate and undergraduate students in algorithm design, applied cryptography, software development and experimental evaluation. Students will work on privacy-preserving computation, build open-source tools and explore applications in areas such as healthcare. Yucheng Fu, a second-year doctoral student studying differential privacy and applied cryptography, said Wang encourages students to develop projects around their own interests. Over time, the research could make privacy-protected data synthesis more reliable and accessible for universities, hospitals, companies and government agencies. If successful, it may help transform sensitive information from a locked resource into a carefully controlled public-science tool—without making the people behind the data pay for progress with their privacy.
Subject of Research: Privacy-preserving artificial intelligence, differential privacy and synthetic data generation.
Article Title: UVA Researcher Wins $677,866 NSF Award to Create Privacy-Protected Synthetic Data for AI and Science
Web References: https://engineering.virginia.edu/faculty/tianhao-wang ; https://www.nsf.gov/funding/opportunities/career-faculty-early-career-development-program ; https://www.nsf.gov/awardsearch/show-award/?AWD_ID=2543284
References: University of Virginia School of Engineering and Applied Science; U.S. National Science Foundation.
Keywords
Differential privacy, synthetic data, privacy-preserving AI, artificial intelligence, cybersecurity, data protection, healthcare technology, machine learning, multimodal data, data synthesis, Tianhao Wang, National Science Foundation CAREER Award
Tags: application of differential privacy in scientific researchartificial data for health and financial recordsdifferential privacy in data synthesisethical considerations in synthetic dataimpact of synthetic data on artificial intelligencemachine learning for data privacymethods to prevent data re-identificationNSF CAREER Award research UVApolicy implications of privacy-preserving data sharingprivacy-preserving synthetic data generationtraining future data privacy scientistsUVA engineering research on data anonymization




