Assistant Professor Jiaqi Ma has received a five-year, $660,307 National Science Foundation CAREER award to investigate one of the most consequential and least understood forces shaping modern artificial intelligence: the data used to train it. His project, “Data Attribution and Curation for Web-Scale AI Systems,” will develop methods for determining how individual training examples influence the behavior of large AI models, potentially changing how researchers build, audit, improve, and govern these systems.
The NSF CAREER award is among the most selective honors available to early-career researchers in the United States. It recognizes faculty members whose work combines strong research potential with a commitment to education and broader public impact. Ma, an assistant professor at the University of Illinois School of Information Sciences, will use the award to study data attribution, a family of techniques designed to estimate the contribution of specific training examples to a model’s predictions, capabilities, errors, or harmful behavior.
The challenge is enormous because today’s AI systems are trained on datasets that can contain billions or even trillions of pieces of information. These collections may include books, websites, images, code, conversations, scientific papers, and other forms of digital content gathered from many sources. Training data can be duplicated, outdated, biased, misleading, copyrighted, or difficult to trace. Once these examples have been processed through a model’s training procedure, identifying which data influenced a particular output becomes extremely difficult.
Data attribution seeks to make that influence measurable. In principle, an attribution method could help determine whether a model’s answer was supported by reliable information, shaped by a particular document, or affected by a problematic example. It could also help researchers identify which training samples improve performance on a task and which ones cause errors. Such information would provide a more precise alternative to treating an entire dataset as equally valuable.
“Data is one of the most important ingredients in modern AI,” Ma said, “but we still have a limited understanding of how individual training examples shape a model’s behavior.” His project will focus on developing attribution tools that remain useful at web scale, where the size and constant evolution of datasets make conventional analysis difficult. The research will also address how attribution can work when models are repeatedly updated with new information rather than trained only once.
Understanding these relationships could transform data curation, the process of deciding what information should be included in a training dataset. Instead of relying primarily on broad filtering rules or simple quality scores, developers could use attribution signals to prioritize examples that improve reasoning, remove data associated with recurring failures, and identify content that contributes to unsafe or unreliable outputs. The same approach could help reveal gaps in a dataset, including areas where a model performs poorly because relevant perspectives or information are missing.
The implications extend beyond accuracy. AI systems increasingly influence search results, recommendations, education, employment, scientific research, and creative work. If training data contains harmful stereotypes, misinformation, or material that should not have been used, those problems can be reflected in model behavior. Attribution could offer researchers a way to trace such effects back to their sources and evaluate whether removing or modifying particular examples changes the system’s responses. It may also support machine unlearning, in which a model is adjusted to reduce the influence of selected data after training.
Ma also sees a possible connection between data attribution and compensation. Much of the content used to develop AI systems is created by people whose contributions may be difficult to identify once the material has been absorbed into a massive dataset. More transparent attribution methods could eventually help determine whether and how particular works contribute to a model’s capabilities. That could create a technical foundation for recognizing contributors and designing more equitable systems for compensating them, although the legal and economic mechanisms for doing so remain unresolved.
The project will combine research with education and public access. Ma plans to create open-source software, conference tutorials, course modules, and other educational resources. The effort will provide research opportunities for graduate and undergraduate students, as well as pre-college learners, with the goal of widening participation in data-centered AI research. By making the tools publicly available, the project could allow academic researchers, technology developers, and policymakers to examine AI systems using shared methods rather than relying only on proprietary evaluations.
Ma’s research spans machine learning and artificial intelligence, with a particular focus on the data foundations of AI. His work examines how training data affect models through attribution, how data-centric algorithms can improve the quality and safety of datasets through curation and synthetic data generation, and how data mediate the societal consequences of AI through compensation and machine unlearning. He received a Best Paper Award at the 2024 International Conference on Learning Representations workshop on navigating and addressing data problems for foundation models and was named a New Faculty Highlight by the Association for the Advancement of Artificial Intelligence in 2025. Ma earned his PhD from the University of Michigan and completed postdoctoral research at Harvard University.
Subject of Research: Data attribution and curation for web-scale artificial intelligence systems
Keywords
Artificial intelligence, machine learning, data attribution, data curation, training data, large language models, foundation models, machine unlearning, synthetic data, AI safety, AI transparency, NSF CAREER award
Tags: AI training data analysischallenges in managing massive training datasetsdata attribution in machine learningdata curation for web-scale AIearly-career AI researcher Illinoisenhancing AI model accuracy through data attributionethical governance of AI training datasetsimpact of training data on AI model behaviorimproving AI model transparency and accountabilitylarge-scale AI system developmentmethods for training data influence estimationNSF CAREER award for AI research



