Artificial intelligence research in India has undergone one of the most dramatic expansions in modern scientific publishing, and a new study has now mapped that growth in unprecedented detail. By combining classical bibliometrics with machine-learning-based text mining, researchers Varun Kumar and Kunwar Singh of Indira Gandhi National Open University and Banaras Hindu University analysed 18,625 open access AI publications from India spanning four decades, from 1986 to 2025. Their findings, published in the journal Discover Artificial Intelligence, reveal not just a quantitative explosion but a profound structural transformation in what Indian AI scientists actually study, from early pattern recognition experiments to today’s deep learning and healthcare applications.
The scale of the dataset alone tells a striking story. The 18,625 publications accumulated 362,236 citations in total, averaging 19.45 citations per paper with a median of 6.0. A single paper drew a maximum of 1,552 citations. But the temporal pattern is where the data becomes truly remarkable. In the late 1980s and early 1990s, output was almost negligible: just four papers appeared between 1986 and 1990, and thirteen between 1991 and 1995. Growth was gradual through the early 2010s, with 50 papers in 2011 rising to 142 by 2015. Then came the inflection point. Annual output jumped from 169 publications in 2017 to 414 in 2018, a single-year increase of 145 percent, followed by 1,155 papers in 2019 and 1,113 in 2020.
The acceleration did not stop there. Publication counts climbed to 2,003 papers in 2021, 3,336 in 2022, and 4,014 in 2023, before peaking at 4,604 publications in 2024. The authors calculate that the 2021 to 2025 interval alone accounts for 14,993 publications, more than 80 percent of the entire four-decade dataset. This trajectory aligns closely with national policy initiatives such as Digital India and the National AI Mission, alongside the growing availability of computational resources and the global shift toward data-driven research. The study suggests the surge reflects a convergence of policy support, technological infrastructure, and researchers’ own enthusiasm for working with large datasets.
Citation patterns within the corpus follow the familiar rhythms of academic life. Papers published in 2018 command the highest average citation rate at 50.5 citations per paper, consistent with having had sufficient time to accumulate scholarly attention. More recent years show much lower averages, with 2024 papers at 5.47 and 2025 papers at just 2.01, a pattern the authors attribute to citation time-lag rather than lower quality. Of the full dataset, 16.2 percent of papers, some 3,025 publications, have received no citations so far, while 3.3 percent, or 615 papers, have crossed the 100-citation mark, and six papers have exceeded 1,000 citations. The high standard deviation of 52.23 relative to the mean of 19.45 reflects the heavily skewed, long-tailed distribution typical of academic citations, where a small number of standout papers drive overall totals.
The methodological heart of the study lies in its use of Latent Dirichlet Allocation, or LDA, a probabilistic topic modeling technique that treats each document as a mixture of underlying themes. The researchers applied LDA to the titles, abstracts, and author keywords of the corpus after extensive natural language processing, including tokenisation, stop-word removal, and lemmatisation using Python libraries such as NLTK, pandas, and NumPy. To determine the optimal number of topics, they evaluated models ranging from three to twelve topics using topic coherence scores computed with the Gensim implementation. The coherence measure peaked at six topics, with a Cv score of 0.58, so six themes were retained. Model parameters were set at alpha and beta values of 0.01, with 1,000 iterations and a fixed random seed to ensure reproducibility.
The six thematic clusters that emerged paint a vivid portrait of Indian AI research. Deep learning and neural networks dominate by publication volume, with 4,215 publications and 118,540 total citations, averaging 28.1 citations per paper. Artificial intelligence in healthcare ranks as the second major theme, reflecting the growing adoption of AI for clinical decision-making, disease prediction, and diagnosis. Computer vision and image processing forms the third cluster, driven by the proliferation of convolutional neural networks for image recognition and object detection. The remaining themes, spanning the Internet of Things and smart systems, cybersecurity and intrusion detection, and natural language processing and text analytics, demonstrate the interdisciplinary spread of AI across application domains far beyond core computer science.
Perhaps the most intriguing finding concerns citation impact across themes. Although deep learning captures the most publications, healthcare-related AI research achieves the highest average citation impact per paper, suggesting that application-oriented, socially relevant work attracts disproportionate scholarly attention. Computer vision and IoT themes show moderate citation impact, while cybersecurity research, being more domain-specific, draws more modest averages. Natural language processing shows the lowest citation impact, which the authors attribute to its relatively recent growth and the time needed for citations to accumulate. The researchers caution that these differences reflect multiple contributing factors, including field maturity, interdisciplinarity, and citation time-lag, rather than any single cause.
The thematic evolution analysis, conducted across eight consecutive five-year intervals, traces a clear intellectual arc. The earliest period, 1986 to 1990, was characterised by foundational themes in early machine learning and pattern recognition. Neural networks and intelligent systems emerged in the 1991 to 1995 window, followed by a shift toward data mining and classification methods from 1996 to 2000. Image processing and computer vision took centre stage in 2001 to 2005, with sensor networks and early IoT foundations appearing from 2006 to 2010. The 2011 to 2015 period was dominated by big data analytics and machine learning, generating 12,005 citations at an average of 30.32 per paper. The most significant transformation occurred in the 2016 to 2020 interval, where deep learning became the prevailing paradigm, and the final 2021 to 2025 window saw AI applications expand rapidly into healthcare, natural language processing, and cybersecurity.
Keyword analysis reinforces this narrative of transformation. Before 2015, general and foundational terms such as neural networks, artificial neural networks, and classification prevailed, indicating a field still centred on classic machine learning methods. After 2015, keywords like machine learning and deep learning surged to become leading research areas, and computer vision terms grew in importance with the rise of convolutional neural networks after 2018. One striking exception appears in 2020, when the keyword COVID-19 suddenly emerged on the charts, reflecting the research community’s rapid response to the global pandemic. Co-word analysis and hierarchical clustering further revealed a well-organised intellectual structure, with artificial intelligence, machine learning, and deep learning forming a tightly connected methodological core linked to application branches in healthcare, COVID-19 research, and cybersecurity. The authors note that keyword trends were not normalised by annual publication counts, so short-term events like the pandemic can appear more salient than they would under normalised measures.
The study is careful about its limits, and those caveats matter for interpreting the headline numbers. Because the dataset was restricted to open access publications via a Scopus search filter, with no non-OA comparison group included, the authors cannot attribute citation patterns causally to open access publishing itself; observed patterns are consistent with topic prominence, field maturity, and citation time-lag as well. The title-based search strategy may also miss relevant studies that discuss AI only in abstracts or keywords, and topic modeling was performed on metadata rather than full texts, with multi-word expressions treated as separate single words. Nevertheless, the practical implications are substantial. The authors call for a unified national open access policy in India, financial support to offset article processing charges that currently push researchers toward repository-based green open access, and better interoperability between institutional repositories. They also predict that generative and explainable AI, biomedical applications, and governance will shape the next wave of Indian AI research, domains likely to attract intense global collaboration as the country’s open access AI ecosystem continues its extraordinary ascent.
Subject of Research: Thematic structure and citation impact of open access artificial intelligence research in India analysed with bibliometrics and LDA topic modeling
Article Title: Exploring thematic structures and emerging trends in open access AI research in India using LDA
Article References: Kumar, V., & Singh, K. (2026). Exploring thematic structures and emerging trends in open access AI research in India using LDA. Discover Artificial Intelligence, 6(1), Article 1344. https://doi.org/10.1007/s44163-026-02428-0
Image Credits: AI Generated
DOI: 10.1007/s44163-026-02428-0
Keywords: artificial intelligence, open access, India, LDA, topic modeling, bibliometrics, deep learning, neural networks, healthcare AI, citation analysis, Scopus, research trends
News Source: Blake Davidson. (October 7, 2026). India’s Open Access AI Research Explodes Past 18,000 Papers, Machine Learning Themes Revealed. Scienmag.



