Artificial intelligence has quietly crossed a threshold. The most capable systems no longer see the world through a single sense: they read text, interpret images, parse audio, and watch video, all within one unified model. A new open-access review published in Discover Informatics by Azhar A. Hadi and K. P. Supreethi of Jawaharlal Nehru Technological University Hyderabad offers the most systematic map yet of this transformation, dissecting thirteen pretrained multimodal deep learning models released between 2020 and 2025 and tracing how they are reshaping fields from radiology to wildlife monitoring.
The review, which screened more than 400 candidate papers down to 120 rigorously assessed studies, arrives at a moment when the field is accelerating faster than most surveys can track. Earlier reviews, the authors argue, tended to focus narrowly on single domains such as medicine or language processing, leaving researchers without a comparative view of the architectures, datasets, and training objectives that define the current generation of models. By covering eleven application domains, including healthcare, education, autonomous transportation, e-commerce, and industrial manufacturing, the paper aims to give both newcomers and specialists a coherent picture of what these systems can actually do.
At the technical heart of the review lies a family of architectures built on the transformer, the attention-based design that revolutionized natural language processing before conquering vision. The Vision Transformer, or ViT, broke with convolutional tradition by slicing images into fixed 16-by-16-pixel patches, flattening them into vectors, and processing them exactly like words in a sentence. This simple reframing allowed vision models to inherit the scaling behavior of language models, and it became the visual backbone for much of what followed. CLIP, developed by OpenAI, went further by training an image encoder and a text encoder jointly on 400 million image-text pairs scraped from the internet, learning to pull matching pairs together in a shared latent space while pushing mismatched ones apart. That contrastive objective gave machines a rudimentary form of grounded understanding, and CLIP now serves as the visual front end for many multimodal large language models.
Subsequent generations refined this recipe in strikingly different directions. BLIP introduced a flexible encoder-decoder design and a bootstrapping mechanism called CapFilt, which generates synthetic captions and filters noisy training pairs to improve data quality. Its successor, BLIP-2, took a radically efficient path: rather than retraining everything, it froze both a pretrained image encoder and a large language model, connecting them through a lightweight Querying Transformer that translates visual features into a form the language model can consume. DeepMind’s Flamingo achieved few-shot multimodal learning by threading visual inputs into a frozen language model through a Perceiver Resampler and gated cross-attention layers, allowing it to answer visual questions from just a handful of examples. Google’s PaLI unified multilingual text generation via mT5 with high-capacity ViT encoders, while LLaVA pioneered visual instruction tuning, using GPT-4 to synthesize training dialogues and then aligning a CLIP vision encoder with a Vicuna language decoder.
The review also charts the frontier models that have captured public attention. GPT-4V extended OpenAI’s flagship language model to accept images, achieving human-level performance on many professional benchmarks while remaining weaker than humans on certain real-world tasks. GPT-4o, released in May 2024, went further still, processing text, images, audio, and video in a single network and becoming the first large language model to perform real-time emotion recognition from video. Florence-2, built on a Dual Attention Vision Transformer and trained on the FLD-5B dataset with 5.4 billion annotations across 126 million images, handles captioning, detection, segmentation, and grounding through one prompt-based sequence-to-sequence framework. SigLIP-2 replaced the conventional softmax with a sigmoid loss, decoupling batch size from training efficacy and supporting batches of up to one million examples while maintaining strong performance in more than 100 languages.
Efficiency has emerged as the defining engineering challenge, and the review documents how the newest models answer it. PaliGemma 2 pairs a SigLIP vision encoder with Gemma language models in sizes from 3 billion to 28 billion parameters, trained in three stages that progressively raise image resolution. Gemma 3 interleaves a single global attention layer with every five local layers using a restricted 1024-token sliding window, taming the memory demands of its 128,000-token context. Llama 4 adopts a sparse Mixture-of-Experts design in which a routing mechanism activates only a fraction of expert networks per token, expanding effective capacity while keeping inference cheap, and stretches context to roughly 10 million tokens for document-scale reasoning. These techniques, alongside parameter-efficient fine-tuning methods such as Low-Rank Adaptation, form what the authors call a crucial pathway to deployable multimodal AI.
The application evidence assembled in the review is striking in its breadth. In healthcare, fine-tuned combinations of vision transformers and language models reach up to 98.5 percent accuracy on lung disease diagnosis, while the PaliGemma-CXR system interprets tuberculosis chest X-rays with 90.32 percent accuracy. ChatIOS, which couples 3D point-cloud encoders with GPT-4V, achieved 93 percent intersection-over-union in automatic tooth segmentation from intraoral scans. Yet the same section catalogues sobering failures: GPT-4V identified anatomy correctly in 87.1 percent of radiology cases but pathology in only 35.2 percent, and diagnostic hallucination rates exceeding 40 percent were reported in some evaluations. In transportation, a fine-tuned PaliGemma model read license plates with 97.66 percent character accuracy, while GPT-4o managed 67 percent on pedestrian behavior prediction. In manufacturing, a CLIP-based defect classifier hit 99.9 percent AUROC with 6.6-millisecond inference, fast enough for real-time production lines.
Against these successes, the review is candid about the field’s structural weaknesses. Data remains the first bottleneck: many studies rely on small, single-center, or single-vendor datasets, annotations from lone experts, and corpora so noisy that models learn spurious correlations, the authors’ memorable example being polar bears on ice. Computational cost is the second, with models scaling to 120 billion parameters and demanding specialized hardware that restricts access for smaller research groups. Reliability is the third: models hallucinate plausible but incorrect findings, struggle with composite figures and micro-expressions, misclassify small pedestrians at low pixel density, and remain vulnerable to adversarial attacks such as projected gradient descent. Ethical concerns compound the technical ones, since multimodal datasets encode societal biases that models can perpetuate in sensitive domains like healthcare and education, and training on medical images raises privacy stakes that demand techniques such as federated learning and differential privacy.
The authors close with a research agenda that reads as a diagnosis of the field’s growing pains. They call for multi-center, human-annotated datasets that pair imaging with genetic and clinical data; for encoder refinements that can handle long audio and raw high-resolution images directly; for larger but memory-efficient transformers; and for explainability to be treated as a core evaluation requirement rather than an afterthought. Techniques like model distillation, which compresses large multimodal systems into deployable smaller versions, and Mixture-of-Experts scaling are highlighted as the most promising routes to accessible systems. What emerges from the 120 studies is a field in transition: the architectural foundations are largely settled, the benchmarks are impressive, but trust, efficiency, and generalization to the messy diversity of real-world data remain the unfinished work. For researchers deciding where to invest the next five years, this review functions as both a map and a warning.
Subject of Research: Pretrained multimodal deep learning models, their architectures, applications, and future research directions
Article Title: A comprehensive review on recent pretrained multimodal deep learning models from architectures to future directions
Article References: Hadi, A. A., & Supreethi, K. P. (2026). A comprehensive review on recent pretrained multimodal deep learning models from architectures to future directions. Discover Informatics, 1(1), Article 21. https://doi.org/10.1007/s44564-026-00022-1
Image Credits: AI Generated
DOI: 10.1007/s44564-026-00022-1
Keywords: multimodal deep learning, CLIP, vision transformers, GPT-4o, BLIP, Flamingo, PaliGemma, Llama 4, vision-language models, healthcare AI, Mixture-of-Experts, model hallucination
Cite Scienmag News
APA
MLA
Chicago
Denise Maddox. (October 1, 2026). From CLIP to Llama 4: A sweeping review maps the rise of pretrained multimodal AI. Scienmag. https://scienmag.com/from-clip-to-llama-4-a-sweeping-review-maps-the-rise-of-pretrained-multimodal-ai/
Denise Maddox. “From CLIP to Llama 4: A sweeping review maps the rise of pretrained multimodal AI.” Scienmag, 1 October 2026, https://scienmag.com/from-clip-to-llama-4-a-sweeping-review-maps-the-rise-of-pretrained-multimodal-ai/. Accessed 1 October 2026.
Denise Maddox. “From CLIP to Llama 4: A sweeping review maps the rise of pretrained multimodal AI.” Scienmag. October 1, 2026. https://scienmag.com/from-clip-to-llama-4-a-sweeping-review-maps-the-rise-of-pretrained-multimodal-ai/
Copy citation
Download RIS
Tags: advancements from CLIP to Llama 4AI applications in wildlife monitoringAI in autonomous transportationAI in industrial manufacturingBLIPCLIPcross-modal data interpretationdeep learning models in healthcareFlamingoGPT-4ohealthcare AIinterdisciplinary AI model analysisLlama 4Mixture of Expertsmodel hallucinationmultimodal deep learningmultimodal model architecturesmultimodal training datasets and objectivesopen-access AI research reviewPaliGemmaPretrained multimodal AIsystematic review of multimodal modelsVision Transformersvision-language models



