ChatGPT-4 has passed one of its most demanding real-world tests yet: planning the safest route for a needle to reach a tumor deep inside a living lung. In a feasibility study published in CVIR Oncology, researchers at Memorial Sloan Kettering Cancer Center showed that a multimodal large language model could look at a CT scan, reason about the anatomy, and propose an entry angle for a lung biopsy that was not only workable in more than nine out of ten cases, but actually shorter and potentially gentler than the path chosen by experienced human specialists. The finding offers an early glimpse of how generative artificial intelligence might move from answering questions in a chat window to guiding steel needles through some of the most delicate anatomy in the human body.
Lung needle biopsies are among the most common and consequential procedures in interventional radiology. When a radiologist spots a suspicious nodule on a chest scan, the definitive diagnosis usually requires extracting a sliver of tissue through the chest wall, guided by computed tomography. The challenge is geometric as much as medical. The needle must reach the lesion while dodging ribs, avoiding the mediastinum and its great vessels, steering clear of pulmonary fissures, and traversing as little healthy lung parenchyma as possible. Every extra millimeter of lung crossed raises the risk of pneumothorax, the collapsed lung that occurs when air leaks through a puncture, as well as bleeding and hemoptysis. Yet despite decades of advances in imaging, the choice of trajectory still varies considerably from one radiologist to the next, a variability that AI and robotics have long been expected to tame.
The new study, led by Taylor Hoffman and Francois H. Cornelis with colleagues including Stephen B. Solomon and Debkumar Sarkar, set out to test whether a general-purpose generative AI model could shoulder part of that planning burden. The team conducted a retrospective, institutional review board-approved analysis of de-identified CT images from 30 consecutive lung biopsies performed by three interventional radiologists, each with at least a decade of experience. The cohort included 17 men and 13 women with a median age of 65 years, and the lesions ranged from 5 to 45 millimeters across, deliberately spanning technically difficult territory: five apical lesions at the top of the lung, seven subpleural nodules hugging the chest wall, three para-mediastinal lesions, and four near the lung base.
The technical setup was elegantly simple. Each axial CT image, acquired on a GE Healthcare scanner and exported in DICOM format, was cropped and magnified but not rescaled. Using the open-source image editor GIMP, the team overlaid a radial grid on the image and converted it to PNG. One author annotated the target lesion with a blue circle and outlined the sensitive structures in red, including the mediastinum and major vessels. The annotated image was then uploaded to ChatGPT-4 with a standardized prompt specifying five key parameters for an optimal entry angle: reaching the target, minimizing the distance traveled through lung tissue, and avoiding bones, fissures, and the mediastinum. The model was asked to read the radial grid and return the angle at which the needle should enter.
Before running the full analysis, the researchers validated the approach by comparing a pre-biopsy planning CT image with a procedural CT image captured during the actual biopsy, acquired at the same anatomical level and showing the needle, skin, and target in place. A separate validation dataset of 30 paired planning and procedural images was then used for data collection. Each case was processed independently in a fresh chat session, with iterations allowed whenever the model failed to incorporate all criteria. The median number of iterations needed to arrive at a complete answer was just two, with an interquartile range of one to three, suggesting the model could converge on a usable plan with minimal back-and-forth.
The results were striking. Independent reviewers, two experienced interventional radiologists who had not performed the original procedures and who assessed the AI-generated paths in blinded fashion, judged 28 of 30 paths, or 93.3 percent, safe and adequate for performing the biopsy. The two rejected paths were deemed theoretically feasible but not representative of best clinical practice: in one case the AI path crossed 20 millimeters of lung tissue, four times longer than the manual route, and in another it unnecessarily punctured lung parenchyma to reach a subpleural nodule, raising concerns about needle stability. More remarkably, the AI paths crossed a median of just 5 millimeters of lung tissue, compared with 23 millimeters for the paths the human radiologists actually chose, a difference that reached statistical significance with a p-value of 0.0063 on a Wilcoxon signed-rank test. Paths with zero lung traversal corresponded exclusively to subpleural nodules, where a direct pleural puncture enters the lesion immediately.
On the question of whether the AI simply rediscovered what the humans had already chosen, the answer was nuanced. The mean entry angle suggested by the AI was 179 degrees, against 174 degrees for the manual procedures, a difference that was not statistically significant. Using a definition of path match requiring entry angles within 10 degrees of each other and similar trajectories through tissue planes, 8 of 30 cases, or 26.7 percent, showed good concordance. The authors note that this match rate is lower than in earlier work, such as a study by Too and colleagues that used a three-dimensional convolutional neural network for lesion detection and path planning and reported 82 percent concordance with an average angular deviation of 2.30 degrees, alongside 93.5 percent sensitivity and 93.2 percent specificity in lesion detection. That earlier system, however, was purpose-built for the task rather than adapted from a general-purpose chatbot, and it defined path match more narrowly, within 5 degrees.
Why should a shorter path matter? The authors argue that reduced traversal through lung tissue could lessen tissue manipulation and lower the risk of pneumothorax and hemoptysis, potentially shortening procedure time, improving patient comfort, reducing anesthesia use, and increasing operating room efficiency. In the 30 procedures studied, no major complications occurred, and minor events were self-limiting: minimal pneumothorax in four cases, 13.3 percent, which resolved spontaneously, and mild hemoptysis in two cases, 6.7 percent. But the authors also flag a clinical caveat. With a coaxial technique, in which an outer introducer needle anchors the pathway for repeated sampling, a very short path through lung tissue might provide insufficient purchase, increasing the risk that the needle dislodges during the procedure. Shorter, in other words, is not automatically better in every scenario.
The study is candid about its limits. With only 30 cases and a retrospective design, generalizability remains unproven. The analysis was performed in two dimensions on single axial slices rather than in three dimensions, a simplification that may underestimate the true spatial complexity of biopsy planning. The prompt did not require the AI to maintain a standard 5-millimeter safety margin from the subclavian and axillary neurovascular bundles, though those structures were kept out of the imaging field because patients were scanned with their arms raised. Iterative follow-up prompts were not scripted, which limits reproducibility, and the study measured technical feasibility only, not clinical outcomes such as diagnostic yield, complication rates, or procedure time. The authors also caution that GPT-4 itself may soon be superseded, arguing that the field should move toward specialized generative models trained specifically on interventional radiology rather than adapting general-purpose systems.
Even so, the trajectory of the technology is hard to ignore. The authors envision generative AI integrated with robotic biopsy platforms to automate trajectory planning, real-time procedural guidance in which the model continuously analyzes intraoperative imaging and suggests adjustments as the patient moves or breathes, pre-procedural virtual rehearsals that let clinicians simulate multiple approaches before committing to one, and eventual three-dimensional reconstruction that would eliminate manual image annotation altogether. Combined with augmented reality, AI-recommended paths could be projected directly into the clinician’s field of view. The same logic extends beyond the lung to tumor ablation, where precise positioning determines whether the entire tumor is covered, and to vascular interventions, where AI could recommend optimal access routes based on patient-specific anatomy. For now, the message of this small feasibility study is measured but provocative: a model built to converse with humans can, with a radial grid and a well-crafted prompt, plan a needle path through the lung that is measurably less invasive than the one a seasoned specialist picked. The needle, for the moment, remains in human hands, but the map may increasingly be drawn by machines.
Subject of Research: Feasibility of using a multimodal large language model to plan entry angles for CT-guided lung needle biopsies
Article Title: Feasibility study on using multimodal large language model for CT-guided lung biopsy trajectory planning
Article References: Hoffman, T., Worthington, M., Sarkar, D., Solomon, S. B., & Cornelis, F. H. (2025). Feasibility study on using multimodal large language model for CT-guided lung biopsy trajectory planning. CVIR Oncology, 1(1), Article 5. https://doi.org/10.1007/s44343-025-00005-3
Image Credits: AI Generated
DOI: 10.1007/s44343-025-00005-3
Keywords: generative AI, ChatGPT-4, lung biopsy, CT-guided biopsy, trajectory planning, interventional radiology, pneumothorax, needle path planning, multimodal large language model, Memorial Sloan Kettering, CVIR Oncology, medical imaging
News Source: Nathaniel Bowman. (October 4, 2026). AI Plans Safer Needle Paths for Lung Biopsies, Feasibility Study Finds. Scienmag.



