Deep beneath the hills of central Armenia, in sedimentary rocks laid down during the late Devonian period some 375 million years ago, microscopic fossil spores lie scattered through ancient shale. For decades, identifying these tiny capsules has been the painstaking work of trained palynologists, who peer through light microscopes and judge each specimen by its shape, symmetry, and surface ornamentation. Now a team of researchers from Armenia, New Zealand, Belgium, and France has shown that artificial intelligence can take over a large share of that labor, detecting and classifying Devonian miospores with an accuracy that rivals expert assessment while dramatically cutting the time each slide demands.
The study, published in the Journal of Micropalaeontology, is, according to its authors, the first attempt to apply deep learning to the automated identification of Devonian miospores. The team, led by Vitalina Lokteva and corresponding author Vahram Serobyan of the Institute of Geological Sciences of the National Academy of Sciences of Armenia, together with Martin Tetard, Pierre Breuer, and Taniel Danelian, focused on three species of enormous biostratigraphic value: Teichertospora torquata, Geminospora lemurata, and Samarisporites triangulatus. These spores are widely used to define global biozones across the ancient continents of Gondwana and Laurussia, and they play a particularly important role in hydrocarbon exploration, notably in Saudi Arabia, where Devonian spore-based stratigraphy helps geologists date and correlate the rock layers that host oil and gas reservoirs.
The material came from the Ertych section in central Armenia, one of the region’s best-known localities for palynological research. Fourteen beds of upper Frasnian age were sampled during several field campaigns, and productive miospore assemblages were recovered from twelve of them. Standard palynological preparation yielded three to five slides per sample, which were examined with a Leica DM2700 transmitted-light microscope. Every specimen that a specialist qualitatively identified as one of the three target species was photographed individually with a Flexacam C5 camera, producing a library of 209 microphotographs that formed the raw material for the machine learning pipeline.
Each of those 209 images was then annotated by hand using the web-based platform Roboflow. An operator drew a rectangular bounding box around every individual spore and recorded the image identifier, class label, and microscope magnification as metadata. The resulting detection dataset contained five classes: the three target species, which together accounted for roughly 57 percent of the annotations, an “other” class comprising non-target miospore taxa, and a background class dominated by images without miospores, filled largely with fragments of plant debris known as phytoclasts. The images were divided into training, validation, and test subsets in a 70/20/10 split, with the test set kept strictly separate so that no image used in evaluation had ever been seen during training.
The architecture of the workflow reflects a growing consensus in automated microscopy: detection first, classification second. In the detection stage, the team deployed YOLOv11, a one-stage object detection network that predicts bounding-box coordinates, objectness scores, and class probabilities in a single forward pass. Initialized with weights pretrained on the COCO image dataset and trained with a single class label of “miospore”, the detector was applied to the test micrographs using a confidence threshold of 0.7. Its performance was striking. Precision, the fraction of predicted boxes that were correct, reached 0.98; recall, the fraction of true objects the model actually found, was a perfect 1.0; and the mean average precision at an intersection-over-union threshold of 0.5 came in at 0.995, averaging 0.906 across the stricter range of thresholds from 0.5 to 0.95.
Crucially, the detection stage did more than simply find the spores. For each detection, the pipeline stored the bounding-box coordinates, a cropped image patch stripped of its cluttered background, and the area of the crop, measured in pixels and calibrated to square micrometers. This inference step produced 159 background-free spore crops, which became the classification dataset, split 70/15/15 into training, validation, and test subsets. The bounding-box area was retained as an auxiliary numeric feature, a simple piece of morphometric information that would later prove surprisingly valuable when taxa differed mainly in size.
For the classification stage, the researchers compared three convolutional neural networks, all implemented in PyTorch and trained through transfer learning, in which the pretrained convolutional feature extractors were frozen and only the task-specific heads were fine-tuned on the miospore data. VGG16, a 16-layer network that stacks small 3×3 convolutions, served as a stable, well-understood baseline. ResNet-18, an 18-layer network with residual skip connections that prevent optimization problems in deeper models, offered a lightweight and efficient feature extractor. EfficientNet-B0, the smallest member of the EfficientNet family, provided a compact architecture suited to limited data and compute. During training, images were resized to 224 by 224 pixels and augmented on the fly with flips, rotations of up to 15 degrees, color jitter, normalization, and random erasing, while validation images underwent only resizing and normalization to keep the evaluation unbiased.
The results, averaged over ten independent training runs, showed a clear hierarchy. ResNet-18 achieved the highest mean test accuracy at 87.9 percent, along with a Matthews correlation coefficient of 0.85, a metric that condenses the entire multiclass confusion matrix into a single value between minus one and one and is robust to the class imbalance typical of palynological datasets. VGG16 posted the highest macro-F1 score of 87.3 percent with a comparable test accuracy of 87.5 percent, while EfficientNet-B0 trailed on every metric, with 79.6 percent accuracy, 76.4 percent macro-F1, and an MCC of 0.75. In the best single runs, both ResNet-18 and VGG16 reached a test accuracy of 95.8 percent. The confusion matrices revealed that all three classifiers identified T. torquata flawlessly, with an F1 score of 1.00, likely because that species is large, reaching 85 to 140 micrometers, and bears a distinctive cavate structure, sub-triangular outline, and clear trilete mark. G. lemurata, a cavate spore of roughly 40 to 50 micrometers with a rounded outline, scored between 0.88 and 0.92. S. triangulatus proved hardest, with F1 scores of 0.73 to 0.80, probably because poorer preservation in the Armenian material obscures its diagnostic cone-shaped sculpture and indistinct trilete mark.
The authors are candid about the limits of the proof of concept. The classification dataset of 159 crops is small by deep learning standards, and the heterogeneous “other” class, which lumps together many different taxa, inevitably adds within-class variability. All images were acquired with a single microscope and a single preparation protocol, so performance may shift when the method meets material prepared and imaged under different conditions. Even so, the pipeline’s average per-class accuracies of 84.5 to 86.25 percent fall comfortably within the range of 83.3 to 99.8 percent reported in recent palynological and microfossil classification studies, a notable achievement for fossil material, which is consistently harder to classify than the pristine modern pollen on which most earlier models were trained. Prior work has shown that accuracy drops when models trained on modern pollen are applied to damaged or fossil specimens, making the team’s direct focus on fossil Devonian spores a meaningful step forward.
The researchers are careful to position the tool as an assistant rather than a replacement for expert knowledge. Supervised AI classification remains dependent on expert-defined taxonomic concepts and carefully curated training data, and the workflow is best understood as a screening and decision-support system that standardizes routine identifications, reduces observer bias, and improves reproducibility. The path forward is clear: enlarging the image dataset, broadening taxonomic coverage, incorporating material from additional sections and laboratories, and coupling the detector with automated slide-scanning systems to raise throughput. The code and annotated dataset have been released on Zenodo, inviting the community to build on the work. If the approach scales as hoped, the same two-step logic could extend to other fossil groups, including Paleozoic marine phytoplankton such as acritarchs and prasinophytes, and even scolecodonts, bringing the routine dating of sedimentary rocks, from academic biostratigraphy to oil exploration, into the age of machine perception.
Subject of Research: Automated detection and identification of Devonian miospores using convolutional neural networks
Article Title: Artificial intelligence applied to the automated detection and identification of Devonian miospores
Article References: Lokteva, V., Serobyan, V., Tetard, M., Breuer, P., & Danelian, T. (2026). Artificial intelligence applied to the automated detection and identification of Devonian miospores. Journal of Micropalaeontology, 45(1), 405-413. https://doi.org/10.5194/jm-45-405-2026
Image Credits: AI Generated
Keywords: artificial intelligence, convolutional neural networks, Devonian miospores, palynology, biostratigraphy, YOLOv11, ResNet-18, microfossils, upper Frasnian, Armenia, fossil spores, automated microscopy
News Source: Blake Davidson. (October 9, 2026). AI Learns to Spot 380-Million-Year-Old Fossil Spores on Microscope Slides. Scienmag.



