On a dairy farm in central Thailand, a handful of ordinary security cameras now do the work that once demanded round-the-clock human observation. Researchers at Chulalongkorn University have built a hierarchical computer vision framework that watches a free-stall barn continuously, identifying what each cow is doing—drinking, feeding, resting, standing, or walking—without a single wearable device or moment of handling. The study, published in Smart Agricultural Technology, demonstrates that a carefully layered pipeline of deep learning models can push through the visual chaos of a working barn: cubicle partitions, feeding rails, shifting daylight, and herds of nearly identical black-and-white animals crowding into overlapping piles.
The motivation is animal welfare. Subtle behavioral changes—less time at the feed bunk, more time lying down, fewer visits to the water trough—often signal declining health days before clinical symptoms appear. But capturing those shifts reliably has proven difficult. Manual observation is intermittent and labor-intensive, and wearable accelerometers, while useful, infer behavior indirectly, require battery maintenance, and offer no visual confirmation of how animals interact with barn infrastructure. The researchers argue that contactless, camera-based monitoring preserves natural behavior while providing individual-level surveillance, provided the underlying algorithms can survive real-world conditions rather than the tidy, unobstructed views common in laboratory studies.
The team’s solution is a three-stage hierarchy. First, object detection models scan each video frame to locate cows and classify their behavior. Second, a segmentation model refines those coarse rectangular boxes into pixel-precise outlines of each animal’s body. Third, a multi-object tracking algorithm stitches detections across frames, maintaining each cow’s identity over time. The researchers deliberately benchmarked competing architectures at every stage, comparing the convolutional YOLOv11m detector against the transformer-based RT-DETR, and testing whether segmentation-assisted tracking genuinely outperforms conventional detection-based tracking under obstruction—a question that had not been systematically answered before.
The data foundation was substantial. Over two consecutive 24-hour periods in August 2025, three fixed high-definition CCTV cameras mounted 4.5 meters above the barn floor recorded nine focal Holstein-Friesian crossbred cows from a herd of roughly 25. The cameras captured a gradient of visual difficulty: one monitored the feeding-resting area where partitions created the worst occlusion, one observed the partially obstructed front resting area, and one offered a wide-angle view with minimal obstruction. From 72 camera-hours of footage, the team built a detection dataset of more than 214,000 images carrying over one million behavior-labeled bounding boxes, plus a separate polygon-mask benchmark for evaluating segmentation under daytime and nighttime illumination and varying crowding levels.
Detection results revealed a striking pattern. On the held-out test set, both detectors achieved identical mean average precision—65.4 percent at the lenient IoU threshold of 0.50 and 50.9 percent averaged across stricter thresholds—despite their fundamentally different architectures. Feeding, resting, and standing were recognized reliably, with per-class AP50 values exceeding 95 percent, because these behaviors present distinctive whole-body postures in predictable locations. Walking fared poorly, with AP50 around 33 percent, and drinking collapsed to just 2 percent for both models. The reason is instructive: single-frame detection cannot access motion cues that distinguish a walking cow from a standing one, and drinking depends on subtle muzzle-to-water contact invisible in a whole-body bounding box. Confusion matrices confirmed that most walking instances were mislabeled as standing, and drinking instances vanished into the standing class or background.
The segmentation stage exposed a genuine architectural trade-off. When detector-generated boxes were fed as prompts to the Segment Anything Model 2 (SAM2), the RT-DETR-based pipeline delivered higher mean IoU and mask recall—capturing a larger share of visible cows, especially in crowded scenes—while the YOLOv11m-based pipeline achieved far higher mask precision, minimizing false-positive masks that could trigger spurious behavioral alerts. Notably, nighttime artificial illumination did not uniformly degrade performance; the YOLOv11m-SAM2 pipeline actually improved on most metrics at night, and the two pipelines’ nighttime instance F1-scores were statistically indistinguishable. As scene density increased from one or two visible cows to six or more, RT-DETR-L-SAM2 held its overlap metrics steady while YOLOv11m-SAM2’s declined, suggesting the transformer prompts were more robust under dense inter-animal overlap.
Tracking results told perhaps the most consequential story. ByteTrack, the conventional detection-based tracker, achieved the highest overall multi-object tracking accuracy—53.0 percent with RT-DETR-L and 49.5 percent with YOLOv11m—but paid for it with identity chaos, racking up 50 and 75 identity switches respectively. The SAM2-assisted pipelines, which propagate pixel-level masks across frames, recorded dramatically fewer identity switches—just 8 and 9 overall—while sacrificing some aggregate accuracy. In the most obstructed camera view, the contrast was stark: on daytime footage, RT-DETR-L-ByteTrack committed 40 identity switches against 5 for the SAM2-assisted pipeline. For welfare monitoring, where a cow’s behavioral time budget must be attributed to the correct individual over hours, that identity stability may matter more than raw frame-level accuracy.
Camera placement emerged as an underappreciated variable. The wide-angle, least-obstructed view produced near-perfect daytime tracking—MOTA of 98.2 to 99.9 percent for ByteTrack pipelines—while the most obstructed view dragged performance to negative values in some configurations, meaning errors exceeded the number of ground-truth cows. Yet the relationship was not simple: wider views pack more animals into each frame, increasing inter-animal occlusion even as structural obstruction falls. The authors caution that minimizing structural occlusion alone is an insufficient criterion for camera placement, and that occlusion, crowding, and viewing geometry interact in ways that demand site-specific evaluation before deployment recommendations can be made.
Computational realities temper the enthusiasm. YOLOv11m trained in roughly 33 hours and infers at over 320 frames per second on server hardware, while RT-DETR-L required nearly 147 hours of training with more than 60 percent more parameters for essentially identical detection accuracy. On a consumer laptop GPU, the SAM2 segmentation stage became the bottleneck, processing at well under one frame per second—far from real-time streaming on edge devices. The framework, the researchers stress, is a within-farm proof of concept: one barn, one breed, two days of footage, and no cross-camera identity linkage. Generalization across commercial facilities, extended recording periods, and dedicated edge hardware remains untested.
Even so, the study charts a credible path toward fully non-invasive welfare monitoring. The authors envision validating vision-derived metrics—sustained drops in feeding duration as an early flag for subclinical ketosis, rising lying time as a lameness indicator—against accelerometers, milk yield records, and veterinary logs. They also propose lightweight temporal modules, such as optical flow or compact recurrent units, to resolve the walking-versus-standing ambiguity that single-frame detection cannot. If those steps succeed, the humble barn security camera could evolve from a passive recorder into a continuous, individual-level health sentinel, catching trouble in the herd before a farmer ever notices a sick animal.
Subject of Research: Computer vision-based continuous monitoring of dairy cow behavior in free-stall barns
Article Title: A Hierarchical Computer Vision Framework for Continuous Dairy Cow Behavior Monitoring in an Obstructed Free-Stall Barns
Article References: Raza, A., Abbas, K., Jamil, M. U., Hogeveen, H., & Inchaisri, C. (2026). A Hierarchical Computer Vision Framework for Continuous Dairy Cow Behavior Monitoring in an Obstructed Free-Stall Barns. Smart Agricultural Technology, 15, Article 102622. https://doi.org/10.1016/j.atech.2026.102622
Image Credits: AI Generated
DOI: 10.1016/j.atech.2026.102622
Keywords: dairy cattle, computer vision, animal welfare, object detection, YOLO, RT-DETR, SAM2, instance segmentation, multi-object tracking, precision livestock farming, deep learning, behavior monitoring
News Source: Alan Morgan. (October 11, 2026). AI Watches the Herd: Vision System Tracks Every Cow in a Crowded Barn. Scienmag.



