• HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
Tuesday, October 6, 2026
BIOENGINEER.ORG
No Result
View All Result
  • Login
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
No Result
View All Result
Bioengineer.org
No Result
View All Result
Home NEWS Science News Technology

New AI Pipeline Makes Searching Hours of Surveillance Video as Easy as Typing a Sentence

by
October 6, 2026
in Technology
Reading Time: 5 mins read
0
New AI Pipeline Makes Searching Hours of Surveillance Video as Easy as Typing a Sentence

New AI Pipeline Makes Searching Hours of Surveillance Video as Easy as Typing a Sentence

Share on FacebookShare on TwitterShare on LinkedinShare on RedditShare on Telegram

Every second, more than six hours of video land on YouTube, while city-wide camera networks churn out petabytes of footage each day. Buried inside those archives are people: suspects to locate, missing persons to find, customers to track across stores. The problem is that nobody has time to watch it all. A team of researchers at the University of Cagliari and the data engineering firm AgileLab has now built a system that promises to change that, letting users find any individual in a mountain of video simply by typing a description like “a young man with glasses and an olive-green parka” or by uploading a single snapshot. The work, published in Multimedia Tools and Applications, reports state-of-the-art results on a standard person-retrieval benchmark while keeping the whole architecture modular enough to scale to industrial workloads.

The core insight of the new study is that the bottleneck in video search is no longer the accuracy of individual AI models but the lack of a coherent system that stitches them together. Deep multimodal models such as CLIP and its successors have made it possible to align images and text in a shared mathematical space, so that a sentence and a photograph of the same scene end up close together as high-dimensional vectors. Yet, as the authors argue, most research focuses on squeezing extra accuracy out of a single model while ignoring the messy engineering questions: how to chunk hours of raw footage, how to avoid indexing the same person thousands of times, how to keep track of who appeared when and where, and how to serve queries in milliseconds. Their answer is a two-part pipeline that cleanly separates offline indexing from online retrieval.

On the indexing side, the system ingests raw video files and first splits them into roughly ten-minute chunks using FFmpeg, preserving the original encoding so each segment can be processed independently and in parallel. A frame-sampling module then skips redundant frames, since typical surveillance footage recorded at 30 or 60 frames per second offers far more temporal detail than person search actually needs. The sampled frames flow into YOLO11, the latest generation of Ultralytics’ real-time object detector, chosen for its favorable speed-accuracy trade-off, paired with the BoT-SORT tracker, which combines motion cues with appearance features to assign each detected person a persistent identity across frames. On the MOT17 benchmark, BoT-SORT reports 80.5 MOTA and 80.2 IDF1, meaning it rarely confuses one pedestrian with another, a property the authors consider essential for generating temporally coherent metadata.

Detection alone would still flood the database with near-duplicate images, so the pipeline adds a quality-filtering stage that is among the paper’s most practical contributions. Every cropped person image receives a score combining the detector’s confidence with an edge penalty: bounding boxes that touch the frame border, indicating a partially visible subject, are down-weighted. A perceptual hashing step then compares consecutive crops within the same tracked trajectory, discarding any whose hash differs by fewer than a threshold number of bits, a signal that the image is essentially identical to the previous one. What survives is a compact, diverse gallery of high-quality person crops rather than a bloated collection of blurry, truncated, or repetitive snapshots.

Each surviving crop is then transformed into a semantic fingerprint by SigLIP2, a multilingual vision-language encoder that extends the original SigLIP architecture with a sigmoid-based training loss, self-supervised objectives, and careful data curation. Because the model was trained on paired images and text, its visual embeddings live in the same latent space as its text embeddings, which is precisely what allows a typed sentence to retrieve a photograph. The system also computes a second, complementary embedding for each person using InsightFace: the RetinaFace detector localizes the face and its landmarks even under challenging poses and lighting, and the ArcFace recognizer converts the aligned face patch into a highly discriminative identity vector. Both embeddings, along with rich metadata such as source video, timestamps, track identity, and bounding-box coordinates, are stored in a Qdrant vector database, so every search result can be traced back to the exact frame in the original footage.

At query time, the retrieval pipeline mirrors the indexing representation. A user can submit an image, a natural-language description, or both; the same SigLIP2 encoders embed the inputs, guaranteeing that queries and database entries are directly comparable. When both modalities are present, a tunable balance parameter blends the image and text vectors, letting users slide between purely visual similarity and purely linguistic matching. A face-based mode restricts the search to identity embeddings when a face is detectable, and a negative-prompt field excludes unwanted attributes. The back-end runs an approximate nearest-neighbor search with cosine similarity, then consolidates results so that only the single best match per tracked object per video is returned, each accompanied by the retrieved frame, a similarity score, and metadata including camera identifier, first-appearance timestamp, and duration in the scene.

The quantitative centerpiece of the paper is an evaluation on RSTPReid, a benchmark of 20,505 pedestrian images covering 4,101 individuals captured by 15 cameras, each image paired with two text descriptions. The researchers fine-tuned the SigLIP2-SO400M-Patch14-384 checkpoint on a curated mixture of person-centric image-text datasets, including CUHK-PEDES, ICFG-PEDES, RSTPReid itself, an attribute-rich multiview re-identification set, and the video-based TVPReid corpus, totaling roughly 316,000 training pairs. Training used the native contrastive loss on four NVIDIA A100 GPUs on the Cineca Leonardo supercomputer, with early stopping based on validation mean Average Precision. The result: a mean Average Precision of 0.68, comfortably ahead of the best competing value of 0.54 among recent text-based person search methods such as OCDL, UP-Person, RDE, CTGI, WoRA, MARS, MRA, and CONQUER, while Recall@1 remained competitive within one standard deviation of the leaders.

Notably, the gains did not come from inventing new architectures. When the model was fine-tuned only on the three standard datasets used by the comparison methods, it already achieved the best mAP, indicating that the SigLIP2 backbone itself provides a stronger shared embedding space for matching descriptions to pedestrian attributes. Adding the larger, more diverse corpus then improved all metrics further. Ablation experiments on a subset of the Wildtrack multi-camera surveillance dataset underscored the value of the system’s design choices: removing face embeddings dropped Recall@1 from about 51 percent to 33 percent, while replacing quality-scored crops with uniformly sampled ones collapsed mAP from 0.66 to 0.39, showing that both identity cues and careful crop selection matter substantially.

The scalability analysis offers a sober, engineering-minded picture. Processing 128 ten-minute videos, about 256 gigabytes of footage, took roughly 90 minutes on a four-GPU server, with tracking saturating the CPU and encoding bottlenecked by memory traffic and model replication across worker processes rather than raw computation. Extrapolating linearly, the authors estimate the current setup could index about a terabyte of video in six hours, though they are candid that GPU utilization during encoding remains low and that a centralized inference server would be needed for truly efficient large deployments. They also acknowledge limitations: the system is person-centric, does not yet model complex temporal activities or events, and the auxiliary training data included AI-generated descriptions that were only spot-checked rather than fully validated by humans.

Even with those caveats, the study reads as a template for how modern AI systems should be built: not as isolated models chasing leaderboard numbers, but as orchestrated pipelines where detection, tracking, filtering, embedding, metadata, and search reinforce one another. The authors envision extending the framework beyond people to vehicles and other objects, adding temporal and multi-view reasoning, and eventually integrating large language models and knowledge graphs for conversational, explainable video search. If surveillance archives, streaming platforms, and personal devices keep growing at their current pace, systems that turn unwatchable oceans of footage into a simple search box may soon become as unremarkable, and as indispensable, as the web search bar itself.

Subject of Research: A scalable semantic pipeline for video indexing and text-based person retrieval using multimodal vision-language embeddings

Article Title: A scalable and semantic pipeline for efficient video indexing and person retrieval

Article References: Hmaidan, R., Milardo, S., Donato, I., Greco, D., Ingargiola, A., & Reforgiato Recupero, D. (2026). A scalable and semantic pipeline for efficient video indexing and person retrieval. Multimedia Tools and Applications, 85(9), Article 727. https://doi.org/10.1007/s11042-026-21882-7

Image Credits: AI Generated

DOI: 10.1007/s11042-026-21882-7

Keywords: video retrieval, person re-identification, multimodal embeddings, SigLIP2, YOLO11, BoT-SORT, face recognition, vector database, surveillance video, computer vision, text-based person search, RSTPReid

News Source: Blake Davidson. (October 6, 2026). New AI Pipeline Makes Searching Hours of Surveillance Video as Easy as Typing a Sentence. Scienmag.

Tags: BoT-SORTComputer Visionface recognitionmultimodal embeddingsperson re-identificationRSTPReidSigLIP2surveillance videotext-based person searchvector databasevideo retrievalYOLO11
Share12Tweet7Share2ShareShareShare1

Related Posts

Bubbles Reshape the Hydraulic Jump: CFD Reveals a Threshold Where Air Transforms Dam Spillway Flows

Bubbles Reshape the Hydraulic Jump: CFD Reveals a Threshold Where Air Transforms Dam Spillway Flows

October 6, 2026
One-Pot Mixed-Phase Indium Sulfide Outperforms Single-Phase Catalysts in Dye Breakdown

One-Pot Mixed-Phase Indium Sulfide Outperforms Single-Phase Catalysts in Dye Breakdown

October 6, 2026

South Korea Shows Why East Asia’s AI Ethics Problem Is Structural, Not Confucian

October 6, 2026

Hidden Rhythms in Paralyzed Gait Revealed by New Spatiotemporal Analysis

October 6, 2026

POPULAR NEWS

  • Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

    Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

    29 shares
    Share 12 Tweet 7
  • Endurance Exercise Reshapes the Liver in Males and Females Through Distinct Molecular Routes

    29 shares
    Share 12 Tweet 7
  • Single Transcription Factor PU.1 Rapidly Converts Fibroblasts into Macrophage-Lineage Cells

    29 shares
    Share 12 Tweet 7
  • New Scale Measures How Ready Nurse Educators Really Are for the AI Era

    29 shares
    Share 12 Tweet 7

About

We bring you the latest biotechnology news from best research centers and universities around the world. Check our website.

Follow us

Recent News

Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

Endurance Exercise Reshapes the Liver in Males and Females Through Distinct Molecular Routes

Single Transcription Factor PU.1 Rapidly Converts Fibroblasts into Macrophage-Lineage Cells

Subscribe to Blog via Email

Success! An email was just sent to confirm your subscription. Please find the email now and click 'Confirm' to start subscribing.

Join 85 other subscribers
  • Contact Us

Bioengineer.org © Copyright 2023 All Rights Reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Homepages
    • Home Page 1
    • Home Page 2
  • News
  • National
  • Business
  • Health
  • Lifestyle
  • Science

Bioengineer.org © Copyright 2023 All Rights Reserved.