• HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
Wednesday, October 7, 2026
BIOENGINEER.ORG
No Result
View All Result
  • Login
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
No Result
View All Result
Bioengineer.org
No Result
View All Result
Home NEWS Science News Technology

Small AI Model Learns When to Freeze Traffic Lights and Save Pedestrians

by
October 7, 2026
in Technology
Reading Time: 5 mins read
0
Small AI Model Learns When to Freeze Traffic Lights and Save Pedestrians

Small AI Model Learns When to Freeze Traffic Lights and Save Pedestrians

Share on FacebookShare on TwitterShare on LinkedinShare on RedditShare on Telegram

Every year, thousands of pedestrians are killed or injured at signalized intersections, often in fleeting moments that conventional traffic controllers never even register. A pedestrian hesitating at a curb, a car edging over the stop bar on a red light, a jaywalker stepping into a crosswalk while conflicting traffic holds a green signal: these are transient, spatially limited events that aggregate traffic measurements simply mask. Now, a research team at the New Jersey Institute of Technology has built and tested a lightweight artificial intelligence system that watches live video of an intersection, reasons about what it sees in plain language, and can order every signal to hold red for a few critical seconds before a collision unfolds. The study, published in Machine Learning with Applications, demonstrates the system in a high-fidelity simulator and reports a striking result: a small, fine-tuned vision-language model cut the time pedestrians spend exposed to conflicting green lights by roughly two-thirds to three-quarters under most tested conditions.

The core insight of the work is that the models that power modern chatbots, known as vision-language models, can do more than describe a scene. They can be taught to make bounded, auditable safety decisions. But the researchers found that off-the-shelf versions of these models fail at the task in a dangerous and instructive way. When general-purpose models were asked, without any task-specific training, to watch short video clips of an intersection and recommend an action, the larger seven-billion-parameter model intervened on nearly every clip it saw, including clips of completely normal traffic where the correct answer was to do nothing. Its apparent accuracy on hazardous events was an artifact of always sounding the alarm. The smaller three-billion-parameter model, meanwhile, produced outputs in the wrong format, so the signal controller could not execute them at all. For a safety layer, the decisive failure is not missing a hazard but crying wolf on every quiet street corner, eroding both traffic flow and trust.

To fix this, the team built a training dataset of 1,020 video clips generated inside the CARLA driving simulator, each roughly two to three seconds long and captured by two fixed virtual cameras whose views are stacked into a single composite image of the intersection. The clips cover nine combinations of weather and lighting, from clear noon to hard rain at night, and fall into three classes: jaywalking, red-light running, and normal traffic. Every clip was labeled automatically by a deterministic expert policy that maps each event and signal state to an ideal intervention. A jaywalker triggers a long all-red hold of eight to ten seconds, calibrated to the geometry of the crosswalks, which span about 14.5 to 15 meters; a red-light runner triggers a shorter hold of five seconds for cars and six for heavy vehicles, enough to clear the intersection before conflicting movements are released; normal traffic triggers no action at all.

One of the most technically interesting parts of the pipeline addresses a subtle data problem. The object detector that anchors the system, a YOLOv10-l model, performs well in daylight but nearly collapses at night and in heavy rain. In the first collection pass, the detector found around 40 jaywalking pedestrians per daytime condition but only six on a clear night and a single one under hard rain at night. Because the pedestrians were scripted actors whose positions the simulator knows exactly, the researchers projected ground-truth positions onto the camera images using pinhole camera geometry, recovering correct labels for 128 of the 348 jaywalking samples. Crucially, the projection changes only the labels, not the video itself, so the model still has to learn to recognize faint, barely visible hazards from raw pixels alone.

The fine-tuning itself is remarkably efficient. The researchers started with Qwen2.5-VL-3B-Instruct, a compact open model, and applied low-rank adaptation, a technique that freezes the original network and injects a small number of trainable update matrices into its attention and feed-forward layers. Only 37.15 million parameters, less than one percent of the model, were trained, and the vision encoder was left untouched. Training took about 40 minutes on a single workstation GPU, with loss decaying smoothly from 1.13 to 0.004 over three passes through the data. The model was trained to emit a strict JSON decision document: whether to intervene and why, how long to hold all-red, which legal signal combination to run, per-phase green times, safety notes, and a confidence level. Every field is range-checked by a separate rule-constrained controller, which rejects anything malformed, out of bounds, or inconsistent with traffic-engineering constraints. The AI never touches the signal directly.

The evaluation was deliberately harsh. Three of the nine weather conditions, wet sunset, clear night, and hard rain at night, were withheld entirely from training, and the fine-tuned model was tested on 313 clips from those unseen conditions. It achieved perfect schema validity, perfect adherence to controller bounds, and a perfect intervene-or-do-nothing decision on every single clip, with a mean hold-time error of just 0.15 seconds. Its false-positive rate on the 108 normal-traffic clips was zero, with a statistical upper bound of 3.4 per hundred. The zero-shot seven-billion-parameter baseline, by contrast, scored only 0.655 on decision accuracy, exactly the fraction of event clips in the test set, because it intervened on all of them, and its hold-time errors averaged over four seconds. Notably, a seven-billion-parameter model fine-tuned with the identical recipe performed no better than the three-billion one, and the smaller model was also faster, with a median decision time of 2.77 seconds versus 3.20 seconds for its larger zero-shot rival. Adaptation, not scale, is what makes these models usable inside a control loop.

The system also generalized beyond its training environment. At a second simulated intersection in a different virtual town, with different road geometry, surroundings, and camera positions, and with no retraining whatsoever, the fine-tuned model again made the correct decision on all 135 test clips. The only degradation was in hold timing: it recommended nine-second holds where the site-specific expert target was eight seconds, a one-second error that still fell within the mandated eight-to-ten-second band. Textual explanations produced at the new site referred to its actual crosswalks and approaches rather than echoing the training scene, suggesting the model carried across an imitated policy rather than a memory of one location’s appearance.

In closed-loop operation, where the full sensing, reasoning, and control stack ran live against scripted violation scenarios, the safety benefits were substantial but came with a measurable cost. Compared with blind fixed-time signal control, the ground-truth pedestrian exposure fraction, the share of total crosswalk-occupancy time overlapping a conflicting green, was 67.6 percent lower at low demand and 73.0 percent lower at medium demand, with every random seed favoring the AI layer. Against a stronger actuated pedestrian-clearance baseline, the reductions were 55.4 and 71.6 percent. An analysis showed the gains came not merely from serving less conflicting green but from shifting that green away from moments when pedestrians were actually present. The price, at medium demand, was a 105.9 percent increase in average stopped delay and a 13.8 percent throughput drop, though the researchers caution that the scripted scenarios inject hazardous events far more often than real intersections experience, inflating the intervention frequency and its delay burden by construction.

The study is equally candid about its limits, and the most consequential one lies upstream of the AI. At night, the object detector recovers only about 10.7 percent of ground-truth crosswalk occupancy from the same imagery, compared with 62.6 percent in daylight, and under clear-night conditions the closed-loop safety improvement shrank to a statistically insignificant 15.2 percent. A semantic layer cannot protect a pedestrian it cannot see. All results come from simulation, with safety-critical events injected by script rather than observed in natural traffic, so no claim of real-world transfer is made. Still, the authors argue the central conclusion stands: domain adaptation, not model scale, is what allows a vision-language model to operate reliably within a controller’s constraints, and the path to deployment now runs through weather-robust perception, sensor fusion, and shadow-mode validation on real traffic cameras. If that perception gap closes, the vision of intersections that watch, understand, and briefly freeze to protect the most vulnerable road users moves a significant step closer to reality.

Subject of Research: A lightweight vision-language AI safety layer for pedestrian protection at signalized intersections

Article Title: A Fine-Tuned Lightweight Vision-Language Safety Layer with Weather-Generalising Decisions at Signalized Intersections

Article References: Afshari, A., Lee, J., & Mashal, Y. (2026). A Fine-Tuned Lightweight Vision-Language Safety Layer with Weather-Generalising Decisions at Signalized Intersections. Machine Learning with Applications, 26, Article 101032. https://doi.org/10.1016/j.mlwa.2026.101032

Image Credits: AI Generated

DOI: 10.1016/j.mlwa.2026.101032

Keywords: vision-language models, traffic signal control, pedestrian safety, fine-tuning, LoRA, CARLA simulator, computer vision, intelligent transportation systems, weather generalization, all-red hold, machine learning, road safety

News Source: Blake Davidson. (October 7, 2026). Small AI Model Learns When to Freeze Traffic Lights and Save Pedestrians. Scienmag.

Tags: all-red holdCARLA simulatorComputer Visionfine-tuningintelligent transportation systemsLoRAMachine LearningPedestrian Safetyroad safetytraffic signal controlvision-language modelsweather generalization
Share12Tweet7Share2ShareShareShare1

Related Posts

AI Learns to Run Renewable Microgrids 326 Times Faster Than Traditional Solvers

AI Learns to Run Renewable Microgrids 326 Times Faster Than Traditional Solvers

October 7, 2026
Simple Models Win: New Framework Turns Churn Prediction Into Profitable Retention Decisions

Simple Models Win: New Framework Turns Churn Prediction Into Profitable Retention Decisions

October 7, 2026

When Algorithms and Humans Shape Each Other: Inside the New Science of Entanglement

October 7, 2026

Chaotic Printing Turns Simple Static Mixers Into Tools for Microarchitected Materials

October 7, 2026

POPULAR NEWS

  • Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

    Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

    29 shares
    Share 12 Tweet 7
  • Endurance Exercise Reshapes the Liver in Males and Females Through Distinct Molecular Routes

    29 shares
    Share 12 Tweet 7
  • Single Transcription Factor PU.1 Rapidly Converts Fibroblasts into Macrophage-Lineage Cells

    29 shares
    Share 12 Tweet 7
  • New Scale Measures How Ready Nurse Educators Really Are for the AI Era

    29 shares
    Share 12 Tweet 7

About

We bring you the latest biotechnology news from best research centers and universities around the world. Check our website.

Follow us

Recent News

Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

Endurance Exercise Reshapes the Liver in Males and Females Through Distinct Molecular Routes

Single Transcription Factor PU.1 Rapidly Converts Fibroblasts into Macrophage-Lineage Cells

Subscribe to Blog via Email

Success! An email was just sent to confirm your subscription. Please find the email now and click 'Confirm' to start subscribing.

Join 85 other subscribers
  • Contact Us

Bioengineer.org © Copyright 2023 All Rights Reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Homepages
    • Home Page 1
    • Home Page 2
  • News
  • National
  • Business
  • Health
  • Lifestyle
  • Science

Bioengineer.org © Copyright 2023 All Rights Reserved.