• HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
Thursday, October 8, 2026
BIOENGINEER.ORG
No Result
View All Result
  • Login
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
No Result
View All Result
Bioengineer.org
No Result
View All Result
Home NEWS Science News Technology

AI Agents Match Human Averages but Diverge in How They Respond to Rewards and Penalties

by
October 8, 2026
in Technology
Reading Time: 5 mins read
0
AI Agents Match Human Averages but Diverge in How They Respond to Rewards and Penalties

AI Agents Match Human Averages but Diverge in How They Respond to Rewards and Penalties

Share on FacebookShare on TwitterShare on LinkedinShare on RedditShare on Telegram

Large language models are increasingly being drafted as stand-ins for human subjects, populating virtual societies and simulating policy responses that would be too slow, costly, or impractical to test on real people. But a new study published in iScience delivers a cautionary finding for this fast-growing practice: when humans and AI agents face the same sequential reward-and-penalty incentives, they can land on nearly identical average outcomes while getting there in strikingly different ways. The research, led by Sihan Tao, Linghao Wang, Zheng Zhu, and Der-Horng Lee, suggests that matching aggregate behavior is not enough to certify an AI as a faithful proxy for human decision-making.

The team built their experiment around a real-world policy problem: greener last-mile delivery in e-commerce logistics. Online retail has made urban delivery a significant source of carbon emissions, and platforms such as Amazon have experimented with one-shot incentives like vouchers to nudge shoppers toward slower, lower-emission shipping. The researchers instead tested a credit-charge-cum-reward mechanism, in which participants accumulate carbon credits across repeated decisions and face a periodic settlement. Choosing standard, faster delivery deducts credits; choosing slow delivery earns them. Positive balances convert into cash rewards, negative balances trigger penalties, and the account resets after each 20-day round. The design forces an intertemporal trade-off between immediate speed and long-term credit, much like the cumulative incentive structures found in real environmental policy.

To anchor the comparison, the researchers first ran a baseline scenario with no incentives at all. Participants, whether human or artificial, simply chose between standard delivery, which took three days but could stretch under simulated congestion, and slow delivery, fixed at seven days. Each participant was assigned a value of time, a concept borrowed from transportation engineering that quantifies how much money a person will pay to save a unit of time, set at high, medium, or low levels. Theory predicted that everyone should pick the fast option, and that is precisely what happened. Humans chose standard delivery 98.9 percent of the time, DeepSeek-v.3.2 agents 99.7 percent, and Qwen3-Max agents 98.1 percent, with no statistically detectable differences among the three groups after correction for multiple comparisons.

Then came the carbon credit mechanism, and with it the study’s central test. The theoretical equilibrium under the new rules called for an overall standard delivery proportion of 51.39 percent, stratified sharply by value of time: 100 percent for high-value participants, 54.17 percent for the medium group, and zero for the low group. In aggregate, all three participant types responded dramatically and almost identically. Humans cut their standard delivery proportion from 0.989 to 0.463, DeepSeek-v.3.2 from 0.997 to 0.513, and Qwen3-Max from 0.981 to 0.450. The adjusted effect sizes ranged from roughly 48 to 53 percentage points, none of the pairwise differences between participant types was statistically significant, and all pooled estimates were compatible with the theoretical benchmark. On the surface, the AI agents looked like perfect miniature humans.

The illusion dissolved once the researchers stratified the results by value of time. Humans retained a stubborn residual of standard delivery choices even in the low-value group, where theory predicted none, holding at 9.2 percent. DeepSeek-v.3.2 produced a much smaller residual of 2.3 percent, while every Qwen3-Max system converged to exactly zero. At the high end, DeepSeek-v.3.2 hugged the theoretical ceiling most closely, whereas humans showed the largest deviation. Dispersion patterns diverged too: human systems remained heterogeneous across all strata, DeepSeek-v.3.2 systems clustered tightly around the interior benchmark in the medium group, and Qwen3-Max systems collapsed to the boundary in the low group. The same average masked three different behavioral signatures.

Longitudinal analysis over 120 virtual decision days exposed further fault lines. Humans adjusted gradually and sometimes incompletely, with their best response rate, a measure of how often a choice minimized one-step generalized cost given available information, improving significantly only in the high-value group. DeepSeek-v.3.2 stayed near its theoretical benchmarks from the outset and maintained the highest overall best response rate at 0.960, compared with 0.666 for humans and 0.679 for Qwen3-Max. Qwen3-Max displayed delayed adjustment, with its standard delivery probability in the medium-value group climbing from 0.018 in the first round to 0.493 in the sixth, a convergence that ultimately exceeded that of the other participants but followed a qualitatively different path. Because the language models generate responses without parameter updating, these trajectories reflect in-context conditioning and stochastic generation rather than the experience-based learning that shapes human behavior.

Conditional choice persistence revealed perhaps the strangest divergence. In the high-value group, whether an agent had chosen standard delivery on the previous purchase day strongly predicted its current choice for both AI models, with adjusted probability increases of 0.218 for DeepSeek-v.3.2 and 0.405 for Qwen3-Max, while humans showed essentially no such carryover at 0.010. In the medium-value group the pattern inverted, with DeepSeek-v.3.2 showing a strong negative persistence of −0.471 while humans and Qwen3-Max showed none. In other words, the AI agents’ decisions were structured by their recent histories in ways that no aggregate statistic would ever reveal.

An exploratory clustering analysis added a final layer of evidence. Using a Gaussian mixture model over allocation frequencies on one-item and two-item purchase days, the researchers identified four strategy modes: standard dominant, slow dominant, mixed allocation, and variable allocation. Human portraits spread across multiple modes within every value-of-time group, reflecting genuine individual diversity. DeepSeek-v.3.2 followed a clean ordered transition from standard dominant at high value of time, to mixed allocation at medium, to slow dominant at low. Qwen3-Max instead concentrated overwhelmingly in the variable allocation mode at high and medium value of time before shifting entirely to slow dominant at low. Notably, when decision-state features such as credit balances were added to the clustering, most of Qwen3-Max’s variable-mode assignments dissolved, indicating that frequency-based portraits can compress distinct state-conditioned behaviors into a single misleading category.

The authors are careful about scope. Only two models were tested, no prompt or parameter sensitivity analysis was performed, and the human sample consisted primarily of Zhejiang University students, so the findings speak to the specific configurations examined rather than to language models in general. The experimental environment also simplifies real logistics, omitting promotional discounts, product heterogeneity, and seasonal demand spikes. Still, the pattern is consistent with a growing body of work showing that surface similarity between human and machine outputs can conceal deep differences in the strategies and cognitive structures that produce them.

The implications reach well beyond delivery apps. As generative agents become a scalable substitute for human experiments in policy simulation, this study provides a concrete methodological warning: validating an AI proxy by its aggregate response alone can certify a simulator that gets the subgroup patterns, the temporal dynamics, and the underlying choice architecture wrong. For policymakers designing cumulative incentive mechanisms, whether carbon credits, congestion charges, or energy-use rewards, the practical effectiveness of a policy depends not just on average uptake but on how individuals perceive, learn, and adapt over time. The study suggests LLM simulations can reliably reproduce the direction of an aggregate policy response under well-defined rules, but researchers and regulators should demand disaggregate evidence, stratified trajectories, and strategy-level portraits before trusting an artificial crowd to stand in for a real one.

Subject of Research: Comparing human and large language model decision-making under sequential reward-penalty incentives in e-commerce delivery choices

Article Title: Similar aggregate but divergent disaggregate responses in human and AI under sequential reward-penalty incentives

Article References: Tao, S., Wang, L., Zhu, Z., & Lee, D.-H. (2026). Similar aggregate but divergent disaggregate responses in human and AI under sequential reward-penalty incentives. iScience, 29(11), Article 117800. https://doi.org/10.1016/j.isci.2026.117800

Image Credits: AI Generated

DOI: 10.1016/j.isci.2026.117800

Keywords: large language models, behavioral simulation, carbon credits, e-commerce logistics, decision-making, incentive design, generative agents, value of time, policy simulation, last-mile delivery, behavioral economics, DeepSeek

News Source: Denise Maddox. (October 8, 2026). AI Agents Match Human Averages but Diverge in How They Respond to Rewards and Penalties. Scienmag.

Tags: behavioral economicsbehavioral simulationcarbon creditsDecision-MakingDeepSeeke-commerce logisticsgenerative agentsincentive designLarge Language Modelslast-mile deliverypolicy simulationvalue of time
Share12Tweet7Share2ShareShareShare1

Related Posts

Privacy-First AI Learns to Spot DDoS Attacks Across IoT Networks Without Sharing Raw Data

Privacy-First AI Learns to Spot DDoS Attacks Across IoT Networks Without Sharing Raw Data

October 8, 2026
Molybdenum Boosts Catalyst That Scrubs Two Pollutants at Once and Shrugs Off Poisoning

Molybdenum Boosts Catalyst That Scrubs Two Pollutants at Once and Shrugs Off Poisoning

October 8, 2026

New LoRA-Based Method Steers Language Models Toward Specific Human Values

October 8, 2026

Contrastive Learning Tames Out-of-Distribution Actions in Offline Reinforcement Learning

October 8, 2026

POPULAR NEWS

  • Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

    Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

    29 shares
    Share 12 Tweet 7
  • Endurance Exercise Reshapes the Liver in Males and Females Through Distinct Molecular Routes

    29 shares
    Share 12 Tweet 7
  • Single Transcription Factor PU.1 Rapidly Converts Fibroblasts into Macrophage-Lineage Cells

    29 shares
    Share 12 Tweet 7
  • New Scale Measures How Ready Nurse Educators Really Are for the AI Era

    29 shares
    Share 12 Tweet 7

About

We bring you the latest biotechnology news from best research centers and universities around the world. Check our website.

Follow us

Recent News

Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

Endurance Exercise Reshapes the Liver in Males and Females Through Distinct Molecular Routes

Single Transcription Factor PU.1 Rapidly Converts Fibroblasts into Macrophage-Lineage Cells

Subscribe to Blog via Email

Success! An email was just sent to confirm your subscription. Please find the email now and click 'Confirm' to start subscribing.

Join 85 other subscribers
  • Contact Us

Bioengineer.org © Copyright 2023 All Rights Reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Homepages
    • Home Page 1
    • Home Page 2
  • News
  • National
  • Business
  • Health
  • Lifestyle
  • Science

Bioengineer.org © Copyright 2023 All Rights Reserved.