• HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
Tuesday, August 25, 2026
BIOENGINEER.ORG
No Result
View All Result
  • Login
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
No Result
View All Result
Bioengineer.org
No Result
View All Result
Home NEWS Science News Technology

Generative AI’s STEM Learning Impact and Limits Revealed by Systematic Review

Bioengineer by Bioengineer
August 25, 2026
in Technology
Reading Time: 6 mins read
0
Generative AI’s STEM Learning Impact and Limits Revealed by Systematic Review
Share on FacebookShare on TwitterShare on LinkedinShare on RedditShare on Telegram

Generative artificial intelligence may be transforming STEM education—but not in the simple, universally positive way suggested by many early headlines. A new systematic review and meta-analysis of research published since the arrival of modern generative AI finds that tools such as ChatGPT can improve externally assessed cognitive learning outcomes in science, technology, engineering and mathematics. Yet the apparent overall benefit largely disappears when possible publication bias is taken into account. What remains is a more complicated picture: AI seems most promising when it strengthens students’ own thinking, particularly when the goal is building knowledge, while replacing essential learning activities can produce weak or even negative results.

The study, published in Artificial Intelligence Review, examined peer-reviewed quantitative research involving a comparison or control group. The researchers searched ERIC, PsycINFO and the Web of Science Core Collection, then expanded their results through forward and backward citation tracking. Their search identified 85 eligible studies, most of them conducted in higher education and involving text-based generative AI systems. Of these, 49 studies provided 59 effect sizes suitable for statistical meta-analysis. The analysis focused specifically on cognitive outcomes—what students learned and how well they performed on externally assessed measures—rather than satisfaction, usability or attitudes toward AI.

At first glance, the results appeared encouraging. A conventional random-effects meta-analysis indicated an overall positive effect of generative AI on STEM learning. In this model, the effect was expressed as Hedges’ g, a standardized measure designed to compare results across studies using different tests and scales. However, the studies differed dramatically from one another. Statistical heterogeneity was extremely high, with an I² value of 96.32 percent. In practical terms, this means that most of the variation in reported effects was not simply random sampling error; it reflected major differences among studies, including the learners, tasks, AI systems, teaching designs and outcome measures. The prediction interval ranged from g = −1.52 to g = 3.20, indicating that a future comparable study could plausibly find a substantial harm, little effect or a very large benefit.

The researchers also found signs that publication bias may be influencing the literature. Publication bias occurs when studies with striking positive findings are more likely to be published, noticed or included than studies showing no effect or negative outcomes. Funnel-plot asymmetry, a common diagnostic in meta-analysis, suggested that the available evidence may not represent the full distribution of results. To test how sensitive the headline finding was to this problem, the team used Robust Bayesian Meta-Analysis, or RoBMA. This approach compares models that allow for publication bias with models that do not, while estimating the underlying effect and uncertainty.

The RoBMA results substantially changed the interpretation. After accounting for the possibility that positive studies were overrepresented, the overall effect was estimated at approximately μ = 0.076 with a standard deviation of 0.254—a result close to zero and highly uncertain. The authors conclude that the positive average effect found in the conventional model can largely be attributed to publication bias. Importantly, however, the disappearance of a reliable overall effect does not mean that generative AI is ineffective in every educational setting. The analysis still revealed very large residual heterogeneity, with an estimated τ of approximately 1.190. In other words, the key scientific question is not whether AI “works” in general, but under what conditions, for which learners and through which kinds of learning activity it works.

Two factors explained part of this variation: the type of learning outcome and the way AI changed the cognitive activity required from students. The first distinction was between knowledge and skills. Knowledge outcomes included conceptual understanding and related forms of learning, while skills involved abilities such as problem-solving, procedures or performance. The second factor was captured by the ISAR framework, which compares the cognitive activity performed by students in an AI-supported intervention with that performed by students in the control condition. The framework draws on the ICAP model, which ranks learning activities from passive to active, constructive and interactive. In simplified terms, AI can substitute for a student’s activity, augment it, or enable a redefinition of the task that changes what students do cognitively.

The moderator analyses suggested that knowledge-focused interventions generally produced larger effects than skill-focused interventions. The meta-regression estimated that studies targeting skills had an effect approximately 0.73 standard deviations lower than studies targeting knowledge, relative to the reference category. ISAR level also mattered: substitution was associated with an estimated effect about 1.05 standard deviations higher than the reference category of redefinition in the statistical model, a result the authors interpret cautiously because many of the largest apparent effects arose when the intervention and control groups were not performing comparable cognitive activities. The moderators explained only part of the variance—about 12.5 percent and 13 percent in separate analyses, and 18.66 percent in a combined meta-regression—leaving most of the heterogeneity unexplained.

That warning is central to the study. The review identified 33 studies reporting large effects, but the authors found that these effects often reflected unequal learning conditions rather than a clean test of AI’s added value. In some cases, students using AI were allowed to engage in a substantially different activity from those in the control group. Comparing an AI-assisted learner who receives automated explanations, solution suggestions or generated examples with a control learner who receives none of these resources may measure the advantage of additional support, not the unique educational value of generative AI. Without carefully matched tasks, time, feedback and instructional support, a large effect can be statistically real while still being difficult to interpret.

A further analysis reinforced the limits of simple conclusions. Among 59 effect sizes, 23 were classified as large, meaning Hedges’ g exceeded 0.6. Nineteen of those large effects came from knowledge-oriented studies, while four came from skill-oriented studies. Large effects appeared most frequently when AI use was categorized at the redefinition level, especially in knowledge-focused settings. Yet no combination of outcome type and ISAR level met the researchers’ threshold for a sufficient condition that reliably produced a large effect. Knowledge was an almost, but not strictly, necessary condition: 19 of the 23 large effects involved knowledge outcomes, yielding a necessity consistency of about 0.83, but only 19 of 41 knowledge-focused studies produced large effects. The pattern points to a tendency, not a guarantee.

The review also examined learner challenges and instructional interventions, although the evidence was too inconsistent to support pooled quantitative estimates. Across the literature, researchers frequently failed to report variables that could determine whether AI helps or harms learning, including students’ AI literacy, metacognitive skills, prompt quality, verification behavior and the extent to which learners delegated tasks to the system. These omissions matter because generative AI can produce fluent but inaccurate explanations, incomplete reasoning or fabricated references. A student who treats an answer as authoritative may achieve short-term task completion without developing durable understanding. By contrast, a student who critiques, checks and revises AI output may use the same system as a cognitive partner rather than a replacement for thinking.

The authors propose six testable hypotheses and an integrative framework for future research, urging investigators to measure the learning process rather than focusing only on final scores. Stronger studies should compare equivalent activities, preregister outcomes, include negative and null findings, document the exact AI system and prompts used, and examine how students verify or challenge generated content. They should also distinguish between immediate performance and lasting learning, because AI may improve a student’s answer during an intervention without improving what the student can independently recall or apply later. For educators, the emerging message is neither an unqualified endorsement nor a rejection of generative AI: use it to augment students’ cognitive work, make verification visible and target clearly defined learning goals. Generative AI may be a powerful tool for STEM education, but the evidence suggests that its benefits depend less on the technology itself than on whether it preserves—and deepens—the intellectual work students must do.

Subject of Research: Generative artificial intelligence and cognitive learning outcomes in STEM education

Article Title: Evidence of impact and interpretational limits of generative AI in STEM education: a systematic review and meta-analysis on cognitive learning outcomes

Article References: Boolzen, C., Kuhn, J., Flegr, S. et al. “Evidence of impact and interpretational limits of generative AI in STEM education: a systematic review and meta-analysis on cognitive learning outcomes.” Artificial Intelligence Review (2026).

Image Credits: AI Generated

DOI: 10.1007/s10462-026-11665-9

Keywords: Generative artificial intelligence; STEM education; learning outcomes; educational technology

Tags: AI’s role in enhancing student knowledge constructionCritical analysis of AI’s educational benefits and risksEffectiveness of generative AI in STEM assessmentsGenerative AI in STEM educationImpact of ChatGPT on cognitive learning outcomesLimitations of AI replacing traditional learning activitiesMeta-analysis of AI tools in science and engineeringPublication bias in AI education researchQuantitative research on AIResearch methodologies in AI educational studiesSystematic review of AI in higher educationText-based AI systems in university settings

Share12Tweet7Share2ShareShareShare1

Related Posts

Deep Learning-Enhanced Bimodal Sensor Enables Intelligent Recognition and Navigation

Deep Learning-Enhanced Bimodal Sensor Enables Intelligent Recognition and Navigation

August 25, 2026
Flux-driven ligand exchange reshapes metal–organic framework glasses

Flux-driven ligand exchange reshapes metal–organic framework glasses

August 25, 2026

Oral Nanodelivery of Gut Microbial Metabolite Boosts T-Cell Stemness in Cancer Immunotherapy

August 25, 2026

In-place dry repair and passivation deliver size-independent performance in III-nitride micro-LEDs

August 25, 2026

POPULAR NEWS

  • Copper Catalyst Directly Converts Aliphatic Primary Amines into Nitriles Through Deaminative Cyanation

    29 shares
    Share 12 Tweet 7
  • Balanced mRNA and ribosome levels govern growth in eukaryotic cells

    29 shares
    Share 12 Tweet 7
  • Non-native plants’ mycorrhizal strategies shift across biomes and disturbance levels

    29 shares
    Share 12 Tweet 7
  • Japanese Colorectal Cancer Reveals Prevalence, Timing, and Microbiome Signatures of Colibactin Mutations

    29 shares
    Share 12 Tweet 7

About

We bring you the latest biotechnology news from best research centers and universities around the world. Check our website.

Follow us

Recent News

Copper Catalyst Directly Converts Aliphatic Primary Amines into Nitriles Through Deaminative Cyanation

Balanced mRNA and ribosome levels govern growth in eukaryotic cells

Non-native plants’ mycorrhizal strategies shift across biomes and disturbance levels

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 85 other subscribers
  • Contact Us

Bioengineer.org © Copyright 2023 All Rights Reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Homepages
    • Home Page 1
    • Home Page 2
  • News
  • National
  • Business
  • Health
  • Lifestyle
  • Science

Bioengineer.org © Copyright 2023 All Rights Reserved.