• HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
Sunday, September 27, 2026
BIOENGINEER.ORG
No Result
View All Result
  • Login
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
No Result
View All Result
Bioengineer.org
No Result
View All Result
Home NEWS Science News Technology

Screens, Seconds and Faces: New Study Probes Why Test Scores Mislead

Bioengineer by Bioengineer
September 27, 2026
in Technology
Reading Time: 6 mins read
0
Screens, Seconds and Faces: New Study Probes Why Test Scores Mislead
Share on FacebookShare on TwitterShare on LinkedinShare on RedditShare on Telegram

A final score on a computer-based test tells you far less than it appears to. That is the central premise of a new study published in Mobile Networks and Applications, in which a team of Serbian and Croatian researchers set out to disentangle what actually drives performance during electronic knowledge assessment. Their conclusion is sobering for anyone who treats a percentage on a screen as a pure measure of learning: the number a candidate produces reflects not only what they know, but also how the test items are constructed, how those items are rendered on screen, how long the candidate spends wrestling with them, and even the fleeting expressions that cross their face in the first seconds of reading a question.

The research, led by Danilo Strugarevic of the Academy of Applied Preschool Teaching and Health Studies in Kruševac, Serbia, together with Vojkan Nikolic of the University of Criminal Investigation and Police Studies in Belgrade, Dragan Perakovic of the University of Zagreb and Aleksandar Jevremovic of Singidunum University, took an unusually broad view of the testing process. Rather than examining a single signal in isolation, the team combined several distinct streams of information: textual characteristics of the test items, entropy-based measures computed from screenshots of the rendered displays, response times, candidates’ self-reported knowledge, and facial-expression features captured during the exam. The goal was to see whether these process indicators, taken together, could support a more accurate and more honest interpretation of electronic test performance than scores alone.

One of the most technically interesting elements of the study is its use of Shannon entropy, the foundational information-theoretic quantity that measures the unpredictability or information content of a signal. In this context, entropy served as a compact descriptor of how informationally dense a rendered test item appeared on screen. The researchers computed entropy measures such as Entropy_AVG, Entropy_SUM and logEntropy from screenshots of the displays that candidates actually saw. The intuition is straightforward: a visually or textually cluttered item carries more information per unit of screen space, and that density may shape how long a candidate needs to process it. The data bore this out in part. Both Entropy_AVG and QA_Total_Length, a measure of question and answer text length, were positively associated with response time, with correlation coefficients of 0.298 and 0.288 respectively, both statistically significant at p < 0.001.

Yet the relationship between time and performance proved far more elusive. Response time was not significantly associated with the final test result at all, showing a correlation of just −0.029 with a p-value of 0.435. In other words, spending longer on a question neither helped nor hurt the score in any statistically detectable way. This finding cuts against a common intuition in assessment design, where slow responses are sometimes read as a sign of struggle or guessing and fast responses as a sign of confident mastery. The study suggests that response time is a genuine process indicator, sensitive to the structural demands of the item, but it is not a reliable proxy for knowledge.

When the team turned to prediction, the picture became more modest. They compared global and segmented regression models that attempted to predict test outcomes from the combined feature set, and found that predictive performance was limited, with no single algorithm emerging as superior across all evaluation measures. This is an important negative result. In an era when machine-learning vendors promise to read competence out of behavioural traces, the study offers a reminder that the mapping from interaction data to knowledge is noisy and algorithm-dependent. Entropy_SUM and logEntropy showed weak positive associations with the final result, hinting that items with higher informational content may relate in some small way to performance, while Question_Length stood out as the highest-ranked negative structural indicator, meaning longer questions were associated with lower results.

The facial-expression analysis produced perhaps the most nuanced findings of all. Using automated facial-expression recognition, the researchers extracted features corresponding to emotional categories over the course of the test. As standalone predictors, these features had limited value: classification agreement was low and regression correlations were near zero. But a subtler pattern emerged when the team examined timing. Features computed from the first five seconds of each item, covering neutral, disgust and anger expressions, were more informative than averages computed over the full response interval. This suggests that the initial emotional reaction to reading a question, the moment a candidate first registers its demands, may carry more signal than the emotional state sustained while answering. Intriguingly, this early-window advantage did not hold across all expression categories, indicating that different emotions unfold on different timescales during problem solving.

The study’s framing draws on a rich tradition in psychometrics and learning analytics. Classical test theory and item response theory have long recognised that item characteristics shape observed scores, and decades of research on guessing in multiple-choice and true-false formats have shown that partial knowledge and informed guessing can inflate results. What is new here is the attempt to bring information theory into the item-review pipeline. Prior work has used entropy-based algorithms to quantify textual comprehension difficulty and to cluster visual question difficulty, and the present study extends that logic to the actual pixels a candidate sees, treating the rendered display as a measurable object in its own right rather than a neutral window onto the question text.

The multimodal approach also connects to a growing body of work on emotion in learning. Pekrun’s control-value theory of achievement emotions has established that emotional states are not incidental to performance but causally intertwined with it, and earlier studies have monitored facial expressions to gauge learner attention in e-learning environments. The Serbian team’s contribution is to test whether such signals can be harvested unobtrusively, from a standard webcam during a real exam, and to report honestly on how much predictive weight they can actually bear. Their answer, based on the data, is: some, but not much on their own, and only when timed carefully.

Methodologically, the study was conducted at the Department of Health Studies in Ćuprija, Serbia, with ethical approval from the institutional Ethics Committee and written informed consent from all participants, including consent to be video-recorded during testing. The authors note that the underlying datasets, which include performance records, behavioural traces, and image and video material, cannot be released publicly because of privacy and ethical restrictions, though anonymised and processed data may be shared on reasonable request subject to institutional approvals. All four authors contributed equally to the work, with Strugarevic leading conceptualisation and methodology and the team sharing data collection, analysis and interpretation.

The practical implications reach toward the next generation of digital assessment platforms. The authors argue that jointly considering item-structure measures, screenshot-derived entropy, response time, self-reported knowledge and facial-expression features could support smarter testing environments that combine conventional scores with interpreted process indicators, helping item-quality review and potentially helping to distinguish knowledge-supported responses from those involving guessing. Crucially, they resist the temptation to crown any single indicator as definitive. No webcam, no timer and no entropy calculation can certify knowledge, detect guessing, or read an emotional state on its own. What the study demonstrates is that these signals are complementary pieces of a larger puzzle, and that the honest path to fairer electronic assessment runs not through any single magic metric but through the disciplined, statistically cautious fusion of many imperfect ones. In a world increasingly examined by algorithm, that caution may be the study’s most valuable result.

Subject of Research: Multimodal human-computer interaction analysis for reducing error in electronic knowledge assessment

Article Title: Analysis of Human-Computer Interaction Aimed at Reducing Error in Assessing Students’ Knowledge

Article References: Strugarevic, D., Nikolic, V., Perakovic, D., & Jevremovic, A. (2026). Analysis of Human-Computer Interaction Aimed at Reducing Error in Assessing Students’ Knowledge. Mobile Networks and Applications. https://doi.org/10.1007/s11036-026-02538-0

Image Credits: AI Generated

DOI: 10.1007/s11036-026-02538-0

Keywords: human-computer interaction, electronic testing, Shannon entropy, facial expression recognition, psychometrics, response time, learning analytics, multimodal assessment, machine learning, item quality, guessing detection, e-learning

Cite Scienmag News

APA
MLA
Chicago

Denise Maddox. (September 27, 2026). Screens, Seconds and Faces: New Study Probes Why Test Scores Mislead. Scienmag. https://scienmag.com/screens-seconds-and-faces-new-study-probes-why-test-scores-mislead/

Denise Maddox. “Screens, Seconds and Faces: New Study Probes Why Test Scores Mislead.” Scienmag, 27 September 2026, https://scienmag.com/screens-seconds-and-faces-new-study-probes-why-test-scores-mislead/. Accessed 27 September 2026.

Denise Maddox. “Screens, Seconds and Faces: New Study Probes Why Test Scores Mislead.” Scienmag. September 27, 2026. https://scienmag.com/screens-seconds-and-faces-new-study-probes-why-test-scores-mislead/

Copy citation
Download RIS

Tags: comprehensive approach to understanding test performanceComputer-based test performance analysise-learningeffect of candidate facial expressions on test resultselectronic testingevaluation of test question complexity and presentationfacial expression recognitionfactors affecting accuracy of electronic knowledge testsfactors influencing test scores beyond knowledgeguessing detectionhuman-computer interactionimpact of test item design on assessment outcomesimplications for digital assessment validityinfluence of test timing and interaction on scoresitem qualitylearning analyticslimitations of percentage scores in measuring learningMachine learningmulti-stream data analysis in educational assessmentmultimodal assessmentpsychometricsresponse timerole of screen rendering in electronic testingShannon entropy

Share12Tweet7Share2ShareShareShare1

Related Posts

Robot Therapies for Autistic Children Face a Hard Ethical Reckoning

Robot Therapies for Autistic Children Face a Hard Ethical Reckoning

September 27, 2026
How a Plant Hormone Switches On Immunity: New Clues From Salicylic Acid Receptors

How a Plant Hormone Switches On Immunity: New Clues From Salicylic Acid Receptors

September 27, 2026

AI Model Watches Bitcoin’s Underworld as Illicit Transactions Evolve Into Hidden Networks

September 27, 2026

Gut Infections Nearly Double the Odds of Childhood Stunting, Global Meta-Analysis Finds

September 27, 2026

POPULAR NEWS

  • Trunk Control May Hold a Key to Balance and Fall Risk in Older Adults

    29 shares
    Share 12 Tweet 7
  • AI Tells Doctors When It Is Unsure About Cancer Treatment Success

    29 shares
    Share 12 Tweet 7
  • Engineered mini CRISPR enzyme gets a 60-fold power boost for gene editing

    29 shares
    Share 12 Tweet 7
  • When Autism Diagnoses Fade: Early Intervention and Milder Symptoms Mark Children Who Lose the Label

    29 shares
    Share 12 Tweet 7

About

We bring you the latest biotechnology news from best research centers and universities around the world. Check our website.

Follow us

Recent News

Trunk Control May Hold a Key to Balance and Fall Risk in Older Adults

AI Tells Doctors When It Is Unsure About Cancer Treatment Success

Engineered mini CRISPR enzyme gets a 60-fold power boost for gene editing

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 85 other subscribers
  • Contact Us

Bioengineer.org © Copyright 2023 All Rights Reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Homepages
    • Home Page 1
    • Home Page 2
  • News
  • National
  • Business
  • Health
  • Lifestyle
  • Science

Bioengineer.org © Copyright 2023 All Rights Reserved.