What if a computer could read decades of animal welfare science and distil it into a single, defensible number that tells us how good — or how bad — a particular housing system really is for the animals inside it? That is the promise of semantic modelling, a formalised method that converts written scientific statements into weighted welfare scores. Now, researchers at the Friedrich-Loeffler-Institute in Germany, together with a colleague at Wageningen University & Research, have unveiled ANyWEL — short for ANimal WELfare assessment of anY farm animal — a generalised framework designed to work for any livestock species and any production direction, published in the journal Archives Animal Breeding.
Semantic modelling is not a meta-analysis in the conventional statistical sense. It does not demand effect sizes, confidence intervals, or the numerical precision that many studies fail to report. Instead, it systematically extracts statements from the peer-reviewed literature that describe how housing conditions affect welfare, decomposes them according to their biological meaning, and converts them into model variables. A housing system is described by its welfare-relevant attributes — floor type, water provision, space allowance, enrichment material — each of which has attribute levels, such as drinking nipples versus open water bowls. Scientific evidence then ranks those levels from best to worst and assigns them scores, producing an overall welfare score on a scale from 0 to 1.
The intellectual lineage of the approach runs back to the SOWEL model for pregnant sows, developed in the early 2000s, and through a family of successors: FOWEL for laying hens, COWEL for dairy cows, SWIM for Atlantic salmon, RICHPIG for enrichment materials, and PIGTAIL for tail-biting risk. Each of these models, however, was built for a single species or production direction. ANyWEL breaks that constraint. It is deliberately species-agnostic, meaning a single set of modelling rules can generate welfare assessments for pigs, poultry, cattle, fish, or entirely new farmed species, provided the scientific knowledge base has been modelled.
To make the method transparent, the team illustrated it with a deliberately fictitious creature called anYmal, an imaginary species farmed under three systems: intensive, semi-intensive, and extensive. The three systems differ only in group size, space allowance, and enrichment material — chains in the intensive system, balls in the semi-intensive and extensive ones. The example is engineered to deliver a shock: when the postulated science about anYmals is fed into the model, the extensive system turns out to be the worst for welfare, scoring just 0.08, while the intensive system scores 0.84 and the semi-intensive system 0.58. The lesson is that evidence, not intuition, drives the assessment — a deliberate demonstration that semantic modelling can overrule the common assumption that extensive systems must always be better.
The machinery behind that counter-intuitive result is worth unpacking. The model rests on welfare needs — for anYmal, just three: social contact, exploration, and movement. Scientific statements are decomposed into comparisons between attribute levels, and each welfare measure cited in a statement is classified into one of twelve weighting categories, or WCats, spanning ethology, physiology, and veterinary science. Nine are negative — pain, illness, survival, fitness, activation of the HPA axis, sympathetic adrenal medullary responses, aggression, abnormal behaviour, and frustration and avoidance — while three are positive: natural behaviour, preference, and demand. Each category carries its own scoring range, with high-impact categories such as demand and pain assigned scores up to ±5, and lower-impact ones capped at ±3.
From these weighted comparisons, the model computes attribute level scores between 0 and 1, proportional to each level’s welfare rank, and weighting factors that capture how important one attribute is relative to the others. In the anYmal example, the enrichment attribute — with rope as the best level, ball intermediate, and chain worst — earns a weighting factor of 5.2, meaning the best available enrichment counts 5.2 times in the final score. Group size, driven by fictitious findings of heavy fighting and elevated mortality in small groups, dominates the model entirely. The overall welfare score for a housing system is then simply the weighted average of the attribute level scores that apply to it, divided by the sum of all weighting factors.
ANyWEL also refines the arithmetic in ways its developers argue make assessments more accurate. In the older SOWEL rules, each additional distinct type of welfare measure within a category added a flat 0.2 to the weighting factor, regardless of how large or small the underlying effect was. ANyWEL instead scales that bonus by the actual weighting-category level score, so a serious finding such as life-threatening illness contributes more than a marginal one. The scale for frustration and avoidance has also been widened from −1 to −5, mirroring the positive scale for demand, on the reasoning that an animal’s motivation to avoid a frightening or painful stimulus can be just as powerful as its motivation to seek a reward.
Perhaps the most pragmatic innovation is how ANyWEL handles knowledge gaps. Innovative housing systems often lack published studies, and previous models quietly filled the void with modeller judgement. The new framework formalises this: expert opinions can enter the model at reduced scores of ±0.10, ±0.25, or ±0.40, so that roughly ten expert opinions carry the weight of one statistically significant experimental finding. Statements reporting only statistical tendencies (P<0.10) can be included at ±0.5, which the authors say may also soften publication bias. And when a proper study later supersedes an expert assessment, a score of zero retires the opinion from the calculations without deleting it, preserving a transparent audit trail.
The authors are candid about limitations. Subjectivity has not been eliminated — modellers still decide how to classify and score statements, and no welfare ‘thermometer’ exists to validate the output directly. Knowledge gaps in the literature can distort weighting factors, and the model’s output should be cross-checked against on-farm indicators such as those in the Welfare Quality protocol, judgement bias tests, or sensitivity analyses. Yet the framework’s transparency is its defence: every decision is documented, errors can be identified and corrected, and alternative calculations can be run to settle disputes. Notably, sensitivity analyses of earlier models suggest that scores calculated with and without weighting factors are highly similar, hinting that disputes over individual scores matter less than identifying the right attributes in the first place.
The implications reach well beyond a single paper. ANyWEL was developed within the German InKalkTier project to assess existing housing systems, and its statement database has already fed an artificial-intelligence tool that extracts welfare statements from literature automatically — a step toward models that update themselves as science advances. The framework could extend to zoo, laboratory, and companion animals, to different life stages within a species, and even to broader sustainability assessments that treat welfare as one component among environmental and economic ones. For a society increasingly demanding welfare-friendly food systems, a method that makes the science behind a welfare score fully visible — and sometimes uncomfortably counter-intuitive — may prove to be one of the more consequential tools in the barn.
Subject of Research: Semantic modelling framework for calculating overall animal welfare scores in farm husbandry systems
Article Title: Semantic modelling of animal welfare explained – Part 1: Calculating overall welfare scores for husbandry systems using the ANyWEL model framework
Article References: Benthin, J., Kauselmann, K., Krause, E. T., Bracke, M. B. M., & Vonholdt-Wenker, M. L. (2026). Semantic modelling of animal welfare explained – Part 1: Calculating overall welfare scores for husbandry systems using the ANyWEL model framework. Archives Animal Breeding, 69(3), 363-382. https://doi.org/10.5194/aab-69-363-2026
Image Credits: AI Generated
Keywords: animal welfare, semantic modelling, ANyWEL, husbandry systems, farm animals, welfare assessment, housing systems, weighting factors, SOWEL, evidence-based assessment, livestock, artificial intelligence
News Source: Alan Morgan. (October 9, 2026). New ANyWEL Model Turns Scientific Literature Into Animal Welfare Scores for Any Farm Species. Scienmag.



