For more than a century, quantitative genetics has rested on a quiet but powerful assumption: that the individuals being compared belong to a single population, or that the statistical relationships among them are the same across populations. From Sir Ronald Fisher’s landmark 1918 paper on the correlation between relatives to the modern machinery of genomic prediction, covariances between relatives, between traits, and between environments have been the workhorses of plant and animal improvement. Yet breeders rarely work with a single, homogeneous population. Maize lines fall into distinct heterotic groups, rice into the indica and japonica subspecies, beef cattle into breeds as divergent as Angus and Texas Longhorn, and coffee into separate species. A new study published in Heredity by InĂ©s Rebollo and Rex Bernardo of the University of Minnesota argues that the field has lacked something fundamental: a rigorous way to define and estimate the genetic covariance between two different populations. Their paper delivers exactly that, along with a validation exercise spanning computer simulations and real-world commercial maize data.
The concept of covariance measures how two random variables vary together, and it is the statistical glue that connects relatives, traits, and environments in classical theory. Within a single population, the covariance between relatives arises because they share alleles that are identical by descent. But when an individual comes from one biparental cross, say an indica rice cross A/B, and another from a completely different cross, such as a japonica cross C/D, there is no established framework for asking how much their genotypic values co-vary. Rebollo and Bernardo propose an elegant answer: shared marker alleles between individuals belonging to two different populations can serve as the basis for defining and estimating a between-population covariance. In other words, even when two populations share no recent ancestry, the molecular markers they have in common provide a measurable bridge for quantifying joint genetic variation.
The analytical core of the method extends earlier formulas developed by Osthushenrich and colleagues in 2017, which estimated the genetic variance within a recombinant inbred line population from marker effects. The authors define the genotypic value of an individual as a random variable whose variance can be decomposed across chromosomes and loci. Under a purely additive genetic model, the heterozygote effect is zero, and the probability that a locus carries the alternative or reference homozygote depends on allele frequencies and the inbreeding coefficient F, which ranges from zero in an F2 population to one in fully inbred doubled haploids. Crucially, the covariance between loci depends on the linkage disequilibrium coefficient D, which in turn is a function of the effective recombination rate between markers. Working through the algebra with allele frequencies of one half, the authors arrive at a strikingly compact result: the genotypic variance equals one plus F multiplied by a sum over chromosome pairs of twice the linkage disequilibrium coefficient times the product of the two marker effects.
This compact formula carries a subtle but important feature: linkage phase is captured by the signs of the marker effects themselves, with same signs indicating coupling and opposite signs indicating repulsion, rather than being embedded in the disequilibrium coefficient as in the classical formulation of Lynch and Walsh. The authors verified that their expression reduces to the classical result when F equals zero, and they checked the algebra with small-scale numerical examples. From the variance formula, the covariance between two populations follows from a simple identity: the variance of the sum of two populations’ genotypic values equals the sum of their individual variances plus twice their covariance. By computing the variance of the combined population using summed marker effects, and subtracting the two individual variances, the between-population covariance drops out analytically.
One practical obstacle is that the method nominally requires a linkage map to compute recombination rates between markers. Maps can be inaccurate, based on the wrong population, inflated by generations of random mating, or simply unavailable when only physical positions are known. The authors provide a workaround: for fully inbred populations and for F2 populations, the disequilibrium coefficient can be estimated directly from observed marker and haplotype frequencies as the absolute difference between the observed frequency of double-minor-allele homozygous haplotypes and the product of the individual allele frequencies. When phasing information is unavailable and the population is not fully inbred, published approximations from Ragsdale and Gravel can supply an unbiased estimate. This flexibility means the framework can be applied even when map resources are poor, a meaningful advantage for orphan crops and under-resourced breeding programs.
To validate the approach, the authors simulated four nested populations of one thousand individuals each, built so that each successive population segregated for only half as many quantitative trait loci as the one before. Populations were generated at three inbreeding levels, F2, F3, and doubled haploid, using the AlphaSimR simulation package with maize-like genome parameters: ten chromosomes of two Morgans each, carrying twenty quantitative trait loci apiece. Across one thousand independent simulation repeats, the analytical estimates tracked the true variances and covariances with remarkable fidelity. The correlation between observed and estimated variance exceeded 0.8 in every repeat, with a mean of 0.99, whether the known linkage map or the estimated disequilibrium coefficients were used. As theory predicted, variances and covariances in F2 populations were half those of doubled haploids, and F3 populations fell at seventy-five percent, confirming that the framework behaves correctly across the full range of inbreeding.
The real-world test came from a commercial maize breeding program at AgReliant Genetics. Four biparental populations of doubled haploid lines, each crossed to two of four tester inbreds, were evaluated in multi-environment field trials across four to ten locations in the U.S. Corn Belt over three years. Five traits were recorded: grain yield, grain moisture at harvest, plant height, ear height, and test weight. Marker effects for each population were estimated separately using ridge regression best linear unbiased prediction, and the authors discovered that the choice of shrinkage factor mattered enormously. When shrinkage was based on the raw number of markers, analytical variance estimates were consistently biased downward; when based on the number of fifty-centimorgan chromosome segments, they were biased upward. Using the effective number of independent tests, calculated from the eigenvalues of the marker correlation matrix following Li and Ji, yielded near-unbiased estimates that correlated strongly, above 0.97, with classical restricted maximum likelihood estimates.
The estimated covariances between maize populations revealed trait-specific and population-specific structure. For grain yield, covariances between population pairs ranged from negative values up to 9.52 square tonnes per hectare, with a mean of 1.67. The correlations between marker effects across populations were generally low to moderate, averaging 0.10 for yield, and the highest covariances appeared, unsurprisingly, between the same population tested with different testers. The lowest covariances occurred between populations sharing no parents. Notably, the correlation between the estimated genetic covariance and the marker-effect correlation was significant for most traits, reaching 0.77 for test weight, underscoring that between-population covariance reflects both relatedness and the genetic architecture of each trait. The authors also showed that their approach scales gracefully: where a joint multitrait analysis of fifty populations with five thousand markers would require estimating 250,000 marker effects and 1,275 covariance parameters simultaneously, their method decomposes the problem into fifty independent analyses.
The implications stretch well beyond the maize field. In breeding, heterogeneous genetic covariances could be folded into prediction models to improve accuracy when training and target populations differ, a chronic challenge in genomic selection. In quantitative trait locus mapping, the framework may allow genetically related populations to be pooled into larger discovery panels, boosting the power and resolution of gene detection. Beyond agriculture, the authors point to ecology and evolutionary genetics, where quantifying genetic covariance between subpopulations could illuminate hybrid zones, inform conservation of genetic diversity, and sharpen studies of population structure and adaptation. The method does carry limitations: it assumes a uniform inbreeding level within each population, and it requires marker effects, which are population-specific and depend on phenotypic data that may not yet exist for untested or uncreated populations.
What makes the study conceptually significant is that it fills a gap that has persisted since the founding of quantitative genetics. Covariance within populations has been understood since Fisher; covariance between traits and environments since the 1930s and 1950s. But the covariance between populations, the quantity that governs how much genetic information flows across breeding program boundaries, subspecies divides, and heterotic group walls, had no formal definition or estimator. By grounding between-population covariance in shared marker alleles and validating it against both simulated truth and commercial field data, Rebollo and Bernardo have given geneticists a new measurement tool that is analytically tractable, computationally scalable, and applicable at any level of inbreeding. As genomic datasets continue to accumulate across crops, livestock, and wild species, a framework for measuring how populations co-vary genetically may prove as foundational as the one Fisher built for measuring how relatives do.
Subject of Research: Estimation of genetic covariance between plant biparental populations using molecular marker effects
Article Title: Genetic covariance between plant biparental populations: concept and estimation
Article References: Rebollo, I., & Bernardo, R. (2026). Genetic covariance between plant biparental populations: concept and estimation. Heredity. https://doi.org/10.1038/s41437-026-00887-w
Image Credits: AI Generated
DOI: 10.1038/s41437-026-00887-w
Keywords: genetic covariance, quantitative genetics, plant breeding, biparental populations, marker effects, linkage disequilibrium, maize, genomic prediction, inbreeding coefficient, doubled haploids, RR-BLUP, Heredity
News Source: Juliet Wilcox. (October 8, 2026). New Framework Measures Genetic Covariance Between Plant Breeding Populations. Scienmag.



