Plant height may seem like a simple trait, but in soybean it sits at the heart of one of agriculture’s most consequential trade-offs. Tall plants can shade weeds and capture light, yet they topple in wind and rain, complicate harvest, and force farmers to sow at lower densities. Short, sturdy plants tolerate crowding and machine harvesting, but the genetic switches that set stem length have remained only partly mapped. A new study published in BMC Plant Biology by Hongmiao Jin, Shanshan Huang, Rui Ding and colleagues at Zhejiang A&F University and Beidahuang KenFeng Seed Co. has now combined two powerful genomic techniques to narrow the search for the genes that control how tall a soybean plant grows, delivering a prioritized shortlist of eight candidate genes that breeders and molecular geneticists can begin to interrogate in earnest.
The team’s strategy rested on a clever experimental design: two soybean varieties that were genetically nearly identical except for their height. Because the lines differed so little across the rest of the genome, any consistent molecular differences between them were likely to be connected to the height difference itself. This near-isogenic setup dramatically reduces the background noise that usually plagues genetic studies of complex traits, where hundreds of genes influencing flowering time, seed composition, disease resistance, and dozens of other characteristics all vary simultaneously and muddy the statistical signal.
To find the genomic neighborhoods harboring height-related variants, the researchers turned to bulked segregant analysis sequencing, or BSA-seq. In this approach, DNA from individuals at the extremes of a trait distribution, here tall versus short plants, is pooled and sequenced, and the allele frequencies of the two pools are compared across the whole genome. Regions genuinely linked to the trait show up as stretches where the pools differ sharply. The team applied three complementary statistical measures to detect these regions: the G′ value, the ΔSNP-index, and the Euclidean distance, or ED. Requiring convergence among independent metrics is a safeguard against statistical artifacts, since each method has its own sensitivities and failure modes.
The analysis converged on two candidate regions, one spanning 1.2 million bases on chromosome 10 and another spanning 3.3 million bases on chromosome 18. Together these intervals contained 108 genes, a number far too large for convenient functional testing but small enough to filter with additional evidence. Crucially, within these regions the team found three genes carrying missense mutations, DNA changes that alter the amino acid sequence of the encoded proteins and therefore have a plausible route to functional impact. The affected genes were GmWRI1, GmEXPA23, and GmPPR. The identification of GmEXPA23 is particularly intriguing from a mechanistic standpoint: expansins are proteins that loosen plant cell walls, and cell wall loosening in internode tissues is a direct physical determinant of how much stems elongate.
Sequencing where the genome differs is only half the story, however. A variant can exist without being switched on in the tissues that matter. To capture the regulatory dimension of the trait, the researchers performed RNA-seq, transcriptome sequencing that measures how actively every gene is expressed. Comparing the tall and short varieties, they identified 5,019 differentially expressed genes. Functional enrichment analysis showed that these genes clustered in plant hormone signaling pathways, MAPK signaling cascades, and a range of metabolic processes. Hormone signaling is a biologically coherent hit for a height trait, since gibberellin, auxin, and brassinosteroid pathways are the canonical regulators of stem elongation in flowering plants, and MAPK pathways frequently act as signaling relays connecting hormonal and environmental cues to growth responses.
The real power of the study came from integrating the two datasets. Genes that both sat inside the BSA-seq candidate regions and showed differential expression between the tall and short lines represented the strongest candidates, because they satisfied both a positional and a transcriptional criterion. Adding tissue-specific expression analysis, which asks whether a gene is active in the stems and other organs where height is actually determined, allowed the team to nominate five further putative genes: GmSEC22, GmFLS2-1, GmFLS2-2, GmPSYR, and GmPHO1-H9. Together with the three missense-mutated genes, this produced a final list of eight candidates, a manageable set for follow-up experiments such as CRISPR knockouts, transgenic complementation, or association mapping in breeding panels.
To further prioritize this list, the researchers brought in population-level evidence. Haplotype analysis across a natural soybean population examines whether different versions of a gene, defined by characteristic combinations of SNPs, correlate with measurable differences in the trait among diverse varieties. The team also performed haplotype sequencing of the two parental lines to confirm which haplotypes each parent carried. The result was striking: four of the eight candidate genes showed significant differences in plant height among haplotypes in the natural population. In other words, farmers’ fields and gene banks already contain natural variants of these genes that produce measurably different plants, which is exactly the kind of standing variation that breeders can exploit through marker-assisted selection without any genetic engineering.
The methodological lesson of the study may prove as influential as the gene list itself. BSA-seq alone can localize genomic regions but cannot distinguish the causal gene among dozens in an interval, and it says nothing about whether a variant is actually expressed. RNA-seq alone identifies expression differences but cannot tell whether they are causes or downstream consequences of the trait. By intersecting the two, filtering through tissue specificity, and validating with haplotype associations, the researchers demonstrated a pipeline that converts a crude quantitative trait locus into a short, testable candidate list at a fraction of the cost of traditional fine-mapping, which can require thousands of progeny and years of field phenotyping.
For soybean improvement, the implications are concrete. Lodging resistance and mechanized harvest efficiency are increasingly urgent as soybean production expands into regions where combine harvesting is the norm, and semi-dwarf architecture has historically delivered yield gains in wheat and rice through the Green Revolution. A validated set of height genes in soybean could allow breeders to tune plant architecture precisely, selecting allele combinations that shorten stems just enough to prevent lodging and permit denser planting without sacrificing the biomass and pod number that drive yield. The study’s authors describe their results as a valuable foundation for elucidating the complex genetic architecture of soybean plant height, and the prioritized candidate set they provide gives the research community a clear starting point for functional validation.
The work, published open access on 12 September 2026 and supported by the National Natural Science Foundation of China, the Zhejiang Provincial Natural Science Foundation, and a Zhejiang A&F University research development program, arrives as genomics-driven breeding accelerates across staple crops. As sequencing costs continue to fall and reference pan-genomes expand, integrated approaches that fuse positional mapping with transcriptomics and population haplotype evidence are likely to become the standard route from trait to gene, not only for height in soybean but for the many architectural and stress-resilience traits on which global food security depends.
Subject of Research: Identification of candidate genes controlling plant height in soybean through integrated BSA-seq and RNA-seq analysis
Article Title: Integrated BSA-seq and RNA-seq analysis to identify candidate genes controlling plant height in soybean
Article References: Jin, H., Huang, S., Ding, R., Nan, X., Chen, N., Hu, M., Zheng, Y., Zheng, Z., Hu, X., & Pan, T. (2026). Integrated BSA-seq and RNA-seq analysis to identify candidate genes controlling plant height in soybean. BMC Plant Biology. https://doi.org/10.1186/s12870-026-09944-2
Image Credits: AI Generated
DOI: 10.1186/s12870-026-09944-2
Keywords: soybean, plant height, BSA-seq, RNA-seq, candidate genes, haplotype analysis, plant architecture, lodging resistance, Glycine max, quantitative trait loci, plant hormone signaling, missense mutations
Cite Scienmag News
APA
MLA
Chicago
Juliet Wilcox. (September 12, 2026). Scientists Combine Genome and Gene Expression Data to Pinpoint Genes That Set Soybean Height. Scienmag. https://scienmag.com/scientists-combine-genome-and-gene-expression-data-to-pinpoint-genes-that-set-soybean-height/
Juliet Wilcox. “Scientists Combine Genome and Gene Expression Data to Pinpoint Genes That Set Soybean Height.” Scienmag, 12 September 2026, https://scienmag.com/scientists-combine-genome-and-gene-expression-data-to-pinpoint-genes-that-set-soybean-height/. Accessed 12 September 2026.
Juliet Wilcox. “Scientists Combine Genome and Gene Expression Data to Pinpoint Genes That Set Soybean Height.” Scienmag. September 12, 2026. https://scienmag.com/scientists-combine-genome-and-gene-expression-data-to-pinpoint-genes-that-set-soybean-height/
Copy citation
Download RIS
Tags: BSA-seqcandidate gene identification for soybean crop improvementcandidate genescombining genomic techniques for crop trait researchgenetic mapping of soybean height traitsgenome and gene expression analysis in soybeansGlycine maxhaplotype analysisidentifying genes controlling soybean stem lengthlodging resistancemissense mutationsmolecular markers for soybean plant architecturenear-isogenic soybean lines for trait studyplant architectureplant heightplant height trade-offs in soybean breedingplant hormone signalingquantitative trait lociRNA-seqsoybeansoybean plant height genetics


