• HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
Wednesday, August 26, 2026
BIOENGINEER.ORG
No Result
View All Result
  • Login
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
No Result
View All Result
Bioengineer.org
No Result
View All Result
Home NEWS Science News Agriculture

New Microhaplotype Databases and Tools Reveal Crop Genetic Diversity

Bioengineer by Bioengineer
August 26, 2026
in Agriculture
Reading Time: 6 mins read
0
New Microhaplotype Databases and Tools Reveal Crop Genetic Diversity
Share on FacebookShare on TwitterShare on LinkedinShare on RedditShare on Telegram

Tiny DNA Signatures Give Crop Breeders a Sharper View of Hidden Genetic Diversity

Plant breeders have gained a new way to see genetic variation that standard DNA markers can miss. An international research team has assembled standardized databases of “microhaplotypes”—short stretches of DNA containing several tightly linked variants—for eight important crops. The resources cover alfalfa, blueberry, cranberry, cucumber, pecan, potato, strawberry and sweetpotato, bringing together more than 60,000 samples and tens of thousands of distinct sequence variants. The researchers say the framework can improve population analysis, genetic mapping and breeding decisions, particularly in crops with duplicated genomes, high heterozygosity or multiple chromosome sets. Their study, published in Theoretical and Applied Genetics, also introduces HapApp, a no-code software tool designed to help breeders add newly discovered variants to the growing databases without writing scripts.

Most modern crop-genotyping systems rely on single-nucleotide polymorphisms, or SNPs. A SNP records a single DNA-letter difference and usually has two possible states, making it a relatively simple biallelic marker. SNPs are abundant, inexpensive to measure and easy to analyze, but that simplicity can become a weakness in complex crops. A short genomic region may contain several nearby variants that are inherited together. Treating each position as an isolated, two-option marker throws away the information contained in the precise combination of variants. A microhaplotype preserves that combination as a single, multiallelic signature. Because the variants lie within a short sequence, they are read together and are generally assumed to remain linked, allowing researchers to determine which changes occur on the same chromosome copy. The result can be a marker with three, ten or even dozens of distinguishable allelic forms rather than only two.

The team built its resources from DArTag targeted-genotyping data. Unlike random reduced-representation methods such as genotyping-by-sequencing, which sample a different subset of the genome from one project to another, DArTag uses designed primers to amplify the same target regions repeatedly. Sequencing reads typically span 50 to 250 base pairs and include the chosen variant as well as neighboring polymorphisms. The researchers focused on the sequence portions of about 54 or 81 base pairs, depending on the panel design. Each DArTag report contains read counts for several sequence classes, including known reference and alternative alleles and newly observed “RefMatch” or “AltMatch” sequences that resemble those alleles but contain additional changes in the flanking DNA. These unanticipated combinations are precisely where microhaplotypes can reveal diversity that a conventional SNP call would overlook.

A central problem was that the same sequence could appear under different identifiers in separate projects. Without a stable name, a breeder could not reliably compare a marker analyzed in one laboratory with the apparently similar marker generated elsewhere. The researchers therefore created a species-agnostic pipeline that maps markers to chromosome or scaffold coordinates and assigns fixed identifiers such as “Chr01_000589632.” The workflow aligns amplicon sequences to 180-to-300-base-pair regions of a reference genome using BLAST, checks strand orientation and removes ambiguity in the original reports. A sequence is linked to an existing database identity only when it matches the full reference entry at 100 percent identity and coverage. A sequence that fails to match exactly but reaches at least 90 percent identity across at least 90 percent of the comparable region is treated as a candidate new variant. Sequences below that threshold are discarded as likely technical artifacts or off-target products.

The database-building effort also preserves a distinction between true alleles and sequences generated from duplicated genomic regions. In polyploid crops, closely related chromosome copies can be difficult to separate, and primers may amplify paralogous loci—related but nonallelic regions—alongside the intended target. The team retained these signals but annotated them rather than silently deleting them, because the same sequences may recur in future experiments. That choice creates a comprehensive catalog while allowing downstream analyses to exclude questionable markers. In a biparental population, for example, biological segregation limits the number of genuine alleles expected at a locus. Linkage software can also identify markers with abnormal inheritance or poor recombination behavior. This layered strategy lets the global database remain inclusive while giving breeders tools to isolate orthologous, biologically interpretable variation.

The eight crop databases vary dramatically in size and genetic richness. The alfalfa resource contains 35,259 microhaplotypes from 3,000 loci and more than 15,600 samples gathered across 25 breeding projects. Its average marker carries about 12 alleles, and the database has approached a plateau near 35,000 sequences, suggesting that the current panel has captured much of the diversity present in the sampled US breeding populations. The blueberry database contains 28,653 sequences from roughly 8,930 samples and 3,000 loci, with an average of about ten alleles per locus. Potato contributes 35,506 sequences from 3,913 loci and more than 3,100 samples, while strawberry contains 38,227 sequences from 5,000 loci and 1,880 samples. The strawberry panel targets all 28 chromosomes of its octoploid genome, including its four subgenomes, and half of its loci contain six or fewer alleles.

The remaining databases illustrate how sampling and genome biology shape the apparent amount of diversity. The cranberry resource contains 14,380 alleles from 3,050 loci and 4,146 samples. Cucumber has 9,823 microhaplotypes from 3,059 loci and 8,272 samples, but a mean of only about three alleles per marker; the authors caution that this may reflect the narrow breeding material sampled rather than a universally low level of cucumber diversity. Pecan contains 26,073 alleles from 3,100 loci and 6,768 samples. Sweetpotato, a globally distributed hexaploid crop with six copies of each chromosome set, contains 37,593 sequences from 3,120 loci and 9,212 samples. Its data span breeding programs in North and South America, Africa, Asia and the Caribbean, allowing the database to capture geographically structured variation rather than diversity from a single breeding population.

Two case studies tested whether microhaplotypes actually improve genetic analyses rather than simply generating larger databases. In pecan, the researchers constructed linkage maps from 188 offspring using three data types: microhaplotypes, only the target SNPs, and all SNPs detected within the amplified regions. The microhaplotype data retained more markers that were informative about recombination in both parents. By contrast, the SNP datasets were dominated by markers informative in only one parent, making it difficult to connect the two parental maps. When the researchers ordered markers using multidimensional scaling of genetic distances, the SNP-based maps placed some markers far from their expected physical positions. Microhaplotypes produced fewer gaps and more stable ordering because the linked variants supplied additional allele combinations and made parental phase easier to determine directly from sequencing reads.

The second test used 4,087 matched sweetpotato samples from seven international projects. After quality filtering, the researchers compared 27,220 microhaplotypes at 2,772 loci with 2,534 target SNPs and 20,935 SNPs extracted from the same regions. Principal-component analysis produced tighter, more clearly separated clusters with microhaplotypes, especially for the Taiwanese population. The first principal component explained 15.72 percent of the variation with microhaplotypes, compared with 10.5 percent for target SNPs and 5 percent for all SNPs. Discriminant analysis of principal components showed that the first linear discriminant axis captured 86.3 percent of between-group variance with microhaplotypes, versus about 60 percent with either SNP dataset. All three approaches ultimately achieved mean classification accuracy above 98 percent, but microhaplotypes remained near a 99 percent success plateau as more principal components were included, suggesting greater stability in a high-dimensional analysis.

The researchers emphasize that the databases are not a universal census of crop diversity. Marker panels were designed mainly in genic regions, so they preferentially sample potentially functional portions of the genome rather than random DNA. Database sizes are also influenced by panel length, genome complexity, ploidy, sample number and the breeding populations that collaborators were able to share. Rare-variant counts are especially sensitive to sample size: in sweetpotato, similarly sized US and Taiwanese cohorts contained vastly different numbers of private microhaplotypes, likely reflecting differences in breeding history and gene-pool structure rather than sampling alone. Even so, standardized records could help genebanks detect duplicated accessions, uncover mislabeled material, identify gaps in collections and assemble representative core sets for pre-breeding. The sequences and scripts are being released under FAIR data principles, with the databases available through Zenodo and the software through GitHub.

HapApp translates the computational workflow into a point-and-click interface. Users upload a DArTag report, select the crop and panel characteristics, and receive a filtered report in which every accepted microhaplotype has a stable identity, along with an updated FASTA sequence file and, when necessary, a new database version. The authors are also developing HapSearch, a planned platform for finding alleles by crop, locus, sequence similarity or project keyword and for identifying germplasm associated with unusual variants. At present, the system is built around DArTag’s proprietary MADC reports, so other targeted platforms such as GT-seq, AgriSeq and FlexSeq would require format-conversion pipelines. Still, the researchers argue that the underlying principle is portable: preserve linked sequence information, give recurring variants durable names and use multiallelic data to make breeding genomes easier to read. In crops facing climate stress, disease and changing production demands, that extra resolution could help breeders find useful genetic variation before it disappears into the noise of a two-allele marker system.

Subject of Research: Standardized microhaplotype databases and genetic-diversity analysis for eight crop species

Article Title: Standardized microhaplotype databases and frameworks for assessing and mining crop genetic diversity

Article References: Zhao, D., Lin, M., Taniguti, C. H. et al. “Standardized microhaplotype databases and frameworks for assessing and mining crop genetic diversity.” Theoretical and Applied Genetics 139, 244 (2026). Original research article

Image Credits: AI Generated

DOI: 10.1007/s00122-026-05340-4

Keywords: microhaplotypes, crop genetics, plant breeding, genetic diversity, polyploid crops, DArTag genotyping, linkage mapping, population structure, HapApp

Tags: crop genetic diversitycrop genome analysis toolscrop genotyping methodsDNA variation in cropsduplicated crop genomesgenetic mapping in agriculturehigh heterozygosity in cropsmicrohaplotypesmulti-variant DNA signaturesno-code genetic analysis softwareplant breeding genetic markersplant population genetics databases

Share12Tweet7Share2ShareShareShare1

Related Posts

Tree peony transcription factor PrIDD7 suppresses seed oil accumulation

Tree peony transcription factor PrIDD7 suppresses seed oil accumulation

August 26, 2026
Scientists uncover how natural products may combat chronic obstructive pulmonary disease

Scientists uncover how natural products may combat chronic obstructive pulmonary disease

August 26, 2026

Phenolic Acids and Protein Levels in Wheat Landraces and Cultivars Without Nitrogen

August 26, 2026

Host selection shapes microbial communities and functions, determining mangrove soil carbon fate

August 26, 2026

POPULAR NEWS

  • Study Evaluates Face-Validity Measures for Automatically Detected Dog Sleep Spindles

    29 shares
    Share 12 Tweet 7
  • New Microhaplotype Databases and Tools Reveal Crop Genetic Diversity

    29 shares
    Share 12 Tweet 7
  • Tree peony transcription factor PrIDD7 suppresses seed oil accumulation

    29 shares
    Share 12 Tweet 7
  • How CsgD Regulates Biofilm Formation and Quorum Sensing in Salmonella Typhimurium

    29 shares
    Share 12 Tweet 7

About

We bring you the latest biotechnology news from best research centers and universities around the world. Check our website.

Follow us

Recent News

Study Evaluates Face-Validity Measures for Automatically Detected Dog Sleep Spindles

New Microhaplotype Databases and Tools Reveal Crop Genetic Diversity

Tree peony transcription factor PrIDD7 suppresses seed oil accumulation

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 85 other subscribers
  • Contact Us

Bioengineer.org © Copyright 2023 All Rights Reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Homepages
    • Home Page 1
    • Home Page 2
  • News
  • National
  • Business
  • Health
  • Lifestyle
  • Science

Bioengineer.org © Copyright 2023 All Rights Reserved.