条形banner-03

News

FAS Blog_1

Reference Genome Bias: Why Choosing the Right Reference Genome Matters

Whether you are performing whole-genome sequencing (WGS), RNA sequencing (RNA-seq), or reduced-representation sequencing such as SLAF-seq, one of the first steps in data analysis is to align sequencing reads to a reference genome. Although this process may seem straightforward, the choice of reference genome can significantly affect the results.

What Is Reference Genome Bias?

A reference genome provides a coordinate system for mapping sequencing reads. However, a reference genome typically represents only one individual, one haplotype, or a limited set of genomes and therefore cannot fully capture the genetic diversity within a species.

Reference bias begins with read mappability—the ability of a sequencing read to be mapped accurately to a reference genome. Reads that closely match the reference sequence are generally more likely to map correctly and receive a high mapping quality score. In contrast, reads that differ from the reference may be more difficult to map, may receive lower mapping quality scores, or may fail to map altogether. Reads that cannot be mapped or do not meet downstream quality criteria may therefore be excluded, resulting in the loss of potentially valuable genetic information.

This bias can occur even when reads are successfully mapped. For example, reads carrying the reference allele may map more efficiently than those carrying an alternative allele because the reference allele matches the reference genome directly. This can reduce the mapping quality of alternative-allele reads and result in their removal during downstream filtering. Likewise, sequences present in the sample but absent from the reference genome—including some insertions and structural variants—may have no corresponding reference sequence to which they can be aligned and may therefore be missed.

Consequently, the reads retained after reference-based alignment may not fully represent the sequences and alleles present in the original sample (Paten et al. 2017). This biased representation of genomic information is a key mechanism underlying reference genome bias.

Is Reference Bias Avoidable?

Reference bias is difficult to avoid completely because the individual being sequenced is almost never the same individual used to generate the reference genome. Individuals within the same species can carry distinct genetic variants, including SNPs, insertions, deletions, and other structural variants. As the genetic differences between the sample and the reference increase, reads carrying these differences may become more difficult to align. As a result, genomic regions or alleles that differ from the reference may be underrepresented in the mapped data.

Although reference bias cannot be completely eliminated, its impact can be minimised by selecting an appropriate, high-quality reference genome that is genetically representative of the study samples.

How Can Reference Bias Affect Your Study?

DNA Sequencing Applications

In whole-genome sequencing and other DNA-sequencing applications, reference bias can affect variant detection. Alternative alleles may receive less read support, while sequences absent from the reference may be missed. This is particularly important for studies of wild populations, genetically diverse species, hybrids, and polyploid organisms. Reference bias can therefore influence population genetic analyses and the resulting biological conclusions. A recent study showed that reference genome choice affected SNP discovery, as well as estimates of nucleotide diversity, population diversity, and effective population size (Akopyan et al. 2025).

RNA-seq Studies

In RNA-seq studies, reference genome bias can affect gene expression quantification because genetic variation can influence the mapping of RNA-seq reads. Reads carrying a non-reference allele may be less likely to map correctly to the reference genome, resulting in biased read counts for some genes or transcripts and, consequently, inaccurate gene expression estimates. These effects are particularly relevant for studies of allele-specific expression, hybrids, genetically diverse populations, and cross-species comparisons, in which genetic differences between the samples and the reference genome may be substantial (Stevenson et al. 2013).

How to Choose the Right Reference Genome

Although reference bias cannot be eliminated, careful reference genome selection can greatly reduce its impact. Where possible, consider the following points:

  • Use a reference from the same species. A conspecific reference is generally preferable, particularly for population genomic studies. Recent evidence shows that using a more closely related reference can substantially improve read mapping and variant detection.
  • Choose a high-quality, complete assembly. Chromosome-level, haplotype-resolved assemblies generated using modern long-read technologies can provide a more comprehensive genomic representation than fragmented draft assemblies generated with short reads.
  • Consider genetic distance. For genetically diverse populations, choose a reference that is representative of, or genetically close to, the study population.
  • Consider ploidy and haplotype structure. These factors are particularly important when working with polyploid or highly heterozygous species.
  • Consider a pangenome for highly diverse datasets. When multiple high-quality genomes are available, a pangenome or pangenome graph can represent sequence diversity that is absent from a single linear reference genome (Ashraf et al. 2026).

BMKGENE’s Recommendation

At BMKGENE, we provide a range of reference genomes in our database to support different research needs. Depending on the species, population, and study objectives, clients can select the reference genome that best aligns with their research goals.

When multiple suitable references are available, selecting a genetically appropriate, high-quality reference can improve read mapping and increase the reliability of downstream analyses, including variant detection, gene expression quantification, and population genetic analysis.

References

  1. Akopyan M, Genchev M, Armstrong EE, Mooney JA. (2025). Reference genome choice compromises population genetic analyses. Cell, 188(24), 6939–6952.e11. DOI: 10.1016/j.cell.2025.08.034
  2. Stevenson KR, Coolon JD, Wittkopp PJ. (2013). Sources of bias in measures of allele-specific expression derived from RNA-seq data aligned to a single reference genome. BMC Genomics, 14, 536. DOI: 10.1186/1471-2164-14-536
  3. Paten B, Novak AM, Eizenga JM, Garrison E. (2017). Genome graphs and the evolution of genome inference. Genome Research, 27(5), 665–676.
  4. Ashraf H, Doerr D, Ebler J, et al. (2026). Building and applying pangenome references to capture genetic diversity. Nature Reviews Genetics. DOI: 10.1038/s41576-026-00987-7

Post time: Aug-20-2026

Send your message to us: