The de novo genome assembly workflow is summarized in Fig. 1. A hybrid assembly strategy integrating PacBio HiFi, ONT UL, Hi-C and NGS sequencing data was employed. Initial assemblies were generated using hifiasm (v0.24.0-r703) (Cheng et al., 2024) and Verkko (v2.2.1) (Rautiainen et al., 2023) with combined HiFi and ONT UL reads (length > 100 kb, quality value > 10). In addition, hifiasm assemblies were generated independently using HiFi-only and ONT-only datasets. Organelle-derived contigs were identified by aligning assembled contigs to reference chloroplast and mitochondrial genomes of Triticum aestivum (NCBI) using minimap2 (Li et al., 2018), and contigs with ≥ 95% organelle sequence content and ≤ 5% divergence were removed.
For scaffolding, Hi-C data were processed using HiC-Pro (v3.1.0) (Servant et al., 2015), and valid interaction pairs were used by YAHS (v1.2.2) (Zhou et al., 2023) to scaffold multiple assembly versions. Scaffolds were further manually curated using Juicebox (v1.11.08) (Durand et al., 2016). Gap filling was performed using the combined HiFi+ONT assembly (hifiasm) as the backbone. Whole-genome collinearity between alternative assemblies and the backbone was established using AnchorWave genoAli (v1.2.6) (Song et al., 2022). Gaps in the backbone assembly were systematically replaced with contiguous sequences from alternative assemblies based on the intervals between collinear anchor pairs. This procedure was iteratively applied across three alternative assemblies to maximize gap closure. Finally, the assembly was polished using NextPolish2 (v0.2.1) (Hu et al., 2024) with HiFi and NGS reads for error correction.
- Cheng, H., Asri, M., Lucas, J., Koren, S., and Li, H. (2024). Scalable telomere-to-telomere assembly for diploid and polyploid genomes with double graph. Nat. Methods 21(6), 967-970. https://doi.org/10.1038/s41592-024-02269-8.
- Rautiainen, M., Nurk, S., Walenz, B.P., Logsdon, G.A., Porubsky, D., Rhie, A., Eichler, E.E., Phillippy, A.M., and Koren, S. (2023). Telomere-to-telomere assembly of diploid chromosomes with Verkko. Nat. Biotechnol. 41(10), 1474-1482. https://doi.org/10.1038/s41587-023-01662-6
- Li, H. (2018). Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics 34(18), 3094-3100. https://doi.org/10.1093/bioinformatics/bty191.
- Servant, N., Varoquaux, N., Lajoie, B.R., Viara, E., Chen, C.J., Vert, J.P., Heard, E., Dekker, J., and Barillot, E. (2015). HiC-Pro: an optimized and flexible pipeline for Hi-C data processing. Genome Biol. 16(1), 259. https://doi.org/10.1186/s13059-015-0831-x.
- Zhou, C., McCarthy, S.A., and Durbin, R. (2023). YaHS: yet another Hi-C scaffolding tool. Bioinformatics 39(1), btac808. https://doi.org/10.1093/bioinformatics/btac808.
- Durand, N.C., Robinson, J.T., Shamim, M.S., Machol, I., Mesirov, J.P., Lander, E.S., and Aiden, E.L. (2016). Juicebox provides a visualization system for Hi-C contact maps with unlimited zoom. Cell Syst. 3(1), 99-101. https://doi.org/10.1016/j.cels.2015.07.012.
- AnchorWave: Sensitive alignment of genomes with high sequence diversity, extensive structural polymorphism, and whole-genome duplication. Proc. Natl. Acad. Sci. USA 119(1), e2113075119. https://doi.org/10.1073/pnas.2113075119.
- Hu, J., Wang, Z., Liang, F., Liu, S.L., Ye, K., and Wang, D.P. (2024). NextPolish2: a repeat-aware polishing tool for genomes assembled using HiFi long reads. Genomics Proteomics Bioinformatics 22(1), qzad009. https://doi.org/10.1093/gpbjnl/qzad009.
