Showing posts with label diploidocus. Show all posts
Showing posts with label diploidocus. Show all posts

Thursday, 19 September 2024

#PAGAustralia Poster 14: Synteny-Guided Semi-Automated Curation of Chromosome-Level Genome Assemblies

If you are attending PAG Australia 2024, come and have a chat at Poster 14 about easing the burden of chromosome-level assembly curation.

Abstract Text

Reference genomes are fundamental resources that underpin research across most aspects of modern biology. Technological improvements in the length and accuracy of long-read sequencing platforms, combined with Hi-C proximity ligation sequencing, has enabled the routine generation of highly contiguous phased assemblies, scaffolded to chromosome-level. Nevertheless, automated generation of perfect gapless “telomere-to-telomere” assemblies remains out of reach for most eukaryotic organisms. Scaffolding errors and false duplications can still occur, and assembly curation is now the main bottleneck for large-scale assembly projects. Here, I present a streamlined data workflow and scaffolding assessment for high-throughput manual curation of chromosome-level genome assemblies. Synteny between haplotypes, or closely related species, is combined with read mapping to orient, pair and visualise assembled chromosomes. Assembly gaps are classified according to scaffolding confidence, highlighting candidates for simple scaffolding corrections, such as inversions. Synteny visualisation, gap classification, and HiC contact maps are then combined to identify and document scaffolding edits with increased speed, precision and confidence. This accelerates the production of curated chromosome-level assemblies, and enables the identification of regions of the assembly that may require further attention. Individual tools used in the workflow (ChromSyn, Telociraptor, SynBad, PAFScaff, DepthKopy and DepthCharge) are available at https://github.com/slimsuite/.

Thursday, 29 June 2023

Three Pawsey Internship projects available for the Ocean Genomes Project

We have three Pawsey student internships available this summer with the Ocean Genomes Laboratory in the Minderoo OceanOmics Centre at UWA. Closing date: 07 August, 2023 at 17:00 AWST (Perth time). This is a 10-week, paid program open to exceptional undergrad (2nd/3rd year), Honours, Master’s and PhD students. Apply at the CSIRO Application page. Please get in touch if you want to know more and/or are interested in a student research project in the lab.

Optimising workflows for whole genome assembly for marine vertebrates (Project #04)

The biodiversity of marine vertebrates is critical for the health of our ocean’s ecosystem, but is under immediate threat from climate change, pollution, overfishing and habitat destruction. To advance our understanding of how best to protect and sustain our ocean life, global efforts are underway (such as the Vertebrate Genome Project; VGP) to establish a complete library of high-quality reference genomes for all ~22,000 marine vertebrates.

Reference genomes are pivotal not only for answering fundamental questions in marine biology and evolution, but also for guiding the conservation of species most at risk within our changing oceans, and for accurately monitoring biodiversity.

This project utilizes data generated in-house, either by Illumina short-read or PacBio high-fidelity long-read sequencing of Australian marine vertebrate species. The primary objective is to optimize analysis workflows on Pawsey, encompassing the entire life cycle of the data from its raw format to the ultimate outcome of a high-quality assembled genome. We have data across a diverse range of species covering small to large genome sizes.

A containerised Pawsey workflow for Diploidocus (Project #10)

Bioinformatics in general, and genomics specifically, is replete with complex workflows that do not translate easily to HPC. Frequently, genomics pipelines will incorporate many different tools and/or in-built functions with very different computational requirements in terms of multithreading, memory requirements and IO pressures. The Diploidocus genome curation pipeline exemplifies this problem with some lengthy single-processor steps building on data produced by highly parallelised tools, such as minimap2. As well as adapting a specific mission-critical tool, this project will help identify and establish some general principles for optimising genomics code/workflows for running on Setonix.

Diploidocus is a published genome curation and clean-up tool that utilises several different underlying bioinformatics tools and in-built algorithms. Different steps (and tools) in the pipeline have markedly different CPU, IO and memory requirements, including some lengthy non-parallelised portions. This makes it hard to run efficiently on HPC without wasting resource allocation and/or failing to take advantage of parallelisation when available.

The expected outcome of this project is a Nextflow workflow for the deployment of the Diploidocus pipeline on HPC. This will (a) increase in-house efficiency of HPC usage, and (b) make Diploidocus more attractive as a tool to other research groups.

A containerised Pawsey workflow high throughput phylogenomics (Project #14)

This project aims to produce a robust and efficient phylogenomics workflow for whole genome sequencing data.

One important application of genome assemblies is to test and improve the taxonomic classification of species using large-scale genome-wide phylogenetics, known as phylogenomics. There is a previously developed Snakemake workflow for the rapid generation of phylogenomic trees from low- to mid-coverage whole genome shotgun sequencing data. This pipeline (1) creates multiple rapid draft assemblies; (2) identifies an optimal set of orthologous genes per species using BUSCO and BUSCOMP; (3) generates a multiple sequence alignment per gene; (4) generates a phylogenetic tree per gene; and (5) generates a consensus tree from all the individual gene trees.

There is now a requirement to (1) update the pipeline to be optimised for the high-coverage draft and reference genomes created by the Ocean Genomes Project, and (2) convert this pipeline from PBS/Snakemake to SLURM/Nextflow in-line with other genomics workflows being developed at the Minderoo OceanOmics Centre at UWA.

This project will adapt the wgs2tree workflow to optionally start from a set of existing genome assemblies and BUSCO orthologue annotations and implement a Nextflow/SLURM workflow optimised to run efficiently on Pawsey.

Friday, 14 January 2022

The Waratah genome paper is out!

The final version of the waratah genome paper now out in Molecular Ecology Resources. This was a fun collaboration with the Royal Botanic Gardens and Domain Trust as one of the pilot genomes for BioPlatforms Australia’s Genomics for Australian Plants (GAP) initiative.

You can read the press release here, or our piece in the Conversation, We’ve unveiled the waratah’s genetic secrets, helping preserve this Australian icon for the future.

In this paper, we present a chromosome-level assembly for the NSW State Floral Emblem, the New South Wales waratah, Telopea speciosissima. This joins macadamia as the 2nd reference genome for the Proteaceae family & should help future studies for the remaining ca. 1700 species.

The genome was assembled from a ONT chassis, scaffolded with 10x Genomics linked reads and Phase Genomics HiC - made possible thanks to quality data from AGRF and the Ramaciotti Centre for Genomics. The final assembly was chromosome-level, with 94.1% on the 11 chromosomes (2n = 22).

As well as the assembly itself, the paper presents a three genomics tools that we hope will be helpful for other assemblies:

1. DepthSizer uses long-read depths and BUSCO predictions to estimate genome size. We estimated the waratah genome to be ca. 900 Mbp - bigger than kmer estimates, but smaller than flow cytometry of Tasmanian waratah.

2. Diploidocus builds on Purge Haplotigs, combining read depths, kmer frequencies & BUSCO predictions to classify and curate/filter assembly scaffolds. This decreases false duplications & contamination, and flags collapsed repeats for closer inspection.

3. DepthKopy uses BUSCO Complete genes to establish sequencing depth (like DepthSizer) and then estimates copy number for regions (e.g. genes), scaffolds & sliding windows of the assembly. This showed that most “Duplicated” BUSCOs are real duplicates.


Chen SH, Rossetto M, van der Merwe M, Lu-Irving P, Yap JS, Sauquet H, Bourke G, Amos TG, Bragg JG & Edwards RJ (accepted): Chromosome-level de novo genome assembly of Telopea speciosissima (New South Wales waratah) using long-reads, linked-reads and Hi-C. Molecular Ecology Resources.
[Mol Ecol Res] [bioRxiv]

Abstract

Telopea speciosissima, the New South Wales waratah, is an Australian endemic woody shrub in the family Proteaceae. Waratahs have great potential as a model clade to better understand processes of speciation, introgression and adaptation, and are significant from a horticultural perspective. Here, we report the first chromosome-level genome for T. speciosissima. Combining Oxford Nanopore long-reads, 10x Genomics Chromium linked-reads and Hi-C data, the assembly spans 823 Mb (scaffold N50 of 69.0 Mb) with 97.8% of Embryophyta BUSCOs “Complete”. We present a new method in Diploidocus (https://github.com/slimsuite/diploidocus) for classifying, curating and QC-filtering scaffolds, which combines read depths, k-mer frequencies and BUSCO predictions. We also present a new tool, DepthSizer (https://github.com/slimsuite/depthsizer), for genome size estimation from the read depth of single-copy orthologues and estimate the genome size to be approximately 900 Mb. The largest 11 scaffolds contained 94.1% of the assembly, conforming to the expected number of chromosomes (2n = 22). Genome annotation predicted 40,158 protein-coding genes, 351 rRNAs and 728 tRNAs. We investigated CYCLOIDEA (CYC) genes, which have a role in determination of floral symmetry, and confirm the presence of two copies in the genome. Read depth analysis of 180 “Duplicated” BUSCO genes using a new tool, DepthKopy (https://github.com/slimsuite/depthkopy), suggests almost all are real duplications, increasing confidence in the annotation and highlighting a possible need to revise the BUSCO set for this lineage. The chromosome-level T. speciosissima reference genome (Tspe_v1) provides an important new genomic resource of Proteaceae to support the conservation of flora in Australia and further afield.

If you want a read and don’t have access, please get it touch or check out the bioRxiv preprint.

Thursday, 3 June 2021

Chromosome-level de novo genome assembly of Telopea speciosissima (New South Wales waratah) using long-reads, linked-reads and Hi-C

The latest genomics paper from the lab is now out on bioRvix. This is the first paper from Stephanie Chen’s PhD project in collaboration with the Royal Botanic Gardens and Domain Trust (RBGDT), Sydney. In this paper, Stephanie reports on the chromosome-level assembly of the New South Wales Waratah, the floral emblem of NSW. This is the first of the pilot reference genomes to be released from the Genomics for Australian Plants initiative.

In addition to the genome itself, this paper describes a couple of genomics tools from the lab. DepthSizer (https://github.com/slimsuite/depthsizer) uses BUSCO predictions to establish the single-copy read depth of sequencing data, from which the genome size can be estimated in a way that is hopefully quite robust to assembly quality. Diploidocus (https://github.com/slimsuite/diploidocus) has been used for our previous Dog genome assemblies to help eliminate “haplotigs” (heterozygous regions of the genome that appear in the assembly twice), and low-quality sequences, in addition to flagging possible collapsed repeats or contaminants for further investigation. Here, the Diploidocus “tidy” pipeline is considerably extended for a much more nuanced classification and filtering of scaffolds, using a combination of read depths, homology, kmer analysis and BUSCO predictions.


Chen SH, Rossetto M, van der Merwe M, Lu-Irving P, Yap JS, Sauquet H, Bourke G, Bragg JG & Edwards RJ (preprint): Chromosome-level de novo genome assembly of Telopea speciosissima (New South Wales waratah) using long-reads, linked-reads and Hi-C. bioRxiv 2021.06.02.444084; doi: 10.1101/2021.06.02.444084.
[bioRxiv]

Abstract

Background: Telopea speciosissima, the New South Wales waratah, is Australian endemic woody shrub in the family Proteaceae. Waratahs have great potential as a model clade to better understand processes of speciation, introgression and adaptation, and are significant from a horticultural perspective. Findings: Here, we report the first chromosome-level reference genome for T. speciosissima. Combining Oxford Nanopore long-reads, 10x Genomics Chromium linked-reads and Hi-C data, the assembly spans 823 Mb (scaffold N50 of 69.0 Mb) with 91.2 % of Embryophyta BUSCOs complete. We introduce a new method in Diploidocus (https://github.com/slimsuite/diploidocus) for classifying, curating and QC-filtering assembly scaffolds. We also present a new tool, DepthSizer (https://github.com/slimsuite/depthsizer), for genome size estimation from the read depth of single copy orthologues and find that the assembly is 93.9 % of the estimated genome size. The largest 11 scaffolds contained 94.1 % of the assembly, conforming to the expected number of chromosomes (2n = 22). Genome annotation predicted 40,158 protein-coding genes, 351 rRNAs and 728 tRNAs. Our results indicate that the waratah genome is highly repetitive, with a repeat content of 62.3 %. Conclusions: The T. speciosissima genome (Tspe_v1) will accelerate waratah evolutionary genomics and facilitate marker assisted approaches for breeding. Broadly, it represents an important new genomic resource of Proteaceae to support the conservation of flora in Australia and further afield.

Thursday, 18 March 2021

Chromosome-length genome assembly and structural variations of the primal Basenji dog (Canis lupus familiaris) genome

Our latest genome paper is now out at BMC Genomics. This was a second collaboration in the team behind the German Shepherd Dog genome last year, led by Bill Ballard. This time, we used a combination of BGI short reads, ONT long reads, and Hi-C scaffolding to generate a chromosome-length assembly. This is one of the most intact and complete dog genomes generated to date, and joins only a handful of published breed-specific chromosome-length assemblies.

The Basenji is particularly interesting as it sits at the base of the dog breed family tree, making it a good unbiased reference for future comparisons between breeds.

The paper also has a few nice nuggets for those interesting in genome assembly. Of particular interest, our initial assembly had an artefact where the entire mitochondrial genome got assembled (in two copies) into the middle of one of the nuclear chromosomes. It is not entirely clear why this happened, but it was inserted into a NUMT (nuclear mitochondrial DNA insertion) fragment at that location. To make finding such things easier, we’ve released a new NUMT finding tool, NUMTFinder.

As with our previous dog genome, the German Shepherd Dog, we also observe that the tandem repeat of Amy2B (Amylase Alpha 2B) genes, was assembled intact but with fewer copies than are present in the actual genome. (This gene is of interest for dog domestication and adaptations to a starch-rich diet.) Crucially, without looking at the raw sequencing data, it would not have been clear that the assembly under-represents Amy2B copy number. This kind of analysis can be repeated using the regcheck or regcnv run modes of Diploidocus, which estimates the copy number of a region based on its read depth versus the single-copy read depth determined from BUSCO single-copy complete genes.

Overall, this presents a nice case study of the need for a bit of TLC and manual curation, even when you have some very impressive completeness and contiguity statistics.


Edwards RJ, Field MA, Ferguson JM, Dudchenko O, Keilwagen K, Rosen BD, Johnson GS, Rice ES, Hillier L, Hammond JM, Towarnicki SG, Omer A, Khan R, Skvortsova K, Bogdanovic O, Zammit RA, Lieberman Aiden E, Warren WC & Ballard JWO (2021): Chromosome-length genome assembly and structural variations of the primal Basenji dog (Canis lupus familiaris) genome. BMC Genomics 22:188

Abstract

Background: Basenjis are considered an ancient dog breed of central African origins that still live and hunt with tribesmen in the African Congo. Nicknamed the barkless dog, Basenjis possess unique phylogeny, geographical origins and traits, making their genome structure of great interest. The increasing number of available canid reference genomes allows us to examine the impact the choice of reference genome makes with regard to reference genome quality and breed relatedness.

Results: Here, we report two high quality de novo Basenji genome assemblies: a female, China (CanFam_Bas), and a male, Wags. We conduct pairwise comparisons and report structural variations between assembled genomes of three dog breeds: Basenji (CanFam_Bas), Boxer (CanFam3.1) and German Shepherd Dog (GSD) (CanFam_GSD). CanFam_Bas is superior to CanFam3.1 in terms of genome contiguity and comparable overall to the high quality CanFam_GSD assembly. By aligning short read data from 58 representative dog breeds to three reference genomes, we demonstrate how the choice of reference genome significantly impacts both read mapping and variant detection.

Conclusions: The growing number of high-quality canid reference genomes means the choice of reference genome is an increasingly critical decision in subsequent canid variant analyses. The basal position of the Basenji makes it suitable for variant analysis for targeted applications of specific dog breeds. However, we believe more comprehensive analyses across the entire family of canids is more suited to a pangenome approach. Collectively this work highlights the importance the choice of reference genome makes in all variation studies.