Showing posts with label starlings. Show all posts
Showing posts with label starlings. Show all posts

Friday, 24 February 2023

Contrasting Patterns of Single Nucleotide Polymorphisms and Structural Variation Across Multiple Invasions

Stuart KC, Edwards RJ, Sherwin WB & Rollins LA (2023): Contrasting patterns of single nucleotide polymorphisms and structural variations across multiple invasions. Mol. Biol. Evol. 40:msad046. [Mol. Biol. Evol.] [PubMed] [bioRxiv]

Genetic divergence is the fundamental process that drives evolution and ultimately speciation. Structural variants (SVs) are large-scale genomic differences within a species or population and can cause functionally important phenotypic differences. Characterizing SVs across invasive species will fill knowledge gaps regarding how patterns of genetic diversity and genetic architecture shape rapid adaptation under new selection regimes. Here, we seek to understand patterns in genetic diversity within the globally invasive European starling, Sturnus vulgaris. Using whole genome sequencing of eight native United Kingdom (UK), eight invasive North America (NA), and 33 invasive Australian (AU) starlings, we examine patterns in genome-wide SNPs and SVs between populations and within Australia. Our findings detail the landscape of standing genetic variation across recently diverged continental populations of this invasive avian. We demonstrate that patterns of genetic diversity estimated from SVs do not necessarily reflect relative patterns from SNP data, either when considering patterns of diversity along the length of the organism’s chromosomes (owing to enrichment of SVs in subtelomeric repeat regions), or interpopulation diversity patterns (possibly a result of altered selection regimes or introduction history). Finally, we find that levels of balancing selection within the native range differ across SNP and SV of different classes and outlier classifications. Overall, our results demonstrate that the processes that shape allelic diversity within populations is complex and support the need for further investigation of SVs across a range of taxa to better understand correlations between often well-studied SNP diversity and that of SVs.

Thursday, 5 January 2023

Evolutionary genomics: Insights from the invasive European starlings

Happy New Year, starling lovers! Our latest paper, looking at evolutionary insights gleaned from the starling genome during Kat Stuart's PhD, is now out in Frontiers in Genetics:

Stuart KC, Sherwin WB, Edwards RJ & Rollins LA (2023): Evolutionary genomics: Insights from the invasive European starlings. Frontiers in Genetics 13:1010456. [Front Genet] [PubMed]

Two fundamental questions for evolutionary studies are the speed at which evolution occurs, and the way that this evolution may present itself within an organism’s genome. Evolutionary studies on invasive populations are poised to tackle some of these pressing questions, including understanding the mechanisms behind rapid adaptation, and how it facilitates population persistence within a novel environment. Investigation of these questions are assisted through recent developments in experimental, sequencing, and analytical protocols; in particular, the growing accessibility of next generation sequencing has enabled a broader range of taxa to be characterised. In this perspective, we discuss recent genetic findings within the invasive European starlings in Australia, and outline some critical next steps within this research system. Further, we use discoveries within this study system to guide discussion of pressing future research directions more generally within the fields of population and evolutionary genetics, including the use of historic specimens, phenotypic data, non-SNP genetic variants (e.g., structural variants), and pan-genomes. In particular, we emphasise the need for exploratory genomics studies across a range of invasive taxa so we can begin understanding broad mechanisms that underpin rapid adaptation in these systems. Understanding how genetic diversity arises and is maintained in a population, and how this contributes to adaptability, requires a deep understanding of how evolution functions at the molecular level, and is of fundamental importance for the future studies and preservation of biodiversity across the globe.

Wednesday, 29 June 2022

The starling genome is out!

See the pre-print post for details.

Stuart KC*, Edwards RJ*, Cheng Y, Warren WC, Burt DW, Sherwin WB, Hofmeister NR, Werner SJ, Ball GF, Bateson M, Brandley MC, Buchanan KL, Cassey P, Clayton DF, De Meyer T, Meddle SL & Rollins LA (2022): Transcript- and annotation-guided genome assembly of the European starling. Molecular Ecology 22(8):3141-3160. doi: 10.1111/1755-0998.13679. [*Joint first authors] [Mol Ecol Res] [PubMed] [bioRxiv]

The European starling, Sturnus vulgaris, is an ecologically significant, globally invasive avian species that is also suffering from a major decline in its native range. Here, we present the genome assembly and long-read transcriptome of an Australian-sourced European starling (S. vulgaris vAU), and a second, North American, short-read genome assembly (S. vulgaris vNA), as complementary reference genomes for population genetic and evolutionary characterization. S. vulgaris vAU combined 10× genomics linked-reads, low-coverage Nanopore sequencing, and PacBio Iso-Seq full-length transcript scaffolding to generate a 1050 Mb assembly on 6222 scaffolds (7.6 Mb scaffold N50, 94.6% busco completeness). Further scaffolding against the high-quality zebra finch (Taeniopygia guttata) genome assigned 98.6% of the assembly to 32 putative nuclear chromosome scaffolds. Species-specific transcript mapping and gene annotation revealed good gene-level assembly and high functional completeness. Using S. vulgaris vAU, we demonstrate how the multifunctional use of PacBio Iso-Seq transcript data and complementary homology-based annotation of sequential assembly steps (assessed using a new tool, saaga) can be used to assess, inform, and validate assembly workflow decisions. We also highlight some counterintuitive behaviour in traditional busco metrics, and present buscomp, a complementary tool for assembly comparison designed to be robust to differences in assembly size and base-calling quality. This work expands our knowledge of avian genomes and the available toolkit for assessing and improving genome quality. The new genomic resources presented will facilitate further global genomic and transcriptomic analysis on this ecologically important species.

Sunday, 13 February 2022

Edwards Lab at #LorneGenome 2022

Lorne Genome 2022 (the 43rd Annual Lorne Genome Conference 2022) kicks off today in Lorne and online. I wasn’t able to make it in person this year due to Omicron and teaching commitments, but happily the lab is still well represented. As well as an online talk, we have two in-person posters, so please check these out if you are lucky enough to be attending in the flesh.

Details below.


A chromosome-level reference genome for Telopea speciosissima (New South Wales waratah) provides insight into waratah evolution (#138)

Stephanie H Chen, Jason G Bragg, Richard J Edwards

Telopea is an eastern Australian genus of five species of long-lived shrubs in the family Proteaceae. Previous work has characterised population structure and patterns of introgression between Telopea species. These studies were performed using a limited set of genetic markers, but point to the great potential of waratah as a model clade for understanding the processes of divergence, environmental adaptation and speciation, when enhanced by a genome-wide perspective enabled by a reference genome. However, few Proteaceae genomes and no waratah genomes are available.

We assembled the first chromosome-level reference genome for T. speciosissima (New South Wales waratah; 2n = 22) using Nanopore long-reads, 10x Chromium linked-reads and Hi-C data. The assembly spans 823 Mb (scaffold N50 of 69.0 Mb) with 97.8 % of Embryophyta universal single-copy orthologues (BUSCOs; n = 1,614) complete. Read depth analysis of 140 ‘Duplicated’ BUSCO genes reveals that almost all are real duplications, increasing confidence in protein family analysis using annotated protein-coding genes, highlighting a possible need to revise the BUSCO set for this lineage. Genome annotation predicted 34,706 genes and pseudogenes, including 27,481 protein-coding genes. We examined the evolutionary dynamics of Telopea using the reference genome in conjunction with DArTseq (n = 244) and whole genome shotgun sequencing (n = 14) of each of the seven lineages; there are three lineages of T. speciosissima – coastal, upland and southern.

Here, I will discuss the population structure and demographic history of the genus. We also examined phylogenomic relationships and developed a scalable method of rapidly generating species trees from short-read data to maximise the recovery of informative data from genomic datasets. The waratah reference genome represents an important new genomic resource in Proteaceae to accelerate our understanding of the origins and evolutionary dynamics of the Australian flora.

[Read more about the waratah genome, here.]


Small but mitey: high-quality long-read assembly of a streamlined mite genome from contaminated sequencing data (#17)

Richard J Edwards, Stephanie H Chen, Jason G Bragg.

As pilot data for project on myrtle rust resistance, we previously assembled two Myrtaceae genomes using 10x Chromium linked reads: Rhodamnia argentea (silver malletwood) and Syzygium oleosum (blue lilly pilly). Both draft genomes achieved scaffolding (N50 > 850 kb) and completeness (BUSCOv3 embryophyta_odb9 > 90 %) of sufficient quality to be annotated by NCBI RefSeq. However, signs of arthropod sequence contamination were subsequently found in the Rhodamnia argentea assembly. We therefore sought to identify and eliminate this contamination during improvement and curation of the genome for publication.

A risk-averse analysis highlighted 49.6 Mb (11.95%) on 2,996 of 15,781 scaffolds of possible arthropod origin. An improved assembly of the same tree, incorporating ~50X long-read (ONT) sequencing, has confirmed this contamination as 11 scaffolds (34.6 Mb) that are distinct from 75 R. argentea assembly scaffolds (346.7 Mb), increasing the likelihood of contamination over the integration of horizontally transferred genes. Taxonomic analysis of predicted protein-coding genes using Taxolotl (https://github.com/slimsuite/taxolotl) suggested that the contamination most likely originates from some form of mite (Order: Trombidiformes), but limited NCBInr mite sequences precluded better taxonomic resolution. Curiously, these contamination scaffolds showed a high depth of coverage (~36X), but a fairly low BUSCO completeness of 58.1% (v5 Augustus, metazoa_odb10 n=954), apparently inconsistent with typical mite genomes.

Phylogenomic analysis with available mite genomes identified the closest relative as Aculops lycopersici, a microscopic (0.2 mm long) eriophyoid mite with a heavily streamlined 32.5 Mb genome. Original low completeness appears to be from a combination of genome reduction and poor performance of that BUSCO version; BUSCO v5 MetaEuk eukaryota_odb10 (n=255) reports 82.8% completeness, which is approaching the 86.3% of A. lycopersici. Here, we discuss the evidence that we have assembled a highly complete but streamlined genome from an unknown eriophyoid mite, plus the need to improve genomic representation of contaminating pest species.


A genetic perspective on rapid adaptation in the globally invasive European starling (Sturnus vulgaris) (#255)

Katarina C Stuart, Richard J Edwards, William (Bill) B Sherwin, Lee Ann Rollins.

Few invasive birds are as globally successful or as well-studied as the common starling (Sturnus vulgaris). Native to the Palaearctic, the starling has been a prolific invader in North and South America, southern Africa, Australia, and The Pacific Islands, while facing declines in excess of 50% in in some native regions. Starlings present an invaluable opportunity to test predictions about the evolutionary trajectory of invasive populations, and gain insight into genetic shifts in response to anthropogenic alteration and climate change.

My research focuses primarily on the invasive European starling population in Australia and aims to investigate the genetics underlying their evolution, using a range of genomic approaches. Through historic museum sample sequencing, I examine single nucleotide polymorphism variations shifts between the native range and Australia, and find parallel selection on both continents, possibly resulting from common global selective forces such as exposure to pollutants and carbohydrate exposure. I further examine matched genetic, morphological, and environmental data to reveal patterns of heritability and plasticity across ecologically significant phenotypic traits, revealing that elevation, as well as rainfall and temperature variability plays an important role in shaping morphology and genetics. Finally, I investigated patterns of structural variants, to uncover evolutionarily significant large-scale genetic variants across a global data set, and more specifically characterise their role in rapid starling adaptation across the entirety of the Australian range. Overall, my research seeks to better understand mechanisms and patterns of genetic change within this species, which may be used to inform invasion or native range management. More broadly, this evolutionary research into the starling provide an important perspective on the role of rapid evolution in invasive species persistence, and the global pressures that may shape range shifts and evolution across many similar avian taxa.

Thursday, 2 December 2021

EdwardsLab at Australasian Evolution Society #AUSEVO2021

Look out for some interesting talks by Edwards Lab members at this year’s Australasian Evolution Society 2021 conference, starting today:

Thursday 2nd December: 3-minute talks | 1130-1230

Kelton Cheung - Analysis of mitochondrial DNA reveals significant genetic diversity in invasive Australian cane toads

Kelton Cheung, Mark Richardson, Richard Edwards & Lee Ann Rollins

Mitochondrial DNA haplotype patterns across the native and invaded ranges can reveal the history and evolutionary trajectory of invasions. The invasion of Australia by the cane toad (Rhinella marina) has accelerated as it expanded westward. Despite this success, previous studies reported no mitochondrial genetic diversity in Australia. Here, we assembled a complete mitochondrial reference genome and haplotyped toads (N=119) from the native range and two introduced populations (Hawai’i and Australia), using whole genome and RNA-seq data. The complete R. marina mitochondrial genome consists of 18,152 base pairs with no significant gene arrangement as compared to other bufonid species. Although native range genetic diversity was much higher than that of the introduced ranges, we identified 29 haplotypes in Australia. While we did find evidence of founder effects following introduction, our results suggest there is significant genetic diversity within Australia, which may assist adaptation and invasion success in this species.


Thursday 2nd December: Selection 1 | 1415-1430

Katarina Stuart - A genetic perspective on rapid adaptation in the globally invasive European starling (Sturnus vulgaris)

Abstract coming soon…


Thursday 2nd December: Biogeography and phylogenetics | 1530-1545

Richard Edwards - DepthKopy: copy number prediction using single-copy long-read depth profiles

Richard J Edwards, Stephanie H Chen, Katarina C Stuart, Mark M Tanaka & Jason G Bragg

Gene duplication, followed by functional divergence of gene copies, is a fundamental component of genetic adaption and the evolution of novelty. Increasingly, evolutionary studies make use of the ever-expanding number of high-quality genome assemblies to characterise patterns of gene gain and loss. However, even reference genomes of the highest quality can experience assembly errors at tandemly repeated gene loci, resulting in incorrect inference of gene duplication patterns. Here, we present DepthKopy (https://github.com/slimsuite/depthkopy), which estimates the copy number for a gene, region or sequence of interest, using an estimate of single-copy sequencing depth derived from complete BUSCO genes. This is useful for identifying haplotigs, and collapsed repeat regions during genome assembly curation. Critically, for evolutionary studies of gene families, DepthKopy can identify genes for which the number of genes in the assembly does not seem to match the genome.


Friday 3rd December: Climate change and temperature | 1330-1345

Collin Ahrens - Genomic constraints of drought adaptation

Abstract coming soon…

Wednesday, 6 October 2021

Edwards Lab at Genetics Society of AustralAsia 2021 #GSAA21

Look out for some interesting genomics talks by Edwards Lab members at this year’s Genetics Society of AustralAsia 2021 conference, which started today. Congratulation to Stephanie for winning the Spencer Smith-White Travel Award (shame about the lack of travel!), Cadel for getting a lightning talk as an Honours student. And a shout out to Kat, who is one of the conference organisers.

Thursday 7th October: Genomics and Transcriptomics Session | 1:30-2:00 (Lightning talks)

Cadel Watson - dedUCE: efficient identification of Ultraconserved Elements from multiple genomes

Cadel Watson, Mitchell J. Cummins, Yasir Kusay, Maxine Halbheer, Eric Urng, John S. Mattick and Richard J. Edwards

Ultraconserved elements (UCEs) are DNA sequences which are extremely conserved and found almost unchanged in the genomes of multiple, divergent species [1]. UCEs have been found in a wide variety of organisms, including mammals, fish, insects, birds, and plants. Whilst the evidence suggests that that they are the result of natural selection, indicating biological importance, their function has thus far proven elusive [2]. The recent (and ongoing) explosion in the quality and quantity of reference genomes across multiple taxa provides new opportunities for investigating the prevalence, evolution and role of UCEs. However, the field is hampered by a lack of fast and resource-efficient algorithms to identify UCEs. Furthermore, common alignment-based algorithms fail to identify non-syntenic UCEs.

Here, we present dedUCE, a novel tool for identifying all UCEs in a set of genomes. dedUCE uses a hash-based algorithm to rapidly identify core UCE kmers that are shared by multiple genomes, before extending and merging candidates into a final comprehensive but non-redundant set of UCEs. dedUCE can support UCEs appearing out-of-order due to genetic rearrangements and/or assembly artefacts, and is able to return UCEs with inexact homology. Stringency can be controlled by parameters controlling the length, support (number of genomes) and required sequence identity. Preliminary results show that dedUCE can identify all UCEs in a group of 40 mammalian genomes in 8 hours on a 16-core machine, which is orders of magnitude faster than previous algorithms. Applications of dedUCE will be discussed, including improving the definition of UCEs, and making use of UCE content to assess genome assembly completeness.

  1. Gill Bejerano, Michael Pheasant, Igor Makunin, Stuart Stephen, W. James Kent, John S. Mattick, and David Haussler (2004). Ultraconserved El- ements in the Human Genome. Science, 304(5675):1321–1325.

  2. Konstantinos Kritsas, Samuel E. Wuest, Daniel Hupalo, Andrew D. Kern, Thomas Wicker, and Ueli Grossniklaus (2012). Computational analysis and char- acterization of UCE-like elements (ULEs) in plant genomes. Genome Research, 22(12):2455–2466.


Friday 8th October: Ecological and Evolutionary Genetics Session | 10:45-11:00

Katarina Stuart - A genetic perspective on rapid adaptation in the globally invasive European starling (Sturnus vulgaris)

Stuart KC, Sherwin WB, Edwards RJ & Rollins LA

Few invasive birds are as globally successful or as well-studied as the common starling (Sturnus vulgaris). Native to the Palaearctic, the starling has been a prolific invader in North and South America, southern Africa, Australia, and The Pacific Islands, while facing declines in excess of 50% in in some native regions. Starlings present an invaluable opportunity to test predictions about the evolutionary trajectory of invasive populations, and gain insight into genetic shifts in response to anthropogenic alteration and climate change. My research focuses primarily on the invasive European starling population in Australia and aims to investigate the genetics underlying their evolution, using a range of genomic approaches. Through historic museum sample sequencing, I examine single nucleotide polymorphism variations shifts between the native range and Australia, and find parallel selection on both continents, possibly resulting from common global selective forces such as exposure to pollutants and carbohydrate exposure. I further examine matched genetic, morphological, and environmental data to reveal patterns of heritability and plasticity across ecologically significant phenotypic traits, revealing that elevation, as well as rainfall and temperature variability plays an important role in shaping morphology and genetics. Finally, I investigated patterns of structural variants, to uncover evolutionarily significant large-scale genetic variants across a global data set, and more specifically characterise their role in rapid starling adaptation across the entirety of the Australian range. Overall, my research seeks to better understand mechanisms and patterns of genetic change within this species, which may be used to inform invasion or native range management. More broadly, this evolutionary research into the starling provide an important perspective on the role of rapid evolution in invasive species persistence, and the global pressures that may shape range shifts and evolution across many similar avian taxa.


Friday 8th October: Spencer Smith-White Travel Award recipient | 1:15-1:30

Stephanie Chen - Genomics of speciation and introgression: insights from waratah (Telopea spp.) as a model clade

Telopea is an eastern Australian genus of five species of long-lived shrubs in the family Proteaceae. Previous work has characterised population structure and patterns of introgression between Telopea species. These studies were performed using a limited set of genetic markers, but point to the great potential of waratah as a model clade for understanding the processes of divergence, environmental adaptation and speciation, when enhanced by a genome-wide perspective enabled by a reference genome. However, few Proteaceae genomes and no waratah genomes are available. We assembled the first chromosome-level reference genome for T. speciosissima (New South Wales waratah; 2n = 22) using Nanopore long-reads, 10x Chromium linked-reads and Hi-C data. The assembly spans 823 Mb, representing 93.9 % of the estimated genome size, with a scaffold N50 of 69.1 Mb and 91.3 % of complete Embryophyta universal single-copy orthologs (BUSCOs) are present. We examined the evolutionary dynamics of Telopea using the reference genome in conjunction with DArTseq (n = 244) and whole genome shotgun sequencing (n = 14) of each of the seven lineages; there are three lineages of T. speciosissima – coastal, upland, and southern. Here, I will discuss the population structure and demographic history of the genus. We also examined phylogenomic relationships and developed a scalable method of rapidly generating species trees from short-read data to maximise the recovery of informative data from genomic datasets. The waratah reference genome represents an important new genomic resource in Proteaceae to accelerate our understanding of the origins and evolutionary dynamics of the Australian flora.

Thursday, 8 April 2021

Transcript- and annotation-guided genome assembly of the European starling

Our starling genome paper is now available as a pre-print on bioRxiv! This was some great work by PhD student, Kat Stuart. Kat assembled a new Australian starling genome, using a combination of linked reads, low coverage long reads, and long-read PacBio iso-seq transcriptomics data. A second Illumina assembly of a North American group is also presented. As we saw with our Basenji genome paper, having two (or more) genomes from a species can be really useful for disentangling real difference from assembly artefacts. (No assembly is perfect!)

This paper is a great example of how a bit of TLC and imagination can get the most out of data produced with a limited budget. We were unable to get deep long-read sequencing this time, but instead show the additional power that long-read full-length transcriptome data can provide in assembling a genome - above and beyond the annotation.

This paper also officially describes a couple of genomics tools from the lab BUSCOMP has been in the works for some time, and this paper updates previous results to BUSCO v5 analysis and confirms our previous snake results using starling-derived test data. BUSCO is a powerful and popular tool that estimates genome completeness using gene prediction and curated models of single-copy protein orthologues. However, we demonstrate how results can be counterintuitive: adding/removing scaffolds can alter BUSCO predictions elsewhere in the assembly, while low sequence quality may reduce “completeness” scores and miss genes that are present in the assembly. BUSCOMP (BUSCO Compilation and Comparison) (https://github.com/slimsuite/buscomp) complements BUSCO to identify/overcome these issues by compiling a non-redundant set of the highest-scoring single-copy BUSCO complete sequences and re-searching these against assemblies for consistent completness scoring. SAAGA (https://github.com/slimsuite/saaga) is a new tool for annotation versus reference proteome comparisons. SAAGA can compare different annotations of the same assembly, or be combined with a lightweight annotation tool like GeMoMa to compare different assemblies of the same organism.


Stuart KC, Edwards RJ, Cheng Y, Warren WC, Burt DW, Sherwin WB, Hofmeister NR, Werner SJ, Ball GF, Bateson M, Brandley MC, Buchanan KL, Cassey P, Clayton DF, De Meyer T, Meddle SL & Rollins LA (preprint): Transcript- and annotation-guided genome assembly of the European starling. bioRxiv 2021.04.07.438753; doi: 10.1101/2021.04.07.438753. [*Joint first authors] [bioRxiv]

Abstract

The European starling, Sturnus vulgaris, is an ecologically significant, globally invasive avian species that is also suffering from a major decline in its native range. Here, we present the genome assembly and long-read transcriptome of an Australian-sourced European starling (S. vulgaris vAU), and a second North American genome (S. vulgaris vNA), as complementary reference genomes for population genetic and evolutionary characterisation. S. vulgaris vAU combined 10x Genomics linked-reads, low-coverage Nanopore sequencing, and PacBio Iso-Seq full-length transcript scaffolding to generate a 1050 Mb assembly on 1,628 scaffolds (72.5 Mb scaffold N50). Species-specific transcript mapping and gene annotation revealed high structural and functional completeness (94.6% BUSCO completeness). Further scaffolding against the high-quality zebra finch (Taeniopygia guttata) genome assigned 98.6% of the assembly to 32 putative nuclear chromosome scaffolds. Rapid, recent advances in sequencing technologies and bioinformatics software have highlighted the need for evidence-based assessment of assembly decisions on a case-by-case basis. Using S. vulgaris vAU, we demonstrate how the multifunctional use of PacBio Iso-Seq transcript data and complementary homology-based annotation of sequential assembly steps (assessed using a new tool, SAAGA) can be used to assess, inform, and validate assembly workflow decisions. We also highlight some counter-intuitive behaviour in traditional BUSCO metrics, and present BUSCOMP, a complementary tool for assembly comparison designed to be robust to differences in assembly size and base-calling quality. Finally, we present a second starling assembly, S. vulgaris vNA, to facilitate comparative analysis and global genomic research on this ecologically important species.

Friday, 4 December 2020

EdwardsLab at #AusEvo2020

If you missed his talk at ABACBS2020, Jack will be presenting today at the Australasian Evolution Society 2020 Conference about The role of gene duplication in the evolution of snake venoms. Two conference presentations in two weeks - not a bad way to prepare for your Honours viva post-submission. Well done, Jack!

Also, Kat Stuart will be presenting her work on invasive starlings in Zoom 2 at 13:00 AEDT. Kat’s talks are always great to listen to:

  • Katarina Stuart: What drives invasion success? Using historical museum samples to examine evolution in an invasive passerine.

Wednesday, 25 November 2020

#ABACBS2020: Whole transcripts in genome assembly, annotation, and assessment: the draft genome assembly of the globally invasive common starling, Sturnus vulgaris

Katarina Stuart, Yuanyuan Cheng, Lee Rollins & Richard J. Edwards

Abstract

Native to the Palearctic, the common starling (Sturnus vulgaris) is a near-globally invasive passerine that has now colonised every continent barring Antarctica. Ecological interest in the species is two-fold – they are considered a conservation risk and crop pest within the invasive ranges, while recent decades have brought with them a worrying decline in starling numbers within historical native ranges. Despite the global interest in this species, there are still fundamental knowledge gaps in our understanding of the genetics and population differences of this species across their native and invasive range. We present the Australian S. vulgaris draft genome and transcriptome to be used as a reference for further investigation into evolutionary characterisation of this ecologically significant species. An initial 10x Genomics linked-read assembly was scaffolded and gap-filled with low coverage nanopore sequencing, complemented by PacBio Isoseq full-length transcript data. Isoseq data was incorporated into assembly scaffolding, annotation, and assembly assessment to inform workflow decisions. We produced a draft assembly with a scaffold N50 size of 72.5 Mb, and assess this alongside a North American S. vulgaris draft genome, previously assembled from Illumina data. Lastly, we use these different reference genomes, alongside a non-scaffolded version of the Australian S. vulgaris genome to assess how choice of reference genome affects common population genetic downstream analysis using a global whole genome resequencing data set.

Tuesday, 10 December 2019

Edwards Lab at GIW/ABACBS2019

The Edwards Lab and close affiliates have five posters at GIW-ABACBS2019 this year. Come and check them out during the two poster sesstions:


Poster #26: Comparative performance of long-read whole genome assembly tools in diploid eukaryotes

Åsa Pérez-Bercoff, Paris Thompson & Richard J. Edwards [PDF]

As long-read (single molecule) sequencing from Pacific Biosciences (PacBio) and Oxford Nanopore Technologies (ONT) is getting more common, more long-read assemblers are emerging. Whilst a few benchmarking studies have been performed, there is no clear “best” assembly tool, and the choice is often a combination of compute resource availability and anecdotal reports of relative performance in similar organisms. Here, we have de novo assembled PacBio long-read sequencing data from 13 diploid Saccharomyces cerevisiae (baker’s yeast) strains (11 diploid, 1 haploid and 1 tetraploid) using four different assemblers: Canu v1.8, Flye v2.4.2, WTDBG2 (a.k.a. Redbean) v2.4 and Ra v20181211. Assembly performance statistics have been generated by comparing assemblies to the reference yeast genome (SGD R64.2.1) using QUAST v5.0.2, and rating each assembly for accuracy, coverage, and contiguity.

The latest generation of assembly tools decide how to process reads based on a genome size parameter. For heterozygous diploid organisms, it is not clear whether this should be the haploid or diploid genome size, or something in-between based on heterozygosity. In this study, we used three genome size settings: haploid (13.1Mb), diploid (26.2Mb) and half-way in-between (19.65Mb). This will enable us to establish how the genome size parameter influences the quality of the resulting genome assembly.


Poster #33: Estimating genome size using long read depth profiles and single copy regions of draft genome assemblies

Timothy G Amos, Ziying Zhang & Richard J Edwards [PDF]

Estimating the size of a eukaryote genome is a fundamental task in genome assembly. As well as informing decisions on sequencing technology and depth, greater accuracy in genome size prediction can assist in assessing the completeness and duplication of a genome assembly. Genome sizes can be estimated through both experimental and genome sequencing approaches. Experimental methods include densitometry, flow cytometry analysis of stained nuclei and quantitative PCR (qPCR) of single copy genes. Genome sequencing approaches include short-read k-mer distributions. However, these methods can give variable results, and are prone to inaccurate predictions for genomes with abundant repetitive sequences.

Here, we show that the read depth of long reads in a draft genome can be used to provide a relatively accurate estimate of genome size across various model organisms. In general, diploid genome assemblies will consist of regions of haploid, diploid and incorrect (or multi-copy organelle) depth. Repeats and assembly errors are likely to be highly variable in terms of genome vs assembly copy number and, therefore, coverage. The modal read depth of conserved single copy orthologues should therefore approximate the sequencing depth of the input data, which can be extrapolated to estimate genome size. We found that our method can provide comparable, if not more accurate, estimates than short-read k-mer distributions.


Poster #92: Using genomics to reveal drivers of invasion success

Katarina Stuart, Lee Ann Rollins, William Sherwin, Richard Edwards, Natalie Hofmeister & Yuanyuan Cheng [PDF]

Invasive species are a global concern due to their negative impacts on the economy and local ecosystems. However, well-documented invasions provide a useful system in which to pose biologically interesting questions regarding short time scale evolution. Answering these questions will further our knowledge of evolutionary mechanisms, as well as inform specific management strategies for the invasive population. The European starling (Sturnus vulgaris) is a global pest that was introduced into Australia’s south-eastern states in the 1860’s and has since greatly expanded its range. Previous research on multiple introduced starling populations has demonstrated that their morphology has undergone subtle shifts following colonisation. My research applies a range of sequencing techniques to investigate genomic variation across Australia’s starling population. We are combining long-, short- and linked-read whole genome sequencing to assemble a high-quality starling reference genome. This genome will be annotated using Iso-seq long-read (PacBio) whole transcriptome sequencing in order to properly identify putatively evolving loci. Population sequencing data (whole genome and DArTSeq) will be mapped onto this reference to identify functional SNPs and reveal potential drivers of the rapid phenotypic divergences across the Australian population.


Poster #93: Advancing genomic resources for myrtle rust research and management

Stephanie Chen, Jason Bragg & Richard Edwards [PDF]

Myrtle rust is a plant disease caused by an invasive fungal pathogen (Austropuccinia psidii) first detected in Australia in 2010. Over 350 native species from the family Myrtaceae, which includes eucalypts, paperbarks, tea-trees, and lillipillies, are known hosts. Detailed genetic information needed for an effective and coordinated response encompassing conservation management is lacking. We performed de novo genome assembly and annotation for two species which exhibit a spectrum of resistance – Rhodamnia argentea (malletwood) and Syzygium oleosum (blue lilly pilly). Reference genomes have been generated using 10X linked reads, and are being supplemented with long reads and a hybrid assembly approach. These genomes will complement reduced representation sequencing (DArTseq) and rust resistance assays from different genotypes across the landscape. Together, these data will facilitate the characterisation of genetic structure and rust resistance across space as well as within and among species. This research is crucial for the management of at-risk species through optimising methods of improving disease resistance and adaptation to future climates in addition to increasing our understanding of myrtle rust which is a pressing concern to native biodiversity.


Poster #168: BUSCOMP: BUSCO Compilation and Comparison for Assessing Completeness in Multiple Genome Assemblies

Richard J. Edwards [PDF]

Advances in DNA sequencing technology and bioinformatics tools have placed de novo genome assembly of complex organisms firmly in the domain of individual labs and small consortia. Nevertheless, the assemblies produced are often fragmented and incomplete. Optimal assembly depends on the size, repeat landscape, ploidy and heterozygosity of the genome, which are often unknown. It is therefore common practice to try multiple strategies, and there is a bottleneck in assessing and comparing assemblies.

BUSCO [1] is a powerful and popular tool that estimates genome completeness using gene prediction and curated models of single-copy protein orthologues. However, results can be counterintuitive: adding/removing scaffolds can alter BUSCO predictions elsewhere in the assembly, while low sequence quality may reduce “completeness” scores and miss genes that are present in the assembly [2].

BUSCOMP (BUSCO Compilation and Comparison) complements BUSCO to identify/overcome these issues. BUSCOMP compiles a non-redundant set of the highest-scoring single-copy BUSCO complete sequences, rapidly searches these against assemblies, and robustly re-rates genes as Complete (Single/Duplicated), Fragmented/Partial or Missing. On test data from three organisms (yeast, cane toad and mainland tiger snake), BUSCOMP (1) gives consistent results when re-running the same assembly, (2) is not affected by adding or removing non-BUSCO-containing scaffolds, and (3) is minimally affected by assembly quality. This makes BUSCOMP ideal to run alongside BUSCO when trying to compare and rank genome assemblies, even in the absence of error-correction.

Available at: https://github.com/slimsuite/buscomp.

  1. Simão FA et al. (2015) Bioinformatics 31:3210–3212
  2. Edwards RJ et al. (2018) GigaScience 7:giy095

Wednesday, 28 November 2018

EdwardsLab at #ABACBS2018

For those who missed it, there’s a (slightly old) poster version of my ABACBS 2018 talk - Sequencing snakes: Pseudodiploid pseudo-long-read whole genome sequencing and assembly of Pseudonaja textilis (eastern brown snake) and Notechis scutatus (mainland tiger snake). If anything in the talk (except the repeat stuff) looks useful to you, this is a citeable poster:

Edwards RJ et al. Pseudodiploid pseudo-long-read whole genome sequencing and assembly of Pseudonaja textilis (eastern brown snake) and Notechis scutatus (mainland tiger snake) [version 1; not peer reviewed]. F1000Research 2018, 7:753 (poster) (doi: 10.7490/f1000research.1115550.1)

We’re still developing the genome size prediction and BUSCO comparison/compilation tools, so get in touch if either of these look useful to you.

ABACBS2018 Posters

We have three lab posters in Poster session 2 this morning:

  • Poster #16. Ã…sa Pérez-Bercoff, Using structural variant detection to resolve difficult regions of a genome assembly.

  • Poster #21. Kirsti Paulsen, Optimising intrinsic protein disorder prediction for short linear motif discovery.

  • Poster #26. Katarina Stuart, Evolution in invasive populations: using genomics to reveal drivers of invasion success in the Australian European starling (Sturnus vulgaris) introduction across Australia.

Also check out the posters of our UNSW neighbours from the Wilkins lab:

  • Poster #29. Chi Nam Ignatius (Igy) Pang, Benchmarking Protein Correlation Profiling datasets against reference protein complexes: case studies in S. cerevisiae.

  • Poster #44. Susan Corley, QuantSeq 3’ sequencing paired with Salmon quantification provides a fast reliable approach for high throughput transcriptomic analysis.

  • Poster #49. Xabier Vázquez-Campos, OTUreporter: an automated pipeline for the analysis and report of amplicon sequencing data.

Monday, 5 February 2018

Katarina Stuart (PhD Student)

Katarina Stuart completed her undergrad at the University of Sydney, completing an honours thesis on the evolution and trait plasticity of the invasive cane toad.

She commenced her PhD at UNSW in February 2018 under the primary supervision of Lee Ann Rollins. Her thesis project aims to use genomics to investigate evolution in the European Starlings (Sturnus vulgaris). The global nature of starling invasions provides an important opportunity to examine evolutionary questions about invasion success, and how this is linked to similarities or emerging differences in their genome post colonisation and during range expansion. Major components of her thesis involve updating the starling genome, and exploring the population genetics of Australia’s starling invasion in relation to their native range counterparts, both modern and historic.

Katarina’s general interest are in invasive species, and their dynamic evolutionary change over geographical and temporal ranges.

[LinkedIn | Twitter]