Showing posts with label cane toad. Show all posts
Showing posts with label cane toad. Show all posts

Saturday, 16 November 2024

Repeat-Rich Regions Cause False-Positive Detection of NUMTs: A Case Study in Amphibians Using an Improved Cane Toad Reference Genome

Version 3* of the cane toad reference genome, aRhiMar1.3 is now officially out and published in Genome Biology and Evolution. [*It’s only the second published genome but for internal reasons the original draft genome was version 2!].

The focus of the paper itself is confirming the lack of Nuclear Mitochondrial fragments (a.k.a. NUMTs) in the cane toad genome, which could impact whole-mitogenome analysis of genetic diversity in cane toads. We were pretty surprised when we first looked for NUMTs in the cane toad genome and could not find any! The draft genome is pretty drafty, especially in terms of missing repetitive regions, so an updated long-read assembly was important to rule out a false negative result.

The new genome is in much better shape, with an extra 922 Mbp (>95% repeats) and a 15x increase in scaffold N50 (2.5 Mbp). (My biggest regret with the original paper was not sticking to my guns and including an early DepthSizer genome size estimate of 3.5 Mbp, which has subsequently turned out to be correct.) The cane toad genome remains a tough nut to crack, and we didn’t quite reach the magic 1Mb contig N50 (860 kb), but the functional completeness was markedly improved and we are pretty confident that the continued absence of NUMT detection is a real phenomenom and does not simply reflect technical limitations.

Watch this space for a chromosome-level cane toad genome, which is still in the works.

Cheung K, Rollins LA, Hammond JM, Barton K, Ferguson JM, Eyck HJF, Shine R & Edwards RJ (2024): Repeat-rich regions cause false positive detection of NUMTs - a case study in amphibians using an improved cane toad reference genome. Genome Biology and Evolution evae246. [Gen Biol Evol] [bioRxiv] [PubMed]

Mitochondrial DNA (mtDNA) has been widely used in genetics research for decades. Contamination from nuclear DNA of mitochondrial origin (NUMTs) can confound studies of phylogenetic relationships and mtDNA heteroplasmy. Homology searches with mtDNA are widely used to detect NUMTs in the nuclear genome. Nevertheless, false-positive detection of NUMTs is common when handling repeat-rich sequences, while fragmented genomes might result in missing true NUMTs. In this study, we investigated different NUMT detection methods and how the quality of the genome assembly affects them. We presented an improved nuclear genome assembly (aRhiMar1.3) of the invasive cane toad (Rhinella marina) with additional long-read Nanopore and 10× linked-read sequencing. The final assembly was 3.47 Gb in length with 91.3% of tetrapod universal single-copy orthologs (n = 5,310), indicating the gene-containing regions were well assembled. We used 3 complementary methods (NUMTFinder, dinumt, and PALMER) to study the NUMT landscape of the cane toad genome. All 3 methods yielded consistent results, showing very few NUMTs in the cane toad genome. Furthermore, we expanded NUMT detection analyses to other amphibians and confirmed a weak relationship between genome size and the number of NUMTs present in the nuclear genome. Amphibians are repeat-rich, and we show that the number of NUMTs found in highly repetitive genomes is prone to inflation when using homology-based detection without filters. Together, this study provides an exemplar of how to robustly identify NUMTs in complex genomes when confounding effects on mtDNA analyses are a concern.

Thursday, 5 September 2024

Origin and maintenance of large ribosomal RNA gene repeat size in mammals

Our latest paper is out as a Featured Article in the journal Genetics, featuring ONT from both the cane toad and BABS Genome snake genomes. This paper looks at how ribosomal RNA gene repeats (a.k.a. rDNA repeats) have evolved in vertebrates to expand in size in mammals. For something so fundamental to the function of an organism - literally every process of every cell ultimately relies on rRNA - there is surprising diversity. These regions are traditionally hard to assemble with short reads, and still provide challenges for long-read assemblies, so the new era of high-quality long-read assemblies is likely to reveal a lot about their evolution.

  • Macdonald E, Whibley A, Waters PD, Patel H, Edwards RJ & Ganley ARD (2024): Origin and maintenance of large ribosomal RNA gene repeat size in mammals. Genetics 228(1): iyae121 [Genetics] [PubMed]

Abstract

The genes encoding ribosomal RNA are highly conserved across life and in almost all eukaryotes are present in large tandem repeat arrays called the rDNA. rDNA repeat unit size is conserved across most eukaryotes but has expanded dramatically in mammals, principally through the expansion of the intergenic spacer region that separates adjacent rRNA coding regions. Here, we used long-read sequence data from representatives of the major amniote lineages to determine where in amniote evolution rDNA unit size increased. We find that amniote rDNA unit sizes fall into two narrow size classes: “normal” (∼11–20 kb) in all amniotes except monotreme, marsupial, and eutherian mammals, which have “large” (∼35–45 kb) sizes. We confirm that increases in intergenic spacer length explain much of this mammalian size increase. However, in stark contrast to the uniformity of mammalian rDNA unit size, mammalian intergenic spacers differ greatly in sequence. These results suggest a large increase in intergenic spacer size occurred in a mammalian ancestor and has been maintained despite substantial sequence changes over the course of mammalian evolution. This points to a previously unrecognized constraint on the length of the intergenic spacer, a region that was thought to be largely neutral. We finish by speculating on possible causes of this constraint.

Thursday, 7 March 2024

Is developmental plasticity triggered by DNA methylation changes in the invasive cane toad (Rhinella marina)?

Second cane toad paper of the week! This time, we're revisiting invasive epigenomics.

Yagound B, Sarma RR, Edwards RJ, Richardson MF, Rodriguez Lopez CM, Crossland MR, Brown GP, DeVore JL, Shine R & Rollins LA (2024): Is developmental plasticity triggered by DNA methylation changes in the invasive cane toad (Rhinella marina)? Ecology and Evolution 14:e11127. [Ecol Evol] [PubMed] [bioRxiv]

Many organisms can adjust their development according to environmental conditions, including the presence of conspecifics. Although this developmental plasticity is common in amphibians, its underlying molecular mechanisms remain largely unknown. Exposure during development to either ‘cannibal cues’ from older conspecifics, or ‘alarm cues’ from injured conspecifics, causes reduced growth and survival in cane toad (Rhinella marina) tadpoles. Epigenetic modifications, such as changes in DNA methylation patterns, are a plausible mechanism underlying these developmental plastic responses. Here we tested this hypothesis, and asked whether cannibal cues and alarm cues trigger the same DNA methylation changes in developing cane toads. We found that exposure to both cannibal cues and alarm cues was associated with local changes in DNA methylation patterns. These DNA methylation changes affected genes putatively involved in developmental processes, but in different genomic regions for different conspecific-derived cues. Genetic background explains most of the epigenetic variation among individuals. Overall, the molecular mechanisms triggered by exposure to cannibal cues seem to differ from those triggered by alarm cues. Studies linking epigenetic modifications to transcriptional activity are needed to clarify the proximate mechanisms that regulate developmental plasticity in cane toads.

Monday, 4 March 2024

Whole-mitogenome analysis unveils previously undescribed genetic diversity in cane toads across their invasion trajectory

Congratulations to Kelton Cheung for getting her first PhD paper out. This one has been a long time brewing and involved quite a lot of data wrangling, but we got there in the end. Invasive cane toads might be a little more complex than we thought.

Cheung K, Amos TG, Shine R, DeVore JL, S Ducatez S, Edwards RJ & Rollins LA (2024): Whole-mitogenome analysis unveils previously undescribed genetic diversity in cane toads across their invasion trajectory. Ecology and Evolution 14:e11115. [Ecol Evol] [PubMed] [bioRxiv]

Invasive species offer insights into rapid adaptation to novel environments. The iconic cane toad (Rhinella marina) is an excellent model for studying rapid adaptation during invasion. Previous research using the mitochondrial NADH dehydrogenase 3 (ND3) gene in Hawai’ian and Australian invasive populations found a single haplotype, indicating an extreme genetic bottleneck following introduction. Nuclear genetic diversity also exhibited reductions across the genome in these two populations. Here, we investigated the mitochondrial genomics of cane toads across this invasion trajectory. We created the first reference mitochondrial genome for this species using long-read sequence data. We combined whole-genome resequencing data of 15 toads with published transcriptomic data of 125 individuals to construct nearly complete mitochondrial genomes from the native (French Guiana) and introduced (Hawai’i and Australia) ranges for population genomic analyses. In agreement with previous investigations of these populations, we identified genetic bottlenecks in both Hawai’ian and Australian introduced populations, alongside evidence of population expansion in the invasive ranges. Although mitochondrial genetic diversity in introduced populations was reduced, our results revealed that it had been underestimated: we identified 45 mitochondrial haplotypes in Hawai’ian and Australian samples, none of which were found in the native range. Additionally, we identified two distinct groups of haplotypes from the native range, separated by a minimum of 110 base pairs (0.6%). These findings enhance our understanding of how invasion has shaped the genetic landscape of this species.

Thursday, 2 December 2021

EdwardsLab at Australasian Evolution Society #AUSEVO2021

Look out for some interesting talks by Edwards Lab members at this year’s Australasian Evolution Society 2021 conference, starting today:

Thursday 2nd December: 3-minute talks | 1130-1230

Kelton Cheung - Analysis of mitochondrial DNA reveals significant genetic diversity in invasive Australian cane toads

Kelton Cheung, Mark Richardson, Richard Edwards & Lee Ann Rollins

Mitochondrial DNA haplotype patterns across the native and invaded ranges can reveal the history and evolutionary trajectory of invasions. The invasion of Australia by the cane toad (Rhinella marina) has accelerated as it expanded westward. Despite this success, previous studies reported no mitochondrial genetic diversity in Australia. Here, we assembled a complete mitochondrial reference genome and haplotyped toads (N=119) from the native range and two introduced populations (Hawai’i and Australia), using whole genome and RNA-seq data. The complete R. marina mitochondrial genome consists of 18,152 base pairs with no significant gene arrangement as compared to other bufonid species. Although native range genetic diversity was much higher than that of the introduced ranges, we identified 29 haplotypes in Australia. While we did find evidence of founder effects following introduction, our results suggest there is significant genetic diversity within Australia, which may assist adaptation and invasion success in this species.


Thursday 2nd December: Selection 1 | 1415-1430

Katarina Stuart - A genetic perspective on rapid adaptation in the globally invasive European starling (Sturnus vulgaris)

Abstract coming soon…


Thursday 2nd December: Biogeography and phylogenetics | 1530-1545

Richard Edwards - DepthKopy: copy number prediction using single-copy long-read depth profiles

Richard J Edwards, Stephanie H Chen, Katarina C Stuart, Mark M Tanaka & Jason G Bragg

Gene duplication, followed by functional divergence of gene copies, is a fundamental component of genetic adaption and the evolution of novelty. Increasingly, evolutionary studies make use of the ever-expanding number of high-quality genome assemblies to characterise patterns of gene gain and loss. However, even reference genomes of the highest quality can experience assembly errors at tandemly repeated gene loci, resulting in incorrect inference of gene duplication patterns. Here, we present DepthKopy (https://github.com/slimsuite/depthkopy), which estimates the copy number for a gene, region or sequence of interest, using an estimate of single-copy sequencing depth derived from complete BUSCO genes. This is useful for identifying haplotigs, and collapsed repeat regions during genome assembly curation. Critically, for evolutionary studies of gene families, DepthKopy can identify genes for which the number of genes in the assembly does not seem to match the genome.


Friday 3rd December: Climate change and temperature | 1330-1345

Collin Ahrens - Genomic constraints of drought adaptation

Abstract coming soon…

Monday, 19 April 2021

Intergenerational effects of manipulating DNA methylation in the early life of an iconic invader

Another cane toad paper has hit the shelves! This is another paper from our ongoing collaboration with Lee Ann Rollins and her great team of invasion biologists and molecular ecologists. This paper once again uses our draft cane toad genome* and builds on the previous cane toad methylation analysis by PhD student Roshmi Sarma to look at some really interesting intergenerational effects. [*Genome update coming soon - watch this space!]

Could this be epigenetic inheritance? Maybe. But it could also be some kind of parental germline thing. Either way, it’s a fascinating result and further evidence to support the fact that whilst our genome may establish our genetic potential, it does not control our destiny.

And if you want to know what it takes to identify effects like this, check out the mind-boggling experiment design Figure! (PDF available on request!)


This article is part of the theme issue ‘How does epigenetics influence the course of evolution?’

Sarma RR, Crossland MR, Eyck HJF, Edwards RJ, DeVore JL, Cocomazzo M, Zhou J, Brown GP, Shine R & Rollins LA (2021): Intergenerational effects of manipulating DNA methylation in the early life of an iconic invader. Philosophical Transactions of the Royal Society B 376:20200125. [Phil Trans Roy Soc B] [PubMed]

Abstract

In response to novel environments, invasive populations often evolve rapidly. Standing genetic variation is an important predictor of evolutionary response but epigenetic variation may also play a role. Here, we use an iconic invader, the cane toad (Rhinella marina), to investigate how manipulating epigenetic status affects phenotypic traits. We collected wild toads from across Australia, bred them, and experimentally manipulated DNA methylation of the subsequent two generations (G1, G2) through exposure to the DNA methylation inhibitor zebularine and/or conspecific tadpole alarm cues. Direct exposure to alarm cues (an indicator of predation risk) increased the potency of G2 tadpole chemical cues, but this was accompanied by reductions in survival. Exposure to alarm cues during G1 also increased the potency of G2 tadpole cues, indicating intergenerational plasticity in this inducible defence. In addition, the negative effects of alarm cues on tadpole viability (i.e. the costs of producing the inducible defence) were minimized in the second generation. Exposure to zebularine during G1 induced similar intergenerational effects, suggesting a role for alteration in DNA methylation. Accordingly, we identified intergenerational shifts in DNA methylation at some loci in response to alarm cue exposure. Substantial demethylation occurred within the sodium channel epithelial 1 subunit gamma gene (SCNN1G) in alarm cue exposed individuals and their offspring. This gene is a key to the regulation of sodium in epithelial cells and may help to maintain the protective epidermal barrier. These data suggest that early life experiences of tadpoles induce intergenerational effects through epigenetic mechanisms, which enhance larval fitness.

Tuesday, 23 June 2020

Do epigenetic changes drive corticosterone responses to alarm cues in larvae of an invasive amphibian?

Our latest cane toad paper is now online at Integrative and Comparative Biology, using our draft cane toad genome as the reference for differential methylation analysis. (Thanks to coronavirus delays, the updated cane toad genome was not quite ready for this analysis but watch this space!)

Sarma RR, Edwards RJ, Crino OL, Eyck HJF, Waters PD, Crossland MR, Shine R & Rollins LA (accepted): Do epigenetic changes drive corticosterone responses to alarm cues in larvae of an invasive amphibian? Integrative and Comparative Biology icaa082

Abstract

The developmental environment can exert powerful effects on animal phenotype. Recently epigenetic modifications have emerged as one mechanism that can modulate developmentally plastic responses to environmental variability. For example, the DNA methylation profile at promoters of hormone receptor genes can affect their expression and patterns of hormone release. Across taxonomic groups, epigenetic alterations have been linked to changes in glucocorticoid (GC) physiology. GCs are metabolic hormones that influence growth, development, transitions between life-history stages, and thus fitness. To date, relatively few studies have examined epigenetic effects on phenotypic traits in wild animals, especially in amphibians. Here, we examined the effects of exposure to predation threat and experimentally manipulated DNA methylation on corticosterone (CORT) levels in tadpoles and metamorphs of the invasive cane toad (Rhinella marina). We included offspring of toads sampled from populations across the species’ Australian range. In these animals, exposure to chemical cues from injured conspecifics induces shifts in developmental trajectories, putatively as an adaptive response that lessens vulnerability to predation. We exposed tadpoles to these alarm cues, and measured changes in DNA methylation and CORT levels, both of which are mechanisms that have been implicated in the control of phenotypically plastic responses in tadpoles. To test the idea that DNA methylation drives shifts in GC physiology, we also experimentally manipulated methylation levels with the drug zebularine. We found differentially methylated regions between control tadpoles and their full-siblings exposed to alarm cues, zebularine or both treatments. However, the effects of these manipulations on methylation patterns were weaker than clutch (e.g. genetic, maternal, etc.) effects. CORT levels were higher in larval cane toads exposed to alarm cues and zebularine. We found little evidence of changes in DNA methylation across the glucocorticoid receptor gene (NR3C1) promoter region in response to alarm cue or zebularine exposure. In both alarm cue and zebularine-exposed individuals, we found differentially methylated DNA in the suppressor of cytokine signaling 3 gene (SOCS3), which may be involved in predator avoidance behavior. In total, our data reveal that alarm cues have significant impacts on tadpole physiology, but show only weak links between DNA methylation and CORT levels. We also identify genes containing differentially methylated regions in tadpoles exposed to alarm cues and zebularine, particularly in range-edge populations, that warrant further investigation.

Friday, 31 January 2020

Research snapshot - January 2020

One of the most important, interesting and challenging questions in biology is how new traits evolve at the molecular level. My lab employs sequence analysis techniques to interrogate DNA and protein sequences for the signals left behind by evolution. We are a bioinformatics lab but like to incorporate bench/field data through collaboration wherever possible.

Main Research

Building on a solid foundation of bioinformatics and evolutionary theory, we apply genomics, transcriptomics, proteomics and interactomics and systems analysis to understand complex biological systems. Core research activities can be broadly divided into two main themes:

  1. Evolutionary Genomics, with a focus on applying de novo genome assembly and population genomics to problems in ecology and biotechnology.

  2. Protein-protein interactions, with a focus on the prediction of short linear interaction motifs and their role in human health and disease, including host-pathogen interactions.

These are explored in more detail, below.

1. Evolutionary Genomics.

The main research focus of the lab is the exploitation of genomic and post-genomic data to understand biological function and adaptation to novel environments. We work closely with the Ramaciotti Centre for Genomics and are involved in numerous de novo whole genome sequencing and assembly projects, using short read (Illumina), long read (PacBio & Nanopore) and linked read (10x Chromium) sequencing. One of the biggest of these is leading the bioinformatics and assembly effort in a consortium to sequence the cane toad genome, and leading the BABS Genome project to sequence two iconic Australian snakes. We are a member of the Oz Mammals Genomics initiative, assisting with the sequencing and assembly of Australia’s unique marsupial fauna. In 2018, we were selected as part of a team to sequence the Waratah genome as part of the pilot phase for the new Genomics of Australian Plants initiative.

We enjoy bringing our bioinformatics to bear on a variety of collaborative research projects. Most notably, we have an ARC Linkage Grant with Microbiogen Pty Ltd to understand how a strain of Saccharomyces cerevisiae has evolved to efficiently use xylose as a sole carbon source: something vital for second generation biofuel production that wild yeast cannot do. We are combining comparative genomics, evolutionary genetics, RNA-Seq transcriptomics, and competition assays to understand how the novel metabolism evolved. Through deep Illumina resequencing of evolving populations, and assembling reliable complete genomes of the founding ancestors, the ultimate goal is to trace how mutations have interacted with existing genetic variation during adaptive evolution. More recently, we have received an ARC Linkage Grant with the Royal Botanic Gardens and Domain Trust, to apply genomics approaches to the challenges of rainforest tree conservation in the face of climate change and invasive pathogens. We are also collaborating with industrial and academic partners to de novo sequence, assemble, annotate and interrogate the genomes of a selection of microbes with interesting metabolic abilities.

2. Short Linear Motifs (SLiMs).

Many protein-protein interactions are mediated by Short Linear Motifs (SLiMs): short stretches of proteins (5-15 amino acids long), of which only a few positions are critical to function. These motifs are vital for biological processes of fundamental importance, acting as ligands for molecular signalling, post-translational modifications and subcellular targeting. SLiMs have extremely compact protein interaction interfaces, generally encoded by less than 4 major affinity-/specificity-determining residues. Their small size enables high functional density and evolutionary plasticity, making them frequent products of convergent “ex nihilo” evolution. It also makes them challenging to identify, both experimentally and computationally.

A major focus of the lab is the computational prediction of SLiMs from protein sequences. This research originated with Rich’s postdoctoral research, during which he developed a sequence analysis methods for the rational design of biologically active short peptides. He subsequently developed SLiMDisc, one of the first algorithms for successfully predicting novel SLiMs from sequence data - and coined the term “SLiM” into the bargain. This subsequently lead to the development of SLiMFinder, the first SLiM prediction algorithm able to estimate the statistical significance of motif predictions. SLiMFinder greatly increased the reliability of predictions. SLiMFinder has since spawned a number of motif discovery tools and webservers and is still arguably the most successful SLiM prediction tool on benchmarking data. Methods are made available through the SLiMSuite bioinformatics package and webservers.

Current research is looking to develop these SLiM prediction tools further and apply them to important biological questions. Of particular interest is the molecular mimicry employed by viruses to interact with host proteins and the role of SLiMs in other diseases, such as cancer. Other work is concerned with the evolutionary dynamics of SLiMs within protein interaction networks.

OTHER RESEARCH PROJECTS

In addition to the main research in the lab, the lab has a number of interdisciplinary collaborative projects applying bioinformatics tools and molecular evolution theory to experimental biology, often using large genomic, transcriptomic and/or proteomic datasets. These projects often involve the development of bespoke bioinformatics pipelines and a number of open source bioinformatics tools have been generated as a result. Please see the Publications and Lab software pages for more detail, or get in touch if something catches your eye and you want to find out more. We frequently have small collaborations and/or undergraduate student research projects. Many of these are “on hold” waiting for the right person, or sometimes data, to come along. If you think that you have what it needs, get in touch!

Tuesday, 10 December 2019

Edwards Lab at GIW/ABACBS2019

The Edwards Lab and close affiliates have five posters at GIW-ABACBS2019 this year. Come and check them out during the two poster sesstions:


Poster #26: Comparative performance of long-read whole genome assembly tools in diploid eukaryotes

Åsa Pérez-Bercoff, Paris Thompson & Richard J. Edwards [PDF]

As long-read (single molecule) sequencing from Pacific Biosciences (PacBio) and Oxford Nanopore Technologies (ONT) is getting more common, more long-read assemblers are emerging. Whilst a few benchmarking studies have been performed, there is no clear “best” assembly tool, and the choice is often a combination of compute resource availability and anecdotal reports of relative performance in similar organisms. Here, we have de novo assembled PacBio long-read sequencing data from 13 diploid Saccharomyces cerevisiae (baker’s yeast) strains (11 diploid, 1 haploid and 1 tetraploid) using four different assemblers: Canu v1.8, Flye v2.4.2, WTDBG2 (a.k.a. Redbean) v2.4 and Ra v20181211. Assembly performance statistics have been generated by comparing assemblies to the reference yeast genome (SGD R64.2.1) using QUAST v5.0.2, and rating each assembly for accuracy, coverage, and contiguity.

The latest generation of assembly tools decide how to process reads based on a genome size parameter. For heterozygous diploid organisms, it is not clear whether this should be the haploid or diploid genome size, or something in-between based on heterozygosity. In this study, we used three genome size settings: haploid (13.1Mb), diploid (26.2Mb) and half-way in-between (19.65Mb). This will enable us to establish how the genome size parameter influences the quality of the resulting genome assembly.


Poster #33: Estimating genome size using long read depth profiles and single copy regions of draft genome assemblies

Timothy G Amos, Ziying Zhang & Richard J Edwards [PDF]

Estimating the size of a eukaryote genome is a fundamental task in genome assembly. As well as informing decisions on sequencing technology and depth, greater accuracy in genome size prediction can assist in assessing the completeness and duplication of a genome assembly. Genome sizes can be estimated through both experimental and genome sequencing approaches. Experimental methods include densitometry, flow cytometry analysis of stained nuclei and quantitative PCR (qPCR) of single copy genes. Genome sequencing approaches include short-read k-mer distributions. However, these methods can give variable results, and are prone to inaccurate predictions for genomes with abundant repetitive sequences.

Here, we show that the read depth of long reads in a draft genome can be used to provide a relatively accurate estimate of genome size across various model organisms. In general, diploid genome assemblies will consist of regions of haploid, diploid and incorrect (or multi-copy organelle) depth. Repeats and assembly errors are likely to be highly variable in terms of genome vs assembly copy number and, therefore, coverage. The modal read depth of conserved single copy orthologues should therefore approximate the sequencing depth of the input data, which can be extrapolated to estimate genome size. We found that our method can provide comparable, if not more accurate, estimates than short-read k-mer distributions.


Poster #92: Using genomics to reveal drivers of invasion success

Katarina Stuart, Lee Ann Rollins, William Sherwin, Richard Edwards, Natalie Hofmeister & Yuanyuan Cheng [PDF]

Invasive species are a global concern due to their negative impacts on the economy and local ecosystems. However, well-documented invasions provide a useful system in which to pose biologically interesting questions regarding short time scale evolution. Answering these questions will further our knowledge of evolutionary mechanisms, as well as inform specific management strategies for the invasive population. The European starling (Sturnus vulgaris) is a global pest that was introduced into Australia’s south-eastern states in the 1860’s and has since greatly expanded its range. Previous research on multiple introduced starling populations has demonstrated that their morphology has undergone subtle shifts following colonisation. My research applies a range of sequencing techniques to investigate genomic variation across Australia’s starling population. We are combining long-, short- and linked-read whole genome sequencing to assemble a high-quality starling reference genome. This genome will be annotated using Iso-seq long-read (PacBio) whole transcriptome sequencing in order to properly identify putatively evolving loci. Population sequencing data (whole genome and DArTSeq) will be mapped onto this reference to identify functional SNPs and reveal potential drivers of the rapid phenotypic divergences across the Australian population.


Poster #93: Advancing genomic resources for myrtle rust research and management

Stephanie Chen, Jason Bragg & Richard Edwards [PDF]

Myrtle rust is a plant disease caused by an invasive fungal pathogen (Austropuccinia psidii) first detected in Australia in 2010. Over 350 native species from the family Myrtaceae, which includes eucalypts, paperbarks, tea-trees, and lillipillies, are known hosts. Detailed genetic information needed for an effective and coordinated response encompassing conservation management is lacking. We performed de novo genome assembly and annotation for two species which exhibit a spectrum of resistance – Rhodamnia argentea (malletwood) and Syzygium oleosum (blue lilly pilly). Reference genomes have been generated using 10X linked reads, and are being supplemented with long reads and a hybrid assembly approach. These genomes will complement reduced representation sequencing (DArTseq) and rust resistance assays from different genotypes across the landscape. Together, these data will facilitate the characterisation of genetic structure and rust resistance across space as well as within and among species. This research is crucial for the management of at-risk species through optimising methods of improving disease resistance and adaptation to future climates in addition to increasing our understanding of myrtle rust which is a pressing concern to native biodiversity.


Poster #168: BUSCOMP: BUSCO Compilation and Comparison for Assessing Completeness in Multiple Genome Assemblies

Richard J. Edwards [PDF]

Advances in DNA sequencing technology and bioinformatics tools have placed de novo genome assembly of complex organisms firmly in the domain of individual labs and small consortia. Nevertheless, the assemblies produced are often fragmented and incomplete. Optimal assembly depends on the size, repeat landscape, ploidy and heterozygosity of the genome, which are often unknown. It is therefore common practice to try multiple strategies, and there is a bottleneck in assessing and comparing assemblies.

BUSCO [1] is a powerful and popular tool that estimates genome completeness using gene prediction and curated models of single-copy protein orthologues. However, results can be counterintuitive: adding/removing scaffolds can alter BUSCO predictions elsewhere in the assembly, while low sequence quality may reduce “completeness” scores and miss genes that are present in the assembly [2].

BUSCOMP (BUSCO Compilation and Comparison) complements BUSCO to identify/overcome these issues. BUSCOMP compiles a non-redundant set of the highest-scoring single-copy BUSCO complete sequences, rapidly searches these against assemblies, and robustly re-rates genes as Complete (Single/Duplicated), Fragmented/Partial or Missing. On test data from three organisms (yeast, cane toad and mainland tiger snake), BUSCOMP (1) gives consistent results when re-running the same assembly, (2) is not affected by adding or removing non-BUSCO-containing scaffolds, and (3) is minimally affected by assembly quality. This makes BUSCOMP ideal to run alongside BUSCO when trying to compare and rank genome assemblies, even in the absence of error-correction.

Available at: https://github.com/slimsuite/buscomp.

  1. Simão FA et al. (2015) Bioinformatics 31:3210–3212
  2. Edwards RJ et al. (2018) GigaScience 7:giy095

Friday, 5 July 2019

Research snapshot - July 2019

One of the most important, interesting and challenging questions in biology is how new traits evolve at the molecular level. My lab employs sequence analysis techniques to interrogate protein and DNA sequences for the signals left behind by evolution. We are a bioinformatics lab but like to incorporate bench data through collaboration wherever possible.

Main Research

The core research in the lab is broadly divided into two main themes:

1. Evolutionary Genomics.

Since moving to UNSW, a major focus of the lab has been the exploitation of genomic and post-genomic data to understand biological function and adaptation to novel environments. We work closely with the Ramaciotti Centre for Genomics and are involved in numerous de novo whole genome sequencing and assembly projects, using short read (Illumina), long read (PacBio & Nanopore) and linked read (10x Chromium) sequencing. The biggest of these is leading the bioinformatics and assembly effort in a consortium to sequence the cane toad genome, and leading the BABS Genome project to sequence two iconic Australian snakes. We are a member of the Oz Mammals Genomics initiative, assisting with the sequencing and assembly of Australia’s unique marsupial fauna. In 2018, we were selected as part of a team to sequence the Waratah genome as part of the pilot phase for the new Genomics of Australian Plants initiative.

We enjoy bringing our bioinformatics to bear on a variety of collaborative research projects. Most notably, we have an ARC Linkage Grant with Microbiogen Pty Ltd to understand how a strain of Saccharomyces cerevisiae has evolved to efficiently use xylose as a sole carbon source: something vital for second-generation biofuel production that wild yeast cannot do. We are combining comparative genomics, evolutionary genetics, RNA-Seq transcriptomics, and competition assays to understand how the novel metabolism evolved. Through deep Illumina resequencing of evolving populations, and assembling reliable complete genomes of the founding ancestors, the ultimate goal is to trace how mutations have interacted with existing genetic variation during adaptive evolution. More recently, we have received an ARC Linkage Grant with the Royal Botanic Gardens and Domain Trust, to apply genomics approaches to the challenges of rainforest tree conservation in the face of climate change and invasive pathogens. We are also collaborating with industrial and academic partners to de novo sequence, assemble, annotate and interrogate the genomes of a selection of microbes with interesting metabolic abilities.

2. Short Linear Motifs (SLiMs).

Many protein-protein interactions are mediated by Short Linear Motifs (SLiMs): short stretches of proteins (5-15 amino acids long), of which only a few positions are critical to function. These motifs are vital for biological processes of fundamental importance, acting as ligands for molecular signalling, post-translational modifications and subcellular targeting. SLiMs have extremely compact protein interaction interfaces, generally encoded by less than 4 major affinity-/specificity-determining residues. Their small size enables high functional density and evolutionary plasticity, making them frequent products of convergent “ex nihilo” evolution. It also makes them challenging to identify, both experimentally and computationally.

A major focus of the lab is the computational prediction of SLiMs from protein sequences. This research originated with Rich’s postdoctoral research, during which he developed a sequence analysis methods for the rational design of biologically active short peptides. He subsequently developed SLiMDisc, one of the first algorithms for successfully predicting novel SLiMs from sequence data - and coined the term “SLiM” into the bargain. This subsequently lead to the development of SLiMFinder, the first SLiM prediction algorithm able to estimate the statistical significance of motif predictions. SLiMFinder greatly increased the reliability of predictions. SLiMFinder has since spawned a number of motif discovery tools and webservers and is still arguably the most successful SLiM prediction tool on benchmarking data. Methods are made available through the SLiMSuite bioinformatics package and webservers.

Current research is looking to develop these SLiM prediction tools further and apply them to important biological questions. Of particular interest is the molecular mimicry employed by viruses to interact with host proteins and the role of SLiMs in other diseases, such as cancer. Other work is concerned with the evolutionary dynamics of SLiMs within protein interaction networks.

OTHER RESEARCH PROJECTS

In addition to the main research in the lab, the lab has a number of interdisciplinary collaborative projects applying bioinformatics tools and molecular evolution theory to experimental biology, often using large genomic, transcriptomic and/or proteomic datasets. These projects often involve the development of bespoke bioinformatics pipelines and a number of open source bioinformatics tools have been generated as a result. Please see the Publications and Lab software pages for more detail, or get in touch if something catches your eye and you want to find out more. We frequently have small collaborations and/or undergraduate student research projects. Many of these are “on hold” waiting for the right person, or sometimes data, to come along. If you think that you have what it needs, get in touch!

Tuesday, 2 July 2019

#GSA2019 - BUSCOMP: BUSCO Compilation and Comparison for Assessing Completeness in Multiple Genome Assemblies

Richard J. Edwards

If you are at the Genetics Society of Australasia Conference 2019, then come and hear me talk at 11:30 in Symposium 4B – Genomics & Bioinformatics (1). If you cannot make it, or loved the talk so much you want to look at it again, the slides are available on F1000Research:

  • Edwards RJ (2019): BUSCOMP: BUSCO compilation and comparison – Assessing completeness in multiple genome assemblies [version 1; not peer reviewed]. F1000Research 8:995 (slides)
    (doi: 10.7490/f1000research.1116972.1)

Abstract

Advances in DNA sequencing technology and free availability of bioinformatics tools have placed de novo genome assembly of complex organisms firmly in the domain of individual labs and small consortia. Nevertheless, the assemblies produced are often fragmented and incomplete. Optimal assembly depends on the size, repeat landscape, ploidy and heterozygosity of the genome, which are often unknown. It is therefore common practice to try multiple strategies, and there is a bottleneck in assessing and comparing assemblies.

BUSCO [1] is a powerful and popular tool that estimates genome completeness using gene prediction and curated models of single-copy protein orthologues. BUSCO assessments combine genome completeness, contiguity, and accuracy to rate genes as “Complete (Single Copy)”, “Duplicated”, “Fragmented” or “Missing”. However, results can be counterintuitive and lack robustness when comparing multiple assemblies of the same genome. Adding/removing scaffolds can alter the BUSCO genes returned by the rest of the assembly [2], while low sequence quality may reduce “completeness” scores and miss genes that are present in the assembly [3].

BUSCOMP (BUSCO Compilation and Comparison) is designed to complement BUSCO and identify/overcome these issues. BUSCOMP first compiles a non-redundant maximal set of the highest-scoring single-copy complete sequences for as many BUSCO genes as possible. These are then searched against assemblies using Minimap2 [4], converted into global alignment statistics, and used to robustly re-rate genes as Complete (Single/Duplicated), Fragmented/Partial or Missing. On test data from three organisms (yeast, cane toad and mainland tiger snake), BUSCOMP (1) gives consistent results when re-running the same assembly, (2) is not affected by adding or removing non-BUSCO-containing scaffolds, and (3) is minimally affected by assembly quality. This makes BUSCOMP ideal to run alongside BUSCO when trying to compare and rank genome assemblies, even in the absence of error-correction.

BUSCOMP is freely available at https://github.com/slimsuite/buscomp under a GNU GPL v3 license.

  1. Simão FA et al. (2015) Bioinformatics 31:3210–3212
  2. Edwards RJ et al. (2018) F1000Research 7:753
  3. Edwards RJ et al. (2018) GigaScience 7:giy095
  4. Li H (2018) Bioinformatics 34:3094-3100

Thursday, 6 June 2019

PhD available: Developing genomic resources to advance the molecular ecology of invasions

Expressions of interest are now open for a Scientia PhD Scholarship, in collaboration between the Edwards Lab, Lee Ann Rollins, and Marc Wilkins:

Developing genomic resources to advance the molecular ecology of invasions

These are exciting four-year scholarships to start in 2020 with full fees covered, a generous stipend, and career development funds. Please click on the toad or get in touch if you want to know more!

Closing date: 12 July 2019.
Location: University of New South Wales, Sydney, Australia

PROJECT DESCRIPTION

Invasive species pose a major challenge to biodiversity worldwide but also provide the unique opportunity to study evolution in action. Rapid changes are often associated with invaders’ introduction to novel environments. Understanding how molecular mechanisms drive these changes enables the creation of innovative solutions to controlling invasions and managing native species’ response to climatic change. The iconic Australian cane toad invasion is one of the best studied globally and is an emerging model for invasion genomics. This project will use whole genome sequencing, novel bioinformatic approaches and proteomics to identify molecular drivers of invasion success.

IDEAL CANDIDATE

We seek a highly motivated, curiosity-driven student with an interest in evolutionary biology and bioinformatics, who would like to understand why invasive species flourish. Ideally, candidates will have demonstrated computer literacy and be willing to learn new approaches to analysing genomic and proteomic data. Strong writing skills will be an asset. We will consider applicants coming from either a computing background who want to work in evolutionary biology or those with evolutionary biology backgrounds and a keen interest in bioinformatics. The supervisory team offer a high level of support in the fields of evolutionary biology, genomics, proteomics and bioinformatics. This project provides the opportunity for the successful applicant to develop the most current skills and build a successful career in these fields while contributing to solutions for two major global issues: loss of biodiversity and species’ response to climate change.

SUPERVISORY TEAM:

  • Dr Lee Ann Rollins, UNSW Scientia Fellow
  • Dr Richard Edwards, Senior Lecturer in Genomics & Bioinformatics
  • Prof Marc Wilkins, Professor of Systems Biology

HOW TO APPLY:

Complete an expression of interest at: https://www.scientia.unsw.edu.au/scientia-phd-scholarships/developing-genomic-resources-advance-molecular-ecology-invasions

The strongest expressions will receive an invitation to submit a full application to the scholarship competition.

CONTACT

To discuss the project and related opportunities, please contact Lee Ann Rollins (l.rollins[at]unsw.edu.au / @rollins_lee) or Rich Edwards (richard.edwards[at]unsw.edu.au / @slimsuite).

Monday, 18 March 2019

Term 2 Honours projects available

For any students completing their undegrad studies in UNSW Term 1, or external students finishing before June 2019, the UNSW Term 2 honours student application is now open. Deadline for Term 2 applications is 5pm Friday 12th April 2019. For how to apply please check the Faculty link: http://www.science.unsw.edu.au/honours-apply.

As usual, the EdwardsLab is looking to recruit enthusiastic students in genomics, in two main areas:

  1. Comparative genomics and molecular evolution in yeast, using long-read PacBio sequencing. We are particularly keen to get a student to work on aspects of our ARC Linkage grant, investigating the evolution of a novel biochemical pathway in yeast.

  2. De novo whole genome sequencing of vertebrates, including two snakes and the cane toad. In addition to general assembly and annotation activities, the lab has a few analytical tools that would benefit from

More details of honours can be found on the BABS website, or please get in touch if you have questions about specific projects. We welcome students interested in any of our Research areas, not just the projects listed. Applications from non-UNSW students are also encouraged.

Term 3 applications

Please also note that T3 honours application open date and deadline below, should you apply for T3 intake:

  • Applications open from Friday 17th May 2019

  • Deadline for Term 3 applications is 5pm Friday 26th July 2019

Monday, 18 February 2019

Harry Eyck (PhD student)

Harry Eyck commenced his PhD at UNSW in February 2019, after having completed his honours at Deakin University studying developmental stress in Zebra finches (Taeniopygia guttata).

His PhD work, under the primary supervision of Lee Ann Rollins, focuses on host-parasite interactions in invasive systems. When invasive species colonise new habitat, they often carry their parasites with them. This can lead to an evolutionary arms-race, causing major selection pressure for both parasites and their hosts. Despite this, they remain understudied. To investigate this, Harry studies the nematode lungworm (Rhabdias pseudosphaerocephala), which was brought to Australia by the infamous Cane toad (Rhinella marina).

His projects include assembling the lungworm genome, exploring how its genome has changed between populations with vastly different host-parasite interactions, and doing fieldwork and experiments investigating its behavioural responses and infection dynamics.

[LinkedIn]

Wednesday, 3 October 2018

Research snapshot - October 2018

One of the most important, interesting and challenging questions in biology is how new traits evolve at the molecular level. My lab employs sequence analysis techniques to interrogate protein and DNA sequences for the signals left behind by evolution. We are a bioinformatics lab but like to incorporate bench data through collaboration wherever possible.

Main Research

The core research in the lab is broadly divided into three main themes:

1. Short Linear Motifs (SLiMs)

Many protein-protein interactions are mediated by Short Linear Motifs (SLiMs): short stretches of proteins (5-15 amino acids long), of which only a few positions are critical to function. These motifs are vital for biological processes of fundamental importance, acting as ligands for molecular signalling, post-translational modifications and subcellular targeting. SLiMs have extremely compact protein interaction interfaces, generally encoded by less than 4 major affinity-/specificity-determining residues. Their small size enables high functional density and evolutionary plasticity, making them frequent products of convergent "ex nihilo" evolution. It also makes them challenging to identify, both experimentally and computationally.

A major focus of the lab is the computational prediction of SLiMs from protein sequences. This research originated with Rich’s postdoctoral research, during which he developed a sequence analysis methods for the rational design of biologically active short peptides. He subsequently developed SLiMDisc, one of the first algorithms for successfully predicting novel SLiMs from sequence data - and coined the term “SLiM” into the bargain. This subsequently lead to the development of SLiMFinder, the first SLiM prediction algorithm able to estimate the statistical significance of motif predictions. SLiMFinder greatly increased the reliability of predictions. SLiMFinder has since spawned a number of motif discovery tools and webservers and is still arguably the most successful SLiM prediction tool on benchmarking data.

Current research is looking to develop these SLiM prediction tools further and apply them to important biological questions. Of particular interest is the molecular mimicry employed by viruses to interact with host proteins and the role of SLiMs in other diseases, such as cancer. Other work is concerned with the evolutionary dynamics of SLiMs within protein interaction networks.

2. The evolution of novel functions.

Previous work in the lab has focused on the evolution of functional specificity following gene duplication. Since moving to UNSW, activities have shifted more towards the use of PacBio long read sequencing and other cutting-edge sequencing technologies, working closely with the Ramaciotti Centre for Genomics. We are collaborating with industrial and academic partners to de novo sequence, assemble, annotate and interrogate the genomes of a selection of microbes with interesting metabolic abilities. Most notably, we have an ARC Linkage Grant with Microbiogen Pty Ltd. to understand how a strain of Saccharomyces cerevisiae has evolved to efficiently use xylose as a sole carbon source: something vital for second generation biofuel production that wild yeast cannot do. We are combining comparative genomics, evolutionary genetics, RNA-Seq transcriptomics, and competition assays to understand how the novel metabolism evolved. Through deep Illumina resequencing of evolving populations, and assembling reliable complete genomes of the founding ancestors, the ultimate goal is to trace how mutations have interacted with existing genetic variation during adaptive evolution.

3. Whole genome sequencing and assembly.

Following our experiences with de novo whole genome assembly in yeast, the lab is getting involved in an increasing number of genome sequencing projects. The biggest of these is leading the bioinformatics and assembly effort in a consortium to sequence the cane toad genome. The lab is also leading the BABS Genome project two iconic Australian snakes for use in teaching and public engagement. We are a member of the Oz Mammals Genomics initiative, assisting with the sequencing and assembly of Australia's unique marsupial fauna. We also have an number of bacterial long-read whole genome sequencing collaborations.

Other Research Projects

In addition to the main research in the lab, the lab has a number of interdisciplinary collaborative projects applying bioinformatics tools and molecular evolution theory to experimental biology, often using large genomic, transcriptomic and/or proteomic datasets. These projects often involve the development of bespoke bioinformatics pipelines and a number of open source bioinformatics tools have been generated as a result. We frequently have small collaborations and/or undergraduate student research projects. Many of these are “on hold” waiting for the right person, or sometimes data, to come along. If you think that you have what it needs, get in touch!

Previous Research

The lab has been involved in a number of interdisciplinary collaborative projects applying bioinformatics tools and molecular evolution theory to experimental biology, often using large genomic, transcriptomic and/or proteomic datasets. These projects often involved the development of bespoke bioinformatics pipelines and a number of open source bioinformatics tools have been generated as a result. Please see the Publications and Lab software pages for more detail, or get in touch if something catches your eye and you want to find out more.

Tuesday, 7 August 2018

Draft genome assembly of the invasive cane toad, Rhinella marina

Richard J Edwards, Daniel Enosi Tuipulotu, Timothy G Amos, Denis O’Meally, Mark F Richardson, Tonia L Russell, Marcelo Vallinoto, Miguel Carneiro, Nuno Ferrand, Marc R Wilkins, Fernando Sequeira, Lee A Rollins, Edward C Holmes, Richard Shine & Peter A White (2018): Draft genome assembly of the invasive cane toad, Rhinella marina. GigaScience 7(9):giy095. [GigaScience] [PubMed] [PDF]

Abstract

Background. The cane toad (Rhinella marina formerly Bufo marinus) is a species native to Central and South America that has spread across many regions of the globe. Cane toads are known for their rapid adaptation and deleterious impacts on native fauna in invaded regions. However, despite an iconic status, there are major gaps in our understanding of cane toad genetics. The availability of a genome would help to close these gaps and accelerate cane toad research.

Findings. We report a draft genome assembly for R. marina, the first of its kind for the Bufonidae family. We used a combination of long read PacBio RS II and short read Illumina HiSeq X sequencing to generate a total of 359.5 Gb of raw sequence data. The final hybrid assembly of 31,392 scaffolds was 2.55 Gb in length with a scaffold N50 of 168 kb. BUSCO analysis revealed that the assembly included full length or partial fragments of 90.6% of tetrapod universal single-copy orthologs (n = 3950), illustrating that the gene-containing regions have been well-assembled. Annotation predicted 25,846 protein coding genes with similarity to known proteins in SwissProt. Repeat sequences were estimated to account for 63.9% of the assembly.

Conclusion. The R. marina draft genome assembly will be an invaluable resource that can be used to further probe the biology of this invasive species. Future analysis of the genome will provide insights into cane toad evolution and enrich our understanding of their interplay with the ecosystem at large.

(More details to follow in future posts.)

Monday, 5 February 2018

Katarina Stuart (PhD Student)

Katarina Stuart completed her undergrad at the University of Sydney, completing an honours thesis on the evolution and trait plasticity of the invasive cane toad.

She commenced her PhD at UNSW in February 2018 under the primary supervision of Lee Ann Rollins. Her thesis project aims to use genomics to investigate evolution in the European Starlings (Sturnus vulgaris). The global nature of starling invasions provides an important opportunity to examine evolutionary questions about invasion success, and how this is linked to similarities or emerging differences in their genome post colonisation and during range expansion. Major components of her thesis involve updating the starling genome, and exploring the population genetics of Australia’s starling invasion in relation to their native range counterparts, both modern and historic.

Katarina’s general interest are in invasive species, and their dynamic evolutionary change over geographical and temporal ranges.

[LinkedIn | Twitter]

Tuesday, 4 July 2017

GEN2017: Mitochondrial variation and heteroplasmy in Australian and Hawai’ian cane toads

If you were attending this year’s Annual Conference of the Genetics Society of Australasia with the NZ Society for Biochemistry & Molecular Biology, hopefully you made it to the oral presentation of Lee Ann Rollins. Although we do not yet have enough PacBio data for a pure long read assembly of the nuclear genome, the mitochondrion is another matter!

Mitochondrial variation and heteroplasmy in Australian and Hawai’ian cane toads.

Lee A Rollins[1], Mark F Richardson[1], Daniel M Selechnik[2], Andrea J West[1], Timothy G Amos[3], Richard J Edwards[3] & Richard Shine[2]

  1. School of Life and Environmental Sciences, Centre for Integrative Ecology, Deakin University, Geelong, VIC, Australia
  2. School of Life and Environmental Sciences, University of Sydney, Sydney, NSW, Australia
  3. School of Biotechnology and Biomolecular Sciences, University of New South Wales, Sydney, NSW, Australia

Abstract

Background/Aims. Invasive species can adapt to new environments despite low levels of standing genetic diversity due to small founding numbers or sequential introductions. The iconic Australian cane toad was sourced from an introduced population in Hawai’i and conflicting evidence exists regarding the level of genetic diversity across these invasions.

Methods. We extracted mitochondrial sequence data from the genome of one individual sequenced using the PacBio RSII and Illumina X10 platforms. From these data, we assembled and annotated the mitochondrial genome. RNAseq data from 18 individuals collected from Hawai’i and 68 individuals from Australia were aligned to the reference sequence. We quantified polymorphism across samples and heteroplasmy (multiple mitochondrial haplotypes within individuals).

Results. A complete, annotated mitochondrial reference genome was constructed consisting of 18154 base pairs (bp), the largest reported bufonid mitochondrial genome. We aligned RNAseq data to the entire reference sequence, with the exception of a 347bp region containing several 104bp repeats. We identified 16 polymorphisms (17 haplotypes); one haplotype was common to 65 individuals sampled in both introductions. Heteroplasmy was detected at most polymorphic sites and also at multiple sites where the predominant haplotype was common to all individuals.

Conclusions. Mitochondrial diversity is low in Australian and Hawai’ian cane toads. Our findings add to the growing body of evidence that heteroplasmy may be ubiquitous across taxa. Selection within heteroplasmic individuals (recently demonstrated in expanding populations) may provide an important source of variation in genetically depauperate populations.

Funding. Australian Research Council DE150101393 (LAR) and FL120100074 (RS)

Friday, 30 June 2017

Research Snapshot - June 2017

Research interests in the Edwards lab stem from a fascination with the molecular basis of evolutionary change and how we can harness the genetic sequence patterns left behind to make useful predictions about contemporary biological systems. We are a bioinformatics lab but like to incorporate bench data through collaboration wherever possible.

Main Research

The core research in the lab is broadly divided into three main themes:

1. Short Linear Motifs (SLiMs)

SLiMs are short regions of proteins that mediate interactions with other proteins. A major focus of the lab is the computational prediction of SLiMs from protein sequences. This research originated with Rich’s postdoctoral research, during which he developed a sequence analysis methods for the rational design of biologically active short peptides. He subsequently developed SLiMDisc, one of the first algorithms for successfully predicting novel SLiMs from sequence data - and coined the term “SLiM” into the bargain. This subsequently lead to the development of SLiMFinder, the first SLiM prediction algorithm able to estimate the statistical significance of motif predictions. SLiMFinder greatly increased the reliability of predictions. SLiMFinder has since spawned a number of motif discovery tools and webservers and is still arguably the most successful SLiM prediction tool on benchmarking data.

Current research is looking to develop these SLiM prediction tools further and apply them to important biological questions. Of particular interest is the molecular mimicry employed by viruses to interact with host proteins and the role of SLiMs in other diseases, such as cancer. Other work is concerned with the evolutionary dynamics of SLiMs within protein interaction networks.

2. The evolution of novel functions.

Previous work in the lab has focused on the evolution of functional specificity following gene duplication. Since moving to UNSW, activities have shifted more towards the use of PacBio long read sequencing and other cutting-edge sequencing technologies, working closely with the Ramaciotti Centre for Genomics. We are collaborating with industrial and academic partners to de novo sequence, assemble, annotate and interrogate the genomes of a selection of microbes with interesting metabolic abilities. Most notably, we have an ARC Linkage Grant with Microbiogen Pty Ltd. to understand how a strain of Saccharomyces cerevisiae has evolved to efficiently use xylose as a sole carbon source: something vital for second generation biofuel production that wild yeast cannot do. This project combines detailed molecular characterisation of highly adapted yeast strains with “molecular palaeontology” to trace the evolutionary process and identify functionally significant loci under selection.

3. Whole genome sequencing and assembly.

Following our experiences with de novo whole genome assembly in yeast, the lab is getting involved in an increasing number of genome sequencing projects. The biggest of these is leading the bioinformatics and assembly effort in a consortium to sequence the cane toad genome. The lab is also leading the BABS Genome project to sequence iconic Australian species for use in teaching and public engagement.

Previous Research

The lab has been involved in a number of interdisciplinary collaborative projects applying bioinformatics tools and molecular evolution theory to experimental biology, often using large genomic, transcriptomic and/or proteomic datasets. These projects often involved the development of bespoke bioinformatics pipelines and a number of open source bioinformatics tools have been generated as a result. Please see the Publications and Lab software pages for more detail, or get in touch if something catches your eye and you want to find out more.

Saturday, 26 November 2016

The cane toad genome project

What are we doing? The Edwards Lab is part of an Australian, Portuguese and Brazilian consortium led by Peter White to sequence and assemble the genome of the cane toad (Rhinella marina). We are leading the bioinformatics component of the assembly effort.

How are we doing it? We are using a combination of Illumina (HiSeq X and NovaSeq) short read sequencing, PacBio (RS II) long read and sequence and 10x Genomics Chromium linked reads.

Details to follow. Please get in touch if you are interested in the project.

Opportunities

Honours and postgraduate* projects are available to work on the assembly and annotation. (*PhD students should have their own scholarship.)

Consortium members

Miguel Carneiro (CIBIO-InBIO), Richard Edwards (UNSW), Nuno Ferrand (CIBIO-InBIO), Eddie Holmes (U Sydney), Craig Moritz (ANU), Lee Ann Rollins (Deakin), Fernando Sequeira (CIBIO-InBIO), Rick Shine (U Sydney), Marcelo Vallinoto de Souza (Federal University of Pará & CIBIO-InBIO), Peter White (UNSW), Marc Wilkins (UNSW).