Showing posts with label pacbio. Show all posts
Showing posts with label pacbio. Show all posts

Tuesday, 2 July 2024

Extant and extinct bilby genomes combined with Indigenous knowledge improve conservation of a unique Australian marsupial

Hogg C*, Edwards RJ*, Farquharson K*, Silver L*, Brandies P, Peel E, Escalona M, Jaya FR, Thavornkanlapachai R, Batley K, Bradford TM, Chang JK, Chen Z, Deshpande N, Dziminski M, Ewart KM, Griffith OW, Marin Gual L, Moon KL, Travouillon KJ, Waters P, Whittington CM, Wilkins MR, Helgen KM, Lo N, Ho SYW, Ruiz Herrera A, Paltridge R, Marshall Graves JA, Renfree M, Shapiro B, Ottewell K, Kiwirrkurra Rangers & Belov K (2024): Extant and extinct bilby genomes combined with Indigenous knowledge improve conservation of a unique Australian marsupial. Nature Ecology & Evolution 8:1311–1326. [*Joint first authors] [Research Square] [Nat Ecol Evol] [PubMed]

Abstract

Ninu (greater bilby, Macrotis lagotis) are desert-dwelling, culturally and ecologically important marsupials. In collaboration with Indigenous rangers and conservation managers, we generated the Ninu chromosome-level genome assembly (3.66 Gbp) and genome sequences for the extinct Yallara (lesser bilby, Macrotis leucura). We developed and tested a scat single-nucleotide polymorphism panel to inform current and future conservation actions, undertake ecological assessments and improve our understanding of Ninu genetic diversity in managed and wild populations. We also assessed the beneficial impact of translocations in the metapopulation (N = 363 Ninu). Resequenced genomes (temperate Ninu, 6; semi-arid Ninu, 6; and Yallara, 4) revealed two major population crashes during global cooling events for both species and differences in Ninu genes involved in anatomical and metabolic pathways. Despite their 45-year captive history, Ninu have fewer long runs of homozygosity than other larger mammals, which may be attributable to their boom-bust life history. Here we investigated the unique Ninu biology using 12 tissue transcriptomes revealing expression of all 115 conserved eutherian chorioallantoic placentation genes in the uterus, an XY1Y2 sex chromosome system and olfactory receptor gene expansions. Together, we demonstrate the holistic value of genomics in improving key conservation actions, understanding unique biological traits and developing tools for Indigenous rangers to monitor remote wild populations.

Friday, 15 December 2023

A high-quality pseudo-phased genome for Melaleuca quinquenervia shows allelic diversity of NLR-type resistance genes

Chen SH, Martino AM, Luo Z, Schwessinger B, Jones A, Tolessa T, Bragg JG, Tobias PA, Edwards RJ (2023): A high-quality pseudo-phased genome for Melaleuca quinquenervia shows allelic diversity of NLR-type resistance genes. GigaScience 12:giad102. [Gigascience] [PubMed]

Background. Melaleuca quinquenervia (broad-leaved paperbark) is a coastal wetland tree species that serves as a foundation species in eastern Australia, Indonesia, Papua New Guinea, and New Caledonia. While extensively cultivated for its ornamental value, it has also become invasive in regions like Florida, USA. Long-lived trees face diverse pest and pathogen pressures, and plant stress responses rely on immune receptors encoded by the nucleotide-binding leucine-rich repeat (NLR) gene family. However, the comprehensive annotation of NLR encoding genes has been challenging due to their clustering arrangement on chromosomes and highly repetitive domain structure; expansion of the NLR gene family is driven largely by tandem duplication. Additionally, the allelic diversity of the NLR gene family remains largely unexplored in outcrossing tree species, as many genomes are presented in their haploid, collapsed state.

Results. We assembled a chromosome-level pseudo-phased genome for M. quinquenervia and described the allelic diversity of plant NLRs using the novel FindPlantNLRs pipeline. Analysis reveals variation in the number of NLR genes on each haplotype, distinct clustering patterns, and differences in the types and numbers of novel integrated domains.

Conclusions. The high-quality M. quinquenervia genome assembly establishes a new framework for functional and evolutionary studies of this significant tree species. Our findings suggest that maintaining allelic diversity within the NLR gene family is crucial for enabling responses to environmental stress, particularly in long-lived plants.

Wednesday, 29 March 2023

The Australasian dingo archetype: De novo chromosome-length genome assembly, DNA methylome, and cranial morphology

Ballard JWO, Field MA, Edwards RJ, Wilson LAB, Koungoulos LG, Rosen BD, Chernoff B, Dudchenko O, Omer A, Keilwagen J, Skvortsova K, Bogdanovic O, Chan E, Zammit R, Hayes V & Aiden EL (2023): The Australasian dingo archetype: De novo chromosome-length genome assembly, DNA methylome, and cranial morphology. Gigascience 12:giad018. [Gigascience] [PubMed]

Background

One difficulty in testing the hypothesis that the Australasian dingo is a functional intermediate between wild wolves and domesticated breed dogs is that there is no reference specimen. Here we link a high-quality de novo long-read chromosomal assembly with epigenetic footprints and morphology to describe the Alpine dingo female named Cooinda. It was critical to establish an Alpine dingo reference because this ecotype occurs throughout coastal eastern Australia where the first drawings and descriptions were completed.

Findings

We generated a high-quality chromosome-level reference genome assembly (Canfam_ADS) using a combination of Pacific Bioscience, Oxford Nanopore, 10X Genomics, Bionano, and Hi-C technologies. Compared to the previously published Desert dingo assembly, there are large structural rearrangements on chromosomes 11, 16, 25, and 26. Phylogenetic analyses of chromosomal data from Cooinda the Alpine dingo and 9 previously published de novo canine assemblies show dingoes are monophyletic and basal to domestic dogs. Network analyses show that the mitochondrial DNA genome clusters within the southeastern lineage, as expected for an Alpine dingo. Comparison of regulatory regions identified 2 differentially methylated regions within glucagon receptor GCGR and histone deacetylase HDAC4 genes that are unmethylated in the Alpine dingo genome but hypermethylated in the Desert dingo. Morphologic data, comprising geometric morphometric assessment of cranial morphology, place dingo Cooinda within population-level variation for Alpine dingoes. Magnetic resonance imaging of brain tissue shows she had a larger cranial capacity than a similar-sized domestic dog.

Conclusions

These combined data support the hypothesis that the dingo Cooinda fits the spectrum of genetic and morphologic characteristics typical of the Alpine ecotype. We propose that she be considered the archetype specimen for future research investigating the evolutionary history, morphology, physiology, and ecology of dingoes. The female has been taxidermically prepared and is now at the Australian Museum, Sydney.

Thursday, 6 October 2022

The Ocean Genomes Laboratory is hiring!

The Minderoo OceanOmics Centre at UWA Ocean Genomes Laboratory is now hiring our technical team to support high throughput DNA sequencing and genome assembly. We currently have three "wet" lab positions going: a Sequencing Specialist Scientific Officer, and two Sequencing Technician positions. Both roles will be providing technical support in the lab, particularly with respect to all aspects of DNA sequencing (sample extraction, library preparation and setting up sequencing runs). You'll get to play with the latest sequencing toys, including Illumina NovaSeq 6000, NextSeq 2000 and iSeq 100, the PacBio Sequel IIe, and ONT (probably PromethION and MinION).

The closing date for applications is 11:55 PM AWST on Thursday 27 October 2022.

To learn more about these opportunities, please click on the links above or contact Rich Edwards at rich.edwards@uwa.edu.au. We will also be advertising some bioinformatics positions soon.

About the team

The Minderoo OceanOmics Centre at UWA combines a joint Ocean Genomes Laboratory, an OceanOmics Laboratory, and Computational Biology Services.

Equipped with the latest high-throughput sequencing technology and in collaboration with global partners, the Ocean Genomes Laboratory will generate a comprehensive library of high quality marine vertebrate reference genome assemblies. All such reference genome data will be subject to rigorous QA/QC and all assemblies will be released publicly with open access.

The Ocean Genomes Laboratory will undertake research and development under the direction of Minderoo’s ambitious OceanOmics Program which has the goal of revolutionising ocean conservation through novel marine sampling and genomics approaches and scaling these to significantly advance our knowledge of marine life. The Ocean Genomes Laboratory and Computational Biology Services will include state of the art infrastructure including sample and eDNA preparation areas, flow cytometry, single cell sequencing equipment and the latest bioinformatics and computational biology tools.

Wednesday, 29 June 2022

The starling genome is out!

See the pre-print post for details.

Stuart KC*, Edwards RJ*, Cheng Y, Warren WC, Burt DW, Sherwin WB, Hofmeister NR, Werner SJ, Ball GF, Bateson M, Brandley MC, Buchanan KL, Cassey P, Clayton DF, De Meyer T, Meddle SL & Rollins LA (2022): Transcript- and annotation-guided genome assembly of the European starling. Molecular Ecology 22(8):3141-3160. doi: 10.1111/1755-0998.13679. [*Joint first authors] [Mol Ecol Res] [PubMed] [bioRxiv]

The European starling, Sturnus vulgaris, is an ecologically significant, globally invasive avian species that is also suffering from a major decline in its native range. Here, we present the genome assembly and long-read transcriptome of an Australian-sourced European starling (S. vulgaris vAU), and a second, North American, short-read genome assembly (S. vulgaris vNA), as complementary reference genomes for population genetic and evolutionary characterization. S. vulgaris vAU combined 10× genomics linked-reads, low-coverage Nanopore sequencing, and PacBio Iso-Seq full-length transcript scaffolding to generate a 1050 Mb assembly on 6222 scaffolds (7.6 Mb scaffold N50, 94.6% busco completeness). Further scaffolding against the high-quality zebra finch (Taeniopygia guttata) genome assigned 98.6% of the assembly to 32 putative nuclear chromosome scaffolds. Species-specific transcript mapping and gene annotation revealed good gene-level assembly and high functional completeness. Using S. vulgaris vAU, we demonstrate how the multifunctional use of PacBio Iso-Seq transcript data and complementary homology-based annotation of sequential assembly steps (assessed using a new tool, saaga) can be used to assess, inform, and validate assembly workflow decisions. We also highlight some counterintuitive behaviour in traditional busco metrics, and present buscomp, a complementary tool for assembly comparison designed to be robust to differences in assembly size and base-calling quality. This work expands our knowledge of avian genomes and the available toolkit for assessing and improving genome quality. The new genomic resources presented will facilitate further global genomic and transcriptomic analysis on this ecologically important species.

Saturday, 23 April 2022

The Australian dingo is an early offshoot of modern breed dogs

Field MA, Yadav S, Dudchenko O, Esvaran M, Rosen BD, Skvortsova K, Edwards RJ, Keilwagen J, Cochran BJ, Manandhar B, Bustamante S, Rasmussen JA, Melvin RG, Chernoffl B, Omer A, Colaric Z, Chan EKF, Minoche AE, Smith TPL, Gilbert MTP, Bogdanovic O, Zammit RA, Thomas T, Aiden EL & Ballard JWO (2022): The Australian dingo is an early offshoot of modern breed dogs. Science Advances 8(16):abm5944; DOI: 10.1126/sciadv.abm5944. [Sci Adv] [PubMed] [PDF]

Dogs are uniquely associated with human dispersal and bring transformational insight into the domestication process. Dingoes represent an intriguing case within canine evolution being geographically isolated for thousands of years. Here, we present a high-quality de novo assembly of a pure dingo (CanFam_DDS). We identified large chromosomal differences relative to the current dog reference (CanFam3.1) and confirmed no expanded pancreatic amylase gene as found in breed dogs. Phylogenetic analyses using variant pairwise matrices show that the dingo is distinct from five breed dogs with 100% bootstrap support when using Greenland wolf as the outgroup. Functionally, we observe differences in methylation patterns between the dingo and German shepherd dog genomes and differences in serum biochemistry and microbiome makeup. Our results suggest that distinct demographic and environmental conditions have shaped the dingo genome. In contrast, artificial human selection has likely shaped the genomes of domestic breed dogs after divergence from the dingo.

Wednesday, 24 November 2021

#ABACBS2021 Lightning talk - DepthSizer and DepthKopy: genome size and copy number prediction using single-copy long-read depth profiles

Tune in for the ABACBS 2021 lightning talks this morning to hear about our applications of long reads to genome size and copy number prediction. For more information, you can read the NSW Waratah genome paper pre-print or visit the GitHub pages for DepthSizer and DepthKopy.

Richard J Edwards, Stephanie H Chen, Katarina C Stuart, Mark M Tanaka, Jason G Bragg.

DepthSizer and DepthKopy: genome size and copy number prediction using single-copy long-read depth profiles

A fundamental part of any genome project is establishing the genome size of the organism being sequenced. The gold standard for genome size measurement is flow cytometry, but this is not available to all groups and can give surprisingly variable results. Popular bioinformatic approaches predict genome size using kmer frequency profiles from high-accuracy (e.g. illumina or hifi) sequencing reads, or the mean depth of coverage reads mapped to an assembly. Both of these approaches can be adversely affected by repetitive regions of the genome. Mean sequencing depth is also highly reliant on assembly completeness.

Here, we present DepthSizer (https://github.com/slimsuite/depthsizer), which refines this approach by estimating sequencing depth based on single-copy complete BUSCO genes. DepthSizer works on the principle that genuine single-copy regions will tend towards the same, true, single-copy read depth. In contrast, assembly errors, collapsed repeats within those genes, or incorrect BUSCO predictions, will give inconsistent read depth deviations. The modal read depth across single-copy BUSCO genes, calculated from a depth density profile of these regions, should therefore provide a good estimate of the true depth of coverage. The method is benchmarked on model organism data and corrections for possible contamination, biases/inconsistencies in read mapping and/or raw read insertion/deletion error profiles are discussed. We also present DepthKopy (https://github.com/slimsuite/depthkopy), which uses the same read depth approach to estimate the copy number of assembly regions. This can be useful for identifying haplotigs, and collapsed repeat regions.

Keywords: BUSCO, Genome Assembly, Genomics, ONT, PacBio, copy number variants

Thursday, 8 April 2021

Transcript- and annotation-guided genome assembly of the European starling

Our starling genome paper is now available as a pre-print on bioRxiv! This was some great work by PhD student, Kat Stuart. Kat assembled a new Australian starling genome, using a combination of linked reads, low coverage long reads, and long-read PacBio iso-seq transcriptomics data. A second Illumina assembly of a North American group is also presented. As we saw with our Basenji genome paper, having two (or more) genomes from a species can be really useful for disentangling real difference from assembly artefacts. (No assembly is perfect!)

This paper is a great example of how a bit of TLC and imagination can get the most out of data produced with a limited budget. We were unable to get deep long-read sequencing this time, but instead show the additional power that long-read full-length transcriptome data can provide in assembling a genome - above and beyond the annotation.

This paper also officially describes a couple of genomics tools from the lab BUSCOMP has been in the works for some time, and this paper updates previous results to BUSCO v5 analysis and confirms our previous snake results using starling-derived test data. BUSCO is a powerful and popular tool that estimates genome completeness using gene prediction and curated models of single-copy protein orthologues. However, we demonstrate how results can be counterintuitive: adding/removing scaffolds can alter BUSCO predictions elsewhere in the assembly, while low sequence quality may reduce “completeness” scores and miss genes that are present in the assembly. BUSCOMP (BUSCO Compilation and Comparison) (https://github.com/slimsuite/buscomp) complements BUSCO to identify/overcome these issues by compiling a non-redundant set of the highest-scoring single-copy BUSCO complete sequences and re-searching these against assemblies for consistent completness scoring. SAAGA (https://github.com/slimsuite/saaga) is a new tool for annotation versus reference proteome comparisons. SAAGA can compare different annotations of the same assembly, or be combined with a lightweight annotation tool like GeMoMa to compare different assemblies of the same organism.


Stuart KC, Edwards RJ, Cheng Y, Warren WC, Burt DW, Sherwin WB, Hofmeister NR, Werner SJ, Ball GF, Bateson M, Brandley MC, Buchanan KL, Cassey P, Clayton DF, De Meyer T, Meddle SL & Rollins LA (preprint): Transcript- and annotation-guided genome assembly of the European starling. bioRxiv 2021.04.07.438753; doi: 10.1101/2021.04.07.438753. [*Joint first authors] [bioRxiv]

Abstract

The European starling, Sturnus vulgaris, is an ecologically significant, globally invasive avian species that is also suffering from a major decline in its native range. Here, we present the genome assembly and long-read transcriptome of an Australian-sourced European starling (S. vulgaris vAU), and a second North American genome (S. vulgaris vNA), as complementary reference genomes for population genetic and evolutionary characterisation. S. vulgaris vAU combined 10x Genomics linked-reads, low-coverage Nanopore sequencing, and PacBio Iso-Seq full-length transcript scaffolding to generate a 1050 Mb assembly on 1,628 scaffolds (72.5 Mb scaffold N50). Species-specific transcript mapping and gene annotation revealed high structural and functional completeness (94.6% BUSCO completeness). Further scaffolding against the high-quality zebra finch (Taeniopygia guttata) genome assigned 98.6% of the assembly to 32 putative nuclear chromosome scaffolds. Rapid, recent advances in sequencing technologies and bioinformatics software have highlighted the need for evidence-based assessment of assembly decisions on a case-by-case basis. Using S. vulgaris vAU, we demonstrate how the multifunctional use of PacBio Iso-Seq transcript data and complementary homology-based annotation of sequential assembly steps (assessed using a new tool, SAAGA) can be used to assess, inform, and validate assembly workflow decisions. We also highlight some counter-intuitive behaviour in traditional BUSCO metrics, and present BUSCOMP, a complementary tool for assembly comparison designed to be robust to differences in assembly size and base-calling quality. Finally, we present a second starling assembly, S. vulgaris vNA, to facilitate comparative analysis and global genomic research on this ecologically important species.

Monday, 2 November 2020

Antarctic desert soil bacteria exhibit high novel natural product potential, evaluated through long-read genome sequencing and comparative genomics

The third of our collaborative “controlled bacterial metagenome” de novo whole genome assembly projects was published in Environmental Microbiology in November. This was a fun collaboration with the Ferrari lab at UNSW trying to maximise bang for buck to sequence some complete bacterial genomes using PacBio sequencing to identify biosynthetic gene clusters. As with a previous paper, we used pooled genomic DNA sequencing and were able to assemble complete genomes (and plasmids) of the 13/17 species that had sufficient depth of coverage. Coolest of all (if you excuse the pun), these were bugs from an Antarctic expedition! Head over to the Ferrari lab website to find out more about their research.

Benaud N, Edwards RJ, Amos TG, D’Agostino PM, Gutiérrez-Cháveza C, Montgomery K, Nicetic I & Ferrari BC (2020). Antarctic desert soil bacteria exhibit high novel natural product potential, evaluated through long-read genome sequencing and comparative genomics. Environmental Microbiology. https://doi.org/10.1111/1462-2920.15300

Abstract

Actinobacteria and Proteobacteria are important producers of bioactive natural products (NP), and these phyla dominate in the arid soils of Antarctica, where metabolic adaptations influence survival under harsh conditions. Biosynthetic gene clusters (BGCs) which encode NPs, are typically long and repetitious high G + C regions difficult to sequence with short‐read technologies. We sequenced 17 Antarctic soil bacteria from multi‐genome libraries, employing the long‐read PacBio platform, to optimize capture of BGCs and to facilitate a comprehensive analysis of their NP capacity. We report 13 complete bacterial genomes of high quality and contiguity, representing 10 different cold‐adapted genera including novel species. Antarctic BGCs exhibited low similarity to known compound BGCs (av. 31%), with an abundance of terpene, non‐ribosomal peptide and polyketide‐encoding clusters. Comparative genome analysis was used to map BGC variation between closely related strains from geographically distant environments. Results showed the greatest biosynthetic differences to be in a psychrotolerant Streptomyces strain, as well as a rare Actinobacteria genus, Kribbella, while two other Streptomyces spp. were surprisingly similar to known genomes. Streptomyces and Kribbella BGCs were predicted to encode antitumour, antifungal, antibacterial and biosurfactant‐like compounds, and the synthesis of NPs with antibacterial, antifungal and surfactant properties was confirmed through bioactivity assays.

Thursday, 2 April 2020

Canfam_GSD: De novo chromosome-length genome assembly of the German Shepherd Dog (Canis lupus familiaris) using a combination of long reads, optical mapping, and Hi-C

Our latest paper is out! This one is a bit more photogenic than the cane toad - a German Shepherd Dog called Nala. This was a big international effort in a collaboration led by Bill Ballard at UNSW that included a dozen institutions across four continents. We threw all the main sequencing technologies at this one and achieved a chromosome-level assembly of better quality than the current “CanFam” reference genome.

You can find out more in the UNSW press release.



Field MA, Rosen BD, Dudchenko O, Chan EKF, Minoche AM, Edwards RJ, Barton K, Lyons RJ, Enosi Tuipulotu D, Hayes VM, Omer AD, Colaric Z, Keilwagen J, Skvortsova K, Bogdanovic O, Smith MA, Lieberman Aiden E, Smith TPL, Zammit RA & Ballard JWO (2020): Canfam_GSD: De novo chromosome-length genome assembly of the German Shepherd Dog (Canis lupus familiaris) using a combination of long reads, optical mapping, and Hi-C. GigaScience 9(4):giaa027. [GigaScience]


Abstract

Background

The German Shepherd Dog (GSD) is one of the most common breeds on earth and has been bred for its utility and intelligence. It is often first choice for police and military work, as well as protection, disability assistance, and search-and-rescue. Yet, GSDs are well known to be susceptible to a range of genetic diseases that can interfere with their training. Such diseases are of particular concern when they occur later in life, and fully trained animals are not able to continue their duties.

Findings

Here, we provide the draft genome sequence of a healthy German Shepherd female as a reference for future disease and evolutionary studies. We generated this improved canid reference genome (CanFam_GSD) utilizing a combination of Pacific Bioscience, Oxford Nanopore, 10X Genomics, Bionano, and Hi-C technologies. The GSD assembly is ∼80 times as contiguous as the current canid reference genome (20.9 vs 0.267 Mb contig N50), containing far fewer gaps (306 vs 23,876) and fewer scaffolds (429 vs 3,310) than the current canid reference genome CanFamv3.1. Two chromosomes (4 and 35) are assembled into single scaffolds with no gaps. BUSCO analyses of the genome assembly results show that 93.0% of the conserved single-copy genes are complete in the GSD assembly compared with 92.2% for CanFam v3.1. Homology-based gene annotation increases this value to ∼99%. Detailed examination of the evolutionarily important pancreatic amylase region reveals that there are most likely 7 copies of the gene, indicative of a duplication of 4 ancestral copies and the disruption of 1 copy.

Conclusions

GSD genome assembly and annotation were produced with major improvement in completeness, continuity, and quality over the existing canid reference. This resource will enable further research related to canine diseases, the evolutionary relationships of canids, and other aspects of canid biology.

Photo credit: Outdoor Action Photography.

Friday, 31 January 2020

Research snapshot - January 2020

One of the most important, interesting and challenging questions in biology is how new traits evolve at the molecular level. My lab employs sequence analysis techniques to interrogate DNA and protein sequences for the signals left behind by evolution. We are a bioinformatics lab but like to incorporate bench/field data through collaboration wherever possible.

Main Research

Building on a solid foundation of bioinformatics and evolutionary theory, we apply genomics, transcriptomics, proteomics and interactomics and systems analysis to understand complex biological systems. Core research activities can be broadly divided into two main themes:

  1. Evolutionary Genomics, with a focus on applying de novo genome assembly and population genomics to problems in ecology and biotechnology.

  2. Protein-protein interactions, with a focus on the prediction of short linear interaction motifs and their role in human health and disease, including host-pathogen interactions.

These are explored in more detail, below.

1. Evolutionary Genomics.

The main research focus of the lab is the exploitation of genomic and post-genomic data to understand biological function and adaptation to novel environments. We work closely with the Ramaciotti Centre for Genomics and are involved in numerous de novo whole genome sequencing and assembly projects, using short read (Illumina), long read (PacBio & Nanopore) and linked read (10x Chromium) sequencing. One of the biggest of these is leading the bioinformatics and assembly effort in a consortium to sequence the cane toad genome, and leading the BABS Genome project to sequence two iconic Australian snakes. We are a member of the Oz Mammals Genomics initiative, assisting with the sequencing and assembly of Australia’s unique marsupial fauna. In 2018, we were selected as part of a team to sequence the Waratah genome as part of the pilot phase for the new Genomics of Australian Plants initiative.

We enjoy bringing our bioinformatics to bear on a variety of collaborative research projects. Most notably, we have an ARC Linkage Grant with Microbiogen Pty Ltd to understand how a strain of Saccharomyces cerevisiae has evolved to efficiently use xylose as a sole carbon source: something vital for second generation biofuel production that wild yeast cannot do. We are combining comparative genomics, evolutionary genetics, RNA-Seq transcriptomics, and competition assays to understand how the novel metabolism evolved. Through deep Illumina resequencing of evolving populations, and assembling reliable complete genomes of the founding ancestors, the ultimate goal is to trace how mutations have interacted with existing genetic variation during adaptive evolution. More recently, we have received an ARC Linkage Grant with the Royal Botanic Gardens and Domain Trust, to apply genomics approaches to the challenges of rainforest tree conservation in the face of climate change and invasive pathogens. We are also collaborating with industrial and academic partners to de novo sequence, assemble, annotate and interrogate the genomes of a selection of microbes with interesting metabolic abilities.

2. Short Linear Motifs (SLiMs).

Many protein-protein interactions are mediated by Short Linear Motifs (SLiMs): short stretches of proteins (5-15 amino acids long), of which only a few positions are critical to function. These motifs are vital for biological processes of fundamental importance, acting as ligands for molecular signalling, post-translational modifications and subcellular targeting. SLiMs have extremely compact protein interaction interfaces, generally encoded by less than 4 major affinity-/specificity-determining residues. Their small size enables high functional density and evolutionary plasticity, making them frequent products of convergent “ex nihilo” evolution. It also makes them challenging to identify, both experimentally and computationally.

A major focus of the lab is the computational prediction of SLiMs from protein sequences. This research originated with Rich’s postdoctoral research, during which he developed a sequence analysis methods for the rational design of biologically active short peptides. He subsequently developed SLiMDisc, one of the first algorithms for successfully predicting novel SLiMs from sequence data - and coined the term “SLiM” into the bargain. This subsequently lead to the development of SLiMFinder, the first SLiM prediction algorithm able to estimate the statistical significance of motif predictions. SLiMFinder greatly increased the reliability of predictions. SLiMFinder has since spawned a number of motif discovery tools and webservers and is still arguably the most successful SLiM prediction tool on benchmarking data. Methods are made available through the SLiMSuite bioinformatics package and webservers.

Current research is looking to develop these SLiM prediction tools further and apply them to important biological questions. Of particular interest is the molecular mimicry employed by viruses to interact with host proteins and the role of SLiMs in other diseases, such as cancer. Other work is concerned with the evolutionary dynamics of SLiMs within protein interaction networks.

OTHER RESEARCH PROJECTS

In addition to the main research in the lab, the lab has a number of interdisciplinary collaborative projects applying bioinformatics tools and molecular evolution theory to experimental biology, often using large genomic, transcriptomic and/or proteomic datasets. These projects often involve the development of bespoke bioinformatics pipelines and a number of open source bioinformatics tools have been generated as a result. Please see the Publications and Lab software pages for more detail, or get in touch if something catches your eye and you want to find out more. We frequently have small collaborations and/or undergraduate student research projects. Many of these are “on hold” waiting for the right person, or sometimes data, to come along. If you think that you have what it needs, get in touch!

Friday, 5 July 2019

Research snapshot - July 2019

One of the most important, interesting and challenging questions in biology is how new traits evolve at the molecular level. My lab employs sequence analysis techniques to interrogate protein and DNA sequences for the signals left behind by evolution. We are a bioinformatics lab but like to incorporate bench data through collaboration wherever possible.

Main Research

The core research in the lab is broadly divided into two main themes:

1. Evolutionary Genomics.

Since moving to UNSW, a major focus of the lab has been the exploitation of genomic and post-genomic data to understand biological function and adaptation to novel environments. We work closely with the Ramaciotti Centre for Genomics and are involved in numerous de novo whole genome sequencing and assembly projects, using short read (Illumina), long read (PacBio & Nanopore) and linked read (10x Chromium) sequencing. The biggest of these is leading the bioinformatics and assembly effort in a consortium to sequence the cane toad genome, and leading the BABS Genome project to sequence two iconic Australian snakes. We are a member of the Oz Mammals Genomics initiative, assisting with the sequencing and assembly of Australia’s unique marsupial fauna. In 2018, we were selected as part of a team to sequence the Waratah genome as part of the pilot phase for the new Genomics of Australian Plants initiative.

We enjoy bringing our bioinformatics to bear on a variety of collaborative research projects. Most notably, we have an ARC Linkage Grant with Microbiogen Pty Ltd to understand how a strain of Saccharomyces cerevisiae has evolved to efficiently use xylose as a sole carbon source: something vital for second-generation biofuel production that wild yeast cannot do. We are combining comparative genomics, evolutionary genetics, RNA-Seq transcriptomics, and competition assays to understand how the novel metabolism evolved. Through deep Illumina resequencing of evolving populations, and assembling reliable complete genomes of the founding ancestors, the ultimate goal is to trace how mutations have interacted with existing genetic variation during adaptive evolution. More recently, we have received an ARC Linkage Grant with the Royal Botanic Gardens and Domain Trust, to apply genomics approaches to the challenges of rainforest tree conservation in the face of climate change and invasive pathogens. We are also collaborating with industrial and academic partners to de novo sequence, assemble, annotate and interrogate the genomes of a selection of microbes with interesting metabolic abilities.

2. Short Linear Motifs (SLiMs).

Many protein-protein interactions are mediated by Short Linear Motifs (SLiMs): short stretches of proteins (5-15 amino acids long), of which only a few positions are critical to function. These motifs are vital for biological processes of fundamental importance, acting as ligands for molecular signalling, post-translational modifications and subcellular targeting. SLiMs have extremely compact protein interaction interfaces, generally encoded by less than 4 major affinity-/specificity-determining residues. Their small size enables high functional density and evolutionary plasticity, making them frequent products of convergent “ex nihilo” evolution. It also makes them challenging to identify, both experimentally and computationally.

A major focus of the lab is the computational prediction of SLiMs from protein sequences. This research originated with Rich’s postdoctoral research, during which he developed a sequence analysis methods for the rational design of biologically active short peptides. He subsequently developed SLiMDisc, one of the first algorithms for successfully predicting novel SLiMs from sequence data - and coined the term “SLiM” into the bargain. This subsequently lead to the development of SLiMFinder, the first SLiM prediction algorithm able to estimate the statistical significance of motif predictions. SLiMFinder greatly increased the reliability of predictions. SLiMFinder has since spawned a number of motif discovery tools and webservers and is still arguably the most successful SLiM prediction tool on benchmarking data. Methods are made available through the SLiMSuite bioinformatics package and webservers.

Current research is looking to develop these SLiM prediction tools further and apply them to important biological questions. Of particular interest is the molecular mimicry employed by viruses to interact with host proteins and the role of SLiMs in other diseases, such as cancer. Other work is concerned with the evolutionary dynamics of SLiMs within protein interaction networks.

OTHER RESEARCH PROJECTS

In addition to the main research in the lab, the lab has a number of interdisciplinary collaborative projects applying bioinformatics tools and molecular evolution theory to experimental biology, often using large genomic, transcriptomic and/or proteomic datasets. These projects often involve the development of bespoke bioinformatics pipelines and a number of open source bioinformatics tools have been generated as a result. Please see the Publications and Lab software pages for more detail, or get in touch if something catches your eye and you want to find out more. We frequently have small collaborations and/or undergraduate student research projects. Many of these are “on hold” waiting for the right person, or sometimes data, to come along. If you think that you have what it needs, get in touch!

Thursday, 23 May 2019

Complete genome sequences of pooled genomic DNA from 10 marine bacteria using PacBio long-read sequencing

Song W, Thomas T & Edwards RJ (2019) Complete genome sequences of pooled genomic DNA from 10 marine bacteria using PacBio long-read sequencing. Marine Genomics 48:100687. DOI: 10.1016/j.margen.2019.05.002

Abstract

Background

High-quality, completed genomes are important to understand the functions of marine bacteria. PacBio sequencing technology provides a powerful way to obtain high-quality completed genomes. However individual library production is currently still costly, limiting the utility of the PacBio system for high-throughput genomics. Here we investigate how to generate high-quality genomes from pooled marine bacterial genomes.

Results

Pooled genomic DNA from 10 marine bacteria were subjected to a single library production and sequenced with eight SMRT cells on the PacBio RS II sequencing platform. In total, 7.35 Gbp of long-read data was generated, which is equivalent to an approximate 168× average coverage for the input genomes. Genome assembly showed that eight genomes with average nucleotide identities (ANI) lower than 91.4% can be assembled with high-quality and completion using standard assembly algorithms (e.g. HGAP or Canu). A reference-based reads phasing step was developed and incorporated to assemble the complete genomes of the remaining two marine bacteria that had an ANI > 97% and whose initial assemblies were highly fragmented.

Conclusions

Ten complete high-quality genomes of marine bacteria were generated. The findings and developments made here, including the reference-based read phasing approach for the assembly of highly similar genomes, can be used in the future to design strategies to sequence pooled genomes using long-read sequencing.

Friday, 10 May 2019

Whole genome sequencing of a novel, dichloromethane-fermenting Peptococcaceae from an enrichment culture

Holland​ SI*, Edwards​ RJ*, Ertan H, Wong YK, Russell TL, Deshpande NP, Manefield M & Lee MJ​ (2019) Whole genome sequencing of a novel, dichloromethane-fermenting Peptococcaceae from an enrichment culture. PeerJ 7:e7775. DOI 10.7717/peerj.7775 [*Joint first authors]

Abstract

Bacteria capable of dechlorinating the toxic environmental contaminant dichloromethane (DCM, CH2Cl2) are of great interest for potential bioremediation applications. A novel, strictly anaerobic, DCM-fermenting bacterium, “DCMF”, was enriched from organochlorine-contaminated groundwater near Botany Bay, Australia. The enrichment culture was maintained in minimal, mineral salt medium amended with dichloromethane as the sole energy source. PacBio whole genome SMRTTM sequencing of DCMF allowed de novo, gap-free assembly despite the presence of cohabiting organisms in the culture. Illumina sequencing reads were utilised to correct minor indels. The single, circularised 6.44 Mb chromosome was annotated with the IMG pipeline and contains 5,773 predicted protein-coding genes. Based on 16S rRNA gene and predicted proteome phylogeny, the organism appears to be a novel member of the Peptococcaceae family. The DCMF genome is large in comparison to known DCM-fermenting bacteria and includes 96 predicted methylamine methyltransferases, which may provide clues to the basis of its DCM metabolism. Full annotation has been provided in a custom genome browser and search tool, in addition to multiple sequence alignments and phylogenetic trees for every predicted protein, available at http://www.slimsuite.unsw.edu.au/research/dcmf/.

Monday, 18 February 2019

Paris Thompson (BABS3301 student)

Paris Thompson is an advanced science student at UNSW majoring in Genetics and Microbiology in her final trimester of undergrad coursework. She is doing the Biomolecular Science Laboratory Project (BABS3301) with the Edwards Lab and working on the Yeast genomes project.

Wednesday, 3 October 2018

Research snapshot - October 2018

One of the most important, interesting and challenging questions in biology is how new traits evolve at the molecular level. My lab employs sequence analysis techniques to interrogate protein and DNA sequences for the signals left behind by evolution. We are a bioinformatics lab but like to incorporate bench data through collaboration wherever possible.

Main Research

The core research in the lab is broadly divided into three main themes:

1. Short Linear Motifs (SLiMs)

Many protein-protein interactions are mediated by Short Linear Motifs (SLiMs): short stretches of proteins (5-15 amino acids long), of which only a few positions are critical to function. These motifs are vital for biological processes of fundamental importance, acting as ligands for molecular signalling, post-translational modifications and subcellular targeting. SLiMs have extremely compact protein interaction interfaces, generally encoded by less than 4 major affinity-/specificity-determining residues. Their small size enables high functional density and evolutionary plasticity, making them frequent products of convergent "ex nihilo" evolution. It also makes them challenging to identify, both experimentally and computationally.

A major focus of the lab is the computational prediction of SLiMs from protein sequences. This research originated with Rich’s postdoctoral research, during which he developed a sequence analysis methods for the rational design of biologically active short peptides. He subsequently developed SLiMDisc, one of the first algorithms for successfully predicting novel SLiMs from sequence data - and coined the term “SLiM” into the bargain. This subsequently lead to the development of SLiMFinder, the first SLiM prediction algorithm able to estimate the statistical significance of motif predictions. SLiMFinder greatly increased the reliability of predictions. SLiMFinder has since spawned a number of motif discovery tools and webservers and is still arguably the most successful SLiM prediction tool on benchmarking data.

Current research is looking to develop these SLiM prediction tools further and apply them to important biological questions. Of particular interest is the molecular mimicry employed by viruses to interact with host proteins and the role of SLiMs in other diseases, such as cancer. Other work is concerned with the evolutionary dynamics of SLiMs within protein interaction networks.

2. The evolution of novel functions.

Previous work in the lab has focused on the evolution of functional specificity following gene duplication. Since moving to UNSW, activities have shifted more towards the use of PacBio long read sequencing and other cutting-edge sequencing technologies, working closely with the Ramaciotti Centre for Genomics. We are collaborating with industrial and academic partners to de novo sequence, assemble, annotate and interrogate the genomes of a selection of microbes with interesting metabolic abilities. Most notably, we have an ARC Linkage Grant with Microbiogen Pty Ltd. to understand how a strain of Saccharomyces cerevisiae has evolved to efficiently use xylose as a sole carbon source: something vital for second generation biofuel production that wild yeast cannot do. We are combining comparative genomics, evolutionary genetics, RNA-Seq transcriptomics, and competition assays to understand how the novel metabolism evolved. Through deep Illumina resequencing of evolving populations, and assembling reliable complete genomes of the founding ancestors, the ultimate goal is to trace how mutations have interacted with existing genetic variation during adaptive evolution.

3. Whole genome sequencing and assembly.

Following our experiences with de novo whole genome assembly in yeast, the lab is getting involved in an increasing number of genome sequencing projects. The biggest of these is leading the bioinformatics and assembly effort in a consortium to sequence the cane toad genome. The lab is also leading the BABS Genome project two iconic Australian snakes for use in teaching and public engagement. We are a member of the Oz Mammals Genomics initiative, assisting with the sequencing and assembly of Australia's unique marsupial fauna. We also have an number of bacterial long-read whole genome sequencing collaborations.

Other Research Projects

In addition to the main research in the lab, the lab has a number of interdisciplinary collaborative projects applying bioinformatics tools and molecular evolution theory to experimental biology, often using large genomic, transcriptomic and/or proteomic datasets. These projects often involve the development of bespoke bioinformatics pipelines and a number of open source bioinformatics tools have been generated as a result. We frequently have small collaborations and/or undergraduate student research projects. Many of these are “on hold” waiting for the right person, or sometimes data, to come along. If you think that you have what it needs, get in touch!

Previous Research

The lab has been involved in a number of interdisciplinary collaborative projects applying bioinformatics tools and molecular evolution theory to experimental biology, often using large genomic, transcriptomic and/or proteomic datasets. These projects often involved the development of bespoke bioinformatics pipelines and a number of open source bioinformatics tools have been generated as a result. Please see the Publications and Lab software pages for more detail, or get in touch if something catches your eye and you want to find out more.

Tuesday, 7 August 2018

Draft genome assembly of the invasive cane toad, Rhinella marina

Richard J Edwards, Daniel Enosi Tuipulotu, Timothy G Amos, Denis O’Meally, Mark F Richardson, Tonia L Russell, Marcelo Vallinoto, Miguel Carneiro, Nuno Ferrand, Marc R Wilkins, Fernando Sequeira, Lee A Rollins, Edward C Holmes, Richard Shine & Peter A White (2018): Draft genome assembly of the invasive cane toad, Rhinella marina. GigaScience 7(9):giy095. [GigaScience] [PubMed] [PDF]

Abstract

Background. The cane toad (Rhinella marina formerly Bufo marinus) is a species native to Central and South America that has spread across many regions of the globe. Cane toads are known for their rapid adaptation and deleterious impacts on native fauna in invaded regions. However, despite an iconic status, there are major gaps in our understanding of cane toad genetics. The availability of a genome would help to close these gaps and accelerate cane toad research.

Findings. We report a draft genome assembly for R. marina, the first of its kind for the Bufonidae family. We used a combination of long read PacBio RS II and short read Illumina HiSeq X sequencing to generate a total of 359.5 Gb of raw sequence data. The final hybrid assembly of 31,392 scaffolds was 2.55 Gb in length with a scaffold N50 of 168 kb. BUSCO analysis revealed that the assembly included full length or partial fragments of 90.6% of tetrapod universal single-copy orthologs (n = 3950), illustrating that the gene-containing regions have been well-assembled. Annotation predicted 25,846 protein coding genes with similarity to known proteins in SwissProt. Repeat sequences were estimated to account for 63.9% of the assembly.

Conclusion. The R. marina draft genome assembly will be an invaluable resource that can be used to further probe the biology of this invasive species. Future analysis of the genome will provide insights into cane toad evolution and enrich our understanding of their interplay with the ecosystem at large.

(More details to follow in future posts.)

Friday, 30 June 2017

Research Snapshot - June 2017

Research interests in the Edwards lab stem from a fascination with the molecular basis of evolutionary change and how we can harness the genetic sequence patterns left behind to make useful predictions about contemporary biological systems. We are a bioinformatics lab but like to incorporate bench data through collaboration wherever possible.

Main Research

The core research in the lab is broadly divided into three main themes:

1. Short Linear Motifs (SLiMs)

SLiMs are short regions of proteins that mediate interactions with other proteins. A major focus of the lab is the computational prediction of SLiMs from protein sequences. This research originated with Rich’s postdoctoral research, during which he developed a sequence analysis methods for the rational design of biologically active short peptides. He subsequently developed SLiMDisc, one of the first algorithms for successfully predicting novel SLiMs from sequence data - and coined the term “SLiM” into the bargain. This subsequently lead to the development of SLiMFinder, the first SLiM prediction algorithm able to estimate the statistical significance of motif predictions. SLiMFinder greatly increased the reliability of predictions. SLiMFinder has since spawned a number of motif discovery tools and webservers and is still arguably the most successful SLiM prediction tool on benchmarking data.

Current research is looking to develop these SLiM prediction tools further and apply them to important biological questions. Of particular interest is the molecular mimicry employed by viruses to interact with host proteins and the role of SLiMs in other diseases, such as cancer. Other work is concerned with the evolutionary dynamics of SLiMs within protein interaction networks.

2. The evolution of novel functions.

Previous work in the lab has focused on the evolution of functional specificity following gene duplication. Since moving to UNSW, activities have shifted more towards the use of PacBio long read sequencing and other cutting-edge sequencing technologies, working closely with the Ramaciotti Centre for Genomics. We are collaborating with industrial and academic partners to de novo sequence, assemble, annotate and interrogate the genomes of a selection of microbes with interesting metabolic abilities. Most notably, we have an ARC Linkage Grant with Microbiogen Pty Ltd. to understand how a strain of Saccharomyces cerevisiae has evolved to efficiently use xylose as a sole carbon source: something vital for second generation biofuel production that wild yeast cannot do. This project combines detailed molecular characterisation of highly adapted yeast strains with “molecular palaeontology” to trace the evolutionary process and identify functionally significant loci under selection.

3. Whole genome sequencing and assembly.

Following our experiences with de novo whole genome assembly in yeast, the lab is getting involved in an increasing number of genome sequencing projects. The biggest of these is leading the bioinformatics and assembly effort in a consortium to sequence the cane toad genome. The lab is also leading the BABS Genome project to sequence iconic Australian species for use in teaching and public engagement.

Previous Research

The lab has been involved in a number of interdisciplinary collaborative projects applying bioinformatics tools and molecular evolution theory to experimental biology, often using large genomic, transcriptomic and/or proteomic datasets. These projects often involved the development of bespoke bioinformatics pipelines and a number of open source bioinformatics tools have been generated as a result. Please see the Publications and Lab software pages for more detail, or get in touch if something catches your eye and you want to find out more.

Monday, 9 January 2017

Peter Santosa (SVRS Student)

Peter is a 3rd year Advance Science student who worked in the lab in January-February 2017 on the Summer Vacation Research Scholarship (SVRS). Peter was working on a bacterial sequencing project in collaboration with Mike Manefield at UNSW. We have successfully used PacBio sequencing to fully and contiguously assemble the genome of a new bacterial strain from a mixed culture. Peter’s project was analysing assembled contigs from other organisms in the culture.

Peter is undertaking a double major of molecular and cell biology and microbiology at UNSW.

Saturday, 26 November 2016

The cane toad genome project

What are we doing? The Edwards Lab is part of an Australian, Portuguese and Brazilian consortium led by Peter White to sequence and assemble the genome of the cane toad (Rhinella marina). We are leading the bioinformatics component of the assembly effort.

How are we doing it? We are using a combination of Illumina (HiSeq X and NovaSeq) short read sequencing, PacBio (RS II) long read and sequence and 10x Genomics Chromium linked reads.

Details to follow. Please get in touch if you are interested in the project.

Opportunities

Honours and postgraduate* projects are available to work on the assembly and annotation. (*PhD students should have their own scholarship.)

Consortium members

Miguel Carneiro (CIBIO-InBIO), Richard Edwards (UNSW), Nuno Ferrand (CIBIO-InBIO), Eddie Holmes (U Sydney), Craig Moritz (ANU), Lee Ann Rollins (Deakin), Fernando Sequeira (CIBIO-InBIO), Rick Shine (U Sydney), Marcelo Vallinoto de Souza (Federal University of Pará & CIBIO-InBIO), Peter White (UNSW), Marc Wilkins (UNSW).