Showing posts with label biodiversity. Show all posts
Showing posts with label biodiversity. Show all posts

Saturday, 15 March 2025

A genome assembly and annotation for the Australian alpine skink Bassiana duperreyi using long-read technologies.

Hanrahan BJ, Alreja K, Reis ALM, Chang JK, Dissanayake DSB, Edwards RJ, Bertozzi T, Hammond JM, O’Meally D, Deveson IW, Georges A, Waters P & Patel HR (accepted): A genome assembly and annotation for the Australian alpine skink Bassiana duperreyi using long-read technologies. G3 jkaf046, DOI: 10.1093/g3journal/jkaf046 [G3] [PubMed]

Abstract

The eastern three-lined skink (Bassiana duperreyi) inhabits the Australian high country in the southeast of the continent including Tasmania. It is a distinctive oviparous species because it undergoes sex reversal (from XX genotypic females to phenotypic males) at low incubation temperatures. We present a chromosome-scale genome assembly of a Bassiana duperreyi XY male individual, constructed using PacBio HiFi and ONT long reads scaffolded using Illumina HiC data. The genome assembly length is 1.57 Gbp with a scaffold N50 of 222 Mbp, N90 of 26 Mbp, 200 gaps and 43.10% GC content. Most (95%) of the assembly is scaffolded into 6 macrochromosomes, 8 microchromosomes and the X chromosome, corresponding to the karyotype. Fragmented Y chromosome scaffolds (n=11 > 1 Mbp) were identified using Y-specific contigs generated by genome subtraction. We identified two novel alpha-satellite repeats of 187 bp and 199 bp in the putative centromeres that did not form higher order repeats. The genome assembly exceeds the standard recommended by the Earth Biogenome Project; 0.02% false expansions, 99.63% kmer completeness, 94.66% complete single copy BUSCO genes and an average 98.42% of transcriptome data mappable to the genome assembly. The mitochondrial genome (17,506 bp) and the model rDNA repeat unit (15,154 bp) were assembled. The B. duperreyi genome assembly has high completeness for a skink and will provide a resource for research focused on sex determination and thermolabile sex reversal, as an oviparous foundation species for studies of the evolution of viviparity, and for other comparative genomics studies of the Scincidae.

Friday, 14 March 2025

Chromosome-level genome assembly of the spangled emperor, Lethrinus nebulosus (Forsskål 1775)

The first data note from the Ocean Genomes project is now out in Scientific Data - a first-in-family genome from the emperor breams.

Parata L, Anstiss L, de Jong E, Doran A, Edwards RJ, Newman SJ, Payet SD, Skepper CL, Wakefield CB, OceanOmics Centre, OceanOmics Division & Corrigan S (2025): Chromosome-level genome assembly of the spangled emperor, Lethrinus nebulosus (Forsskål 1775). Scientific Data 12:435. DOI: 10.1038/s41597-025-04690-w.

Abstract

Spangled emperor, Lethrinus nebulosus (Forsskål 1775), is a tropical marine fish of economic and cultural importance throughout the Indo-West Pacific. It is one of the most targeted recreational fishes in the Gascoyne Coast Bioregion of Western Australia where it serves as an indicator species for recreational fishing. Here, we present a highly accurate, near-gapless, chromosome-level, haplotype-phased reference genome assembly of L. nebulosus (Lethrinus nebulosus (Spangled Emperor) genome, fLetNeb1.1; PRJNA1074345), the first for the species and the first high-quality genome representative of the family Lethrinidae. The 1.09 Gb genome was assembled from PacBio HiFi and Dovetail Omni-C proximity ligation sequencing data. The contig N50 is 21–24 Mbp and BUSCO completeness greater than 99%. A preliminary gene annotation identified 24,583 genes with the predicted transcriptome achieving a BUSCO completeness score of 99.1%. This resource will facilitate genomic studies to inform the sustainable management of L. nebulosus and other Lethrinids.

Thursday, 13 February 2025

Small but mitey: a gapless telomere-to-telomere assembly of an unidentified mite with a streamlined genome

A few years ago, we accidentally sequenced an interesting mite when making the Rhodamnia argentea reference genome. We don’t really know what it is, but thanks to nanopore sequencing and a low repeat content, we produced a gapless telomere-to-telomere assembly! As an unfunded passion project, it’s been ticking along in the background since then, but is now published in Genome Biology and Evolution.

One of the most interesting aspects of the paper is the genome reduction - despite being a complete nuclear genome, the assembly is under 35 megabases! This is a couple of orders of magnitude smaller than many other arachnid genomes.

We are yet to do a comprehensive analysis, but just looking at the core set of expected arachnid genes, as defined by BUSCO (e.g. single-copy genes that are expected to be present in most arachnid genomes) revealed that about a third of them were missing. (This low completeness was partly responsible for the time taken to fully recognised and deal with the contamination.) Tellingly, the closest sequenced relatives of this mite also have reduced genomes and are missing many of the same BUSCO genes, revealing a long history of gene loss.

The proportion of “Duplicated” BUSCO genes is also surprisingly high at 4.5%. These are genuine duplications, with consistent diploid read depths and many of the pairs present on both chromosomes. It will be interesting to see if these duplications are replacing some of the lost functions from the other genes, or are novel genes behind such a specialised lifestyle. As an unfunded passion project, it was beyond scope to investigate the full annotation as part of this paper, but get in touch if you would be interested to do this!

Edwards RJ, Chen SH, Halliday B & Bragg JG (2025): Small but mitey: a gapless telomere-to-telomere assembly of an unidentified mite with a streamlined genome. Genome Biology and Evolution Feb 13. [Gen Biol Evol] [PubMed]

Abstract

A draft assembly of the rainforest tree Rhodamnia argentea Benth. (malletwood, Myrtaceae) revealed contaminating DNA sequences that most closely matched those from mites in the family Eriophyidae. Eriophyoid mites are plant parasites that often induce galls or other deformities on their host plants. They are notable for their small size (averaging 200 μm), distinctive four-legged body structure, and heavily streamlined genomes, which are among the smallest known of all arthropods. Contaminating mite sequences were assembled into a high-quality gapless telomere-to-telomere nuclear genome. The entire genome was assembled on two fully contiguous chromosomes, capped with a novel TTTGG or TTTGGTGTTGG telomere sequence, and exhibited clear signs of genome reduction (34.5 Mbp total length, 68.6% arachnid Benchmarking Universal Single-Copy Ortholog completeness). Phylogenomic analysis confirmed that this genome is that of a previously unsequenced eriophyoid mite. Despite its unknown identity, this complete nuclear genome provides a valuable resource to investigate invertebrate genome reduction.

Monday, 27 January 2025

A reference genome for the eastern bettong (Bettongia gaimardi)

Silver LW, Edwards RJ, Neaves L,A Manning, CJ Hogg & S Banks (2025): A reference genome for the eastern bettong (Bettongia gaimardi) [version 2; peer review: 3 approved]. F1000Research 13:1544. [F1000Res] [PubMed]

Abstract

The eastern or Tasmanian bettong (Bettongia gaimardi) is one of four extant bettong species and is listed as ‘Near Threatened’ by the IUCN. We sequenced short read data on the 10x system to generate a reference genome 3.46Gb in size and contig N50 of 87.36Kb and scaffold N50 of 2.93Mb. Additionally, we used GeMoMa to provide and accompanying annotation for the reference genome. The generation of a reference genome for the eastern bettong provides a vital resource for the conservation of the species.

Tuesday, 17 December 2024

Parental assigned chromosomes for cultivated cacao provides insights into genetic architecture underlying resistance to vascular streak dieback

As a big fan of chocolate, it was great to have the opportunity to help out with a cacao genome. This study presents the first diploid, fully scaffolded, and parentally phased genome resource for Theobroma cacao L. to provide insights into the genetic architecture underlying resistance and susceptibility to vascular streak dieback (VSD), a significant threat to cacao production in Southeast Asia and Melanesia. By analyzing NLR gene clusters and other disease response gene candidates in proximity to informative QTLs, the research identifies structural variants within NLRs inherited from resistant and susceptible parents, offering potential breeding targets for VSD resistance.

Tobias PA, Downs J, Epaina P, Singh G, Park RF, Edwards RJ, Brugman E, Zulkifli A, Muhammad J, Purwantara A & Guest DI (2024): Parental assigned chromosomes for cultivated cacao provides insights into genetic architecture underlying resistance to vascular streak dieback. The Plant Genome doi: 10.1002/tpg2.20524. [Plant Genome] [bioRxiv] [PubMed]

Abstract

Diseases of Theobroma cacao L. (Malvaceae) disrupt cocoa bean supply and economically impact growers. Vascular streak dieback (VSD), caused by Ceratobasidium theobromae, is a new encounter disease of cacao currently contained to southeast Asia and Melanesia. Resistance to VSD has been tested with large progeny trials in Sulawesi, Indonesia, and in Papua New Guinea with the identification of informative quantitative trait loci (QTLs). Using a VSD susceptible progeny tree (clone 26), derived from a resistant and susceptible parental cross, we assembled the genome to chromosome-level and discriminated alleles inherited from either resistant or susceptible parents. The parentally phased genomes were annotated for all predicted genes and then specifically for resistance genes of the nucleotide-binding site leucine-rich repeat class (NLR). On investigation, we determined the presence of NLR clusters and other potential disease response gene candidates in proximity to informative QTLs. We identified structural variants within NLRs inherited from parentals. We present the first diploid, fully scaffolded, and parentally phased genome resource for T. cacao L. and provide insights into the genetics underlying resistance and susceptibility to VSD.

Tuesday, 5 November 2024

#ABACBS2024 Poster 102: Improving phased Hifiasm assemblies with 20 kb ONT reads

After a great presentation this morning by Emma de Jong on our High-Quality Genomes for Australian Lutjanidae Species (abstract below), if you’re at ABACBS2024 then please drop by Poster #102 to find out about some of the work we’re doing with ONT data.

Abstracts

Improving phased Hifiasm assemblies with 20 kb ONT reads

Richard J Edwards, Adrianne Doran, Emma de Jong, Lara Parata, Shannon Corrigan

The quality and quantity of genome assembly has improved dramatically over recent years. Many large-scale genome projects combine assembly of HiFi and HiC reads using Hifiasm to produce contiguous phased assemblies, scaffolded to chromosome-level. Nevertheless, HiFi reads are typically under 25 kb and can still struggle to assemble long, low diversity repeat regions. Obtaining ultra-long (100 kb or longer) ONT reads to solve this problem remains a significant challenge due to technical constraints and DNA sample requirements. Here, we explore the utility of using standard ONT long reads (20 kb or more) as “ultra-long” input to improve phased Hifiasm assemblies for 22 species of bony fish (Genome Size, 627 Mb 1.54 Gb). We also explore whether the new --telo-m mode in Hifiasm v0.9.0 improves telomere prediction. Incorporating 20kb+ ONT reads (7.8X 93.5X) significantly increased assembly contiguity. BUSCO Completeness was not significantly altered, although there was some re-partitioning of BUSCO genes between phased haplotypes for some species. Improvement did not strongly correlate with read depth (either HiFi or ONT), suggesting that the underlying read length distributions and/or specific genome features are more important for determining the outcome. Hifiasm --telo-m mode significantly increased telomere recovery, assembling over six times the number of gapless telomere-to-telomere chromosomes when combined with 20kb+ ONT reads. Verification of how these results translate to the ease of curation and/or quality of final HiC-scaffolded chromosome-level assemblies is ongoing, with a goal to determine whether the additional sample preparation and sequencing in the lab is cost-effective.

High-Quality Genomes for Australian Lutjanidae Species

Emma de Jong, Lara Parata, Philipp E Bayer, Shannon Corrigan, Richard J Edwards

Lutjanidae (snappers) are highly valued in commercial and recreational fisheries worldwide and serve as indicator species of the health of marine environments and fishery bioregions in Western Australia. Comprehensive genomic mapping of immune gene families of Lutjanidae species are lacking, but this information is critical for understanding disease vulnerability, the impact of environmental stress, improving aquaculture efforts and to provide insights into the health of wild populations. Despite their importance, only 3 out of 113 Lutjanid species currently have available reference genomes, two of which are highly fragmented (>11,000 and >200,000 contigs), impacting studies on gene families relevant to aquaculture. In this study, we present high-quality chromosome-level reference genomes for 14 Australian Lutjanidae species across seven genera, generated using HiFi and HiC data. We present initial comparative genomic analyses, including immune gene content and chromosomal synteny analyses across species. These analyses provide insights into the genomic architecture and evolutionary relationships within Lutjanidae. Ongoing work aims to comprehensively map and compare the immune gene family repertoire across Lutjanidae genera, as well as Lethrinidae species as an outgroup, to determine genus-specific changes in genes (e.g. loss, selection, duplication) important for pathogen detection, antigen presentation, inflammation, and immune memory. These genome assemblies will serve as a foundational resource to the wider scientific community interested in Lutjanidae

Sunday, 3 November 2024

Chromosome-level genome assembly of the Australian rainforest tree Rhodamnia argentea (malletwood)

Genome projects don’t always go according to plan, and when we first sequenced Rhodamnia argentea with 10x Genomics linked reads, we accidentally sequenced a parasite along with it. This was quite hard to identify from the sequencing data itself, as the depth of sequencing was quite high, and we were unable to identify the guilty bug itself, which is probably microscopic. Getting to the bottom of this took a back seat for a while when the focus of the project shifted to Melaleuca quinquenervia, but with the addition of ONT reads and Hi-C, we have now been able to generate a chromosome-level decontaminated assembly. (An assembly of the contaminating mite will follow…)

Chen SH, Jones A, Lu-Irving P, Yap JYS, van der Merwe M, Bragg JG & Edwards RJ (2024): Chromosome-level genome assembly of the Australian rainforest tree Rhodamnia argentea (malletwood). Genome Biology and Evolution 16(11):evae238. [Gen Biol Evol] [PubMed]

Abstract

Myrtaceae are a large family of woody plants, including hundreds that are currently under threat from the global spread of a fungal pathogen, Austropuccinia psidii (G. Winter) Beenken, which causes myrtle rust. A reference genome for the Australian native rainforest tree Rhodamnia argentea Benth. (malletwood) was assembled from Oxford Nanopore Technologies long-reads, 10x Genomics Chromium linked-reads, and Hi-C data (N50 = 32.3 Mb and BUSCO completeness 98.0%) with 99.0% of the 347 Mb assembly anchored to 11 chromosomes (2n = 22). The R. argentea genome will inform conservation efforts for Myrtaceae species threatened by myrtle rust, against which it shows variable resistance. We observed contamination in the sequencing data, and further investigation revealed an arthropod source. This study emphasizes the importance of checking sequencing data for contamination, especially when working with nonmodel organisms. It also enhances our understanding of a tree that faces conservation challenges, contributing to broader biodiversity initiatives.

Thursday, 31 October 2024

BioDiversity Genomics Conference (BG24)

The global BioDiversity Genomics Conference (BG24) was another great success this year. As well as running a session on Marine Vertebrate Genomics, it was particularly rewarding to see some many quality contributions from lab members and alumni.

Biodiversity Genomics in Australasia

Jessica Pearce (UWA): Reconstructing tiger shark history using genomics

Sharks and rays are a clade of high evolutionary, ecological, economic, and cultural significance, and yet they are one of the most threatened taxa groups in the marine environment. Despite this, there remains a lack of molecular resources for this class to assist with their conservation. The tiger shark (Galeocerdo cuvier) is a near threatened, keystone species distributed circumglobally that is under substantial pressure from human impacts, making it a high priority for management worldwide. We sequenced and characterised a reference quality assembly for the tiger shark, the first genome for this family, and used this to dive deeper into the evolution, adaption, and demographic history of this ancient species. We investigated how its effective population size (Ne), genome-wide heterozygosity and inbreeding has changed over time to infer how this species has responded to past global events, and hence potential responses to ongoing and future accumulating threats. This aims to assist in effective management of this high-profile species. 

Katarina Stuart (University of Auckland): Lifetime fitness is correlated more strongly with structural variant than SNP mutational load in a threatened bird species

Conservation genomics is becoming increasingly interested in whether structural variant (SV) information can help the management of threatened species. The functional consequences of SVs are more complex than for single nucleotide polymorphisms (SNPs) and thus may be more likely to contribute to load. While the impacts of SV-specific genetic load may be less consequential for large populations, the interplay between weakened selection and stochastic processes mean that smaller populations, like those of the threatened Aotearoa hihi/New Zealand stitchbird (Notiomystis cincta), may harbour a high SV load. Hihi were once confined to a single remnant population, but have been reestablished into six sanctuaries and reserves, often via secondary bottlenecks, resulting in low genetic diversity, low adaptive potential and inbreeding depression. In this study, we use whole genome resequencing of 30 individuals from the Tiritiri Matangi population to identify the nature and distribution of both SNPs and SVs within this small avian population. We find that SNP and SV individual mutation load is only moderately correlated, likely because SVs arise in regions of high recombination and reduced evolutionary conservation. Finally, we leverage a long-term monitoring dataset of pedigree and fitness data to assess the impact of SNP and SV mutation load on individual fitness, and demonstrate that SV load correlates more strongly than SNP load with lifetime fitness. The results of this study indicate that only examining SNPs neglects important aspects of intraspecific variation, and that studying SVs has direct implications for linking genetic diversity and genetic health to inform management decisions.

Richard Edwards (UWA): Improving Hifiasm assemblies with 20 kb ONT reads

The quality and quantity of genome assembly has improved dramatically over recent years. Many large-scale genome projects assemble HiFi and HiC reads using Hifiasm to produce contiguous phased assemblies, scaffolded to chromosome-level. Nevertheless, HiFi reads are typically under 25 kb and can still struggle to assemble long, low-diversity repeat regions. Obtaining ‘ultra-long’ (100 kb or longer) ONT reads to solve this problem remains a significant challenge due to technical constraints and DNA sample requirements. Here, we explore the utility of using standard ONT long reads (20 kb or more) as ‘ultra-long’ input to improve phased Hifiasm assemblies for 22 species of bony fish (Genome Size, 627 Mb - 1.54 Gb). We also explore whether the new ‘telo-m’ mode in Hifiasm v0.9.0 improves telomere prediction in these species. Incorporating 20+ kb ONT reads (7.8X - 93.5X) significantly increased assembly contiguity. BUSCO completeness was not significantly altered, although there was some re-partitioning of BUSCO genes between phased haplotypes for some species. Improvement did not strongly correlate with read depth (neither HiFi nor ONT), suggesting that the underlying read length distributions and/or specific genome features are more important for determining the outcome. Hifiasm ‘telo-m’ mode significantly increased telomere recovery, assembling over six times the number of gapless telomere-to-telomere chromosomes when combined with incorporation of ONT reads. Verification of how these results translate to the quality and/or ease of curation of final HiC-scaffolded chromosome-level assemblies is ongoing, with a goal to determine whether the additional sample preparation and sequencing in the lab is cost-effective.

Emma de Jong (UWA): High-Quality Genomes for Australian Lutjanidae Species

Lutjanidae (snappers) are highly valued in commercial and recreational fisheries worldwide and some species serve as fisheries indicator species particularly for bioregions in Western Australia. Comprehensive genomic mapping of immune gene families of Lutjanidae species are lacking, but this information can inform understanding disease vulnerability, the impact of environmental stress, improving aquaculture efforts and to provide insights into the health of wild populations. Despite their importance, only 3 out of 113 Lutjanid species currently have available reference genomes, two of which are highly fragmented (>11,000 and >200,000 contigs), impacting studies on gene families relevant to aquaculture. In this study, we present high-quality chromosome-level reference genomes for 14 Australian lutjanid species across seven genera, generated using PacBio HiFi and Dovetail HiC data. We present initial comparative genomic analyses, including immune gene content and chromosomal synteny analyses across species. These analyses provide insights into the genomic architecture and evolutionary relationships within Lutjanidae. Ongoing work aims to comprehensively map and compare the immune gene family repertoire across genera in Lutjanidae, as well as lethrinid species as an outgroup, to determine genus-specific changes in genes (e.g., loss, selection, duplication) important for pathogen detection, antigen presentation, inflammation, and immune memory. These genome assemblies will serve as a foundational resource to the wider scientific community interested in these species.

Research of ECRs who work in biodiversity genomics

Lara Parata (UWA): Genome Evolution in Marine Ray-Finned Fishes

Approximately half of extant vertebrate species are fishes, with more than 30,000 species classified as ray-finned fishes (Actinopterygii). Actinopterygii represent diverse phenotypes, feeding strategies, life history traits and occupy distinct ecological niches, making them an ideal taxa for studying molecular drivers of diversity and adaptation. Despite their diversity, ecological, and economical importance, only 145 Illumina genome assemblies are available for marine Actinopterygii species. In this study we present 250 new marine Actinopterygii genome assemblies generated using Illumina whole genome sequencing and initial results from a large-scale study of these 395 genomes. Using reference-based annotation tools we determine which fish families have unique patterns of gene family frequency / structure (e.g., losses, expansions, contractions), and correlate these with predicted functional signatures to infer biological and ecological adaptations. We identify fish families with distinct rates of change in the gene families present within their genomes (e.g., more losses / expansions or diversity) and associate these patterns with increased rates of diversification or speciation to further elucidate the genomic attributes contributing to ecological success. The results of this work contribute to the growing understanding of fish genome evolution and provide new insights into the evolutionary history and ecological success of marine Actinopterygii.

Tuesday, 9 July 2024

New pre-print: The Genomics for Australian Plants (GAP) framework initiative – developing genomic resources for understanding the evolution and conservation of the Australian flora

The Bioplatforms Australia Genomics for Australian Plants (GAP) initiative aims to sequence and assemble representative genomes of Australia’s unique flora, which boasts over 24,000 native vascular plant species evolved over millions of years. The program brings together academic groups, herbaria and botanic gardens from across the country to build genomic capacity and create valuable resources for the classification, conservation and utilisation of Australian plants. We were lucky enough to sequence one of the first GAP species, the NSW Waratah. Now, the capstone paper outlining the project and its key findings from multiple species is out as a pre-print at EcoEvoRxiv:

Simpson L, Cantrill DJ, Byrne M, Allnutt TR, King GJ, Lum M, Al Bkhetan Z, Andrew R, Baker WJ, Barrett MD, Batley J, Berry O, Binks RM, Bragg JG, Broadhurst L, Brown G, Bruhl J, Edwards RJ, Ferguson S, Forest F, Gustafsson J, Hammer TA, Holmes GD, Jackson CJ, James EA, Jones A, Kersey PJ, Leitch IJ, Maurin O, McLay TGB, Murphy DJ, Nargar K, Nauheimer L, Sauquet H, Schmidt-Lebuhn AN, Shepherd KA, Syme AE, Waycott M, Wilson TC, Crayn DM (preprint): The Genomics for Australian Plants (GAP) framework initiative – developing genomic resources for understanding the evolution and conservation of the Australian flora. EcoEvoRxiv DOI: https://doi.org/10.32942/X2RP70

The generation and analysis of genome-scale data—genomics—is driving a rapid increase in plant biodiversity knowledge. However, the speed and complexity of technological advance in genomics presents challenges for its widescale use in evolutionary and conservation biology. Here, we introduce and describe a national-scale collaboration conceived to build genomic resources and capability for understanding the Australian flora: the Genomics for Australian Plants (GAP) Framework Initiative. We outline (a) the history of the project including the collaborative framework, partners, and funding; (b) GAP principles such as rigour in design, sample verification and documentation, data management, and data accessibility; and (c) the structure of the consortium and its four activity streams (reference genomes, phylogenomics, conservation genomics, and training), with the rationale and aims for each of them. We show, through discussion of its successes and challenges, the value of this multi-institutional consortium approach and the enablers, such as well-curated collections and national collaborative research infrastructure, all of which have led to a substantial increase in capacity and delivery of biodiversity knowledge outcomes.

The initiative is about more than just reference genomes, with core activity in phylogenomics, conservation genomics and training too. For more information on the project and the resources generated (with more to come), read the paper and/or visit the GAP website.

Tuesday, 2 July 2024

Extant and extinct bilby genomes combined with Indigenous knowledge improve conservation of a unique Australian marsupial

Hogg C*, Edwards RJ*, Farquharson K*, Silver L*, Brandies P, Peel E, Escalona M, Jaya FR, Thavornkanlapachai R, Batley K, Bradford TM, Chang JK, Chen Z, Deshpande N, Dziminski M, Ewart KM, Griffith OW, Marin Gual L, Moon KL, Travouillon KJ, Waters P, Whittington CM, Wilkins MR, Helgen KM, Lo N, Ho SYW, Ruiz Herrera A, Paltridge R, Marshall Graves JA, Renfree M, Shapiro B, Ottewell K, Kiwirrkurra Rangers & Belov K (2024): Extant and extinct bilby genomes combined with Indigenous knowledge improve conservation of a unique Australian marsupial. Nature Ecology & Evolution 8:1311–1326. [*Joint first authors] [Research Square] [Nat Ecol Evol] [PubMed]

Abstract

Ninu (greater bilby, Macrotis lagotis) are desert-dwelling, culturally and ecologically important marsupials. In collaboration with Indigenous rangers and conservation managers, we generated the Ninu chromosome-level genome assembly (3.66 Gbp) and genome sequences for the extinct Yallara (lesser bilby, Macrotis leucura). We developed and tested a scat single-nucleotide polymorphism panel to inform current and future conservation actions, undertake ecological assessments and improve our understanding of Ninu genetic diversity in managed and wild populations. We also assessed the beneficial impact of translocations in the metapopulation (N = 363 Ninu). Resequenced genomes (temperate Ninu, 6; semi-arid Ninu, 6; and Yallara, 4) revealed two major population crashes during global cooling events for both species and differences in Ninu genes involved in anatomical and metabolic pathways. Despite their 45-year captive history, Ninu have fewer long runs of homozygosity than other larger mammals, which may be attributable to their boom-bust life history. Here we investigated the unique Ninu biology using 12 tissue transcriptomes revealing expression of all 115 conserved eutherian chorioallantoic placentation genes in the uterus, an XY1Y2 sex chromosome system and olfactory receptor gene expansions. Together, we demonstrate the holistic value of genomics in improving key conservation actions, understanding unique biological traits and developing tools for Indigenous rangers to monitor remote wild populations.

Friday, 15 March 2024

Towards telomere-to-telomere fish genomes with Oxford Nanopore Technologies gap-filling

It was a pleasure to be invited to the “What You’re Missing Matters” tour at Perth, and present some ongoing work investigating the best way to incorporate Oxford Nanopore Technologies data into our high-quality HiFi+HiC fish genomes. The results are too preliminary to share here (and will soon be superseded) but do get in touch if it sounds interesting to you.

Sunday, 28 January 2024

Toward genome assemblies for all marine vertebrates: current landscape and challenges

The first Ocean Genomes paper is now out! This one is a small commentary piece, but some high-quality genomes are on their way - watch this space. Well, actually, watch this space at Genomes on a Tree!

de Jong E, Parata L, Bayer PE, Corrigan S & Edwards RJ (2024): Toward genome assemblies for all marine vertebrates: current landscape and challenges. Gigascience 13:giad119. [Gigascience] [PubMed]

Marine vertebrate biodiversity is fundamental to ocean ecosystem health but is threatened by climate change, overharvesting, and habitat degradation. High-quality reference genomes are valuable foundational scientific resources that can inform conservation efforts. Consequently, global consortia are striving to produce reference genomes for representatives of all life. Here, we summarize the current landscape of available marine vertebrate reference genomes, including their phylogenetic diversity and geographic hotspots of production. We discuss key logistical and technical challenges that remain to be overcome if we are to realize the vision of a comprehensive reference genome library of all marine vertebrates.

Friday, 15 December 2023

A high-quality pseudo-phased genome for Melaleuca quinquenervia shows allelic diversity of NLR-type resistance genes

Chen SH, Martino AM, Luo Z, Schwessinger B, Jones A, Tolessa T, Bragg JG, Tobias PA, Edwards RJ (2023): A high-quality pseudo-phased genome for Melaleuca quinquenervia shows allelic diversity of NLR-type resistance genes. GigaScience 12:giad102. [Gigascience] [PubMed]

Background. Melaleuca quinquenervia (broad-leaved paperbark) is a coastal wetland tree species that serves as a foundation species in eastern Australia, Indonesia, Papua New Guinea, and New Caledonia. While extensively cultivated for its ornamental value, it has also become invasive in regions like Florida, USA. Long-lived trees face diverse pest and pathogen pressures, and plant stress responses rely on immune receptors encoded by the nucleotide-binding leucine-rich repeat (NLR) gene family. However, the comprehensive annotation of NLR encoding genes has been challenging due to their clustering arrangement on chromosomes and highly repetitive domain structure; expansion of the NLR gene family is driven largely by tandem duplication. Additionally, the allelic diversity of the NLR gene family remains largely unexplored in outcrossing tree species, as many genomes are presented in their haploid, collapsed state.

Results. We assembled a chromosome-level pseudo-phased genome for M. quinquenervia and described the allelic diversity of plant NLRs using the novel FindPlantNLRs pipeline. Analysis reveals variation in the number of NLR genes on each haplotype, distinct clustering patterns, and differences in the types and numbers of novel integrated domains.

Conclusions. The high-quality M. quinquenervia genome assembly establishes a new framework for functional and evolutionary studies of this significant tree species. Our findings suggest that maintaining allelic diversity within the NLR gene family is crucial for enabling responses to environmental stress, particularly in long-lived plants.

Thursday, 21 September 2023

PAG Australia 2023: Exploring Dingo Ecology and Evolution with Chromosome-Level Canid Genomes

Richard J Edwards, Matt F Field and J William O Ballard - PAG Australia 2023

Dogs are uniquely associated with human dispersal and bring novel insight into human migration and the domestication process. Dingoes represent an intriguing case within canine evolution being geographically isolated for thousands of years. The exact origin(s) and people(s) who transported the canines that became dingoes to Australia is debated, but it has been suggested they arrived by boat ~5,000-8,000 BP. Published morphological and genetic evidence has established the presence of at least two dingo lineages. The Alpine dingo is commonly found in south-eastern Australia while the Desert ecotype is found in the north, central and western Australia. The relationship of dingoes to modern dogs, and the ecotypes to each other, has important implications management and protection of this top predator, as well as providing interesting perspectives on human colonisation and canine domestication.

We have generated chromosome-level assemblies of both dingo ecotypes, along with domesticated dogs representing both ancient (Basenji) and derived (German Shepherd) breeds. In each case, long-read sequencing and Hi-C scaffolding have been combined to produce genome assemblies with high contiguity and structural completeness. Comparison of these assemblies with additional dog breeds, using the Greenland wolf as an outgroup, places the dingo as an early offshoot of modern dogs, situated between the grey wolf and the domesticated dogs of today. This is supported by patterns of genetic variation, and chromosome structure. Furthermore, we confirm that dingoes have not experienced the expansion of the AMY2B pancreatic amylase gene that occurred during domestication of modern dogs. This has important implications for dingo ecology and behaviour, and raises the prospect of using AMY2B copy number as a novel and reliable in-field discriminator between dingoes and feral dogs.

Thursday, 29 June 2023

Three Pawsey Internship projects available for the Ocean Genomes Project

We have three Pawsey student internships available this summer with the Ocean Genomes Laboratory in the Minderoo OceanOmics Centre at UWA. Closing date: 07 August, 2023 at 17:00 AWST (Perth time). This is a 10-week, paid program open to exceptional undergrad (2nd/3rd year), Honours, Master’s and PhD students. Apply at the CSIRO Application page. Please get in touch if you want to know more and/or are interested in a student research project in the lab.

Optimising workflows for whole genome assembly for marine vertebrates (Project #04)

The biodiversity of marine vertebrates is critical for the health of our ocean’s ecosystem, but is under immediate threat from climate change, pollution, overfishing and habitat destruction. To advance our understanding of how best to protect and sustain our ocean life, global efforts are underway (such as the Vertebrate Genome Project; VGP) to establish a complete library of high-quality reference genomes for all ~22,000 marine vertebrates.

Reference genomes are pivotal not only for answering fundamental questions in marine biology and evolution, but also for guiding the conservation of species most at risk within our changing oceans, and for accurately monitoring biodiversity.

This project utilizes data generated in-house, either by Illumina short-read or PacBio high-fidelity long-read sequencing of Australian marine vertebrate species. The primary objective is to optimize analysis workflows on Pawsey, encompassing the entire life cycle of the data from its raw format to the ultimate outcome of a high-quality assembled genome. We have data across a diverse range of species covering small to large genome sizes.

A containerised Pawsey workflow for Diploidocus (Project #10)

Bioinformatics in general, and genomics specifically, is replete with complex workflows that do not translate easily to HPC. Frequently, genomics pipelines will incorporate many different tools and/or in-built functions with very different computational requirements in terms of multithreading, memory requirements and IO pressures. The Diploidocus genome curation pipeline exemplifies this problem with some lengthy single-processor steps building on data produced by highly parallelised tools, such as minimap2. As well as adapting a specific mission-critical tool, this project will help identify and establish some general principles for optimising genomics code/workflows for running on Setonix.

Diploidocus is a published genome curation and clean-up tool that utilises several different underlying bioinformatics tools and in-built algorithms. Different steps (and tools) in the pipeline have markedly different CPU, IO and memory requirements, including some lengthy non-parallelised portions. This makes it hard to run efficiently on HPC without wasting resource allocation and/or failing to take advantage of parallelisation when available.

The expected outcome of this project is a Nextflow workflow for the deployment of the Diploidocus pipeline on HPC. This will (a) increase in-house efficiency of HPC usage, and (b) make Diploidocus more attractive as a tool to other research groups.

A containerised Pawsey workflow high throughput phylogenomics (Project #14)

This project aims to produce a robust and efficient phylogenomics workflow for whole genome sequencing data.

One important application of genome assemblies is to test and improve the taxonomic classification of species using large-scale genome-wide phylogenetics, known as phylogenomics. There is a previously developed Snakemake workflow for the rapid generation of phylogenomic trees from low- to mid-coverage whole genome shotgun sequencing data. This pipeline (1) creates multiple rapid draft assemblies; (2) identifies an optimal set of orthologous genes per species using BUSCO and BUSCOMP; (3) generates a multiple sequence alignment per gene; (4) generates a phylogenetic tree per gene; and (5) generates a consensus tree from all the individual gene trees.

There is now a requirement to (1) update the pipeline to be optimised for the high-coverage draft and reference genomes created by the Ocean Genomes Project, and (2) convert this pipeline from PBS/Snakemake to SLURM/Nextflow in-line with other genomics workflows being developed at the Minderoo OceanOmics Centre at UWA.

This project will adapt the wgs2tree workflow to optionally start from a set of existing genome assemblies and BUSCO orthologue annotations and implement a Nextflow/SLURM workflow optimised to run efficiently on Pawsey.

Wednesday, 29 March 2023

The Australasian dingo archetype: De novo chromosome-length genome assembly, DNA methylome, and cranial morphology

Ballard JWO, Field MA, Edwards RJ, Wilson LAB, Koungoulos LG, Rosen BD, Chernoff B, Dudchenko O, Omer A, Keilwagen J, Skvortsova K, Bogdanovic O, Chan E, Zammit R, Hayes V & Aiden EL (2023): The Australasian dingo archetype: De novo chromosome-length genome assembly, DNA methylome, and cranial morphology. Gigascience 12:giad018. [Gigascience] [PubMed]

Background

One difficulty in testing the hypothesis that the Australasian dingo is a functional intermediate between wild wolves and domesticated breed dogs is that there is no reference specimen. Here we link a high-quality de novo long-read chromosomal assembly with epigenetic footprints and morphology to describe the Alpine dingo female named Cooinda. It was critical to establish an Alpine dingo reference because this ecotype occurs throughout coastal eastern Australia where the first drawings and descriptions were completed.

Findings

We generated a high-quality chromosome-level reference genome assembly (Canfam_ADS) using a combination of Pacific Bioscience, Oxford Nanopore, 10X Genomics, Bionano, and Hi-C technologies. Compared to the previously published Desert dingo assembly, there are large structural rearrangements on chromosomes 11, 16, 25, and 26. Phylogenetic analyses of chromosomal data from Cooinda the Alpine dingo and 9 previously published de novo canine assemblies show dingoes are monophyletic and basal to domestic dogs. Network analyses show that the mitochondrial DNA genome clusters within the southeastern lineage, as expected for an Alpine dingo. Comparison of regulatory regions identified 2 differentially methylated regions within glucagon receptor GCGR and histone deacetylase HDAC4 genes that are unmethylated in the Alpine dingo genome but hypermethylated in the Desert dingo. Morphologic data, comprising geometric morphometric assessment of cranial morphology, place dingo Cooinda within population-level variation for Alpine dingoes. Magnetic resonance imaging of brain tissue shows she had a larger cranial capacity than a similar-sized domestic dog.

Conclusions

These combined data support the hypothesis that the dingo Cooinda fits the spectrum of genetic and morphologic characteristics typical of the Alpine ecotype. We propose that she be considered the archetype specimen for future research investigating the evolutionary history, morphology, physiology, and ecology of dingoes. The female has been taxidermically prepared and is now at the Australian Museum, Sydney.

Wednesday, 23 November 2022

Minderoo OceanOmics Centre at UWA Grand Opening

The Grand Opening of the Minderoo OceanOmics Centre at UWA is only a day away! Join the launch of the Centre online from 4:40 to learn more about the inspiration and the vision behind this project, which aims to harness environmental DNA and genomics for marine conservation: https://lnkd.in/gCP4GAhs

You can find out a bit more about the Minderoo OceanOmics Centre at UWA here: https://lnkd.in/gmXKjXNu

And the broader Minderoo OceanOmics program here: https://lnkd.in/gtKHSk7g

Or get in touch if you want to know more!

Friday, 22 July 2022

The Edwards Lab is moving to the UWA Oceans Institute!

More details will follow but, in August, I will be starting a new position at the University of Western Australia Oceans Institute to head up the new Ocean Genomes Laboratory as part of the Minderoo OceanOmics Centre. This exciting project will collaborate closely with the Minderoo Foundation, the Vertebrate Genome Project, and scientists across Australia to create marine vertebrate reference genomes.

The goal of the Ocean Genomes Lab is "building and openly publishing the reference libraries for marine vertebrates ... to accurately detect, monitor and determine the health of these species". The lab is still being setup and we're hiring. Currently available is a Level B postdoc positions: http://bit.ly/OceanOmics. If building genomes is your thing, and you want to help fight the biodiversity crisis in our oceans, come and work with me! (Or pass it on if you know someone who does!) Research Assistant positions will follow.

Look out for a bunch of updates over the next few weeks, both as I update some of the outstanding presentations and posters from this year, and as the website rebrands. In the meantime, please get in touch if any of this sounds interesting!

Wednesday, 29 June 2022

The starling genome is out!

See the pre-print post for details.

Stuart KC*, Edwards RJ*, Cheng Y, Warren WC, Burt DW, Sherwin WB, Hofmeister NR, Werner SJ, Ball GF, Bateson M, Brandley MC, Buchanan KL, Cassey P, Clayton DF, De Meyer T, Meddle SL & Rollins LA (2022): Transcript- and annotation-guided genome assembly of the European starling. Molecular Ecology 22(8):3141-3160. doi: 10.1111/1755-0998.13679. [*Joint first authors] [Mol Ecol Res] [PubMed] [bioRxiv]

The European starling, Sturnus vulgaris, is an ecologically significant, globally invasive avian species that is also suffering from a major decline in its native range. Here, we present the genome assembly and long-read transcriptome of an Australian-sourced European starling (S. vulgaris vAU), and a second, North American, short-read genome assembly (S. vulgaris vNA), as complementary reference genomes for population genetic and evolutionary characterization. S. vulgaris vAU combined 10× genomics linked-reads, low-coverage Nanopore sequencing, and PacBio Iso-Seq full-length transcript scaffolding to generate a 1050 Mb assembly on 6222 scaffolds (7.6 Mb scaffold N50, 94.6% busco completeness). Further scaffolding against the high-quality zebra finch (Taeniopygia guttata) genome assigned 98.6% of the assembly to 32 putative nuclear chromosome scaffolds. Species-specific transcript mapping and gene annotation revealed good gene-level assembly and high functional completeness. Using S. vulgaris vAU, we demonstrate how the multifunctional use of PacBio Iso-Seq transcript data and complementary homology-based annotation of sequential assembly steps (assessed using a new tool, saaga) can be used to assess, inform, and validate assembly workflow decisions. We also highlight some counterintuitive behaviour in traditional busco metrics, and present buscomp, a complementary tool for assembly comparison designed to be robust to differences in assembly size and base-calling quality. This work expands our knowledge of avian genomes and the available toolkit for assessing and improving genome quality. The new genomic resources presented will facilitate further global genomic and transcriptomic analysis on this ecologically important species.

Saturday, 23 April 2022

The Australian dingo is an early offshoot of modern breed dogs

Field MA, Yadav S, Dudchenko O, Esvaran M, Rosen BD, Skvortsova K, Edwards RJ, Keilwagen J, Cochran BJ, Manandhar B, Bustamante S, Rasmussen JA, Melvin RG, Chernoffl B, Omer A, Colaric Z, Chan EKF, Minoche AE, Smith TPL, Gilbert MTP, Bogdanovic O, Zammit RA, Thomas T, Aiden EL & Ballard JWO (2022): The Australian dingo is an early offshoot of modern breed dogs. Science Advances 8(16):abm5944; DOI: 10.1126/sciadv.abm5944. [Sci Adv] [PubMed] [PDF]

Dogs are uniquely associated with human dispersal and bring transformational insight into the domestication process. Dingoes represent an intriguing case within canine evolution being geographically isolated for thousands of years. Here, we present a high-quality de novo assembly of a pure dingo (CanFam_DDS). We identified large chromosomal differences relative to the current dog reference (CanFam3.1) and confirmed no expanded pancreatic amylase gene as found in breed dogs. Phylogenetic analyses using variant pairwise matrices show that the dingo is distinct from five breed dogs with 100% bootstrap support when using Greenland wolf as the outgroup. Functionally, we observe differences in methylation patterns between the dingo and German shepherd dog genomes and differences in serum biochemistry and microbiome makeup. Our results suggest that distinct demographic and environmental conditions have shaped the dingo genome. In contrast, artificial human selection has likely shaped the genomes of domestic breed dogs after divergence from the dingo.