Showing posts with label gene duplication. Show all posts
Showing posts with label gene duplication. Show all posts

Thursday, 13 February 2025

Small but mitey: a gapless telomere-to-telomere assembly of an unidentified mite with a streamlined genome

A few years ago, we accidentally sequenced an interesting mite when making the Rhodamnia argentea reference genome. We don’t really know what it is, but thanks to nanopore sequencing and a low repeat content, we produced a gapless telomere-to-telomere assembly! As an unfunded passion project, it’s been ticking along in the background since then, but is now published in Genome Biology and Evolution.

One of the most interesting aspects of the paper is the genome reduction - despite being a complete nuclear genome, the assembly is under 35 megabases! This is a couple of orders of magnitude smaller than many other arachnid genomes.

We are yet to do a comprehensive analysis, but just looking at the core set of expected arachnid genes, as defined by BUSCO (e.g. single-copy genes that are expected to be present in most arachnid genomes) revealed that about a third of them were missing. (This low completeness was partly responsible for the time taken to fully recognised and deal with the contamination.) Tellingly, the closest sequenced relatives of this mite also have reduced genomes and are missing many of the same BUSCO genes, revealing a long history of gene loss.

The proportion of “Duplicated” BUSCO genes is also surprisingly high at 4.5%. These are genuine duplications, with consistent diploid read depths and many of the pairs present on both chromosomes. It will be interesting to see if these duplications are replacing some of the lost functions from the other genes, or are novel genes behind such a specialised lifestyle. As an unfunded passion project, it was beyond scope to investigate the full annotation as part of this paper, but get in touch if you would be interested to do this!

Edwards RJ, Chen SH, Halliday B & Bragg JG (2025): Small but mitey: a gapless telomere-to-telomere assembly of an unidentified mite with a streamlined genome. Genome Biology and Evolution Feb 13. [Gen Biol Evol] [PubMed]

Abstract

A draft assembly of the rainforest tree Rhodamnia argentea Benth. (malletwood, Myrtaceae) revealed contaminating DNA sequences that most closely matched those from mites in the family Eriophyidae. Eriophyoid mites are plant parasites that often induce galls or other deformities on their host plants. They are notable for their small size (averaging 200 μm), distinctive four-legged body structure, and heavily streamlined genomes, which are among the smallest known of all arthropods. Contaminating mite sequences were assembled into a high-quality gapless telomere-to-telomere nuclear genome. The entire genome was assembled on two fully contiguous chromosomes, capped with a novel TTTGG or TTTGGTGTTGG telomere sequence, and exhibited clear signs of genome reduction (34.5 Mbp total length, 68.6% arachnid Benchmarking Universal Single-Copy Ortholog completeness). Phylogenomic analysis confirmed that this genome is that of a previously unsequenced eriophyoid mite. Despite its unknown identity, this complete nuclear genome provides a valuable resource to investigate invertebrate genome reduction.

Thursday, 31 October 2024

BioDiversity Genomics Conference (BG24)

The global BioDiversity Genomics Conference (BG24) was another great success this year. As well as running a session on Marine Vertebrate Genomics, it was particularly rewarding to see some many quality contributions from lab members and alumni.

Biodiversity Genomics in Australasia

Jessica Pearce (UWA): Reconstructing tiger shark history using genomics

Sharks and rays are a clade of high evolutionary, ecological, economic, and cultural significance, and yet they are one of the most threatened taxa groups in the marine environment. Despite this, there remains a lack of molecular resources for this class to assist with their conservation. The tiger shark (Galeocerdo cuvier) is a near threatened, keystone species distributed circumglobally that is under substantial pressure from human impacts, making it a high priority for management worldwide. We sequenced and characterised a reference quality assembly for the tiger shark, the first genome for this family, and used this to dive deeper into the evolution, adaption, and demographic history of this ancient species. We investigated how its effective population size (Ne), genome-wide heterozygosity and inbreeding has changed over time to infer how this species has responded to past global events, and hence potential responses to ongoing and future accumulating threats. This aims to assist in effective management of this high-profile species. 

Katarina Stuart (University of Auckland): Lifetime fitness is correlated more strongly with structural variant than SNP mutational load in a threatened bird species

Conservation genomics is becoming increasingly interested in whether structural variant (SV) information can help the management of threatened species. The functional consequences of SVs are more complex than for single nucleotide polymorphisms (SNPs) and thus may be more likely to contribute to load. While the impacts of SV-specific genetic load may be less consequential for large populations, the interplay between weakened selection and stochastic processes mean that smaller populations, like those of the threatened Aotearoa hihi/New Zealand stitchbird (Notiomystis cincta), may harbour a high SV load. Hihi were once confined to a single remnant population, but have been reestablished into six sanctuaries and reserves, often via secondary bottlenecks, resulting in low genetic diversity, low adaptive potential and inbreeding depression. In this study, we use whole genome resequencing of 30 individuals from the Tiritiri Matangi population to identify the nature and distribution of both SNPs and SVs within this small avian population. We find that SNP and SV individual mutation load is only moderately correlated, likely because SVs arise in regions of high recombination and reduced evolutionary conservation. Finally, we leverage a long-term monitoring dataset of pedigree and fitness data to assess the impact of SNP and SV mutation load on individual fitness, and demonstrate that SV load correlates more strongly than SNP load with lifetime fitness. The results of this study indicate that only examining SNPs neglects important aspects of intraspecific variation, and that studying SVs has direct implications for linking genetic diversity and genetic health to inform management decisions.

Richard Edwards (UWA): Improving Hifiasm assemblies with 20 kb ONT reads

The quality and quantity of genome assembly has improved dramatically over recent years. Many large-scale genome projects assemble HiFi and HiC reads using Hifiasm to produce contiguous phased assemblies, scaffolded to chromosome-level. Nevertheless, HiFi reads are typically under 25 kb and can still struggle to assemble long, low-diversity repeat regions. Obtaining ‘ultra-long’ (100 kb or longer) ONT reads to solve this problem remains a significant challenge due to technical constraints and DNA sample requirements. Here, we explore the utility of using standard ONT long reads (20 kb or more) as ‘ultra-long’ input to improve phased Hifiasm assemblies for 22 species of bony fish (Genome Size, 627 Mb - 1.54 Gb). We also explore whether the new ‘telo-m’ mode in Hifiasm v0.9.0 improves telomere prediction in these species. Incorporating 20+ kb ONT reads (7.8X - 93.5X) significantly increased assembly contiguity. BUSCO completeness was not significantly altered, although there was some re-partitioning of BUSCO genes between phased haplotypes for some species. Improvement did not strongly correlate with read depth (neither HiFi nor ONT), suggesting that the underlying read length distributions and/or specific genome features are more important for determining the outcome. Hifiasm ‘telo-m’ mode significantly increased telomere recovery, assembling over six times the number of gapless telomere-to-telomere chromosomes when combined with incorporation of ONT reads. Verification of how these results translate to the quality and/or ease of curation of final HiC-scaffolded chromosome-level assemblies is ongoing, with a goal to determine whether the additional sample preparation and sequencing in the lab is cost-effective.

Emma de Jong (UWA): High-Quality Genomes for Australian Lutjanidae Species

Lutjanidae (snappers) are highly valued in commercial and recreational fisheries worldwide and some species serve as fisheries indicator species particularly for bioregions in Western Australia. Comprehensive genomic mapping of immune gene families of Lutjanidae species are lacking, but this information can inform understanding disease vulnerability, the impact of environmental stress, improving aquaculture efforts and to provide insights into the health of wild populations. Despite their importance, only 3 out of 113 Lutjanid species currently have available reference genomes, two of which are highly fragmented (>11,000 and >200,000 contigs), impacting studies on gene families relevant to aquaculture. In this study, we present high-quality chromosome-level reference genomes for 14 Australian lutjanid species across seven genera, generated using PacBio HiFi and Dovetail HiC data. We present initial comparative genomic analyses, including immune gene content and chromosomal synteny analyses across species. These analyses provide insights into the genomic architecture and evolutionary relationships within Lutjanidae. Ongoing work aims to comprehensively map and compare the immune gene family repertoire across genera in Lutjanidae, as well as lethrinid species as an outgroup, to determine genus-specific changes in genes (e.g., loss, selection, duplication) important for pathogen detection, antigen presentation, inflammation, and immune memory. These genome assemblies will serve as a foundational resource to the wider scientific community interested in these species.

Research of ECRs who work in biodiversity genomics

Lara Parata (UWA): Genome Evolution in Marine Ray-Finned Fishes

Approximately half of extant vertebrate species are fishes, with more than 30,000 species classified as ray-finned fishes (Actinopterygii). Actinopterygii represent diverse phenotypes, feeding strategies, life history traits and occupy distinct ecological niches, making them an ideal taxa for studying molecular drivers of diversity and adaptation. Despite their diversity, ecological, and economical importance, only 145 Illumina genome assemblies are available for marine Actinopterygii species. In this study we present 250 new marine Actinopterygii genome assemblies generated using Illumina whole genome sequencing and initial results from a large-scale study of these 395 genomes. Using reference-based annotation tools we determine which fish families have unique patterns of gene family frequency / structure (e.g., losses, expansions, contractions), and correlate these with predicted functional signatures to infer biological and ecological adaptations. We identify fish families with distinct rates of change in the gene families present within their genomes (e.g., more losses / expansions or diversity) and associate these patterns with increased rates of diversification or speciation to further elucidate the genomic attributes contributing to ecological success. The results of this work contribute to the growing understanding of fish genome evolution and provide new insights into the evolutionary history and ecological success of marine Actinopterygii.

Friday, 15 December 2023

A high-quality pseudo-phased genome for Melaleuca quinquenervia shows allelic diversity of NLR-type resistance genes

Chen SH, Martino AM, Luo Z, Schwessinger B, Jones A, Tolessa T, Bragg JG, Tobias PA, Edwards RJ (2023): A high-quality pseudo-phased genome for Melaleuca quinquenervia shows allelic diversity of NLR-type resistance genes. GigaScience 12:giad102. [Gigascience] [PubMed]

Background. Melaleuca quinquenervia (broad-leaved paperbark) is a coastal wetland tree species that serves as a foundation species in eastern Australia, Indonesia, Papua New Guinea, and New Caledonia. While extensively cultivated for its ornamental value, it has also become invasive in regions like Florida, USA. Long-lived trees face diverse pest and pathogen pressures, and plant stress responses rely on immune receptors encoded by the nucleotide-binding leucine-rich repeat (NLR) gene family. However, the comprehensive annotation of NLR encoding genes has been challenging due to their clustering arrangement on chromosomes and highly repetitive domain structure; expansion of the NLR gene family is driven largely by tandem duplication. Additionally, the allelic diversity of the NLR gene family remains largely unexplored in outcrossing tree species, as many genomes are presented in their haploid, collapsed state.

Results. We assembled a chromosome-level pseudo-phased genome for M. quinquenervia and described the allelic diversity of plant NLRs using the novel FindPlantNLRs pipeline. Analysis reveals variation in the number of NLR genes on each haplotype, distinct clustering patterns, and differences in the types and numbers of novel integrated domains.

Conclusions. The high-quality M. quinquenervia genome assembly establishes a new framework for functional and evolutionary studies of this significant tree species. Our findings suggest that maintaining allelic diversity within the NLR gene family is crucial for enabling responses to environmental stress, particularly in long-lived plants.

Friday, 4 December 2020

EdwardsLab at #AusEvo2020

If you missed his talk at ABACBS2020, Jack will be presenting today at the Australasian Evolution Society 2020 Conference about The role of gene duplication in the evolution of snake venoms. Two conference presentations in two weeks - not a bad way to prepare for your Honours viva post-submission. Well done, Jack!

Also, Kat Stuart will be presenting her work on invasive starlings in Zoom 2 at 13:00 AEDT. Kat’s talks are always great to listen to:

  • Katarina Stuart: What drives invasion success? Using historical museum samples to examine evolution in an invasive passerine.

Tuesday, 24 November 2020

#ABACBS2020: The role of gene duplication in the evolution of snake venoms

Jack Clarke, Vicki Thomson & Richard Edwards

Abstract

Snakes are one of the most venomous animals on the planet, using their venom for defence and the capturing of prey. Snake venoms have evolved independently of other venoms in other vertebrates, and there is considerable variation between species in their proteomic composition. One of the primary mechanisms through which snake venoms are thought to evolve is the duplication, recruitment and specialisation of proteins from other tissues. In some cases, this evolution is known to involve the tandem duplication of genes resulting in chromosomal clusters of venom genes in some gene families. We have recently sequenced and assembled the genomes of two highly venomous Australian snakes: Notechis scutatus (mainland tiger snake) and Pseudonaja textilis (eastern brown snake). In conjunction with publicly available proteomes from 10 other venomous snakes and 2 non-venomous snakes, these genomes provide an excellent opportunity to examine the role that duplication and neofunctionalisation has played in snake venom evolution.

We have analysed 43 protein families known to play a role in snake venom and examined their pattern of duplication in snakes, compared to high quality reference genomes of other reptiles and non-venomous vertebrates. We find evidence for extensive duplications across some of these families, but no clear enrichment for duplication in the evolution of venom specifically. Instead, we identify a trend where numerous duplications specific to venomous snakes occur in proteins that seem predisposed to evolve by duplication and specialisation, even in non-venomous vertebrates. A subset of high-quality snake genomes was then used to further explore the nature of duplications. While tandem gene duplication is evident in some larger families, it remains absent in many.

The snake venom metalloproteinase (SVMP) family provides an excellent case study, with multiple duplication events throughout its evolutionary history in vertebrates. Part of the broader ADAM (“a disintegrin and metalloproteinase”) family of single-pass transmembrane and secreted zinc proteases, SVMP appears to have expanded by independent tandem duplications in different snake lineages. We also identify a second ADAM subfamily, ADAM20, with an abundance of venomous snake-specific duplications. Ongoing work in exploring the possible role of ADAM20 proteins in snake venoms and the role that genome assembly quality has played in our ability to robustly detect the presence or absence of gene duplication events.

Tuesday, 18 February 2020

Jack Clarke (Honours student)

Jack worked in the Edwards Lab in 2020 for his Honour years, researching the evolution of snake venoms using the lab’s de novo genome assemblies of two Australian snakes (Eastern brown snake and mainland tiger snake). His researched focused on phylogenetic analysis of venom gene families in order to understand how venom divergence has occurred in different branches of the snake lineage. The work focused on the role that tandem gene duplication events had played in the evolution of specific venom gene families and identified novel sites of tandem duplication shared across multiple snake species.

Jack holds a Bachelor of Advanced Science (Honours) (First Class) with majors in Bioinformatics and Genetics. He is currently completing his PhD at UNSW in collaboration with the Victor Chang institute.

[LinkedIn]

Friday, 1 June 2007

Evolution of specificity and diversity

Shields DC, Johnston CR, Wallace IM & Edwards RJ (2007): Evolution of specificity and diversity. In: Ancestral Sequence Reconstruction Edited by DH Ardell, DA Liberles, G Matassi. Oxford University Press.

Wednesday, 24 January 2007

Evaluation of whether accelerated protein evolution in chordates has occurred before, after, or simultaneously with gene duplication

Johnston CR, O’dushlaine C, Fitzpatrick DA, Edwards RJ & Shields DC (2007): Evaluation Of Whether Accelerated Protein Evolution In Chordates Has Occurred Before, After Or Simultaneously With Gene Duplication. Mol. Biol. Evol. 24:315-323.

Abstract

Gene duplication and loss are predicted to be at least of the order of the substitution rate and are key contributors to the development of novel gene function and overall genome evolution. Although it has been established that proteins evolve more rapidly after gene duplication, we were interested in testing to what extent this reflects causation or association. Therefore, we investigated the rate of evolution prior to gene duplication in chordates. Two patterns emerged; firstly, branches, which are both preceded by a duplication and followed by a duplication, display an elevated rate of amino acid replacement. This is reflected in the ratio of nonsynonymous to synonymous substitution (mean nonsynonymous to synonymous nucleotide substitution rate ratio [Ka:Ks]) of 0.44 compared with branches preceded by and followed by a speciation (mean Ka:Ks of 0.23). The observed patterns suggest that there can be simultaneous alteration in the selection pressures on both gene duplication and amino acid replacement, which may be consistent with co-occurring increases in positive selection, or alternatively with concurrent relaxation of purifying selection. The pattern is largely, but perhaps not completely, explained by the existence of certain families that have elevated rates of both gene duplication and amino acid replacement. Secondly, we observed accelerated amino acid replacement prior to duplication (mean Ka:Ks for postspeciation preduplication branches was 0.27). In some cases, this could reflect adaptive changes in protein function precipitating a gene duplication event. In conclusion, the circumstances surrounding the birth of new proteins may frequently involve a simultaneous change in selection pressures on both gene-copy number and amino acid replacement. More precise modeling of the relative importance of preduplication, postduplication, and simultaneous amino acid replacement will require larger and denser genomic data sets from multiple species, allowing simultaneous estimation of lineage-specific fluctuations in mutation rates and adaptive constraints.

PMID: 17065596

Monday, 15 January 2007

Bioinformatic discovery of novel bioactive peptides

Edwards RJ*, Moran N*, Devocelle M, Kiernan A, Meade G, Signac W, Foy M, Park SDE, Dunne E, Kenny D & Shields DC (2007): Bioinformatic discovery of novel bioactive peptides. Nature Chem. Biol. 3(2):108-112. *Joint first authors

Abstract

Short synthetic oligopeptides based on regions of human proteins that encompass functional motifs are versatile reagents for understanding protein signaling and interactions. They can either mimic or inhibit the parent protein’s activity and have been used in drug development. Peptide studies typically either derive peptides from a single identified protein or (at the other extreme) screen random combinatorial peptides, often without knowledge of the signaling pathways targeted. Our objective was to determine whether rational bioinformatic design of oligopeptides specifically targeted to potentially signaling-rich juxtamembrane regions could identify modulators of human platelet function. High-throughput in vitro platelet function assays of palmitylated cell-permeable oligopeptides corresponding to these regions identified many agonists and antagonists of platelet function. Many bioactive peptides were from adhesion molecules, including a specific CD226-derived inhibitor of inside-out platelet signaling. Systematic screens of this nature are highly efficient tools for discovering short signaling motifs in molecular signaling pathways.

Comment in

A shortcut to peptides to modulate platelets. Nat Chem Biol. 2007.

PMID: 17220901

Wednesday, 16 November 2005

BADASP: predicting functional specificity in protein families using ancestral sequences

Edwards RJ & Shields DC (2005): BADASP: predicting functional specificity in protein families using ancestral sequences. Bioinformatics 21(22):4190-1.

Abstract

SUMMARY: Burst After Duplication with Ancestral Sequence Predictions (BADASP) is a software package for identifying sites that may confer subfamily-specific biological functions in protein families following functional divergence of duplicated proteins. A given protein phylogeny is grouped into subfamilies based on orthology/paralogy relationships and/or user definitions. Ancestral sequences are then predicted from the sequence alignment and the functional specificity is calculated using variants of the Burst After Duplication method, which tests for radical amino acid substitutions following gene duplications that are subsequently conserved. Statistics are output along with subfamily groupings and ancestral sequences for an easy analysis with other packages.

AVAILABILITY: BADASP is freely available from http://www.bioinformatics.rcsi.ie/~redwards/badasp/

PMID: 16159912