Friday, 8 January 2010

Estimation and efficient computation of the true probability of recurrence of short linear protein sequence motifs in unrelated proteins

Davey NE, Edwards RJ & Shields DC (2010): Estimation and efficient computation of the true probability of recurrence of short linear protein sequence motifs in unrelated proteins. BMC Bioinformatics 11: 14.

Abstract

BACKGROUND: Large datasets of protein interactions provide a rich resource for the discovery of Short Linear Motifs (SLiMs) that recur in unrelated proteins. However, existing methods for estimating the probability of motif recurrence may be biased by the size and composition of the search dataset, such that p-value estimates from different datasets, or from motifs containing different numbers of non-wildcard positions, are not strictly comparable. Here, we develop more exact methods and explore the potential biases of computationally efficient approximations.

RESULTS: A widely used heuristic for the calculation of motif over-representation approximates motif probability by assuming that all proteins have the same length and composition. We introduce pv, which calculates the probability exactly. Secondly, the recently introduced SLiMFinder statistic Sig, accounts for multiple testing (across all possible motifs) in motif discovery. However, it approximates the probability of all other possible motifs, occurring with a score of p or less, as being equal to p. Here, we show that the exhaustive calculation of the probability of all possible motif occurrences that are as rare or rarer than the motif of interest, Sig’, may be carried out efficiently by grouping motifs of a common probability (i.e. those which have permuted orders of the same residues). Sig’v, which corrects both approximations, is shown to be uniformly distributed in a random dataset when searching for non-ambiguous motifs, indicating that it is a robust significance measure.

CONCLUSIONS: A method is presented to compute exactly the true probability of a non-ambiguous short protein sequence motif, and the utility of an approximate approach for novel motif discovery across a large number of datasets is demonstrated.

PMID: 20055997

Monday, 16 February 2009

Masking residues using context-specific evolutionary conservation significantly improves short linear motif discovery

Davey NE, Shields DC & Edwards RJ (2009): Masking residues using context-specific evolutionary conservation significantly improves short linear motif discovery. Bioinformatics 25(4): 443-50.

Abstract

MOTIVATION: Short linear motifs (SLiMs) are important mediators of protein-protein interactions. Their short and degenerate nature presents a challenge for computational discovery. We sought to improve SLiM discovery by incorporating evolutionary information, since SLiMs are more conserved than surrounding residues.

RESULTS: We have developed a new method that assesses the evolutionary signal of a residue in its sequence and structural context. Under-conserved residues are masked out prior to SLiM discovery, allowing incorporation into the existing statistical model employed by SLiMFinder. The method shows considerable robustness in terms of both the conservation score used for individual residues and the size of the sequence neighbourhood. Optimal parameters significantly improve return of known functional motifs from benchmarking data, raising the return of significant validated SLiMs from typical human interaction datasets from 20% to 60%, while retaining the high level of stringency needed for application to real biological data. The success of this regime indicates that it could be of general benefit to computational annotation and prediction of protein function at the sequence level.

AVAILABILITY: All data and tools in this article are available at http://bioware.ucd.ie/~slimdisc/slimfinder/conmasking/.

PMID: 19136552

Friday, 16 May 2008

CompariMotif: quick and easy comparisons of sequence motifs

Edwards RJ, Davey NE & Shields DC (2008): CompariMotif: Quick and easy comparisons of sequence motifs. Bioinformatics 24(10):1307-9.

Abstract

CompariMotif is a novel tool for making motif-motif comparisons, identifying and describing similarities between regular expression motifs. CompariMotif can identify a number of different relationships between motifs, including exact matches, variants of degenerate motifs and complex overlapping motifs. Motif relationships are scored using shared information content, allowing the best matches to be easily identified in large comparisons. Many input and search options are available, enabling a list of motifs to be compared to itself (to identify recurring motifs) or to datasets of known motifs.

AVAILABILITY: CompariMotif can be run online at http://bioware.ucd.ie/ and is freely available for academic use as a set of open source Python modules under a GNU General Public License from http://bioinformatics.ucd.ie/shields/software/comparimotif/

PMID: 18375965

Thursday, 4 October 2007

SLiMFinder: a probabilistic method for identifying over-represented, convergently evolved, short linear motifs in proteins

Edwards RJ, Davey NE & Shields DC (2007): SLiMFinder: A probabilistic method for identifying over-represented, convergently evolved, short linear motifs in proteins. PLoS ONE 2(10): e967.

Abstract

BACKGROUND: Short linear motifs (SLiMs) in proteins are functional microdomains of fundamental importance in many biological systems. SLiMs typically consist of a 3 to 10 amino acid stretch of the primary protein sequence, of which as few as two sites may be important for activity, making identification of novel SLiMs extremely difficult. In particular, it can be very difficult to distinguish a randomly recurring “motif” from a truly over-represented one. Incorporating ambiguous amino acid positions and/or variable-length wildcard spacers between defined residues further complicates the matter.

METHODOLOGY/PRINCIPAL FINDINGS: In this paper we present two algorithms. SLiMBuild identifies convergently evolved, short motifs in a dataset of proteins. Motifs are built by combining dimers into longer patterns, retaining only those motifs occurring in a sufficient number of unrelated proteins. Motifs with fixed amino acid positions are identified and then combined to incorporate amino acid ambiguity and variable-length wildcard spacers. The algorithm is computationally efficient compared to alternatives, particularly when datasets include homologous proteins, and provides great flexibility in the nature of motifs returned. The SLiMChance algorithm estimates the probability of returned motifs arising by chance, correcting for the size and composition of the dataset, and assigns a significance value to each motif. These algorithms are implemented in a software package, SLiMFinder. SLiMFinder default settings identify known SLiMs with 100% specificity, and have a low false discovery rate on random test data.

CONCLUSIONS/SIGNIFICANCE: The efficiency of SLiMBuild and low false discovery rate of SLiMChance make SLiMFinder highly suited to high throughput motif discovery and individual high quality analyses alike. Examples of such analyses on real biological data, and how SLiMFinder results can help direct future discoveries, are provided. SLiMFinder is freely available for download under a GNU license from http://bioinformatics.ucd.ie/shields/software/slimfinder/.

PMID: 17912346

Saturday, 1 September 2007

Rich Edwards (Principal Investigator)

Rich Edwards is a Principal Research Fellow in the University of Western Australia (UWA) Oceans Institute, and the lead academic for the Minderoo OceanOmics Centre at UWA. The centre is a collaboration between the Minderoo Foundation and UWA, supporting an ambitious research program to revolutionise the way that environmental DNA (eDNA) is used to monitor and protect marine vertebrate biodiversity. Rich leads the technical team that runs the centre, and an academic research program in evolutionary/conservation genomics. Here, the focus is collaborating with research groups across Australia (and taxa!) to generate high-quality genome assemblies in support of applications in conservation and evolutionary biology. Rich maintains an adjunct Associate Professor position in the School of Biotechnology and Biomolecular Sciences (BABS) at the University of New South Wales.

Originally from southern England, Rich trained a geneticist at the University of Nottingham (UK), studying the population genetics of transposable elements in bacteria for his PhD. He moved to Dublin (Ireland) in 2001 to become a full time bioinformatician in the Shields Lab, developing a sequence analysis methods for rational design of biologically active short peptides based on functional specificity and ancestral sequence prediction. The biological activity of these short peptides started an interest in Short Linear Motifs (SLiMs), which are short regions of proteins that mediate interactions with other proteins.

Rich has developed several tools for the prediction and analysis of SLiMs, distributed in the SLiMSuite package (as well as coining the term “SLiM” to describe this specific type of protein interaction motif). The lab was established in 2007 when Rich moved to the University of Southampton (UK), where he continued to work on SLiMs but diversified to collaborate on numerous projects involving DNA and/or protein sequence analysis.

Rich moved to UNSW in late 2013, where he has built a close working relationship with the Ramaciotti Centre for Genomics and established genomics as a core research activity. Here, he was involved in multiple de novo whole genome sequencing and assembly projects, using short read (Illumina), long read (PacBio & Nanopore) and linked read (10x Chromium) sequencing, and Hi-C proximity ligation. These include yeast, bacteria, invasive cane toads and starlings, venomous Australian snakes through the BABS Genome Project, Aussie marsupials as part of the Oz Mammals Genomics initiative, dogs, dingoes, rainforest trees, pathogenic rust fungi, and the NSW Waratah as part of the Genomics of Australian Plants initiative. In 2022, he moved to Perth to lead the Minderoo OceanOmics Centre at UWA, where he is heavily involved in marine vertebrate genome sequencing assembly with the Ocean Genomes project.

Employment History

  • 2022-present: Principal Research Fellow, OceanOmics Centre and Laboratory Lead, Ocean Genomes Laboratory. UWA Oceans Institute, University of Western Australia.
  • 2022-present: Adjunct Associate Professor in Genomics and Bioinformatics. School of Biotechnology and Biomolecular Sciences, University of New South Wales, Australia.
  • 2013-2022: Senior Lecturer. School of Biotechnology and Biomolecular Sciences, University of New South Wales, Sydney, Australia.
  • 2014-2016: Adjunct Associate Professor. Centre for Biological Sciences, University of Southampton, Southampton SO17 1BJ, UK.
  • 2013-2014: Senior Lecturer. Centre for Biological Sciences, University of Southampton, Southampton SO17 1BJ, UK.
  • 2011-2013: Lecturer. Centre for Biological Sciences, University of Southampton, Southampton SO17 1BJ, UK.
  • 2007-2011: Senior Research Fellow. School of Biological Sciences, University of Southampton, Southampton SO16 7PX, UK.
  • 2005-2007: Postdoctoral Research Fellow. The Conway Institute, University College Dublin, Dublin, Ireland.
  • 2001-2005: Postdoctoral Research Fellow. The Royal College of Surgeons in Ireland, Dublin, Ireland.

Summary of Academic Qualifications

  • 2010: Post-graduate Certificate in Academic Practice. University of Southampton, UK.
  • 2002: PhD, "Adaptive Insertion Mutations in Bacteria". Institute of Genetics, University of Nottingham
  • 1998: BSc (Hons) Genetics. Division of Genetics, University of Nottingham.

Awards

  • 2020 Genetics Society of AustralAsia Award for Excellence in Education.

Main Funding History

  • 2022-2027: Minderoo OceanOmics Centre at UWA (Minderoo Foundation)
  • 2019-2022: ARC Linkage Project (Royal Botanic Gardens and Domain Trust).
  • 2016-2019: ARC Linkage Project (Microbiogen Pty Ltd).
  • 2015-2015: Department of Industry and Science, Research Connections Grant (Microbiogen Pty Ltd).
  • 2011-2014: BBSRC New Investigator Award BB/I006230/1.
  • 2007-2012: University of Southampton Research Fellowship.
  • 2003-2007: SFI Investigator Award (D Shields). [Named Researcher]
  • 2002-2005: HRB Programme Grant, Platelet Biology (D Kenny). [Named Researcher]
  • 2001-2003: PRTLI HEA Cycle 2 (RCSI).

Tuesday, 19 June 2007

The SLiMDisc server: short, linear motif discovery in proteins

Davey NE*, Edwards RJ* & Shields DC (2007): The SLiMDisc server: short, linear motif discovery in proteins. Nucleic Acids Res. 35(Web Server issue):W455-9. *Joint first authors

Abstract

Short, linear motifs (SLiMs) play a critical role in many biological processes, particularly in protein-protein interactions. Overrepresentation of convergent occurrences of motifs in proteins with a common attribute (such as similar subcellular location or a shared interaction partner) provides a feasible means to discover novel occurrences computationally. The SLiMDisc (Short, Linear Motif Discovery) web server corrects for common ancestry in describing shared motifs, concentrating on the convergently evolved motifs. The server returns a listing of the most interesting motifs found within unmasked regions, ranked according to an information content-based scoring scheme. It allows interactive input masking, according to various criteria. Scoring allows for evolutionary relationships in the data sets through treatment of BLAST local alignments. Alongside this ranked list, visualizations of the results improve understanding of the context of suggested motifs, helping to identify true motifs of interest. These visualizations include alignments of motif occurrences, alignments of motifs and their homologues and a visual schematic of the top-ranked motifs. Additional options for filtering and/or re-ranking motifs further permit the user to focus on motifs with desired attributes. Returned motifs can also be compared with known SLiMs from the literature. SLiMDisc is available at: http://bioware.ucd.ie/~slimdisc/.

PMID: 17576682

Friday, 1 June 2007

Evolution of specificity and diversity

Shields DC, Johnston CR, Wallace IM & Edwards RJ (2007): Evolution of specificity and diversity. In: Ancestral Sequence Reconstruction Edited by DH Ardell, DA Liberles, G Matassi. Oxford University Press.