Wednesday, 2 June 2010

Computational identification and analysis of protein short linear motifs

Davey NE, Edwards RJ & Shields DC (2010): Computational identification and analysis of protein short linear motifs. Frontiers in Bioscience 15: 801-25.

Abstract

Short linear motifs (SLiMs) in proteins can act as targets for proteolytic cleavage, sites of post-translational modification, determinants of sub-cellular localization, and mediators of protein-protein interactions. Computational discovery of SLiMs involves assembling a group of proteins postulated to share a potential motif, masking out residues less likely to contain such a motif, down-weighting shared motifs arising through common evolutionary descent, and calculation of statistical probabilities allowing for the multiple testing of all possible motifs. Much of the challenge for motif discovery lies in the assembly and masking of datasets of proteins likely to share motifs, since the motifs are typically short (between 3 and 10 amino acids in length), so that potential signals can be easily swamped by the noise of stochastically recurring motifs. Focusing on disordered regions of proteins, where SLiMs are predominantly found, and masking out non-conserved residues can reduce the level of noise but more work is required to improve the quality of high-throughput experimental datasets (e.g. of physical protein interactions) as input for computational discovery.

PMID: 20515727

Monday, 24 May 2010

SLiMFinder: a web server to find novel, significantly over-represented, short protein motifs

Davey NE, Haslam NJ, Shields DC & Edwards RJ (2010): SLiMFinder: a web server to find novel, significantly over-represented, short protein motifs. Nucleic Acids Research 38: W534-W539.

Abstract

Short, linear motifs (SLiMs) play a critical role in many biological processes, particularly in protein-protein interactions. The Short, Linear Motif Finder (SLiMFinder) web server is a de novo motif discovery tool that identifies statistically over-represented motifs in a set of protein sequences, accounting for the evolutionary relationships between them. Motifs are returned with an intuitive P-value that greatly reduces the problem of false positives and is accessible to biologists of all disciplines. Input can be uploaded by the user or extracted directly from UniProt. Numerous masking options give the user great control over the contextual information to be included in the analyses. The SLiMFinder server combines these with user-friendly output and visualizations of motif context to allow the user to quickly gain insight into the validity of a putatively functional motif. These visualizations include alignments of motif occurrences, alignments of motifs and their homologues and a visual schematic of the top-ranked motifs. Returned motifs can also be compared with known SLiMs from the literature using CompariMotif. All results are available for download. The SLiMFinder server is available at: http://bioware.ucd.ie/slimfinder.html.

PMID: 20497999

Friday, 8 January 2010

Estimation and efficient computation of the true probability of recurrence of short linear protein sequence motifs in unrelated proteins

Davey NE, Edwards RJ & Shields DC (2010): Estimation and efficient computation of the true probability of recurrence of short linear protein sequence motifs in unrelated proteins. BMC Bioinformatics 11: 14.

Abstract

BACKGROUND: Large datasets of protein interactions provide a rich resource for the discovery of Short Linear Motifs (SLiMs) that recur in unrelated proteins. However, existing methods for estimating the probability of motif recurrence may be biased by the size and composition of the search dataset, such that p-value estimates from different datasets, or from motifs containing different numbers of non-wildcard positions, are not strictly comparable. Here, we develop more exact methods and explore the potential biases of computationally efficient approximations.

RESULTS: A widely used heuristic for the calculation of motif over-representation approximates motif probability by assuming that all proteins have the same length and composition. We introduce pv, which calculates the probability exactly. Secondly, the recently introduced SLiMFinder statistic Sig, accounts for multiple testing (across all possible motifs) in motif discovery. However, it approximates the probability of all other possible motifs, occurring with a score of p or less, as being equal to p. Here, we show that the exhaustive calculation of the probability of all possible motif occurrences that are as rare or rarer than the motif of interest, Sig’, may be carried out efficiently by grouping motifs of a common probability (i.e. those which have permuted orders of the same residues). Sig’v, which corrects both approximations, is shown to be uniformly distributed in a random dataset when searching for non-ambiguous motifs, indicating that it is a robust significance measure.

CONCLUSIONS: A method is presented to compute exactly the true probability of a non-ambiguous short protein sequence motif, and the utility of an approximate approach for novel motif discovery across a large number of datasets is demonstrated.

PMID: 20055997

Monday, 16 February 2009

Masking residues using context-specific evolutionary conservation significantly improves short linear motif discovery

Davey NE, Shields DC & Edwards RJ (2009): Masking residues using context-specific evolutionary conservation significantly improves short linear motif discovery. Bioinformatics 25(4): 443-50.

Abstract

MOTIVATION: Short linear motifs (SLiMs) are important mediators of protein-protein interactions. Their short and degenerate nature presents a challenge for computational discovery. We sought to improve SLiM discovery by incorporating evolutionary information, since SLiMs are more conserved than surrounding residues.

RESULTS: We have developed a new method that assesses the evolutionary signal of a residue in its sequence and structural context. Under-conserved residues are masked out prior to SLiM discovery, allowing incorporation into the existing statistical model employed by SLiMFinder. The method shows considerable robustness in terms of both the conservation score used for individual residues and the size of the sequence neighbourhood. Optimal parameters significantly improve return of known functional motifs from benchmarking data, raising the return of significant validated SLiMs from typical human interaction datasets from 20% to 60%, while retaining the high level of stringency needed for application to real biological data. The success of this regime indicates that it could be of general benefit to computational annotation and prediction of protein function at the sequence level.

AVAILABILITY: All data and tools in this article are available at http://bioware.ucd.ie/~slimdisc/slimfinder/conmasking/.

PMID: 19136552

Friday, 16 May 2008

CompariMotif: quick and easy comparisons of sequence motifs

Edwards RJ, Davey NE & Shields DC (2008): CompariMotif: Quick and easy comparisons of sequence motifs. Bioinformatics 24(10):1307-9.

Abstract

CompariMotif is a novel tool for making motif-motif comparisons, identifying and describing similarities between regular expression motifs. CompariMotif can identify a number of different relationships between motifs, including exact matches, variants of degenerate motifs and complex overlapping motifs. Motif relationships are scored using shared information content, allowing the best matches to be easily identified in large comparisons. Many input and search options are available, enabling a list of motifs to be compared to itself (to identify recurring motifs) or to datasets of known motifs.

AVAILABILITY: CompariMotif can be run online at http://bioware.ucd.ie/ and is freely available for academic use as a set of open source Python modules under a GNU General Public License from http://bioinformatics.ucd.ie/shields/software/comparimotif/

PMID: 18375965

Thursday, 4 October 2007

SLiMFinder: a probabilistic method for identifying over-represented, convergently evolved, short linear motifs in proteins

Edwards RJ, Davey NE & Shields DC (2007): SLiMFinder: A probabilistic method for identifying over-represented, convergently evolved, short linear motifs in proteins. PLoS ONE 2(10): e967.

Abstract

BACKGROUND: Short linear motifs (SLiMs) in proteins are functional microdomains of fundamental importance in many biological systems. SLiMs typically consist of a 3 to 10 amino acid stretch of the primary protein sequence, of which as few as two sites may be important for activity, making identification of novel SLiMs extremely difficult. In particular, it can be very difficult to distinguish a randomly recurring “motif” from a truly over-represented one. Incorporating ambiguous amino acid positions and/or variable-length wildcard spacers between defined residues further complicates the matter.

METHODOLOGY/PRINCIPAL FINDINGS: In this paper we present two algorithms. SLiMBuild identifies convergently evolved, short motifs in a dataset of proteins. Motifs are built by combining dimers into longer patterns, retaining only those motifs occurring in a sufficient number of unrelated proteins. Motifs with fixed amino acid positions are identified and then combined to incorporate amino acid ambiguity and variable-length wildcard spacers. The algorithm is computationally efficient compared to alternatives, particularly when datasets include homologous proteins, and provides great flexibility in the nature of motifs returned. The SLiMChance algorithm estimates the probability of returned motifs arising by chance, correcting for the size and composition of the dataset, and assigns a significance value to each motif. These algorithms are implemented in a software package, SLiMFinder. SLiMFinder default settings identify known SLiMs with 100% specificity, and have a low false discovery rate on random test data.

CONCLUSIONS/SIGNIFICANCE: The efficiency of SLiMBuild and low false discovery rate of SLiMChance make SLiMFinder highly suited to high throughput motif discovery and individual high quality analyses alike. Examples of such analyses on real biological data, and how SLiMFinder results can help direct future discoveries, are provided. SLiMFinder is freely available for download under a GNU license from http://bioinformatics.ucd.ie/shields/software/slimfinder/.

PMID: 17912346

Saturday, 1 September 2007

Rich Edwards (Principal Investigator)

Rich Edwards is a Principal Research Fellow in the University of Western Australia (UWA) Oceans Institute, and the lead academic for the Minderoo OceanOmics Centre at UWA. The centre is a collaboration between the Minderoo Foundation and UWA, supporting an ambitious research program to revolutionise the way that environmental DNA (eDNA) is used to monitor and protect marine vertebrate biodiversity. Rich leads the technical team that runs the centre, and an academic research program in evolutionary/conservation genomics. Here, the focus is collaborating with research groups across Australia (and taxa!) to generate high-quality genome assemblies in support of applications in conservation and evolutionary biology. Rich maintains an adjunct Associate Professor position in the School of Biotechnology and Biomolecular Sciences (BABS) at the University of New South Wales.

Originally from southern England, Rich trained a geneticist at the University of Nottingham (UK), studying the population genetics of transposable elements in bacteria for his PhD. He moved to Dublin (Ireland) in 2001 to become a full time bioinformatician in the Shields Lab, developing a sequence analysis methods for rational design of biologically active short peptides based on functional specificity and ancestral sequence prediction. The biological activity of these short peptides started an interest in Short Linear Motifs (SLiMs), which are short regions of proteins that mediate interactions with other proteins.

Rich has developed several tools for the prediction and analysis of SLiMs, distributed in the SLiMSuite package (as well as coining the term “SLiM” to describe this specific type of protein interaction motif). The lab was established in 2007 when Rich moved to the University of Southampton (UK), where he continued to work on SLiMs but diversified to collaborate on numerous projects involving DNA and/or protein sequence analysis.

Rich moved to UNSW in late 2013, where he has built a close working relationship with the Ramaciotti Centre for Genomics and established genomics as a core research activity. Here, he was involved in multiple de novo whole genome sequencing and assembly projects, using short read (Illumina), long read (PacBio & Nanopore) and linked read (10x Chromium) sequencing, and Hi-C proximity ligation. These include yeast, bacteria, invasive cane toads and starlings, venomous Australian snakes through the BABS Genome Project, Aussie marsupials as part of the Oz Mammals Genomics initiative, dogs, dingoes, rainforest trees, pathogenic rust fungi, and the NSW Waratah as part of the Genomics of Australian Plants initiative. In 2022, he moved to Perth to lead the Minderoo OceanOmics Centre at UWA, where he is heavily involved in marine vertebrate genome sequencing assembly with the Ocean Genomes project.

Employment History

  • 2022-present: Principal Research Fellow, OceanOmics Centre and Laboratory Lead, Ocean Genomes Laboratory. UWA Oceans Institute, University of Western Australia.
  • 2022-present: Adjunct Associate Professor in Genomics and Bioinformatics. School of Biotechnology and Biomolecular Sciences, University of New South Wales, Australia.
  • 2013-2022: Senior Lecturer. School of Biotechnology and Biomolecular Sciences, University of New South Wales, Sydney, Australia.
  • 2014-2016: Adjunct Associate Professor. Centre for Biological Sciences, University of Southampton, Southampton SO17 1BJ, UK.
  • 2013-2014: Senior Lecturer. Centre for Biological Sciences, University of Southampton, Southampton SO17 1BJ, UK.
  • 2011-2013: Lecturer. Centre for Biological Sciences, University of Southampton, Southampton SO17 1BJ, UK.
  • 2007-2011: Senior Research Fellow. School of Biological Sciences, University of Southampton, Southampton SO16 7PX, UK.
  • 2005-2007: Postdoctoral Research Fellow. The Conway Institute, University College Dublin, Dublin, Ireland.
  • 2001-2005: Postdoctoral Research Fellow. The Royal College of Surgeons in Ireland, Dublin, Ireland.

Summary of Academic Qualifications

  • 2010: Post-graduate Certificate in Academic Practice. University of Southampton, UK.
  • 2002: PhD, "Adaptive Insertion Mutations in Bacteria". Institute of Genetics, University of Nottingham
  • 1998: BSc (Hons) Genetics. Division of Genetics, University of Nottingham.

Awards

  • 2020 Genetics Society of AustralAsia Award for Excellence in Education.

Main Funding History

  • 2022-2027: Minderoo OceanOmics Centre at UWA (Minderoo Foundation)
  • 2019-2022: ARC Linkage Project (Royal Botanic Gardens and Domain Trust).
  • 2016-2019: ARC Linkage Project (Microbiogen Pty Ltd).
  • 2015-2015: Department of Industry and Science, Research Connections Grant (Microbiogen Pty Ltd).
  • 2011-2014: BBSRC New Investigator Award BB/I006230/1.
  • 2007-2012: University of Southampton Research Fellowship.
  • 2003-2007: SFI Investigator Award (D Shields). [Named Researcher]
  • 2002-2005: HRB Programme Grant, Platelet Biology (D Kenny). [Named Researcher]
  • 2001-2003: PRTLI HEA Cycle 2 (RCSI).