Showing posts with label #programs. Show all posts
Showing posts with label #programs. Show all posts

Monday, 11 October 2021

SLiMSuite Short Linear Motif and Genomics Analysis Tools: BUSCOMP v0.13.0 (MetaEuk) release

SLiMSuite: BUSCOMP v0.13.0 (MetaEuk) release: BUSCOMP v0.13.0 is now on GitHub. This release features updates to parse additional BUSCO v5 outputs, including transcriptome and proteome modes. It has also been updated to be compatible with MetaEuk runs by generating the missing *.fna files where possible.

Wednesday, 22 July 2020

Computational Prediction of Disordered Protein Motifs Using SLiMSuite

Edwards RJ, Paulsen K, Aguilar Gomez CM & Pérez-Bercoff Å (2020): Computational Prediction of Disordered Protein Motifs using SLiMSuite. Methods Mol Biol. 2141:37-72. doi: 10.1007/978-1-0716-0524-0_3. [PubMed]

Abstract

Short linear motifs (SLiMs) are important mediators of interactions between intrinsically disordered regions of proteins and their interaction partners. Here, we detail instructions for the computational prediction of SLiMs in disordered protein regions, using the main tools of the SLiMSuite package: (1) SLiMProb identifies and calculates enrichment of predefined motifs in a set of proteins; (2) SLiMFinder predicts SLiMs de novo in a set of proteins, accounting for evolutionary relationships; (3) QSLiMFinder increases SLiMFinder sensitivity by focusing SLiM prediction on a specific query protein/region; (4) CompariMotif compares predicted SLiMs to known SLiMs or other SLiM predictions to identify common patterns. For each tool, command-line and online server examples are provided. Detailed notes provide additional advice on different applications of SLiMSuite, including batch running of multiple datasets and conservation masking using alignments of predicted orthologues.

Monday, 1 June 2015

New SLiMSuite release and GitHub site

A new download of SLiMSuite (release 2015-06-01) is now available. This is the first release in the new git repository at https://github.com/slimsuite/SLiMSuite. A tarball slimsuite.2015-06-01.tgz is also available, containing the same code. Once unpacked, it should be possible to pull down additional updates with git.

Monday, 23 June 2014

New SLiMSuite release now available

SLiMSuite Short Linear Motif discovery and analysis: New SLiMSuite release now available: A new download of SLiMSuite (release 2014-06-22 ) is now available. As well as fixing the minor GOPHER output bug , a new Taxonomy proces...

Thursday, 24 April 2014

SLiMSuite 2014-04-22 now available

SLiMSuite 2014-04-22 now available: A new download of SLiMSuite (release 2014-04-22) is now available. As well as fixing the gopher.py error, the download page and readme ha...

Wednesday, 4 December 2013

Saturday, 25 September 2010

SLiMSearch: a webserver for finding novel occurrences of short linear motifs in proteins, incorporating sequence context

Davey NE, Haslam NJ, Shields DC & Edwards RJ (2010): SLiMSearch: a webserver for finding novel occurrences of short linear motifs in proteins, incorporating sequence context. In: Pattern Recognition in Bioinformatics Edited by Dijkstra TMH, Tsivtsivadze E, Marchiori E & Heskes T. Springer-Verlag, Berlin. Lecture Notes in Bioinformatics 6282: 50-61.

Abstract

Short, linear motifs (SLiMs) play a critical role in many biological processes. The SLiMSearch (Short, Linear Motif Search) webserver is a flexible tool that enables researchers to identify novel occurrences of predefined SLiMs in sets of proteins. Numerous masking options give the user great control over the contextual information to be included in the analyses, including evolutionary filtering and protein structural disorder. User-friendly output and visualizations of motif context allow the user to quickly gain insight into the validity of a putatively functional motif occurrence. Users can search motifs against the human proteome, or submit their own datasets of UniProt proteins, in which case motif support within the dataset is statistically assessed for over- and under-representation, accounting for evolutionary relationships between input proteins. SLiMSearch is freely available as open source Python modules and all webserver results are available for download. The SLiMSearch server is available at: http://bioware.ucd.ie/slimsearch.html.

Monday, 16 February 2009

Masking residues using context-specific evolutionary conservation significantly improves short linear motif discovery

Davey NE, Shields DC & Edwards RJ (2009): Masking residues using context-specific evolutionary conservation significantly improves short linear motif discovery. Bioinformatics 25(4): 443-50.

Abstract

MOTIVATION: Short linear motifs (SLiMs) are important mediators of protein-protein interactions. Their short and degenerate nature presents a challenge for computational discovery. We sought to improve SLiM discovery by incorporating evolutionary information, since SLiMs are more conserved than surrounding residues.

RESULTS: We have developed a new method that assesses the evolutionary signal of a residue in its sequence and structural context. Under-conserved residues are masked out prior to SLiM discovery, allowing incorporation into the existing statistical model employed by SLiMFinder. The method shows considerable robustness in terms of both the conservation score used for individual residues and the size of the sequence neighbourhood. Optimal parameters significantly improve return of known functional motifs from benchmarking data, raising the return of significant validated SLiMs from typical human interaction datasets from 20% to 60%, while retaining the high level of stringency needed for application to real biological data. The success of this regime indicates that it could be of general benefit to computational annotation and prediction of protein function at the sequence level.

AVAILABILITY: All data and tools in this article are available at http://bioware.ucd.ie/~slimdisc/slimfinder/conmasking/.

PMID: 19136552

Friday, 16 May 2008

CompariMotif: quick and easy comparisons of sequence motifs

Edwards RJ, Davey NE & Shields DC (2008): CompariMotif: Quick and easy comparisons of sequence motifs. Bioinformatics 24(10):1307-9.

Abstract

CompariMotif is a novel tool for making motif-motif comparisons, identifying and describing similarities between regular expression motifs. CompariMotif can identify a number of different relationships between motifs, including exact matches, variants of degenerate motifs and complex overlapping motifs. Motif relationships are scored using shared information content, allowing the best matches to be easily identified in large comparisons. Many input and search options are available, enabling a list of motifs to be compared to itself (to identify recurring motifs) or to datasets of known motifs.

AVAILABILITY: CompariMotif can be run online at http://bioware.ucd.ie/ and is freely available for academic use as a set of open source Python modules under a GNU General Public License from http://bioinformatics.ucd.ie/shields/software/comparimotif/

PMID: 18375965

Thursday, 4 October 2007

SLiMFinder: a probabilistic method for identifying over-represented, convergently evolved, short linear motifs in proteins

Edwards RJ, Davey NE & Shields DC (2007): SLiMFinder: A probabilistic method for identifying over-represented, convergently evolved, short linear motifs in proteins. PLoS ONE 2(10): e967.

Abstract

BACKGROUND: Short linear motifs (SLiMs) in proteins are functional microdomains of fundamental importance in many biological systems. SLiMs typically consist of a 3 to 10 amino acid stretch of the primary protein sequence, of which as few as two sites may be important for activity, making identification of novel SLiMs extremely difficult. In particular, it can be very difficult to distinguish a randomly recurring “motif” from a truly over-represented one. Incorporating ambiguous amino acid positions and/or variable-length wildcard spacers between defined residues further complicates the matter.

METHODOLOGY/PRINCIPAL FINDINGS: In this paper we present two algorithms. SLiMBuild identifies convergently evolved, short motifs in a dataset of proteins. Motifs are built by combining dimers into longer patterns, retaining only those motifs occurring in a sufficient number of unrelated proteins. Motifs with fixed amino acid positions are identified and then combined to incorporate amino acid ambiguity and variable-length wildcard spacers. The algorithm is computationally efficient compared to alternatives, particularly when datasets include homologous proteins, and provides great flexibility in the nature of motifs returned. The SLiMChance algorithm estimates the probability of returned motifs arising by chance, correcting for the size and composition of the dataset, and assigns a significance value to each motif. These algorithms are implemented in a software package, SLiMFinder. SLiMFinder default settings identify known SLiMs with 100% specificity, and have a low false discovery rate on random test data.

CONCLUSIONS/SIGNIFICANCE: The efficiency of SLiMBuild and low false discovery rate of SLiMChance make SLiMFinder highly suited to high throughput motif discovery and individual high quality analyses alike. Examples of such analyses on real biological data, and how SLiMFinder results can help direct future discoveries, are provided. SLiMFinder is freely available for download under a GNU license from http://bioinformatics.ucd.ie/shields/software/slimfinder/.

PMID: 17912346

Thursday, 20 July 2006

SLiMDisc: short, linear motif discovery, correcting for common evolutionary descent

Davey NE, Shields DC & Edwards RJ (2006): SLiMDisc: short, linear motif discovery, correcting for common evolutionary descent. Nucleic Acids Res. 34(12):3546-54.

Abstract

Many important interactions of proteins are facilitated by short, linear motifs (SLiMs) within a protein’s primary sequence. Our aim was to establish robust methods for discovering putative functional motifs. The strongest evidence for such motifs is obtained when the same motifs occur in unrelated proteins, evolving by convergence. In practise, searches for such motifs are often swamped by motifs shared in related proteins that are identical by descent. Prediction of motifs among sets of biologically related proteins, including those both with and without detectable similarity, were made using the TEIRESIAS algorithm. The number of motif occurrences arising through common evolutionary descent were normalized based on treatment of BLAST local alignments. Motifs were ranked according to a score derived from the product of the normalized number of occurrences and the information content. The method was shown to significantly outperform methods that do not discount evolutionary relatedness, when applied to known SLiMs from a subset of the eukaryotic linear motif (ELM) database. An implementation of Multiple Spanning Tree weighting outperformed two other weighting schemes, in a variety of settings.

PMID: 16855291

Wednesday, 16 November 2005

BADASP: predicting functional specificity in protein families using ancestral sequences

Edwards RJ & Shields DC (2005): BADASP: predicting functional specificity in protein families using ancestral sequences. Bioinformatics 21(22):4190-1.

Abstract

SUMMARY: Burst After Duplication with Ancestral Sequence Predictions (BADASP) is a software package for identifying sites that may confer subfamily-specific biological functions in protein families following functional divergence of duplicated proteins. A given protein phylogeny is grouped into subfamilies based on orthology/paralogy relationships and/or user definitions. Ancestral sequences are then predicted from the sequence alignment and the functional specificity is calculated using variants of the Burst After Duplication method, which tests for radical amino acid substitutions following gene duplications that are subsequently conserved. Statistics are output along with subfamily groupings and ancestral sequences for an easy analysis with other packages.

AVAILABILITY: BADASP is freely available from http://www.bioinformatics.rcsi.ie/~redwards/badasp/

PMID: 16159912

Tuesday, 7 September 2004

GASP: Gapped Ancestral Sequence Prediction for proteins

Edwards RJ & Shields DC (2004): GASP: Gapped Ancestral Sequence Prediction for proteins. BMC Bioinformatics 5(1):123.

Abstract

BACKGROUND: The prediction of ancestral protein sequences from multiple sequence alignments is useful for many bioinformatics analyses. Predicting ancestral sequences is not a simple procedure and relies on accurate alignments and phylogenies. Several algorithms exist based on Maximum Parsimony or Maximum Likelihood methods but many current implementations are unable to process residues with gaps, which may represent insertion/deletion (indel) events or sequence fragments.

RESULTS: Here we present a new algorithm, GASP (Gapped Ancestral Sequence Prediction), for predicting ancestral sequences from phylogenetic trees and the corresponding multiple sequence alignments. Alignments may be of any size and contain gaps. GASP first assigns the positions of gaps in the phylogeny before using a likelihood-based approach centred on amino acid substitution matrices to assign ancestral amino acids. Important outgroup information is used by first working down from the tips of the tree to the root, using descendant data only to assign probabilities, and then working back up from the root to the tips using descendant and outgroup data to make predictions. GASP was tested on a number of simulated datasets based on real phylogenies. Prediction accuracy for ungapped data was similar to three alternative algorithms tested, with GASP performing better in some cases and worse in others. Adding simple insertions and deletions to the simulated data did not have a detrimental effect on GASP accuracy.

CONCLUSIONS: GASP (Gapped Ancestral Sequence Prediction) will predict ancestral sequences from multiple protein alignments of any size. Although not as accurate in all cases as some of the more sophisticated maximum likelihood approaches, it can process a wide range of input phylogenies and will predict ancestral sequences for gapped and ungapped residues alike.

PMID: 15350199