Protein sequence similarity searches using patterns as seeds.
AUTOR(ES)
Zhang, Z
RESUMO
Protein families often are characterized by conserved sequence patterns or motifs. A researcher frequently wishes to evaluate the significance of a specific pattern within a protein, or to exploit knowledge of known motifs to aid the recognition of greatly diverged but homologous family members. To assist in these efforts, the pattern-hit initiated BLAST (PHI-BLAST) program described here takes as input both a protein sequence and a pattern of interest that it contains. PHI-BLAST searches a protein database for other instances of the input pattern, and uses those found as seeds for the construction of local alignments to the query sequence. The random distribution of PHI-BLAST alignment scores is studied analytically and empirically. In many instances, the program is able to detect statistically significant similarity between homologous proteins that are not recognizably related using traditional single-pass database search methods. PHI-BLAST is applied to the analysis of CED4-like cell death regulators, HS90-type ATPase domains, archaeal tRNA nucleotidyltransferases and archaeal homologs of DnaG-type DNA primases.
ACESSO AO ARTIGO
http://www.pubmedcentral.nih.gov/articlerender.fcgi?artid=147803Documentos Relacionados
- Expressed Sequence Tags from Developing Castor Seeds.
- Rapid similarity searches of nucleic acid and protein data banks.
- Homology Induction: the use of machine learning to improve sequence similarity searches
- Nucleotide sequence of the canavalin gene from Canavalia gladiata seeds.
- The nucleotide sequence of ribosomal 5S RNA from lettuce seeds.