Using protein domains to improve the accuracy of Ab Initio gene finding

Mihaela Pertea; Steven L. Salzberg

Using protein domains to improve the accuracy of Ab Initio gene finding

Research output: Chapter in Book/Report/Conference proceeding › Conference contribution

Abstract

Background: Protein domains are the common functional elements used by nature to generate tremendous diversity among proteins, and they are used repeatedly in different combinations across all major domains of life. In this paper we address the problem of using similarity to known protein domains in helping with the identification of genes in a DNA sequence. We have adapted the generalized hidden Markov model (GHMM) architecture of the ab intio gene finder GlimmerHMM such that a higher probability is assigned to exons that contain homologues to protein domains. To our knowledge, this domain homology based approach has not been used previously in the context of ab initio gene prediction. Results: GlimmerHMM was augmented with a protein domain module that recognizes gene structures that are similar to Pfam models. The augmented system, GlimmerHMM+, shows 2% improvement in sensitivity and a 1% increase in specificity in predicting exact gene structures compared to GlimmerHMM without this option. These results were obtained on two very different model organisms: Arabidopsis thaliana (mustard wee) and Danio rerio (zebrafish), and together these preliminary results demonstrate the value of using protein domain homology in gene prediction. The results obtained are encouraging, and we believe that a more comprehensive approach including a model that reflects the statistical characteristics of specific sets of protein domain families would result in a greater increase of the accuracy of gene prediction. GlimmerHMM and GlimmerHMM+ are freely available as open source software at http://cbcb.umd.edu/software.

Original language	English (US)
Title of host publication	Algorithms in Bioinformatics - 7th International Workshop, WABI 2007, Proceedings
Pages	208-215
Number of pages	8
State	Published - Dec 24 2007
Externally published	Yes
Event	7th International Workshop on Algorithms in Bioinformatics, WABI 2007 - PhiIadelphia, PA, United States Duration: Sep 8 2007 → Sep 9 2007

Publication series

Name	Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Volume	4645 LNBI
ISSN (Print)	0302-9743
ISSN (Electronic)	1611-3349

Other

Other	7th International Workshop on Algorithms in Bioinformatics, WABI 2007
Country/Territory	United States
City	PhiIadelphia, PA
Period	9/8/07 → 9/9/07

Keywords

GHMM
Pfam
Profile HMM
Protein domain
ab intio gene finding

ASJC Scopus subject areas

Theoretical Computer Science
General Computer Science

Cite this

Using protein domains to improve the accuracy of Ab Initio gene finding. / Pertea, Mihaela ; Salzberg, Steven L.
Algorithms in Bioinformatics - 7th International Workshop, WABI 2007, Proceedings. 2007. p. 208-215 (Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics); Vol. 4645 LNBI).

Research output: Chapter in Book/Report/Conference proceeding › Conference contribution

Pertea, M & Salzberg, SL 2007, Using protein domains to improve the accuracy of Ab Initio gene finding. in Algorithms in Bioinformatics - 7th International Workshop, WABI 2007, Proceedings. Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), vol. 4645 LNBI, pp. 208-215, 7th International Workshop on Algorithms in Bioinformatics, WABI 2007, PhiIadelphia, PA, United States, 9/8/07.

@inproceedings{a6c96a4ab1014a09a68bdf4b3ba8174c,

title = "Using protein domains to improve the accuracy of Ab Initio gene finding",

abstract = "Background: Protein domains are the common functional elements used by nature to generate tremendous diversity among proteins, and they are used repeatedly in different combinations across all major domains of life. In this paper we address the problem of using similarity to known protein domains in helping with the identification of genes in a DNA sequence. We have adapted the generalized hidden Markov model (GHMM) architecture of the ab intio gene finder GlimmerHMM such that a higher probability is assigned to exons that contain homologues to protein domains. To our knowledge, this domain homology based approach has not been used previously in the context of ab initio gene prediction. Results: GlimmerHMM was augmented with a protein domain module that recognizes gene structures that are similar to Pfam models. The augmented system, GlimmerHMM+, shows 2% improvement in sensitivity and a 1% increase in specificity in predicting exact gene structures compared to GlimmerHMM without this option. These results were obtained on two very different model organisms: Arabidopsis thaliana (mustard wee) and Danio rerio (zebrafish), and together these preliminary results demonstrate the value of using protein domain homology in gene prediction. The results obtained are encouraging, and we believe that a more comprehensive approach including a model that reflects the statistical characteristics of specific sets of protein domain families would result in a greater increase of the accuracy of gene prediction. GlimmerHMM and GlimmerHMM+ are freely available as open source software at http://cbcb.umd.edu/software.",

keywords = "GHMM, Pfam, Profile HMM, Protein domain, ab intio gene finding",

author = "Mihaela Pertea and Salzberg, {Steven L.}",

year = "2007",

month = dec,

day = "24",

language = "English (US)",

isbn = "9783540741251",

series = "Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)",

pages = "208--215",

booktitle = "Algorithms in Bioinformatics - 7th International Workshop, WABI 2007, Proceedings",

note = "7th International Workshop on Algorithms in Bioinformatics, WABI 2007 ; Conference date: 08-09-2007 Through 09-09-2007",

}

TY - GEN

T1 - Using protein domains to improve the accuracy of Ab Initio gene finding

AU - Pertea, Mihaela

AU - Salzberg, Steven L.

PY - 2007/12/24

Y1 - 2007/12/24

N2 - Background: Protein domains are the common functional elements used by nature to generate tremendous diversity among proteins, and they are used repeatedly in different combinations across all major domains of life. In this paper we address the problem of using similarity to known protein domains in helping with the identification of genes in a DNA sequence. We have adapted the generalized hidden Markov model (GHMM) architecture of the ab intio gene finder GlimmerHMM such that a higher probability is assigned to exons that contain homologues to protein domains. To our knowledge, this domain homology based approach has not been used previously in the context of ab initio gene prediction. Results: GlimmerHMM was augmented with a protein domain module that recognizes gene structures that are similar to Pfam models. The augmented system, GlimmerHMM+, shows 2% improvement in sensitivity and a 1% increase in specificity in predicting exact gene structures compared to GlimmerHMM without this option. These results were obtained on two very different model organisms: Arabidopsis thaliana (mustard wee) and Danio rerio (zebrafish), and together these preliminary results demonstrate the value of using protein domain homology in gene prediction. The results obtained are encouraging, and we believe that a more comprehensive approach including a model that reflects the statistical characteristics of specific sets of protein domain families would result in a greater increase of the accuracy of gene prediction. GlimmerHMM and GlimmerHMM+ are freely available as open source software at http://cbcb.umd.edu/software.

AB - Background: Protein domains are the common functional elements used by nature to generate tremendous diversity among proteins, and they are used repeatedly in different combinations across all major domains of life. In this paper we address the problem of using similarity to known protein domains in helping with the identification of genes in a DNA sequence. We have adapted the generalized hidden Markov model (GHMM) architecture of the ab intio gene finder GlimmerHMM such that a higher probability is assigned to exons that contain homologues to protein domains. To our knowledge, this domain homology based approach has not been used previously in the context of ab initio gene prediction. Results: GlimmerHMM was augmented with a protein domain module that recognizes gene structures that are similar to Pfam models. The augmented system, GlimmerHMM+, shows 2% improvement in sensitivity and a 1% increase in specificity in predicting exact gene structures compared to GlimmerHMM without this option. These results were obtained on two very different model organisms: Arabidopsis thaliana (mustard wee) and Danio rerio (zebrafish), and together these preliminary results demonstrate the value of using protein domain homology in gene prediction. The results obtained are encouraging, and we believe that a more comprehensive approach including a model that reflects the statistical characteristics of specific sets of protein domain families would result in a greater increase of the accuracy of gene prediction. GlimmerHMM and GlimmerHMM+ are freely available as open source software at http://cbcb.umd.edu/software.

KW - GHMM

KW - Pfam

KW - Profile HMM

KW - Protein domain

KW - ab intio gene finding

UR - http://www.scopus.com/inward/record.url?scp=37249023663&partnerID=8YFLogxK

UR - http://www.scopus.com/inward/citedby.url?scp=37249023663&partnerID=8YFLogxK

M3 - Conference contribution

AN - SCOPUS:37249023663

SN - 9783540741251

T3 - Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)

SP - 208

EP - 215

BT - Algorithms in Bioinformatics - 7th International Workshop, WABI 2007, Proceedings

T2 - 7th International Workshop on Algorithms in Bioinformatics, WABI 2007

Y2 - 8 September 2007 through 9 September 2007

ER -

Using protein domains to improve the accuracy of Ab Initio gene finding

Abstract

Publication series

Other

Keywords

ASJC Scopus subject areas

Other files and links

Fingerprint

Cite this