flâneur

Annotating proteins based on critical functions · Arcadia Science

research.arcadiascience.com · saved by 1 readers

DNA sequencing methods are rapidly advancing and producing mountains of high-quality protein sequence data. However, protein sequences are frequently only as useful as their functional annotations, and predicting a protein’s function remains a bottleneck in the meaningful analysis of protein sequence data [1]. Cellular functions can define protein identity. We are developing a framework to computationally predict and validate annotations based on protein functions. These annotations should help scientists uncover novelty and generate hypotheses. Most annotations are based on information gathered in a small selection of organisms. In fact, 85% of Gene Ontology annotations are based on information from just ten species, including humans and other typical model organisms [1]. Even in some of the most well-studied organisms, annotation remains difficult. For example, in E. coli, a third of the genome remains un-annotated [2] and in both fission and budding yeast, roughly 20% of the genome

DNA sequencing methods are rapidly advancing and producing mountains of high-quality protein sequence data. However, protein sequences are frequently only as useful as their functional annotations, and predicting a protein’s function remains a bottleneck in the meaningful analysis of protein sequence data [1]. Cellular functions can define protein identity. We are developing a framework to computationally predict and validate annotations based on protein functions. These annotations should help scientists uncover novelty and generate hypotheses. Most annotations are based on information gather