None declared

None declared. == REFERENCES ==. continue to be decided at an exponential rate, both due to improvements in structure determination techniques and the increase of involved investigators (13). Additionally , these structures are increasingly of larger complexes such as ribosomes and viral capsids, mostly due to the rise of cryo-electron microscopy (cryoEM) (4, 5). Analysis BYK 49187 of domains within newly released protein structures can lead to hypotheses about function and evolutionary origins, but classification of these domains can be time-consuming. Reducing the burden of the domain classification process has typically been achieved by judicious selection of representatives to classify or by implementing automated Rabbit polyclonal to ANKRA2 procedures to supplement, aid and replace elements of the manual curation process (6). We developed the Evolutionary Classification Of protein Domains (or ECOD) (7) as a hierarchal classification, which emphasizes distantly related homologs that are difficult to detect (H- and X-groups) and takes into account closer sequence-based relationships between protein domains that are placed in family members. ECOD exclusively classifies residues in the Protein Data Financial BYK 49187 institution (PDB), i. e. any given residue appears once and only once in the classification. A distinctive feature of ECOD is that it explicitly classifies domains by topology at a lower level (T-groups), while classifying domains by evolutionary relatedness at a higher level (H-groups) that accounts for significant structural changes in protein evolution. ECOD releases are coupled to PDB releases (with a few weeks delay), such that for every week that there is a PDB release, there is also an ECOD release. We achieve this accelerated update schedule through a combination of automatic and manuals updates: a pipeline, which fully partitions and assigns nearly all proteins in any given week, leaving a fraction that is classified by a manual curator. We have previously discussed challenging examples of manual curation, and shown how strict reliance on structural or sequence similarity scores is insufficient to achieve accurate classification in most difficult cases (8). == BYK 49187 Improved performance of automated ECOD updates == Following the initial release of ECOD (v22), we implemented a weekly upgrade pipeline intended for ECOD coupled to the release of new structures from the PDB. By quickly and efficiently classifying known structures, and through dedicated manual curation, we are able to classify all depositions in the PDB without overburdening manual curators or relying on solely classifying a reduced set of representative structures. Briefly, the ECOD upgrade pipeline separates a weekly PDB release into a set of individual protein queries based on peptide chains within the PDB depositions (putative fragments or peptides are removed early in the process and BYK 49187 either placed into the peptide/fragment categories or reincorporated into ECOD as segments of multi-chain domains). Each member of this set of peptide chains is then individually queried against ECOD reference libraries using a combination of sequence (BLAST, HHsearch) and structural (DALI) aligners (912). BLAST alignments to well-covered ECOD reference proteins are used to directly partition query proteins in many (90%) cases. Where well-scoring full-protein alignments are not available, individual highly-covered domain hits by BLAST or HHsearch are used to partition unassigned regions of the query. We initiated weekly updates in February 2014 with ECOD version 33. In the following 123 weekly PDB releases, 187 51 structures were classified each week on average (Figure1A). These classified structures were incorporated into subsequent ECOD releases. Where the automatic upgrade pipeline could generate a putative domain architecture that significantly (> 90% and <20 residues uncovered) covered the query chain, those putative domains were assigned to the hierarchy with their hit domains. Where no hits were found, or only a partial domain solution could be resolved, chains were passed along with alignment data to the manual curator (Figure1B). Curated chains were either assigned to ECOD using a combination of alignment data, functional considerations and/or topological similarities to known domains, or assigned to one of several special architectures, which annotate those residues that are either unclassifiable by our current methodology, or lack sufficient data to be classified in any case (i. e. low resolution structures, peptides, fragments). Total chains partitioned and assigned by manual curation declined over time, and the fraction of manual curation dedicated to assigning unclassified chains to special architectures, rather than the domain hierarchy, increased (Figure1C). Nearly all representative peptide chains (89%) in PDB.