{"id":18854,"date":"2018-03-25T09:48:36","date_gmt":"2018-03-25T09:48:36","guid":{"rendered":"http:\/\/biomedpharmajournal.org\/?p=18854"},"modified":"2020-04-23T06:00:55","modified_gmt":"2020-04-23T06:00:55","slug":"webbioret-a-webtool-for-accessing-multiple-biological-sequences-with-features","status":"publish","type":"post","link":"https:\/\/biomedpharmajournal.org\/staging\/vol11no1\/webbioret-a-webtool-for-accessing-multiple-biological-sequences-with-features\/","title":{"rendered":"WEBBIORET \u2013 A WebTool for Accessing Multiple Biological Sequences with Features"},"content":{"rendered":"<p><strong>Introduction<\/strong><\/p>\n<p>Bioinformatics is an interdisciplinary field which is mainly utilized to develop methods and software tools for understanding different types of \u00a0biological data[1]. Biological data can be broadly classified as Genomics and Proteomics Data.<\/p>\n<p>Drug discovery [2] is a crucial implementation area of Bioinformatics and ever going research is taking place in that domain for years. One such sub area is the studies relating to cellular activities and disease states in humans \u00a0and other organisms[1]. Identification of DNA and protein sequences, protein domains and protein structures are very crucial at this point. So it can be understood that protein sequences are mandatory for these research activities.<\/p>\n<p>The challenge faced by the researcher is to retrieve protein sequences in bulk numbers from well known databases such as NCBI, PDB, Uniprot, etc. In most of the cases the researcher need to search with the protein id and copy the sequence from the global databases.<\/p>\n<p>Another issue is with the different formats of files given by the protein datasets. FASTA [3] is such a file type. It has a specific structure. The proposed tool bridges problems such as a) retrieve the sequence from the FASTA format and all other features such as sequence id, sequence description etc b) derive essential features according to the user attribute requirements c) give appropriate presentation of the data.<\/p>\n<p><strong>Background<\/strong><\/p>\n<p>In this section we discuss the important background information pertinent to this proposed work like Uniprot, AJAX and LAMP Server.<\/p>\n<p><strong>UniProt<\/strong><\/p>\n<p>The Universal Protein Resource (UniProt) is a commonly used data base for protein research. The data set is available in three categories such as UniProt Knowledgebase (UniProtKB), the\u00a0UniProt Reference Clusters (UniRef), and the\u00a0UniProt Archive (UniParc)[3] [4]. Uniprot\u2019s development was closely tied up with TrEMBL and Swiss-Prot.<\/p>\n<p>The reason for TrEMBL (Translated EMBL Nucleotide Sequence Data Library) development was due to the fact that the data was generated in a speed which couldn\u2019t be managed by Swiss-Prot database alone. In 2002 the three institutes decided to combine their resources and expertise and formed the UniProt consortium[3] [4].<\/p>\n<p><strong>AJAX<\/strong><\/p>\n<p>Ajax expands to\u201cAsynchronous\u00a0JavaScript\u00a0and\u00a0XML\u201d. This is a web technology. It is a group of interrelated\u00a0programming \u00a0techniques applied in the\u00a0client browser side\u00a0to create Web applications. With the help of Ajax, web programs can send transmit data without refresing page[6]\u00a0.<\/p>\n<p><strong>LAMP Server<\/strong><\/p>\n<p>LAMP is an application server platform used to develop websites and web tools. The powerful PHP Web Application server in combination with the powerful Relational Database Management System makes a unique combination. Figure 1 clearly shows the different layers in the architecture. LAMP consists of Apache Webserver, MySql database and PHP Application server.<\/p>\n<p>As a solution stack, LAMP is suitable for building interactive web softwares\u00a0[7].<\/p>\n<table style=\"width: 70%;\" border=\"1\" cellpadding=\"5\">\n<tbody>\n<tr>\n<td>\u00a0<img decoding=\"async\" class=\"alignnone size-thumbnail wp-image-18856\" src=\"https:\/\/biomedpharmajournal.org\/wp-content\/uploads\/2018\/02\/Vol11No1_Web_Saj_fig1-150x150.jpg\" alt=\"Figure 1: LAMP Architecture\" width=\"150\" height=\"150\" srcset=\"https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2018\/02\/Vol11No1_Web_Saj_fig1-150x150.jpg 150w, https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2018\/02\/Vol11No1_Web_Saj_fig1-256x256.jpg 256w, https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2018\/02\/Vol11No1_Web_Saj_fig1.jpg 658w\" sizes=\"(max-width: 150px) 100vw, 150px\" \/><\/td>\n<td><strong>Figure 1: LAMP Architecture<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p><a href=\"http:\/\/biomedpharmajournal.org\/wp-content\/uploads\/2018\/02\/Vol11No1_Web_Saj_fig1.jpg\" target=\"_blank\">Click here to View figure<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>&nbsp;<\/p>\n<p><strong>AAIndex<\/strong><\/p>\n<p>AAindex is knows as a database of numerical values on behalf of different physicochemical attributes of amino acids [8]. AAindex comprises of three sections now:, AAindex1, AAindex2 and AAindex3 [8].<\/p>\n<p>The complete database can be retrieved through the DBGET\/LinkDB system at GenomeNet (http:\/\/www.genome.jp\/dbget-bin\/www b_nd?aaindex) [9] or downloaded by anonymous FTP (ftp:\/\/ftp.genome.jp\/pub\/db\/community\/aaindex\/) [9]. Structures of Protein sequences and it\u2019s functions are given by the combinations of physicochemical and biochemical \u00a0attributes of 20 amino acids that are the building blocks of proteins[9]. Out of these properties an Amino acid Index can be created. From 222 amino acid indices Nakai et al. has done some research work\u00a0 to unearth the relationships among them using hierarchical cluster mechanisms [8][10]. Additionally, they released AAIndex2 after collection of 42 amino acid substitution matrices taken from literature.. Scientists are updating this AAIndex database in this manner [8] [12][13].<\/p>\n<p>AAindex is in wide use especially in research of various protein analysis of organisms [8] , such as Protein subcellular localization prediction, hub protein prediction, membrane protein prediction etc[8]. AAindex has become a really notable resource in bioinformatics research[8]. The AAindex is released almost every year. The latest version which is available is the 9.0 release.[8].<\/p>\n<p>The AAIndex1 currently contains 544 amino acid indices with its explanations[8].\u00a0 For\u00a0 20 amino acids each entry consists of a number called accession number, a short explanation of the index, the reference data and the numerical values for the protein properties[8] .<\/p>\n<p><strong>Proposed Tool<\/strong><\/p>\n<p>A web application is developed which takes an excel file as the maininput. This excel file contains the proteins ids where \u00a0protein sequences are to be retrieved. The\u00a0\u00a0\u00a0 Figure 2 contains a few protein ids in a spread sheet.<\/p>\n<table style=\"width: 70%;\" border=\"1\" cellpadding=\"5\">\n<tbody>\n<tr>\n<td>\u00a0<img decoding=\"async\" class=\"alignnone size-thumbnail wp-image-18857\" src=\"https:\/\/biomedpharmajournal.org\/wp-content\/uploads\/2018\/02\/Vol11No1_Web_Saj_fig2-150x150.jpg\" alt=\"Figure 2: Sample data as protein ids in a spread sheet\" width=\"150\" height=\"150\" srcset=\"https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2018\/02\/Vol11No1_Web_Saj_fig2-150x150.jpg 150w, https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2018\/02\/Vol11No1_Web_Saj_fig2-256x256.jpg 256w, https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2018\/02\/Vol11No1_Web_Saj_fig2.jpg 291w\" sizes=\"(max-width: 150px) 100vw, 150px\" \/><\/td>\n<td><strong>Figure 2: Sample data as protein ids in a spread sheet<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p><a href=\"http:\/\/biomedpharmajournal.org\/wp-content\/uploads\/2018\/02\/Vol11No1_Web_Saj_fig2.jpg\" target=\"_blank\">Click here to View figure<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>Using the web application this file is selected and uploaded. Once the file is uploaded the web program will start reading the protein ids one by one using PHP excel API and start retrieving the sequence using UNIPROT web service along with the required feature set.<\/p>\n<table style=\"width: 70%;\" border=\"1\" cellpadding=\"5\">\n<tbody>\n<tr>\n<td>\u00a0<img decoding=\"async\" class=\"alignnone size-thumbnail wp-image-18858\" src=\"https:\/\/biomedpharmajournal.org\/wp-content\/uploads\/2018\/02\/Vol11No1_Web_Saj_fig3-150x150.jpg\" alt=\"Figure 3: Site Homepage\" width=\"150\" height=\"150\" srcset=\"https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2018\/02\/Vol11No1_Web_Saj_fig3-150x150.jpg 150w, https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2018\/02\/Vol11No1_Web_Saj_fig3-256x256.jpg 256w, https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2018\/02\/Vol11No1_Web_Saj_fig3.jpg 835w\" sizes=\"(max-width: 150px) 100vw, 150px\" \/><\/td>\n<td><strong>Figure 3: Site Homepage<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p><a href=\"http:\/\/biomedpharmajournal.org\/wp-content\/uploads\/2018\/02\/Vol11No1_Web_Saj_fig3.jpg\" target=\"_blank\">Click here to View figure<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>&nbsp;<\/p>\n<table style=\"width: 70%;\" border=\"1\" cellpadding=\"5\">\n<tbody>\n<tr>\n<td>\u00a0<img decoding=\"async\" class=\"alignnone size-thumbnail wp-image-18859\" src=\"https:\/\/biomedpharmajournal.org\/wp-content\/uploads\/2018\/02\/Vol11No1_Web_Saj_fig4-150x150.jpg\" alt=\"Figure 4: Retrieved sequences along with sample features\" width=\"150\" height=\"150\" srcset=\"https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2018\/02\/Vol11No1_Web_Saj_fig4-150x150.jpg 150w, https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2018\/02\/Vol11No1_Web_Saj_fig4-256x256.jpg 256w, https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2018\/02\/Vol11No1_Web_Saj_fig4.jpg 706w\" sizes=\"(max-width: 150px) 100vw, 150px\" \/><\/td>\n<td><strong>Figure 4: Retrieved sequences along with sample features<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p><a href=\"http:\/\/biomedpharmajournal.org\/wp-content\/uploads\/2018\/02\/Vol11No1_Web_Saj_fig4.jpg\" target=\"_blank\">Click here to View figure<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>The web service retrieves the sequence in the form of FASTA file. \u00a0The web application retrieves the sequence from the file and displays in the screen in a neat tabular format as given in the Figure 3. Here along with the sequence 20 amino acid frequencies are also extracted from the sequence as they are considered as important biomarkers which describe the physicochemical property of the protein[12].<\/p>\n<p><strong>Conclusion<\/strong><\/p>\n<p>The programmed tool retrieves any number of sequences with the help of the protein ids stored in the spreadsheet file. But the retrieved data is represented in the form of html table data along with feature set.<\/p>\n<p>As a next step in this line the date retrieved will be stored in spreadsheet format along with the retrieved features. More than thousand amino acid features are relevant in the domain of proteomics research.\u00a0 All such features can be incorporated in the file and downloaded to the local file system of the researcher\u2019s computer for further analysis and studies. Wavelet features are also in our list for the next version of our tool. The site can be viewed in the address\u00a0 http:\/\/www.snit.ac.in\/research\/.<\/p>\n<p><strong>Conflict of interest <\/strong><\/p>\n<p>There is no conflict of interest between authors.<\/p>\n<p><strong>References<\/strong><\/p>\n<ol>\n<li>Opinion in Biotechnology, Volume 5, Issue 6, December 1994<\/li>\n<li>http:\/\/en.wikipedia.org\/wiki\/FASTA_format dated 18\/4\/2017 9.00 a.m.<\/li>\n<li>http:\/\/uniprot.org<\/li>\n<li>http:\/\/www.ebi.ac.uk\/<\/li>\n<li>http:\/\/en.wikipedia.org\/wiki\/Ajax_(prorammin) dated 18\/4\/2017 9.00 a.m.<\/li>\n<li>http:\/\/en.wikipedia.org\/wiki\/LAMP_(28software_bundle)<\/li>\n<li>Kenta Nakai, Akinori Kidera, and Minoru Kanehisa. Cluster analysis of amino acid indices for prediction of protein structure and function. Protein Engineering, 2(2):93{100, 1988.<\/li>\n<li>http:\/\/www.nar.oxfordjournals.org<\/li>\n<li>Kawashima, P.Pokarowski,M.Pokaro S. Kawashima, P. Pokarowski, M. Pokarowska, A. Kolinski, T. Katayama, M. Kanehisa. &#8220;AAindex: amino acid index database, progress report 2008&#8221;, Nucleic Acids Research, 2007<br \/>\n<a href=\"https:\/\/doi.org\/10.1093\/nar\/gkm998\" target=\"_blank\">CrossRef<\/a><\/li>\n<li>Kentaro Tomii and Minoru Kanehisa. Analysis of amino acid indices and mutation matrices for sequence comparison and structure prediction of proteins. Protein Engineering, 9(1):27{36, 1996.<\/li>\n<li>Shuichi Kawashima and Minoru Kanehisa. Aaindex: amino acid index database. Nucleic acids research, 28(1):374{374, 2000.<\/li>\n<li>Shuichi Kawashima, Hiroyuki Ogata, and Minoru Kanehisa. Aaindex: amino acid index database. Nucleic Acids Research, 27(1):368{369, 1999.<\/li>\n<li>Esmaeil Ebrahimie, Mansour Ebrahimi, Mahdi Ebrahimi, \u201cAmino acid features: a missing compartment of prediction of protein function\u201d, Nature Proceedings, doi:10.1038.\/npre.2011.6693.1.<\/li>\n<\/ol>\n","protected":false},"excerpt":{"rendered":"<p>Introduction Bioinformatics is an interdisciplinary field which is mainly utilized  [&#8230;]<\/p>\n","protected":false},"author":9,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[55],"tags":[],"class_list":["post-18854","post","type-post","status-publish","format-standard","hentry","category-vol11no1"],"_links":{"self":[{"href":"https:\/\/biomedpharmajournal.org\/staging\/wp-json\/wp\/v2\/posts\/18854","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/biomedpharmajournal.org\/staging\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/biomedpharmajournal.org\/staging\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/biomedpharmajournal.org\/staging\/wp-json\/wp\/v2\/users\/9"}],"replies":[{"embeddable":true,"href":"https:\/\/biomedpharmajournal.org\/staging\/wp-json\/wp\/v2\/comments?post=18854"}],"version-history":[{"count":5,"href":"https:\/\/biomedpharmajournal.org\/staging\/wp-json\/wp\/v2\/posts\/18854\/revisions"}],"predecessor-version":[{"id":32195,"href":"https:\/\/biomedpharmajournal.org\/staging\/wp-json\/wp\/v2\/posts\/18854\/revisions\/32195"}],"wp:attachment":[{"href":"https:\/\/biomedpharmajournal.org\/staging\/wp-json\/wp\/v2\/media?parent=18854"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/biomedpharmajournal.org\/staging\/wp-json\/wp\/v2\/categories?post=18854"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/biomedpharmajournal.org\/staging\/wp-json\/wp\/v2\/tags?post=18854"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}