{"id":13448,"date":"2017-03-25T10:36:58","date_gmt":"2017-03-25T10:36:58","guid":{"rendered":"http:\/\/biomedpharmajournal.org\/?p=13448"},"modified":"2020-04-23T10:13:02","modified_gmt":"2020-04-23T10:13:02","slug":"classification-of-histopathological-images-of-breast-cancerous-and-non-cancerous-cells-based-on-morphological-features","status":"publish","type":"post","link":"https:\/\/biomedpharmajournal.org\/staging\/vol10no1\/classification-of-histopathological-images-of-breast-cancerous-and-non-cancerous-cells-based-on-morphological-features\/","title":{"rendered":"Classification of Histopathological images of Breast Cancerous and Non Cancerous Cells Based on Morphological features"},"content":{"rendered":"<p><strong>Introduction<\/strong><\/p>\n<p>The extensive use of computer aided diagnosis (CAD) these days can be traced back to the appearance of digital histopathology. Now a day, CAD has become a part of routine clinical detection methods for cancer diagnosis using digitized histological images at various screening centre and hospitals and hence it has become one of the most important key research subjects in histopathological imaging and diagnostic<sup>1<\/sup>. There is an imperative need for CAD to minimize the human error. Cancer is a vital health issue, which can cause the death of the living being. In 2015 an average 1,658,370 cancer cases along with 589,430 cancer deaths were reported in the US<sup>2<\/sup>. Among the developing countries, India suffers leading cause of death from cancer. These include mostly cervix, breast, skin and oral cancer. Cancer begins, when a gene in a cell becomes abnormal, and the cell starts to grow and divide in an uncontrolled manner. Cancerous cells replicate much faster than normal cells. Cells divide and multiply to form a tumor that may be cancerous and non-cancerous<sup>3<\/sup>. Cellular morphology is one of the best and preferred methods in identifying the abnormalities and the physiological state of the cells present in the tissue. Several biological functions seem to be related to significant changes in the geometry of the cells and nuclei<sup>4<\/sup>. Human error occurs due to several reasons while observing the image overlapping, blurriness, artifacts, weak boundary detection, and uneven dying. Visualizing the cells particularly consists of observing minute structures, composition, functions, distribution of cells and regularities of cell shapes across the tissue which helps the pathologist to make a decision of cells whether it is cancerous and non-cancerous. This whole process is very time consuming and cumbersome that requires much experience<sup>5<\/sup>. To overcome these problems, CAD is an opinion to the pathologist for the development of automation in biological image enhancement, feature extraction and classification for disease identification. Therefore, automation prevails over limitations of manual microscopy based detection method to help pathology practice. Several images analysis software like Cell profiler, Mac Biophotonics Image J and Scion has played a major role in the analysis of cellular images. CAD has been successfully applied in breast cancer, cervical cancer, prostate cancer and lung cancer analysis<sup>6<\/sup>. Arevalo et al.<sup>7<\/sup> have presented a review on automatic image analysis tasks and its current trends like digital pathology in histopathology image. Dermir and Yerner<sup>8 <\/sup>presented a cellular level automatic diagnosis of biopsy image using image processing techniques. Singh et al.<sup>9<\/sup> applied contrast limited adaptive histogram equalization (CLAHE) methods in biopsy image of cancer cells detection and classification. Wilkinson et al.<sup>10<\/sup> proposed segmentation method using robust automatic threshold selection (RATS) for microbes\u2019 image analysis, in that they have reported that RATS is suitable for thresholding a noisy image with the variable background. He Le et al.<sup>11<\/sup> presented an algorithm for Gaussian mixture modeling thresholding for hematoxylin and eosin stained histology image segmentation. Tissue constituents such as nuclei, stroma, and connecting contents from the background are extracted using these models. Bredfeldt et\u00a0 al.<sup>12<\/sup> have demonstrated a protocol that allowed consistent scoring throughout large patient cohorts in two steps; the first step involves the use of Trainable Weka Segmentation (TWS) Image J plugin for finding epithelial cell nuclei and other involves the application of a cascaded matched filter, threshold operation to identify clusters and boundaries. Sheha et al.<sup>13<\/sup> difference between Malignant Melanoma and Melanocytic Nevi based classifications are proposed on gray level co-occurrence matrix (GLCM) by using multi layer perceptron (MLP). For discrimination of melanocytic skin tumors, texture analysis can be used for high accuracy. Amaral et al.<sup>14<\/sup> presented a computational pipeline for automatically classifying and scoring breast cancer TMA spots mapped onto an ordinal scale used by pathologists. MLP classifier is compared with support vector machines and latent topic models for spot classification and with Gaussian process ordinal regression and linear models for scoring. Landwehr et al.<sup>15<\/sup> have developed algorithms for accurate and compact classifiers by evaluating the performance of logistic modal tree (LMT) on 36 datasets collected from the UCI repository. Zhang et al.<sup>16<\/sup> worked on breast cancer images with combined multiple features using the curvelet transform, statistics of completed local binary patterns (CLBP), and GLCM with a classifier Random Subspace Ensemble (RSE), with classification rate 95.22%. Nguyen et al.<sup>17<\/sup> proposed a method, to calculate the tubule percentage (TP), i.e., the ratio of the tubule area to the total glandular area for 353 Hematoxylin and Eosin images of the three TSs, and plot the distribution of these TP values. This plot shows the clear division among these three scores, suggesting that the proposed algorithm is useful in distinguishing images of these TSs by using a random forest classifier. George et al.<sup>18<\/sup> evaluated datasets 92 fine needle aspiration cytology (FNAC) images to classify the benign and malignant of breast tumor. The predictive ability of support vector machine (SVM) and probabilistic neural networks (PNN) are stronger than the MLP using back-propagation algorithm and learning vector quantization (LVQ).<\/p>\n<p>The aim of present work is to develop an automated system for detecting the cancer cells using shape and morphological features extracted from the segmented images. Segmentation is done using mixture modeling thresholding (MMT), simple interactive object extraction (SIOX), RATS, and TWS methods. Classification is done by MLP, LMT, sequential minimal optimization (SMO), Na\u00efve Bayes, Random Forest, Rotation Forest, J-Rip and PART which is trained using histopathology images of cancerous and non-cancerous categories. The efficiency of these classifiers is compared to each other for the identification of best classifier.<\/p>\n<p><strong>Materials and Methods<\/strong><\/p>\n<p><strong>Image Acquisition<\/strong><\/p>\n<p>Breast cancer cellular datasets used in present work has been obtained from www.bioimage.ucsb.edu. The study consists of 70 histopathology images (35 non-cancerous and 35 cancerous). Structural and intensity based 16 features are acquired to classify non-cancerous and cancerous cells. The images are hematoxylin and eosin stained to visualize various parts, cellular structures such as cells, nuclei, and cytoplasm of the tissue. The nuclei are stained blue with hematoxylin while cytoplasm and extra-cellular components are in pink due to eosin staining. Flowchart for the present work is shown in figure 1 which describes basic steps involved in the cells morphology image analysis.<\/p>\n<table style=\"width: 70%;\" border=\"1\" cellpadding=\"5\">\n<tbody>\n<tr>\n<td><img decoding=\"async\" class=\"size-thumbnail wp-image-13449 alignleft\" src=\"https:\/\/biomedpharmajournal.org\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_fig1-150x150.jpg\" alt=\"Figure 1: Schematic flowchart of the proposed method\" width=\"150\" height=\"150\" srcset=\"https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_fig1-150x150.jpg 150w, https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_fig1-256x256.jpg 256w, https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_fig1.jpg 349w\" sizes=\"(max-width: 150px) 100vw, 150px\" \/><\/td>\n<td><strong>Figure 1: Schematic flowchart of the proposed method<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p><a href=\"http:\/\/biomedpharmajournal.org\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_fig1.jpg\" target=\"_blank\">Click here to View figure<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>&nbsp;<\/p>\n<p><strong>Image Pre-Processing <\/strong><\/p>\n<p><strong>\u00a0<\/strong>In histopathology images, the blurriness, artifacts, weak boundary detection and overlapping problem occurred due to uneven staining of the slide as a result of human error. To eradicate these types of irregularities or uneven staining, the CLAHE method is proposed. CLAHE algorithm improves the image contrast by improving the local contrast present in an image and also by enhancing the weak boundary edges in each pixel of an image through limited amplification<sup>19<\/sup>. Digital image processing techniques interpret the result in a much better way than conventional methods. So it is well suited for features enhancement of histopathology images. CLAHE method has been used for pre-processing of images in figure 2.<\/p>\n<table style=\"width: 70%;\" border=\"1\" cellpadding=\"5\">\n<tbody>\n<tr>\n<td style=\"text-align: center;\"><img decoding=\"async\" class=\"size-thumbnail wp-image-13451 alignleft\" src=\"https:\/\/biomedpharmajournal.org\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_fig2-150x150.jpg\" alt=\"Figure 2: (A) Original non-cancerous image and (B) Enhanced non-cancerous image using CLAHE method\" width=\"150\" height=\"150\" srcset=\"https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_fig2-150x150.jpg 150w, https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_fig2-256x256.jpg 256w, https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_fig2.jpg 406w\" sizes=\"(max-width: 150px) 100vw, 150px\" \/><\/td>\n<td><strong>\u00a0Figure 2: (A) Original non-cancerous image and (B) Enhanced non-cancerous image using CLAHE method<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p><a href=\"http:\/\/biomedpharmajournal.org\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_fig2.jpg\" target=\"_blank\">Click here to View figure<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>&nbsp;<\/p>\n<p><strong>Segmentation<\/strong><\/p>\n<p>In digital pathology, segmentation of histopathology sections is a ubiquitous requirement due to the large variability of histopathology tissue. Further machine learning techniques play a vital role in delivering superior performance over standard image processing methods. During image analysis, the segmentation process is an essential domain. It is used to locate objects and boundaries in an image<sup>20<\/sup>. The proposed method, pre-processing steps involve, removing noise and enhancing the contrast for segmentation purpose. The basic purpose of segmentation is to extract the important features from the image and perceive the information. Selection of appropriate segmentation methods depends on the type of features that has to be maintained for detection. Segmentation methods like MMT, SIOX, RATS, and TWS has proposed from Fiji open access free software for image analysis<sup>21<\/sup>. Mixture modeling algorithm uses Gaussian model to separate the histogram of an image into two Gaussian classes based on average, standard deviation and thresholding<sup>22<\/sup>. SIOX is a method used for extracting foreground information from a colored (RGB) image<sup>23<\/sup>. RATS measure the threshold map of an image based on pixels value and the corresponding gradients value<sup>24<\/sup>. TWS is a pixel-based segmentation method which combines machine learning algorithms with selected set of image features<sup>25<\/sup>.<\/p>\n<p>The performance of various segmentations is quantified regarding the global consistency error (GCE), variation of information (VI) and probabilistic rand index (PRI) of the segmented image with the ground truth image. The brief description of these performance measures is as follows.<\/p>\n<p>GCE is calculated as follows: let us assume segments \ud835\udc60<sub>\ud835\udc56<\/sub> and \ud835\udc54\ud835\udc57 contain a pixel, say \ud835\udc58, such that \ud835\udc60 \u03f5 \ud835\udc46, \ud835\udc54 \u03f5 \ud835\udc3a where \ud835\udc46 represents the set of segments that are generated by the segmentation algorithm being evaluated and \ud835\udc3a denotes the set of reference segments. To start with, a measure of local refinement error is estimated using Eq. (1) and then local and global consistency errors are computed, where denotes the set of difference operation and \ud835\udc45 (\ud835\udc65,\ud835\udc66)\u00a0 represents the set of pixels corresponding to a region <em>x<\/em> that includes pixel \ud835\udc66. Using eq. (2) GCE forces all local refinements in the same direction<sup>26 <\/sup>which are computed using Eq. (2) where \ud835\udc5b denotes the total number of pixels of the image. GCE quantify the amount of error in the segmentation (0 signifies no error and 1 indicates no agreement):<\/p>\n<p><img decoding=\"async\" class=\"alignnone size-full wp-image-13453\" src=\"https:\/\/biomedpharmajournal.org\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_f12.jpg\" alt=\"Formula 1,2\" width=\"590\" height=\"106\" srcset=\"https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_f12-300x54.jpg 300w, https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_f12.jpg 590w\" sizes=\"(max-width: 590px) 100vw, 590px\" \/><\/p>\n<p>VI is a measure of the distance between two clusters (partitions of elements)<sup>27<\/sup>. Clustering with clusters is denoted by a random variable \ud835\udc4b, \ud835\udc4b = {1, . . , \ud835\udc58} such that \ud835\udc43<sub>\ud835\udc56<\/sub> = |\ud835\udc4b <sub>\ud835\udc56<\/sub> |\/\ud835\udc5b, \ud835\udc56 \u00a0\u03f5 \ud835\udc4b, and \ud835\udc5b = \u03a3<sub>\ud835\udc56<\/sub> \ud835\udc4b<sub>\ud835\udc56<\/sub> is the variation of information between two clusters \u00a0\ud835\udc4b and \ud835\udc4c. \u00a0Thus VI (\ud835\udc4b, \ud835\udc4c) is represented using<\/p>\n<p>VI (\ud835\udc4b, \ud835\udc4c) = \ud835\udc3b (\ud835\udc4b) = \ud835\udc3b (\ud835\udc4c) \u2212 2\ud835\udc3c (\ud835\udc4b, \ud835\udc4c), \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0(3)<\/p>\n<p>Where \ud835\udc3b (\ud835\udc4b) is entropy of \ud835\udc4b and \ud835\udc3c (\ud835\udc4b, \ud835\udc4c) is common information between \ud835\udc4b and \ud835\udc4c. VI (\ud835\udc4b, \ud835\udc4c) measures reduction in cluster assignment in clustering \ud835\udc4b into the uncertainty of item\u2019s cluster in clustering \ud835\udc4c. PRI is the nonparametric measure of goodness of segmentation algorithms<sup>28<\/sup>. Rand index between test (<em>S<\/em>) and ground truth (<em>G<\/em>) is estimated by adding the number of pixel pairs with the same label and some pixel pairs having different labels in both <em>S<\/em> and \ud835\udc3a then dividing it by a total number of pixel pairs. Given a set of ground truth segmentations <em>G<\/em><sub>\ud835\udc58<\/sub>, the PRI is estimated using Eq. (4) such that <em>c<\/em><sub>\ud835\udc56\ud835\udc57<\/sub> is an event that describes a pixel pair (\ud835\udc56, \ud835\udc57) having the same or different label in the test image test<\/p>\n<p><img decoding=\"async\" class=\"alignnone size-full wp-image-13454\" src=\"https:\/\/biomedpharmajournal.org\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_f4.jpg\" alt=\"Formula 4\" width=\"638\" height=\"39\" srcset=\"https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_f4-300x18.jpg 300w, https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_f4.jpg 638w\" sizes=\"(max-width: 638px) 100vw, 638px\" \/><\/p>\n<p>GCE and VI should be low, where as PRI should be high for a better segmented cells in image. The MMT, SIOX, and RATS method have high GCE and VI where as low PRI in comparison to TWS, which shows an edge of proposed TWS method over conventional methods. TWS gives better result which is shown in figure 3(F) for non-cancerous cells and figures 4(F) for cancerous cells because TWS uses random forest machine learning algorithm for image segmentation. There is no overlapping in the cells and shows cells separated well from each other. This is providing the most accurate shape of the cells as compare to other methods.<\/p>\n<table style=\"width: 70%;\" border=\"1\" cellpadding=\"5\">\n<tbody>\n<tr>\n<td style=\"text-align: center;\"><img decoding=\"async\" class=\"size-thumbnail wp-image-13455 alignleft\" src=\"https:\/\/biomedpharmajournal.org\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_fig3-150x150.jpg\" alt=\"Figure 3: Segmentation of normal cells from histopathology images by different methods (A) Original image, (B) Ground truth image, (C) MMT, (D) SIOX, (E) RATS, (F) TWS\" width=\"150\" height=\"150\" srcset=\"https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_fig3-150x150.jpg 150w, https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_fig3-256x256.jpg 256w, https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_fig3.jpg 466w\" sizes=\"(max-width: 150px) 100vw, 150px\" \/><\/td>\n<td><strong>Figure 3: Segmentation of normal cells from histopathology images by different methods (A) Original image, (B) Ground truth image, (C) MMT, (D) SIOX, (E) RATS, (F) TWS<\/strong><\/p>\n<p><a href=\"http:\/\/biomedpharmajournal.org\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_fig3.jpg\" target=\"_blank\">Click here to View figure<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>&nbsp;<\/p>\n<table style=\"width: 70%;\" border=\"1\" cellpadding=\"5\">\n<tbody>\n<tr>\n<td style=\"text-align: center;\"><img decoding=\"async\" class=\"size-thumbnail wp-image-13456 alignleft\" src=\"https:\/\/biomedpharmajournal.org\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_fig4-150x150.jpg\" alt=\"Figure 4: Segmentation of cancerous cells from histopathology images by different methods (A) Original image, (B) Ground truth image, (C) MMT, (D) SIOX, (E) RATS, (F) TWS\" width=\"150\" height=\"150\" srcset=\"https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_fig4-150x150.jpg 150w, https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_fig4-256x256.jpg 256w, https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_fig4.jpg 471w\" sizes=\"(max-width: 150px) 100vw, 150px\" \/><\/td>\n<td><strong>Figure 4: Segmentation of cancerous cells from histopathology images by different methods (A) Original image, (B) Ground truth image, (C) MMT, (D) SIOX, (E) RATS, (F) TWS<\/strong><\/p>\n<p><a href=\"http:\/\/biomedpharmajournal.org\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_fig4.jpg\" target=\"_blank\">Click here to View figure<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>&nbsp;<\/p>\n<p>The ROI of the segmented histopathology image is compared to ground truth images for the quantitative assessment of different segmentation methods by GCE, VI and PRI, for 25 sample images from histopathology dataset. Hence, TWS is associated with the lower value of GCE and VI and higher value of PRI in comparison to others methods which perform better regarding all parameters as shown in table 1 and graphical representation in figure 5. So it is chosen as the segmentation method in the proposed work for cancer detection from histopathology images.<\/p>\n<p><strong>Table 1:<\/strong> <strong>Quantitative comparison of segmentation methods on the basis of average values of 25 images.<\/strong><\/p>\n<table style=\"width: 95%;\" border=\"1\" cellspacing=\"0\" cellpadding=\"4\">\n<tbody>\n<tr>\n<td style=\"text-align: center;\"><strong>Segmentation Methods<\/strong><\/td>\n<td style=\"text-align: center;\" width=\"114\"><strong>\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0 PRI<\/strong><\/p>\n<p><strong>(Higher better)<\/strong><\/td>\n<td style=\"text-align: center;\" width=\"114\"><strong>GCE<\/strong><\/p>\n<p><strong>(Lower better)<\/strong><\/td>\n<td style=\"text-align: center;\" width=\"120\"><strong>\u00a0\u00a0\u00a0\u00a0\u00a0 VI<\/strong><\/p>\n<p><strong>(Lower better)<\/strong><\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: center;\">Mixture modeling<\/td>\n<td style=\"text-align: center;\" width=\"114\">0.95038<\/td>\n<td style=\"text-align: center;\" width=\"114\">0.028408<\/td>\n<td style=\"text-align: center;\" width=\"120\">0.303852<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: center;\">SIOX<\/td>\n<td style=\"text-align: center;\" width=\"114\">0.9734<\/td>\n<td style=\"text-align: center;\" width=\"114\">0.01608<\/td>\n<td style=\"text-align: center;\" width=\"120\">0.209856<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: center;\">RATS<\/td>\n<td style=\"text-align: center;\" width=\"114\">0.975016<\/td>\n<td style=\"text-align: center;\" width=\"114\">0.015312<\/td>\n<td style=\"text-align: center;\" width=\"120\">0.201652<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: center;\">Trainable Weka Segmentation.<\/td>\n<td style=\"text-align: center;\" width=\"114\">\u00a0\u00a0\u00a0\u00a0 0.976124<\/td>\n<td style=\"text-align: center;\" width=\"114\">0.013844<\/td>\n<td style=\"text-align: center;\" width=\"120\">0.19144<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>PRI- Probabilistic Rand Index, GCE- Global Consistency Error, VI- Variation of Information<\/p>\n<p>&nbsp;<\/p>\n<table style=\"width: 70%;\" border=\"1\" cellpadding=\"5\">\n<tbody>\n<tr>\n<td><img decoding=\"async\" class=\"alignnone size-thumbnail wp-image-13462\" src=\"https:\/\/biomedpharmajournal.org\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_fig5-150x150.jpg\" alt=\"Figure 5: Comparison of segmentation methods on the basis of average values of (A) PRI, (B) GCE and (C) VI for 25 sample images\" width=\"150\" height=\"150\" srcset=\"https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_fig5-150x150.jpg 150w, https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_fig5-256x256.jpg 256w, https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_fig5.jpg 841w\" sizes=\"(max-width: 150px) 100vw, 150px\" \/><\/td>\n<td><strong>Figure 5: Comparison of segmentation methods on the basis of average values of (A) PRI, (B) GCE and (C) VI for 25 sample images<\/strong><\/p>\n<p><a href=\"http:\/\/biomedpharmajournal.org\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_fig5.jpg\" target=\"_blank\">Click here to View figure<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>&nbsp;<\/p>\n<p><strong>Feature Extraction<\/strong><\/p>\n<p>Image morphology is a very powerful tool for analyzing the shapes of the objects and to extract the image features, which are necessary for object recognition<sup>29<\/sup>. The most significant portion of this work is the computation of features. Morphological and shape based features have been extracted after segmentation of image for further classification purpose. These features provide information regarding the size and shape of cells<sup>30<\/sup>. TWS method is considered for features extraction from the segmented cells of the images as shown in figure 3(F) and 4(F). Total 16 features have been used in this paper. The quantification of these features helps to differentiate the cancerous cells from non-cancerous cells. The features used in this paper are explained from F1 to F 16 in table 2. count,\u00a0 total area, average size, area fraction, perimeter, major axis length, minor axis length, angle, circularity, solidity, feret, feret X, feret Y, feret angle, min feret, and integrated density.<\/p>\n<table style=\"width: 70%;\" border=\"1\" cellpadding=\"5\">\n<tbody>\n<tr>\n<td><img decoding=\"async\" class=\"alignnone size-thumbnail wp-image-13464\" src=\"https:\/\/biomedpharmajournal.org\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_table2-150x150.jpg\" alt=\"Table 2: Description of morphological features\" width=\"150\" height=\"150\" srcset=\"https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_table2-150x150.jpg 150w, https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_table2-256x256.jpg 256w, https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_table2.jpg 873w\" sizes=\"(max-width: 150px) 100vw, 150px\" \/><\/td>\n<td><strong>Table 2: Description of morphological features<\/strong><\/p>\n<p><a href=\"http:\/\/biomedpharmajournal.org\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_table2.jpg\" target=\"_blank\">Click here to View\u00a0table<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>&nbsp;<\/p>\n<p><strong>Classification<\/strong><\/p>\n<p><em>\u00a0<\/em>Classification of non-cancerous and cancerous cells can be done based on the extracted features. Factors such as staining, artifact, noise, and blurriness cause variation in the image and result in misclassification. Hence, a good classifier should be capable of overcomes these flaws<sup>31<\/sup>. Moreover, the choice to classifier must be made by fast computation and proficient enough to meet good classification. Experiments are carried out using the Weka data mining for classification purpose. Supervised machine learning approaches have been used on the dataset of cancerous and non-cancerous histopathology images for classification. In this work 16 features of cells are extracted. The features obtained are the area, perimeter, major axis, etc. After features extraction, a dataset of order 70&#215;16 in arff (attribute relation file format) are prepared. In which 70 instances and 16 attributes are available. For classification, selected features get feed into various classifiers as mention above.<\/p>\n<p>A MLP is a classifier based on feed forward artificial neural network modal that uses back propagation to classify instances. It has much triumphant application in data classification. It consists of different layers having various nodes, which represents directed graph and every layer is fully connected with the further layer. The supervised learning process consists of input data y and target P, requires the objective function (Z, P) in order to evaluate the divergence of the predicted output values, Z=MLP(Y; K) from the observed data values P and employ that evaluation for the convergence towards an optimal set of weights k. Many MLP training algorithms use \u00a0radiant information whether directly or indirectly <sup>32<\/sup>.<\/p>\n<p>The LMT is a classification replica, which has an affiliated supervised learning algorithm that amalgamates logistic regression (LR) and decision tree learning. It is made of standard decision tree having logistic regression functions at the leave nodes, which is based on the concept of a modal tree. The leave nodes contain two child nodes. One of the child nodes represents left branch and other represent right branch by threshold values. Feature value which is smaller than a threshold is sorted to left and greater than a threshold is sorted to right branch<sup>33<\/sup>.<\/p>\n<p>Random forest proposed by Breiman is one type of ensemble learning process for classification and regression. A random forest is a multiway classifier composed of some trees, and each tree grows using randomization. The leaf nodes of each tree are labeled by approximation of the posterior distribution over the classes of image. This test has been done to split the space of data to be classified every interior node that contains a test that best splits it. Classification of an image takes place by sending it down every tree and after that aggregating the reached leaf distributions. Randomness can be inserted at two points during training and testing.\u00a0 This concept is used so that training process can be done by using different data subset. Randomness can be injected in selecting the node tests<sup>34<\/sup>. Large scale sample sets are trained that is based on decomposition and iteration. These methods decrease accuracy.<\/p>\n<p>SMO was introduced by John Pitt in 1998 at Microsoft research to solve this problem. It is used to solve the quadratic programming (QP) problem that appears during the training of support vector machines. SMO disintegrates the (QP) problem into sub problems, using Osuna\u2019s theorem which selects to resolve the smallest feasible optimization problem at every step. The smallest feasible optimization problem for the standard SVM-QP problem involves two large range multipliers must obey a linear equality constraint. It selects two langrange multipliers jointly to optimize at every step and tries to find the optimal values for these multipliers.\u00a0 After that updates, the SVM reflects the new optimal values which solves langrange multipliers analytically<sup>35<\/sup>.<\/p>\n<p>Rotation forest is assembled with independent decision trees. Each tree is trained with complete information system with a rotated feature space. It uses hyperplanes parallel to the feature axes and a small rotation of the axes guide to diverse trees. Rodriguez et al.<sup>36 <\/sup>done the comparative study and proved that rotation forest performs better than random forest, bagging and AdaBoost. It is devised that rotation forest produces more accurate classifiers than AdaBoost which are also more diverse than bagging.<\/p>\n<p>Na\u00efve Bayes classifiers are based on a probabilistic approach for classification hinged on Bufes&#8217;theorem with strong independence assumptions between the features. These classifiers are highly scalable. Na\u00efve Bayes nearest neighbor classifier (NBNN) is a not-parametric approach for image classification introduced by Bioman<sup>3<\/sup><sup>7<\/sup>. J-Rip is used to learn propositional rules by frequently developing rules and trimming them. Precursors are appended greedily until a termination condition is satisfied during the growth phase. After that antecedent is pruned in the upcoming phase on a pruning metric on one occasion, the rule set is generated. Optimization is required for the rules, which are evaluated by some criteria and deleted by their performance against those criteria on randomized data<sup>38<\/sup>. PART produces rules through frequently creating decision trees from data. The algorithm acquires a separate and conquers strategy in that. It abolishes instances covered by the ongoing rule set during processing. Essentially a rule is generated by constructing a pruned tree for the present set of instances; the leaf with the maximum coverage is converted into a rule<sup>39<\/sup>.<\/p>\n<p><strong>Performance Measure of Classifier<\/strong><\/p>\n<p>Performance evaluation of each classifier is considered using confusion matrix (2 \u00d7 2) of size. The value of True Positive (TP), True Negative (TN), False Positive (FP) and False Negative (FN) is calculated. TP is a condition where the system correctly identifies an abnormality, FN is a condition of system that incorrectly identifies abnormality as normality, TN is a condition where the system correctly identifies normality and FP is a condition where the system incorrectly identifies normality as an abnormality are shown in table 3.<\/p>\n<p><strong>Table 3: The diagram of FP, FN, TP and TN.<\/strong><\/p>\n<table style=\"width: 95%;\" border=\"1\" cellspacing=\"0\" cellpadding=\"4\">\n<tbody>\n<tr>\n<td width=\"136\"><\/td>\n<td style=\"text-align: center;\" width=\"104\"><strong>\u00a0<\/strong><\/td>\n<td style=\"text-align: center;\" colspan=\"2\" width=\"217\"><strong>System Decision<\/strong><\/td>\n<\/tr>\n<tr>\n<td width=\"136\"><\/td>\n<td style=\"text-align: center;\" width=\"104\"><strong>\u00a0<\/strong><\/td>\n<td style=\"text-align: center;\" width=\"92\"><strong>Abnormal<\/strong><\/p>\n<p><strong>(Cancerous)<\/strong><\/td>\n<td style=\"text-align: center;\" width=\"125\"><strong>Normal<\/strong><\/p>\n<p><strong>(Non-Cancerous)<\/strong><\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: center;\" rowspan=\"2\" width=\"136\">&nbsp;<\/p>\n<p>Truth of Clinical Situation<\/td>\n<td style=\"text-align: center;\" width=\"104\">Abnormal<\/td>\n<td style=\"text-align: center;\" width=\"92\">TP<\/td>\n<td style=\"text-align: center;\" width=\"125\">FN<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: center;\" width=\"104\">Normal<\/td>\n<td style=\"text-align: center;\" width=\"92\">FP<\/td>\n<td style=\"text-align: center;\" width=\"125\">TN<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>&nbsp;<\/p>\n<p>The performance parameters like accuracy, sensitivity, specificity, balanced classification rate (BCR), F-measure (F-m), Matthews&#8217; correlation coefficient (MCC) and area under the curve (AUC) are defined to assess the success of the diagnostic system and can be calculated using Eq.(5)-(10).The definition of these performance measures have been illustrated as follows<\/p>\n<p>Accuracy of a classification is technique that depends on the number of correctly classified samples (i.e. true negative and true positive) and calculated as follows<\/p>\n<p><img decoding=\"async\" class=\"alignnone size-full wp-image-13465\" src=\"https:\/\/biomedpharmajournal.org\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_f5.jpg\" alt=\"formula 5\" width=\"698\" height=\"50\" srcset=\"https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_f5-300x21.jpg 300w, https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_f5.jpg 698w\" sizes=\"(max-width: 698px) 100vw, 698px\" \/><\/p>\n<p>Where <em>N<\/em> is the total number of sample present in the histopathological images for testing.<br \/>\nSensitivity is the probability of a positive diagnosis test among persons that have the disease and it is defined as,<\/p>\n<p><img decoding=\"async\" class=\"alignnone wp-image-13466 size-full\" src=\"https:\/\/biomedpharmajournal.org\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_f6.jpg\" alt=\"formula 6\" width=\"626\" height=\"47\" srcset=\"https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_f6-300x23.jpg 300w, https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_f6.jpg 626w\" sizes=\"(max-width: 626px) 100vw, 626px\" \/><\/p>\n<p>Where, the value of sensitivity ranges between 0 (mean worst) and 1 (best classification) respectively.<br \/>\nSpecificity is the probability of a negative diagnosis test among persons that do not have the disease and it is defined as,<\/p>\n<p><img decoding=\"async\" class=\"alignnone size-full wp-image-13467\" src=\"https:\/\/biomedpharmajournal.org\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_f7.jpg\" alt=\"formula 7\" width=\"629\" height=\"51\" srcset=\"https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_f7-300x24.jpg 300w, https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_f7.jpg 629w\" sizes=\"(max-width: 629px) 100vw, 629px\" \/><\/p>\n<p>Its value ranges between 0 and 1, where 0 and 1, respectively, mean worst and best classification.<br \/>\nBCR is the geometric mean of sensitivity and specificity is measured as balance classification rate.\u00a0 It is represented by<\/p>\n<p><img decoding=\"async\" class=\"alignnone size-full wp-image-13468\" src=\"https:\/\/biomedpharmajournal.org\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_f8.jpg\" alt=\"formula 8\" width=\"542\" height=\"42\" srcset=\"https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_f8-300x23.jpg 300w, https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_f8.jpg 542w\" sizes=\"(max-width: 542px) 100vw, 542px\" \/><\/p>\n<p>The F-m is a measure of harmonic mean of precision and recall. It is defined by using<\/p>\n<p><img decoding=\"async\" class=\"alignnone size-full wp-image-13469\" src=\"https:\/\/biomedpharmajournal.org\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_for1.jpg\" alt=\"Vol10No1_Clas_Anur_for1\" width=\"375\" height=\"47\" srcset=\"https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_for1-300x38.jpg 300w, https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_for1.jpg 375w\" sizes=\"(max-width: 375px) 100vw, 375px\" \/><\/p>\n<p><img decoding=\"async\" class=\"alignnone size-full wp-image-13470\" src=\"https:\/\/biomedpharmajournal.org\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_f9.jpg\" alt=\"formula 9\" width=\"671\" height=\"53\" srcset=\"https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_f9-300x24.jpg 300w, https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_f9.jpg 671w\" sizes=\"(max-width: 671px) 100vw, 671px\" \/><\/p>\n<p>The value of measure ranges between 0 and 1, where 0 means the worst classification and 1 means the best classification.<br \/>\nMCC is a measure of the distinction of binary class classifications.\u00a0 It can be calculated using the following formula:<\/p>\n<p><img decoding=\"async\" class=\"alignnone size-full wp-image-13471\" src=\"https:\/\/biomedpharmajournal.org\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_f10.jpg\" alt=\"formula 10\" width=\"586\" height=\"53\" srcset=\"https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_f10-300x27.jpg 300w, https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_f10.jpg 586w\" sizes=\"(max-width: 586px) 100vw, 586px\" \/><\/p>\n<p>Its value ranges between \u22121 and +1, where \u22121, +1, and 0, respectively, correspond to worst, best, at random prediction.<br \/>\nAUC is used to measure the performance of the system. The AUC ranges from 0 to 1. The higher AUC is, the greater the probability of a true decision.<\/p>\n<p><strong>Results and Discussion<\/strong><\/p>\n<p>The proposed methodologies are implemented with image analysis software Fiji (www.fiji.net) for enhancement, segmentation and feature extraction on the dataset of digitized at 40x, magnification on PC with 3.4 GHz Intel Core i7 processor, 2GB RAM, and Windows 8.1 platform. For experimentation purposes, a total of 70 histopathology images are used. The dataset includes cancerous and non-cancerous images. The given methodology for diagnosis of cancer from histopathology images consists of image enhancement, segmentation, feature extraction, and classification. The CLAHE is used for enhancement of histopathology images because it has shown to better results. It highlights the region of interests in the images as tested through experimentation. The original image has been processed through following two preprocessing steps l. contrast enhancement, 2. bilateral filtering to remove the artifact, blurriness that has been introduced during the staining process and to produce a better contrast image of good quality as shown in figure 2.The segmentation has been done by following methods MMT, SIOX, RATS, TWS and afterward there results have been compared. TWS performs better in comparison to other methods as shown in figure 3(F) for non-cancerous cells and figures 4(F) for cancerous cells. In other segmentation techniques cells are overlapping but in TWS, there is no overlapping has been visualized.<br \/>\nIn feature extraction phase, various shape and morphology based features as shown in table 2 are extracted from the segmented images. Finally, a 2D matrix of order [70 \u00d7 16] feature is formed using all the feature sets, where 70 microscopic images in the dataset and 20 total numbers of features are extracted, further these features used for classification. The experiment is performed using 10-fold cross validation approach. The proposed framework for different histopathology images containing cancerous and non-cancerous features of cells are tested using eight popular classifiers like MLP, LMT, Random forest, Rotation forest, SMO, Na\u00efve Bayes, J-Rip and PART as shown in table 4 and graphical presentation shown in figure 6. Among all these classification methods rotation forest differentiated better between cancerous and non-cancerous cells with the accuracy of 85.7% and with maximum BCR value 0.806. The superiority of rotation forest measure lies in the application of rotation matrix, created by linear transformed subsets.<\/p>\n<p><strong>Table 4: Comparative performances of various classifiers<\/strong><\/p>\n<table style=\"width: 95%;\" border=\"1\" cellspacing=\"0\" cellpadding=\"4\">\n<tbody>\n<tr>\n<td style=\"text-align: center;\" width=\"14%\"><strong>Classifier<\/strong><\/td>\n<td style=\"text-align: center;\" width=\"14%\"><strong>Accuracy<\/strong><\/td>\n<td style=\"text-align: center;\" width=\"15%\"><strong>Sensitivity<\/strong><\/td>\n<td style=\"text-align: center;\" width=\"15%\"><strong>Specificity<\/strong><\/td>\n<td style=\"text-align: center;\" width=\"9%\"><strong>BCR<\/strong><\/td>\n<td style=\"text-align: center;\" width=\"9%\"><strong>F- m<\/strong><\/td>\n<td style=\"text-align: center;\" width=\"10%\"><strong>MCC<\/strong><\/td>\n<td style=\"text-align: center;\" width=\"9%\"><strong>AUC<\/strong><\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: center;\" width=\"14%\">MLP<\/td>\n<td style=\"text-align: center;\" width=\"14%\">0.800<\/td>\n<td style=\"text-align: center;\" width=\"15%\">0.829<\/td>\n<td style=\"text-align: center;\" width=\"15%\">0.771<\/td>\n<td style=\"text-align: center;\" width=\"9%\">0.701<\/td>\n<td style=\"text-align: center;\" width=\"9%\">0.794<\/td>\n<td style=\"text-align: center;\" width=\"10%\">0.601<\/td>\n<td style=\"text-align: center;\" width=\"9%\">0.892<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: center;\" width=\"14%\">LMT<\/td>\n<td style=\"text-align: center;\" width=\"14%\">0.829<\/td>\n<td style=\"text-align: center;\" width=\"15%\">0.914<\/td>\n<td style=\"text-align: center;\" width=\"15%\">0.743<\/td>\n<td style=\"text-align: center;\" width=\"9%\">0.710<\/td>\n<td style=\"text-align: center;\" width=\"9%\">0.813<\/td>\n<td style=\"text-align: center;\" width=\"10%\">0.667<\/td>\n<td style=\"text-align: center;\" width=\"9%\">0.920<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: center;\" width=\"14%\">Random Forest<\/td>\n<td style=\"text-align: center;\" width=\"14%\">0.800<\/td>\n<td style=\"text-align: center;\" width=\"15%\">0.829<\/td>\n<td style=\"text-align: center;\" width=\"15%\">0.771<\/td>\n<td style=\"text-align: center;\" width=\"9%\">0.701<\/td>\n<td style=\"text-align: center;\" width=\"9%\">0.794<\/td>\n<td style=\"text-align: center;\" width=\"10%\">0.601<\/td>\n<td style=\"text-align: center;\" width=\"9%\">0.886<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: center;\" width=\"14%\">Rotation Forest<\/td>\n<td style=\"text-align: center;\" width=\"14%\">0.857<\/td>\n<td style=\"text-align: center;\" width=\"15%\">0.829<\/td>\n<td style=\"text-align: center;\" width=\"15%\">0.886<\/td>\n<td style=\"text-align: center;\" width=\"9%\">0.806<\/td>\n<td style=\"text-align: center;\" width=\"9%\">0.861<\/td>\n<td style=\"text-align: center;\" width=\"10%\">0.715<\/td>\n<td style=\"text-align: center;\" width=\"9%\">0.884<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: center;\" width=\"14%\">Na\u00efve Bayes<\/td>\n<td style=\"text-align: center;\" width=\"14%\">0.829<\/td>\n<td style=\"text-align: center;\" width=\"15%\">0.857<\/td>\n<td style=\"text-align: center;\" width=\"15%\">0.800<\/td>\n<td style=\"text-align: center;\" width=\"9%\">0.740<\/td>\n<td style=\"text-align: center;\" width=\"9%\">0.824<\/td>\n<td style=\"text-align: center;\" width=\"10%\">0.658<\/td>\n<td style=\"text-align: center;\" width=\"9%\">0.855<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: center;\" width=\"14%\">SMO<\/td>\n<td style=\"text-align: center;\" width=\"14%\">0.857<\/td>\n<td style=\"text-align: center;\" width=\"15%\">0.914<\/td>\n<td style=\"text-align: center;\" width=\"15%\">0.800<\/td>\n<td style=\"text-align: center;\" width=\"9%\">0.764<\/td>\n<td style=\"text-align: center;\" width=\"9%\">0.848<\/td>\n<td style=\"text-align: center;\" width=\"10%\">0.719<\/td>\n<td style=\"text-align: center;\" width=\"9%\">0.857<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: center;\" width=\"14%\">J Rip<\/td>\n<td style=\"text-align: center;\" width=\"14%\">0.829<\/td>\n<td style=\"text-align: center;\" width=\"15%\">0.857<\/td>\n<td style=\"text-align: center;\" width=\"15%\">0.800<\/td>\n<td style=\"text-align: center;\" width=\"9%\">0.740<\/td>\n<td style=\"text-align: center;\" width=\"9%\">0.824<\/td>\n<td style=\"text-align: center;\" width=\"10%\">0.658<\/td>\n<td style=\"text-align: center;\" width=\"9%\">0.821<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: center;\" width=\"14%\">PART<\/td>\n<td style=\"text-align: center;\" width=\"14%\">0.771<\/td>\n<td style=\"text-align: center;\" width=\"15%\">0.771<\/td>\n<td style=\"text-align: center;\" width=\"15%\">0.771<\/td>\n<td style=\"text-align: center;\" width=\"9%\">0.676<\/td>\n<td style=\"text-align: center;\" width=\"9%\">0.771<\/td>\n<td style=\"text-align: center;\" width=\"10%\">0.543<\/td>\n<td style=\"text-align: center;\" width=\"9%\">0.749<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>BCR- Balanced Classification Rate, F-m- F-measure, MCC- Matthews\u2019s Correlation Coefficient and AUC-Area under the Curve<\/p>\n<p>&nbsp;<\/p>\n<table style=\"width: 70%;\" border=\"1\" cellpadding=\"5\">\n<tbody>\n<tr>\n<td><img decoding=\"async\" class=\"alignnone size-thumbnail wp-image-13472\" src=\"https:\/\/biomedpharmajournal.org\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_fig6-150x150.jpg\" alt=\"Figure 6: Graph for comparative performances of various classifiers\" width=\"150\" height=\"150\" srcset=\"https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_fig6-150x150.jpg 150w, https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_fig6-256x256.jpg 256w, https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_fig6.jpg 565w\" sizes=\"(max-width: 150px) 100vw, 150px\" \/><\/td>\n<td><strong>Figure 6: Graph for <\/strong><strong>comparative performances of various classifiers<\/strong><\/p>\n<p><a href=\"http:\/\/biomedpharmajournal.org\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_fig6.jpg\" target=\"_blank\">Click here to View figure<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>&nbsp;<\/p>\n<p>Features play a major role for the classification purpose. Ranking of the feature was done by their importance for the classification purpose. Ranks of all the features have obtained in the features vector by applying Releif F algorithms<sup>40<\/sup> in weka 3.8. Relief-F is to draw instances at random, compute their nearest neighbours, and change a feature weighting vector to give more weight to features that differentiate the instance from neighbours of different classes. In particular, it tries to get a better estimate of the following probability to allocate as the weight for each feature f.<\/p>\n<p>w<sub>f <\/sub>= <em>P<\/em> (different value of f | different class) \u2013 <em>P <\/em>(different value of f | same class)<\/p>\n<p>In the experiments, the ranks of the features of cells have been investigated, which are given in table 5 and graph in figure 7.<\/p>\n<table style=\"width: 70%;\" border=\"1\" cellpadding=\"5\">\n<tbody>\n<tr>\n<td>\u00a0<img decoding=\"async\" class=\"alignnone size-thumbnail wp-image-13475\" src=\"https:\/\/biomedpharmajournal.org\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_fig7-150x150.jpg\" alt=\"Figure 7: Graph for ranking of maximal relevance factor\" width=\"150\" height=\"150\" srcset=\"https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_fig7-150x150.jpg 150w, https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_fig7-256x256.jpg 256w, https:\/\/biomedpharmajournal.org\/staging\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_fig7.jpg 475w\" sizes=\"(max-width: 150px) 100vw, 150px\" \/><\/td>\n<td><strong>Figure 7: Graph for ranking of maximal relevance factor<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p><a href=\"http:\/\/biomedpharmajournal.org\/wp-content\/uploads\/2017\/02\/Vol10No1_Clas_Anur_fig7.jpg\" target=\"_blank\">Click here to View figure<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>&nbsp;<\/p>\n<p>Maximal relevance factor is derived for obtaining feature importance. Based on this factor selection of the important features of cells for the classification is done instead of using a large number of features making the computational work complex.<\/p>\n<p><strong>Conclusion<\/strong><\/p>\n<p>In this paper, an effective and automatic computer aided technique is proposed and use for pre-processing, segmentation and classification. The cells are classified by shape and morphological features. This work deals issues related with staining and with color consistency problems. These features contributed significantly to realize quantification, statistical analysis, and computer aided diagnosis, interactive systems to detect cancerous and non-cancerous cells. The method has great potential for assisting in the early detection of cancer. This provides good detection performance, where the background is complex and has the similar appearance with the foreground. The developed technique for automated analysis and evaluation of histopathology images will assist the pathologists and reduce the human error. Such automated cancer diagnosis facilitates improved judgment by the pathologist.<\/p>\n<p><strong>Acknowledgments<\/strong><\/p>\n<p>The first author gratefully acknowledges financial assistance in the form of Rajiv Gandhi National Fellowship.<strong>\u00a0<\/strong><\/p>\n<p><strong>Conflict of interests <\/strong><\/p>\n<p><strong>\u00a0<\/strong>The authors declare that there is no conflict of interests regarding the publication of this paper.<\/p>\n<p><strong>References<\/strong><\/p>\n<ol>\n<li>Chen M., Qu A.P., Wang L.W., Yuan J.P., Yang F., Xiang Q.M. and Li Y., New breast cancer prognostic factors identified by computer-aided image analysis of HE stained histopathology images,\u00a0<em>Scientific reports<\/em>.,5(2015)<\/li>\n<li>Siegel R.L., Miller K.D. and Jemal A., Cancer statistics 2015,\u00a0<em>CA: a cancer journal for clinicians.,<\/em>65(1),5-29(2015)<\/li>\n<li>Rubin R., Strayer D.S. and Rubin E., eds.,\u00a0Rubin&#8217;s pathology: clinicopathologic foundations of medicine, Lippincott Williams &amp; Wilkins,(2008)<\/li>\n<li>Chen S., Zhao M., Wu G., Yao C. and Zhang J., Recent advances in morphological cell image analysis,\u00a0<em>Computational and Mathematical Methods in Medicine<\/em>., (2012)<\/li>\n<li>Veta M., Pluim J.P., Van Diest P.J. and Viergever M., Breast cancer histopathology image analysis: A review,\u00a0<em>Biomedical Engineering IEEE Transactions on.,<\/em>61(5),1400-1411(2014)<\/li>\n<li>Zhang X., Liu W., Dundar M., Badve S. and Zhang S., Towards large-scale histopathological image analysis: Hashing-based image retrieval, <em>Medical Imaging IEEE Transactions on.<\/em>,34(2), 496-506 (2015)<\/li>\n<li>Arevalo J. and Cruz-Roa A., Histopathology image representation for automatic analysis: a state-of-the-art review,<em>Revista Med.<\/em>,22(2),79-91(2014)<\/li>\n<li>Demir C., Yener B., Automated cancer diagnosis based on histopathological images: a systematic survey,\u00a0Rensselaer Polytechnic Institute Tech. Rep,(2005)<\/li>\n<li>Singh S., Cancer cells detection and classification in biopsy image, <em>International Journal of Computer Applications.,<\/em>38 (3),15-21(2012)<\/li>\n<li>Wilkinson M. H., Wijbenga T., De Vries G. and Westenberg M. A., Blood vessel segmentation using moving-window robust automatic threshold selection. InImage Processing, ICIP 2003 International Conference on,IEEE.2,II-1093.(2003)<\/li>\n<li>He L., Rodney L., Antani S. and Thomas G.R., Local and global Gaussian mixture models for hematoxylin and eosin stained histology image segmentation, International Conference on Hybrid Intelligent Systems.,223\u2013228(2010)<\/li>\n<li>Bredfeldt J. S., Liu Y., Conklin M.W., Keely P.J., Mackie T.R. and Eliceiri K.W., Automated quantification of aligned collagen for human breast carcinoma prognosis,<em>Journal of pathology informatics.<\/em>,5(1),28(2014)<\/li>\n<li>Sheha M.A., Mabrouk M.S. and Sharawy A., Automatic detection of melanoma skin cancer using texture analysis,<em>International Journal of Computer Applications.<\/em>,42(20),22-26(2012)<\/li>\n<li>Amaral T., McKenna S.J., Robertson K. and Thompson A., Classification and immunohistochemical scoring of breast tissue microarray spots.<em>IEEE Transactions on Biomedical Engineering.<\/em>,60(10), 2806-2814(2013)<\/li>\n<li>Landwehr, N., Mark H., and Eibe F., Logistic model trees, <em>Machine Learning.,<\/em>59(1-2),161-205(2005)<\/li>\n<li>Zhang Y., Zhang B., Lu W., Pham T.D., Zhou , Tanaka H., Oyama-Higa M.,\u00a0 Jiang X., Sun\u00a0 C., Kowalski J. and Jia X.,. Breast cancer classification from histological images with multiple features and random subspace classifier ensemble, In\u00a0AIP Conference Proceedings-American Institute of Physics.,1371(1),19(2011)<\/li>\n<li>Nguyen K., Barnes M., Srinivas C., and Chefd\u2019hotel C., Automatic glandular and tubule region segmentation in histological grading of breast cancer, InSPIE Medical Imaging\u00a0(94200G-94200G). International Society for Optics and Photonics,(2015)<\/li>\n<li>George Y.M., Elbagoury B.M., Zayed H.H. and Roushdy M.I., Breast fine needle tumor classification using neural networks,<em>IJCSI International Journal of Computer Science<\/em>,9(5), 247-56 (2012)<\/li>\n<li>Zuiderveld K., Contrast limited adaptive histogram equalization,\u00a0<em>Graphics gems IV<\/em>, Academic Press Professional, Inc.,474-485(1994)<\/li>\n<li>Sharma N., Ray A.K., Sharma S., Shukla K.K., Aggarwal L. and Pradhan S., Segmentation of medical images using simulated annealing based fuzzy C Means algorithm, <em>International Journal of Biomedical Engineering and Technology<\/em>.,2(3),260-278(2009)<\/li>\n<li>Schindelin J., Arganda-Carreras I., Frise E., Kaynig V., Longair M. and Pietzsch T., Fiji: an open-source platform for biological-image analysis, <em>Nat Methods.,<\/em>9,676\u201382(2012)<\/li>\n<li>Huang Z. K., and Chau K. W., A new image thresholding method based on Gaussian mixture model,\u00a0<em>Applied Mathematics and Computation<\/em>,205(2),899-907(2008)<\/li>\n<li>Friedland G., Jantz K. and Rojas R., Siox: Simple interactive object extraction in still images. In\u00a0Seventh IEEE International Symposium on Multimedia (ISM&#8217;05)\u00a0IEEE,(7)(2005)<\/li>\n<li>, Optimizing edge detectors for robust automatic threshold selection, <em>Graph. Models Image Proc.,<\/em>60: M.H.F.(1998)<\/li>\n<li>Arganda-Carreras I., Kaynig V., Schindelin , Cardona A. and Seung H.S., Trainable weka segmentation: a machine learning tool for microscopy image segmentation.,(2014)<\/li>\n<li>Martin D., Fowlkes C., Tal D. and Malik J., A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics, <em>In <\/em>Computer Vision. ICCV. Proceedings. Eighth IEEE International Conference on<em>.<\/em>2,416-423(2001)<\/li>\n<li>Unnikrishnan R., Pantofaru C., and Hebert M., Toward objective evaluation of image segmentation algorithms, <em>Pattern Analysis and Machine Intelligence, IEEE Transactions on,<\/em> 29,929-944(2007)<\/li>\n<li>Unnikrishnan R., Pantofaru C. and Hebert M., A measure for objective evaluation of image segmentation algorithms, In: Computer Vision and Pattern Recognition-Workshops, 2005. CVPR Workshops. IEEE Computer Society Conference on 25, 34-34,(2005)<\/li>\n<li>Zhao M., Chen L., Bian L., Zhang J., Yao C. and Zhang J., Feature Quantification and Abnormal Detection on Cervical Squamous Epithelial Cells,\u00a0<em>Computational and mathematical methods in medicine<\/em>.,(2015)<\/li>\n<li>, Saxena S., Shukla K.K. and Sharma S., Cellular Image Segmentation using Morphological Operators and Extraction of Features for Quantitative Measurement.\u00a0<em>Biosciences Biotechnology Research Asia.<\/em>,13(2),1101-1112,(2016)<\/li>\n<li>Spanhol F., Oliveira L., Petitjean C. and Heutte L., A Dataset for Breast Cancer Histopathological Image Classification.,(2015)<\/li>\n<li>Silva L.M., de S\u00e1 J.M. and Alexandre L.A., Data classification with multilayer perceptrons using a generalized error function,\u00a0<em>Neural Networks.<\/em>,21(9),1302-1310(2008)<\/li>\n<li>Mahesh V., Kandaswamy A., Vimal C. and Sathish B., ECG arrhythmia classification based on logistic model tree,\u00a0<em>Journal of Biomedical Science and Engineering<\/em>,2(06),405(2009)<\/li>\n<li>Breiman L., Random forests,\u00a0<em>Machine learning.,<\/em>45(1),5-32(2001)<\/li>\n<li>Platt J., Sequential minimal optimization: A fast algorithm for training support vector machines., (1998)<\/li>\n<li>Rodriguez J.J., Kuncheva L.I. and Alonso C.J., Rotation forest: A new classifier ensemble method.\u00a0<em>IEEE transactions on pattern analysis and machine intelligence<\/em>,28(10),1619-1630(2006).<\/li>\n<li>Timofte R., Tuytelaars T. and Van Gool L., Naive bayes image classification: beyond nearest neighbors. InAsian Conference on Computer Vision, Springer Berlin Heidelberg.689-703(2012)<\/li>\n<li>Cohen W.W., Fast effective rule induction. In Machine Learning: Proceedings of the 12th International Conference, Morgan Kaufmann,115-123(1995)<\/li>\n<li>Witten I. H. and Frank E., Generating accurate rule sets without global optimization. In Machine Learning: Proceedings of 15th International Conference,San Francisco: Morgan Kaufmann,(1998)<\/li>\n<li>Kira K., Rendell L.A., A Practical Approach to Feature Selection. In: Sleeman D, Edwards P (eds) Machine Learning Proceedings, Morgan Kaufmann, San Francisco (CA),249-256(1992)<\/li>\n<\/ol>\n","protected":false},"excerpt":{"rendered":"<p>Introduction The extensive use of computer aided diagnosis (CAD) these  [&#8230;]<\/p>\n","protected":false},"author":9,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[46],"tags":[],"class_list":["post-13448","post","type-post","status-publish","format-standard","hentry","category-vol10no1"],"_links":{"self":[{"href":"https:\/\/biomedpharmajournal.org\/staging\/wp-json\/wp\/v2\/posts\/13448","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/biomedpharmajournal.org\/staging\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/biomedpharmajournal.org\/staging\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/biomedpharmajournal.org\/staging\/wp-json\/wp\/v2\/users\/9"}],"replies":[{"embeddable":true,"href":"https:\/\/biomedpharmajournal.org\/staging\/wp-json\/wp\/v2\/comments?post=13448"}],"version-history":[{"count":5,"href":"https:\/\/biomedpharmajournal.org\/staging\/wp-json\/wp\/v2\/posts\/13448\/revisions"}],"predecessor-version":[{"id":32358,"href":"https:\/\/biomedpharmajournal.org\/staging\/wp-json\/wp\/v2\/posts\/13448\/revisions\/32358"}],"wp:attachment":[{"href":"https:\/\/biomedpharmajournal.org\/staging\/wp-json\/wp\/v2\/media?parent=13448"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/biomedpharmajournal.org\/staging\/wp-json\/wp\/v2\/categories?post=13448"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/biomedpharmajournal.org\/staging\/wp-json\/wp\/v2\/tags?post=13448"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}