Scopus İndeksli Yayınlar Koleksiyonu

Permanent URI for this collectionhttps://hdl.handle.net/20.500.12573/395

Browse

Search Results

Now showing 1 - 10 of 11
  • Article
    Citation - Scopus: 3
    Prediction of Colorectal Cancer Based on Taxonomic Levels of Microorganisms and Discovery of Taxonomic Biomarkers Using the Grouping-Scoring (G-S-M) Approach
    (Elsevier Ltd, 2025-03) Bakir-Güngör, Burcu; Temiz, Mustafa; Canakcimaksutoglu, Beyza; Yousef, Malik
    Colorectal cancer (CRC) is one of the most prevalent forms of cancer globally. The human gut microbiome plays an important role in the development of CRC and serves as a biomarker for early detection and treatment. This research effort focuses on the identification of potential taxonomic biomarkers of CRC using a grouping-based feature selection method. Additionally, this study investigates the effect of incorporating biological domain knowledge into the feature selection process while identifying CRC-associated microorganisms. Conventional feature selection techniques often fail to leverage existing biological knowledge during metagenomic data analysis. To address this gap, we propose taxonomy-based Grouping Scoring Modeling (G-S-M) method that integrates biological domain knowledge into feature grouping and selection. In this study, using metagenomic data related to CRC, classification is performed at three taxonomic levels (genus, family and order). The MetaPhlAn tool is employed to determine the relative abundance values of species in each sample. Comparative performance analyses involve six feature selection methods and four classification algorithms. When experimented on two CRC associated metagenomics datasets, the highest performance metric, yielding an AUC of 0.90, is observed at the genus taxonomic level. At this level, 7 out of top 10 groups (Parvimonas, Peptostreptococcus, Fusobacterium, Gemella, Streptococcus, Porphyromonas and Solobacterium) were commonly identified for both datasets. Moreover, the identified microorganisms at genus, family, and order levels are thoroughly discussed via refering to CRC-related metagenomic literature. This study not only contributes to our understanding of CRC development, but also highlights the applicability of taxonomy-based G-S-M method in tackling various diseases. © 2025 Elsevier B.V., All rights reserved.
  • Conference Object
    Citation - Scopus: 5
    Identifying Taxonomic Biomarkers of Colorectal Cancer in Human Intestinal Microbiota Using Multiple Feature Selection Methods
    (Institute of Electrical and Electronics Engineers Inc., 2022-09-07) Jabeer, Amhar; Kocak, Aysegul; Akkaş, Huseyin; Yenisert, Ferhan; Nalbantoĝlu, Özkan Ufuk; Yousef, Malik; Bakir-Güngör, Burcu; Bakir Gungor, Burcu
    A variety of bacterial species called gut microbiota work together to maintain a steady intestinal environment. The gastrointestinal tract contains tremendous amount of different species including archaea, bacteria, fungi, and viruses. While these organisms are crucial immune system stabilizers, the dysbiosis of the intestinal flora has been related to gastrointestinal disorders including Colorectal cancer (CRC), intestinal cancer, irritable bowel syndrome and inflammatory bowel disease. In the last decade, next-generation sequencing (NGS) methods have accelerated the identification of human gut flora. CRC is a deathly condition that has been on the rise in the last century, affecting half a million people each year. Since early CRC diagnosis is critical for an effective treatment, there is an immediate requirement for a classification system that can expedite CRC diagnosis. In this study, via analyzing the available metagenomics data on CRC, we aim to facilitate the CRC diagnosis via finding biomarkers linked with CRC, and via building a classification model. We have obtained the metagenomic sequencing data of the healthy individuals and CRC patients from a metagenome-wide association analysis and we have classified this data according to the disease stages. Conditional Mutual Information Maximization (CMIM), Fast Correlation Based Filter (FCBF), Extreme Gradient Boosting (XGBoost), min redundancy max relevance (mRMR), Information Gain (IG) and Select K Best (SKB) feature selection algorithms were utilized to cope with the complexity of the features. We observed that the SKB, IG, and XGBoost techniques made significant contributions to decrease the microbiota in use for CRC diagnosis, thereby reducing cost and time. We realized that our Random Forest classifier outperformed Adaboost, Support Vector Machine, Decision Tree, Logitboost and stacking ensemble classifiers in terms of CRC classification performance. Our results reiterated some known and some potential microbiome associated mechanisms in CRC, which could aid the design of new diagnostics based on the microbiome. © 2022 Elsevier B.V., All rights reserved.
  • Conference Object
    Citation - WoS: 23
    Citation - Scopus: 52
    Evaluation of Classification Algorithms, Linear Discriminant Analysis and a New Hybrid Feature Selection Methodology for the Diagnosis of Coronary Artery Disease
    (Institute of Electrical and Electronics Engineers Inc., 2018-12) Kolukisa, Burak; Hacilar, Hilal; Göy, Gökhan; Kus, Mustafa; Bakir-Güngör, Burcu; Aral, Atilla; Güngör, Vehbi Çağrı
    According to the World Health Organization (WHO), 31% of the world's total deaths in 2016 (17.9 million) was due to cardiovascular diseases (CVD). With the development of information technologies, it has become possible to predict whether people have heart diseases or not by checking certain physical and biochemical values at a lower cost. In this study, we have evalated a set of different classification algorithms, linear discriminant analysis and proposed a new hybrid feature selection methodology for the diagnosis of coronary heart diseases (CHD). Throughout this research effort, using three publicly available Heart Disease diagnosis datasets (UCI Machine Learning Repository), we have conducted comparative performance evaluations in terms of accuracy, sensitivity, specificity, F-measure, AUC and running time. © 2023 Elsevier B.V., All rights reserved.
  • Article
    Citation - WoS: 41
    Citation - Scopus: 69
    Ensemble Feature Selection and Classification Methods for Machine Learning-Based Coronary Artery Disease Diagnosis
    (Elsevier, 2023-03) Kolukisa, Burak; Bakir-Gungor, Burcu
    Coronary artery disease (CAD) is a condition in which the heart is not fed sufficiently as a result of the accumulation of fatty matter. As reported by the World Health Organization, around 32% of the total deaths in the world are caused by CAD, and it is estimated that approximately 23.6 million people will die from this disease in 2030. CAD develops over time, and the diagnosis of this disease is difficult until a blockage or a heart attack occurs. In order to bypass the side effects and high costs of the current methods, researchers have proposed to diagnose CADs with computer-aided systems, which analyze some physical and biochemical values at a lower cost. In this study, for the CAD diagnosis, (i) seven different computational feature selection (FS) methods, one domain knowledge-based FS method, and different classification algorithms have been evaluated; (ii) an exhaustive ensemble FS method and a probabilistic ensemble FS method have been proposed. The proposed approach is tested on three publicly available CAD data sets using six different classification algorithms and four different variants of voting algorithms. The performance metrics have been comparatively evaluated with numerous combinations of classifiers and FS methods. The multi-layer perceptron classifier obtained satisfactory results on three data sets. Performance evaluations show that the proposed approach resulted in 91.78%, 85.55%, and 85.47% accuracy for the Z-Alizadeh Sani, Statlog, and Cleveland data sets, respectively.
  • Article
    Citation - WoS: 7
    Citation - Scopus: 5
    Effect of Interpolation on Specular Reflections in Texture-Based Automatic Colonic Polyp Detection
    (Wiley, 2020-06-26) Kacmaz, Rukiye Nur; Yilmaz, Bulent; Aydin, Zafer
    Reflections of LED light cause unwanted noise effects called specular reflection (SR) on colonoscopic images. The aim of this study was to seek answers to the following two questions. (a) How are the texture features used in automatic detection of polyps affected by the interpolation on specular reflections? (b) If they are affected does it really affect the classification performance? In order to answer these questions, we used 610 colonoscopy images, and divided each image into tiles whose sizes were 32-by-32 pixels. From these tiles, we selected the ones without any specular reflection. We added different shape and size specular reflections cropped from real images onto the reflection-free tiles. We then used the nearest neighbors, bilinear and bicubic interpolation techniques on the tiles on which SRs were added. On these tiles we extracted 116 texture features using 3 second-order approaches, and 4 first-order statistics. First, we used paired samplettest. Second, we performed automatic classification of polyps and background using random forest and k nearest neighbors (k-NN) approaches using the texture features for different combinations of specular reflections added on the tiles from the polyp or background. The results showed that depending on the size of specular reflection, interpolation can cause a significant difference between the texture features that were coming from reflection-free tiles and the same tiles on which interpolation was performed. In addition, we note that bicubic interpolation may be preferred to eliminate specular reflection when texture features are used for background and polyp discrimination.
  • Article
    Citation - Scopus: 10
    Building a Challenging Medical Dataset for Comparative Evaluation of Classifier Capabilities
    (Elsevier Ltd, 2024-08) Bozkurt, Berat; Coskun, Kerem; Bakal, Gokhan
    Since the 2000s, digitalization has been a crucial transformation in our lives. Nevertheless, digitalization brings a bulk of unstructured textual data to be processed, including articles, clinical records, web pages, and shared social media posts. As a critical analysis, the classification task classifies the given textual entities into correct categories. Categorizing documents from different domains is straightforward since the instances are unlikely to contain similar contexts. However, document classification in a single domain is more complicated due to sharing the same context. Thus, we aim to classify medical articles about four common cancer types (Leukemia, Non-Hodgkin Lymphoma, Bladder Cancer, and Thyroid Cancer) by constructing machine learning and deep learning models. We used 383,914 medical articles about four common cancer types collected by the PubMed API. To build classification models, we split the dataset into 70% as training, 20% as testing, and 10% as validation. We built widely used machine-learning (Logistic Regression, XGBoost, CatBoost, and Random Forest Classifiers) and modern deep-learning (convolutional neural networks - CNN, long short-term memory - LSTM, and gated recurrent unit - GRU) models. We computed the average classification performances (precision, recall, F-score) to evaluate the models over ten distinct dataset splits. The best-performing deep learning model(s) yielded a superior F1 score of 98%. However, traditional machine learning models also achieved reasonably high F1 scores, 95% for the worst-performing case. Ultimately, we constructed multiple models to classify articles, which compose a hard-to-classify dataset in the medical domain. © 2024 Elsevier B.V., All rights reserved.
  • Article
    Citation - WoS: 7
    Citation - Scopus: 8
    Autonomic Workload Performance Tuning in Large-Scale Data Repositories
    (Springer London Ltd, 2018-09-04) Raza, Basit; Sher, Asma; Afzal, Sana; Malik, Ahmad Kamran; Anjum, Adeel; Kumar, Yogan Jaya; Faheem, Muhammad
    The workload in large-scale data repositories involves concurrent users and contains homogenous and heterogeneous data. The large volume of data, dynamic behavior and versatility of large-scale data repositories is not easy to be managed by humans. This requires computational power for managing the load of current servers. Autonomic technology can support predicting the workload type; decision support system or online transaction processing can help servers to autonomously adapt to the workloads. The intelligent system could be designed by knowing the type of workload in advance and predict the performance of workload that could autonomically adapt the changing behavior of workload. Workload management involves effectively monitoring and controlling the workflow of queries in large-scale data repositories. This work presents a taxonomy through systematic analysis of workload management in large-scale data repositories with respect to autonomic computing (AC) including database management systems and data warehouses. The state-of-the-art practices in large-scale data repositories are reviewed with respect to AC for characterization, performance prediction and adaptation of workload. Current issues are highlighted at the end with future directions.
  • Conference Object
    Citation - Scopus: 3
    A Hybrid Adaptive Neuro-Fuzzy Inference System (Anfis) Approach for Professional Bloggers Classification
    (IEEE, 2019-11) Asim, Yousra; Raza, Basit; Malik, Ahmad Kamran; Shahid, Ahmad R.; Faheem, Muhammad; Kumar, Yogan Jaya
    Despite their small numbers, some users of the online social networks demonstrate the ability to influence others. Bloggers are one of such kind of users that through their ideas and opinions on different topics, influence other users. Their identification may be beneficial for several purposes, such as online marketing for products. Much effort has been expanded towards finding the impact of such bloggers within the blogging community. We have expanded on their work by identifying influential bloggers using labeled data. We have improved upon the accuracy of the classification of professional and non-professional bloggers. We have made use of Adaptive Neuro-Fuzzy Inference System (ANFIS), and the Fuzzy Inference System (FIS) models. Their performance has been gauged and compared with the existing techniques and approaches, such as an Artificial Neural Network (ANN), Alternating Decision Tree (ADTree) algorithm, and Classification Based on Associations (CBA) algorithm. Adaptive techniques (ANFIS and ANN) are found better than the aforementioned rule-based classifiers. The FIS model outperformed the CBA algorithm, but showed similar performance to the ADTree algorithm. Our proposed ANFIS model showed improved results in terms of performance measures with 93% accuracy for blogger classification.
  • Conference Object
    Citation - Scopus: 3
    A Hybrid Adaptive Neuro-Fuzzy Inference System (ANFIS) Approach for Professional Bloggers Classification
    (Institute of Electrical and Electronics Engineers Inc., 2019-11) Asim, Yousra; Raza, Basit; Malik, Ahmad Kamran Kamran; Shahid, Ahmad Raza; Faheem, Muhammed Yasir; Kumar, Y. J.
    Despite their small numbers, some users of the online social networks demonstrate the ability to influence others. Bloggers are one of such kind of users that through their ideas and opinions on different topics, influence other users. Their identification may be beneficial for several purposes, such as online marketing for products. Much effort has been expanded towards finding the impact of such bloggers within the blogging community. We have expanded on their work by identifying influential bloggers using labeled data. We have improved upon the accuracy of the classification of professional and nonprofessional bloggers. We have made use of Adaptive Neuro-Fuzzy Inference System (ANFIS), and the Fuzzy Inference System (FIS) models. Their performance has been gauged and compared with the existing techniques and approaches, such as an Artificial Neural Network (ANN), Alternating Decision Tree (ADTree) algorithm, and Classification Based on Associations (CBA) algorithm. Adaptive techniques (ANFIS and ANN) are found better than the aforementioned rule-based classifiers. The FIS model outperformed the CBA algorithm, but showed similar performance to the ADTree algorithm. Our proposed ANFIS model showed improved results in terms of performance measures with 93% accuracy for blogger classification. © 2020 Elsevier B.V., All rights reserved.
  • Conference Object
    Population Specific Classification of Colorectal Cancer With Meta-Analysis of Metagenomic Data
    (Institute of Electrical and Electronics Engineers Inc., 2023-10-11) Temiz, Mustafa; Yousef, Malik; Bakir-Güngör, Burcu
    Advances in next-generation sequencing and '-omics' technologies makes it possible to characterize the human gut microbiome. While some of these microorganisms are important regulators of our immune system, modulation of the microbiota leads to a variety of diseases. Colorectal cancer (CRC), the third most common cancer worldwide, is caused by genetic mutations, environmental conditions, and abnormalities in the gut microbiota. Using various machine learning methods and meta-analysis techniques, this study aims to build a classification model that can help in CRC diagnosis by analyzing metagenomic datasets of different populations obtained at the species level. Using 8 different countries and 9 different metagenomic datasets, 3 different meta-analyzes are performed: within-population, cross-population, and one population is selected for testing and the rest is used as a training dataset (LODO). For CRC classification, 4 different classification algorithms (Random Forest (RF), Logitboost, Adaboost, and Decision Tree (DT)) are used. The best performance among these methods was obtained with the Random Forest algorithm with an AUC of 0.98 by using JP for the training data set and JPN populations for the test data set in the cross-population performance evaluation. © 2023 Elsevier B.V., All rights reserved.