Scopus İndeksli Yayınlar Koleksiyonu

Permanent URI for this collectionhttps://hdl.handle.net/20.500.12573/395

Browse

Search Results

Now showing 1 - 6 of 6
  • Conference Object
    Citation - Scopus: 1
    Words Speak Louder Than Actions: Decoding Emotions Through NLP
    (Institute of Electrical and Electronics Engineers Inc., 2024-10-26) Paksoy, Melda; Bakal, Gokhan
    Emotion detection in text remains a significant challenge in Natural Language Processing due to human emotions' complexity and subtle nuances. This paper presents multiple experimental models for emotion classification using an up-to-date dataset curated to address 13 emotions implied in Twitter posts. We evaluated various machine learning (ML) models, including Logistic Regression, Random Forest, SVM, and XGBoost, alongside deep learning (DL) architectures such as LSTM and CNN. Our results demonstrate the efficacy of deep learning models, particularly the CNN model by achieving an impressive F1 score of 0.99. This study contributes to emotion detection capabilities, paving the way for more nuanced and accurate sentiment analysis (SA) in various text analysis applications. © 2025 Elsevier B.V., All rights reserved.
  • Conference Object
    Citation - Scopus: 1
    Improving Salary Offer Processes With Classification Based Machine Learning Models
    (Institute of Electrical and Electronics Engineers Inc., 2024-09-21) Kaya, Rukiye; Saatci, Mehtap; Bakal, Gokhan; Bakal, Mehmet Gokhan
    In job applications, salary is major motivational factor for employees and making accurate salary prediction is crucial for both employers and employees. Utilizing advanced technologies can significantly enhance the accuracy and efficiency of salary prediction process. In this study, we explore Machine Learning (ML) methods to enhance salary prediction process. We evaluated seven classification models for predicting salary categories, with the Artificial Neural Network (ANN) model achieving the highest accuracy at 58.2% on the test dataset, followed by the K-Nearest Neighbors (KNN) model with an accuracy of 56.8%. Additionally, we employed ensemble models to further enhance prediction accuracy. Among these, the Majority Voting Classifier using Hard Voting achieved the highest accuracy at 59.3%, demonstrating the potential of ensemble techniques in refining salary predictions. The developed salary prediction tool estimates the most appropriate salary category for each candidate and help mitigate potential biases in manual salary assessments, hence enables a more objective and consistent compensation system. ∗CRITICAL: Do Not Use Symbols, Special Characters, or Math in Paper Title or Abstract, and do not cite other papers in the abstract. © 2024 Elsevier B.V., All rights reserved.
  • Conference Object
    Citation - Scopus: 1
    From Traditional to Deep: Evaluating Sentiment Analysis Models on a Large-Scale Tweet Dataset
    (Institute of Electrical and Electronics Engineers Inc., 2024-10-26) Mammadov, Alisahib; Bakal, Gokhan
    This study investigates the effectiveness of various machine learning (ML) and deep learning (DL) techniques for large-scale sentiment analysis on Twitter data. We leverage a publicly available dataset of one million tweets, annotated with four sentiment labels (positive, negative, uncertainty, and liti-gious), to train and evaluate a range of models. Our experiments demonstrate that traditional ML algorithms, particularly XG-Boost, achieve high performance, with the best F1 score reaching 95.81% using a combination of unigrams and bigrams. Among DL models, a hybrid CNN-BiGRU architecture yields the highest average F1 score of 95.42%. Our findings highlight the strengths of different approaches for sentiment analysis on Twitter data and emphasize the importance of data preprocessing and model selection for achieving optimal performance. © 2025 Elsevier B.V., All rights reserved.
  • Article
    Citation - Scopus: 10
    Building a Challenging Medical Dataset for Comparative Evaluation of Classifier Capabilities
    (Elsevier Ltd, 2024-08) Bozkurt, Berat; Coskun, Kerem; Bakal, Gokhan
    Since the 2000s, digitalization has been a crucial transformation in our lives. Nevertheless, digitalization brings a bulk of unstructured textual data to be processed, including articles, clinical records, web pages, and shared social media posts. As a critical analysis, the classification task classifies the given textual entities into correct categories. Categorizing documents from different domains is straightforward since the instances are unlikely to contain similar contexts. However, document classification in a single domain is more complicated due to sharing the same context. Thus, we aim to classify medical articles about four common cancer types (Leukemia, Non-Hodgkin Lymphoma, Bladder Cancer, and Thyroid Cancer) by constructing machine learning and deep learning models. We used 383,914 medical articles about four common cancer types collected by the PubMed API. To build classification models, we split the dataset into 70% as training, 20% as testing, and 10% as validation. We built widely used machine-learning (Logistic Regression, XGBoost, CatBoost, and Random Forest Classifiers) and modern deep-learning (convolutional neural networks - CNN, long short-term memory - LSTM, and gated recurrent unit - GRU) models. We computed the average classification performances (precision, recall, F-score) to evaluate the models over ten distinct dataset splits. The best-performing deep learning model(s) yielded a superior F1 score of 98%. However, traditional machine learning models also achieved reasonably high F1 scores, 95% for the worst-performing case. Ultimately, we constructed multiple models to classify articles, which compose a hard-to-classify dataset in the medical domain. © 2024 Elsevier B.V., All rights reserved.
  • Conference Object
    Citation - Scopus: 4
    A Transfer Learning Application on the Reliability of Psychological Drugs' Comments
    (Institute of Electrical and Electronics Engineers Inc., 2023-07-25) Sen, Tarik Uveys; Bakal, Gokhan
    As digitalization and the Internet stay emerging concepts by gaining popularity, the accuracy of personal reviews/opinions will be a critical issue. This circumstance also particularly applies to patients taking psychological drugs, where accurate information is crucial for other patients and medical professionals. In this study, we analyze drug reviews from drugs.com to determine the effectiveness of reviews for psychological drugs. Our dataset includes over 200,000 drug reviews, which we labeled as positive, negative, or neutral according to their rating scores. We apply machine learning (ML) models, including Logistic Regression, Recurrent Neural Network (RNN), and Long Short-Term Memory (LSTM) algorithms, to predict the sentiment class of each review. Our results demonstrate an F1-Weighted score of 85.3% for the LSTM model. However, by applying the transfer learning technique, we further improved the F1 score (nearly 3% increase) obtained by the LSTM model. Our findings proved that there is no contextual difference between the comments made by the patients suffering from psychological or other diseases. © 2023 Elsevier B.V., All rights reserved.
  • Conference Object
    Citation - Scopus: 3
    A Computational Drug Repositioning Effort Using Patients' Reviews Dataset
    (Institute of Electrical and Electronics Engineers Inc., 2023-07-25) Akkaya, Ali; Bakal, Gokhan
    The drug discovery process is one of the core motivations in both medical and, specifically, pharmaceutical disciplines. Due to the nature of the process, it requires an excessive amount of time, clinical experiments, and budget to cover each discovery phase. In this sense, computational drug discovery efforts can shorten the discovery process by providing plausible candidates since many of the attempts fail for several reasons, such as a lack of participants, financial problems, or ineffective results. In this study, the goal is to identify plausible candidate drugs for diseases. To do that, we utilize a personal experience of drugs dataset generated by patients. Beyond the user-generated comments, the users also give a rate between 1 and 10. Since we want to ensure the dataset quality, we first performed sentiment analysis experiments to prove that the reviews/comments are consistent with the given rating score. Then, only the review pairs having an effectiveness rate of 6 or more are selected as pre-filtered drug-disease pairs. We also build a knowledge graph using treatment-related biomedical relations using predications from Semantic Medline Database to identify drug similarities utilizing the Simrank similarity algorithm. As a result, we reported a list of plausible drugs as repurposing/repositioning candidates for further experiments. © 2023 Elsevier B.V., All rights reserved.