Profile URL: https://hdl.handle.net/20.500.12573/2689
Job Title:Prof. Dr.
Email Address:zafer.aydin@agu.edu.tr
Main Affiliation:02. 04. Bilgisayar Mühendisliği
Status: Current Staff
ORCID:
0000-0003-3802-4211
0000-0003-3802-4211Scopus ID:
7003852510
7003852510YÖK Akademik: CDBF71256EF0830A
Google Scholar:
W1v7UfUAAAAJ
W1v7UfUAAAAJWeb of Science ID:
IRN-6351-2023
IRN-6351-2023Name Variants:
Aydin, Zafer Aydın, Zafer
72 results
Scholarly Output Search Results
Now showing 1 - 10 of 72
Conference Object Citation - WoS: 2Citation - Scopus: 4Data Mining Techniques in Direct Marketing on Imbalanced Data Using Tomek Link Combined With Random Under-Sampling(Assoc Computing Machinery, 2021-05-27) Yilmaz, Umit; Gezer, Cengiz; Aydin, Zafer; Gungor, V. CaGri; Yllmaz, Ümit; Aydln, ZaferDetermining the potential customers is very important in direct marketing. Data mining techniques are one of the most important methods for companies to determine potential customers. However, since the number of potential customers is very low compared to the number of non-potential customers, there is a class imbalance problem that significantly affects the performance of data mining techniques. In this paper, different combinations of basic and advanced resampling techniques such as Synthetic Minority Over-sampling Technique (SMOTE), Tomek Link, RUS, and ROS were evaluated to improve the performance of customer classification. Different feature selection techniques are used in order the decrease the number of non-informative features from the data such as Information Gain, Gain Ratio, Chi-squared, and Relief. Classification performance was compared and utilized using several data mining techniques, such as LightGBM, XGBoost, Gradient Boost, Random Forest, AdaBoost, ANN, Logistic Regression, Decision Trees, SVC, Bagging Classifier based on ROC AUC and sensitivity metrics. A combination of Tomek Link and Random Under-Sampling as a resampling technique and Chi-squared method as feature selection algorithm showed superior performance among the other combinations. Detailed performance evaluations demonstrated that with the proposed approach, LightGBM, which is a gradient boosting algorithm based on decision tree, gave the best results among the other classifiers with 0.947 sensitivity and 0.896 ROC AUC value.Research Project RNA İkincil Yapılarının Çok Boyutlu Gösterimi ve Pre-Mirna Tespiti Için Uygulamaları(TUBİTAK, 2021) Saçar Demirci, Müşerref Duygu; Demirci, Yilmaz MehmetMikroRNA'lar (miRNA'lar), transkripsiyon sonrası gen ekspresyonu düzenleyicileridir. Bir_x000D_ miRNA yüzlerce haberci RNA'yı (mRNA'lar) hedefleyebildiği gibi, bir mRNA farklı miRNA'lar_x000D_ tarafından hedeflenebilir, üstelik tek bir miRNA bir mRNA sekansında çeşitli bağlanma_x000D_ bölgelerine sahip olabilir. Bu nedenle miRNA'ları deneysel olarak araştırmak oldukça_x000D_ karmaşıktır. Bu tür zorlukları aşabilmek için makine öğrenimi (ML) sıklıkla kullanılmaktadır._x000D_ ML analizinin temel kısımları büyük ölçüde giriş verilerinin kalitesine ve verileri tanımlayan_x000D_ özelliklerin kapasitesine bağlıdır. Daha önce miRNA'lar için 1000'den fazla özellik önerilmişti._x000D_ Bu projede, RNA ikincil yapısını temsil eden yeni özellikler ve yüksek doğruluk değerleri_x000D_ sağlayan, dinamik, çok boyutlu grafik gösterimini tanımlamayı hedeflemiştik. Bu çalışmada,_x000D_ ML tabanlı miRNA tahmini için yeni ve kolayca güncellenebilir bir yaklaşım geliştirilmiştir._x000D_ Bilinen insan miRNA'larının ve sözde saç tokalarının random forest (RF), support vector_x000D_ machine (SVM) ve multilayer perceptron (MLP) gibi çeşitli sınıflandırıcılarla_x000D_ sınıflandırılmasıyla binlerce model oluşturulmuştur. Yöntem insan verilerine dayanarak_x000D_ oluşturulmuş olsa da en iyi model miRBase ve MirGeneDB gibi kamu veri tabanlarından_x000D_ insan olmayan saç tokaları üzerinde test edilmiş ve yüksek skorlar üretilmiştir. Ayrıca,_x000D_ yöntemin farklı veriler üzerindeki etkinliğini göstermek için ekspresyon farkları tahmini_x000D_ (differential expression prediction) analizinde de kullanılmıştır. Bu aşamada SARS-CoV-2_x000D_ enfeksiyonunun etkisini ölçen bir veri setinin analizinden elde edilen sonuçlar yayınlanmıştır.Research Project Zenginleştirilmiş Öznitelikler ve Makine Öğrenmesi Yöntemleriyle Protein Yerel Yapı Tahmini(TUBİTAK, 2017) Aydın, ZaferProjenin amacı proteinlerde bulunan ikincil yapı, dihedral açı ve çözücü erişilirlik gibi bir boyutlu yapısal özelliklerin başarılı olarak tahmin edilmesi ve bu tahminleri kullanarak parçacık seçimi yapan yeni bir yöntem geliştirilmesidir. Geliştirilen yöntemler sayesinde proteinlerin üç boyutlu yapısının daha doğru tahmin edilmesi, proteinlerin fonksiyonlarının daha iyi anlaşılması ve daha etkili ilaç tasarımı yapılması mümkün olacaktır. Bir boyutlu yapısal özelliklerin tahmini için yürütücünün daha önce geliştirdiği iki aşamalı hibrit sınıflandırma yöntemi kullanılmıştır. Bu yöntemde bulunan sınıflandırıcılar için dizi tabanlı profiller, yapısal profil matrisleri gibi çeşitli öznitelik vektörleri kullanılmıştır. İkinci aşamadaki sınıflandırıcı için destek vektör makinası, derin KSA, rastgele orman ve topluluk gibi çeşitli öğrenme yöntemleri eğitilmiş ve geliştirilen yöntemlerin tahmin başarı oranları standart veri kümelerinde incelenmiştir. Ayrıca bu aşamada derin otokodlayıcılar ve öznitelik seçme yaklaşımları ile boyut düşürme gerçekleştirilmiştir. Protein parçacık seçimi için verilen iki amino asit dizisi parçacığının yapısal olarak benzer olup olmadığının tahmin eden yöntemler geliştirilmiştir. Bunun için Rosetta programının parçacık veritabanında bulunan proteinlerden parçacık ikilileri örneklenmiş, bu ikililer BCScore yöntemi ile etiketlenmiş, eğitim ve test kümeleri oluşturulmuştur. Ayrıca farklı öznitelik kümeleri konsept hiyerarşi yaklaşımı ile kapsamlı olarak incelenmiş ve en başarılı sonucu veren öznitelik kombinasyonları tespit edilmiştir. Parçacık seçimi probleminde 3 ve 9 amino asitlik parçacıklar üzerinde çalışılmıştır ancak yöntemler diğer uzunluktaki parçacıklar için de kolaylıkla uygulanabilecektir. Projede geliştirilen yöntemler sayesinde ikincil yapı tahmin başarısı en zor tahmin kategorisinde %2.6 iyileşmiş, dihedral açı tahmin başarısı önemli oranda iyileşmiş, çözücü erişilirlik probleminde literatürdeki en başarılı yöntemler ile benzer bir seviye yakalanmıştır. Parçacık seçiminde ise verilen iki parçacığın yapılarının benzer olup olmadıkları 3-mer parçacıklar için %94 ve 9merler içinse %97 oranı ile tahmin edilmiştir. Yapılan çalışmaların neticesinde öznitelik vektörlerinin daha iyi tasarlanmasının ve farklı sınıflandırma yöntemlerinin birleştirilip optimize edilmesinin yapısal özellik tahmin başarısını önemli oranda iyileştirdiği sonucuna varılmıştır.Article Citation - WoS: 15Citation - Scopus: 22A Deep Learning Approach With Bayesian Optimization and Ensemble Classifiers for Detecting Denial of Service Attacks(Wiley, 2020-05-06) Gormez, Yasin; Aydin, Zafer; Karademir, Ramazan; Gungor, Vehbi C.Detecting malicious behavior is important for preventing security threats in a computer network. Denial of Service (DoS) is among the popular cyber attacks targeted at web sites of high-profile organizations and can potentially have high economic and time costs. In this paper, several machine learning methods including ensemble models and autoencoder-based deep learning classifiers are compared and tuned using Bayesian optimization. The autoencoder framework enables to extract new features by mapping the original input to a new space. The methods are trained and tested both for binary and multi-class classification on Digiturk and Labris datasets, which were introduced recently for detecting various types of DDoS attacks. The best performing methods are found to be ensembles though deep learning classifiers achieved comparable level of accuracy.Article Citation - WoS: 14Citation - Scopus: 14A Continuously Benchmarked and Crowdsourced Challenge for Rapid Development and Evaluation of Models to Predict COVID-19 Diagnosis and Hospitalization(Amer Medical Assoc, 2021-10-11) Yan, Yao; Schaffter, Thomas; Bergquist, Timothy; Yu, Thomas; Prosser, Justin; Aydin, Zafer; Mooney, SeanIMPORTANCE Machine learning could be used to predict the likelihood of diagnosis and severity of illness. Lack of COVID-19 patient data has hindered the data science community in developing models to aid in the response to the pandemic. OBJECTIVES To describe the rapid development and evaluation of clinical algorithms to predict COVID-19 diagnosis and hospitalization using patient data by citizen scientists, provide an unbiased assessment of model performance, and benchmark model performance on subgroups. DESIGN, SETTING, AND PARTICIPANTS This diagnostic and prognostic study operated a continuous, crowdsourced challenge using a model-to-data approach to securely enable the use of regularly updated COVID-19 patient data from the University of Washington by participants from May 6 to December 23, 2020. A postchallenge analysis was conducted from December 24, 2020, to April 7, 2021, to assess the generalizability of models on the cumulative data set as well as subgroups stratified by age, sex, race, and time of COVID-19 test. By December 23, 2020, this challenge engaged 482 participants from 90 teams and 7 countries. MAIN OUTCOMES AND MEASURES Machine learning algorithms used patient data and output a score that represented the probability of patients receiving a positive COVID-19 test result or being hospitalized within 21 days after receiving a positive COVID-19 test result. Algorithms were evaluated using area under the receiver operating characteristic curve (AUROC) and area under the precision recall curve (AUPRC) scores. Ensemble models aggregating models from the top challenge teams were developed and evaluated. RESULTS In the analysis using the cumulative data set, the best performance for COVID-19 diagnosis prediction was an AUROC of 0.776 (95% CI, 0.775-0.777) and an AUPRC of 0.297, and for hospitalization prediction, an AUROC of 0.796 (95% CI, 0.794-0.798) and an AUPRC of 0.188. Analysis on top models submitting to the challenge showed consistently better model performance on the female group than the male group. Among all age groups, the best performance was obtained for the 25- to 49-year age group, and the worst performance was obtained for the group aged 17 years or younger. CONCLUSIONS AND RELEVANCE In this diagnostic and prognostic study, models submitted by citizen scientists achieved high performance for the prediction of COVID-19 testing and hospitalization outcomes. Evaluation of challenge models on demographic subgroups and prospective data revealed performance discrepancies, providing insights into the potential bias and limitations in the models.Master Thesis Bilgisayar Ağlarında Anormal Durum Tespiti Yapan Öğrenme Yöntemlerinin Geliştirilmesi(Abdullah Gül Üniversitesi, 2018) MUKHANDI, HABIBU SHOMARI; Mukhandi, Habibu Shomari; Aydın, ZaferMakine öğrenmesi, verilerdeki bilginin bir bilgisayar ya da makina tarafından otomatik olarak öğrenilmesi ve karşılaşılan yeni durumlarda anlamlı bilgi ya da davranışların üretilmesini amaçlar. Bir çok uygulama alanı bulunan makine öğrenmesi daha önce hiç karşılaşılmamış olan sıradışı durumların tespit edilmesi için de kullanılmaktadır. Bilgisayar ağlarındaki siber saldırılar, kredi kartı dolandırıcılığı ve internet sitelerinin linklerine yapılan çok sayıda sahte tıklamalar dünya genelinde ekonomileri ciddi oranda zarara uğratabilecek niteliktedir. Bu tezde üç farklı anormal durum tespiti problemi üzerinde çalışılmıştır: bilgisayar ağlarında saldırı tespiti, kredi kartı dolandırıcılığı tespiti ve internet sitelerdeki linklere sahte tıklama tespiti. Anormal durum tespiti için geliştirilen ve optimize edilen modeller arasında rastgele orman, en yakın komşu, destek vektör makinası, logistic regresyon, karar ağacı, AdaBoost, çantalama ve yığınlama gibi sınıflandırma yöntemleri bulunmaktadır. Yöntemlerin hiper-parametreleri eğitim kümelerinde yapılan çapraz doğrulama deneyleri ile optimize edilmiştir. Bir sonraki aşamada optimum hiper-parametre konfigürasyonları kullanılarak eğitilen modeler ile test verilerinde tahmin sonuçları hesaplanmıştır. Bu deneyler neticesinde genel doğruluk oranı ve F-measure skorlarında yüksek başarı elde edilmiştir. Geliştirilen yöntemler arasında en başarılı sonuçlar topluluk modelleri ile elde edilmiştir.Conference Object Citation - Scopus: 3Ceph-Based Storage Server Application(Institute of Electrical and Electronics Engineers Inc., 2018-03) Azgınoglu, Nuh; Eren, Mehmet Akif; Celik, Mete; Aydin, ZaferCeph is a scalable and high performance distributed file system. In this study, a Ceph-based storage server was implemented and used actively. This storage system has been used as a disk of 40 virtual servers in 4 different Proxmox servers. Performance evaluation of the system has been conducted on virtual servers that holds Windows and Linux based operating systems. © 2018 Elsevier B.V., All rights reserved.Article Probing Adversarial Robustness of Protein Language Models: A Reproducible Case Study of ESM-2 Under Substitution-Based Attacks(AIUB Office of Research and Publication, 2026) Aydin, Zafer; Niloy, Md. Robiul Islam; Moazzam, Md. GolamConference Object Citation - WoS: 3Citation - Scopus: 3Template Scoring Methods for Protein Torsion Angle Prediction(Springer-Verlag Berlin, 2015) Aydin, Zafer; Baker, David; Noble, William StaffordPrediction of backbone torsion angles provides important constraints about the 3D structure of a protein and is receiving a growing interest in the structure prediction community. In this paper, we introduce a three-stage machine learning classifier to predict the 7-state torsion angles of a protein. The first two stages employ dynamic Bayesian and neural networks to produce an ab-initio prediction of torsion angle states starting from sequence profiles. The third stage is a committee classifier, which combines the ab-initio prediction with a structural frequency profile derived from templates obtained by HHsearch. We develop several structural profile models and obtain significant improvements over the Laplacian scoring technique through: (1) scaling templates by integer powers of sequence identity score, (2) incorporating other alignment scores as multiplicative factors (3) adjusting or optimizing parameters of the profile models with respect to the similarity interval of the target. We also demonstrate that the torsion angle prediction accuracy improves at all levels of target-template similarity even when templates are distant from the target. The improvement is at significantly higher rates as template structures gradually get closer to target.Conference Object Citation - Scopus: 5Development of Knowledge Based Response Correction for a Reconfigurable N-Shaped Microstrip Antenna Design(Institute of Electrical and Electronics Engineers Inc., 2015-08) Aoad, Ashrf; Simsek, Murat; Aydin, ZaferThis study presents the use of prior knowledge of inverse artificial neural network (ANN) to model and optimize a reconfigurable N-shaped microstrip antenna. Three accurate prior knowledge inverse ANNs with large amount training data are proposed where the frequency information is incorporated into the structure of ANN. The complexity of the input/output relationship is reduced by using prior knowledge. Three separate methods of incorporating knowledge in the second step of the training process with a multilayer perceptron (MLP) in the first step are demonstrated and their results are compared to EM simulation. © 2023 Elsevier B.V., All rights reserved.
