Scopus İndeksli Yayınlar Koleksiyonu

Permanent URI for this collectionhttps://hdl.handle.net/20.500.12573/395

Browse

Search Results

Now showing 1 - 7 of 7
  • Conference Object
    AI Explainability for Adaptive Mmwave Beam Configuration in Dynamic Vehicular Environments
    (Institute of Electrical and Electronics Engineers Inc., 2026) Foh, Chuan Heng; Kose, Abdulkadir; Yigit, Ugur; Akbas, Ayhan; Shojafar, Mohammad
  • Conference Object
    Citation - Scopus: 1
    Traffic Light Management Systems Using Reinforcement Learning
    (Institute of Electrical and Electronics Engineers Inc., 2022-09-07) Can, Sultan Kubra; Thahir, Adam Rizvi; Cos¸kun, Mustafa; Güngör, Vehbi Çağrı; Coskun, Mustafa
    While reducing traffic congestion and decrease the number of traffic accidents in the intersections, most of the traffic light management approaches cannot adapt well to fast changing traffic dynamics and growing demands of the intersections with modern world developments. To overcome this problem, adaptive traffic controllers are developed, and detectors and sensors are added to systems to enable adoption and dynamism. Recently, reinforcement learning has shown its capability to learn the dynamics of complex environments, such as urban traffic. Although it was studied in single junction systems, one of the problems was the lack of consistency with how the real world system works. Most of the systems assume that the environment is fully observable or actions would be freely executed using simulators. This study aims to merge usefulness of reinforcement learning methods with real-world traffic constraints. Comparative performance evaluations show that the reinforcement learning algorithm (Advantage Actor-Critic (A2C)) converges well while staying stable under changing traffic dynamics. © 2022 Elsevier B.V., All rights reserved.
  • Article
    Citation - WoS: 24
    Citation - Scopus: 30
    Optimal Control of Microgrids With Multi-Stage Mixed-Integer Nonlinear Programming Guided Q-Learning Algorithm
    (State Grid Electric Power Research inst, 2020) Yoldas, Yeliz; Goren, Selcuk; Onen, Ahmet
    This paper proposes an energy management system (EMS) for the real-time operation of a pilot stochastic and dynamic microgrid on a university campus in Malta consisting of a diesel generator, photovoltaic panels, and batteries. The objective is to minimize the total daily operation costs, which include the degradation cost of batteries, the cost of energy bought from the main grid, the fuel cost of the diesel generator, and the emission cost. The optimization problem is modeled as a finite Markov decision process (MDP) by combining network and technical constraints, and Q-learning algorithm is adopted to solve the sequential decision subproblems. The proposed algorithm decomposes a multi-stage mixed-integer nonlinear programming (MINLP) problem into a series of single-stage problems so that each subproblem can be solved by using Bellman's equation. To prove the effectiveness of the proposed algorithm, three case studies are taken into consideration: (1) minimizing the daily energy cost; (2) minimizing the emission cost; (3) minimizing the daily energy cost and emission cost simultaneously. Moreover, each case is operated under different battery operation conditions to investigate the battery lifetime. Finally, performance comparisons are carried out with a conventional Q-learning algorithm.
  • Article
    Citation - WoS: 1
    Citation - Scopus: 1
    Intelligent Traffic Light Systems Using Edge Flow Predictions
    (Elsevier, 2024-01) Thahir, Adam Rizvi; Coskun, Mustafa; Kilic, Sultan Kubra; Gungor, Vehbi Cagri
    In this paper, we propose a novel graph-based semi-supervised learning approach for traffic light management in multiple intersections. Specifically, the basic premise behind our paper is that if we know some of the occupied roads and predict which roads will be congested, we can dynamically change traffic lights at the intersections that are connected to the roads anticipated to be congested. Comparative performance evaluations show that the proposed approach can produce comparable average vehicle waiting time and reduce the training/learning time of learning adequate traffic light configurations for all intersections within a few seconds, while a deep learning-based approach can be trained in a few days for learning similar light configurations.
  • Conference Object
    Citation - Scopus: 8
    Generating Emergency Evacuation Route Directions Based on Crowd Simulations With Reinforcement Learning
    (Institute of Electrical and Electronics Engineers Inc., 2022-09-07) Unal, Ahmet Emin; Gezer, Cengiz; Kuleli Pak, Burcu Kuleli; Güngör, Vehbi Çağrı; Pak, Burcu Kuleli
    In an emergency, it is vital to evacuate individuals from the dangerous environments. Emergency evacuation plan-ning ensures that the evacuation is safe and optimal in terms of evacuation time for all of the people in evacuation. To this end, the computer-enabled evacuation simulation systems are used to generate optimal routes for the evacuees. In this paper, a dynamic emergency evacuation route generator has been proposed based on indoor plans of the building and the locations of the evacuees. To generate the optimal routes in real-time, a reinforcement learning algorithm (proximal policy optimization) is presented. Comparative performance results show that the proposed model is successful for evacuating the individuals from the building in different scenarios. © 2022 Elsevier B.V., All rights reserved.
  • Conference Object
    Evaluating the Impact of Sentiment Analysis on Deep Reinforcement Learning-Based Trading Strategies
    (Institute of Electrical and Electronics Engineers Inc., 2024-10-26) Etcil, Mustafa; Kolukisa, Burak; Bakir-Güngör, Burcu
    Portfolio optimization is a form of investment management that aims to maximize returns while minimizing risks. However, the inherent complexity and unpredictability of financial markets pose a challenge. Recent advancements in machine learning, particularly in deep reinforcement learning (DRL), offer promising solutions by enabling dynamic and adaptive trading strategies. This paper presents a comprehensive evaluation of three actor-critic-based DRL algorithms-Advantage Actor-Critic (A2C), Deep Deterministic Policy Gradient (DDPG), and Proximal Policy Optimization (PPO)-applied to portfolio optimization. These strategies were implemented in both sentiment-aware and non-sentiment-aware versions, allowing for a direct comparison of their performance. The sentiment-aware models incorporated sentiment analysis using FinBERT and knowledge graphs to measure market sentiment from financial news, while the non-sentiment-aware models relied solely on stock prices and technical indicators. Our comparative study demonstrates that incorporating sentiment analysis resulted in consistently superior risk-adjusted returns and portfolio resilience during market fluctuations compared to non-sentiment-aware strategies. © 2025 Elsevier B.V., All rights reserved.
  • Article
    Citation - WoS: 6
    Citation - Scopus: 7
    A Reinforcement Learning-Based Demand Response Strategy Designed From the Aggregator's Perspective
    (Frontiers Media S.A., 2022-09-15) Oh, Seongmun; Jung, Jaesung; Onen, Ahmet; Lee, Chul-Ho
    The demand response (DR) program is a promising way to increase the ability to balance both supply and demand, optimizing the economic efficiency of the overall system. This study focuses on the DR participation strategy in terms of aggregators who offer appropriate DR programs to customers with flexible loads. DR aggregators engage in the electricity market according to customer behavior and must make decisions that increase the profits of both DR aggregators and customers. Customers use the DR program model, which sends its demand reduction capabilities to a DR aggregator that bids aggregate demand reduction to the electricity market. DR aggregators not only determine the optimal rate of incentives to present to the customers but can also serve customers and formulate an optimal energy storage system (ESS) operation to reduce their demands. This study formalized the problem as a Markov decision process (MDP) and used the reinforcement learning (RL) framework. In the RL framework, the DR aggregator and each customer are allocated to each agent, and the agents interact with the environment and are trained to make an optimal decision. The proposed method was validated using actual industrial and commercial customer demand profiles and market price profiles in South Korea. Simulation results demonstrated that the proposed method could optimize decisions from the perspective of the DR aggregator.