Potential-based reward shaping using state–space segmentation for efficiency in reinforcement learning

Publication date

2024-08

Authors

Bal, Melis İlayda
Aydin, Hüseyin
İyigün, Cem
Polat, Faruk

Editors

Advisors

Supervisors

Document Type

Article
Open Access logo

License

taverne

Abstract

Reinforcement Learning (RL) algorithms encounter slow learning in environments with sparse explicit reward structures due to the limited feedback available on the agent's behavior. This problem is exacerbated particularly in complex tasks with large state and action spaces. To address this inefficiency, in this paper, we propose a novel approach based on potential-based reward-shaping using state–space segmentation to decompose the task and to provide more frequent feedback to the agent. Our approach involves extracting state–space segments by formulating the problem as a minimum cut problem on a transition graph, constructed using the agent's experiences during interactions with the environment via the Extended Segmented Q-Cut algorithm. Subsequently, these segments are leveraged in the agent's learning process through potential-based reward shaping. Our experimentation on benchmark problem domains with sparse rewards demonstrated that our proposed method effectively accelerates the agent's learning without compromising computation time while upholding the policy invariance principle.

Keywords

Potential-based reward shaping, Reinforcement learning, Reward shaping, Sparse rewards, State–space segmentation, Taverne, Software, Hardware and Architecture, Computer Networks and Communications

Citation

Bal, M İ, Aydın, H, İyigün, C & Polat, F 2024, 'Potential-based reward shaping using state–space segmentation for efficiency in reinforcement learning', Future Generation Computer Systems, vol. 157, pp. 469-484. https://doi.org/10.1016/j.future.2024.03.057