Journal Article

·2023

Dynamic Programming vs Q-learning for Feedback Motion Planning of Manipulators

Uğur Yıldıran YTU

Abstract

Reinforcement Learning (RL) based methods have became popular for control and motion planning of robots, recently. Unlike sampling based motion planners, optimal policies computed by them provide feedback motion plans which eliminates the need for re-computing (optimal) trajectories when a robot starts from a different initial configuration each time. In related studies, an optimal policy (actor) and the associated value function (critic) are usually calculated preforming training in a simulation environment. During training, RL allows learning by interactions with the environment in a physically realistic manner. However, in a simulation system, it is possible to make physically unimplementable moves. Thus, instead of RL, one can make use of Dynamic Programming approaches such as Value Iteration for computing optimal policies, which does not require an exploration component and known to have better convergence properties. In addition, dimension of a value function is smaller than that of a Q-fuction, thereby lessening the severity of the curse of dimensionality. Motivated by these facts, the aim of this paper is to employ Value Iteration algorithm for motion planning of robot manipulators and elaborate its effectiveness compared to a popular RL method, Q-learning.

Keywords

Reinforcement learning Bellman equation Curse of dimensionality Computer science Convergence (economics) Motion planning Dimension (graph theory) Dynamic programming Robot Motion (physics) Function (biology) Mathematical optimization Q-learning Function approximation Artificial intelligence Artificial neural network Mathematics Algorithm

Subject Areas

Robotic Path Planning Algorithms ·Computer Vision and Pattern Recognition ·Physical Sciences
Reinforcement Learning in Robotics ·Artificial Intelligence ·Physical Sciences
Robot Manipulation and Learning ·Control and Systems Engineering ·Physical Sciences

OpenAlex SDG Match

SDGs auto-classified by OpenAlex (score ≥ 0.4 shown).

Sustainable cities and communities 51%