Convergence of Model-Based Temporal Difference Learning for Control

Publication date

2007

Authors

Hasselt, H. van
Wiering, M.A.

Editors

Advisors

Supervisors

DOI

Document Type

Article in proceedings
Open Access logo

License

Abstract

A theoretical analysis of Model-Based Temporal Difference Learning for Control is given, leading to a proof of convergence. This work differs from earlier work on the convergence of Temporal Difference Learning by proving convergence to the optimal value function. This means that not the values of the current policy are found, but instead the policy is updated in such a manner that ultimately the optimal policy is guaranteed to be reached.

Keywords

Citation