Convergence of Model-Based Temporal Difference Learning for Control
Publication date
2007
Authors
Hasselt, H. van
Wiering, M.A.
Editors
Advisors
Supervisors
DOI
Document Type
Article in proceedings
Metadata
Show full item recordCollections
License
Abstract
A theoretical analysis of Model-Based Temporal
Difference Learning for Control is given, leading to a proof of
convergence. This work differs from earlier work on the convergence
of Temporal Difference Learning by proving convergence
to the optimal value function. This means that not the values of
the current policy are found, but instead the policy is updated in
such a manner that ultimately the optimal policy is guaranteed
to be reached.