Efficient exploration through active learning for value function approximation in reinforcement learning.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 20080026.
- Also identified by DOI 10.1016/j.neunet.2009.12.010.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Appropriately designing sampling policies is highly important for obtaining better control policies in reinforcement learning. In this paper, we first show that the least-squares policy iteration (LSPI) framework allows us to employ statistical active learning methods for linear regression. Then we propose a design method of good sampling policies for efficient exploration, which is particularly useful when the sampling cost of immediate rewards is high. The effectiveness of the proposed method, which we call active policy iteration (API), is demonstrated through simulations with a batting robot.
Medical subject headings
- Artificial Intelligence
- Exploratory Behavior
- Reinforcement, Psychology