Efficient Continuous Control with Double Actors and Regularized Critics

Jiafei Lyu; Xiaoteng Ma; Jiangpeng Yan; Xiu Li

Efficient Continuous Control with Double Actors and Regularized Critics

Jiafei Lyu, Xiaoteng Ma, Jiangpeng Yan, Xiu Li

[AAAI-22] Main Track

Keywords
Poster Session 6 @ Blue 1, Poster Session 12 @ Blue 1, Oral Session 12 @ Blue 1, Poster Session 6, Poster Session 12, Oral Session 12

Download Paper

Enter the Virtual Venue

Abstract: How to obtain good value estimation is a critical problem in Reinforcement Learning (RL). Current value estimation methods in continuous control, such as DDPG and TD3, suffer from unnecessary over- or under- estimation. In this paper, we explore the potential of double actors, which has been neglected for a long time, for better value estimation in the continuous setting. First, we interestingly find that double actors improve the exploration ability of the agent. Next, we uncover the bias alleviation property of double actors in handling overestimation with single critic, and underestimation with double critics respectively. Finally, to mitigate the potentially pessimistic value estimate in double critics, we propose to regularize the critics under double actors architecture. Together, we present Double Actors Regularized Critics (DARC) algorithm. Extensive experiments on challenging continuous control benchmarks, MuJoCo and PyBullet, show that DARC significantly outperforms current baselines with higher average return and better sample efficiency.

Introduction Video

Sessions where this paper appears

Timezone

Poster Session 6

Sat, February 26 8:45 AM - 10:30 AM (+00:00)

Blue 1

Add to Calendar
Apple
Google
iCal File
Microsoft 365
Outlook.com
Yahoo

Poster Session 6
Poster Session 12

Mon, February 28 8:45 AM - 10:30 AM (+00:00)

Blue 1

Add to Calendar
Apple
Google
iCal File
Microsoft 365
Outlook.com
Yahoo

Poster Session 12
Oral Session 12

Mon, February 28 10:30 AM - 11:45 AM (+00:00)

Blue 1

Add to Calendar
Apple
Google
iCal File
Microsoft 365
Outlook.com
Yahoo

Oral Session 12