Modified Actor-Critics

2 Jul 2019 Erinc Merdivan Sten Hanke Matthieu Geist

Recent successful deep reinforcement learning algorithms, such as Trust Region Policy Optimization (TRPO) or Proximal Policy Optimization (PPO), are fundamentally variations of conservative policy iteration (CPI). These algorithms iterate policy evaluation followed by a softened policy improvement step... (read more)

PDF Abstract
No code implementations yet. Submit your code now

Tasks


Results from the Paper


  Submit results from this paper to get state-of-the-art GitHub badges and help the community compare results to other papers.

Methods used in the Paper


METHOD TYPE
Experience Replay
Replay Memory
Entropy Regularization
Regularization
Dense Connections
Feedforward Networks
ReLU
Activation Functions
Adam
Stochastic Optimization
Soft Actor Critic
Policy Gradient Methods
PPO
Policy Gradient Methods