TASK	DATASET	MODEL	METRIC NAME	METRIC VALUE	GLOBAL RANK
Atari Games	Atari 2600 Bowling	RUDDER	Score	179	# 9
Atari Games	Atari 2600 Venture	RUDDER	Score	1350	# 14
Atari Games	Atari 2600 Yars Revenge	RUDDER	Score	60577	# 12

Badge	Markdown
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/rudder-return-decomposition-for-delayed/atari-games-on-atari-2600-bowling)](https://paperswithcode.com/sota/atari-games-on-atari-2600-bowling?p=rudder-return-decomposition-for-delayed)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/rudder-return-decomposition-for-delayed/atari-games-on-atari-2600-yars-revenge)](https://paperswithcode.com/sota/atari-games-on-atari-2600-yars-revenge?p=rudder-return-decomposition-for-delayed)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/rudder-return-decomposition-for-delayed/atari-games-on-atari-2600-venture)](https://paperswithcode.com/sota/atari-games-on-atari-2600-venture?p=rudder-return-decomposition-for-delayed)`

RUDDER: Return Decomposition for Delayed Rewards

NeurIPS 2019 · Jose A. Arjona-Medina, Michael Gillhofer, Michael Widrich, Thomas Unterthiner, Johannes Brandstetter, Sepp Hochreiter ·

We propose RUDDER, a novel reinforcement learning approach for delayed rewards in finite Markov decision processes (MDPs). In MDPs the Q-values are equal to the expected immediate reward plus the expected future rewards. The latter are related to bias problems in temporal difference (TD) learning and to high variance problems in Monte Carlo (MC) learning. Both problems are even more severe when rewards are delayed. RUDDER aims at making the expected future rewards zero, which simplifies Q-value estimation to computing the mean of the immediate reward. We propose the following two new concepts to push the expected future rewards toward zero. (i) Reward redistribution that leads to return-equivalent decision processes with the same optimal policies and, when optimal, zero expected future rewards. (ii) Return decomposition via contribution analysis which transforms the reinforcement learning task into a regression task at which deep learning excels. On artificial tasks with delayed rewards, RUDDER is significantly faster than MC and exponentially faster than Monte Carlo Tree Search (MCTS), TD({\lambda}), and reward shaping approaches. At Atari games, RUDDER on top of a Proximal Policy Optimization (PPO) baseline improves the scores, which is most prominent at games with delayed rewards. Source code is available at \url{https://github.com/ml-jku/rudder} and demonstration videos at \url{https://goo.gl/EQerZV}.

PDF Abstract NeurIPS 2019 PDF NeurIPS 2019 Abstract

Code

Add Remove Mark official

ml-jku/baselines-rudder official

265

ml-jku/rudder official

Tasks

Add Remove

Atari Games

reinforcement-learning

Reinforcement Learning (RL)

Datasets

OpenAI Gym

Arcade Learning Environment

Results from the Paper

Add Remove

Ranked #9 on Atari Games on Atari 2600 Bowling

Get a GitHub badge

Task	Dataset	Model	Metric Name	Metric Value	Global Rank	Benchmark
Atari Games	Atari 2600 Bowling	RUDDER	Score	179	# 9	Compare
Atari Games	Atari 2600 Venture	RUDDER	Score	1350	# 14	Compare
Atari Games	Atari 2600 Yars Revenge	RUDDER	Score	60577	# 12	Compare

Methods

Add Remove

No methods listed for this paper. Add relevant methods here

Edit Social Preview

RUDDER: Return Decomposition for Delayed Rewards

Code Edit Add Remove Mark official

Tasks Edit Add Remove

Datasets Edit

Results from the Paper Edit Add Remove

Methods Edit Add Remove

Code

Add Remove Mark official

Tasks

Add Remove

Datasets

Results from the Paper

Add Remove

Methods

Add Remove