“Policy invariance under reward transformations: Theory and application to reward shaping”
Andrew Ng, Daishi Harada and Stuart Russell · 1999
Earlier work this paper cites.
“Human Instruction-Following with Deep Reinforcement Learning via Transfer-Learning from Text”, 2020
Original
Felix Hill, Sona Mokra, Nathaniel Wong and Tim Harley · 2005
Earlier work this paper cites.
“Learning to interpret natural language navigation instructions from observations”
David Chen and Raymond Mooney · 2011
Earlier work this paper cites.
“Understanding natural language commands for robotic navigation and mobile manipulation”
Stefanie Tellex et al · 2011
Earlier work this paper cites.
“Weakly supervised learning of semantic parsers for mapping instructions to actions”
Yoav Artzi and Luke Zettlemoyer · 2013
Earlier work this paper cites.
“Compression and communication in the cultural evolution of linguistic structure”
Simon Kirby, Monica Tamariz, Hannah Cornish and Kenny Smith · 2015
Earlier work this paper cites.
“Human-level control through deep reinforcement learning”
Volodymyr Mnih et al · 2015
Earlier work this paper cites.
“Universal value function approximators”
Tom Schaul, Daniel Horgan, Karol Gregor and David Silver · 2015
Earlier work this paper cites.
“Neural module networks”
Jacob Andreas, Marcus Rohrbach, Trevor Darrell and Dan Klein · 2016
Earlier work this paper cites.
“Unifying count-based exploration and intrinsic motivation”
Marc Bellemare et al · 2016
Earlier work this paper cites.
“Deep reinforcement learning with double Q-Learning”
Hado Hasselt, Arthur Guez and David Silver · 2016
Earlier work this paper cites.
“Virtual embodiment: A scalable long-term strategy for artificial intelligence research”
Original
Douwe Kiela, Luana Bulat, Anita Vero and Stephen Clark · 2016
Earlier work this paper cites.
“Prioritized experience replay”
Tom Schaul, John Quan, Ioannis Antonoglou and David Silver · 2016
Earlier work this paper cites.
“Dueling Network Architectures for Deep Reinforcement Learning”
Ziyu Wang et al · 2016
Earlier work this paper cites.