Fetching the paper…
Reading the bibliography…
Reinforcement learning is a promising framework for solving control problems, but its use in practical situations is hampered by the fact that reward functions are often difficult to engineer.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian Ziebart, Andrew Maas, Andrew Bagnell, and Anind Dey · 2008
Earlier work this paper cites.
A survey of robot learning from demonstration
Brenna D. Argall, Sonia Chernova, Manuela Veloso, and Brett Browning · 2009
Earlier work this paper cites.
Reinforcement learning for mapping instructions to actions
S. R. K. Branavan, Harr Chen, Luke S. Zettlemoyer, and Regina Barzilay · 2009
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Brian Ziebart · 2010
Earlier work this paper cites.
Understanding natural language commands for robotic navigation and mobile manipulation
Stefanie Tellex, Thomas Kollar, Steven Dickerson, Matthew R. Walter, Ashis Gopal Banerjee, Seth Teller, and Nicholas Roy · 2011
Earlier work this paper cites.
Nonparametric bayesian inverse reinforcement learning for multiple reward functions
Jaedeug Choi and Kee-eung Kim · 2012
Earlier work this paper cites.
European conference on recent advances in reinforcement learning (ewrl)
Christos Dimitrakakis and Constantin A. Rothkopf · 2012
Earlier work this paper cites.
Tell me dave: Context-sensitive grounding of natural language to manipulation instructions
Dipendra Kumar Misra, Jaeyong Sung, Kevin Lee, and Ashutosh Saxena · 2014
Earlier work this paper cites.
Robot programming by demonstration with situated spatial language understanding
Maxwell Forbes, Rajesh P. N. Rao, Luke Zettlemoyer, and Maya Cakmak · 2015
Earlier work this paper cites.
Grounding english commands to reward functions
James MacGlashan, Monica Babes-Vroman, Marie desJardins, Michael L. Littman, Smaranda Muresan, Shawn Squire, Stefanie Tellex, Dilip Arumugam, and Lei Yang · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Cited alongside, same era.
Robobarista: Object part based transfer of manipulation trajectories from crowd-sourcing in 3d pointclouds
Jaeyong Sung, Seok Hyun Jin, and Ashutosh Saxena · 2015
Cited alongside, same era.
Chelsea Finn, Paul Christiano, Pieter Abbeel, and Sergey Levine · 2016
Cited alongside, same era.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon · 2016
Mapping instructions and visual observations to actions with reinforcement learning
Dipendra Kumar Misra, John Langford, and Yoav Artzi · 2017
Later among the works it cites.
Semantic scene completion from a single depth image
Shuran Song, Fisher Yu, Andy Zeng, Angel X Chang, Manolis Savva, and Thomas Funkhouser · 2017
Later among the works it cites.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sünderhauf, Ian D. Reid, Stephen Gould, and Anton van den Hengel · 2018
Later among the works it cites.
Learning to Follow Language Instructions with Adversarial Reward Induction
D. Bahdanau, F. Hill, J. Leike, E. Hughes, P. Kohli, and E. Grefenstette · 2018
Later among the works it cites.
Embodied question answering
Abhishek Das, Samyak Datta, Georgia Gkioxari, Stefan Lee, Devi Parikh, and Dhruv Batra · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Listen, attend, and walk: Neural mapping of navigational instructions to action sequences
Hongyuan Mei, Mohit Bansal, and Matthew R. Walter · 2016
Cited alongside, same era.
Deep reinforcement learning from human preferences
P. Christiano, J. Leike, T. B. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Cited alongside, same era.
Grounded language learning in a simulated 3d world
Karl Moritz Hermann, Felix Hill, Simon Green, Fumin Wang, Ryan Faulkner, Hubert Soyer, David Szepesvari, Wojciech Marian Czarnecki, Max Jaderberg, Denis Teplyashin, Marcus Wainwright, Chris Apps, Demis Hassabis, and Phil Blunsom · 2017
Cited alongside, same era.
Meta inverse reinforcement learning via maximum reward sharing for human motion analysis
Kun Li and Joel W. Burdick · 2017
Cited alongside, same era.
Justin Fu, Katie Luo, and Sergey Levine · 2018
Later among the works it cites.
Film: Visual reasoning with a general conditioning layer
Ethan Perez, Florian Strub, Harm de Vries, Vincent Dumoulin, and Aaron C. Courville · 2018
Later among the works it cites.
Pararth Shah, Marek Fiser, Aleksandra Faust, J. Chase Kew, and Dilek Hakkani-Tür · 2018
Later among the works it cites.
Reward learning from narrated demonstrations
Hsiao-Yu Fish Tung, Adam W. Harley, Liang-Kang Huang, and Katerina Fragkiadaki · 2018
Later among the works it cites.