Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) in long horizon and sparse reward tasks is notoriously difficult and requires a lot of training steps.
The development of categorization in the second year and its relation to other cognitive and linguistic developments
Alison Gopnik and Andrew Meltzoff · 1987
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
A. Y. Ng, D. Harada, and S. Russell · 1999
Earlier work this paper cites.
Gaussian processes for machine learning
C. E. Rasmussen and C. K. I Williams · 2005
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Deep recurrent q-learning for partially observable mdps
M. Hausknecht and P. Stone · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Earlier work this paper cites.
Generative adversarial imitation learning
Stefano Ermon Jonathan Ho · 2016
Earlier work this paper cites.
Listen, attend, and walk: Neural mapping of navigational instructions to action sequences
Hongyuan Mei, Mohit Bansal, and Matthew R. Walter · 2016
Earlier work this paper cites.
Modular multitask reinforcement learning with policy sketches
J. Andreas, D. Klein, and S. Levine · 2017
Earlier work this paper cites.
Grounded language learning in a simulated 3d world
Karl Moritz Hermann, Felix Hill, Simon Green, Fumin Wang, Ryan Faulkner, Hubert Soyer, David Szepesvari, Wojciech Marian Czarnecki, Max Jaderberg, and Denis Teplyashin · 2017
Earlier work this paper cites.
Mapping instructions and visual observations to actions with reinforcement learning
Dipendra Misra, John Langford, and Yoav Artzi · 2017
Earlier work this paper cites.
Film: Visual reasoning with a general conditioning layer
Ethan Perez, Florian Strub, Harm de Vries, Vincent Dumoulin, and Aaron Courville · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sünderhauf, Ian Reid, Stephen Gould, and Anton van den Henge · 2018
Earlier work this paper cites.
Learning to understand goal specifications by modelling reward
Dzmitry Bahdanau, Felix Hill, Jan Leike, Edward Hughes, Arian Hosseini, Pushmeet Kohli, and Edward Grefenstette · 2018
Cited alongside, same era.
Two can play this game: Visual dialog with discriminative question generation and answering
U. Jain, S. Lazebnik, and A. G. Schwing · 2018
Cited alongside, same era.
Reinforcement Learning: An Introduction
R. S. Sutton and A. G. Barto · 2018
Cited alongside, same era.
A survey on intrinsic motivation in reinforcement learning
Arthur Aubret, Laetitia Matignon, and Salima Hassas · 2019
Cited alongside, same era.
Learning to understand goal specifications by modelling reward
Dzmitry Bahdanau, Felix Hill, Edward Hughes Jan Leike, Arian Hosseini, Pushmeet Kohli, and Edward Grefenstette · 2019
Language as a cognitive tool to imagine goals in curiosity-driven exploration
Cédric Colas, Tristan Karch, Nicolas Lair, Jean-Michel Dussoux, Clément Moulin-Frier, Peter Ford Dominey, and Pierre-Yves Oudeyer · 2020
Later among the works it cites.
C. Lynch and P. Sermanet · 2020
Later among the works it cites.
Ride: Rewarding impact-driven exploration for procedurally-generated environments
Roberta Raileanu and Tim Rocktäschel · 2020
Later among the works it cites.
Offline reinforcement learning with implicit q-learning
Ilya Kostrikov, Ashvin Nair, and Sergey Levine · 2021
Later among the works it cites.
Ella: Exploration through learned language abstraction
Suvir Mirchandani, Siddharth Karamcheti, and Dorsa Sadigh · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Babyai: A platform to study the sample efficiency of grounded language learning
Maxime Chevalier-Boisvert, Dzmitry Bahdanau, Salem Lahlou, Lucas Willems, Chitwan Saharia, Thien Huu Nguyen, and Yoshua Bengio · 2019
Cited alongside, same era.
Using natural language for reward shaping in reinforcement learning
Goyal, Niekum, and Mooney · 2019
Cited alongside, same era.
Hierarchical decision making by generating and following natural language instructions
Hengyuan Hu, Denis Yarats, Qucheng Gong, Yuandong Tian, and Mike Lewis · 2019
Cited alongside, same era.
Language as an abstraction for hierarchical deep reinforcement learning
Y. Jiang, S. S. Gu, K. P. Murphy, and C. Finn · 2019
Cited alongside, same era.
A survey of reinforcement learning informed by natural language
Jelena Luketina, Nantas Nardelli, Gregory Farquhar, Jakob Foerster, Jacob Andreas, Edward Grefenstette, Shimon Whiteson, and Tim Rocktäschel · 2019
Cited alongside, same era.
Alfred: A benchmark for interpreting grounded instructions for everyday tasks
Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox · 2019
Cited alongside, same era.
A narration-based reward shaping approach using grounded natural language commands
Nicholas Waytowich, Sean L. Barton, Vernon Lawhern, and Garrett Warnell · 2019
Cited alongside, same era.
Alexander Pashevich, Cordelia Schmid, and Chen Sun · 2021
Later among the works it cites.
Data-QuestEval: A referenceless metric for data-to-text semantic evaluation
Clement Rebuffel, Thomas Scialom, Laure Soulier, Benjamin Piwowarski, Sylvain Lamprier, Jacopo Staiano, Geoffrey Scoutheeten, and Patrick Gallinari · 2021
Later among the works it cites.
Questeval: Summarization asks for fact-based evaluation
Thomas Scialom, Paul-Alexis Dray, Patrick Gallinari, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, and Alex Wang · 2021
Later among the works it cites.
Autotelic agents with intrinsically motivated goal-conditioned reinforcement learning: a short survey
Cédric Colas, Tristan Karch, Olivier Sigaud, and Pierre-Yves Oudeyer · 2022
Closest in time.
Asking for knowledge: Training rl agents to query external knowledge using language
Iou-Jen Liu, Xingdi Yuan, Marc-Alexandre Côté, Pierre-Yves Oudeyer, and Alexander G. Schwing · 2022
Closest in time.
Improving intrinsic exploration with language abstractions
Jesse Mu, Victor Zhong, Roberta Raileanu, Minqi Jiang, Noah Goodman, Tim Rocktäschel, and Edward Grefenstette · 2022
Closest in time.
Semantic exploration from language abstractions and pretrained representations
Allison C Tam, Neil C Rabinowitz, Andrew K Lampinen, Nicholas A Roy, Stephanie CY Chan, DJ Strouse, Jane X Wang, Andrea Banino, and Felix Hill · 2022
Closest in time.
Perceiving the world: Question-guided reinforcement learning for text-based games
Yunqiu Xu, Meng Fang, Ling Chen, Yali Du, Joey Tianyi Zhou, and Chengqi Zhang · 2022
Closest in time.