Fetching the paper…
Reading the bibliography…
As a step towards developing zero-shot task generalization capabilities in reinforcement learning (RL), we introduce a new RL problem where the agent should learn to execute sequences of instructions after learning useful skills that solve subtasks.
The efficient learning of multiple task sequences
S. P. Singh · 1991
Earlier work this paper cites.
Transfer of learning by composing solutions of elemental sequential tasks
S. P. Singh · 1992
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Reinforcement learning with hierarchies of machines
R. Parr and S. J. Russell · 1997
Earlier work this paper cites.
Actor-critic algorithms
V. R. Konda and J. N. Tsitsiklis · 1999
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
R. S. Sutton, D. Precup, and S. Singh · 1999
Earlier work this paper cites.
Programmable reinforcement learning agents
D. Andre and S. J. Russell · 2000
Earlier work this paper cites.
Hierarchical reinforcement learning with the maxq value function decomposition
T. G. Dietterich · 2000
Earlier work this paper cites.
State abstraction for programmable reinforcement learning agents
D. Andre and S. J. Russell · 2002
Earlier work this paper cites.
Autonomous discovery of temporal abstractions from interaction with an environment
A. McGovern and A. G. Barto · 2002
Earlier work this paper cites.
Hierarchical policy gradient algorithms
M. Ghavamzadeh and S. Mahadevan · 2003
Earlier work this paper cites.
Dimensionality reduction by learning an invariant mapping
R. Hadsell, S. Chopra, and Y. LeCun · 2006
Earlier work this paper cites.
Walk the talk: Connecting language, knowledge, and action in route instructions
M. MacMahon, B. Stankiewicz, and B. Kuipers · 2006
Earlier work this paper cites.
Building portable options: Skill transfer in reinforcement learning
G. Konidaris and A. G. Barto · 2007
Earlier work this paper cites.
Reinforcement learning for mapping instructions to actions
S. R. K. Branavan, H. Chen, L. S. Zettlemoyer, and R. Barzilay · 2009
Earlier work this paper cites.
Learning to represent spatial transformations with factored higher-order boltzmann machines
R. Memisevic and G. E. Hinton · 2010
Cited alongside, same era.
Learning to interpret natural language navigation instructions from observations
D. L. Chen and R. J. Mooney · 2011
Cited alongside, same era.
Understanding natural language commands for robotic navigation and mobile manipulation
S. Tellex, T. Kollar, S. Dickerson, M. R. Walter, A. G. Banerjee, S. J. Teller, and N. Roy · 2011
Cited alongside, same era.
Learning parameterized skills
B. C. da Silva, G. Konidaris, and A. G. Barto · 2012
Cited alongside, same era.
Transfer in reinforcement learning via shared features
G. Konidaris, I. Scheidwasser, and A. G. Barto · 2012
Cited alongside, same era.
Exploiting similarities among languages for machine translation
Reinforcement learning neural turing machines
W. Zaremba and I. Sutskever · 2015
Later among the works it cites.
Modular multitask reinforcement learning with policy sketches
J. Andreas, D. Klein, and S. Levine · 2016
Later among the works it cites.
Learning feed-forward one-shot learners
L. Bertinetto, J. F. Henriques, J. Valmadre, P. H. Torr, and A. Vedaldi · 2016
Later among the works it cites.
Learning and transfer of modulated locomotor controllers
N. Heess, G. Wayne, Y. Tassa, T. P. Lillicrap, M. A. Riedmiller, and D. Silver · 2016
Later among the works it cites.
Using task features for zero-shot knowledge transfer in lifelong learning
D. Isele, M. Rostami, and E. Eaton · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Mikolov, Q. V. Le, and I. Sutskever · 2013
Cited alongside, same era.
A. Graves, G. Wayne, and I. Danihelka · 2014
Cited alongside, same era.
A clockwork rnn
J. Koutnik, K. Greff, F. Gomez, and J. Schmidhuber · 2014
Cited alongside, same era.
Asking for help using inverse semantics
S. Tellex, R. A. Knepper, A. Li, D. Rus, and N. Roy · 2014
Cited alongside, same era.
Predicting deep zero-shot convolutional neural networks using textual descriptions
J. Lei Ba, K. Swersky, S. Fidler, et al · 2015
Cited alongside, same era.
Listen, attend, and walk: Neural mapping of navigational instructions to action sequences
H. Mei, M. Bansal, and M. R. Walter · 2015
Cited alongside, same era.
Action-conditional video prediction using deep networks in atari games
J. Oh, X. Guo, H. Lee, R. L. Lewis, and S. Singh · 2015
Cited alongside, same era.
Later among the works it cites.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
T. D. Kulkarni, K. R. Narasimhan, A. Saeedi, and J. B. Tenenbaum · 2016
Later among the works it cites.
Control of memory, active perception, and action in minecraft
J. Oh, V. Chockalingam, S. P. Singh, and H. Lee · 2016
Later among the works it cites.
Actor-mimic: Deep multitask and transfer reinforcement learning
E. Parisotto, J. L. Ba, and R. Salakhutdinov · 2016
Later among the works it cites.
Policy distillation
A. A. Rusu, S. G. Colmenarejo, C. Gulcehre, G. Desjardins, J. Kirkpatrick, R. Pascanu, V. Mnih, K. Kavukcuoglu, and R. Hadsell · 2016
Later among the works it cites.
High-dimensional continuous control using generalized advantage estimation
J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel · 2016
Later among the works it cites.
The option-critic architecture
P.-L. Bacon, J. Harb, and D. Precup · 2017
Closest in time.
Hierarchical multiscale recurrent neural networks
J. Chung, S. Ahn, and Y. Bengio · 2017
Closest in time.
Learning modular neural network policies for multi-task and multi-robot transfer
C. Devin, A. Gupta, T. Darrell, P. Abbeel, and S. Levine · 2017
Closest in time.
Stochastic neural networks for hierarchical reinforcement learning
C. Florensa, Y. Duan, and P. Abbeel · 2017
Closest in time.
A deep hierarchical approach to lifelong learning in minecraft
C. Tessler, S. Givony, T. Zahavy, D. J. Mankowitz, and S. Mannor · 2017
Closest in time.