Self-monitoring navigation agent via auxiliary progress estimation
Original
Chih-Yao Ma, Jiasen Lu, Zuxuan Wu, G. Al-Regib, Z. Kira, R. Socher, and Caiming Xiong. 2019 · 1901
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Original
Colin Raffel, Noam M. Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, W. Li, and Peter J. Liu. 2020 · 1910
Earlier work this paper cites.
A maximization technique occurring in the statistical analysis of probabilistic functions of markov chains
L. Baum, T. Petrie, George W. Soules, and Norman Weiss. 1970 · 1970
Earlier work this paper cites.
Learning and executing generalized robot plans
R. Fikes, P. Hart, and N. Nilsson. 1972 · 1972
Earlier work this paper cites.
Human problem solving
A. Newell. 1973 · 1973
Earlier work this paper cites.
Planning in a hierarchy of abstraction spaces
E. Sacerdoti. 1973 · 1973
Earlier work this paper cites.
The principle of maximum causal entropy for estimating interacting processes
Brian D Ziebart, J Andrew Bagnell, and Anind K Dey. 2013 · 1980
Earlier work this paper cites.
A tutorial on hidden markov models and selected applications
Lawrence R. Rabiner. 1989 · 1989
Earlier work this paper cites.
Feudal reinforcement learning
P. Dayan and Geoffrey E. Hinton. 1992 · 1992
Earlier work this paper cites.
Hybrid models for motion control systems
R. Brockett. 1993 · 1993
Earlier work this paper cites.
Hierarchical reinforcement learning with the MAXQ value function decomposition
Thomas G Dietterich. 1999 · 1999
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
R S Sutton, D Precup, and S Singh. 1999 · 1999
Earlier work this paper cites.
Automatic discovery of subgoals in reinforcement learning using diverse density
A. McGovern and A. Barto. 2001 · 2001
Earlier work this paper cites.
Graphical models, exponential families, and variational inference
Martin J. Wainwright and M.I. Jordan. 2008 · 2008
Earlier work this paper cites.
Reinforcement learning for mapping instructions to actions
S. Branavan, Harr Chen, Luke Zettlemoyer, and R. Barzilay. 2009 · 2009
Earlier work this paper cites.
Alfworld: Aligning text and embodied environments for interactive learning
Original
Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and M. Hausknecht. 2021 · 2010
Earlier work this paper cites.
Learning to interpret natural language navigation instructions from observations
David L. Chen and R. Mooney. 2011 · 2011
Earlier work this paper cites.
Hierarchical task and motion planning in the now
L P Kaelbling and T Lozano-Pérez. 2011 · 2011
Earlier work this paper cites.