Fetching the paper…
Reading the bibliography…
We introduce Language Feedback Models (LFMs) that identify desirable behaviour - actions that help achieve tasks specified in the instruction - for imitation learning in instruction following.
A framework for behavioural cloning
M. Bain and C. Sammut · 1995
Earlier work this paper cites.
Is imitation learning the route to humanoid robots?
S. Schaal · 1999
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
P. Abbeel and A. Y. Ng · 2004
Earlier work this paper cites.
Walk the talk: Connecting language, knowledge, and action in route instructions
M. MacMahon, B. Stankiewicz, and B. Kuipers · 2006
Earlier work this paper cites.
Toward understanding natural language directions
T. Kollar, S. Tellex, D. Roy, and N. Roy · 2010
Earlier work this paper cites.
Learning to Interpret Natural Language Navigation Instructions from Observations
D. Chen and R. Mooney · 2011
Earlier work this paper cites.
A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning
S. Ross, G. J. Gordon, and J. A. Bagnell · 2011
Earlier work this paper cites.
Alignment-based compositional semantics for instruction following
J. Andreas and D. Klein · 2015
Earlier work this paper cites.
Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks
S. Bengio, O. Vinyals, N. Jaitly, and N. Shazeer · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition, 2015
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Earlier work this paper cites.
spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing
M. Honnibal and I. Montani · 2017
Earlier work this paper cites.
Unified Pragmatic Models for Generating and Following Instructions
D. Fried, J. Andreas, and D. Klein · 2018
Earlier work this paper cites.
Reinforcement Learning: An Introduction
R. S. Sutton and A. G. Barto · 2018
Earlier work this paper cites.
Behavioral cloning from observation, 2018
F. Torabi, G. Warnell, and P. Stone · 2018
Earlier work this paper cites.
Learning to Map Natural Language Instructions to Physical Quadcopter Control using Simulated Flight
V. Blukis, Y. Terme, E. Niklasson, R. A. Knepper, and Y. Artzi · 2019
Earlier work this paper cites.
Touchdown: Natural Language Navigation and Spatial Reasoning in Visual Street Environments
H. Chen, A. Suhr, D. Misra, N. Snavely, and Y. Artzi · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Cited alongside, same era.
On Warm-Starting Neural Network Training
J. T. Ash and R. P. Adams · 2020
Cited alongside, same era.
Language Models are Few-Shot Learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Cited alongside, same era.
ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks
M. Shridhar, J. Thomason, D. Gordon, Y. Bisk, W. Han, R. Mottaghi, L. Zettlemoyer, and D. Fox · 2020
Cited alongside, same era.
Learning to summarize from human feedback
N. Stiennon, L. Ouyang, J. Wu, D. M. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. Christiano · 2020
Cited alongside, same era.
Analyzing Generalization of Vision and Language Navigation to Unseen Outdoor Areas
R. Schumann and S. Riezler · 2022
Later among the works it cites.
Scienceworld: Is your agent smarter than a 5th grader?
R. Wang, P. A. Jansen, M.-A. Côté, and P. Ammanabrolu · 2022
Later among the works it cites.
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. Le, and D. Zhou · 2022
Later among the works it cites.
RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control, 2023
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn, P. Florence, C. Fu, M. G. Arenas, K. Gopalakrishnan, K. Han, K. Hausman, A. Herzog, J. Hsu, B. Ichter, A. Irpan, N. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, I. Leal, L. Lee, T.-W. E. Lee, S. Levine, Y. Lu, H. Michalewski, I. Mordatch, K. Pertsch, K. Rao, K. Reymann, M. Ryoo, G. Salazar, P. Sanketi, P. Sermanet, J. Singh, A. Singh, R. Soricut, H. Tran, V. Vanhoucke, Q. Vuong, A. Wahid, S. Welker, P. Wohlhart, J. Wu, F. Xia, T. Xiao, P. Xu, S. Xu, T. Yu, and B. Zitkovich · 2023
Later among the works it cites.
Augmenting autotelic agents with large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
RTFM: Generalising to Novel Environment Dynamics via Reading
V. Zhong, T. Rocktäschel, and E. Grefenstette · 2020
Cited alongside, same era.
Fine-Tuning Language Models from Human Preferences, 2020
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving · 2020
Cited alongside, same era.
Language Models are Few-Shot Butlers
V. Micheli and F. Fleuret · 2021
Cited alongside, same era.
Learning Transferable Visual Models From Natural Language Supervision, 2021
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Cited alongside, same era.
SILG: The Multi-environment Symbolic Interactive Language Grounding Benchmark
V. Zhong, A. W. Hanjie, S. I. Wang, K. Narasimhan, and L. Zettlemoyer · 2021
Cited alongside, same era.
Do As I Can, Not As I Say: Grounding Language in Robotic Affordances, 2022
M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, C. Fu, K. Gopalakrishnan, K. Hausman, A. Herzog, D. Ho, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, E. Jang, R. J. Ruano, K. Jeffrey, S. Jesmonth, N. J. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, K.-H. Lee, S. Levine, Y. Lu, L. Luu, C. Parada, P. Pastor, J. Quiambao, K. Rao, J. Rettinghouse, D. Reyes, P. Sermanet, N. Sievers, C. Tan, A. Toshev, V. Vanhoucke, F. Xia, T. Xiao, P. Xu, S. Xu, M. Yan, and A. Zeng · 2022
Cited alongside, same era.
Constitutional AI: Harmlessness from AI Feedback, 2022
Y. Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon, C. Chen, C. Olsson, C. Olah, D. Hernandez, D. Drain, D. Ganguli, D. Li, E. Tran-Johnson, E. Perez, J. Kerr, J. Mueller, J. Ladish, J. Landau, K. Ndousse, K. Lukosuite, L. Lovitt, M. Sellitto, N. Elhage, N. Schiefer, N. Mercado, N. DasSarma, R. Lasenby, R. Larson, S. Ringer, S. Johnston, S. Kravec, S. E. Showk, S. Fort, T. Lanham, T. Telleen-Lawton, T. Conerly, T. Henighan, T. Hume, S. R. Bowman, Z. Hatfield-Dodds, B. Mann, D. Amodei, N. Joseph, S. McCandlish, T. Brown, and J. Kaplan · 2022
Cited alongside, same era.
C. Colas, L. Teodorescu, P.-Y. Oudeyer, X. Yuan, and M.-A. Côté · 2023
Later among the works it cites.
Mind2Web: Towards a Generalist Agent for the Web
X. Deng, Y. Gu, B. Zheng, S. Chen, S. Stevens, B. Wang, H. Sun, and Y. Su · 2023
Later among the works it cites.
Guiding Pretraining in Reinforcement Learning with Large Language Models
Y. Du, O. Watkins, Z. Wang, C. Colas, T. Darrell, P. Abbeel, A. Gupta, and J. Andreas · 2023
Later among the works it cites.
RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback, 2023
H. Lee, S. Phatale, H. Mansoor, T. Mesnard, J. Ferret, K. Lu, C. Bishop, E. Hall, V. Carbune, A. Rastogi, and S. Prakash · 2023
Later among the works it cites.
SwiftSage: A Generative Agent with Fast and Slow Thinking for Complex Interactive Tasks, 2023
B. Y. Lin, Y. Fu, K. Yang, F. Brahman, S. Huang, C. Bhagavatula, P. Ammanabrolu, Y. Choi, and X. Ren · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Reflexion: language agents with verbal reinforcement learning
N. Shinn, F. Cassano, B. Labash, A. Gopinath, K. Narasimhan, and S. Yao · 2023
Later among the works it cites.
ReAct: Synergizing Reasoning and Acting in Language Models
S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao · 2023
Later among the works it cites.
Motif: Intrinsic Motivation from Artificial Intelligence Feedback
M. Klissarov, P. D’Oro, S. Sodhani, R. Raileanu, P.-L. Bacon, P. Vincent, A. Zhang, and M. Henaff · 2024
Closest in time.
VELMA: Verbalization Embodiment of LLM Agents for Vision and Language Navigation in Street View
R. Schumann, W. Zhu, W. Feng, T.-J. Fu, S. Riezler, and W. Y. Wang · 2024
Closest in time.
Self-Rewarding Language Models, 2024
W. Yuan, R. Y. Pang, K. Cho, S. Sukhbaatar, J. Xu, and J. Weston · 2024
Closest in time.