Fetching the paper…
Reading the bibliography…
We propose and deploy an approach to continually train an instruction-following agent from feedback provided by users during collaborative interactions.
Fine-tuning language models from human preferences
Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 1909
Earlier work this paper cites.
A generalization of sampling without replacement from a finite universe
D. G. Horvitz and D. J. Thompson · 1952
Earlier work this paper cites.
Changes in reference phrases as a function of frequency of usage in social interaction: a preliminary study
Robert M. Krauss and Sidney Weinheimer · 1964
Earlier work this paper cites.
Referring as a collaborative process
Herbert H Clark and Deanna Wilkes-Gibbs · 1986
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 1992
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
User simulation for spoken dialogue systems: Learning and evaluation
Kallirroi Georgila, James Henderson, and Oliver Lemon · 2006
Earlier work this paper cites.
The epoch-greedy algorithm for multi-armed bandits with side information
John Langford and Tong Zhang · 2007
Earlier work this paper cites.
Interactively shaping agents via human reinforcement: the TAMER framework
W. Bradley Knox and Peter Stone · 2009
Earlier work this paper cites.
Bootstrapping semantic parsers from conversations
Yoav Artzi and Luke Zettlemoyer · 2011
Earlier work this paper cites.
Learning to interpret natural language navigation instructions from observations
David L. Chen and Raymond J. Mooney · 2011
Earlier work this paper cites.
Understanding natural language commands for robotic navigation and mobile manipulation
Stephanie Tellex, Thomas Kollar, Steven Dickerson, Matthew R. Walter, Ashis Gopal Banerjee, Seth Teller, and Nicholas Roy · 2011
Earlier work this paper cites.
Weakly supervised learning of semantic parsers for mapping instructions to actions
Yoav Artzi and Luke Zettlemoyer · 2013
Earlier work this paper cites.
Training a robot via human feedback: A case study
W. Bradley Knox, Peter Stone, and Cynthia Breazeal · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Learning to interpret natural language commands through human-robot dialog
Jesse Thomason, Shiqi Zhang, Raymond Mooney, and Peter Stone · 2015
Cited alongside, same era.
PAC reinforcement learning with rich observations
Akshay Krishnamurthy, Alekh Agarwal, and John Langford · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Cited alongside, same era.
Learning language games through interaction
Sida I. Wang, Percy Liang, and Christopher D. Manning · 2016
Cited alongside, same era.
Learning a neural semantic parser from user feedback
Srinivasan Iyer, Ioannis Konstas, Alvin Cheung, Jayant Krishnamurthy, and Luke Zettlemoyer · 2017
Cited alongside, same era.
Dialogue learning with human teaching and feedback in end-to-end trainable task-oriented dialogue systems
Bing Liu, Gokhan Tür, Dilek Hakkani-Tür, Pararth Shah, and Larry Heck · 2018
Later among the works it cites.
Mapping instructions to actions in 3D environments with visual goal prediction
Dipendra Misra, Andrew Bennett, Valts Blukis, Eyvind Niklasson, Max Shatkhin, and Yoav Artzi · 2018
Later among the works it cites.
Learning to map context-dependent sentences to executable formal queries
Alane Suhr, Srinivasan Iyer, and Yoav Artzi · 2018
Later among the works it cites.
Deep TAMER: Interactive agent shaping in high-dimensional state spaces
Garrett Warnell, Nicholas R. Waytowich, Vernon Lawhern, and Peter R Stone · 2018
Later among the works it cites.
Touchdown: Natural language navigation and spatial reasoning in visual street environments
Howard Chen, Alane Suhr, Dipendra Misra, Noah Snavely, and Yoav Artzi · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
AI2-THOR: An interactive 3D environment for visual AI
Eric Kolve, Roozbeh Mottaghi, Winson Han, Eli VanderBilt, Luca Weihs, Alvaro Herrasti, Daniel Gordon, Yuke Zhu, Abhinav Gupta, and Ali Farhadi · 2017
Cited alongside, same era.
Interactive learning from policy-dependent human feedback
James MacGlashan, Mark K. Ho, Robert Tyler Loftin, Bei Peng, David L. Roberts, Matthew E. Taylor, and Michael L. Littman · 2017
Cited alongside, same era.
Mapping instructions and visual observations to actions with reinforcement learning
Dipendra Misra, John Langford, and Yoav Artzi · 2017
Cited alongside, same era.
Reinforcement learning for bandit neural machine translation with simulated human feedback
Khanh Nguyen, Hal Daumé III, and Jordan Boyd-Graber · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sünderhauf, Ian Reid, Stephen Gould, and Anton van den Hengel · 2018
Cited alongside, same era.
Executing instructions in situated collaborative interactions
Alane Suhr, Claudia Yan, Charlotte Schluger, Stanley Yu, Hadi Khader, Marwa Mouallem, Iris Zhang, and Yoav Artzi · 2019
Later among the works it cites.
The EMPATHIC framework for task learning from implicit human feedback
Yuchen Cui, Qiping Zhang, Alessandro Allievi, Peter Stone, Scott Niekum, and W Bradley Knox · 2020
Later among the works it cites.
Continual adaptation for efficient machine communication
Robert Hawkins, Minae Kwon, Dorsa Sadigh, and Noah Goodman · 2020
Later among the works it cites.
ALFRED: A benchmark for interpreting grounded instructions for everyday tasks
Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox · 2020
Later among the works it cites.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F. Christiano · 2020
Later among the works it cites.
An imitation game for learning semantic parsers from user interaction
Ziyu Yao, Yiqi Tang, Wen-tau Yih, Huan Sun, and Yu Su · 2020
Later among the works it cites.
Analysis of language change in collaborative instruction following
Anna Effenberger, Rhia Singh, Eva Yan, Alane Suhr, and Yoav Artzi · 2021
Later among the works it cites.
Continual learning for grounded instruction generation by observing human following behavior
Noriyuki Kojima, Alane Suhr, and Yoav Artzi · 2021
Later among the works it cites.
Simulating bandit learning from user feedback for extractive question answering
Ge Gao, Eunsol Choi, and Yoav Artzi · 2022
Closest in time.
Is reinforcement learning (not) for natural language processing: Benchmarks, baselines, and building blocks for natural language policy optimization
Rajkumar Ramamurthy, Prithviraj Ammanabrolu, Kianté Brantley, Jack Hessel, Rafet Sifa, Christian Bauckhage, Hannaneh Hajishirzi, and Yejin Choi · 2023
Closest in time.