Fetching the paper…
Reading the bibliography…
We introduce a new benchmark, LLF-Bench (Learning from Language Feedback Benchmark; pronounced as "elf-bench"), to evaluate the ability of AI agents to interactively learn from natural language feedback and instructions.
Focus on formative feedback
V. J. Shute · 2008
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. Gordon, and D. Bagnell · 2011
Earlier work this paper cites.
Language understanding for text-based games using deep reinforcement learning
K. Narasimhan, T. D. Kulkarni, and R. Barzilay · 2015
Earlier work this paper cites.
Concrete problems in AI safety
D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mané · 2016
Earlier work this paper cites.
Natural language communication with robots
Y. Bisk, D. Yuret, and D. Marcu · 2016
Earlier work this paper cites.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Earlier work this paper cites.
Learning language games through interaction
S. I. Wang, P. Liang, and C. D. Manning · 2016
Earlier work this paper cites.
From language to programs: Bridging reinforcement learning and maximum marginal likelihood
K. Guu, P. Pasupat, E. Liu, and P. Liang · 2017
Earlier work this paper cites.
Mapping instructions and visual observations to actions with reinforcement learning
D. Misra, J. Langford, and Y. Artzi · 2017
Earlier work this paper cites.
Active preference-based learning of reward functions
D. Sadigh, A. D. Dragan, S. Sastry, and S. A. Seshia · 2017
Earlier work this paper cites.
World of Bits: An open-domain platform for web-based agents
T. Shi, A. Karpathy, L. Fan, J. Hernandez, and P. Liang · 2017
Earlier work this paper cites.
Batch active preference-based learning of reward functions
E. Biyik and D. Sadigh · 2018
Earlier work this paper cites.
Gated-attention architectures for task-oriented language grounding
D. S. Chaplot, K. M. Sathyendra, R. K. Pasumarthi, D. Rajagopal, and R. Salakhutdinov · 2018
Earlier work this paper cites.
Understanding back-translation at scale
S. Edunov, M. Ott, M. Auli, and D. Grangier · 2018
Earlier work this paper cites.
Representation learning for grounded spatial reasoning
M. Janner, K. Narasimhan, and R. Barzilay · 2018
Earlier work this paper cites.
Reinforcement learning on web interfaces using workflow-guided exploration
E. Z. Liu, K. Guu, P. Pasupat, and P. Liang · 2018
Earlier work this paper cites.
Mapping instructions to actions in 3D environments with visual goal prediction
D. Misra, A. Bennett, V. Blukis, E. Niklasson, M. Shatkhin, and Y. Artzi · 2018
Earlier work this paper cites.
Semantically equivalent adversarial rules for debugging nlp models
M. T. Ribeiro, S. Singh, and C. Guestrin · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Cited alongside, same era.
Learning to understand goal specifications by modelling reward
D. Bahdanau, F. Hill, J. Leike, E. Hughes, P. Kohli, and E. Grefenstette · 2019
Cited alongside, same era.
Active learning of reward dynamics from hierarchical queries
C. Basu, E. Bıyık, Z. He, M. Singhal, and D. Sadigh · 2019
Cited alongside, same era.
Touchdown: Natural language navigation and spatial reasoning in visual street environments
H. Chen, A. Suhr, D. Misra, N. Snavely, and Y. Artzi · 2019
Cited alongside, same era.
BabyAI: First steps towards grounded language learning with a human in the loop
M. Chevalier-Boisvert, D. Bahdanau, S. Lahlou, L. Willems, C. Saharia, T. H. Nguyen, and Y. Bengio · 2019
Safe reinforcement learning with natural language constraints
T.-Y. Yang, M. Hu, Y. Chow, P. Ramadge, and K. R. Narasimhan · 2021
Later among the works it cites.
SILG: The multi-domain symbolic interactive language grounding benchmark
V. Zhong, H. A. Wang, S. Wang, K. R. Narasimhan, and L. Zettlemoyer · 2021
Later among the works it cites.
Lila: Language-informed latent actions
S. Karamcheti, M. Srivastava, P. Liang, and D. Sadigh · 2022
Later among the works it cites.
WebShop: Towards scalable real-world web interaction with grounded language agents
S. Yao, H. Chen, J. Yang, and K. R. Narasimhan · 2022
Later among the works it cites.
LMRL Gym: Benchmarks for multi-turn reinforcement learning with language models
M. Abdulhai, I. White, C. Snell, C. Sun, J. Hong, Y. Zhai, K. Xu, and S. Levine · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On the measure of intelligence
F. Chollet · 2019
Cited alongside, same era.
Textworld: A learning environment for text-based games
M.-A. Côté, A. Kádár, X. Yuan, B. Kybartas, T. Barnes, E. Fine, J. Moore, M. Hausknecht, L. El Asri, M. Adada, W. Tay, and A. Trischler · 2019
Cited alongside, same era.
Universal adversarial triggers for attacking and analyzing NLP
E. Wallace, S. Feng, N. Kandpal, M. Gardner, and S. Singh · 2019
Cited alongside, same era.
Meta-World: A benchmark and evaluation for multi-task and meta reinforcement learning
T. Yu, D. Quillen, Z. He, R. Julian, A. Narayan, H. Shively, A. Bellathur, K. Hausman, C. Finn, and S. Levine · 2019
Cited alongside, same era.
The NetHack learning environment
H. Küttler, N. Nardelli, A. H. Miller, R. Raileanu, M. Selvatici, E. Grefenstette, and T. Rocktäschel · 2020
Cited alongside, same era.
ALFRED: A benchmark for interpreting grounded instructions for everyday tasks
M. Shridhar, J. Thomason, D. Gordon, Y. Bisk, W. Han, R. Mottaghi, L. Zettlemoyer, and D. Fox · 2020
Cited alongside, same era.
S. Arora, A. Narayan, M. F. Chen, L. Orr, N. Guha, K. Bhatia, I. Chami, and C. Re · 2023
Closest in time.
No, to the right: Online language corrections for robotic manipulation via shared autonomy
Y. Cui, S. Karamcheti, R. Palleti, N. Shivakumar, P. Liang, and D. Sadigh · 2023
Closest in time.
OpenAGI: When LLM meets domain experts
Y. Ge, W. Hua, J. Ji, J. Tan, S. Xu, and Y. Zhang · 2023
Closest in time.
Gemini: A family of highly capable multimodal models
Gemini Team · 2023
Closest in time.
Language-driven representation learning for robotics
S. Karamcheti, S. Nair, A. S. Chen, T. Kollar, C. Finn, D. Sadigh, and P. Liang · 2023
Closest in time.
The language of prompting: What linguistic properties make a prompt successful?
A. Leidinger, R. van Rooij, and E. Shutova · 2023
Closest in time.
GPT-4 technical report, 2023
OpenAI · 2023
Closest in time.
Toolformer: Language models can teach themselves to use tools
T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, L. Zettlemoyer, N. Cancedda, and T. Scialom · 2023
Closest in time.
Fundamental limitations of alignment in large language models
Y. Wolf, N. Wies, Y. Levine, and A. Shashua · 2023
Closest in time.
Autogen: Enabling next-gen llm applications via multi-agent conversation framework
Q. Wu, G. Bansal, J. Zhang, Y. Wu, S. Zhang, E. Zhu, B. Li, L. Jiang, X. Zhang, and C. Wang · 2023
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
S. Yao, D. Yu, J. Zhao, I. Shafran, T. L. Griffiths, Y. Cao, and K. Narasimhan · 2023
Closest in time.
PromptBench: Towards evaluating the robustness of large language models on adversarial prompts
K. Zhu, J. Wang, J. Zhou, Z. Wang, H. Chen, Y. Wang, L. Yang, W. Ye, Y. Zhang, N. Z. Gong, and X. Xie · 2023
Closest in time.