Fetching the paper…
Reading the bibliography…
While large language models (LLMs) show impressive decision-making abilities, current methods lack a mechanism for automatic self-improvement from errors during task execution.
The complexity of markov decision processes
Christos H Papadimitriou and John N Tsitsiklis · 1987
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Leslie Pack Kaelbling, Michael L Littman, and Anthony R Cassandra · 1998
Earlier work this paper cites.
Off-road obstacle avoidance through end-to-end learning
Urs Muller, Jan Ben, Eric Cosatto, Beat Flepp, and Yann Cun · 2005
Earlier work this paper cites.
Alfworld: Aligning text and embodied environments for interactive learning
Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and J Andrew Bagnell · 2011
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer · 2015
Earlier work this paper cites.
Learning using privileged information: similarity control and knowledge transfer
Vladimir Vapnik, Rauf Izmailov, et al · 2015
Earlier work this paper cites.
Learning deep control policies for autonomous aerial vehicles with mpc-guided policy search
Tianhao Zhang, Gregory Kahn, Sergey Levine, and Pieter Abbeel · 2016
Earlier work this paper cites.
Imitating driver behavior with generative adversarial networks
Alex Kuefler, Jeremy Morton, Tim Wheeler, and Mykel Kochenderfer · 2017
Earlier work this paper cites.
Chauffeurnet: Learning to drive by imitating the best and synthesizing the worst
Mayank Bansal, Alex Krizhevsky, and Abhijit Ogale · 2018
Earlier work this paper cites.
Data-driven planning via imitation learning
Sanjiban Choudhury, Mohak Bhardwaj, Sankalp Arora, Ashish Kapoor, Gireeja Ranade, Sebastian Scherer, and Debadeepta Dey · 2018
Earlier work this paper cites.
Nl2bash: A corpus and semantic parser for natural language interface to the linux operating system
Xi Victoria Lin, Chenglong Wang, Luke Zettlemoyer, and Michael D Ernst · 2018
Earlier work this paper cites.
Exploring the limitations of behavior cloning for autonomous driving
Felipe Codevilla, Eder Santana, Antonio M López, and Adrien Gaidon · 2019
Earlier work this paper cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi · 2019
Earlier work this paper cites.
Learning by cheating
Dian Chen, Brady Zhou, Vladlen Koltun, and Philipp Krähenbühl · 2020
Earlier work this paper cites.
Conservative q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Earlier work this paper cites.
Learning quadrupedal locomotion over challenging terrain
Joonho Lee, Jemin Hwangbo, Lorenz Wellhausen, Vladlen Koltun, and Marco Hutter · 2020
Earlier work this paper cites.
On exposure bias, hallucination and domain shift in neural machine translation
Chaojun Wang and Rico Sennrich · 2020
Earlier work this paper cites.
Fighting copycat agents in behavioral cloning from observation histories
Chuan Wen, Jierui Lin, Trevor Darrell, Dinesh Jayaraman, and Yang Gao · 2020
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Cited alongside, same era.
Shaking the foundations: delusions in sequence models for interaction and control
Pedro A Ortega, Markus Kunesch, Grégoire Delétang, Tim Genewein, Jordi Grau-Moya, Joel Veness, Jonas Buchli, Jonas Degrave, Bilal Piot, Julien Perolat, et al · 2021
Cited alongside, same era.
On covariate shift of latent confounders in imitation and reinforcement learning
Guy Tennenholtz, Assaf Hallak, Gal Dalal, Shie Mannor, Gal Chechik, and Uri Shalit · 2021
Cited alongside, same era.
Bridging the imitation gap by adaptive insubordination
Refiner: Reasoning feedback on intermediate representations
Debjit Paul, Mete Ismayilzada, Maxime Peyrard, Beatriz Borges, Antoine Bosselut, Robert West, and Boi Faltings · 2023
Later among the works it cites.
Code llama: Open foundation models for code
Baptiste Roziere, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Romain Sauvestre, Tal Remez, et al · 2023
Later among the works it cites.
Reflexion: Language agents with verbal reinforcement learning.(2023)
Noah Shinn, Federico Cassano, Beck Labash, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao · 2023
Later among the works it cites.
Shepherd: A critic for language model generation
Tianlu Wang, Ping Yu, Xiaoqing Ellen Tan, Sean O’Brien, Ramakanth Pasunuru, Jane Dwivedi-Yu, Olga Golovneva, Luke Zettlemoyer, Maryam Fazel-Zarandi, and Asli Celikyilmaz · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Luca Weihs, Unnat Jain, Iou-Jen Liu, Jordi Salvador, Svetlana Lazebnik, Aniruddha Kembhavi, and Alex Schwing · 2021
Cited alongside, same era.
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al · 2022
Cited alongside, same era.
Leveraging fully observable policies for learning under partial observability
Hai Nguyen, Andrea Baisero, Dian Wang, Christopher Amato, and Robert Platt · 2022
Cited alongside, same era.
Sequence model imitation learning with unobserved contexts
Gokul Swamy, Sanjiban Choudhury, J Bagnell, and Steven Z Wu · 2022
Cited alongside, same era.
Impossibly good experts and how to follow them
Aaron Walsman, Muru Zhang, Sanjiban Choudhury, Dieter Fox, and Ali Farhadi · 2022
Cited alongside, same era.
Fighting fire with fire: avoiding dnn shortcuts through priming
Chuan Wen, Jianing Qian, Jierui Lin, Jiaye Teng, Dinesh Jayaraman, and Yang Gao · 2022
Cited alongside, same era.
Star: Bootstrapping reasoning with reasoning
Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah Goodman · 2022
Cited alongside, same era.
The capacity for moral self-correction in large language models
Deep Ganguli, Amanda Askell, Nicholas Schiefer, Thomas I Liao, Kamilė Lukošiūtė, Anna Chen, Anna Goldie, Azalia Mirhoseini, Catherine Olsson, Danny Hernandez, et al · 2023
Cited alongside, same era.
Sean Welleck, Ximing Lu, Peter West, Faeze Brahman, Tianxiao Shen, Daniel Khashabi, and Yejin Choi · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2023
Later among the works it cites.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Closest in time.
Kto: Model alignment as prospect theoretic optimization
Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky, and Douwe Kiela · 2024
Closest in time.
V-star: Training verifiers for self-taught reasoners
Arian Hosseini, Xingdi Yuan, Nikolay Malkin, Aaron Courville, Alessandro Sordoni, and Rishabh Agarwal · 2024
Closest in time.
Privileged sensing scaffolds reinforcement learning
Edward S Hu, James Springer, Oleh Rybkin, and Dinesh Jayaraman · 2024
Closest in time.
Large language models cannot self-correct reasoning yet
Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou · 2024
Closest in time.
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al · 2024
Closest in time.
Iterative reasoning preference optimization
Richard Yuanzhe Pang, Weizhe Yuan, Kyunghyun Cho, He He, Sainbayar Sukhbaatar, and Jason Weston · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2024
Closest in time.
Step: Stacked llm policies for web actions
Paloma Sodhi, SRK Branavan, Yoav Artzi, and Ryan McDonald · 2024
Closest in time.
Spin: Simultaneous perception interaction and navigation
Shagun Uppal, Ananye Agarwal, Haoyu Xiong, Kenneth Shaw, and Deepak Pathak · 2024
Closest in time.
Intercode: Standardizing and benchmarking interactive coding with execution feedback
John Yang, Akshara Prabhakar, Karthik Narasimhan, and Shunyu Yao · 2024
Closest in time.