nle-sample-factory-baseline, 2022
Miffyli · 2022
Later among the works it cites.
Improving intrinsic exploration with language abstractions
Jesse Mu, Victor Zhong, Roberta Raileanu, Minqi Jiang, Noah D. Goodman, Tim Rocktäschel, and Edward Grefenstette · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
A dataset perspective on offline reinforcement learning
Kajetan Schweighofer, Marius-constantin Dinu, Andreas Radler, Markus Hofmarcher, Vihang Prakash Patil, Angela Bitto-Nemling, Hamid Eghbal-zadeh, and Sepp Hochreiter · 2022
Later among the works it cites.
Defining and characterizing reward hacking
Original
Joar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, and David Krueger · 2022
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Later among the works it cites.
React: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao · 2022
Later among the works it cites.
Language reward modulation for pretraining reinforcement learning
Original
Ademi Adeniji, Amber Xie, Carmelo Sferrazza, Younggyo Seo, Stephen James, and Pieter Abbeel · 2023
Closest in time.
Learning about progress from experts
Jake Bruce, Ankit Anand, Bogdan Mazoure, and Rob Fergus · 2023
Closest in time.
Grounding large language models in interactive environments with online reinforcement learning, 2023
Thomas Carta, Clément Romac, Thomas Wolf, Sylvain Lamprier, Olivier Sigaud, and Pierre-Yves Oudeyer · 2023
Closest in time.
Mind2web: Towards a generalist agent for the web
Original
Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Samuel Stevens, Boshi Wang, Huan Sun, and Yu Su · 2023
Closest in time.
Towards a unified agent with foundation models
Norman Di Palo, Arunkumar Byravan, Leonard Hasenclever, Markus Wulfmeier, Nicolas Heess, and Martin Riedmiller · 2023
Closest in time.
Guiding pretraining in reinforcement learning with large language models
Yuqing Du, Olivia Watkins, Zihan Wang, Cédric Colas, Trevor Darrell, Pieter Abbeel, Abhishek Gupta, and Jacob Andreas · 2023
Closest in time.
A study of global and episodic bonuses for exploration in contextual mdps
Mikael Henaff, Minqi Jiang, and Roberta Raileanu · 2023
Closest in time.
Language models can solve computer tasks
Original
Geunwoo Kim, Pierre Baldi, and Stephen McAleer · 2023
Closest in time.
Deep laplacian-based options for temporally-extended exploration
Martin Klissarov and Marlos C. Machado · 2023
Closest in time.
Efficient memory management for large language model serving with pagedattention
Original
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E Gonzalez, Hao Zhang, and Ion Stoica · 2023
Closest in time.
Rlaif: Scaling reinforcement learning from human feedback with ai feedback, 2023
Harrison Lee, Samrat Phatale, Hassan Mansoor, Kellie Lu, Thomas Mesnard, Colton Bishop, Victor Carbune, and Abhinav Rastogi · 2023
Closest in time.
Steve-1: A generative model for text-to-behavior in minecraft, 2023
Shalev Lifshitz, Keiran Paster, Harris Chan, Jimmy Ba, and Sheila McIlraith · 2023
Closest in time.
LIV: language-image representations and rewards for robotic control
Yecheng Jason Ma, Vikash Kumar, Amy Zhang, Osbert Bastani, and Dinesh Jayaraman · 2023
Closest in time.
Mapl: Parameter-efficient adaptation of unimodal pre-trained models for vision-language few-shot prompting
Oscar Mañas, Pau Rodriguez Lopez, Saba Ahmadi, Aida Nematzadeh, Yash Goyal, and Aishwarya Agrawal · 2023
Closest in time.
Nethack is hard to hack
Original
Ulyana Piterbarg, Lerrel Pinto, and Rob Fergus · 2023
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Original
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D Manning, and Chelsea Finn · 2023
Closest in time.
Policy optimization in a noisy neighborhood: On return landscapes in continuous control
Nathan Rahn, Pierluca D’Oro, Harley Wiltzer, Pierre-Luc Bacon, and Marc G Bellemare · 2023
Closest in time.
Proxy objectives in reinforcement learning from human feedback
John Schulman · 2023
Closest in time.
From pixels to ui actions: Learning to follow instructions via graphical user interfaces
Original
Peter Shaw, Mandar Joshi, James Cohan, Jonathan Berant, Panupong Pasupat, Hexiang Hu, Urvashi Khandelwal, Kenton Lee, and Kristina Toutanova · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Original
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Closest in time.
Voyager: An open-ended embodied agent with large language models
Original
Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar · 2023
Closest in time.
Spring: Gpt-4 out-performs rl algorithms by studying papers and reasoning
Original
Yue Wu, So Yeon Min, Shrimai Prabhumoye, Yonatan Bisk, Ruslan Salakhutdinov, Amos Azaria, Tom Mitchell, and Yuanzhi Li · 2023
Closest in time.
Language to rewards for robotic skill synthesis
Original
Wenhao Yu, Nimrod Gileadi, Chuyuan Fu, Sean Kirmani, Kuang-Huei Lee, Montse Gonzalez Arenas, Hao-Tien Lewis Chiang, Tom Erez, Leonard Hasenclever, Jan Humplik, et al · 2023
Closest in time.
Discovering policies with domino: Diversity optimization maintaining near optimality
Tom Zahavy, Yannick Schroecker, Feryal M. P. Behbahani, Kate Baumli, Sebastian Flennerhag, Shaobo Hou, and Satinder Singh · 2023
Closest in time.
Omni: Open-endedness via models of human notions of interestingness
Original
Jenny Zhang, Joel Lehman, Kenneth Stanley, and Jeff Clune · 2023
Closest in time.