Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have emerged as powerful and general solutions to many natural language tasks.
D4rl: Datasets for deep data-driven reinforcement learning
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine · 2004
Earlier work this paper cites.
On-line policy optimisation of spoken dialogue systems via live interaction with human subjects
Milica Gašić, Filip Jurčíček, Blaise Thomson, Kai Yu, and Steve Young · 2011
Earlier work this paper cites.
Finale Doshi-Velez and George Dimitri Konidaris · 2013
Earlier work this paper cites.
Distilling the knowledge in a neural network, 2015
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Sequence-level knowledge distillation
Yoon Kim and Alexander M. Rush · 2016
Earlier work this paper cites.
Hafez: an interactive poetry generation system
Marjan Ghazvininejad, Xing Shi, Jay Priyadarshi, and Kevin Knight · 2017
Earlier work this paper cites.
Trainable greedy decoding for neural machine translation
Jiatao Gu, Kyunghyun Cho, and Victor O.K. Li · 2017
Earlier work this paper cites.
Learning to decode for future success, 2017
Jiwei Li, Will Monroe, and Dan Jurafsky · 2017
Earlier work this paper cites.
A deep reinforced model for abstractive summarization, 2017
Romain Paulus, Caiming Xiong, and Richard Socher · 2017
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2018
Earlier work this paper cites.
Decoupling strategy and generation in negotiation dialogues, 2018
He He, Derek Chen, Anusha Balakrishnan, and Percy Liang · 2018
Earlier work this paper cites.
Learning to write with cooperative discriminators
Ari Holtzman, Jan Buys, Maxwell Forbes, Antoine Bosselut, David Golub, and Yejin Choi · 2018
Earlier work this paper cites.
Learning to extract coherent summary via deep reinforcement learning, 2018
Yuxiang Wu and Baotian Hu · 2018
Earlier work this paper cites.
Better rewards yield better summaries: Learning to summarise without references
Florian Böhm, Yang Gao, Christian M. Meyer, Ori Shapira, Ido Dagan, and Iryna Gurevych · 2019
Earlier work this paper cites.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, Craig Ferguson, Àgata Lapedriza, Noah Jones, Shixiang Gu, and Rosalind W. Picard · 2019
Earlier work this paper cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
Aviral Kumar, Justin Fu, Matthew Soh, George Tucker, and Sergey Levine · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Multiwoz – a large-scale multi-domain wizard-of-oz dataset for task-oriented dialogue modelling, 2020
Paweł Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, Iñigo Casanueva, Stefan Ultes, Osman Ramadan, and Milica Gašić · 2020
Cited alongside, same era.
Rl unplugged: Benchmarks for offline reinforcement learning
Caglar Gulcehre, Ziyu Wang, Alexander Novikov, Tom Le Paine, Sergio Gómez Colmenarejo, Konrad Zolna, Rishabh Agarwal, Josh Merel, Daniel Mankowitz, Cosmin Paduraru, et al · 2020
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Cited alongside, same era.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Cited alongside, same era.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Defining and characterizing reward hacking
Joar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, and David Krueger · 2022
Later among the works it cites.
Offline rl for natural language generation with implicit language q learning
Charlie Snell, Ilya Kostrikov, Yi Su, Mengjiao Yang, and Sergey Levine · 2022
Later among the works it cites.
Chai: A chatbot ai for task-oriented dialogue with offline reinforcement learning, 2022
Siddharth Verma, Justin Fu, Mengjiao Yang, and Sergey Levine · 2022
Later among the works it cites.
Grounding large language models in interactive environments with online reinforcement learning, 2023
Thomas Carta, Clément Romac, Thomas Wolf, Sylvain Lamprier, Olivier Sigaud, and Pierre-Yves Oudeyer · 2023
Closest in time.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Fine-tuning language models from human preferences, 2020
Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 2020
Cited alongside, same era.
Offline reinforcement learning with fisher divergence critic regularization
Ilya Kostrikov, Jonathan Tompson, Rob Fergus, and Ofir Nachum · 2021
Cited alongside, same era.
Recursively summarizing books with human feedback
Jeff Wu, Long Ouyang, Daniel M Ziegler, Nisan Stiennon, Ryan Lowe, Jan Leike, and Paul Christiano · 2021
Cited alongside, same era.
FUDGE: Controlled text generation with future discriminators
Kevin Yang and Dan Klein · 2021
Cited alongside, same era.
Improving alignment of dialogue agents via targeted human judgements
Amelia Glaese, Nat McAleese, Maja Trebacz, John Aslanides, Vlad Firoiu, Timo Ewalds, Maribeth Rauh, Laura Weidinger, Martin Chadwick, Phoebe Thacker, et al · 2022
Cited alongside, same era.
GPT-critic: Offline reinforcement learning for end-to-end task-oriented dialogue systems
Youngsoo Jang, Jongmin Lee, and Kee-Eung Kim · 2022
Cited alongside, same era.
When should we prefer offline reinforcement learning over behavioral cloning?, 2022
Aviral Kumar, Joey Hong, Anikait Singh, and Sergey Levine · 2022
Cited alongside, same era.
Closest in time.
Deep reinforcement learning from human preferences, 2023
Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2023
Closest in time.
Reinforced self-training (rest) for language modeling, 2023
Caglar Gulcehre, Tom Le Paine, Srivatsan Srinivasan, Ksenia Konyushkova, Lotte Weerts, Abhishek Sharma, Aditya Siddhant, Alex Ahern, Miaosen Wang, Chenjie Gu, Wolfgang Macherey, Arnaud Doucet, Orhan Firat, and Nando de Freitas · 2023
Closest in time.
Generative agents: Interactive simulacra of human behavior, 2023
Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein · 2023
Closest in time.
Direct preference optimization: Your language model is secretly a reward model, 2023
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn · 2023
Closest in time.
Is reinforcement learning (not) for natural language processing: Benchmarks, baselines, and building blocks for natural language policy optimization
Rajkumar Ramamurthy, Prithviraj Ammanabrolu, Kianté Brantley, Jack Hessel, Rafet Sifa, Christian Bauckhage, Hannaneh Hajishirzi, and Yejin Choi · 2023
Closest in time.
Code llama: Open foundation models for code, 2023
Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, Artyom Kozhevnikov, Ivan Evtimov, Joanna Bitton, Manish Bhatt, Cristian Canton Ferrer, Aaron Grattafiori, Wenhan Xiong, Alexandre Défossez, Jade Copet, Faisal Azhar, Hugo Touvron, Louis Martin, Nicolas Usunier, Thomas Scialom, and Gabriel Synnaeve · 2023
Closest in time.
Personality traits in large language models, 2023
Greg Serapio-García, Mustafa Safdari, Clément Crepy, Luning Sun, Stephen Fitz, Peter Romero, Marwa Abdulhai, Aleksandra Faust, and Maja Matarić · 2023
Closest in time.
Offline rl for natural language generation with implicit language q learning
Charlie Snell, Ilya Kostrikov, Yi Su, Mengjiao Yang, and Sergey Levine · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Closest in time.
Chain-of-thought prompting elicits reasoning in large language models, 2023
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou · 2023
Closest in time.
A study on robustness and reliability of large language model code generation, 2023
Li Zhong and Zilong Wang · 2023
Closest in time.