Continuous vs. discrete optimization of deep neural networks
Omer Elkabetz and Nadav Cohen · 2021
Later among the works it cites.
Revisiting the weaknesses of reinforcement learning for neural machine translation
Samuel Kiegeland and Julia Kreutzer · 2021
Later among the works it cites.
Softmax policy gradient methods can take exponential time to converge
Gen Li, Yuting Wei, Yuejie Chi, Yuantao Gu, and Yuxin Chen · 2021
Later among the works it cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Original
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Later among the works it cites.
Scaling instruction-finetuned language models
Original
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al · 2022
Later among the works it cites.
An alternate policy gradient estimator for softmax policies
Shivam Garg, Samuele Tosatto, Yangchen Pan, Martha White, and Rupam Mahmood · 2022
Later among the works it cites.
Quark: Controllable text generation with reinforced unlearning
Ximing Lu, Sean Welleck, Jack Hessel, Liwei Jiang, Lianhui Qin, Peter West, Prithviraj Ammanabrolu, and Yejin Choi · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
Lamda: Language models for dialog applications
Original
Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al · 2022
Later among the works it cites.
Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks
Original
Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Anjana Arunkumar, Arjun Ashok, Arut Selvan Dhanasekaran, Atharva Naik, David Stap, et al · 2022
Later among the works it cites.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le · 2022
Later among the works it cites.
Raft: Reward ranked finetuning for generative foundation model alignment
Original
Hanze Dong, Wei Xiong, Deepanshu Goyal, Rui Pan, Shizhe Diao, Jipeng Zhang, Kashun Shum, and Tong Zhang · 2023
Closest in time.
Alpacafarm: A simulation framework for methods that learn from human feedback
Original
Yann Dubois, Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto · 2023
Closest in time.
Benign overfitting in linear classifiers and leaky relu networks from kkt conditions for margin maximization
Spencer Frei, Gal Vardi, Peter Bartlett, and Nathan Srebro · 2023
Closest in time.
Reinforced self-training (rest) for language modeling
Original
Caglar Gulcehre, Tom Le Paine, Srivatsan Srinivasan, Ksenia Konyushkova, Lotte Weerts, Abhishek Sharma, Aditya Siddhant, Alex Ahern, Miaosen Wang, Chenjie Gu, et al · 2023
Closest in time.
Aligning language models with offline reinforcement learning from human feedback
Original
Jian Hu, Li Tao, June Yang, and Chandler Zhou · 2023
Closest in time.
The flan collection: Designing data and methods for effective instruction tuning
Original
Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, Yi Tay, Denny Zhou, Quoc V Le, Barret Zoph, Jason Wei, et al · 2023
Closest in time.
Gpt-4 technical report
Original
OpenAI · 2023
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Original
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D Manning, and Chelsea Finn · 2023
Closest in time.
Is reinforcement learning (not) for natural language processing?: Benchmarks, baselines, and building blocks for natural language policy optimization
Rajkumar Ramamurthy, Prithviraj Ammanabrolu, Kianté Brantley, Jack Hessel, Rafet Sifa, Christian Bauckhage, Hannaneh Hajishirzi, and Yejin Choi · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Original
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Closest in time.
A survey of large language models
Original
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al · 2023
Closest in time.