Fetching the paper…
Reading the bibliography…
Aligning language models (LMs) with preferences is an important problem in natural language generation.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi · 1904
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 1907
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry · 1952
Earlier work this paper cites.
The analysis of permutations
Robin L Plackett · 1975
Earlier work this paper cites.
Likelihood ratio gradient estimation for stochastic systems
Peter W Glynn · 1990
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Support vector learning for ordinal regression
Ralf Herbrich, Thore Graepel, and Klaus Obermayer · 1999
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, and Pascal Vincent · 2000
Earlier work this paper cites.
Exploiting cloze questions for few shot text classification and natural language inference
Timo Schick and Hinrich Schütze · 2001
Earlier work this paper cites.
An efficient boosting algorithm for combining preferences
Yoav Freund, Raj Iyer, Robert E Schapire, and Yoram Singer · 2003
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
Learning to rank using gradient descent
Chris Burges, Tal Shaked, Erin Renshaw, Ari Lazier, Matt Deeds, Nicole Hamilton, and Greg Hullender · 2005
Earlier work this paper cites.
Gradient estimation
Michael C Fu · 2006
Earlier work this paper cites.
Listwise approach to learning to rank: theory and algorithm
Fen Xia, Tie-Yan Liu, Jue Wang, Wensheng Zhang, and Hang Li · 2008
Earlier work this paper cites.
It’s not just size that matters: Small language models are also few-shot learners
Timo Schick and Hinrich Schütze · 2009
Earlier work this paper cites.
Individual choice behavior: A theoretical analysis
R Duncan Luce · 2012
Earlier work this paper cites.
Framework of automatic text summarization using reinforcement learning
Seonggi Ryang and Takeshi Abekawa · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Teaching machines to read and comprehend
Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom · 2015
Earlier work this paper cites.
Sequence level training with recurrent neural networks
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun · 2015
Earlier work this paper cites.
Reward augmented maximum likelihood for neural structured prediction
Mohammad Norouzi, Samy Bengio, Navdeep Jaitly, Mike Schuster, Yonghui Wu, Dale Schuurmans, et al · 2016
Earlier work this paper cites.
Hindsight experience replay
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, OpenAI Pieter Abbeel, and Wojciech Zaremba · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Reinforcement Learning with Deep Energy-Based Policies
Tuomas Haarnoja, Haoran Tang, P. Abbeel, and Sergey Levine · 2017
Earlier work this paper cites.
Adversarial ranking for language generation
Kevin Lin, Dianqi Li, Xiaodong He, Zhengyou Zhang, and Ming-Ting Sun · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
A deep reinforced model for abstractive summarization
Romain Paulus, Caiming Xiong, and Richard Socher · 2017
Earlier work this paper cites.
Self-critical sequence training for image captioning
Steven J Rennie, Etienne Marcheret, Youssef Mroueh, Jerret Ross, and Vaibhava Goel · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Seqgan: Sequence generative adversarial nets with policy gradient
Lantao Yu, Weinan Zhang, Jun Wang, and Yong Yu · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Long text generation via adversarial training with leaked information
Jiaxian Guo, Sidi Lu, Han Cai, Weinan Zhang, Yong Yu, and Jun Wang · 2018
Cited alongside, same era.
Approaching neural grammatical error correction as a low-resource machine translation task
Marcin Junczys-Dowmunt, Roman Grundkiewicz, Shubha Guha, and Kenneth Heafield · 2018
Cited alongside, same era.
Shashi Narayan, Shay B Cohen, and Mirella Lapata · 2018
Cited alongside, same era.
Improving Language Understanding by Generative Pre-Training
Alec Radford and Ilya Sutskever · 2018
Cross-task generalization via natural language crowdsourcing instructions
Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi · 2021
Later among the works it cites.
True few-shot learning with language models
Ethan Perez, Douwe Kiela, and Kyunghyun Cho · 2021
Later among the works it cites.
Learning how to ask: Querying lms with mixtures of soft prompts
Guanghui Qin and Jason Eisner · 2021
Later among the works it cites.
Causal-aware safe policy improvement for task-oriented dialogue
Govardana Sachithanandam Ramachandran, Kazuma Hashimoto, and Caiming Xiong · 2021
Later among the works it cites.
Multitask prompted training enables zero-shot task generalization
Victor Sanh, Albert Webson, Colin Raffel, Stephen H Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, et al · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Toward diverse text generation with inverse reinforcement learning
Zhan Shi, Xinchi Chen, Xipeng Qiu, and Xuanjing Huang · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Unsupervised text style transfer using language models as discriminators
Zichao Yang, Zhiting Hu, Chris Dyer, Eric P Xing, and Taylor Berg-Kirkpatrick · 2018
Cited alongside, same era.
Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations
Daniel Brown, Wonjoon Goo, Prabhat Nagarajan, and Scott Niekum · 2019
Cited alongside, same era.
Learning from dialogue after deployment: Feed yourself, chatbot!
Braden Hancock, Antoine Bordes, Pierre-Emmanuel Mazare, and Jason Weston · 2019
Cited alongside, same era.
Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog
Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, Craig Ferguson, À. Lapedriza, Noah J. Jones, S. Gu, and Rosalind W. Picard · 2019
Cited alongside, same era.
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer · 2019
Cited alongside, same era.
Later among the works it cites.
Reward optimization for neural machine translation with learned metrics
Raphael Shu, Kang Min Yoo, and Jung-Woo Ha · 2021
Later among the works it cites.
Process for adapting language models to society (PALMS) with values-targeted datasets
Irene Solaiman and Christy Dennison · 2021
Later among the works it cites.
On transferability of prompt tuning for natural language understanding
Yusheng Su, Xiaozhi Wang, Yujia Qin, Chi-Min Chan, Yankai Lin, Zhiyuan Liu, Peng Li, Juanzi Li, Lei Hou, Maosong Sun, et al · 2021
Later among the works it cites.
Recursively summarizing books with human feedback
Jeff Wu, Long Ouyang, Daniel M Ziegler, Nisan Stiennon, Ryan Lowe, Jan Leike, and Paul Christiano · 2021
Later among the works it cites.
Factual probing is [mask]: Learning vs. learning to recall
Zexuan Zhong, Dan Friedman, and Danqi Chen · 2021
Later among the works it cites.
Input-tuning: Adapting unfamiliar inputs to frozen pretrained models
Shengnan An, Yifei Li, Zeqi Lin, Qian Liu, Bei Chen, Qiang Fu, Weizhu Chen, Nanning Zheng, and Jian-Guang Lou · 2022
Later among the works it cites.
Robust preference learning for storytelling via contrastive reinforcement learning
Louis Castricato, Alexander Havrilla, Shahbuland Matiana, Michael Pieler, Anbang Ye, Ian Yang, Spencer Frazier, and Mark Riedl · 2022
Later among the works it cites.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al · 2022
Later among the works it cites.
Rlprompt: Optimizing discrete text prompts with reinforcement learning
Mingkai Deng, Jianyu Wang, Cheng-Ping Hsieh, Yihan Wang, Han Guo, Tianmin Shu, Meng Song, Eric P Xing, and Zhiting Hu · 2022
Later among the works it cites.
Black-box prompt learning for pre-trained language models
Shizhe Diao, Xuechun Li, Yong Lin, Zhichao Huang, and Tong Zhang · 2022
Later among the works it cites.
Efficient (soft) q-learning for text generation with limited good data
Han Guo, Bowen Tan, Zhengzhong Liu, Eric Xing, and Zhiting Hu · 2022
Later among the works it cites.
Tomasz Korbak, Hady Elsahar, Germán Kruszewski, and Marc Dymetman · 2022
Later among the works it cites.
Coderl: Mastering code generation through pretrained models and deep reinforcement learning
Hung Le, Yue Wang, Akhilesh Deepak Gotmare, Silvio Savarese, and Steven Chu Hong Hoi · 2022
Later among the works it cites.
Quark: Controllable text generation with reinforced unlearning
Ximing Lu, Sean Welleck, Liwei Jiang, Jack Hessel, Lianhui Qin, Peter West, Prithviraj Ammanabrolu, and Yejin Choi · 2022
Later among the works it cites.
Teaching language models to support answers with verified quotes
Jacob Menick, Maja Trebacz, Vladimir Mikulik, John Aslanides, Francis Song, Martin Chadwick, Mia Glaese, Susannah Young, Lucy Campbell-Gillingham, Geoffrey Irving, et al · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
Reward gaming in conditional text generation
Yuanzhe Richard Pang, Vishakh Padmakumar, Thibault Sellam, Ankur P Parikh, and He He · 2022
Later among the works it cites.
Grips: Gradient-free, edit-based instruction search for prompting large language models
Archiki Prasad, Peter Hase, Xiang Zhou, and Mohit Bansal · 2022
Later among the works it cites.
Rajkumar Ramamurthy, Prithviraj Ammanabrolu, Kianté Brantley, Jack Hessel, Rafet Sifa, Christian Bauckhage, Hannaneh Hajishirzi, and Yejin Choi · 2022
Later among the works it cites.
Multitask prompted training enables zero-shot task generalization
Victor Sanh, Albert Webson, Colin Raffel, Stephen Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Teven Le Scao, Stella Biderman, Leo Gao, Thomas Wolf, and Alexander M Rush · 2022
Later among the works it cites.
Training language models with language feedback, 2022
Jérémy Scheurer, Jon Ander Campos, Jun Shern Chan, Angelica Chen, Kyunghyun Cho, and Ethan Perez · 2022
Later among the works it cites.
Offline rl for natural language generation with implicit language q learning
Charlie Snell, Ilya Kostrikov, Yi Su, Mengjiao Yang, and Sergey Levine · 2022
Later among the works it cites.
Black-box tuning for language-model-as-a-service
Tianxiang Sun, Yunfan Shao, Hong Qian, Xuanjing Huang, and Xipeng Qiu · 2022
Later among the works it cites.
Fantastic rewards and how to tame them: A case study on reward learning for task-oriented dialogue systems
Yihao Feng*, Shentao Yang*, Shujian Zhang, Jianguo Zhang, Caiming Xiong, Mingyuan Zhou, and Huan Wang · 2023
Closest in time.
Aligning language models with preferences through f-divergence minimization
Dongyoung Go, Tomasz Korbak, Germán Kruszewski, Jos Rozen, Nahyeon Ryu, and Marc Dymetman · 2023
Closest in time.
Preference transformer: Modeling human preferences using transformers for RL
Changyeon Kim, Jongjin Park, Jinwoo Shin, Honglak Lee, Pieter Abbeel, and Kimin Lee · 2023
Closest in time.
Pretraining language models with human preferences
Tomasz Korbak, Kejian Shi, Angelica Chen, Rasika Bhalerao, Christopher L Buckley, Jason Phang, Samuel R Bowman, and Ethan Perez · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
A dense reward view on aligning text-to-image diffusion with preference
Shentao Yang, Tianqi Chen, and Mingyuan Zhou · 2024
Closest in time.
Segmenting text and learning their rewards for improved rlhf in language model, 2025
Yueqin Yin, Shentao Yang, Yujia Xie, Ziyi Yang, Yuting Sun, Hany Awadalla, Weizhu Chen, and Mingyuan Zhou · 2025
Closest in time.