Fetching the paper…
Reading the bibliography…
While large-scale unsupervised language models (LMs) learn broad world knowledge and some reasoning skills, achieving precise control of their behavior is difficult due to the completely unsupervised nature of their training.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 1901
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
R. A. Bradley and M. E. Terry · 1952
Earlier work this paper cites.
The analysis of permutations
R. L. Plackett · 1975
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Reinforcement learning by reward-weighted regression for operational space control
J. Peters and S. Schaal · 2007
Earlier work this paper cites.
Learning word vectors for sentiment analysis
A. L. Maas, R. E. Daly, P. T. Pham, D. Huang, A. Y. Ng, and C. Potts · 2011
Earlier work this paper cites.
The k-armed dueling bandits problem
Y. Yue, J. Broder, R. Kleinberg, and T. Joachims · 2011
Earlier work this paper cites.
Individual choice behavior: A theoretical analysis
R. D. Luce · 2012
Earlier work this paper cites.
Learning trajectory preferences for manipulators via iterative improvement
A. Jain, B. Wojcik, T. Joachims, and A. Saxena · 2013
Earlier work this paper cites.
Preference-based reinforcement learning: evolutionary direct policy search using a preference-based racing algorithm
R. Busa-Fekete, B. Szörényi, P. Weng, W. Cheng, and E. Hüllermeier · 2014
Earlier work this paper cites.
Contextual dueling bandits
M. Dudík, K. Hofmann, R. E. Schapire, A. Slivkins, and M. Zoghi · 2015
Earlier work this paper cites.
Sequence level training with recurrent neural networks
M. Ranzato, S. Chopra, M. Auli, and W. Zaremba · 2015
Earlier work this paper cites.
Abstractive text summarization using sequence-to-sequence RNNs and beyond
R. Nallapati, B. Zhou, C. dos Santos, Ç. Gulçehre, and B. Xiang · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Sequence tutor: Conservative fine-tuning of sequence generation models with kl-control
N. Jaques, S. Gu, D. Bahdanau, J. M. Hernández-Lobato, R. E. Turner, and D. Eck · 2017
Earlier work this paper cites.
Active preference-based learning of reward functions
D. Sadigh, A. D. Dragan, S. Sastry, and S. A. Seshia · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms, 2017
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
TL;DR: Mining Reddit to learn automatic summarization
M. Völske, M. Potthast, S. Syed, and B. Stein · 2017
Cited alongside, same era.
Reliability and learnability of human bandit feedback for sequence-to-sequence reinforcement learning
J. Kreutzer, J. Uyheng, and S. Riezler · 2018
Cited alongside, same era.
Learning Dynamic Robot-to-Human Object Handover from Human Feedback , pages 161–176
A. Kupcsik, D. Hsu, and W. S. Lee · 2018
Cited alongside, same era.
Reinforcement learning and control as probabilistic inference: Tutorial and review, 2018
S. Levine · 2018
Cited alongside, same era.
A deep reinforced model for abstractive summarization
R. Paulus, C. Xiong, and R. Socher · 2018
Cited alongside, same era.
Learning to extract coherent summary via deep reinforcement learning
Y. Wu and B. Hu · 2018
Scaling instruction-finetuned language models, 2022
H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, Y. Li, X. Wang, M. Dehghani, S. Brahma, A. Webson, S. S. Gu, Z. Dai, M. Suzgun, X. Chen, A. Chowdhery, A. Castro-Ros, M. Pellat, K. Robinson, D. Valter, S. Narang, G. Mishra, A. Yu, V. Zhao, Y. Huang, A. Dai, H. Yu, S. Petrov, E. H. Chi, J. Dean, J. Devlin, A. Roberts, D. Zhou, Q. V. Le, and J. Wei · 2022
Later among the works it cites.
On reinforcement learning and distribution matching for fine-tuning language models with no catastrophic forgetting
T. Korbak, H. Elsahar, G. Kruszewski, and M. Dymetman · 2022
Later among the works it cites.
Cross-task generalization via natural language crowdsourcing instructions
S. Mishra, D. Khashabi, C. Baral, and H. Hajishirzi · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. F. Christiano, J. Leike, and R. Lowe · 2022
Later among the works it cites.
Multitask prompted training enables zero-shot task generalization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
X. B. Peng, A. Kumar, G. Zhang, and S. Levine · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners, 2019
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever · 2019
Cited alongside, same era.
Neural text generation with unlikelihood training
S. Welleck, I. Kulikov, S. Roller, E. Dinan, K. Cho, and J. Weston · 2019
Cited alongside, same era.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Cited alongside, same era.
Human-centric dialog training via offline reinforcement learning
N. Jaques, J. H. Shen, A. Ghandeharioun, C. Ferguson, A. Lapedriza, N. Jones, S. S. Gu, and R. Picard · 2020
Cited alongside, same era.
Fine-tuning language models from human preferences, 2020
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving · 2020
Cited alongside, same era.
V. Sanh, A. Webson, C. Raffel, S. Bach, L. Sutawika, Z. Alyafeai, A. Chaffin, A. Stiegler, A. Raja, M. Dey, M. S. Bari, C. Xu, U. Thakker, S. S. Sharma, E. Szczechla, T. Kim, G. Chhablani, N. Nayak, D. Datta, J. Chang, M. T.-J. Jiang, H. Wang, M. Manica, S. Shen, Z. X. Yong, H. Pandey, R. Bawden, T. Wang, T. Neeraj, J. Rozen, A. Sharma, A. Santilli, T. Fevry, J. A. Fries, R. Teehan, T. L. Scao, S. Biderman, L. Gao, T. Wolf, and A. M. Rush · 2022
Later among the works it cites.
Learning to summarize from human feedback, 2022
N. Stiennon, L. Ouyang, J. Wu, D. M. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. Christiano · 2022
Later among the works it cites.
Lamda: Language models for dialog applications, 2022
R. Thoppilan, D. D. Freitas, J. Hall, N. Shazeer, A. Kulshreshtha, H.-T. Cheng, A. Jin, T. Bos, L. Baker, Y. Du, Y. Li, H. Lee, H. S. Zheng, A. Ghafouri, M. Menegali, Y. Huang, M. Krikun, D. Lepikhin, J. Qin, D. Chen, Y. Xu, Z. Chen, A. Roberts, M. Bosma, V. Zhao, Y. Zhou, C.-C. Chang, I. Krivokon, W. Rusch, M. Pickett, P. Srinivasan, L. Man, K. Meier-Hellstern, M. R. Morris, T. Doshi, R. D. Santos, T. Duke, J. Soraker, B. Zevenbergen, V. Prabhakaran, M. Diaz, B. Hutchinson, K. Olson, A. Molina, E. Hoffman-John, J. Lee, L. Aroyo, R. Rajakumar, A. Butryna, M. Lamm, V. Kuzmina, J. Fenton, A. Cohen, R. Bernstein, R. Kurzweil, B. Aguera-Arcas, C. Cui, M. Croak, E. Chi, and Q. Le · 2022
Later among the works it cites.
Human preferences as dueling bandits
X. Yan, C. Luo, C. L. A. Clarke, N. Craswell, E. M. Voorhees, and P. Castells · 2022
Later among the works it cites.
Pythia: A suite for analyzing large language models across training and scaling, 2023
S. Biderman, H. Schoelkopf, Q. Anthony, H. Bradley, K. O’Brien, E. Hallahan, M. A. Khan, S. Purohit, U. S. Prashanth, E. Raff, A. Skowron, L. Sutawika, and O. van der Wal · 2023
Closest in time.
Sparks of artificial general intelligence: Early experiments with GPT-4, 2023
S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg, H. Nori, H. Palangi, M. T. Ribeiro, and Y. Zhang · 2023
Closest in time.
Y. Chen, R. Wang, H. Jiang, S. Shi, and R.-L. Xu · 2023
Closest in time.
Aligning language models with preferences through f-divergence minimization
D. Go, T. Korbak, G. Kruszewski, J. Rozen, N. Ryu, and M. Dymetman · 2023
Closest in time.
Is reinforcement learning (not) for natural language processing: Benchmarks, baselines, and building blocks for natural language policy optimization
R. Ramamurthy, P. Ammanabrolu, K. Brantley, J. Hessel, R. Sifa, C. Bauckhage, H. Hajishirzi, and Y. Choi · 2023
Closest in time.
Dueling rl: Reinforcement learning with trajectory preferences
A. Saha, A. Pacchiano, and J. Lee · 2023
Closest in time.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Closest in time.
CarperAI/trlx: v0.6.0: LLaMa (Alpaca), Benchmark Util, T5 ILQL, Tests, Mar. 2023
L. von Werra, J. Tow, reciprocated, S. Matiana, A. Havrilla, cat state, L. Castricato, Alan, D. V. Phung, A. Thakur, A. Bukhtiyarov, aaronrmm, F. Milo, Daniel, D. King, D. Shin, E. Kim, J. Wei, M. Romero, N. Pochinkov, O. Sanseviero, R. Adithyan, S. Siu, T. Simonini, V. Blagojevic, X. Song, Z. Witten, alexandremuzio, and crumb · 2023
Closest in time.