Fetching the paper…
Reading the bibliography…
Reinforcement learning from human feedback (RLHF) can improve the quality of large language model's (LLM) outputs by aligning them with human preferences.
A Markovian decision process
R. Bellman · 1957
Earlier work this paper cites.
Probability of error of some adaptive pattern-recognition machines
H. Scudder · 1965
Earlier work this paper cites.
ALVINN: An autonomous land vehicle in a neural network
D. A. Pomerleau · 1989
Earlier work this paper cites.
Unsupervised word sense disambiguation rivaling supervised methods
D. Yarowsky · 1995
Earlier work this paper cites.
Bootstrapping pos-taggers using unlabelled data
S. Clark, J. R. Curran, and M. Osborne · 2003
Earlier work this paper cites.
Supervised seeded iterated learning for interactive language learning
Y. Lu, S. Singhal, F. Strub, O. Pietquin, and A. Courville · 2010
Earlier work this paper cites.
Batch reinforcement learning
S. Lange, T. Gabel, and M. Riedmiller · 2012
Earlier work this paper cites.
Report on the 11th iwslt evaluation campaign
M. Cettolo, J. Niehues, S. Stüker, L. Bentivogli, and M. Federico · 2014
Earlier work this paper cites.
Iterated learning and the evolution of language
S. Kirby, T. Griffiths, and K. Smith · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Harley, T. P. Lillicrap, D. Silver, and K. Kavukcuoglu · 2016
Earlier work this paper cites.
Exploiting source-side monolingual data in neural machine translation
J. Zhang and C. Zong · 2016
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
L. Espeholt, H. Soyer, R. Munos, K. Simonyan, V. Mnih, T. Ward, Y. Doron, V. Firoiu, T. Harley, I. Dunning, S. Legg, and K. Kavukcuoglu · 2018
Earlier work this paper cites.
Subword regularization: Improving neural network translation models with multiple subword candidates
T. Kudo · 2018
Earlier work this paper cites.
Self-imitation learning
J. Oh, Y. Guo, S. Singh, and H. Lee · 2018
Earlier work this paper cites.
Revisiting self-training for neural sequence generation
J. He, J. Gu, J. Shen, and M. Ranzato · 2019
Earlier work this paper cites.
The DeepMind JAX Ecosystem, 2020
I. Babuschkin, K. Baumli, A. Bell, S. Bhupatiraju, J. Bruce, P. Buchlovsky, D. Budden, T. Cai, A. Clark, I. Danihelka, C. Fantacci, J. Godwin, C. Jones, R. Hemsley, T. Hennigan, M. Hessel, S. Hou, S. Kapturowski, T. Keck, I. Kemaev, M. King, M. Kunesch, L. Martens, H. Merzic, V. Mikulik, T. Norman, J. Quan, G. Papamakarios, R. Ring, F. Ruiz, A. Sanchez, R. Schneider, E. Sezener, S. Spencer, S. Srinivasan, L. Wang, W. Stokowiec, and F. Viola · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
D4RL: Datasets for deep data-driven reinforcement learning
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine · 2020
Earlier work this paper cites.
Findings of the wmt 2020 shared task on parallel corpus filtering and alignment
P. Koehn, V. Chaudhary, A. El-Kishky, N. Goyal, P.-J. Chen, and F. Guzmán · 2020
Cited alongside, same era.
COMET: A neural framework for MT evaluation
R. Rei, C. Stewart, A. C. Farinha, and A. Lavie · 2020
Cited alongside, same era.
BLEURT: Learning robust metrics for text generation
T. Sellam, D. Das, and A. Parikh · 2020
Cited alongside, same era.
V-mpo: On-policy maximum a posteriori policy optimization for discrete and continuous control
H. F. Song, A. Abdolmaleki, J. T. Springenberg, A. Clark, H. Soyer, J. W. Rae, S. Noury, A. Ahuja, S. Liu, D. Tirumala, N. Heess, D. Belov, M. Riedmiller, and M. M. Botvinick · 2020
Cited alongside, same era.
Learning to summarize with human feedback
N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano · 2020
Cited alongside, same era.
Self-training with noisy student improves imagenet classification
Launchpad: A programming model for distributed machine learning research
F. Yang, G. Barth-Maron, P. Stańczyk, M. Hoffman, S. Liu, M. Kroiss, A. Pope, and A. Rrustemi · 2021
Later among the works it cites.
Beyond tabula rasa: Reincarnating reinforcement learning
R. Agarwal, M. Schwarzer, P. S. Castro, A. Courville, and M. G. Bellemare · 2022
Later among the works it cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, et al · 2022
Later among the works it cites.
When does return-conditioned supervised learning work for offline reinforcement learning?
D. Brandfonbrener, A. Bietti, J. Buckman, R. Laroche, and J. Bruna · 2022
Later among the works it cites.
Mad for robust reinforcement learning in machine translation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Q. Xie, M.-T. Luong, E. Hovy, and Q. V. Le · 2020
Cited alongside, same era.
The DeepMind chinese–english document translation system at wmt2020
L. Yu, L. Sartran, P.-S. Huang, W. Stokoweic, D. Donato, S. Srinivasan, A. Andreev, W. Ling, S. Mokra, A. D. Lago, Y. Doron, S. Young, P. Blunsom, and C. Dyer · 2020
Cited alongside, same era.
On multi-objective policy optimization as a tool for reinforcement learning
A. Abdolmaleki, S. Huang, G. Vezzani, B. Shahriari, J. T. Springenberg, S. Mishra, D. Tirumala, A. Byravan, K. Bousmalis, A. György, et al · 2021
Cited alongside, same era.
Expert iteration
T. W. Anthony · 2021
Cited alongside, same era.
Decision transformer: Reinforcement learning via sequence modeling
L. Chen, K. Lu, A. Rajeswaran, K. Lee, A. Grover, M. Laskin, P. Abbeel, A. Srinivas, and I. Mordatch · 2021
Cited alongside, same era.
Results of the wmt21 metrics shared task: Evaluating metrics with expert-based human evaluations on ted and news domain
M. Freitag, R. Rei, N. Mathur, C.-k. Lo, C. Stewart, G. Foster, A. Lavie, and O. Bojar · 2021
Cited alongside, same era.
Scaling laws for neural machine translation
B. Ghorbani, O. Firat, M. Freitag, A. Bapna, M. Krikun, X. Garcia, C. Chelba, and C. Cherry · 2021
Cited alongside, same era.
D. Donato, L. Yu, W. Ling, and C. Dyer · 2022
Later among the works it cites.
Results of WMT22 metrics shared task: Stop using BLEU – neural metrics are better and more robust
M. Freitag, R. Rei, N. Mathur, C.-k. Lo, C. Stewart, E. Avramidis, T. Kocmi, G. Foster, A. Lavie, and A. F. T. Martins · 2022
Later among the works it cites.
Scaling laws for reward model overoptimization
L. Gao, J. Schulman, and J. Hilton · 2022
Later among the works it cites.
Improving alignment of dialogue agents via targeted human judgements
A. Glaese, N. McAleese, M. Trębacz, J. Aslanides, V. Firoiu, T. Ewalds, M. Rauh, L. Weidinger, M. Chadwick, P. Thacker, et al · 2022
Later among the works it cites.
An empirical analysis of compute-optimal large language model training
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, et al · 2022
Later among the works it cites.
Competition-level code generation with alphacode
Y. Li, D. Choi, J. Chung, N. Kushman, J. Schrittwieser, R. Leblond, T. Eccles, J. Keeling, F. Gimeno, A. Dal Lago, et al · 2022
Later among the works it cites.
Red teaming language models with language models
E. Perez, S. Huang, F. Song, T. Cai, R. Ring, J. Aslanides, A. Glaese, N. McAleese, and G. Irving · 2022
Later among the works it cites.
Defining and characterizing reward hacking
J. Skalse, N. H. R. Howe, D. Krasheninnikov, and D. Krueger · 2022
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
A. Srivastava, A. Rastogi, A. Rao, A. A. M. Shoeb, A. Abid, A. Fisch, A. R. Brown, A. Santoro, A. Gupta, A. Garriga-Alonso, et al · 2022
Later among the works it cites.
Solving math word problems with process-and outcome-based feedback
J. Uesato, N. Kushman, R. Kumar, F. Song, N. Siegel, L. Wang, A. Creswell, G. Irving, and I. Higgins · 2022
Later among the works it cites.
Star: Bootstrapping reasoning with reasoning
E. Zelikman, Y. Wu, J. Mu, and N. Goodman · 2022
Later among the works it cites.
Sparks of artificial general intelligence: Early experiments with GPT-4
S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg, et al · 2023
Closest in time.
Raft: Reward ranked finetuning for generative foundation model alignment
H. Dong, W. Xiong, D. Goyal, R. Pan, S. Diao, J. Zhang, K. Shum, and T. Zhang · 2023
Closest in time.
J. Jung, P. West, L. Jiang, F. Brahman, X. Lu, J. Fisher, T. Sorensen, and Y. Choi · 2023
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning, and C. Finn · 2023
Closest in time.
The curse of recursion: Training on generated data makes models forget
I. Shumailov, Z. Shumaylov, Y. Zhao, Y. Gal, N. Papernot, and R. Anderson · 2023
Closest in time.