Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A., M. E. Terry · 1952
Earlier work this paper cites.
Training deep neural networks on noisy labels with bootstrapping
Original
Reed, S., H. Lee, D. Anguelov, et al · 2014
Earlier work this paper cites.
Proximal policy optimization algorithms
Original
Schulman, J., F. Wolski, P. Dhariwal, et al · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., J. Leike, T. Brown, et al · 2017
Earlier work this paper cites.
Tl; dr: Mining reddit to learn automatic summarization
Völske, M., M. Potthast, S. Syed, et al · 2017
Earlier work this paper cites.
Scalable agent alignment via reward modeling: a research direction, 2018
Leike, J., D. Krueger, T. Everitt, et al · 2018
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation, 2018
Schulman, J., P. Moritz, S. Levine, et al · 2018
Earlier work this paper cites.
Fine-tuning language models from human preferences
Original
Ziegler, D. M., N. Stiennon, J. Wu, et al · 2019
Earlier work this paper cites.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
Original
Jaques, N., A. Ghandeharioun, J. H. Shen, et al · 2019
Earlier work this paper cites.
When does label smoothing help?
Müller, R., S. Kornblith, G. E. Hinton · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
Original
Ziegler, D. M., N. Stiennon, J. Wu, et al · 2019
Earlier work this paper cites.
Learning to summarize from human feedback
Original
Stiennon, N., L. Ouyang, J. Wu, et al · 2020
Earlier work this paper cites.
Unsupervised learning of visual features by contrasting cluster assignments
Caron, M., I. Misra, J. Mairal, et al · 2020
Earlier work this paper cites.
The curious case of neural text degeneration, 2020
Holtzman, A., J. Buys, L. Du, et al · 2020
Earlier work this paper cites.
Alignment of language agents
Original
Kenton, Z., T. Everitt, L. Weidinger, et al · 2021
Earlier work this paper cites.