Fetching the paper…
Reading the bibliography…
Reinforcement Learning (RL) from Human Preference-based feedback is a popular paradigm for fine-tuning generative models, which has produced impressive models such as GPT-4 and Claude3 Opus.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G. (2019) · 1909
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E. (1952) · 1952
Earlier work this paper cites.
A natural policy gradient
Kakade, S. M. (2001) · 2001
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J. (2002) · 2002
Earlier work this paper cites.
Covariant policy search
Bagnell, J. A. and Schneider, J. (2003) · 2003
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Kakade, S. M. (2003) · 2003
Earlier work this paper cites.
Learning decisions: Robustness, uncertainty, and approximation
Bagnell, J. A. (2004) · 2004
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y. (2004) · 2004
Earlier work this paper cites.
Learning as search optimization: Approximate large margin methods for structured prediction
Daumé III, H. and Marcu, D. (2005) · 2005
Earlier work this paper cites.
Search-based structured prediction
Daumé, H., Langford, J., and Marcu, D. (2009) · 2009
Earlier work this paper cites.
Reinforcement learning with a near optimal rate of convergence
Azar, M. G., Munos, R., Ghavamzadeh, M., and Kappen, H. (2011) · 2011
Earlier work this paper cites.
The k-armed dueling bandits problem
Yue, Y., Broder, J., Kleinberg, R., and Joachims, T. (2012) · 2012
Earlier work this paper cites.
Relative confidence sampling for efficient on-line ranker evaluation
Zoghi, M., Whiteson, S. A., De Rijke, M., and Munos, R. (2014) · 2014
Earlier work this paper cites.
Contextual dueling bandits
Dudík, M., Hofmann, K., Schapire, R. E., Slivkins, A., and Zoghi, M. (2015) · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P. (2015) · 2015
Earlier work this paper cites.
Contextual decision processes with low bellman rank are pac-learnable
Jiang, N., Krishnamurthy, A., Agarwal, A., Langford, J., and Schapire, R. E. (2016) · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al. (2016) · 2016
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., and Munos, R. (2017) · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D. (2017) · 2017
Earlier work this paper cites.
Interactive learning from policy-dependent human feedback
MacGlashan, J., Ho, M. K., Loftin, R., Peng, B., Wang, G., Roberts, D. L., Taylor, M. E., and Littman, M. L. (2017) · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017) · 2017
Cited alongside, same era.
Get to the point: Summarization with pointer-generator networks
See, A., Liu, P. J., and Manning, C. D. (2017) · 2017
Cited alongside, same era.
A survey of preference-based reinforcement learning methods
Wirth, C., Akrour, R., Neumann, G., Fürnkranz, J., et al. (2017) · 2017
Cited alongside, same era.
Overcoming exploration in reinforcement learning with demonstrations
Nair, A., McGrew, B., Andrychowicz, M., Zaremba, W., and Abbeel, P. (2018) · 2018
Cited alongside, same era.
Learning montezuma’s revenge from a single demonstration
Salimans, T. and Chen, R. (2018) · 2018
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. (2022) · 2022
Later among the works it cites.
Ramamurthy, R., Ammanabrolu, P., Brantley, K., Hessel, J., Sifa, R., Bauckhage, C., Hajishirzi, H., and Choi, Y. (2022) · 2022
Later among the works it cites.
Hybrid rl: Using both offline and online data can make rl efficient
Song, Y., Zhou, Y., Sekhari, A., Bagnell, J. A., Krishnamurthy, A., and Sun, W. (2022) · 2022
Later among the works it cites.
Efficient local planning with linear function approximation
Yin, D., Hao, B., Abbasi-Yadkori, Y., Lazić, N., and Szepesvári, C. (2022) · 2022
Later among the works it cites.
Offline reinforcement learning with realizability and single-policy concentrability
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep tamer: Interactive agent shaping in high-dimensional state spaces
Warnell, G., Waytowich, N., Lawhern, V., and Stone, P. (2018) · 2018
Cited alongside, same era.
Reinforcement learning: Theory and algorithms
Agarwal, A., Jiang, N., Kakade, S. M., and Sun, W. (2019) · 2019
Cited alongside, same era.
Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations
Brown, D., Goo, W., Nagarajan, P., and Niekum, S. (2019) · 2019
Cited alongside, same era.
Provably efficient reinforcement learning with linear function approximation
Jin, C., Yang, Z., Wang, Z., and Jordan, M. I. (2020) · 2020
Cited alongside, same era.
Dueling posterior sampling for preference-based reinforcement learning
Novoseller, E., Wei, Y., Sui, Y., Yue, Y., and Burdick, J. (2020) · 2020
Cited alongside, same era.
Learning to summarize with human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F. (2020) · 2020
Cited alongside, same era.
Preference-based reinforcement learning with finite-time guarantees
Xu, Y., Wang, R., Yang, L., Singh, A., and Dubrawski, A. (2020) · 2020
Cited alongside, same era.
Zhan, W., Huang, B., Huang, A., Jiang, N., and Lee, J. (2022) · 2022
Later among the works it cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al. (2023) · 2023
Later among the works it cites.
Efficient online reinforcement learning with offline data
Ball, P. J., Smith, L., Kostrikov, I., and Levine, S. (2023) · 2023
Later among the works it cites.
Pythia: A suite for analyzing large language models across training and scaling
Biderman, S., Schoelkopf, H., Anthony, Q. G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., et al. (2023) · 2023
Later among the works it cites.
Learning to generate better than your llm
Chang, J. D., Brantley, K., Ramamurthy, R., Misra, D., and Sun, W. (2023) · 2023
Later among the works it cites.
https://contextual.ai/better-cheaper-faster-llm-alignment-with-kto/
Contextual.ai (2023) · 2023
Later among the works it cites.
Aligning text-to-image models using human feedback
Lee, K., Liu, H., Ryu, M., Watkins, O., Du, Y., Boutilier, C., Abbeel, P., Ghavamzadeh, M., and Gu, S. S. (2023) · 2023
Later among the works it cites.
Reinforcement learning with human feedback: Learning dynamic choices via pessimism
Li, Z., Yang, Z., and Wang, M. (2023) · 2023
Later among the works it cites.
Languages are rewards: Hindsight finetuning using human feedback
Liu, H., Sferrazza, C., and Abbeel, P. (2023) · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C. (2023) · 2023
Later among the works it cites.
Benchmarks and algorithms for offline preference-based reward learning
Shin, D., Dragan, A. D., and Brown, D. S. (2023) · 2023
Later among the works it cites.
Demonstration-regularized rl
Tiapkin, D., Belomestny, D., Calandriello, D., Moulines, E., Naumov, A., Perrault, P., Valko, M., and Menard, P. (2023) · 2023
Later among the works it cites.
Jump-start reinforcement learning
Uchendu, I., Xiao, T., Lu, Y., Zhu, B., Yan, M., Simon, J., Bennice, M., Fu, C., Ma, C., Jiao, J., et al. (2023) · 2023
Later among the works it cites.
Making rl with preference-based feedback efficient via randomization
Wu, R. and Sun, W. (2023) · 2023
Later among the works it cites.
Pairwise proximal policy optimization: Harnessing relative feedback for llm alignment
Wu, T., Zhu, B., Zhang, R., Wen, Z., Ramchandran, K., and Jiao, J. (2023) · 2023
Later among the works it cites.
Self-rewarding language models
Yuan, W., Pang, R. Y., Cho, K., Sukhbaatar, S., Xu, J., and Weston, J. (2024) · 2024
Closest in time.