Fetching the paper…
Reading the bibliography…
KL-regularized reinforcement learning (RL) is a popular alignment framework to control the language model responses towards high reward outcomes.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
Karimi, H., Nutini, J., and Schmidt, M · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Learning to decode for future success
Li, J., Monroe, W., and Jurafsky, D · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., Van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., and Silver, D · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Earlier work this paper cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Peng, X. B., Kumar, A., Zhang, G., and Levine, S · 2019
Earlier work this paper cites.
Learning to summarize with human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F · 2020
Earlier work this paper cites.
Neural text generation with unlikelihood training
Welleck, S., Kulikov, I., Roller, S., Dinan, E., Cho, K., and Weston, J · 2020
Earlier work this paper cites.
GeDi: Generative discriminator guided sequence generation
Krause, B., Gotmare, A. D., McCann, B., Keskar, N. S., Joty, S., Socher, R., and Rajani, N. F · 2021
Earlier work this paper cites.
WebGPT: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al · 2021
Earlier work this paper cites.
To beam or not to beam: That is a question of cooperation for language gans
Scialom, T., Dray, P.-A., Staiano, J., Lamprier, S., and Piwowarski, B · 2021
Earlier work this paper cites.
FUDGE: Controlled text generation with future discriminators
Yang, K. and Klein, D · 2021
Earlier work this paper cites.
The cringe loss: Learning what language not to model
Adolphs, L., Gao, T., Xu, J., Shuster, K., Sukhbaatar, S., and Weston, J · 2022
Cited alongside, same era.
Director: Generator-classifiers for supervised language modeling
Arora, K., Shuster, K., Sukhbaatar, S., and Weston, J · 2022
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al · 2022
Cited alongside, same era.
PPL-MCTS: Constrained textual generation through discriminator-guided MCTS decoding
Chaffin, A., Claveau, V., and Kijak, E · 2022
Cited alongside, same era.
Improving alignment of dialogue agents via targeted human judgements
A general theoretical paradigm to understand learning from human preferences
Azar, M. G., Rowland, M., Piot, B., Guo, D., Calandriello, D., Valko, M., and Munos, R · 2023
Closest in time.
Accelerating large language model decoding with speculative sampling
Chen, C., Borgeaud, S., Irving, G., Lespiau, J.-B., Sifre, L., and Jumper, J · 2023
Closest in time.
Helping or herding? reward model ensembles mitigate but do not eliminate reward hacking
Eisenstein, J., Nagpal, C., Agarwal, A., Beirami, A., D’Amour, A., Dvijotham, D., Fisch, A., Heller, K., Pfohl, S., Ramachandran, D., et al · 2023
Closest in time.
Scaling laws for reward model overoptimization
Gao, L., Schulman, J., and Hilton, J · 2023
Closest in time.
Critic-guided decoding for controlled text generation
Kim, M., Lee, H., Yoo, K. M., Park, J., Lee, H., and Jung, K · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Glaese, A., McAleese, N., Trebacz, M., Aslanides, J., Firoiu, V., Ewalds, T., Rauh, M., Weidinger, L., Chadwick, M., Thacker, P., et al · 2022
Cited alongside, same era.
RL with KL penalties is better viewed as Bayesian inference
Korbak, T., Perez, E., and Buckley, C · 2022
Cited alongside, same era.
NeuroLogic a*esque decoding: Constrained text generation with lookahead heuristics
Lu, X., Welleck, S., West, P., Jiang, L., Kasai, J., Khashabi, D., Le Bras, R., Qin, L., Yu, Y., Zellers, R., Smith, N. A., and Choi, Y · 2022
Cited alongside, same era.
Controllable text generation with neurally-decomposed oracle
Meng, T., Lu, S., Peng, N., and Chang, K.-W · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
COLD decoding: Energy-based constrained text generation with langevin dynamics
Qin, L., Welleck, S., Khashabi, D., and Choi, Y · 2022
Cited alongside, same era.
Convergent and efficient deep Q network algorithm
Wang, Z. T. and Ueda, M · 2022
Cited alongside, same era.
Discup: Discriminator cooperative unlikelihood prompt-tuning for controllable text generation
Zhang, H. and Song, D · 2022
Cited alongside, same era.
Closest in time.
Fast inference from transformers via speculative decoding
Leviathan, Y., Kalman, M., and Matias, Y · 2023
Closest in time.
DSTC8 Reddit Corpus
Microsoft · 2023
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2023
Closest in time.
Offline rl for natural language generation with implicit language q learning
Snell, C. V., Kostrikov, I., Su, Y., Yang, S., and Levine, S · 2023
Closest in time.
SpecTr: Fast speculative decoding via optimal transport
Sun, Z., Suresh, A. T., Ro, J. H., Beirami, A., Jain, H., and Yu, F · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Closest in time.
Theoretical guarantees on the best-of-n alignment policy
Beirami, A., Agarwal, A., Berant, J., D’Amour, A., Eisenstein, J., Nagpal, C., and Suresh, A. T · 2024
Closest in time.
WARM: On the benefits of weight averaged reward models
Ramé, A., Vieillard, N., Hussenot, L., Dadashi, R., Cideron, G., Bachem, O., and Ferret, J · 2024
Closest in time.
Asymptotics of language model alignment
Yang, J. Q., Salamatian, S., Sun, Z., Suresh, A. T., and Beirami, A · 2024
Closest in time.