Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) training is inherently unstable due to factors such as moving targets and high gradient variance.
Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 1910
Earlier work this paper cites.
Ohio supercomputer center, 1987
Center, O. S · 1987
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y · 2000
Earlier work this paper cites.
Normalized loss functions for deep learning with noisy labels
Ma, X., Huang, H., Wang, Y., Romano, S., Erfani, S. M., and Bailey, J · 2006
Earlier work this paper cites.
Learning to summarize from human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D. M., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F · 2009
Earlier work this paper cites.
Box2d, a 2d physics engine for games, 2011
Catto, E · 2011
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. A · 2013
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 2014
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Earlier work this paper cites.
Deep reinforcement learning with double q-learning, 2015
van Hasselt, H., Guez, A., and Silver, D · 2015
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T. P., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Robust loss functions under label noise for deep neural networks, 2017
Ghosh, A., Kumar, H., and Sastry, P. S · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
TL;DR: Mining Reddit to learn automatic summarization
Völske, M., Potthast, M., Syed, S., and Stein, B · 2017
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor, 2018
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation, 2018
Schulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P · 2018
Cited alongside, same era.
Reinforcement learning with perturbed rewards
Wang, J., Liu, Y., and Li, B · 2018
Cited alongside, same era.
Generalized cross entropy loss for training deep neural networks with noisy labels
Zhang, Z. and Sabuncu, M. R · 2018
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Cited alongside, same era.
Symmetric cross entropy for robust learning with noisy labels
Wang, Y., Ma, X., Chen, Z., Luo, Y., Yi, J., and Bailey, J · 2019
Cited alongside, same era.
The phenomenon of policy churn, 2022
Schaul, T., Barreto, A., Quan, J., and Ostrovski, G · 2022
Later among the works it cites.
Tril: Transformers reinforcement and imitation learning library
Chang, J. D., Brantley, K., Ramamurthy, R., Misra, D., and Sun, W · 2023
Later among the works it cites.
Palm-e: An embodied multimodal language model, 2023
Driess, D., Xia, F., Sajjadi, M. S. M., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., Yu, T., Huang, W., Chebotar, Y., Sermanet, P., Duckworth, D., Levine, S., Vanhoucke, V., Hausman, K., Toussaint, M., Greff, K., Zeng, A., Mordatch, I., and Florence, P · 2023
Later among the works it cites.
trlX: A framework for large scale reinforcement learning from human feedback
Havrilla, A., Zhuravinskyi, M., Phung, D., Tiwari, A., Tow, J., Biderman, S., Anthony, Q., and Castricato, L · 2023
Later among the works it cites.
Let’s verify step by step, 2023
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Clinicalbert: Modeling clinical notes and predicting hospital readmission, 2020
Huang, K., Altosaar, J., and Ranganath, R · 2020
Cited alongside, same era.
Rl baselines3 zoo
Raffin, A · 2020
Cited alongside, same era.
Discretizing continuous action space for on-policy optimization, 2020
Tang, Y. and Agrawal, S · 2020
Cited alongside, same era.
Trl: Transformer reinforcement learning
von Werra, L., Belkada, Y., Tunstall, L., Beeching, E., Thrush, T., Lambert, N., and Huang, S · 2020
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., and Chen, W · 2021
Cited alongside, same era.
Stable-baselines3: Reliable reinforcement learning implementations
Raffin, A., Hill, A., Gleave, A., Kanervisto, A., Ernestus, M., and Dormann, N · 2021
Cited alongside, same era.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Wang, B. and Komatsuzaki, A · 2021
Cited alongside, same era.
Small batch deep reinforcement learning, 2023
Obando-Ceron, J., Bellemare, M. G., and Castro, P. S · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model, 2023
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2023
Later among the works it cites.
Is reinforcement learning (not) for natural language processing: Benchmarks, baselines, and building blocks for natural language policy optimization, 2023
Ramamurthy, R., Ammanabrolu, P., Brantley, K., Hessel, J., Sifa, R., Bauckhage, C., Hajishirzi, H., and Choi, Y · 2023
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models, 2023
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q., and Zhou, D · 2023
Later among the works it cites.
Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms, 2024
Ahmadian, A., Cremer, C., Gallé, M., Fadaee, M., Kreutzer, J., Pietquin, O., Üstün, A., and Hooker, S · 2024
Closest in time.
Normality-guided distributional reinforcement learning for continuous control, 2024
Byun, J.-S. and Perrault, A · 2024
Closest in time.
Byun, J.-S., Chun, J., Kil, J., and Perrault, A · 2024
Closest in time.
Maxmin-rlhf: Towards equitable alignment of large language models with diverse human preferences, 2024
Chakraborty, S., Qiu, J., Yuan, H., Koppel, A., Huang, F., Manocha, D., Bedi, A. S., and Wang, M · 2024
Closest in time.
The distributional reward critic architecture for perturbed-reward reinforcement learning, 2024
Chen, X., Zhu, Z., and Perrault, A · 2024
Closest in time.
Kto: Model alignment as prospect theoretic optimization, 2024
Ethayarajh, K., Xu, W., Muennighoff, N., Jurafsky, D., and Kiela, D · 2024
Closest in time.
OpenAI · 2024
Closest in time.
Math-shepherd: Verify and reinforce llms step-by-step without human annotations, 2024
Wang, P., Li, L., Shao, Z., Xu, R. X., Dai, D., Li, Y., Chen, D., Wu, Y., and Sui, Z · 2024
Closest in time.