Fetching the paper…
Reading the bibliography…
Specifying rewards for reinforcement learned (RL) agents is challenging.
Bradley, R. A., & Terry, M. E. (1952). Rank analysis of incomplete block designs: I. The method of paired comparisons. Biometrika, 39(3/4), 324-345
1952
Earlier work this paper cites.
Watkins, C. J., & Dayan, P. (1992). Q-learning. Machine learning, 8(3), 279-292
1992
Earlier work this paper cites.
Rossi, F., Van Beek, P., & Walsh, T. (Eds.). (2006). Handbook of constraint programming. Elsevier
2006
Earlier work this paper cites.
Christiano, Paul F., Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. ”Ddeep reinforcement learning from human preferences.” Advances in neural information processing systems 30 (2017)
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Tang, H., Houthooft, R., Foote, D., Stooke, A., Xi Chen, O., Duan, Y., … & Abbeel, P. (2017). # exploration: A study of count-based exploration for deep reinforcement learning. Advances in neural information processing systems, 30
2017
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., … & Polosukhin, I. (2017). Attention is all you need. Advances in neural information processing systems, 30
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Ha, D., & Schmidhuber, J. (2018). Recurrent world models facilitate policy evolution. Advances in neural information processing systems, 31
2018
Cited alongside, same era.
Ferret, J., Marinier, R., Geist, M., & Pietquin, O. (2019). Self-attentional credit assignment for transfer in reinforcement learning. Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence (IJCAI-20)
2019
Cited alongside, same era.
2020
Later among the works it cites.
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., … & Christiano, P. F. (2020). Learning to summarize with human feedback. Advances in Neural Information Processing Systems, 33, 3008-3021
2020
Later among the works it cites.
2021
Later among the works it cites.
2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
Kambhampati, Subbarao, Sarath Sreedharan, Mudit Verma, Yantian Zha, and Lin Guan. ”Symbols as a lingua franca for bridging human-ai chasm for explainable and advisable ai systems.” In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, no. 11, pp. 12262-12267. 2022
2022
Closest in time.