Fetching the paper…
Reading the bibliography…
We define a novel neuro-symbolic framework, argumentative reward learning, which combines preference-based argumentation with existing approaches to reinforcement learning from human feedback.
On the acceptability of arguments…
Dung, P. M · 1995
Earlier work this paper cites.
Reasoning about preferences in argumentation frameworks
Modgil, S · 2009
Earlier work this paper cites.
Preference-based policy learning
Akrour, R., Schoenauer, M., and Sebag, M · 2011
Earlier work this paper cites.
April: Active preference learning-based reinforcement learning
Akrour, R., Schoenauer, M., and Sebag, M · 2012
Earlier work this paper cites.
A bayesian approach for policy learning from trajectory preference queries
Wilson, A., Fern, A., and Tadepalli, P · 2012
Earlier work this paper cites.
On the acceptability of arguments in preference-based argumentation
Amgoud, L. and Cayrol, C · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. A · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Cooperative inverse reinforcement learning
Hadfield-Menell, D., Russell, S. J., Abbeel, P., and Dragan, A. D · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Earlier work this paper cites.
Learning language games through interaction
Wang, S. I., Liang, P., and Manning, C. D · 2016
Cited alongside, same era.
Towards Artificial Argumentation
Atkinson, K., Baroni, P., Giacomin, M., Hunter, A., Prakken, H., Reed, C., Simari, G., Thimm, M., and Villata, S · 2017
Cited alongside, same era.
Deep rl from human preferences
Christiano, Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Cited alongside, same era.
Reward learning from human preferences and demonstrations in atari, 2018
Ibarz, B., Leike, J., Pohlen, T., Irving, G., Legg, S., and Amodei, D · 2018
Cited alongside, same era.
Scalable agent alignment via reward modeling…
Leike, Krueger, D., Everitt, T., Martic, M., Maini, V., and Legg, S · 2018
Cited alongside, same era.
An extensible interactive interface for agent design, 2019
Rahtz, M., Fang, J., Dragan, A. D., and Hadfield-Menell, D · 2019
Later among the works it cites.
Grandmaster level in starcraft II using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., Oh, J., Horgan, D., Kroiss, M., Danihelka, I., Huang, A., Sifre, L., Cai, T., Agapiou, J. P., Jaderberg, M., Vezhnevets, A. S., Leblond, R., Pohlen, T., Dalibard, V., Budden, D., Sulsky, Y., Molloy, J., Paine, T. L., Gülçehre, Ç., Wang, Z., Pfaff, T., Wu, Y., Ring, R., Yogatama, D., Wünsch, D., McKinney, K., Smith, O., Schaul, T., Lillicrap, T. P., Kavukcuoglu, K., Hassabis, D., Apps, C., and Silver, D · 2019
Later among the works it cites.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 2019
Later among the works it cites.
Learning to summarize from human feedback, 2020
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D. M., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Marcus, G · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Cited alongside, same era.
Parenting: Safe reinforcement learning from human input, 2019
Frye, C. and Feige, I · 2019
Cited alongside, same era.
Visual concept-metaconcept learning
Han, C., Mao, J., Gan, C., Tenenbaum, J., and Wu, J · 2019
Cited alongside, same era.
Explainable and contextual preferences based decision making with assumption-based argumentation for diagnostics and prognostics of alzheimer’s disease
Zeng, Z., Shen, Z., Chin, J. J., Leung, C., Wang, Y., Chi, Y., and Miao, C · 2020
Later among the works it cites.
Scaling language models: Methods, analysis & insights from training gopher
Rae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, F., Aslanides, J., Henderson, S., Ring, R., Young, S., et al · 2021
Later among the works it cites.
The societal implications of deep reinforcement learning
Whittlestone, J., Arulkumaran, K., and Crosby, M · 2021
Later among the works it cites.
Recursively summarizing books with human feedback
Wu, J., Ouyang, L., Ziegler, D. M., Stiennon, N., Lowe, R., Leike, J., and Christiano, P · 2021
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Closest in time.