Fetching the paper…
Reading the bibliography…
In machine learning for sequential decision-making, an algorithmic agent learns to interact with an environment while receiving feedback in the form of a reward signal.
A survey of robot learning from demonstration
Argall, B. D., Chernova, S., Veloso, M., and Browning, B · 2009
Earlier work this paper cites.
Preference-based policy learning
Akrour, R., Schoenauer, M., and Sebag, M · 2011
Earlier work this paper cites.
The Malmo platform for artificial intelligence experimentation
Johnson, M., Hofmann, K., Hutton, T., and Bignell, D · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
Van Hasselt, H., Guez, A., and Silver, D · 2016
Earlier work this paper cites.
Do you want your autonomous car to drive like you?
Basu, C., Yang, Q., Hungerman, D., Singhal, M., and Dragan, A. D · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Accurately interpreting clickthrough data as implicit feedback
Joachims, T., Granka, L., Pan, B., Hembrooke, H., and Gay, G · 2017
Earlier work this paper cites.
Active preference-based learning of reward functions
Sadigh, D., Dragan, A. D., Sastry, S., and Seshia, S. A · 2017
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
Reward learning from human preferences and demonstrations in Atari
Ibarz, B., Leike, J., Pohlen, T., Irving, G., Legg, S., and Amodei, D · 2018
Cited alongside, same era.
Advancements in dueling bandits
Sui, Y., Zoghi, M., Hofmann, K., and Yue, Y · 2018
Cited alongside, same era.
Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations
Brown, D., Goo, W., Nagarajan, P., and Niekum, S · 2019
Cited alongside, same era.
MineRL: a large-scale dataset of Minecraft demonstrations
Guss, W. H., Houghton, B., Topin, N., Wang, P., Codel, C., Veloso, M., and Salakhutdinov, R · 2019
Cited alongside, same era.
Asking easy questions: A user-friendly approach to active reward learning
PEBBLE: Feedback-efficient interactive reinforcement learning via relabeling experience and unsupervised pre-training
Lee, K., Smith, L. M., and Abbeel, P · 2021
Later among the works it cites.
The MineRL BASALT competition on learning from human feedback
Shah, R., Wild, C., Wang, S. H., Alex, N., Houghton, B., Guss, W., Mohanty, S., Kanervisto, A., Milani, S., Topin, N., et al · 2021
Later among the works it cites.
Improving sample efficiency in model-free reinforcement learning from images
Yarats, D., Zhang, A., Kostrikov, I., Amos, B., Pineau, J., and Fergus, R · 2021
Later among the works it cites.
Video pretraining (VPT): Learning to act by watching unlabeled online videos
Baker, B., Akkaya, I., Zhokov, P., Huizinga, J., Tang, J., Ecoffet, A., Houghton, B., Sampedro, R., and Clune, J · 2022
Later among the works it cites.
Retrospective on the 2021 MineRL BASALT competition on learning from human feedback
Shah, R., Wang, S. H., Wild, C., Milani, S., Kanervisto, A., Goecks, V. G., Waytowich, N., Watkins-Valls, D., Prakash, B., Mills, E., et al · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bıyık, E., Palan, M., Landolfi, N. C., Losey, D. P., Sadigh, D., et al · 2020
Cited alongside, same era.
SQIL: Imitation learning via reinforcement learning with sparse rewards
Reddy, S., Dragan, A. D., and Levine, S · 2020
Cited alongside, same era.
Goecks, V. G., Waytowich, N., Watkins, D., and Prakash, B · 2021
Cited alongside, same era.
Safe imitation learning via fast Bayesian reward inference from preferences
Brown, D., Coleman, R., Srinivasan, R., and Niekum, S
Cited in the paper.
Better-than-demonstrator imitation learning via automatically-ranked demonstrations
Brown, D. S., Goo, W., and Niekum, S
Cited in the paper.
Later among the works it cites.
Milani, S., Kanervisto, A., Ramanauskas, K., Schulhoff, S., Houghton, B., Mohanty, S., Galbraith, B., Chen, K., Song, Y., Zhou, T., et al · 2023
Closest in time.
MineRLTreechop-v0 environment
MineRL documentation · 2023
Closest in time.