Fetching the paper…
Reading the bibliography…
It is notoriously difficult to control the behavior of reinforcement learning agents.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 1910
Earlier work this paper cites.
Multitask learning
Caruana, R · 1997
Earlier work this paper cites.
Guiding a reinforcement learner with natural language advice: Initial results in robocup soccer
Kuhlmann, G., Stone, P., Mooney, R., and Shavlik, J · 2004
Earlier work this paper cites.
A survey of multi-objective sequential decision-making
Roijers, D. M., Vamplew, P., Whiteson, S., and Dazeley, R · 2013
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Cho, K., Van Merriënboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D · 2015
Cited alongside, same era.
Q-learning for robust satisfaction of signal temporal logic specifications
Aksaray, D., Jones, A., Kong, Z., Schwager, M., and Belta, C · 2016
Cited alongside, same era.
Multi-objective deep reinforcement learning
Mossalam, H., Assael, Y. M., Roijers, D. M., and Whiteson, S · 2016
Cited alongside, same era.
Hindsight experience replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, O. P., and Zaremba, W · 2017
Cited alongside, same era.
Reinforcement learning with temporal logic rewards
Li, X., Vasile, C.-I., and Belta, C · 2017
Cited alongside, same era.
Dynamic weights in multi-objective deep reinforcement learning
Generalizing across multi-objective reward functions in deep reinforcement learning
Friedman, E. and Fontaine, F · 2018
Later among the works it cites.
Soft actor-critic algorithms and applications
Haarnoja, T., Zhou, A., Hartikainen, K., Tucker, G., Ha, S., Tan, J., Kumar, V., Zhu, H., Gupta, A., Abbeel, P., et al · 2018
Later among the works it cites.
Structured Reward Shaping using Signal Temporal Logic specifications
Balakrishnan, A. and Deshmukh, J. V · 2019
Closest in time.
From language to goals: Inverse reinforcement learning for vision-based instruction following
Fu, J., Korattikara, A., Levine, S., and Guadarrama, S · 2019
Closest in time.
Using natural language for reward shaping in reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Abels, A., Roijers, D. M., Lenaerts, T., Nowé, A., and Steckelmacher, D · 2018
Cited alongside, same era.
Goyal, P., Niekum, S., and Mooney, R. J · 2019
Closest in time.
Temporal logic guided safe reinforcement learning using control barrier functions
Li, X. and Belta, C · 2019
Closest in time.