Fetching the paper…
Reading the bibliography…
This paper proposes Self-Imitation Learning (SIL), a simple off-policy actor-critic algorithm that learns to reproduce the agent's past good decisions.
Adaptive confidence and adaptive curiosity
Schmidhuber, J · 1991
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, L. J · 1992
Earlier work this paper cites.
Memory-based reinforcement learning: Efficient computation with prioritized sweeping
Moore, A. W. and Atkeson, C. G · 1992
Earlier work this paper cites.
The role of exploration in learning control
Thrun, S. B · 1992
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y · 1999
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. R. and Tsitsiklis, J. N · 2000
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Precup, D., Sutton, R. S., and Singh, S. P · 2000
Earlier work this paper cites.
Off-policy temporal difference learning with function approximation
Precup, D., Sutton, R. S., and Dasgupta, S · 2001
Earlier work this paper cites.
Hippocampal contributions to control: the third way
Lengyel, M. and Dayan, P · 2008
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Strehl, A. L. and Littman, M. L · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Ziebart, B. D., Maas, A. L., Bagnell, J. A., and Dey, A. K · 2008
Earlier work this paper cites.
Modeling Purposeful Adaptive Behavior with the Principle of Maximum Causal Entropy
Ziebart, B. D · 2010
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Cited alongside, same era.
Incentivizing exploration in reinforcement learning with deep predictive models
Stadie, B. C., Levine, S., and Abbeel, P · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R · 2016
Cited alongside, same era.
Blundell, C., Uria, B., Pritzel, A., Li, Y., Ruderman, A., Leibo, J. Z., Rae, J., Wierstra, D., and Hassabis, D · 2016
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Later among the works it cites.
Learning to play in a day: Faster deep reinforcement learning by optimality tightening
He, F. S., Liu, Y., Schwing, A. G., and Peng, J · 2017
Later among the works it cites.
Bridging the gap between value and policy based reinforcement learning
Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D · 2017
Later among the works it cites.
Overcoming exploration in reinforcement learning with demonstrations
Nair, A., McGrew, B., Andrychowicz, M., Zaremba, W., and Abbeel, P · 2017
Later among the works it cites.
Combining policy gradient and q-learning
O’Donoghue, B., Munos, R., Kavukcuoglu, K., and Mnih, V · 2017
Later among the works it cites.
Count-based exploration with neural density models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Neural symbolic machines: Learning semantic parsers on freebase with weak supervision
Liang, C., Berant, J., Le, Q., Forbus, K. D., and Lao, N · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T. P., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Safe and efficient off-policy reinforcement learning
Munos, R., Stepleton, T., Harutyunyan, A., and Bellemare, M. G · 2016
Cited alongside, same era.
Deep exploration via bootstrapped dqn
Osband, I., Blundell, C., Pritzel, A., and Van Roy, B · 2016
Cited alongside, same era.
Prioritized experience replay
Schaul, T., Quan, J., Antonoglou, I., and Silver, D · 2016
Cited alongside, same era.
Ostrovski, G., Bellemare, M. G., Oord, A. v. d., and Munos, R · 2017
Later among the works it cites.
Neural episodic control
Pritzel, A., Uria, B., Srinivasan, S., Puigdomènech, A., Vinyals, O., Hassabis, D., Wierstra, D., and Blundell, C · 2017
Later among the works it cites.
#exploration: A study of count-based exploration for deep reinforcement learning
Tang, H., Houthooft, R., Foote, D., Stooke, A., Chen, O. X., Duan, Y., Schulman, J., DeTurck, F., and Abbeel, P · 2017
Later among the works it cites.
Sample efficient actor-critic with experience replay
Wang, Z., Bapst, V., Heess, N., Mnih, V., Munos, R., Kavukcuoglu, K., and de Freitas, N · 2017
Later among the works it cites.
Neural program synthesis with priority queue training
Abolafia, D. A., Norouzi, M., and Le, Q. V · 2018
Closest in time.
The reactor: A sample-efficient actor-critic architecture
Gruslys, A., Azar, M. G., Bellemare, M. G., and Munos, R · 2018
Closest in time.
Deep q-learning from demonstrations
Hester, T., Vecerik, M., Pietquin, O., Lanctot, M., Schaul, T., Piot, B., Sendonaris, A., Dulac-Arnold, G., Osband, I., Agapiou, J., et al · 2018
Closest in time.
Reinforcement learning from imperfect demonstrations
Xu, H., Gao, Y., Lin, J., Yu, F., Levine, S., and Darrell, T · 2018
Closest in time.