Deep Reinforcement Learning that Matters
Original
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., and Meger, D · 2017
Later among the works it cites.
Deep q-learning from demonstrations
Original
Hester, T., Vecerik, M., Pietquin, O., Lanctot, M., Schaul, T., Piot, B., Horgan, D., Quan, J., Sendonaris, A., Dulac-Arnold, G., et al · 2017
Later among the works it cites.
Imitation learning: A survey of learning methods
Hussein, A., Gaber, M. M., Elyan, E., and Jayne, C · 2017
Later among the works it cites.
Deeply aggrevated: Differentiable imitation learning for sequential prediction
Sun, W., Venkatraman, A., Gordon, G. J., Boots, B., and Bagnell, J. A · 2017
Later among the works it cites.
Non-parametric policy search with limited information loss
Van Hoof, H., Neumann, G., and Peters, J · 2017
Later among the works it cites.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
Original
Večerík, M., Hester, T., Scholz, J., Wang, F., Pietquin, O., Piot, B., Heess, N., Rothörl, T., Lampe, T., and Riedmiller, M · 2017
Later among the works it cites.
A deeper look at experience replay
Original
Zhang, S. and Sutton, R. S · 2017
Later among the works it cites.
Efficient exploration through bayesian deep q-networks
Original
Azizzadenesheli, K., Brunskill, E., and Anandkumar, A · 2018
Closest in time.
Sample-efficient reinforcement learning with stochastic ensemble value expansion
Buckman, J., Hafner, D., Tucker, G., Brevdo, E., and Lee, H · 2018
Closest in time.
Fast policy learning through imitation and reinforcement
Original
Cheng, C.-A., Yan, X., Wagener, N., and Boots, B · 2018
Closest in time.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Chua, K., Calandra, R., McAllister, R., and Levine, S · 2018
Closest in time.
Addressing function approximation error in actor-critic methods
Fujimoto, S., van Hoof, H., and Meger, D · 2018
Closest in time.
Reinforcement learning from imperfect demonstrations
Original
Gao, Y., Lin, J., Yu, F., Levine, S., and Darrell, T · 2018
Closest in time.
Synthesizing neural network controllers with probabilistic model based reinforcement learning
Original
Higuera, J. C. G., Meger, D., and Dudek, G · 2018
Closest in time.
Selective experience replay for lifelong learning
Original
Isele, D. and Cosgun, A · 2018
Closest in time.
Residual reinforcement learning for robot control
Original
Johannink, T., Bahl, S., Nair, A., Luo, J., Kumar, A., Loskyll, M., Ojea, J. A., Solowjow, E., and Levine, S · 2018
Closest in time.
Representation balancing mdps for off-policy policy evaluation
Liu, Y., Gottesman, O., Raghu, A., Komorowski, M., Faisal, A. A., Doshi-Velez, F., and Brunskill, E · 2018
Closest in time.
Non-delusional q-learning and value-iteration
Lu, T., Schuurmans, D., and Boutilier, C · 2018
Closest in time.
Overcoming exploration in reinforcement learning with demonstrations
Nair, A., McGrew, B., Andrychowicz, M., Zaremba, W., and Abbeel, P · 2018
Closest in time.
The uncertainty Bellman equation and exploration
O’Donoghue, B., Osband, I., Munos, R., and Mnih, V · 2018
Closest in time.
Randomized prior functions for deep reinforcement learning
Osband, I., Aslanides, J., and Cassirer, A · 2018
Closest in time.
Constrained exploration and recovery from experience shaping
Original
Pham, T.-H., De Magistris, G., Agravante, D. J., Chaudhury, S., Munawar, A., and Tachibana, R · 2018
Closest in time.
Residual policy learning
Original
Silver, T., Allen, K., Tenenbaum, J., and Kaelbling, L · 2018
Closest in time.
Truncated horizon policy search: Combining reinforcement learning & imitation learning
Original
Sun, W., Bagnell, J. A., and Boots, B · 2018
Closest in time.
Randomized value functions via multiplicative normalizing flows
Original
Touati, A., Satija, H., Romoff, J., Pineau, J., and Vincent, P · 2018
Closest in time.
Algorithmic framework for model-based reinforcement learning with theoretical guarantees
Original
Xu, H., Li, Y., Tian, Y., Darrell, T., and Ma, T · 2018
Closest in time.