Learning end-to-end goal-oriented dialog with multiple answers
Rajendran, J., Ganhotra, J., Singh, S., and Polymenakos, L. (2018) · 2018
Later among the works it cites.
Meta-gradient reinforcement learning
Xu, Z., van Hasselt, H. P., and Silver, D. (2018) · 2018
Later among the works it cites.
On learning intrinsic rewards for policy gradient methods
Zheng, Z., Oh, J., and Singh, S. (2018) · 2018
Later among the works it cites.
RUDDER: return decomposition for delayed rewards
Arjona-Medina, J. A., Gillhofer, M., Widrich, M., Unterthiner, T., Brandstetter, J., and Hochreiter, S. (2019) · 2019
Later among the works it cites.
Tackling sparse rewards in real-time games with statistical forward planning methods
Gaina, R. D., Lucas, S. M., and Pérez-Liébana, D. (2019) · 2019
Later among the works it cites.
Benchmarking neural network robustness to common corruptions and perturbations
Hendrycks, D., and Dietterich, T. G. (2019) · 2019
Later among the works it cites.
Obstacle Tower: A Generalization Challenge in Vision, Control, and Planning
Juliani, A., Khalifa, A., Berges, V., Harper, J., Teng, E., Henry, H., Crespi, A., Togelius, J., and Lange, D. (2019) · 2019
Later among the works it cites.
Self-paced contextual reinforcement learning
Klink, P., Abdulsamad, H., Belousov, B., and Peters, J. (2019) · 2019
Later among the works it cites.
Behaviour suite for reinforcement learning
Osband, I., Doron, Y., Hessel, M., Aslanides, J., Sezener, E., Saraiva, A., McKinney, K., Lattimore, T., Szepezvari, C., Singh, S., Roy, B. V., Sutton, R., Silver, D., and Hasselt, H. V. (2019) · 2019
Later among the works it cites.
Learning to design rna
Runge, F., Stoll, D., Falkner, S., and Hutter, F. (2019) · 2019
Later among the works it cites.
Exploratory not explanatory: Counterfactual analysis of saliency maps for deep reinforcement learning
Atrey, A., Clary, K., and Jensen, D. D. (2020) · 2020
Later among the works it cites.
Implementation matters in deep RL: A case study on PPO and TRPO
Engstrom, L., Ilyas, A., Santurkar, S., Tsipras, D., Janoos, F., Rudolph, L., and Madry, A. (2020) · 2020
Later among the works it cites.
Learning to utilize shaping rewards: A new approach of reward shaping
Hu, Y., Wang, W., Jia, H., Wang, Y., Chen, Y., Hao, J., Wu, F., and Fan, C. (2020) · 2020
Later among the works it cites.
Action space shaping in deep reinforcement learning
Kanervisto, A., Scheller, C., and Hautamäki, V. (2020) · 2020
Later among the works it cites.
Reward tweaking: Maximizing the total reward while planning for short horizons.
Tessler, C., and Mannor, S. (2020) · 2020
Later among the works it cites.
What matters for on-policy deep actor-critic methods? A large-scale study
Andrychowicz, M., Raichuk, A., Stanczyk, P., Orsini, M., Girgin, S., Marinier, R., Hussenot, L., Geist, M., Pietquin, O., Michalski, M., Gelly, S., and Bachem, O. (2021) · 2021
Later among the works it cites.
TempoRL: Learning when to act
Biedenkapp, A., Rajan, R., Hutter, F., and Lindauer, M. (2021) · 2021
Later among the works it cites.
The benchmark lottery
Dehghani, M., Tay, Y., Gritsenko, A. A., Zhao, Z., Houlsby, N., Diaz, F., Metzler, D., and Vinyals, O. (2021) · 2021
Later among the works it cites.
On the Importance of Hyperparameter Optimization for Model-based Reinforcement Learning
Zhang, B., Rajan, R., Pineda, L., Lambert, N., Biedenkapp, A., Chua, K., Hutter, F., and Calandra, R. (2021) · 2021
Later among the works it cites.
Learning task-distribution reward shaping with meta-learning
Zou, H., Ren, T., Yan, D., Su, H., and Zhu, J. (2021) · 2021
Later among the works it cites.
Leveraging procedural generation to benchmark reinforcement learning
Cobbe, K., Hesse, C., Hilton, J., and Schulman, J. (2020) · 2056
Later among the works it cites.