Fetching the paper…
Reading the bibliography…
A Markovian decision process
Bellman, R · 1957
Earlier work this paper cites.
Optimal control of Markov processes with incomplete state information i
Åström, K. J · 1965
Earlier work this paper cites.
Temporal credit assignment in reinforcement learning
Sutton, R. S · 1984
Earlier work this paper cites.
The complexity of Markov decision processes
Papadimitriou, C. H. and Tsitsiklis, J. N · 1987
Earlier work this paper cites.
Evolutionary principles in self-referential learning, or on learning how to learn: the meta-meta-… hook
Schmidhuber, J · 1987
Earlier work this paper cites.
Finding structure in time
Elman, J. L · 1990
Earlier work this paper cites.
Active perception and reinforcement learning
Whitehead, S. D. and Ballard, D. H · 1990
Earlier work this paper cites.
Reinforcement learning in Markovian and non-Markovian environments
Schmidhuber, J · 1991
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Bengio, Y., Simard, P. Y., and Frasconi, P · 1994
Earlier work this paper cites.
Acting optimally in partially observable stochastic domains
Cassandra, A. R., Kaelbling, L. P., and Littman, M. L · 1994
Earlier work this paper cites.
On the computational power of neural nets
Siegelmann, H. T. and Sontag, E. D · 1995
Earlier work this paper cites.
Robust and optimal control
Khalil, I. S., Doyle, J., and Glover, K · 1996
Earlier work this paper cites.
Algorithms for sequential decision-making
Littman, M. L · 1996
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Kaelbling, L. P., Littman, M. L., and Cassandra, A. R · 1998
Earlier work this paper cites.
Solving uncertain Markov decision processes
Bagnell, J. A., Ng, A. Y., and Schneider, J. G · 2001
Earlier work this paper cites.
Reinforcement learning with long short-term memory
Bakker, B · 2001
Earlier work this paper cites.
Learning to learn using gradient descent
Hochreiter, S., Younger, A. S., and Conwell, P. R · 2001
Earlier work this paper cites.
Markov decision processes with delays and asynchronous cost collection
Katsikopoulos, K. V. and Engelbrecht, S. E · 2003
Earlier work this paper cites.
Robust reinforcement learning
Morimoto, J. and Doya, K · 2005
Earlier work this paper cites.
Robust control of Markov decision processes with uncertain transition matrices
Nilim, A. and Ghaoui, L. E · 2005
Earlier work this paper cites.
Recurrent neural networks are universal approximators
Schäfer, A. M. and Zimmermann, H. G · 2006
Earlier work this paper cites.
Solving deep memory pomdps with recurrent policy gradients
Wierstra, D., Förster, A., Peters, J., and Schmidhuber, J · 2007
Earlier work this paper cites.
Multi-task reinforcement learning: a hierarchical bayesian approach
Wilson, A., Fern, A., Ray, S., and Tadepalli, P · 2007
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Ziebart, B. D · 2010
Earlier work this paper cites.
Learning to grasp under uncertainty
Stulp, F., Theodorou, E. A., Buchli, J., and Schaal, S · 2011
Earlier work this paper cites.
Protecting against evaluation overfitting in empirical reinforcement learning
Whiteson, S., Tanner, B., Taylor, M. E., and Stone, P · 2011
Earlier work this paper cites.
Learning to learn
Thrun, S. and Pratt, L · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Cho, K., van Merrienboer, B., Gülçehre, Ç., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y · 2014
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Chung, J., Gülçehre, Ç., Cho, K., and Bengio, Y · 2014
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 2014
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S · 2014
Earlier work this paper cites.
Deep recurrent q-learning for partially observable mdps
Hausknecht, M. J. and Stone, P · 2015
Earlier work this paper cites.
Memory-based control with recurrent neural networks
Heess, N., Hunt, J. J., Lillicrap, T. P., and Silver, D · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Embed to control: A locally linear latent dynamics model for control from raw images
Watter, M., Springenberg, J. T., Boedecker, J., and Riedmiller, M. A · 2015
Earlier work this paper cites.
PyBullet, a python module for physics simulation for games, robotics and machine learning
Coumans, E. and Bai, Y · 2016
Earlier work this paper cites.
Hybrid computing using a neural network with dynamic external memory
Graves, A., Wayne, G., Reynolds, M., Harley, T., Danihelka, I., Grabska-Barwinska, A., Colmenarejo, S. G., Grefenstette, E., Ramalho, T., Agapiou, J. P., Badia, A. P., Hermann, K. M., Zwols, Y., Ostrovski, G., Cain, A., King, H., Summerfield, C., Blunsom, P., Kavukcuoglu, K., and Hassabis, D · 2016
Earlier work this paper cites.
Control of memory, active perception, and action in minecraft
Oh, J., Chockalingam, V., Singh, S., and Lee, H · 2016
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Cited alongside, same era.
Adversarial attacks on neural network policies
Huang, S. H., Papernot, N., Goodfellow, I. J., Duan, Y., and Abbeel, P · 2017
Cited alongside, same era.
Tactics of adversarial attack on deep reinforcement learning agents
Lin, Y., Hong, Z., Liao, Y., Shih, M., Liu, M., and Sun, M · 2017
Cited alongside, same era.
Learning to navigate in complex environments
Mirowski, P., Pascanu, R., Viola, F., Soyer, H., Ballard, A., Banino, A., Denil, M., Goroshin, R., Sifre, L., Kavukcuoglu, K., Kumaran, D., and Hadsell, R · 2017
Cited alongside, same era.
Robust adversarial reinforcement learning
Pinto, L., Davidson, J., Sukthankar, R., and Gupta, A · 2017
Sequence modeling of temporal credit assignment for episodic reinforcement learning
Liu, Y., Luo, Y., Zhong, Y., Chen, X., Liu, Q., and Peng, J · 2019
Later among the works it cites.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Rakelly, K., Zhou, A., Finn, C., Levine, S., and Quillen, D · 2019
Later among the works it cites.
Action robust reinforcement learning and applications in continuous control
Tessler, C., Efroni, Y., and Mannor, S · 2019
Later among the works it cites.
Robust reinforcement learning in pomdps with incomplete and noisy observations
Wang, Y., He, H., and Tan, X · 2019
Later among the works it cites.
Meta-World: A benchmark and evaluation for multi-task and meta reinforcement learning
Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Epopt: Learning robust neural network policies using model ensembles
Rajeswaran, A., Ghotra, S., Ravindran, B., and Levine, S · 2017
Cited alongside, same era.
Towards generalization and simplicity in continuous control
Rajeswaran, A., Lowrey, K., Todorov, E., and Kakade, S. M · 2017
Cited alongside, same era.
Domain randomization for transferring deep neural networks from simulation to the real world
Tobin, J., Fong, R., Ray, A., Schneider, J., Zaremba, W., and Abbeel, P · 2017
Cited alongside, same era.
Learning to reinforcement learn
Wang, J., Kurth-Nelson, Z., Soyer, H., Leibo, J. Z., Tirumala, D., Munos, R., Blundell, C., Kumaran, D., and Botvinick, M. M · 2017
Cited alongside, same era.
Target-driven visual navigation in indoor scenes using deep reinforcement learning
Zhu, Y., Mottaghi, R., Kolve, E., Lim, J. J., Gupta, A., Fei-Fei, L., and Farhadi, A · 2017
Cited alongside, same era.
Planning with trust for human-robot collaboration
Chen, M., Nikolaidis, S., Soh, H., Hsu, D., and Srinivasa, S. S · 2018
Cited alongside, same era.
Investigating generalisation in continuous deep reinforcement learning
Zhao, C., Sigaud, O., Stulp, F., and Hospedales, T. M · 2019
Later among the works it cites.
Leveraging procedural generation to benchmark reinforcement learning
Cobbe, K., Hesse, C., Hilton, J., and Schulman, J · 2020
Later among the works it cites.
Offline meta learning of exploration
Dorfman, R., Shenfeld, I., and Tamar, A · 2020
Later among the works it cites.
Implementation matters in deep RL: A case study on PPO and TRPO
Engstrom, L., Ilyas, A., Santurkar, S., Tsipras, D., Janoos, F., Rudolph, L., and Madry, A · 2020
Later among the works it cites.
Meta-q-learning
Fakoor, R., Chaudhari, P., Soatto, S., and Smola, A. J · 2020
Later among the works it cites.
Self-attentional credit assignment for transfer in reinforcement learning
Ferret, J., Marinier, R., Geist, M., and Pietquin, O · 2020
Later among the works it cites.
Learning guidance rewards with trajectory-space smoothing
Gangwani, T., Zhou, Y., and Peng, J · 2020
Later among the works it cites.
Adversarial policies: Attacking deep reinforcement learning
Gleave, A., Dennis, M., Wild, C., Kant, N., Levine, S., and Russell, S · 2020
Later among the works it cites.
Variational recurrent models for solving partially observable control tasks
Han, D., Doya, K., and Tani, J · 2020
Later among the works it cites.
Stochastic latent actor-critic: Deep reinforcement learning with a latent variable model
Lee, A. X., Nagabandi, A., Abbeel, P., and Levine, S · 2020
Later among the works it cites.
Network randomization: A simple technique for generalization in deep reinforcement learning
Lee, K., Lee, K., Shin, J., and Lee, H · 2020
Later among the works it cites.
Robust reinforcement learning for continuous control with model misspecification
Mankowitz, D. J., Levine, N., Jeong, R., Abdolmaleki, A., Springenberg, J. T., Shi, Y., Kay, J., Hester, T., Mann, T. A., and Riedmiller, M. A · 2020
Later among the works it cites.
Stabilizing transformers for reinforcement learning
Parisotto, E., Song, H. F., Rae, J. W., Pascanu, R., Gülçehre, Ç., Jayakumar, S. M., Jaderberg, M., Kaufman, R. L., Clark, A., Noury, S., Botvinick, M. M., Heess, N., and Hadsell, R · 2020
Later among the works it cites.
Rl baselines3 zoo
Raffin, A · 2020
Later among the works it cites.
Observational overfitting in reinforcement learning
Song, X., Jiang, Y., Tu, S., Du, Y., and Neyshabur, B · 2020
Later among the works it cites.
Meta-gradient reinforcement learning with an objective discovered online
Xu, Z., van Hasselt, H. P., Hessel, M., Oh, J., Singh, S., and Silver, D · 2020
Later among the works it cites.
Robust deep reinforcement learning against adversarial perturbations on state observations
Zhang, H., Chen, H., Xiao, C., Li, B., Liu, M., Boning, D. S., and Hsieh, C · 2020
Later among the works it cites.
Episodic reinforcement learning with associative memory
Zhu, G., Lin, Z., Yang, G., and Zhang, C · 2020
Later among the works it cites.
Varibad: A very good method for bayes-adaptive deep RL via meta-learning
Zintgraf, L. M., Shiarlis, K., Igl, M., Schulze, S., Gal, Y., Hofmann, K., and Whiteson, S · 2020
Later among the works it cites.
What matters for on-policy deep actor-critic methods? A large-scale study
Andrychowicz, M., Raichuk, A., Stanczyk, P., Orsini, M., Girgin, S., Marinier, R., Hussenot, L., Geist, M., Pietquin, O., Michalski, M., Gelly, S., and Bachem, O · 2021
Closest in time.
Monotonic robust policy optimization with model discrepancy
Jiang, Y., Li, C., Dai, W., Zou, J., and Xiong, H · 2021
Closest in time.
A survey of generalisation in deep reinforcement learning
Kirk, R., Zhang, A., Grefenstette, E., and Rocktäschel, T · 2021
Closest in time.
Memory-based deep reinforcement learning for pomdps
Meng, L., Gorbet, R., and Kulic, D · 2021
Closest in time.
Counterfactual credit assignment in model-free reinforcement learning
Mesnard, T., Weber, T., Viola, F., Thakoor, S., Saade, A., Harutyunyan, A., Dabney, W., Stepleton, T. S., Heess, N., Guez, A., Moulines, E., Hutter, M., Buesing, L., and Munos, R · 2021
Closest in time.
Smooth exploration for robotic reinforcement learning
Raffin, A., Kober, J., and Stulp, F · 2021
Closest in time.
Decoupling value and policy for generalization in reinforcement learning
Raileanu, R. and Fergus, R · 2021
Closest in time.
Synthetic returns for long-term credit assignment
Raposo, D., Ritter, S., Santoro, A., Wayne, G., Weber, T., Botvinick, M. M., van Hasselt, H., and Song, H. F · 2021
Closest in time.
Learning long-term reward redistribution via randomized return decomposition
Ren, Z., Guo, R., Zhou, Y., and Peng, J · 2021
Closest in time.
Towards optimal attacks on reinforcement learning policies
Russo, A. and Proutière, A · 2021
Closest in time.
Safe exploration by solving early terminated MDP
Sun, H., Xu, Z., Fang, M., Peng, Z., Guo, J., Dai, B., and Zhou, B · 2021
Closest in time.
Tianshou: a highly modularized deep reinforcement learning library
Weng, J., Chen, H., Yan, D., You, K., Duburcq, A., Zhang, M., Su, H., and Zhu, J · 2021
Closest in time.
Recurrent off-policy baselines for memory-based continuous control
Yang, Z. and Nguyen, H · 2021
Closest in time.
Robust reinforcement learning on state observations with learned optimal adversary
Zhang, H., Chen, H., Boning, D. S., and Hsieh, C · 2021
Closest in time.
Varibad: Variational bayes-adaptive deep RL via meta-learning
Zintgraf, L. M., Schulze, S., Lu, C., Feng, L., Igl, M., Shiarlis, K., Gal, Y., Hofmann, K., and Whiteson, S · 2021
Closest in time.