Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) agents are widely used for solving complex sequential decision making tasks, but still exhibit difficulty in generalizing to scenarios not seen during training.
On stochastic limit and order relationships
H. B. Mann and A. Wald · 1943
Earlier work this paper cites.
Asymptotic minimax character of the sample distribution function and of the classical multinomial estimator
A. Dvoretzky, J. Kiefer, and J. Wolfowitz · 1956
Earlier work this paper cites.
Q-learning
C. J. Watkins and P. Dayan · 1992
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
P. Dayan · 1993
Earlier work this paper cites.
Generalization in reinforcement learning: Safely approximating the value function
J. Boyan and A. W. Moore · 1995
Earlier work this paper cites.
The sample complexity of pattern classification with neural networks: the size of the weights is more important than the size of the network
P. L. Bartlett · 1998
Earlier work this paper cites.
Asymptotic statistics. cambridge series in statistical and probabilistic mathematics, 1998
A. W. van der Vaart · 1998
Earlier work this paper cites.
A survey of pomdp solution techniques
K. P. Murphy · 2000
Earlier work this paper cites.
Metrics for finite markov decision processes
N. Ferns, P. Panangaden, and D. Precup · 2004
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
D. Ernst, P. Geurts, and L. Wehenkel · 2005
Earlier work this paper cites.
State abstraction discovery from irrelevant state variables
N. K. Jong and P. Stone · 2005
Earlier work this paper cites.
Towards a unified theory of state abstraction for mdps
L. Li, T. J. Walsh, and M. L. Littman · 2006
Earlier work this paper cites.
A universal representation transformer layer for few-shot image classification
L. Liu, W. Hamilton, G. Long, J. Jiang, and H. Larochelle · 2006
Earlier work this paper cites.
Using bisimulation for policy transfer in mdps
P. S. Castro and D. Precup · 2010
Earlier work this paper cites.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
R. S. Sutton, J. Modayil, M. Delp, T. Degris, P. M. Pilarski, A. White, and D. Precup · 2011
Earlier work this paper cites.
Batch reinforcement learning
S. Lange, T. Gabel, and M. Riedmiller · 2012
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Earlier work this paper cites.
Batch learning from logged bandit feedback through counterfactual risk minimization
A. Swaminathan and T. Joachims · 2015
Earlier work this paper cites.
Successor features for transfer in reinforcement learning
A. Barreto, W. Dabney, R. Munos, J. J. Hunt, T. Schaul, H. Van Hasselt, and D. Silver · 2016
Earlier work this paper cites.
Darla: Improving zero-shot transfer in reinforcement learning
I. Higgins, A. Pal, A. Rusu, L. Matthey, C. Burgess, A. Pritzel, M. Botvinick, C. Blundell, and A. Lerchner · 2017
Earlier work this paper cites.
J. Oh, S. Singh, and H. Lee · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
L. Espeholt, H. Soyer, R. Munos, K. Simonyan, V. Mnih, T. Ward, Y. Doron, V. Firoiu, T. Harley, I. Dunning, et al · 2018
Earlier work this paper cites.
Learning deep representations by mutual information estimation and maximization
R. D. Hjelm, A. Fedorov, S. Lavoie-Marchildon, K. Grewal, P. Bachman, A. Trischler, and Y. Bengio · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
A. v. d. Oord, Y. Li, and O. Vinyals · 2018
Cited alongside, same era.
Truncated horizon policy search: Combining reinforcement learning & imitation learning
W. Sun, J. A. Bagnell, and B. Boots · 2018
Cited alongside, same era.
Quantifying generalization in reinforcement learning
K. Cobbe, O. Klimov, C. Hesse, T. Kim, and J. Schulman · 2019
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
S. Fujimoto, D. Meger, and D. Precup · 2019
Cited alongside, same era.
Reward decomposition with representation decomposition
Z. Lin, D. Yang, L. Zhao, T. Qin, G. Yang, and T.-Y. Liu · 2020
Later among the works it cites.
Count-based exploration with the successor representation
M. C. Machado, M. G. Bellemare, and M. Bowling · 2020
Later among the works it cites.
Deep reinforcement and infomax learning
B. Mazoure, R. T. d. Combes, T. Doan, P. Bachman, and R. D. Hjelm · 2020
Later among the works it cites.
Kinematic state abstraction and provably efficient rich-observation reinforcement learning
D. Misra, M. Henaff, A. Krishnamurthy, and J. Langford · 2020
Later among the works it cites.
Data-efficient reinforcement learning with self-predictive representations
M. Schwarzer, A. Anand, R. Goel, R. D. Hjelm, A. Courville, and P. Bachman · 2020
Later among the works it cites.
Multi-label contrastive predictive coding
J. Song and S. Ermon · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deepmdp: Learning continuous latent space models for representation learning
C. Gelada, S. Kumar, J. Buckman, O. Nachum, and M. G. Bellemare · 2019
Cited alongside, same era.
Multi-task deep reinforcement learning with popart
M. Hessel, H. Soyer, L. Espeholt, W. Czarnecki, S. Schmitt, and H. van Hasselt · 2019
Cited alongside, same era.
Observational overfitting in reinforcement learning
X. Song, Y. Jiang, S. Tu, Y. Du, and B. Neyshabur · 2019
Cited alongside, same era.
Meta-dataset: A dataset of datasets for learning to learn from few examples
E. Triantafillou, T. Zhu, V. Dumoulin, P. Lamblin, U. Evci, K. Xu, R. Goroshin, C. Gelada, K. Swersky, P.-A. Manzagol, et al · 2019
Cited alongside, same era.
Behavior regularized offline reinforcement learning
Y. Wu, G. Tucker, and O. Nachum · 2019
Cited alongside, same era.
Improving sample efficiency in model-free reinforcement learning from images
D. Yarats, A. Zhang, I. Kostrikov, B. Amos, J. Pineau, and R. Fergus · 2019
Cited alongside, same era.
Pc-pg: Policy cover directed exploration for provable policy gradient learning
A. Agarwal, M. Henaff, S. Kakade, and W. Sun · 2020
Cited alongside, same era.
Later among the works it cites.
Curl: Contrastive unsupervised representations for reinforcement learning
A. Srinivas, M. Laskin, and P. Abbeel · 2020
Later among the works it cites.
Generalization bounds for deep learning
G. Valle-Pérez and A. A. Louis · 2020
Later among the works it cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
T. Yu, D. Quillen, Z. He, R. Julian, K. Hausman, C. Finn, and S. Levine · 2020
Later among the works it cites.
Learning invariant representations for reinforcement learning without reconstruction
A. Zhang, R. McAllister, R. Calandra, Y. Gal, and S. Levine · 2020
Later among the works it cites.
Contrastive behavioral similarity embeddings for generalization in reinforcement learning
R. Agarwal, M. C. Machado, P. S. Castro, and M. G. Bellemare · 2021
Closest in time.
Augmented world models facilitate zero-shot dynamics generalization from a single offline environment
P. J. Ball, C. Lu, J. Parker-Holder, and S. Roberts · 2021
Closest in time.
Learning generalizable robotic reward functions from” in-the-wild” human videos
A. S. Chen, S. Nair, and C. Finn · 2021
Closest in time.
Heuristic-guided reinforcement learning
C.-A. Cheng, A. Kolobov, and A. Swaminathan · 2021
Closest in time.
TF-Agents: A library for reinforcement learning in tensorflow
S. Guadarrama, A. Korattikara, O. Ramirez, P. Castro, E. Holly, S. Fishman, K. Wang, E. Gonina, N. Wu, E. Kokiopoulou, L. Sbaiz, J. Smith, G. Bartók, J. Berent, C. Harris, V. Vanhoucke, and E. Brevdo · 2021
Closest in time.
Offline reinforcement learning with fisher divergence critic regularization
I. Kostrikov, J. Tompson, R. Fergus, and O. Nachum · 2021
Closest in time.
Provable representation learning for imitation with contrastive fourier features
O. Nachum and M. Yang · 2021
Closest in time.
Decoupling value and policy for generalization in reinforcement learning
R. Raileanu and R. Fergus · 2021
Closest in time.
Pretraining representations for data-efficient reinforcement learning
M. Schwarzer, N. Rajkumar, M. Noukhovitch, A. Anand, L. Charlin, D. Hjelm, P. Bachman, and A. Courville · 2021
Closest in time.
S4rl: Surprisingly simple self-supervision for offline reinforcement learning
S. Sinha and A. Garg · 2021
Closest in time.
The distracting control suite–a challenging benchmark for reinforcement learning from pixels
A. Stone, O. Ramirez, K. Konolige, and R. Jonschkowski · 2021
Closest in time.
Decoupling representation learning from reinforcement learning
A. Stooke, K. Lee, P. Abbeel, and M. Laskin · 2021
Closest in time.
Learning one representation to optimize all rewards
A. Touati and Y. Ollivier · 2021
Closest in time.