Fetching the paper…
Reading the bibliography…
A highly desirable property of a reinforcement learning (RL) agent -- and a major difficulty for deep RL approaches -- is the ability to generalize policies learned on a few tasks over a high-dimensional observation space to similar tasks not seen during training.
Silhouettes: a graphical aid to the interpretation and validation of cluster analysis
P. J. Rousseeuw · 1987
Earlier work this paper cites.
Metrics for finite Markov decision processes
N. Ferns, P. Panangaden, and D. Precup · 2004
Earlier work this paper cites.
Optimal transport: old and new
C. Villani · 2008
Earlier work this paper cites.
Sinkhorn distances: Lightspeed computation of optimal transportation distances
M. Cuturi · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Earlier work this paper cites.
Successor features for transfer in reinforcement learning
A. Barreto, W. Dabney, R. Munos, J. J. Hunt, T. Schaul, H. Van Hasselt, and D. Silver · 2016
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Earlier work this paper cites.
Generalization and exploration via randomized value functions
B. Van Roy and Z. Wen · 2016
Earlier work this paper cites.
Generating visual representations for zero-shot classification
M. Bucher, S. Herbin, and F. Jurie · 2017
Earlier work this paper cites.
DARLA: Improving zero-shot transfer in reinforcement learning
I. Higgins, A. Pal, A. A. Rusu, L. Matthey, C. P. Burgess, A. Pritzel, M. Botvinick, C. Blundell, and A. Lerchner · 2017
Earlier work this paper cites.
Reinforcement learning with unsupervised auxiliary tasks
M. Jaderberg, V. Mnih, W. M. Czarnecki, T. Schaul, J. Z. Leibo, D. Silver, and K. Kavukcuoglu · 2017
Earlier work this paper cites.
Zero-shot task generalization with multi-task deep reinforcement learning
J. Oh, S. Singh, H. Lee, and P. Kohli · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Unsupervised meta-learning for reinforcement learning
A. Gupta, B. Eysenbach, C. Finn, and S. Levine · 2018
Earlier work this paper cites.
Learning deep representations by mutual information estimation and maximization
R. D. Hjelm, A. Fedorov, S. Lavoie-Marchildon, K. Grewal, P. Bachman, A. Trischler, and Y. Bengio · 2018
Earlier work this paper cites.
Film: Visual reasoning with a general conditioning layer
E. Perez, F. Strub, H. de Vries, V. Dumoulin, and A. Courville · 2018
Earlier work this paper cites.
Hierarchical reinforcement learning for zero-shot generalization with subtask dependencies
S. Sohn, J. Oh, and H. Lee · 2018
Earlier work this paper cites.
Unsupervised state representation learning in Atari
A. Anand, E. Racah, S. Ozair, Y. Bengio, M.-A. Côté, and R. D. Hjelm · 2019
Cited alongside, same era.
Learning representations by maximizing mutual information across views
P. Bachman, R. D. Hjelm, and W. Buchwalter · 2019
Cited alongside, same era.
Diversity is all you need: Learning skills without a reward function
B. Eysenbach, A. Gupta, J. Ibarz, and S. Levine · 2019
Cited alongside, same era.
DeepMDP: Learning continuous latent space models for representation learning
C. Gelada, S. Kumar, J. Buckman, O. Nachum, and M. G. Bellemare · 2019
Cited alongside, same era.
Learning to drive in a day
A. Kendall, J. Hawke, D. Janz, P. Mazur, D. Reda, J.-M. Allen, V.-D. Lam, A. Bewley, and A. Shah · 2019
Cited alongside, same era.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
K. Rakelly, A. Zhou, C. Finn, S. Levine, and D. Quillen · 2019
Locality and compositionality in zero-shot learning
T. Sylvain, L. Petrini, and D. Hjelm · 2020
Later among the works it cites.
Plannable approximations to mdp homomorphisms: Equivariance under actions
E. van der Pol, T. Kipf, F. A. Oliehoek, and M. Welling · 2020
Later among the works it cites.
Self-supervised domain-aware generative network for generalized zero-shot learning
J. Wu, T. Zhang, Z.-J. Zha, J. Luo, Y. Zhang, and F. Wu · 2020
Later among the works it cites.
Learning robust state abstractions for hidden-parameter block mdps
A. Zhang, S. Sodhani, K. Khetarpal, and J. Pineau · 2020
Later among the works it cites.
Contrastive behavioral similarity embeddings for generalization in reinforcement learning
R. Agarwal, M. C. Machado, P. S. Castro, and M. G. Bellemare · 2021
Closest in time.
Mine your own view: Self-supervised learning through across-sample prediction
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Representation learning with contrastive predictive coding
A. van den Oord, Y. Li, and O. Vinyals · 2019
Cited alongside, same era.
Pc-pg: Policy cover directed exploration for provable policy gradient learning
A. Agarwal, M. Henaff, S. Kakade, and W. Sun · 2020
Cited alongside, same era.
Self-labelling via simultaneous clustering and representation learning
Y. M. Asano, C. Rupprecht, and A. Vedaldi · 2020
Cited alongside, same era.
Unsupervised learning of visual features by contrasting cluster assignments
M. Caron, I. Misra, J. Mairal, P. Goyal, P. Bojanowski, and A. Joulin · 2020
Cited alongside, same era.
A simple framework for contrastive learning of visual representations
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton · 2020
Cited alongside, same era.
Leveraging procedural generation to benchmark reinforcement learning
K. Cobbe, C. Hesse, J. Hilton, and J. Schulman · 2020
Cited alongside, same era.
M. Azabou, M. G. Azar, R. Liu, C.-H. Lin, E. C. Johnson, K. Bhaskaran-Nair, M. Dabagia, K. B. Hengen, W. Gray-Roncal, M. Valko, and E. L. Dyer · 2021
Closest in time.
Phasic policy gradient
K. Cobbe, J. Hilton, O. Klimov, and J. Schulman · 2021
Closest in time.
The value-improvement path: Towards better representations for reinforcement learning
W. Dabney, A. Barreto, M. Rowland, R. Dadashi, J. Quan, M. G. Bellemare, and D. Silver · 2021
Closest in time.
Domain adversarial reinforcement learning
B. Li, V. François-Lavet, T. Doan, and J. Pineau · 2021
Closest in time.
S. Mohanty, J. Poonganam, A. Gaidon, A. Kolobov, B. Wulfe, D. Chakraborty, G. Šemetulskis, J. Schapke, J. Kubilius, J. Pašukonis, et al · 2021
Closest in time.
Decoupling value and policy for generalization in reinforcement learning
R. Raileanu and R. Fergus · 2021
Closest in time.
Data-efficient reinforcement learning with self-predictive representations
M. Schwarzer, A. Anand, R. Goel, R. D. Hjelm, A. Courville, and P. Bachman · 2021
Closest in time.
Decoupling representation learning from reinforcement learning
A. Stooke, K. Lee, P. Abbeel, and M. Laskin · 2021
Closest in time.
Learning one representation to optimize all rewards
A. Touati and Y. Ollivier · 2021
Closest in time.
Representation matters: Offline pretraining for sequential decision making
M. Yang and O. Nachum · 2021
Closest in time.
Reinforcement learning with prototypical representations
D. Yarats, R. Fergus, A. Lazaric, and L. Pinto · 2021
Closest in time.
Learning invariant representations for reinforcement learning without reconstruction
A. Zhang, R. McAllister, R. Calandra, Y. Gal, and S. Levine · 2021
Closest in time.