Fetching the paper…
Reading the bibliography…
Mutual information maximization provides an appealing formalism for learning representations of data.
Aggregation in dynamic programming
J. C. Bean, J. R. Birge, and R. L. Smith · 1987
Earlier work this paper cites.
Adaptive aggregation methods for infinite horizon dynamic programming
D. P. Bertsekas, D. A. Castanon, et al · 1988
Earlier work this paper cites.
Self-organization in a perceptual network
R. Linsker · 1988
Earlier work this paper cites.
Self-organizing neural network that discovers surfaces in random-dot stereograms
S. Becker and G. E. Hinton · 1992
Earlier work this paper cites.
An information-maximization approach to blind separation and blind deconvolution
A. J. Bell and T. J. Sejnowski · 1995
Earlier work this paper cites.
Reinforcement Learning with Selective Perception and Hidden State
A. K. McCallum · 1996
Earlier work this paper cites.
Model minimization in markov decision processes
T. Dean and R. Givan · 1997
Earlier work this paper cites.
Elements of information theory
T. M. Cover · 1999
Earlier work this paper cites.
Equivalence notions and model minimization in markov decision processes
R. Givan, T. Dean, and M. Greig · 2003
Earlier work this paper cites.
Smdp homomorphisms: an algebraic approach to abstraction in semi-markov decision processes
B. Ravindran and A. G. Barto · 2003
Earlier work this paper cites.
State abstraction discovery from irrelevant state variables
N. K. Jong and P. Stone · 2005
Earlier work this paper cites.
Empowerment: A universal agent-centric measure of control
A. S. Klyubin, D. Polani, and C. L. Nehaniv · 2005
Earlier work this paper cites.
Methods for computing state similarity in markov decision processes
N. Ferns, P. S. Castro, D. Precup, and P. Panangaden · 2006
Earlier work this paper cites.
Towards a unified theory of state abstraction for mdps
L. Li, T. J. Walsh, and M. L. Littman · 2006
Earlier work this paper cites.
Keep your options open: An information-based driving principle for sensorimotor systems
A. S. Klyubin, D. Polani, and C. L. Nehaniv · 2008
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
M. Gutmann and A. Hyvärinen · 2010
Earlier work this paper cites.
An information-theoretic approach to curiosity-driven reinforcement learning
S. Still and D. Precup · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Earlier work this paper cites.
Learning state representations with robotic priors
R. Jonschkowski and O. Brock · 2015
Earlier work this paper cites.
Variational information maximisation for intrinsically motivated reinforcement learning
S. Mohamed and D. J. Rezende · 2015
Earlier work this paper cites.
Embed to control: A locally linear latent dynamics model for control from raw images
M. Watter, J. Springenberg, J. Boedecker, and M. Riedmiller · 2015
Earlier work this paper cites.
Near optimal behavior via approximate state abstraction
D. Abel, D. E. Hershkowitz, and M. L. Littman · 2016
Earlier work this paper cites.
Learning to poke by poking: Experiential learning of intuitive physics
P. Agrawal, A. V. Nair, P. Abbeel, J. Malik, and S. Levine · 2016
Cited alongside, same era.
Learning to act by predicting the future
A. Dosovitskiy and V. Koltun · 2016
Cited alongside, same era.
Deep spatial autoencoders for visuomotor learning
C. Finn, X. Y. Tan, Y. Duan, T. Darrell, S. Levine, and P. Abbeel · 2016
Cited alongside, same era.
K. Gregor, D. J. Rezende, and D. Wierstra · 2016
Cited alongside, same era.
Reinforcement learning with unsupervised auxiliary tasks
M. Jaderberg, V. Mnih, W. M. Czarnecki, T. Schaul, J. Z. Leibo, D. Silver, and K. Kavukcuoglu · 2016
Cited alongside, same era.
A geometric perspective on optimal representations for reinforcement learning
M. Bellemare, W. Dabney, R. Dadashi, A. A. Taiga, P. S. Castro, N. Le Roux, D. Schuurmans, T. Lattimore, and C. Lyle · 2019
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Later among the works it cites.
Is a good representation sufficient for sample efficient reinforcement learning?
S. S. Du, S. M. Kakade, R. Wang, and L. F. Yang · 2019
Later among the works it cites.
Hyperbolic discounting and learning over multiple horizons
W. Fedus, C. Gelada, Y. Bengio, M. G. Bellemare, and H. Larochelle · 2019
Later among the works it cites.
Garage: A toolkit for reproducible reinforcement learning research
T. garage contributors · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Karl, M. Soelch, J. Bayer, and P. van der Smagt · 2016
Cited alongside, same era.
Loss is its own reward: Self-supervision for reinforcement learning
E. Shelhamer, P. Mahmoudieh, M. Argus, and T. Darrell · 2016
Cited alongside, same era.
Is maximum likelihood useful for representation learning?, 2017
F. Huszár · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Cited alongside, same era.
Independently controllable factors
V. Thomas, J. Pondard, E. Bengio, M. Sarfati, P. Beaudoin, M.-J. Meurs, J. Pineau, D. Precup, and Y. Bengio · 2017
Cited alongside, same era.
Fixing a broken elbo
A. A. Alemi, B. Poole, I. Fischer, J. V. Dillon, R. A. Saurous, and K. Murphy · 2018
Cited alongside, same era.
Mine: mutual information neural estimation
M. I. Belghazi, A. Baratin, S. Rajeswar, S. Ozair, Y. Bengio, A. Courville, and R. D. Hjelm · 2018
Cited alongside, same era.
Deepmdp: Learning continuous latent space models for representation learning
C. Gelada, S. Kumar, J. Buckman, O. Nachum, and M. G. Bellemare · 2019
Later among the works it cites.
Learning latent dynamics for planning from pixels
D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson · 2019
Later among the works it cites.
Contrastive learning of structured world models
T. Kipf, E. van der Pol, and M. Welling · 2019
Later among the works it cites.
Learning with good feature representations in bandits and in rl with a generative model
T. Lattimore and C. Szepesvari · 2019
Later among the works it cites.
Stochastic latent actor-critic: Deep reinforcement learning with a latent variable model
A. X. Lee, A. Nagabandi, P. Abbeel, and S. Levine · 2019
Later among the works it cites.
A unified bellman optimality principle combining reward maximization and empowerment
F. Leibfried, S. Pascual-Díaz, and J. Grau-Moya · 2019
Later among the works it cites.
Learning to navigate
P. Mirowski · 2019
Later among the works it cites.
On variational bounds of mutual information
B. Poole, S. Ozair, A. v. d. Oord, A. A. Alemi, and G. Tucker · 2019
Later among the works it cites.
Understanding the limitations of variational mutual information estimators
J. Song and S. Ermon · 2019
Later among the works it cites.
Contrastive bidirectional transformer for temporal representation learning
C. Sun, F. Baradel, K. Murphy, and C. Schmid · 2019
Later among the works it cites.
Comments on the du-kakade-wang-yang lower bounds
B. Van Roy and S. Dong · 2019
Later among the works it cites.
Discovery of useful questions as auxiliary tasks
V. Veeriah, M. Hessel, Z. Xu, J. Rajendran, R. L. Lewis, J. Oh, H. P. van Hasselt, D. Silver, and S. Singh · 2019
Later among the works it cites.
Unsupervised visuomotor control through distributional planning networks
T. Yu, G. Shevchuk, D. Sadigh, and C. Finn · 2019
Later among the works it cites.
Deep reinforcement and infomax learning
B. Mazoure, R. T. d. Combes, T. Doan, P. Bachman, and R. D. Hjelm · 2020
Later among the works it cites.
Predictive coding for locally-linear control
R. Shu, T. Nguyen, Y. Chow, T. Pham, K. Than, M. Ghavamzadeh, S. Ermon, and H. Bui · 2020
Later among the works it cites.
Curl: Contrastive unsupervised representations for reinforcement learning
A. Srinivas, M. Laskin, and P. Abbeel · 2020
Later among the works it cites.
Decoupling representation learning from reinforcement learning
A. Stooke, K. Lee, P. Abbeel, and M. Laskin · 2020
Later among the works it cites.
A framework for efficient robotic manipulation
A. Zhan, P. Zhao, L. Pinto, P. Abbeel, and M. Laskin · 2020
Later among the works it cites.