Fetching the paper…
Reading the bibliography…
Human decision-making often involves combining similar states into categories and reasoning at the level of the categories rather than the actual states.
Model-based reinforcement learning for atari
Kaiser, L.; Babaeizadeh, M.; Milos, P.; Osinski, B.; Campbell, R. H.; Czechowski, K.; Erhan, D.; Finn, C.; Kozakowski, P.; Levine, S.; et al. 2019 · 1903
Earlier work this paper cites.
Towards interpretable reinforcement learning using attention augmented agents
Mott, A.; Zoran, D.; Chrzanowski, M.; Wierstra, D.; and Rezende, D. J. 2019 · 1906
Earlier work this paper cites.
Stochastic latent actor-critic: Deep reinforcement learning with a latent variable model
Lee, A. X.; Nagabandi, A.; Abbeel, P.; and Levine, S. 2019a · 1907
Earlier work this paper cites.
Graying the black box: Understanding dqns
Zahavy, T.; Ben-Zrihem, N.; and Mannor, S. 2016 · 1908
Earlier work this paper cites.
Generalization in reinforcement learning with selective noise injection and information bottleneck
Igl, M.; Ciosek, K.; Li, Y.; Tschiatschek, S.; Zhang, C.; Devlin, S.; and Hofmann, K. 2019 · 1910
Earlier work this paper cites.
Network randomization: A simple technique for generalization in deep reinforcement learning
Lee, K.; Lee, K.; Shin, J.; and Lee, H. 2019b · 1910
Earlier work this paper cites.
Neuronlike adaptive elements that can solve difficult learning control problems
Barto, A. G.; Sutton, R. S.; and Anderson, C. W. 1983 · 1983
Earlier work this paper cites.
Salt-and-Pepper Noise Removal by Median-Type Noise Detectors and Detail-Preserving Regularization
Chan, R. H.; Ho, C.; and Nikolova, M. 2005 · 2005
Earlier work this paper cites.
Transformer vq-vae for unsupervised unit discovery and speech synthesis: Zerospeech 2020 challenge
Tjandra, A.; Sakti, S.; and Nakamura, S. 2020 · 2005
Earlier work this paper cites.
An object-oriented representation for efficient reinforcement learning
Diuk, C.; Cohen, A.; and Littman, M. L. 2008 · 2008
Earlier work this paper cites.
Visualizing data using t-SNE
Van der Maaten, L.; and Hinton, G. 2008 · 2008
Earlier work this paper cites.
Discrete Latent Space World Models for Reinforcement Learning
Robine, J.; Uelwer, T.; and Harmeling, S. 2020 · 2010
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P.; and Welling, M. 2013 · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A. A.; Veness, J.; Bellemare, M. G.; Graves, A.; Riedmiller, M.; Fidjeland, A. K.; Ostrovski, G.; et al. 2015 · 2015
Earlier work this paper cites.
Reinforcement learning with unsupervised auxiliary tasks
Jaderberg, M.; Mnih, V.; Czarnecki, W. M.; Schaul, T.; Leibo, J. Z.; Silver, D.; and Kavukcuoglu, K. 2016 · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Levine, S.; Finn, C.; Darrell, T.; and Abbeel, P. 2016 · 2016
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
Silver, D.; Huang, A.; Maddison, C. J.; Guez, A.; Sifre, L.; Van Den Driessche, G.; Schrittwieser, J.; Antonoglou, I.; Panneershelvam, V.; Lanctot, M.; et al. 2016 · 2016
Cited alongside, same era.
Deep reinforcement learning: A brief survey
Arulkumaran, K.; Deisenroth, M. P.; Brundage, M.; and Bharath, A. A. 2017 · 2017
Cited alongside, same era.
A distributional perspective on reinforcement learning
Bellemare, M. G.; Dabney, W.; and Munos, R. 2017 · 2017
Cited alongside, same era.
Quantifying generalization in reinforcement learning
Cobbe, K.; Klimov, O.; Hesse, C.; Kim, T.; and Schulman, J. 2019 · 2019
Later among the works it cites.
Combined reinforcement learning via abstract representations
François-Lavet, V.; Bengio, Y.; Precup, D.; and Pineau, J. 2019 · 2019
Later among the works it cites.
Low bit-rate speech coding with VQ-VAE and a WaveNet decoder
Gârbacea, C.; van den Oord, A.; Li, Y.; Lim, F. S.; Luebs, A.; Vinyals, O.; and Walters, T. C. 2019 · 2019
Later among the works it cites.
Low Bit-rate Speech Coding with VQ-VAE and a WaveNet Decoder
Gârbacea, C.; van den Oord, A.; Li, Y.; Lim, F. S. C.; Luebs, A.; Vinyals, O.; and Walters, T. C. 2019 · 2019
Later among the works it cites.
Human-level performance in 3D multiplayer games with population-based reinforcement learning
Jaderberg, M.; Czarnecki, W. M.; Dunning, I.; Marris, L.; Lever, G.; Castaneda, A. G.; Beattie, C.; Rabinowitz, N. C.; Morcos, A. S.; Ruderman, A.; et al. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Oord, A. v. d.; Vinyals, O.; and Kavukcuoglu, K. 2017 · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017 · 2017
Cited alongside, same era.
Minimalistic Gridworld Environment for OpenAI Gym
Chevalier-Boisvert, M.; Willems, L.; and Pal, S. 2018 · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
Oord, A. v. d.; Li, Y.; and Vinyals, O. 2018 · 2018
Cited alongside, same era.
Assessing generalization in deep reinforcement learning
Packer, C.; Gao, K.; Kos, J.; Krähenbühl, P.; Koltun, V.; and Song, D. 2018 · 2018
Cited alongside, same era.
An atari model zoo for analyzing, visualizing, and comparing deep reinforcement learning agents
Such, F. P.; Madhavan, V.; Liu, R.; Wang, R.; Castro, P. S.; Li, Y.; Zhi, J.; Schubert, L.; Bellemare, M. G.; Clune, J.; et al. 2018 · 2018
Cited alongside, same era.
Programmatically interpretable reinforcement learning
Verma, A.; Murali, V.; Singh, R.; Kohli, P.; and Chaudhuri, S. 2018 · 2018
Cited alongside, same era.
SDRL: interpretable and data-efficient deep reinforcement learning leveraging symbolic planning
Lyu, D.; Yang, F.; Liu, B.; and Gustafson, S. 2019 · 2019
Later among the works it cites.
Generating diverse high-fidelity images with vq-vae-2
Razavi, A.; van den Oord, A.; and Vinyals, O. 2019 · 2019
Later among the works it cites.
Action robust reinforcement learning and applications in continuous control
Tessler, C.; Efroni, Y.; and Mannor, S. 2019 · 2019
Later among the works it cites.
Bootstrap latent-predictive representations for multitask reinforcement learning
Guo, Z. D.; Pires, B. A.; Piot, B.; Grill, J.-B.; Altché, F.; Munos, R.; and Azar, M. G. 2020 · 2020
Later among the works it cites.
Certified adversarial robustness for deep reinforcement learning
Lütjens, B.; Everett, M.; and How, J. P. 2020 · 2020
Later among the works it cites.
Reinforcement learning with perturbed rewards
Wang, J.; Liu, Y.; and Li, B. 2020 · 2020
Later among the works it cites.
Decoupling representation learning from reinforcement learning
Stooke, A.; Lee, K.; Abbeel, P.; and Laskin, M. 2021 · 2021
Later among the works it cites.
VideoGPT: Video Generation using VQ-VAE and Transformers
Yan, W.; Zhang, Y.; Abbeel, P.; and Srinivas, A. 2021 · 2021
Later among the works it cites.
Leveraging Procedural Generation to Benchmark Reinforcement Learning
Cobbe, K.; Hesse, C.; Hilton, J.; and Schulman, J. 2020 · 2056
Closest in time.