Fetching the paper…
Reading the bibliography…
In many sequential decision-making tasks, the agent is not able to model the full complexity of the world, which consists of multitudes of relevant and irrelevant information.
Learning policies for partially observable environments: Scaling up
Michael L. Littman, Anthony R. Cassandra, and Leslie Pack Kaelbling · 1995
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Leslie Pack Kaelbling, Michael L. Littman, and Anthony R. Cassandra · 1998
Earlier work this paper cites.
System Identification
Lennart Ljung · 1998
Earlier work this paper cites.
A solution to the simultaneous localization and map building (slam) problem
MWM Gamini Dissanayake, Paul Newman, Steve Clark, Hugh F Durrant-Whyte, and Michael Csorba · 2001
Earlier work this paper cites.
Predictive representations of state
Michael L. Littman, Richard S. Sutton, and Satinder Singh · 2001
Earlier work this paper cites.
All else being equal be empowered
Alexander S. Klyubin, Daniel Polani, and Chrystopher L. Nehaniv · 2005
Earlier work this paper cites.
Simultaneous localization and mapping: part i
Hugh Durrant-Whyte and Tim Bailey · 2006
Earlier work this paper cites.
What is intrinsic motivation? A typology of computational approaches
Pierre-Yves Oudeyer and Frédéric Kaplan · 2007
Earlier work this paper cites.
Autonomous reinforcement learning on raw visual input data in a real world application
Sascha Lange, Martin A. Riedmiller, and Arne Voigtländer · 2012
Earlier work this paper cites.
Comparison of co-expression measures: mutual information, correlation, and model based indices
Lin Song, Peter Langfelder, and Steve Horvath · 2012
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Kingma Diederik, Ba Jimmy, et al · 2014
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Earlier work this paper cites.
Variational information maximisation for intrinsically motivated reinforcement learning
Shakir Mohamed and Danilo Jimenez Rezende · 2015
Earlier work this paper cites.
Incentivizing exploration in reinforcement learning with deep predictive models
Bradly C. Stadie, Sergey Levine, and Pieter Abbeel · 2015
Earlier work this paper cites.
From pixels to torques: Policy learning with deep dynamical models
Niklas Wahlström, Thomas B. Schön, and Marc Peter Deisenroth · 2015
Earlier work this paper cites.
Embed to control: A locally linear latent dynamics model for control from raw images
Manuel Watter, Jost Tobias Springenberg, Joschka Boedecker, and Martin A. Riedmiller · 2015
Earlier work this paper cites.
Past, present, and future of simultaneous localization and mapping: Toward the robust-perception age
Cesar Cadena, Luca Carlone, Henry Carrillo, Yasir Latif, Davide Scaramuzza, José Neira, Ian Reid, and John J Leonard · 2016
Earlier work this paper cites.
Deep Learning , volume 1
Ian Goodfellow, Yoshua Bengio, Aaron Courville, and Yoshua Bengio · 2016
Earlier work this paper cites.
Gaussian Error Linear Units (GELUs), 2016
Dan Hendrycks and Kevin Gimpel · 2016
Cited alongside, same era.
Deep variational information bottleneck
Alexander A Alemi, Ian Fischer, Joshua V Dillon, and Kevin Murphy · 2017
Cited alongside, same era.
Matterport3D: Learning from RGB-D data in indoor environments
Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Niessner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang · 2017
Cited alongside, same era.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2017
Cited alongside, same era.
Markov chains and mixing times , volume 107
David A Levin and Yuval Peres · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell · 2017
Cited alongside, same era.
Vector quantization - pytorch
Phil Wang · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Later among the works it cites.
Learning task informed abstractions
Xiang Fu, Ge Yang, Pulkit Agrawal, and Tommi Jaakkola · 2021
Later among the works it cites.
Necessary and sufficient conditions for causal feature selection in time series with latent common causes
Atalanti A Mastakouri, Bernhard Schölkopf, and Dominik Janzing · 2021
Later among the works it cites.
Which mutual-information representation learning objectives are sufficient for control?
Kate Rakelly, Abhishek Gupta, Carlos Florensa, and Sergey Levine · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Neural discrete representation learning
Aaron van den Oord, Oriol Vinyals, and koray kavukcuoglu · 2017
Cited alongside, same era.
Neural discrete representation learning
Aaron Van Den Oord, Oriol Vinyals, et al · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Discovering and removing exogenous state variables and rewards for reinforcement learning
Thomas G. Dietterich, George Trimponias, and Zhitang Chen · 2018
Cited alongside, same era.
Dota 2 with large scale deep reinforcement learning
Christopher Berner, Greg Brockman, Brooke Chan, Vicki Cheung, Przemysław Dębiak, Christy Dennison, David Farhi, Quirin Fischer, Shariq Hashme, Chris Hesse, et al · 2019
Cited alongside, same era.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2019
Cited alongside, same era.
Nino Scherrer, Olexa Bilaniuk, Yashas Annadani, Anirudh Goyal, Patrick Schwab, Bernhard Schölkopf, Michael C Mozer, Yoshua Bengio, Stefan Bauer, and Nan Rosemary Ke · 2021
Later among the works it cites.
Data-efficient reinforcement learning with self-predictive representations
Max Schwarzer, Ankesh Anand, Rishab Goel, R. Devon Hjelm, Aaron C. Courville, and Philip Bachman · 2021
Later among the works it cites.
Pretraining representations for data-efficient reinforcement learning
Max Schwarzer, Nitarshan Rajkumar, Michael Noukhovitch, Ankesh Anand, Laurent Charlin, R. Devon Hjelm, Philip Bachman, and Aaron C. Courville · 2021
Later among the works it cites.
Mlp-mixer: An all-mlp architecture for vision
Ilya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, Mario Lucic, and Alexey Dosovitskiy · 2021
Later among the works it cites.
Mastering atari games with limited data
Weirui Ye, Shaohuai Liu, Thanard Kurutach, Pieter Abbeel, and Yang Gao · 2021
Later among the works it cites.
Soundstream: An end-to-end neural audio codec, 2021
Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi · 2021
Later among the works it cites.
Learning invariant representations for reinforcement learning without reconstruction
Amy Zhang, Rowan Thomas McAllister, Roberto Calandra, Yarin Gal, and Sergey Levine · 2021
Later among the works it cites.
Video pretraining (vpt): Learning to act by watching unlabeled online videos
Bowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga, Jie Tang, Adrien Ecoffet, Brandon Houghton, Raul Sampedro, and Jeff Clune · 2022
Closest in time.
Information prioritization through empowerment in visual model-based RL
Homanga Bharadhwaj, Mohammad Babaeizadeh, Dumitru Erhan, and Sergey Levine · 2022
Closest in time.
Byol-explore: Exploration by bootstrapped prediction
Zhaohan Daniel Guo, Shantanu Thakoor, Miruna Pîslar, Bernardo Avila Pires, Florent Altché, Corentin Tallec, Alaa Saade, Daniele Calandriello, Jean-Bastien Grill, Yunhao Tang, et al · 2022
Closest in time.
Uniqueness and complexity of inverse mdp models, 2022
Marcus Hutter and Steven Hansen · 2022
Closest in time.
The file size of every core legend of zelda game, 2022
Teddy Michel · 2022
Closest in time.
Learning latent structural causal models
Jithendaraa Subramanian, Yashas Annadani, Ivaxi Sheth, Nan Rosemary Ke, Tristan Deleu, Stefan Bauer, Derek Nowrouzezahrai, and Samira Ebrahimi Kahou · 2022
Closest in time.