Fetching the paper…
Reading the bibliography…
Intelligent agents can cope with sensory-rich environments by learning task-agnostic state abstractions.
Optimal control of markov processes with incomplete state information
Karl J Aastrom · 1965
Earlier work this paper cites.
Adaptive aggregation for infinite horizon dynamic programming
Dimitri Bertsekas and David Castanon · 1989
Earlier work this paper cites.
Inferring statistical complexity
James P Crutchfield and Karl Young · 1989
Earlier work this paper cites.
Bisimulation through probabilistic testing (preliminary report)
Kim G. Larsen and Arne Skou · 1989
Earlier work this paper cites.
Learning algorithms for networks with internal and external feedback
Jürgen Schmidhuber · 1990
Earlier work this paper cites.
Extracting refined rules from knowledge-based neural networks
Geoffrey G Towell and Jude W Shavlik · 1993
Earlier work this paper cites.
Acting optimally in partially observable stochastic domains
Anthony R Cassandra, Leslie Pack Kaelbling, and Michael L Littman · 1994
Earlier work this paper cites.
Rule generation from neural networks
LiMin Fu · 1994
Earlier work this paper cites.
Effective data mining using neural networks
Hongjun Lu, Rudy Setiono, and Huan Liu · 1996
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Leslie Pack Kaelbling, Michael L Littman, and Anthony R Cassandra · 1998
Earlier work this paper cites.
Predictive representations of state
Michael L Littman and Richard S Sutton · 2001
Earlier work this paper cites.
Computational mechanics: Pattern and prediction, structure and simplicity
Cosma Rohilla Shalizi and James P Crutchfield · 2001
Earlier work this paper cites.
Equivalence notions and model minimization in markov decision processes
Robert Givan, Thomas L. Dean, and Matthew Greig · 2003
Earlier work this paper cites.
Metrics for finite markov decision processes
Norm Ferns, Prakash Panangaden, and Doina Precup · 2004
Earlier work this paper cites.
Blind construction of optimal nonlinear recursive predictors for discrete sequences
Cosma Rohilla Shalizi and Kristina Lisa Shalizi · 2004
Earlier work this paper cites.
Predictive state representations: A new theory for modeling dynamical systems
Satinder Singh, Michael R James, and Matthew R Rudary · 2004
Earlier work this paper cites.
Finding approximate POMDP solutions through belief compression
Nicholas Roy, Geoffrey Gordon, and Sebastian Thrun · 2005
Earlier work this paper cites.
Plan2vec: Unsupervised representation learning by latent plans
Ge Yang, Amy Zhang, Ari S. Morcos, Joelle Pineau, Pieter Abbeel, and Roberto Calandra · 2005
Earlier work this paper cites.
Towards a unified theory of state abstraction for MDPs
Lihong Li, Thomas Walsh, and Michael Littman · 2006
Earlier work this paper cites.
Performance loss bounds for approximate value iteration with state aggregation
Benjamin Van Roy · 2006
Earlier work this paper cites.
Learning invariant representations for reinforcement learning without reconstruction
Amy Zhang, Rowan McAllister, Roberto Calandra, Yarin Gal, and Sergey Levine · 2006
Earlier work this paper cites.
Interventions and causal inference
Frederick Eberhardt and Richard Scheines · 2007
Earlier work this paper cites.
Representing systems with hidden state
Christopher Hundt, Prakash Panangaden, Joelle Pineau, and Doina Precup · 2007
Earlier work this paper cites.
Equivalence relations in fully and partially observable markov decision processes
Pablo Samuel Castro, Prakash Panangaden, and Doina Precup · 2009
Cited alongside, same era.
Causality: Models, Reasoning and Inference
Judea Pearl · 2009
Cited alongside, same era.
Information-theoretic approach to interactive learning
Susanne Still · 2009
Cited alongside, same era.
Bounding performance loss in approximate MDP homomorphisms
Jonathan Taylor, Doina Precup, and Prakash Panagaden · 2009
Cited alongside, same era.
Using bisimulation for policy transfer in MDPs
Pablo Samuel Castro and Doina Precup · 2010
Cited alongside, same era.
Synchronization and control in intrinsic and designed computation: An information-theoretic analysis of competing models of stochastic computation
James P Crutchfield, Christopher J Ellison, Ryan G James, and John R Mahoney · 2010
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Later among the works it cites.
Rule extraction algorithm for deep neural networks: A review
Tameru Hailesilassie · 2016
Later among the works it cites.
Improving predictive state representations via gradient descent
Nan Jiang, Alex Kulesza, and Satinder Singh · 2016
Later among the works it cites.
ViZDoom: A Doom-based AI research platform for visual reinforcement learning
Michał Kempka, Marek Wydmuch, Grzegorz Runc, Jakub Toczek, and Wojciech Jaśkowski · 2016
Later among the works it cites.
Causal inference using invariant prediction: identification and confidence intervals
Jonas Peters, Peter Bühlmann, and Nicolai Meinshausen · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Bisimulation metrics for continuous markov decision processes
Norm Ferns, Prakash Panangaden, and Doina Precup · 2011
Cited alongside, same era.
A monte-carlo aixi approximation
Joel Veness, Kee Siong Ng, Marcus Hutter, William Uther, and David Silver · 2011
Cited alongside, same era.
On causal and anticausal learning
Bernhard Schölkopf, Dominik Janzing, Jonas Peters, Eleni Sgouritsa, Kun Zhang, and Joris Mooij · 2012
Cited alongside, same era.
An information-theoretic approach to curiosity-driven reinforcement learning
Susanne Still and Doina Precup · 2012
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Marc G. Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Cited alongside, same era.
Hilbert space embeddings of predictive state representations
Byron Boots, Geoffrey Gordon, and Arthur Gretton · 2013
Cited alongside, same era.
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
Matteo Hessel, Joseph Modayil, Hado van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Gheshlaghi Azar, and David Silver · 2017
Later among the works it cites.
Neural discrete representation learning
Aaron van den Oord, Oriol Vinyals, and koray kavukcuoglu · 2017
Later among the works it cites.
An empirical evaluation of recurrent neural network rule extraction
Qinglong Wang, Kaixuan Zhang, II Ororbia, G Alexander, Xinyu Xing, Xue Liu, and C Lee Giles · 2017
Later among the works it cites.
Extracting automata from recurrent neural networks using queries and counterexamples
Gail Weiss, Yoav Goldberg, and Eran Yahav · 2017
Later among the works it cites.
Efficient model-based deep reinforcement learning with variational state tabulation
Dane S. Corneil, Wulfram Gerstner, and Johanni Brea · 2018
Later among the works it cites.
Neural predictive belief representations
Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, Bilal Piot, Bernardo A. Pires, Toby Pohlen, and Rémi Munos · 2018
Later among the works it cites.
Recurrent predictive state policy networks
Ahmed Hefny, Zita Marinho, Wen Sun, Siddhartha Srinivasa, and Geoffrey Gordon · 2018
Later among the works it cites.
Deep variational reinforcement learning for POMDPs
Maximilian Igl, Luisa Zintgraf, Tuan Anh Le, Frank Wood, and Shimon Whiteson · 2018
Later among the works it cites.
Learning finite state representations of recurrent policy networks
Anurag Koul, Sam Greydanus, and Alan Fern · 2018
Later among the works it cites.
Variational inference for data-efficient model learning in POMDPs
Sebastian Tschiatschek, Kai Arulkumaran, Jan Stühmer, and Katja Hofmann · 2018
Later among the works it cites.
Qinglong Wang, Kaixuan Zhang, II Ororbia, G Alexander, Xinyu Xing, Xue Liu, and C Lee Giles · 2018
Later among the works it cites.
Composable planning with attributes
Amy Zhang, Adam Lerer, Sainbayar Sukhbaatar, Rob Fergus, and Arthur Szlam · 2018
Later among the works it cites.
Invariant Risk Minimization
Martin Arjovsky, Leon Bottou, Ishaan Gulrajani, and David Lopez-Paz · 2019
Closest in time.
DeepMDP: Learning continuous latent space models for representation learning
Carles Gelada, Saurabh Kumar, Jacob Buckman, Ofir Nachum, and Marc G. Bellemare · 2019
Closest in time.
Causality for machine learning
Bernhard Schölkopf · 2019
Closest in time.
Scalable methods for computing state similarity in deterministic Markov decision processes
Pablo Samuel Castro · 2020
Closest in time.
Mastering atari with discrete world models, 2020
Danijar Hafner, Timothy Lillicrap, Mohammad Norouzi, and Jimmy Ba · 2020
Closest in time.