Fetching the paper…
Reading the bibliography…
Many reinforcement learning (RL) environments consist of independent entities that interact sparsely.
The development and generalization of contingency awareness in early infancy: Some hypotheses
John S. Watson · 1966
Earlier work this paper cites.
Investigating causal relations by econometric models and cross-spectral methods
C. W. J. Granger · 1969
Earlier work this paper cites.
Curious model-building control systems
J. Schmidhuber · 1991
Earlier work this paper cites.
Reinforcement driven information acquisition in non-deterministic environments
Jan Storck, Sepp Hochreiter, and Jürgen Schmidhuber · 1995
Earlier work this paper cites.
Causation, Prediction, and Search
P. Spirtes, C. Glymour, and R. Scheines · 2000
Earlier work this paper cites.
Measuring information transfer
Thomas Schreiber · 2000
Earlier work this paper cites.
Information-theoretic approach to the study of control systems
Hugo Touchette and Seth Lloyd · 2003
Earlier work this paper cites.
Estimating mutual information
Alexander Kraskov, Harald Stögbauer, and Peter Grassberger · 2004
Earlier work this paper cites.
Empowerment: a universal agent-centric measure of control
A.S. Klyubin, D. Polani, and C.L. Nehaniv · 2005
Earlier work this paper cites.
Geometric robustness theory and biological networks
Nihat Ay and David C. Krakauer · 2006
Earlier work this paper cites.
Elements of Information Theory
Thomas M. Cover and Joy A. Thomas · 2006
Earlier work this paper cites.
Approximating the Kullback Leibler divergence between gaussian mixture models
John R. Hershey and Peder A. Olsen · 2007
Earlier work this paper cites.
Information flows in causal networks
N. Ay and D. Polani · 2008
Earlier work this paper cites.
Causality: Models, Reasoning and Inference
Judea Pearl · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Identifiability of causal graphs using functional models
J. Peters, J. Mooij, D. Janzing, and B. Schölkopf · 2011
Earlier work this paper cites.
Causal Inference in Time Series Analysis
Michael Eichler · 2012
Earlier work this paper cites.
Lower and upper bounds for approximation of the kullback-leibler divergence between gaussian mixture models
Jean-Louis Durrieu, J. Thiran, and Finnian Kelly · 2012
Earlier work this paper cites.
MuJoCo: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
The local information dynamics of distributed computation in complex systems
Joseph T. Lizier · 2013
Earlier work this paper cites.
Quantifying causal influences
D. Janzing, D. Balduzzi, M. Grosse-Wentrup, and B. Schölkopf · 2013
Earlier work this paper cites.
Learning and exploration in action-perception loops
Daniel Little and Friedrich Sommer · 2013
Earlier work this paper cites.
Information driven self-organization of complex robotic behaviors
Georg Martius, Ralf Der, and Nihat Ay · 2013
Earlier work this paper cites.
Linear combination of one-step predictive information with an external reward in an episodic policy gradient setting: a critical analysis
Keyan Zahedi, Georg Martius, and Nihat Ay · 2013
Cited alongside, same era.
Empowerment–An Introduction , pages 67–114
Christoph Salge, Cornelius Glackin, and Daniel Polani · 2014
Cited alongside, same era.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M. Saxe, James L. McClelland, and S. Ganguli · 2014
Cited alongside, same era.
Bandits with unobserved confounders: A causal approach
E. Bareinboim, A. Forney, and J. Pearl · 2015
Cited alongside, same era.
Variational information maximisation for intrinsically motivated reinforcement learning
Shakir Mohamed and Danilo Jimenez Rezende · 2015
Cited alongside, same era.
Nonparametric von Mises estimators for entropies, divergences and mutual informations
Multi-goal reinforcement learning: Challenging robotics environments and request for research
Matthias Plappert, Marcin Andrychowicz, Alex Ray, Bob McGrew, Bowen Baker, Glenn Powell, J. Schneider, Joshua Tobin, Maciek Chociej, P. Welinder, V. Kumar, and W. Zaremba · 2018
Later among the works it cites.
Energy-based hindsight experience prioritization
Rui Zhao and Volker Tresp · 2018
Later among the works it cites.
Spectral normalization for generative adversarial networks
Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Y. Yoshida · 2018
Later among the works it cites.
Learning neural causal models from unknown interventions
Nan Rosemary Ke, Olexa Bilaniuk, Anirudh Goyal, Stefan Bauer, Hugo Larochelle, Bernhard Schölkopf, Michael C. Mozer, Chris Pal, and Yoshua Bengio · 2019
Later among the works it cites.
Woulda, Coulda, Shoulda: Counterfactually-Guided Policy Search
Lars Buesing, T. Weber, Yori Zwols, Sébastien Racanière, A. Guez, J. Lespiau, and N. Heess · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kirthevasan Kandasamy, Akshay Krishnamurthy, Barnabas Poczos, Larry Wasserman, and james m robins · 2015
Cited alongside, same era.
Efficient Estimation of Mutual Information for Strongly Dependent Variables
Shuyang Gao, Greg Ver Steeg, and Aram Galstyan · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Vime: Variational information maximizing exploration
Rein Houthooft, Xi Chen, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2016
Cited alongside, same era.
Actual Causality
Joseph Y. Halpern · 2016
Cited alongside, same era.
Causal Inference in Statistics: A Primer
J. Pearl, M. Glymour, and N. Jewell · 2016
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, J. Schneider, John Schulman, Jie Tang, and W. Zaremba · 2016
Cited alongside, same era.
Later among the works it cites.
Near-optimal reinforcement learning in dynamic treatment regimes
J. Zhang and E. Bareinboim · 2019
Later among the works it cites.
Control What You Can: Intrinsically motivated task-planning agent
Sebastian Blaes, Marin Vlastelica, Jia-Jie Zhu, and Georg Martius · 2019
Later among the works it cites.
CURIOUS: intrinsically motivated modular multi-goal reinforcement learning
Cédric Colas, Pierre-Yves Oudeyer, Olivier Sigaud, Pierre Fournier, and Mohamed Chetouani · 2019
Later among the works it cites.
Contingency-aware exploration in reinforcement learning
Jongwook Choi, Yijie Guo, Marcin Moczulski, Junhyuk Oh, Neal Wu, Mohammad Norouzi, and Honglak Lee · 2019
Later among the works it cites.
On variational bounds of mutual information
Ben Poole, Sherjil Ozair, Aaron Van Den Oord, Alex Alemi, and George Tucker · 2019
Later among the works it cites.
Self-supervised exploration via disagreement
Deepak Pathak, Dhiraj Gandhi, and Abhinav Gupta · 2019
Later among the works it cites.
Exploration via hindsight goal generation
Zhizhou Ren, Kefan Dong, Yuan Zhou, Qiang Liu, and Jian Peng · 2019
Later among the works it cites.
Learning dexterous in-hand manipulation
Marcin Andrychowicz, Bowen Baker, Maciek Chociej, Rafal Jozefowicz, Bob McGrew, Jakub Pachocki, Arthur Petron, Matthias Plappert, Glenn Powell, Alex Ray, et al · 2020
Later among the works it cites.
A meta-transfer objective for learning to dis entangle causal mechanisms
Yoshua Bengio, Tristan Deleu, Nasim Rahaman, Nan Rosemary Ke, Sebastien Lachapelle, Olexa Bilaniuk, Anirudh Goyal, and Christopher Pal · 2020
Later among the works it cites.
Weakly-supervised disentanglement without compromises
Francesco Locatello, Ben Poole, Gunnar Raetsch, Bernhard Schölkopf, Olivier Bachem, and Michael Tschannen · 2020
Later among the works it cites.
Causally Correct Partial Models for Reinforcement Learning
Danilo Jimenez Rezende, Ivo Danihelka, George Papamakarios, N. Ke, Ray Jiang, T. Weber, K. Gregor, Hamza Merzic, Fabio Viola, J. Wang, Jovana Mitrovic, F. Besse, Ioannis Antonoglou, and Lars Buesing · 2020
Later among the works it cites.
Counterfactual data augmentation using locally factored dynamics
Silviu Pitis, Elliot Creager, and Animesh Garg · 2020
Later among the works it cites.
Mega-Reward: Achieving human-level play without extrinsic rewards
Yuhang Song, Jianyi Wang, Thomas Lukasiewicz, Zhenghua Xu, Shangtong Zhang, Andrzej Wojcicki, and Mai Xu · 2020
Later among the works it cites.
Inductive biases for deep learning of higher-level cognition
Anirudh Goyal and Yoshua Bengio · 2020
Later among the works it cites.
Towards causal representation learning
B. Schölkopf, F. Locatello, S. Bauer, R. Nan Ke, N. Kalchbrenner, A. Goyal, and Y. Bengio · 2021
Closest in time.
Recurrent independent mechanisms
A. Goyal, A. Lamb, J. Hoffmann, S. Sodhani, S. Levine, Y. Bengio, and B. Schölkopf · 2021
Closest in time.
Causal curiosity: Rl agents discovering self-supervised experiments for causal representation learning
S. Sontakke, A. Mehrjou, L. Itti, and B. Schölkopf · 2021
Closest in time.
Mutual information state intrinsic control
Rui Zhao, Yang Gao, Pieter Abbeel, Volker Tresp, and Wei Xu · 2021
Closest in time.