Fetching the paper…
Reading the bibliography…
Empowerment is an information-theoretic method that can be used to intrinsically motivate learning agents.
A mathematical theory of communication
C. E. Shannon · 1948
Earlier work this paper cites.
Dynamic Programming
R. E. Bellman · 1957
Earlier work this paper cites.
Coding theorems for a discrete source with a fidelity criterion
C. E. Shannon · 1959
Earlier work this paper cites.
An algorithm for computing the capacity of arbitrary discrete memoryless channels
S. Arimoto · 1972
Earlier work this paper cites.
Computation of channel capacity and rate-distortion functions
R. Blahut · 1972
Earlier work this paper cites.
The bidirectional communication theory–a generalization of information theory
H. Marko · 1973
Earlier work this paper cites.
Information geometry and alternating minimization procedures
I. Csiszar and G. Tusnady · 1984
Earlier work this paper cites.
Reinforcement learning for robots using neural networks
L.-J. Lin · 1993
Earlier work this paper cites.
The Arimoto-Blahut algorithm for finding channel capacity
R. G. Gallager · 1994
Earlier work this paper cites.
Neuro-Dynamic Programming
D. P. Bertsekas and J. N. Tsitsiklis · 1996
Earlier work this paper cites.
Directed information for channels with feedback
G. Kramer · 1998
Earlier work this paper cites.
Reinforcement Learning: An Introduction
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour · 2000
Earlier work this paper cites.
Predictability, complexity, and learning
W. Bialek, I. Nemenman, and N. Tishby · 2001
Earlier work this paper cites.
Implications of rational inattention
C. A. Sims · 2003
Earlier work this paper cites.
Empowerment: A universal agent-centric measure of control
A. S. Klyubin, D. Polani, and C. L. Nehaniv · 2005
Earlier work this paper cites.
Conversion of mutual and directed information
J. L. Massey and P. C. Massey · 2005
Earlier work this paper cites.
Elements of Information Theory
T. M. Cover and J. A. Thomas · 2006
Earlier work this paper cites.
Evolving spatiotemporal coordination in a modular robotic system
M. Prokopenko, V. Gerasimov, and I. Tanev · 2006
Earlier work this paper cites.
Keep your options open: An information-based driving principle for sensorimotor systems
A. S. Klyubin, D. Polani, and C. L. Nehaniv · 2008
Earlier work this paper cites.
Double Q-learning
H. van Hasselt · 2010
Earlier work this paper cites.
Higher coordination with less control-A result of information maximization in the sensorimotor loop
K. Zahedi, N. Ay, and R. Der · 2010
Earlier work this paper cites.
Modeling purposeful adaptive behavior wih the principle of maximum causal entropy
B. D. Ziebart · 2010
Earlier work this paper cites.
Dynamic policy programming with function approximation
M. G. Azar, V. Gomez, and H. J. Kappen · 2011
Earlier work this paper cites.
Empowerment for continuous agent-environment systems
T. Jung, D. Polani, and P. Stone · 2011
Cited alongside, same era.
Information theory of decisions and actions
N. Tishby and D. Polani · 2011
Cited alongside, same era.
Off-policy actor-critic
T. Degris, M. White, and R. S. Sutton · 2012
Cited alongside, same era.
Trading value and information in MDPs
J. Rubin, O. Shamir, and N. Tishby · 2012
Cited alongside, same era.
An information-theoretic approach to curiosity-driven reinforcement learning
S. Still and D. Precup · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Cited alongside, same era.
Artificial Intelligence: A Modern Approach
S. J. Russell and P. Norvig · 2016
Later among the works it cites.
Information-theoretic neuro-correlates boost evolution of cognitive systems
J. Schossau, C. Adami, and A. Hintze · 2016
Later among the works it cites.
Deep reinforcement learning with double Q-learning
H. van Hasselt, A. Guez, and D. Silver · 2016
Later among the works it cites.
Reinforcement learning with deep energy-based policies
T. Haarnoja, H. Tang, P. Abbeel, and S. Levine · 2017
Later among the works it cites.
Bridging the gap between value and policy based reinforcement learning
O. Nachum, M. Norouzi, K. Xu, and D. Schuurmans · 2017
Later among the works it cites.
A unified view of entropy-regularized Markov decision processes
G. Neu, V. Gomez, and A. Jonsson · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. A. Ortega and D. A. Braun · 2013
Cited alongside, same era.
Auto-encoding variational Bayes
D. P. Kingma and M. Welling · 2014
Cited alongside, same era.
Stochastic backpropagation and approximate inference in deep generative models
D. J. Rezende, S. Mohamed, and D. Wierstra · 2014
Cited alongside, same era.
Empowerment–an introduction
C. Salge, C. Glackin, and D. Polani · 2014
Cited alongside, same era.
Bounded rationality, abstraction, and hierarchical decision-making: An information-theoretic optimality principle
T. Genewein, F. Leibfried, J. Grau-Moya, and D. A. Braun · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Cited alongside, same era.
Equivalence between policy gradients and soft Q-learning
J. Schulman, P. Abbeel, and X. Chen · 2017
Later among the works it cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Later among the works it cites.
Maximum a posteriori policy optimisation
A. Abdolmaleki, J. T. Springenberg, Y. Tassa, R. Munos, N. Heess, and M. Riedmiller · 2018
Later among the works it cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
K. Chua, R. Calandra, R. McAllister, and S. Levine · 2018
Later among the works it cites.
A unified strategy for implementing curiosity and empowerment driven reinforcement learning
I. M. de Abril and R. Kanai · 2018
Later among the works it cites.
Adressing function approximation error in actor-critic methods
S. Fujimoto, H. van Hoof, and D. Meger · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Later among the works it cites.
Empowerment-driven exploration using mutual information estimation
N. M. Kumar · 2018
Later among the works it cites.
An information-theoretic optimality principle for deep reinforcement learning
F. Leibfried, J. Grau-Moya, and H. Bou-Ammar · 2018
Later among the works it cites.
Reinforcement learning and control as probabilistic inference: Tutorial and review
S. Levine · 2018
Later among the works it cites.
RLlib: Abstractions for distributed reinforcement learning
E. Liang, R. Liaw, P. Moritz, R. Nishihara, R. Fox, K. Goldberg, J. E. Gonzalez, M. I. Jordan, and I. Stoica · 2018
Later among the works it cites.
Reproducible, reusable, and robust reinforcement learning
J. Pineau · 2018
Later among the works it cites.
A unified Bellman equation for causal information and value in Markov decision processes
S. Tiomkin and N. Tishby · 2018
Later among the works it cites.
Soft Q-learning with mutual-information regularization
J. Grau-Moya, F. Leibfried, and P. Vrancx · 2019
Closest in time.
Soft actor-critic algorithms and applications
T. Haarnoja, A. Zhou, K. Hartikainen, G. Tucker, S. Ha, J. Tan, V. Kumar, H. Zhu, A. Gupta, P. Abbeel, and S. Levine · 2019
Closest in time.
Mutual-information regularization in Markov decision processes and actor-critic learning
F. Leibfried and J. Grau-Moya · 2019
Closest in time.
Adversarial imitation via variational inverse reinforcement learning
A. H. Qureshi, B. Boots, and M. C. Yip · 2019
Closest in time.