Efficient Exploration via State Marginal Matching
Original
L Lee, B Eysenbach, E Parisotto, E Xing, S Levine, R Salakhutdinov · 1906
Earlier work this paper cites.
Stochastic Latent Actor-Critic: Deep Reinforcement Learning With a Latent Variable Model
Original
AX Lee, A Nagabandi, P Abbeel, S Levine · 1907
Earlier work this paper cites.
Dream to Control: Learning Behaviors by Latent Imagination
Original
D Hafner, T Lillicrap, J Ba, M Norouzi · 1912
Earlier work this paper cites.
Some Tests of Significance, Treated by the Theory of Probability
H Jeffreys · 1935
Earlier work this paper cites.
What Is Life? The Physical Aspect of the Living Cell and Mind
E Schrödinger · 1944
Earlier work this paper cites.
An Essentially Complete Class of Admissible Decision Functions
A Wald · 1947
Earlier work this paper cites.
A Mathematical Theory of Communication
CE Shannon · 1948
Earlier work this paper cites.
Cybernetics or Control and Communication in the Animal and the Machine
N Wiener · 1948
Earlier work this paper cites.
On Information and Sufficiency
S Kullback RA Leibler · 1951
Earlier work this paper cites.
A Method for the Construction of Minimum-Redundancy Codes
DA Huffman · 1952
Earlier work this paper cites.
Theory of Games and Economic Behavior
O Morgenstern J Von Neumann · 1953
Earlier work this paper cites.
On a Measure of the Information Provided by an Experiment
DV Lindley et al · 1956
Earlier work this paper cites.
Information Theory and Statistical Mechanics
ET Jaynes · 1957
Earlier work this paper cites.
The Principles of Quantum Mechanics
PAM Dirac · 1958
Earlier work this paper cites.
A New Approach to Linear Filtering and Prediction Problems
RE Kalman · 1960
Earlier work this paper cites.
Markov’s Conditional Processes
R Stratonovich · 1960
Earlier work this paper cites.
An Introduction to Cybernetics
WR Ashby · 1961
Earlier work this paper cites.
Risk Aversion in the Small and in the Large
JW Pratt · 1964
Earlier work this paper cites.
Complex Analysis, 1966
R Rudin · 1966
Earlier work this paper cites.
A Theory of Adaptive Pattern Classifiers
S Amari · 1967
Earlier work this paper cites.
Risk-Sensitive Markov Decision Processes
RA Howard JE Matheson · 1972
Earlier work this paper cites.
Perceptions as Hypotheses
RL Gregory · 1980
Earlier work this paper cites.
A Complete Class Theorem for Statistical Problems With Finite Sample Spaces
LD Brown · 1981
Earlier work this paper cites.
The Science of Structure: Synergetics
H Haken · 1981
Earlier work this paper cites.
Large Automatic Learning, Rule Extraction, and Generalization
J Denker, D Schwartz, B Wittner, S Solla, R Howard, L Jackel, J Hopfield · 1987
Earlier work this paper cites.
A Mean Field Theory Learning Algorithm for Neural Networks
C Peterson · 1987
Earlier work this paper cites.
Stochastic Multisensory Data Fusion for Mobile Robot Location and Environment Modelling. 5th Int
P Moutarlier R Chatila · 1989
Earlier work this paper cites.
Curious Model-Building Control Systems
J Schmidhuber · 1991
Earlier work this paper cites.
Dyna, an Integrated Architecture for Learning, Planning, and Reacting
RS Sutton · 1991
Earlier work this paper cites.
Function Optimization Using Connectionist Reinforcement Learning Algorithms
RJ Williams J Peng · 1991
Earlier work this paper cites.
Ockham’s Razor and Bayesian Analysis
WH Jefferys JO Berger · 1992
Earlier work this paper cites.
Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning
RJ Williams · 1992
Earlier work this paper cites.
Keeping the Neural Networks Simple by Minimizing the Description Length of the Weights
GE Hinton D Van Camp · 1993
Earlier work this paper cites.
The Helmholtz Machine
P Dayan, GE Hinton, RM Neal, RS Zemel · 1995
Earlier work this paper cites.
Learning From Incomplete Data, 1995
Z Ghahramani MI Jordan · 1995
Earlier work this paper cites.
Bayes Factors
RE Kass AE Raftery · 1995
Earlier work this paper cites.
Causal Diagrams for Empirical Research
J Pearl · 1995
Earlier work this paper cites.
An Introduction to Variational Methods for Graphical Models
MI Jordan, Z Ghahramani, TS Jaakkola, LK Saul · 1999
Earlier work this paper cites.
Between Mdps and Semi-Mdps: A Framework for Temporal Abstraction in Reinforcement Learning
RS Sutton, D Precup, S Singh · 1999
Earlier work this paper cites.
An Analysis of Stochastic Game Theory for Multiagent Reinforcement Learning
M Bowling M Veloso · 2000
Earlier work this paper cites.
Policy Gradient Methods for Reinforcement Learning With Function Approximation
RS Sutton, DA McAllester, SP Singh, Y Mansour · 2000
Earlier work this paper cites.
The IM Algorithm: a Variational Approach to Information Maximization
D Barber FV Agakov · 2003
Earlier work this paper cites.
Information Projections Revisited
I Csiszár F Matus · 2003
Earlier work this paper cites.
Information Theory, Inference and Learning Algorithms
DJ MacKay · 2003
Earlier work this paper cites.
Empowerment: A Universal Agent-Centric Measure of Control
AS Klyubin, D Polani, CL Nehaniv · 2005
Earlier work this paper cites.
Pattern Recognition and Machine Learning
CM Bishop · 2006
Earlier work this paper cites.
A Fast Learning Algorithm for Deep Belief Nets
GE Hinton, S Osindero, YW Teh · 2006
Earlier work this paper cites.
A Tutorial on Energy-Based Learning
Y LeCun, S Chopra, R Hadsell, M Ranzato, F Huang · 2006
Earlier work this paper cites.
Intrinsic Motivation Systems for Autonomous Mental Development
PY Oudeyer, F Kaplan, VV Hafner · 2007
Earlier work this paper cites.
General Duality Between Optimal Control and Estimation
E Todorov · 2008
Earlier work this paper cites.