Fetching the paper…
Reading the bibliography…
We humans seem to have an innate understanding of the asymmetric progression of time, which we use to efficiently and safely perceive and manipulate our environment.
The nature of the physical world / by A.S. Eddington
Arthur Stanley Eddington · 1929
Earlier work this paper cites.
Smoothing and differentiation of data by simplified least squares procedures
Abraham Savitzky and Marcel JE Golay · 1964
Earlier work this paper cites.
Dissipative dynamical systems part i: General theory
Jan C Willems · 1972
Earlier work this paper cites.
Time, structure, and fluctuations
Ilya Prigogine · 1978
Earlier work this paper cites.
Sample estimate of the entropy of a random vector
LF Kozachenko and Nikolai N Leonenko · 1987
Earlier work this paper cites.
Algorithmic randomness and physical entropy
Wojciech H Zurek · 1989
Earlier work this paper cites.
The variational formulation of the fokker–planck equation
Richard Jordan, David Kinderlehrer, and Felix Otto · 1998
Earlier work this paper cites.
Decoherence, chaos, quantum-classical correspondence, and the algorithmic arrow of time
Wojciech H Zurek · 1998
Earlier work this paper cites.
Entropy production fluctuation theorem and the nonequilibrium work relation for free energy differences
Gavin E. Crooks · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y Ng and Stuart Russell · 2000
Earlier work this paper cites.
Estimating mutual information
Alexander Kraskov, Harald Stögbauer, and Peter Grassberger · 2004
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
Nuttapong Chentanez, Andrew G Barto, and Satinder P Singh · 2005
Earlier work this paper cites.
Dynamical billiards
L. Bunimovich · 2007
Earlier work this paper cites.
On the entropy production of time series with unidirectional linearity
Dominik Janzing · 2010
Earlier work this paper cites.
Formal theory of creativity, fun, and intrinsic motivation (1990–2010)
Jürgen Schmidhuber · 2010
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2011
Cited alongside, same era.
Stochastic thermodynamics, fluctuation theorems and molecular machines
Udo Seifert · 2012
Cited alongside, same era.
Safe exploration in markov decision processes
Teodor Mihai Moldovan and Pieter Abbeel · 2012
Cited alongside, same era.
Stability of nonlinear stochastic discrete-time systems
Yan Li, Weihai Zhang, and Xikui Liu · 2013
Cited alongside, same era.
Seeing the arrow of time
Lyndsey C Pickup, Zheng Pan, Donglai Wei, YiChang Shih, Changshui Zhang, Andrew Zisserman, Bernhard Scholkopf, and William T Freeman · 2014
Cited alongside, same era.
A damped oscillator as a hamiltonian system
Kirk T McDonald · 2015
Cited alongside, same era.
Leave no trace: Learning to reset for safe and autonomous reinforcement learning
Benjamin Eysenbach, Shixiang Gu, Julian Ibarz, and Sergey Levine · 2017
Later among the works it cites.
Jan Leike, Miljan Martic, Victoria Krakovna, Pedro A Ortega, Tom Everitt, Andrew Lefrancq, Laurent Orseau, and Shane Legg · 2017
Later among the works it cites.
Low impact artificial intelligences
Stuart Armstrong and Benjamin Levinstein · 2017
Later among the works it cites.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A. Efros, and Trevor Darrell · 2017
Later among the works it cites.
Elements of causal inference: foundations and learning algorithms
Jonas Peters, Dominik Janzing, and Bernhard Schölkopf · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Efficient estimation of mutual information for strongly dependent variables
Shuyang Gao, Greg Ver Steeg, and Aram Galstyan · 2015
Cited alongside, same era.
Dueling network architectures for deep reinforcement learning
Ziyu Wang, Tom Schaul, Matteo Hessel, Hado Van Hasselt, Marc Lanctot, and Nando De Freitas · 2015
Cited alongside, same era.
Algorithmic independence of initial condition and dynamical law in thermodynamics and causal inference
Dominik Janzing, Rafael Chaves, and Bernhard Schölkopf · 2016
Cited alongside, same era.
The arrow of time in multivariate time series
Stefan Bauer, Bernhard Schölkopf, and Jonas Peters · 2016
Cited alongside, same era.
Concrete problems in ai safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Cited alongside, same era.
Safe and efficient off-policy reinforcement learning
Rémi Munos, Tom Stepleton, Anna Harutyunyan, and Marc Bellemare · 2016
Cited alongside, same era.
Later among the works it cites.
Recall traces: Backtracking models for efficient reinforcement learning
Anirudh Goyal, Philemon Brakel, William Fedus, Soumye Singhal, Timothy Lillicrap, Sergey Levine, Hugo Larochelle, and Yoshua Bengio · 2018
Later among the works it cites.
Time reversal as self-supervision
Suraj Nair, Mohammad Babaeizadeh, Chelsea Finn, Sergey Levine, and Vikash Kumar · 2018
Later among the works it cites.
Measuring and avoiding side effects using relative reachability
Victoria Krakovna, Laurent Orseau, Miljan Martic, and Shane Legg · 2018
Later among the works it cites.
Learning and using the arrow of time
Donglai Wei, Joseph J Lim, Andrew Zisserman, and William T Freeman · 2018
Later among the works it cites.
David Ha and Jürgen Schmidhuber · 2018
Later among the works it cites.
Episodic curiosity through reachability
Nikolay Savinov, Anton Raichuk, Raphaël Marinier, Damien Vincent, Marc Pollefeys, Timothy Lillicrap, and Sylvain Gelly · 2018
Later among the works it cites.
Large-scale study of curiosity-driven learning
Yuri Burda, Harri Edwards, Deepak Pathak, Amos Storkey, Trevor Darrell, and Alexei A Efros · 2018
Later among the works it cites.
gym-sokoban
Max-Philipp B. Schrader · 2018
Later among the works it cites.
Modularized implementation of deep rl algorithms in pytorch
Zhang Shangtong · 2018
Later among the works it cites.