Fetching the paper…
Reading the bibliography…
We propose to learn to distinguish reversible from irreversible actions for better informed decision-making in Reinforcement Learning (RL).
The contingency model for the selection of decision strategies: An empirical test of the effects of significance, accountability, and reversibility
D. W. McAllister, T. R. Mitchell, and L. R. Beach · 1979
Earlier work this paper cites.
Neuronlike adaptive elements that can solve difficult learning control problems
A. G. Barto, R. S. Sutton, and C. W. Anderson · 1983
Earlier work this paper cites.
Markov chains
J. R. Norris · 1998
Earlier work this paper cites.
Anticipated regret, expected feedback and behavioral decision making
M. Zeelenberg · 1999
Earlier work this paper cites.
Don’t do things you can’t undo: Reversibility models for generating safe behaviours
M. Kruusmaa, Y. Gavshin, and A. Eppendahl · 2007
Earlier work this paper cites.
Safe exploration for reinforcement learning
A. Hans, D. Schneegaß, A. M. Schäfer, and S. Udluft · 2008
Earlier work this paper cites.
Photo sequencing
T. Basha, Y. Moses, and S. Avidan · 2012
Earlier work this paper cites.
Safe exploration of state and action spaces in reinforcement learning
J. García and F. Fernández · 2012
Earlier work this paper cites.
Safe exploration in markov decision processes
T. M. Moldovan and P. Abbeel · 2012
Earlier work this paper cites.
Seeing the arrow of time
L. C. Pickup, Z. Pan, D. Wei, Y. C. Shih, C. Zhang, A. Zisserman, B. Scholkopf, and W. T. Freeman · 2014
Earlier work this paper cites.
Modeling video evolution for action recognition
B. Fernando, E. Gavves, J. M. Oramas, A. Ghodrati, and T. Tuytelaars · 2015
Earlier work this paper cites.
Unsupervised learning of spatiotemporally coherent metrics
R. Goroshin, J. Bruna, J. Tompson, D. Eigen, and Y. LeCun · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Earlier work this paper cites.
Learning temporal embeddings for complex video analysis
V. Ramanathan, K. Tang, G. Mori, and L. Fei-Fei · 2015
Earlier work this paper cites.
Concrete problems in AI safety
D. Amodei, C. Olah, J. Steinhardt, P. F. Christiano, J. Schulman, and D. Mané · 2016
Earlier work this paper cites.
Safe learning of regions of attraction for uncertain, nonlinear systems with gaussian processes
F. Berkenkamp, R. Moriconi, A. Schoellig, and A. Krause · 2016
Cited alongside, same era.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Cited alongside, same era.
Shuffle and learn: unsupervised learning using temporal order verification
I. Misra, C. L. Zitnick, and M. Hebert · 2016
Cited alongside, same era.
Self-supervised video representation learning with odd-one-out networks
B. Fernando, H. Bilen, E. Gavves, and S. Gould · 2017
Cited alongside, same era.
J. Leike, M. Martic, V. Krakovna, P. A. Ortega, T. Everitt, A. Lefrancq, L. Orseau, and S. Legg · 2017
Cited alongside, same era.
An investigation of model-free planning
A. Guez, M. Mirza, K. Gregor, et al · 2019
Later among the works it cites.
Active Rollouts in MDP with Irreversible Dynamics
O.-A. Maillard, T. Mann, R. Ortner, and S. Mannor · 2019
Later among the works it cites.
Stable baselines3
A. Raffin, A. Hill, M. Ernestus, A. Gleave, A. Kanervisto, and N. Dormann · 2019
Later among the works it cites.
Episodic curiosity through reachability
N. Savinov, A. Raichuk, R. Marinier, D. Vincent, M. Pollefeys, T. Lillicrap, and S. Gelly · 2019
Later among the works it cites.
Self-supervised spatiotemporal learning via video clip order prediction
D. Xu, J. Xiao, Z. Zhao, J. Shao, D. Xie, and Y. Zhuang · 2019
Later among the works it cites.
Can temporal information help with contrastive self-supervised learning?
Y. Bai, H. Fan, I. Misra, G. Venkatesh, Y. Lu, Y. Zhou, Q. Yu, V. Chandra, and A. Yuille · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Cited alongside, same era.
Imagination-augmented agents for deep reinforcement learning
T. Weber, S. Racanière, D. P. Reichert, L. Buesing, A. Guez, D. J. Rezende, A. P. Badia, O. Vinyals, N. Heess, Y. Li, et al · 2017
Cited alongside, same era.
Motion perception in reinforcement learning with dynamic objects
A. Amiranashvili, A. Dosovitskiy, V. Koltun, and T. Brox · 2018
Cited alongside, same era.
Playing hard exploration games by watching youtube
Y. Aytar, T. Pfaff, D. Budden, T. L. Paine, Z. Wang, and N. de Freitas · 2018
Cited alongside, same era.
Learning actionable representations from visual observations
D. Dwibedi, J. Tompson, C. Lynch, and P. Sermanet · 2018
Cited alongside, same era.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
L. Espeholt, H. Soyer, R. Munos, K. Simonyan, V. Mnih, T. Ward, Y. Doron, V. Firoiu, T. Harley, I. Dunning, et al · 2018
Cited alongside, same era.
Neural predictive belief representations
Z. D. Guo, M. G. Azar, B. Piot, B. A. Pires, and R. Munos · 2018
Cited alongside, same era.
Later among the works it cites.
Primal wasserstein imitation learning
R. Dadashi, L. Hussenot, M. Geist, and O. Pietquin · 2020
Later among the works it cites.
Bootstrap latent-predictive representations for multitask reinforcement learning
Z. D. Guo, B. A. Pires, B. Piot, J.-B. Grill, F. Altché, R. Munos, and M. G. Azar · 2020
Later among the works it cites.
Acme: A research framework for distributed reinforcement learning
M. Hoffman, B. Shahriari, J. Aslanides, G. Barth-Maron, et al · 2020
Later among the works it cites.
Time reversal as self-supervision
S. Nair, M. Babaeizadeh, C. Finn, S. Levine, and V. Kumar · 2020
Later among the works it cites.
Learning the arrow of time for problems in reinforcement learning
N. Rahaman, S. Wolf, A. Goyal, R. Remme, and Y. Bengio · 2020
Later among the works it cites.
Learning from heterogeneous eeg signals with differentiable channel reordering
A. Saeed, D. Grangier, O. Pietquin, and N. Zeghidour · 2020
Later among the works it cites.
Curl: Contrastive unsupervised representations for reinforcement learning
A. Srinivas, M. Laskin, and P. Abbeel · 2020
Later among the works it cites.
Munchausen reinforcement learning
N. Vieillard, O. Pietquin, and M. Geist · 2020
Later among the works it cites.
Self-supervised learning of audio representations from permutations with differentiable ranking
A. N. Carr, Q. Berthet, M. Blondel, O. Teboul, and N. Zeghidour · 2021
Closest in time.
Reinforcement learning with prototypical representations
D. Yarats, R. Fergus, A. Lazaric, and L. Pinto · 2021
Closest in time.