Fetching the paper…
Reading the bibliography…
We transform reinforcement learning (RL) into a form of supervised learning (SL) by turning traditional RL on its head, calling this Upside Down RL (UDRL).
Über formal unentscheidbare Sätze der Principia Mathematica und verwandter Systeme I
K. Gödel · 1931
Earlier work this paper cites.
An unsolvable problem of elementary number theory
A. Church · 1936
Earlier work this paper cites.
Finite combinatory processes-formulation 1
E. L. Post · 1936
Earlier work this paper cites.
On computable numbers, with an application to the Entscheidungsproblem
A. M. Turing · 1936
Earlier work this paper cites.
Dynamic Programming
R. Bellman · 1957
Earlier work this paper cites.
Gradient theory of optimal flight paths
H. J. Kelley · 1960
Earlier work this paper cites.
Conditional Markov processes
R. Stratonovich · 1960
Earlier work this paper cites.
A formal theory of inductive inference. Part I
R. J. Solomonoff · 1964
Earlier work this paper cites.
Cybernetic Predicting Devices
A. G. Ivakhnenko and V. G. Lapa · 1965
Earlier work this paper cites.
Three approaches to the quantitative definition of information
A. N. Kolmogorov · 1965
Earlier work this paper cites.
Artificial Intelligence through Simulated Evolution
L. Fogel, A. Owens, and M. Walsh · 1966
Earlier work this paper cites.
The representation of the cumulative rounding error of an algorithm as a Taylor expansion of the local rounding errors
S. Linnainmaa · 1970
Earlier work this paper cites.
K. Zuse · 1970
Earlier work this paper cites.
Evolutionsstrategie - Optimierung technischer Systeme nach Prinzipien der biologischen Evolution. Dissertation, 1971
I. Rechenberg · 1973
Earlier work this paper cites.
Adaptation in Natural and Artificial Systems
J. H. Holland · 1975
Earlier work this paper cites.
Numerische Optimierung von Computer-Modellen. Dissertation, 1974
H. P. Schwefel · 1977
Earlier work this paper cites.
Applications of advances in nonlinear sensitivity analysis
P. J. Werbos · 1982
Earlier work this paper cites.
A dual back-propagation scheme for scalar reinforcement learning
P. W. Munro · 1987
Earlier work this paper cites.
The utility driven dynamic error propagation network
A. J. Robinson and F. Fallside · 1987
Earlier work this paper cites.
Building and understanding adaptive systems: A statistical/numerical approach to factory automation and brain research
P. J. Werbos · 1987
Earlier work this paper cites.
Supervised learning and systems with excess degrees of freedom
M. I. Jordan · 1988
Earlier work this paper cites.
Generalization of backpropagation with application to a recurrent gas market model
P. J. Werbos · 1988
Earlier work this paper cites.
Designing neural networks using genetic algorithms
G. Miller, P. Todd, and S. Hedge · 1989
Earlier work this paper cites.
The truck backer-upper: An example of self learning in neural networks
N. Nguyen and B. Widrow · 1989
Earlier work this paper cites.
Dynamic reinforcement driven error propagation networks with application to game playing
T. Robinson and F. Fallside · 1989
Earlier work this paper cites.
Backpropagation and neurocontrol: A review and prospectus
P. J. Werbos · 1989
Earlier work this paper cites.
Neural networks for control and system identification
P. J. Werbos · 1989
Earlier work this paper cites.
Making the world differentiable: On using fully recurrent self-supervised neural networks for dynamic reinforcement learning and planning in non-stationary environments
J. Schmidhuber · 1990
Cited alongside, same era.
An on-line algorithm for dynamic reinforcement learning and planning in reactive environments
J. Schmidhuber · 1990
Cited alongside, same era.
Learning to generate artificial fovea trajectories for target detection
J. Schmidhuber and R. Huber · 1990
Cited alongside, same era.
Programming robots using reinforcement learning and teaching
L.-J. Lin · 1991
Cited alongside, same era.
Learning to generate sub-goals for action sequences
J. Schmidhuber · 1991
Cited alongside, same era.
Reinforcement learning in Markovian and non-Markovian environments
Pattern Recognition and Machine Learning
C. M. Bishop · 2006
Later among the works it cites.
Accelerated neural evolution through cooperatively coevolved synapses
F. J. Gomez, J. Schmidhuber, and R. Miikkulainen · 2008
Later among the works it cites.
A survey of robot learning from demonstration
B. D. Argall, S. Chernova, M. Veloso, and B. Browning · 2009
Later among the works it cites.
A novel connectionist system for improved unconstrained handwriting recognition
A. Graves, M. Liwicki, S. Fernandez, R. Bertolami, H. Bunke, and J. Schmidhuber · 2009
Later among the works it cites.
The elements of statistical learning
T. Hastie, R. Tibshirani, and J. Friedman · 2009
Later among the works it cites.
Exponential natural evolution strategies
T. Glasmachers, T. Schaul, Y. Sun, D. Wierstra, and J. Schmidhuber · 2010
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Schmidhuber · 1991
Cited alongside, same era.
Learning complex, extended sequences using the principle of history compression
J. Schmidhuber · 1991
Cited alongside, same era.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Cited alongside, same era.
Continual Learning in Reinforcement Environments
M. B. Ring · 1994
Cited alongside, same era.
Evolving virtual creatures
K. Sims · 1994
Cited alongside, same era.
Gradient-based learning algorithms for recurrent networks and their computational complexity
R. J. Williams and D. Zipser · 1994
Cited alongside, same era.
Long Short-Term Memory
S. Hochreiter and J. Schmidhuber · 1995
Cited alongside, same era.
Neural Networks: Tricks of the Trade
G. Montavon, G. Orr, and K. Müller · 2012
Later among the works it cites.
Reinforcement Learning
M. Wiering and M. van Otterlo · 2012
Later among the works it cites.
First experiments with PowerPlay
R. K. Srivastava, B. R. Steunebrink, and J. Schmidhuber · 2013
Later among the works it cites.
Deep learning in neural networks: An overview
J. Schmidhuber · 2014
Later among the works it cites.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Later among the works it cites.
Google voice search: faster and more accurate
H. Sak, A. Senior, K. Rao, F. Beaufays, and J. Schalkwyk · 2015
Later among the works it cites.
J. Schmidhuber · 2015
Later among the works it cites.
Bringing the Magic of Amazon AI and Alexa to Apps on AWS
W. Vogels · 2016
Later among the works it cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Y. Wu, M. Schuster, Z. Chen, Q. V. Le, M. Norouzi, W. Macherey, M. Krikun, Y. Cao, Q. Gao, K. Macherey, J. Klingner, A. Shah, M. Johnson, X. Liu, L. Kaiser, S. Gouws, Y. Kato, T. Kudo, H. Kazawa, K. Stevens, G. Kurian, N. Patil, W. Wang, C. Young, J. Smith, J. Riesa, A. Rudnick, O. Vinyals, G. Corrado, M. Hughes, and J. Dean · 2016
Later among the works it cites.
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, P. Abbeel, and W. Zaremba · 2017
Later among the works it cites.
Learning to act by predicting the future
A. Dosovitskiy and V. Koltun · 2017
Later among the works it cites.
One-shot imitation learning
Y. Duan, M. Andrychowicz, B. Stadie, O. J. Ho, J. Schneider, I. Sutskever, P. Abbeel, and W. Zaremba · 2017
Later among the works it cites.
Transitioning entirely to neural machine translation
J. Pino, A. Sidorov, and N. Ayan · 2017
Later among the works it cites.
P. Rauber, F. Mutz, and J. Schmidhuber · 2017
Later among the works it cites.
Time-contrastive networks: Self-supervised learning from multi-view observation
P. Sermanet, C. Lynch, J. Hsu, and S. Levine · 2017
Later among the works it cites.
Progressive reinforcement learning with distillation for multi-skilled motion control
G. Berseth, C. Xie, P. Cernek, and M. V. de Panne · 2018
Later among the works it cites.
J. Schmidhuber · 2018
Later among the works it cites.
Rudder: Return decomposition for delayed rewards
J. A. Arjona-Medina, M. Gillhofer, M. Widrich, T. Unterthiner, J. Brandstetter, and S. Hochreiter · 2019
Closest in time.
Reinforcement learning upside down: Don’t predict rewards–just map them to actions
J. Schmidhuber · 2019
Closest in time.
Training agents using upside-down reinforcement learning
R. K. Srivastava, P. Shyam, F. Mutz, W. Jaskowski, and J. Schmidhuber · 2019
Closest in time.