Fetching the paper…
Reading the bibliography…
I review unsupervised or self-supervised neural networks playing minimax games in game-theoretic settings: (i) Artificial Curiosity (AC, 1990) is based on two such networks.
Theory of games and economic behavior
O. Morgenstern and J. Von Neumann · 1944
Earlier work this paper cites.
Chess programs, in The Plankalkuel
K. Zuse · 1945
Earlier work this paper cites.
A mathematical theory of communication (parts I and II)
C. E. Shannon · 1948
Earlier work this paper cites.
On information and sufficiency
S. Kullback and R. A. Leibler · 1951
Earlier work this paper cites.
Some aspects of the sequential design of experiments
H. Robbins · 1952
Earlier work this paper cites.
Some studies in machine learning using the game of checkers
A. L. Samuel · 1959
Earlier work this paper cites.
Cybernetics or Control and Communication in the Animal and the Machine
N. Wiener · 1965
Earlier work this paper cites.
The representation of the cumulative rounding error of an algorithm as a Taylor expansion of the local rounding errors
S. Linnainmaa · 1970
Earlier work this paper cites.
Rate-distortion theory and application
L. D. Davisson · 1972
Earlier work this paper cites.
Theory of optimal experiments
V. V. Fedorov · 1972
Earlier work this paper cites.
Taylor expansion of the accumulated rounding error
S. Linnainmaa · 1976
Earlier work this paper cites.
Applications of advances in nonlinear sensitivity analysis
P. J. Werbos · 1982
Earlier work this paper cites.
A dual back-propagation scheme for scalar reinforcement learning
P. W. Munro · 1987
Earlier work this paper cites.
Learning how the world works: Specifications for predictive networks in robots and brains
P. J. Werbos · 1987
Earlier work this paper cites.
Parasites and sex
J. Seger and W. Hamilton · 1988
Earlier work this paper cites.
On the use of backpropagation in associative reinforcement learning
R. J. Williams · 1988
Earlier work this paper cites.
Unsupervised learning
H. B. Barlow · 1989
Earlier work this paper cites.
Finding minimum entropy codes
H. B. Barlow, T. P. Kaushal, and G. J. Mitchison · 1989
Earlier work this paper cites.
Multi-armed Bandit Allocation Indices
J. C. Gittins · 1989
Earlier work this paper cites.
The truck backer-upper: An example of self learning in neural networks
N. Nguyen and B. Widrow · 1989
Earlier work this paper cites.
Neural networks for control and system identification
P. J. Werbos · 1989
Earlier work this paper cites.
Co-evolving parasites improve simulated evolution as an optimization procedure
W. D. Hillis · 1990
Earlier work this paper cites.
Supervised learning with a distal teacher
M. I. Jordan and D. E. Rumelhart · 1990
Earlier work this paper cites.
Making the world differentiable: On using fully recurrent self-supervised neural networks for dynamic reinforcement learning and planning in non-stationary environments
J. Schmidhuber · 1990
Earlier work this paper cites.
An on-line algorithm for dynamic reinforcement learning and planning in reactive environments
J. Schmidhuber · 1990
Earlier work this paper cites.
Learning to generate artificial fovea trajectories for target detection
J. Schmidhuber and R. Huber · 1990
Earlier work this paper cites.
Curious model-building control systems
J. Schmidhuber · 1991
Earlier work this paper cites.
Learning factorial codes by predictability minimization
J. Schmidhuber · 1991
Earlier work this paper cites.
A possibility for implementing curiosity and boredom in model-building neural controllers
J. Schmidhuber · 1991
Earlier work this paper cites.
Reinforcement learning in Markovian and non-Markovian environments
J. Schmidhuber · 1991
Earlier work this paper cites.
Personal Communication
P. Dayan, R. Zemel, and A. Pouget, 1992 · 1992
Earlier work this paper cites.
Learning factorial codes by predictability minimization
J. Schmidhuber · 1992
Earlier work this paper cites.
The application of the genetic algorithm to the minimization of potential energy functions
S. M. Le Grand and K. M. Merz · 1993
Earlier work this paper cites.
Netzwerkarchitekturen, Zielfunktionen und Kettenregel. (Network architectures, objective functions, and chain rule.)
J. Schmidhuber · 1993
Earlier work this paper cites.
On learning how to learn learning strategies
J. Schmidhuber · 1994
Earlier work this paper cites.
Gambling in a rigged casino: The adversarial multi-armed bandit problem
P. Auer, N. Cesa-Bianchi, Y. Freund, and R. E. Schapire · 1995
Cited alongside, same era.
Reinforcement driven information acquisition in non-deterministic environments
J. Storck, S. Hochreiter, and J. Schmidhuber · 1995
Cited alongside, same era.
Reinforcement learning: a survey
L. P. Kaelbling, M. L. Littman, and A. W. Moore · 1996
Cited alongside, same era.
Semilinear predictability minimization produces well-known feature detectors
J. Schmidhuber, M. Eldracher, and B. Foltin · 1996
Cited alongside, same era.
Stochastic approximation with two time scales
V. S. Borkar · 1997
Cited alongside, same era.
Reinforcement learning with self-modifying policies
J. Schmidhuber, J. Zhao, and N. Schraudolph · 1997
Cited alongside, same era.
Conditional generative adversarial nets
M. Mirza and S. Osindero · 2014
Later among the works it cites.
Stochastic backpropagation and approximate inference in deep generative models
D. J. Rezende, S. Mohamed, and D. Wierstra · 2014
Later among the works it cites.
Deep learning in neural networks: An overview
J. Schmidhuber · 2014
Later among the works it cites.
Deep generative image models using a Laplacian pyramid of adversarial networks
E. L. Denton, S. Chintala, R. Fergus, et al · 2015
Later among the works it cites.
How (not) to train your generative model: Scheduled sampling, likelihood, adversary?
F. Huszár · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shifting inductive bias with success-story algorithm, adaptive Levin search, and incremental self-improvement
J. Schmidhuber, J. Zhao, and M. Wiering · 1997
Cited alongside, same era.
What’s interesting?
J. Schmidhuber · 1998
Cited alongside, same era.
Reinforcement learning: An introduction
R. Sutton and A. Barto · 1998
Cited alongside, same era.
Artificial curiosity based on discovering novel algorithmic predictability through coevolution
J. Schmidhuber · 1999
Cited alongside, same era.
Neural predictors for detecting and removing redundant information
J. Schmidhuber · 1999
Cited alongside, same era.
Processing images by semi-linear predictability minimization
N. N. Schraudolph, M. Eldracher, and J. Schmidhuber · 1999
Cited alongside, same era.
Later among the works it cites.
A. Makhzani, J. Shlens, N. Jaitly, I. Goodfellow, and B. Frey · 2015
Later among the works it cites.
Unsupervised representation learning with deep convolutional generative adversarial networks
A. Radford, L. Metz, and S. Chintala · 2015
Later among the works it cites.
J. Schmidhuber · 2015
Later among the works it cites.
Domain separation networks
K. Bousmalis, G. Trigeorgis, N. Silberman, D. Krishnan, and D. Erhan · 2016
Later among the works it cites.
InfoGAN: Interpretable Representation Learning by Information Maximizing Generative Adversarial Nets
X. Chen, Y. Duan, R. Houthooft, J. Schulman, I. Sutskever, and P. Abbeel · 2016
Later among the works it cites.
J. Donahue, P. Krähenbühl, and T. Darrell · 2016
Later among the works it cites.
Adversarially learned inference
V. Dumoulin, I. Belghazi, B. Poole, A. Lamb, M. Arjovsky, O. Mastropietro, and A. Courville · 2016
Later among the works it cites.
Domain-adversarial training of neural networks
Y. Ganin, E. Ustinova, H. Ajakan, P. Germain, H. Larochelle, F. Laviolette, M. Marchand, and V. Lempitsky · 2016
Later among the works it cites.
f-GAN: training generative neural samplers using variational divergence minimization
S. Nowozin, B. Cseke, and R. Tomioka · 2016
Later among the works it cites.
Connecting generative adversarial networks and actor-critic methods
D. Pfau and O. Vinyals · 2016
Later among the works it cites.
M. Arjovsky, S. Chintala, and L. Bottou · 2017
Later among the works it cites.
A neural representation of sketch drawings
D. Ha and D. Eck · 2017
Later among the works it cites.
GANs trained by a two time-scale update rule converge to a local Nash equilibrium
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter · 2017
Later among the works it cites.
GANs trained by a two time-scale update rule converge to a Nash equilibrium
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, G. Klambauer, and S. Hochreiter · 2017
Later among the works it cites.
Two time-scale stochastic approximation with controlled Markov noise and off-policy temporal-difference learning
P. Karmakar and S. Bhatnagar · 2017
Later among the works it cites.
Least squares generative adversarial networks
X. Mao, Q. Li, H. Xie, R. Y. Lau, Z. Wang, and S. Paul Smolley · 2017
Later among the works it cites.
Curiosity-driven exploration by self-supervised prediction
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Later among the works it cites.
Coulomb GANs: provably optimal Nash equilibria via potential fields
T. Unterthiner, B. Nessler, G. Klambauer, M. Heusel, H. Ramsauer, and S. Hochreiter · 2017
Later among the works it cites.
Large-scale study of curiosity-driven learning
Y. Burda, H. Edwards, D. Pathak, A. Storkey, T. Darrell, and A. A. Efros · 2018
Later among the works it cites.
Synthesizing programs for images using reinforced adversarial learning
Y. Ganin, T. Kulkarni, I. Babuschkin, S. Eslami, and O. Vinyals · 2018
Later among the works it cites.
D. Ha and J. Schmidhuber · 2018
Later among the works it cites.
J. Schmidhuber · 2018
Later among the works it cites.
Learning to paint with model-based deep reinforcement learning
Z. Huang, W. Heng, and S. Zhou · 2019
Closest in time.
A style-based generator architecture for generative adversarial networks
T. Karras, S. Laine, and T. Aila · 2019
Closest in time.
TLDR: Schmidhuber’s Lab did it first
S. Le Grand · 2019
Closest in time.
On finding local Nash equilibria (and only local Nash equilibria) in zero-sum games
E. V. Mazumdar, M. I. Jordan, and S. S. Sastry · 2019
Closest in time.
Neural painters: A learned differentiable constraint for generating brushstroke paintings
R. Nakano · 2019
Closest in time.
J. Schmidhuber really had GANs in 1990. Online discussion, 2019
Reddit/ML · 2019
Closest in time.
Strokenet: A neural painting environment
N. Zheng, Y. Jiang, and D. Huang · 2019
Closest in time.