Fetching the paper…
Reading the bibliography…
This paper is a review of the evolutionary history of deep learning models.
Elementary principles in statistical mechanics
Gibbs J Willard · 1902
Earlier work this paper cites.
A logical calculus of the ideas immanent in nervous activity
Warren S McCulloch and Walter Pitts · 1943
Earlier work this paper cites.
The organization of behavior: A neuropsychological theory
Donald Olding Hebb · 1949
Earlier work this paper cites.
Equilibrium points in n-person games
John F Nash et al · 1950
Earlier work this paper cites.
Iterative solution of games by fictitious play
George W Brown · 1951
Earlier work this paper cites.
Non-cooperative games
John Nash · 1951
Earlier work this paper cites.
The perceptron: a probabilistic model for information storage and organization in the brain
Frank Rosenblatt · 1958
Earlier work this paper cites.
Receptive fields of single neurones in the cat’s striate cortex
David H Hubel and Torsten N Wiesel · 1959
Earlier work this paper cites.
Adaptive” adaline” Neuron Using Chemical” memistors.”
Bernard Widrow et al · 1960
Earlier work this paper cites.
Perceptrons: an introduction to computational geometry
Marvin L Minski and Seymour A Papert · 1969
Earlier work this paper cites.
Theory of spin glasses
Samuel Frederick Edwards and Phil W Anderson · 1975
Earlier work this paper cites.
A statistical-thermodynamic approach to determination of structure amplitude phases
AG Khachaturyan, SV Semenovskaya, and B Vainstein · 1979
Earlier work this paper cites.
Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position
Kunihiko Fukushima · 1980
Earlier work this paper cites.
Neural networks and physical systems with emergent collective computational abilities
John J Hopfield · 1982
Earlier work this paper cites.
Simplified neuron model as a principal component analyzer
Erkki Oja · 1982
Earlier work this paper cites.
A learning algorithm for boltzmann machines
David H Ackley, Geoffrey E Hinton, and Terrence J Sejnowski · 1985
Earlier work this paper cites.
Selective attention gates visual processing in the extrastriate cortex
Jeffrey Moran and Robert Desimone · 1985
Earlier work this paper cites.
On stochastic approximation of the eigenvectors and eigenvalues of the expectation of a random matrix
Erkki Oja and Juha Karhunen · 1985
Earlier work this paper cites.
Learning internal representations by error propagation
David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams · 1985
Earlier work this paper cites.
Separating the polynomial-time hierarchy by oracles
Andrew Chi-Chih Yao · 1985
Earlier work this paper cites.
Almost optimal lower bounds for small depth circuits
Johan Hastad · 1986
Earlier work this paper cites.
Serial order: A parallel distributed processing approach
Michael I Jordan · 1986
Earlier work this paper cites.
Information processing in dynamical systems: Foundations of harmony theory
Paul Smolensky · 1986
Earlier work this paper cites.
The utility driven dynamic error propagation network
AJ Robinson and Frank Fallside · 1987
Earlier work this paper cites.
Simulated annealing and boltzmann machines
Emile Aarts and Jan Korst · 1988
Earlier work this paper cites.
Continuous valued neural networks with two hidden layers are sufficient
G Cybenko · 1988
Earlier work this paper cites.
How neural nets work
Alan S Lapedes and Robert M Farber · 1988
Earlier work this paper cites.
Generalization of backpropagation with application to a recurrent gas market model
Paul J Werbos · 1988
Earlier work this paper cites.
Back propagation fails to separate where perceptrons succeed
Martin L Brady, Raghu Raghavan, and Joseph Slawny · 1989
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
George Cybenko · 1989
Earlier work this paper cites.
The cascade-correlation learning architecture
Scott E Fahlman and Christian Lebiere · 1989
Earlier work this paper cites.
Theory of the backpropagation neural network
Robert Hecht-Nielsen · 1989
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Kurt Hornik, Maxwell Stinchcombe, and Halbert White · 1989
Earlier work this paper cites.
Learning in feedforward layered networks: The tiling algorithm
Marc Mézard and Jean-P Nadal · 1989
Earlier work this paper cites.
A focused back-propagation algorithm for temporal pattern recognition
Michael C Mozer · 1989
Earlier work this paper cites.
Finding structure in time
Jeffrey L Elman · 1990
Earlier work this paper cites.
The upstart algorithm: A method for constructing and training feedforward neural networks
Marcus Frean · 1990
Earlier work this paper cites.
The self-organizing map
Teuvo Kohonen · 1990
Earlier work this paper cites.
Handwritten digit recognition with a back-propagation network
B Boser Le Cun, John S Denker, D Henderson, Richard E Howard, W Hubbard, and Lawrence D Jackel · 1990
Earlier work this paper cites.
Spin glass theory and beyond
Marc Mézard, Giorgio Parisi, and Miguel-Angel Virasoro · 1990
Earlier work this paper cites.
A survey of decision tree classifier methodology
S Rasoul Safavian and David Landgrebe · 1990
Earlier work this paper cites.
Backpropagation through time: what it does and how to do it
Paul J Werbos · 1990
Earlier work this paper cites.
Distributed optimization by ant colonies
Alberto Colorni, Marco Dorigo, Vittorio Maniezzo, et al · 1991
Earlier work this paper cites.
Deep Learning via Stacked Sparse Autoencoders for Automated Voxel-Wise Brain Parcellation Based on Functional Connectivity
Céline Gravelines · 1991
Earlier work this paper cites.
On the problem of local minima in backpropagation
Marco Gori and Alberto Tesi · 1992
Earlier work this paper cites.
A direct adaptive method for faster backpropagation learning: The rprop algorithm
Martin Riedmiller and Heinrich Braun · 1993
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Yoshua Bengio, Patrice Simard, and Paolo Frasconi · 1994
Earlier work this paper cites.
The” wake-sleep” algorithm for unsupervised neural networks
Geoffrey E Hinton, Peter Dayan, Brendan J Frey, and Radford M Neal · 1995
Earlier work this paper cites.
Foundations of vision
Brian A Wandell · 1995
Earlier work this paper cites.
Neural network fundamentals with graphs, algorithms, and applications
Nirmal K Bose et al · 1996
Earlier work this paper cites.
Tutorial in biostatistics: Using the general linear mixed model to analyse unbalanced repeated measures and longitudinal data
Avital Cnaan, NM Laird, and Peter Slasor · 1997
Earlier work this paper cites.
An introduction to neural networks
Kevin Gurney · 1997
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Machine learning. wcb, 1997
Tom M Mitchell et al · 1997
Earlier work this paper cites.
Constructive neural network learning algorithms for multi-category real-valued pattern classification
Rajesh G Parekh, Jihoon Yang, and Vasant Honavar · 1997
Earlier work this paper cites.
Bidirectional recurrent neural networks
Mike Schuster and Kuldip K Paliwal · 1997
Earlier work this paper cites.
Increasing the capacity of a hopfield network without sacrificing functionality
Amos Storkey · 1997
Earlier work this paper cites.
Bain on neural networks
Alan L Wilkes and Nicholas J Wade · 1997
Earlier work this paper cites.
A neural model of contour integration in the primary visual cortex
Zhaoping Li · 1998
Earlier work this paper cites.
An introduction to genetic algorithms
Melanie Mitchell · 1998
Earlier work this paper cites.
Maximizing generalized linear mixed model likelihoods with an automated monte carlo em algorithm
James G Booth and James P Hobert · 1999
Earlier work this paper cites.
Self organizing maps
Tom Germano · 1999
Earlier work this paper cites.
Error tolerant associative memory
Cheng-Yuan Liou and Shao-Kuo Yuan · 1999
Earlier work this paper cites.
Object recognition from local scale-invariant features
David G Lowe · 1999
Earlier work this paper cites.
Neural and adaptive systems: fundamentals through simulations with CD-ROM
Jose C Principe, Neil R Euliano, and W Curt Lefebvre · 1999
Cited alongside, same era.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, Yishay Mansour, et al · 1999
Cited alongside, same era.
Evolving artificial neural networks
Xin Yao · 1999
Cited alongside, same era.
Talking nets: An oral history of neural networks
James A Anderson and Edward Rosenfeld · 2000
Cited alongside, same era.
Psychology: the beginnings
CG Boeree · 2000
Cited alongside, same era.
Neural networks with a continuous squashing function in the output are universal approximators
Juan Luis Castro, Carlos Javier Mantas, and JM Benıtez · 2000
Cited alongside, same era.
Dropout training for support vector machines
Ning Chen, Jun Zhu, Jianfei Chen, and Bo Zhang · 2014
Later among the works it cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Later among the works it cites.
Neural networks and neuroscience-inspired computer vision
David Daniel Cox and Thomas Dean · 2014
Later among the works it cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Later among the works it cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik · 2014
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Recurrent nets that time and count
Felix A Gers and Jürgen Schmidhuber · 2000
Cited alongside, same era.
The distributed human neural system for face perception
James V Haxby, Elizabeth A Hoffman, and M Ida Gobbini · 2000
Cited alongside, same era.
A multithreaded software model for backpropagation neural network applications
Kiyoshi Kawaguchi · 2000
Cited alongside, same era.
John von neumann’s conception of the minimax theorem: a journey through different mathematical contexts
Tinne Hoff Kjeldsen · 2001
Cited alongside, same era.
A tutorial on bayesian belief networks
Mark L Krieg · 2001
Cited alongside, same era.
Generalized linear mixed models
Charles E McCulloch and John M Neuhaus · 2001
Cited alongside, same era.
Later among the works it cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Later among the works it cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Later among the works it cites.
Semi-supervised learning with deep generative models
Diederik P Kingma, Shakir Mohamed, Danilo Jimenez Rezende, and Max Welling · 2014
Later among the works it cites.
An introduction to brain and behavior , volume 1273
Bryan Kolb, Ian Q Whishaw, and G Campbell Teskey · 2014
Later among the works it cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Later among the works it cites.
On the saddle point problem for non-convex optimization
Razvan Pascanu, Yann N Dauphin, Surya Ganguli, and Yoshua Bengio · 2014
Later among the works it cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Later among the works it cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey E Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Later among the works it cites.
Importance weighted autoencoders
Yuri Burda, Roger Grosse, and Ruslan Salakhutdinov · 2015
Later among the works it cites.
Describing multimedia content using attention-based encoder-decoder networks
Kyunghyun Cho, Aaron Courville, and Yoshua Bengio · 2015
Later among the works it cites.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2015
Later among the works it cites.
Attention-based models for speech recognition
Jan K Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio · 2015
Later among the works it cites.
The power of depth for feedforward neural networks
Ronen Eldan and Ohad Shamir · 2015
Later among the works it cites.
A neural algorithm of artistic style
Leon A Gatys, Alexander S Ecker, and Matthias Bethge · 2015
Later among the works it cites.
Fast r-cnn
Ross Girshick · 2015
Later among the works it cites.
Klaus Greff, Rupesh Kumar Srivastava, Jan Koutník, Bas R Steunebrink, and Jürgen Schmidhuber · 2015
Later among the works it cites.
Learning structure in gene expression data using deep architectures, with an application to gene clustering
Aman Gupta, Haohan Wang, and Madhavi Ganapathiraju · 2015
Later among the works it cites.
Hypercolumns for object segmentation and fine-grained localization
Bharath Hariharan, Pablo Arbeláez, Ross Girshick, and Jitendra Malik · 2015
Later among the works it cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Later among the works it cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Later among the works it cites.
Deep convolutional inverse graphics network
Tejas D Kulkarni, William F Whitney, Pushmeet Kohli, and Josh Tenenbaum · 2015
Later among the works it cites.
Human-level concept learning through probabilistic program induction
Brenden M Lake, Ruslan Salakhutdinov, and Joshua B Tenenbaum · 2015
Later among the works it cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Later among the works it cites.
A critical review of recurrent neural networks for sequence learning
Zachary C Lipton, John Berkowitz, and Charles Elkan · 2015
Later among the works it cites.
Deep neural networks are easily fooled: High confidence predictions for unrecognizable images
Anh Nguyen, Jason Yosinski, and Jeff Clune · 2015
Later among the works it cites.
Complex recurrent neural networks for denoising speech signals
Keiichi Osako, Rita Singh, and Bhiksha Raj · 2015
Later among the works it cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik · 2015
Later among the works it cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Later among the works it cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Later among the works it cites.
Deep learning in neural networks: An overview
Jürgen Schmidhuber · 2015
Later among the works it cites.
Learning structured output representation using deep conditional generative models
Kihyuk Sohn, Honglak Lee, and Xinchen Yan · 2015
Later among the works it cites.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Later among the works it cites.
The human kernel
Andrew G Wilson, Christoph Dann, Chris Lucas, and Eric P Xing · 2015
Later among the works it cites.
Convolutional lstm network: A machine learning approach for precipitation nowcasting
SHI Xingjian, Zhourong Chen, Hao Wang, Dit-Yan Yeung, Wai-Kin Wong, and Wang-chun Woo · 2015
Later among the works it cites.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhutdinov, Richard S Zemel, and Yoshua Bengio · 2015
Later among the works it cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Later among the works it cites.
An analysis of deep neural network models for practical applications
Alfredo Canziani, Adam Paszke, and Eugenio Culurciello · 2016
Later among the works it cites.
Infogan: Interpretable representation learning by information maximizing generative adversarial nets
Xi Chen, Yan Duan, Rein Houthooft, John Schulman, Ilya Sutskever, and Pieter Abbeel · 2016
Later among the works it cites.
Dynamic filter networks
Bert De Brabandere, Xu Jia, Tinne Tuytelaars, and Luc Van Gool · 2016
Later among the works it cites.
Tutorial on variational autoencoders
Carl Doersch · 2016
Later among the works it cites.
Nips 2016 tutorial: Generative adversarial networks
Ian Goodfellow · 2016
Later among the works it cites.
Identity mappings in deep residual networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Later among the works it cites.
Harnessing deep neural networks with logic rules
Zhiting Hu, Xuezhe Ma, Zhengzhong Liu, Eduard Hovy, and Eric Xing · 2016
Later among the works it cites.
Nal Kalchbrenner, Aaron van den Oord, Karen Simonyan, Ivo Danihelka, Oriol Vinyals, Alex Graves, and Koray Kavukcuoglu · 2016
Later among the works it cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2016
Later among the works it cites.
End-to-end sequence labeling via bi-directional lstm-cnns-crf
Xuezhe Ma and Eduard Hovy · 2016
Later among the works it cites.
Dropout with expectation-linear regularization
Xuezhe Ma, Yingkai Gao, Zhiting Hu, Yaoliang Yu, Yuntian Deng, and Eduard Hovy · 2016
Later among the works it cites.
Pixel recurrent neural networks
Aaron van den Oord, Nal Kalchbrenner, and Koray Kavukcuoglu · 2016
Later among the works it cites.
Improved techniques for training gans
Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen · 2016
Later among the works it cites.
Conditional image generation with pixelcnn decoders
Aaron van den Oord, Nal Kalchbrenner, Lasse Espeholt, Oriol Vinyals, Alex Graves, et al · 2016
Later among the works it cites.
Residual networks behave like ensembles of relatively shallow networks
Andreas Veit, Michael J Wilber, and Serge Belongie · 2016
Later among the works it cites.
Multiple confounders correction with regularized linear mixed effect models, with application in biological processes
Haohan Wang and Jingkang Yang · 2016
Later among the works it cites.
Select-additive learning: Improving cross-individual generalization in multimodal sentiment analysis
Haohan Wang, Aaksha Meghawat, Louis-Philippe Morency, and Eric P Xing · 2016
Later among the works it cites.
Multi-task cross-lingual sequence tagging from scratch
Zhilin Yang, Ruslan Salakhutdinov, and William Cohen · 2016
Later among the works it cites.
Sergey Zagoruyko and Nikos Komodakis · 2016
Later among the works it cites.
Energy-based generative adversarial network
Junbo Zhao, Michael Mathieu, and Yann LeCun · 2016
Later among the works it cites.
Martin Arjovsky, Soumith Chintala, and Léon Bottou · 2017
Closest in time.
Calibrating energy-based generative adversarial networks
Zihang Dai, Amjad Almahairi, Bachman Philip, Eduard Hovy, and Aaron Courville · 2017
Closest in time.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Closest in time.