Fetching the paper…
Reading the bibliography…
In this paper we study a constraint-based representation of neural network architectures.
D. E. Rumelhart, G. E. Hinton, R. J. Williams et al. , “Learning representations by back-propagating errors,” Cognitive modeling , vol. 5, no. 3, p. 1, 1988
1988
Earlier work this paper cites.
Y. LeCun, D. Touresky, G. Hinton, and T. Sejnowski, “A theoretical framework for back-propagation,” in Proceedings of the 1988 connectionist models summer school , vol. 1. CMU, Pittsburgh, Pa: Morgan Kaufmann, 1988, pp. 21–28
1988
Earlier work this paper cites.
J. C. Platt and A. H. Barr, “Constrained differential optimization,” in Neural Information Processing Systems , 1988, pp. 612–621
1988
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE , vol. 86, no. 11, pp. 2278–2324, 1998
1998
Earlier work this paper cites.
S. Hochreiter, Y. Bengio, P. Frasconi, J. Schmidhuber et al. , “Gradient flow in recurrent nets: the difficulty of learning long-term dependencies,” 2001
2001
Earlier work this paper cites.
F. Scarselli, M. Gori, A. C. Tsoi, M. Hagenbuchner, and G. Monfardini, “The graph neural network model,” IEEE Transactions on Neural Networks , vol. 20, no. 1, pp. 61–80, 2009
2009
Earlier work this paper cites.
M. Carreira-Perpinan and W. Wang, “Distributed optimization of deeply nested systems,” in Artificial Intelligence and Statistics , 2014, pp. 10–19
2014
Earlier work this paper cites.
D. P. Bertsekas, Constrained optimization and Lagrange multiplier methods . Academic press, 2014
2014
Earlier work this paper cites.
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting,” The Journal of Machine Learning Research , vol. 15, no. 1, pp. 1929–1958, 2014
2014
Cited alongside, same era.
M. Fernández-Delgado, E. Cernadas, S. Barro, and D. Amorim, “Do we need hundreds of classifiers to solve real world classification problems?” The Journal of Machine Learning Research , vol. 15, no. 1, pp. 3133–3181, 2014
2014
Cited alongside, same era.
J. Schmidhuber, “Deep learning in neural networks: An overview,” Neural networks , vol. 61, pp. 85–117, 2015
2015
Cited alongside, same era.
G. Gnecco, M. Gori, S. Melacci, and M. Sanguineti, “Foundations of support constraint machines,” Neural computation , vol. 27, no. 2, pp. 388–480, 2015
2015
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Later among the works it cites.
——, “Identity mappings in deep residual networks,” in European conference on computer vision . Springer, 2016, pp. 630–645
2016
Later among the works it cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems , 2017, pp. 5998–6008
2017
Later among the works it cites.
M. Gori, Machine Learning: A Constraint-based Approach . Morgan Kaufmann, 2017
2017
Later among the works it cites.
D. Dua and E. Karra Taniskidou, “UCI machine learning repository,” 2017. [Online]. Available: http://archive.ics.uci.edu/ml
2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
2015
Cited alongside, same era.
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola, “Stacked attention networks for image question answering,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , June 2016
2016
Cited alongside, same era.
H. Xu and K. Saenko, “Ask, attend and answer: Exploring question-guided spatial attention for visual question answering,” in European Conference on Computer Vision . Springer, 2016, pp. 451–466
2016
Cited alongside, same era.
G. Taylor, R. Burmeister, Z. Xu, B. Singh, A. Patel, and T. Goldstein, “Training neural networks without gradients: A scalable admm approach,” in International conference on machine learning , 2016, pp. 2722–2731
2016
Cited alongside, same era.
Later among the works it cites.
A. Gotmare, V. Thomas, J. Brea, and M. Jaggi, “Decoupling backpropagation using constrained optimization methods,” in Workshop on Efficient Credit Assignment in Deep Learning and Deep Reinforement Learning, ICML 2018. , 2018, pp. 1–11
2018
Later among the works it cites.
2018
Later among the works it cites.
M. Olson, A. Wyner, and R. Berk, “Modern neural networks generalize on small data sets,” in Advances in Neural Information Processing Systems 31 , S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, Eds. Curran Associates, Inc., 2018, pp. 3619–3628. [Online]. Available: http://papers.nips.cc/paper/7620-modern-neural-networks-generalize-on-small-data-sets.pdf
2018
Later among the works it cites.