Fetching the paper…
Reading the bibliography…
Deep neural networks excel at learning the training data, but often provide incorrect and confident predictions when evaluated on slightly different test examples.
Approximation by superpositions of a sigmoidal function
Cybenko, G · 1989
Earlier work this paper cites.
Lower bounds on the vc dimension of smoothly parameterized function classes
Lee, W. S., Bartlett, P. L., and Williamson, R. C · 1995
Earlier work this paper cites.
Generalization performance of support vector machines and other pattern classifiers, 1998
Bartlett, P. and Shawe-taylor, J · 1998
Earlier work this paper cites.
A theory of learning from different domains
Ben-David, S., Blitzer, J., Crammer, K., Kulesza, A., Pereira, F., and Vaughan, J. W · 2010
Earlier work this paper cites.
Improving neural networks by preventing co-adaptation of feature detectors
Hinton, G. E., Srivastava, N., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2012
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Mikolov, T., Chen, K., Corrado, G., and Dean, J · 2013
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Zeiler, M. D. and Fergus, R · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., and Fergus, R · 2014
Earlier work this paper cites.
Explaining and Harnessing Adversarial Examples
Goodfellow, I. J., Shlens, J., and Szegedy, C · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Deep learning
LeCun, Y., Bengio, Y., and Hinton, G · 2015
Earlier work this paper cites.
Difference target propagation
Lee, D.-H., Zhang, S., Fischer, A., and Bengio, Y · 2015
Earlier work this paper cites.
Deep learning and the information bottleneck principle
Tishby, N. and Zaslavsky, N · 2015
Cited alongside, same era.
MixUp as Locally Linear Out-Of-Manifold Regularization
Guo, H., Mao, Y., and Zhang, R · 2016
Cited alongside, same era.
Identity mappings in deep residual networks
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Improved techniques for training gans
Salimans, T., Goodfellow, I., Zaremba, W., Cheung, V., Radford, A., and Chen, X · 2016
Cited alongside, same era.
Deep variational information bottleneck
Alemi, A. A., Fischer, I., Dillon, J. V., and Murphy, K · 2017
Cited alongside, same era.
Wasserstein generative adversarial networks
Arjovsky, M., Chintala, S., and Bottou, L · 2017
Information dropout: Learning optimal representations through noisy computation
Achille, A. and Soatto, S · 2018
Closest in time.
Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples
Athalye, A., Carlini, N., and Wagner, D · 2018
Closest in time.
Assessing the scalability of biologically-motivated deep learning algorithms and architectures
Bartunov, S., Santoro, A., Richards, B. A., Hinton, G. E., and Lillicrap, T · 2018
Closest in time.
MINE: mutual information neural estimation
Belghazi, I., Rajeswar, S., Baratin, A., Hjelm, R. D., and Courville, A. C · 2018
Closest in time.
Transfer and exploration via the information bottleneck
Goyal, A., Islam, R., Strouse, D., Ahmed, Z., Larochelle, H., Botvinick, M., Levine, S., and Bengio, Y · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Improved regularization of convolutional neural networks with cutout
Devries, T. and Taylor, G. W · 2017
Cited alongside, same era.
Improved training of wasserstein gans
Gulrajani, I., Ahmed, F., Arjovsky, M., Dumoulin, V., and Courville, A. C · 2017
Cited alongside, same era.
Equilibrium propagation: Bridging the gap between energy-based models and backpropagation
Scellier, B. and Bengio, Y · 2017
Cited alongside, same era.
Opening the black box of deep neural networks via information
Shwartz-Ziv, R. and Tishby, N · 2017
Cited alongside, same era.
An approximation of the error backpropagation algorithm in a predictive coding network with local hebbian synaptic plasticity
Whittington, J. C. and Bogacz, R · 2017
Cited alongside, same era.
MixUp as Locally Linear Out-Of-Manifold Regularization
Guo, H., Mao, Y., and Zhang, R
Cited in the paper.
Towards deep learning models resistant to adversarial attacks
Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A · 2018
Closest in time.
Spectral normalization for generative adversarial networks
Miyato, T., Kataoka, T., Koyama, M., and Yoshida, Y · 2018
Closest in time.
Between-class learning for image classification
Tokozume, Y., Ushiku, Y., and Harada, T · 2018
Closest in time.
Deep learning with data dependent implicit activation function
Wang, B., Luo, X., Li, Z., Zhu, W., Shi, Z., and Osher, S. J · 2018
Closest in time.
mixup: Beyond empirical risk minimization
Zhang, H., Cisse, M., Dauphin, Y. N., and Lopez-Paz, D · 2018
Closest in time.
Zhao, J. and Cho, K · 2018
Closest in time.