Fetching the paper…
Reading the bibliography…
In this paper, we propose a simple but effective method for training neural networks with a limited amount of training data.
Multilayer feedforward networks are universal approximators
K. Hornik, M. Stinchcombe, and H. White · 1989
Earlier work this paper cites.
Learning many related tasks at the same time with backpropagation
R. Caruana · 1994
Earlier work this paper cites.
Algorithm 778: L-bfgs-b: Fortran subroutines for large-scale bound-constrained optimization
C. Zhu, R. H. Byrd, P. Lu, and J. Nocedal · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Model compression
C. Buciluǎ, R. Caruana, and A. Niculescu-Mizil · 2006
Earlier work this paper cites.
A unifying view of sparse approximate gaussian process regression
J. Quinonero Candela and C. E. Rasmussen · 2006
Earlier work this paper cites.
Sparse Gaussian processes using pseudo-inputs
E. Snelson and Z. Ghahramani · 2006
Earlier work this paper cites.
Bayesian multicategory support vector machines
Z. Zhang and M. Jordan · 2006
Earlier work this paper cites.
The variational Gaussian approximation revisited
M. Opper and C. Archambeau · 2009
Earlier work this paper cites.
Variational learning of inducing variables in sparse Gaussian processes
M. Titsias · 2009
Earlier work this paper cites.
Gaussian processes for big data
J. Hensman, N. Fusi, and N. D. Lawrence · 2013
Earlier work this paper cites.
Do deep nets really need to be deep?
J. Ba and R. Caruana · 2014
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Semi-supervised learning with deep generative models
D. P Kingma, S. Mohamed, D. Jimenez Rezende, and M. Welling · 2014
Earlier work this paper cites.
Explaining and harnessing adversarial examples
I. Goodfellow, J. Shlens, and C. Szegedy · 2015
Cited alongside, same era.
Scalable variational Gaussian process classification
J. Hensman, A. Matthews, and Z. Ghahramani · 2015
Cited alongside, same era.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, and J. Dean · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Cited alongside, same era.
Siamese neural networks for one-shot image recognition
G. Koch, R. Zemel, and R. Salakhutdinov · 2015
Cited alongside, same era.
Fitnets: Hints for thin deep nets
A. Romero, N. Ballas, S. E. Kahou, A. Chassang, C. Gatta, and Y. Bengio · 2015
Cited alongside, same era.
Matching networks for one shot learning
O. Vinyals, C. Blundell, T. Lillicrap, K. Kavukcuoglu, and D. Wierstra · 2016
Later among the works it cites.
Adversarial examples, uncertainty, and transfer testing robustness in Gaussian process hybrid deep networks
J. Bradshaw, A. G. d. G. Matthews, and Z. Ghahramani · 2017
Later among the works it cites.
Fidelity-weighted learning
M. Dehghani, A. Mehrjou, S. Gouws, J. Kamps, and B. Schölkopf · 2017
Later among the works it cites.
Born again neural networks
T. Furlanello, Z. C. Lipton, L. Itti, and A. Anandkumar · 2017
Later among the works it cites.
Data-free knowledge distillation for deep neural networks
R. Gontijo Lopes, S. Fenu, and T. Starner · 2017
Later among the works it cites.
Knowledge distillation using unlabeled mismatched images
M. Kulkarni, K. Patil, and S. S. Karande · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Cited alongside, same era.
Learning feed-forward one-shot learners
L. Bertinetto, J. F. Henriques, J. Valmadre, P. Torr, and A. Vedaldi · 2016
Cited alongside, same era.
Incorporating Nesterov momentum into Adam
T. Dozat · 2016
Cited alongside, same era.
Dropout as a Bayesian approximation: Representing model uncertainty in deep learning
Y. Gal and Z. Ghahramani · 2016
Cited alongside, same era.
Distributional smoothing by virtual adversarial examples
T. Miyato, S. Maeda, M. Koyama, K. Nakae, and S. Ishii · 2016
Cited alongside, same era.
Transferability in Machine Learning: from Phenomena to Black-Box Attacks using Adversarial Samples
N. Papernot, P. McDaniel, and I. Goodfellow · 2016
Cited alongside, same era.
GPflow: A Gaussian process library using TensorFlow
A. G. d. G. Matthews, M. van der Wilk, T. Nickson, K. Fujii, A. Boukouvalas, P. León-Villagrá, Z. Ghahramani, and J. Hensman · 2017
Later among the works it cites.
Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results
A. Tarvainen and H. Valpola · 2017
Later among the works it cites.
Fashion-MNIST: A novel image dataset for benchmarking machine learning algorithms
H. Xiao, K. Rasul, and R. Vollgraf · 2017
Later among the works it cites.
A gift from knowledge distillation: Fast optimization, network minimization and transfer learning
J. Yim, D. Joo, J. Bae, and J. Kim · 2017
Later among the works it cites.
Learning from multiple teacher networks
S. You, C. Xu, C. Xu, and D. Tao · 2017
Later among the works it cites.
Generating natural adversarial examples
Z. Zhao, D. Dua, and S. Singh · 2017
Later among the works it cites.
Spatially transformed adversarial examples
C. Xiao, J.-Y. Zhu, B. Li, W. He, M. Liu, and D. Song · 2018
Closest in time.