Fetching the paper…
Reading the bibliography…
Data augmentation is commonly applied to improve performance of deep learning by enforcing the knowledge that certain transformations on the input preserve the output.
Neocognitron: A self-organizing neural network model for a mechanism of visual pattern recognition
Kunihiko Fukushima and Sei Miyake · 1982
Earlier work this paper cites.
A practical bayesian framework for backpropagation networks
David JC MacKay · 1992
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Occam’s razor
Carl Edward Rasmussen and Zoubin Ghahramani · 2001
Earlier work this paper cites.
Information theory, inference and learning algorithms
David JC MacKay · 2003
Earlier work this paper cites.
Nineteen dubious ways to compute the exponential of a matrix, twenty-five years later
Cleve Moler and Charles Van Loan · 2003
Earlier work this paper cites.
Pattern recognition and machine learning
Christopher M Bishop · 2006
Earlier work this paper cites.
Gaussian processes for machine learning
Carl Edward Rasmussen and Christopher KI Williams · 2006
Earlier work this paper cites.
The minimum description length principle
Peter D Grünwald · 2007
Earlier work this paper cites.
Group theoretical methods in machine learning, 2008
Imre Risi Kondor · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Léon Bottou · 2010
Earlier work this paper cites.
MNIST handwritten digit database
Yann LeCun and Corinna Cortes · 2010
Earlier work this paper cites.
Machine learning: a probabilistic perspective
Kevin P Murphy · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Weight uncertainty in neural networks
Charles Blundell, Julien Cornebise, Koray Kavukcuoglu, and Daan Wierstra · 2015
Earlier work this paper cites.
Spatial transformer networks
Max Jaderberg, Karen Simonyan, Andrew Zisserman, et al · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Earlier work this paper cites.
Group equivariant convolutional networks
Taco Cohen and Max Welling · 2016
Earlier work this paper cites.
Pac-bayesian theory meets bayesian inference
Pascal Germain, Francis Bach, Alexandre Lacoste, and Simon Lacoste-Julien · 2016
Cited alongside, same era.
On degeneracy and invariances of random fields paths with applications in gaussian process modelling
David Ginsbourger, Olivier Roustant, and Nicolas Durrande · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Wide residual networks
Sergey Zagoruyko and Nikos Komodakis · 2016
Cited alongside, same era.
Practical Gauss-Newton optimisation for deep learning
Aleksandar Botev, Hippolyt Ritter, and David Barber · 2017
Cited alongside, same era.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Cited alongside, same era.
Fixup initialization: Residual learning without normalization
Hongyi Zhang, Yann N Dauphin, and Tengyu Ma · 2019
Later among the works it cites.
Learning invariances in neural networks
Gregory Benton, Marc Finzi, Pavel Izmailov, and Andrew Gordon Wilson · 2020
Later among the works it cites.
On the marginal likelihood and cross-validation
Edwin Fong and CC Holmes · 2020
Later among the works it cites.
Optimizing millions of hyperparameters by implicit differentiation
Jonathan Lorraine, Paul Vicol, and David Duvenaud · 2020
Later among the works it cites.
New insights and perspectives on the natural gradient method
James Martens · 2020
Later among the works it cites.
Scalable and practical natural gradient for large-scale deep learning
Kazuki Osawa, Yohei Tsuji, Yuichiro Ueno, Akira Naruse, Chuan-Sheng Foo, and Rio Yokota · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms, 2017
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Cited alongside, same era.
Taco S Cohen, Mario Geiger, Jonas Köhler, and Max Welling · 2018
Cited alongside, same era.
Autoaugment: Learning augmentation policies from data
Ekin D Cubuk, Barret Zoph, Dandelion Mane, Vijay Vasudevan, and Quoc V Le · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Clebsch–gordan nets: a fully fourier space spherical convolutional neural network
Risi Kondor, Zhen Lin, and Shubhendu Trivedi · 2018
Cited alongside, same era.
A scalable laplace approximation for neural networks
Hippolyt Ritter, Aleksandar Botev, and David Barber · 2018
Cited alongside, same era.
Later among the works it cites.
Mdp homomorphic networks: Group symmetries in reinforcement learning
Elise van der Pol, Daniel Worrall, Herke van Hoof, Frans Oliehoek, and Max Welling · 2020
Later among the works it cites.
How good is the bayes posterior in deep neural networks really?
Florian Wenzel, Kevin Roth, Bastiaan S Veeling, Jakub Światkowski, Linh Tran, Stephan Mandt, Jasper Snoek, Tim Salimans, Rodolphe Jenatton, and Sebastian Nowozin · 2020
Later among the works it cites.
Meta-learning symmetries by reparameterization
Allan Zhou, Tom Knowles, and Chelsea Finn · 2020
Later among the works it cites.
Se (3)-equivariant graph neural networks for data-efficient and accurate interatomic potentials
Simon Batzner, Albert Musaelian, Lixin Sun, Mario Geiger, Jonathan P Mailoa, Mordechai Kornbluth, Nicola Molinari, Tess E Smidt, and Boris Kozinsky · 2021
Later among the works it cites.
Geometric and physical quantities improve e (3) equivariant message passing
Johannes Brandstetter, Rob Hesselink, Elise van der Pol, Erik Bekkers, and Max Welling · 2021
Later among the works it cites.
Laplace redux-effortless bayesian deep learning
Erik Daxberger, Agustinus Kristiadi, Alexander Immer, Runa Eschenhagen, Matthias Bauer, and Philipp Hennig · 2021
Later among the works it cites.
Data augmentation in bayesian neural networks and the cold posterior effect
Seth Nabarro, Stoil Ganev, Adrià Garriga-Alonso, Vincent Fortuin, Mark van der Wilk, and Laurence Aitchison · 2021
Later among the works it cites.
Global inducing point variational posteriors for bayesian neural networks and deep gaussian processes
Sebastian W Ober and Laurence Aitchison · 2021
Later among the works it cites.
The promises and pitfalls of deep kernel learning
Sebastian W. Ober, Carl E. Rasmussen, and Mark van der Wilk · 2021
Later among the works it cites.
Improving deterministic uncertainty estimation in deep learning for classification and regression
Joost van Amersfoort, Lewis Smith, Andrew Jesson, Oscar Key, and Yarin Gal · 2021
Later among the works it cites.
Learning invariant weights in neural networks
Tycho FA van der Ouderaa and Mark van der Wilk · 2021
Later among the works it cites.
Probing as quantifying inductive bias
Alexander Immer, Lucas Torroba Hennigen, Vincent Fortuin, and Ryan Cotterell · 2022
Closest in time.
Last layer marginal likelihood for invariance learning
Pola Schwöbel, Martin Jørgensen, Sebastian W. Ober, and Mark van der Wilk · 2022
Closest in time.