Fetching the paper…
Reading the bibliography…
The inductive bias of a neural network is largely determined by the architecture and the training algorithm.
Foundations of modern potential theory
N.S. Landkof · 1972
Earlier work this paper cites.
Iterative algorithms for gram-schmidt orthogonalization
Walter Hoffmann · 1989
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
A unified view of the orthogonalization methods
Vipin Srivastava · 2000
Earlier work this paper cites.
Stochastical approximation of smooth convex bodies
Matthias Reitzner · 2004
Earlier work this paper cites.
Extremal systems of points and numerical integration on the sphere
Ian H Sloan and Robert S Womersley · 2004
Earlier work this paper cites.
Minimal riesz energy point configurations for rectifiable d-dimensional manifolds
DP Hardin and EB Saff · 2005
Earlier work this paper cites.
Labeled faces in the wild: A database forstudying face recognition in unconstrained environments
Gary B Huang, Marwan Mattar, Tamara Berg, and Eric Learned-Miller · 2008
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2008
Earlier work this paper cites.
Collective classification in network data
Prithviraj Sen, Galileo Namata, Mustafa Bilgic, Lise Getoor, Brian Galligher, and Tina Eliassi-Rad · 2008
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Singular value decomposition and the centrality of löwdin orthogonalizations
Ramesh Naidu Annavarapu · 2013
Earlier work this paper cites.
Min Lin, Qiang Chen, and Shuicheng Yan · 2013
Earlier work this paper cites.
A feasible method for optimization with orthogonality constraints
Zaiwen Wen and Wotao Yin · 2013
Earlier work this paper cites.
Speeding up convolutional neural networks with low rank expansions
Max Jaderberg, Andrea Vedaldi, and Andrew Zisserman · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Learning face representation from scratch
Dong Yi, Zhen Lei, Shengcai Liao, and Stan Z Li · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Sparse convolutional neural networks
Baoyuan Liu, Min Wang, Hassan Foroosh, Marshall Tappen, and Marianna Pensky · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Earlier work this paper cites.
3d shapenets: A deep representation for volumetric shapes
Zhirong Wu, Shuran Song, Aditya Khosla, Fisher Yu, Linguang Zhang, Xiaoou Tang, and Jianxiong Xiao · 2015
Earlier work this paper cites.
Deep fried convnets
Zichao Yang, Marcin Moczulski, Misha Denil, Nando de Freitas, Alex Smola, Le Song, and Ziyu Wang · 2015
Earlier work this paper cites.
Unitary evolution recurrent neural networks
Martin Arjovsky, Amar Shah, and Yoshua Bengio · 2016
Earlier work this paper cites.
Ms-celeb-1m: A dataset and benchmark for large-scale face recognition
Yandong Guo, Lei Zhang, Yuxiao Hu, Xiaodong He, and Jianfeng Gao · 2016
Earlier work this paper cites.
David Ha, Andrew Dai, and Quoc V Le · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Recurrent orthogonal networks and long-memory tasks
Mikael Henaff, Arthur Szlam, and Yann LeCun · 2016
Cited alongside, same era.
Dynamic filter networks
Xu Jia, Bert De Brabandere, Tinne Tuytelaars, and Luc V Gool · 2016
Cited alongside, same era.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Cited alongside, same era.
Semi-supervised classification with graph convolutional networks
Thomas N Kipf and Max Welling · 2016
Cited alongside, same era.
Gradient descent only converges to minimizers
Jason D Lee, Max Simchowitz, Michael I Jordan, and Benjamin Recht · 2016
Random point sets on the sphere—hole radii, covering, and separation
Johann S Brauchart, Alexander B Reznikov, Edward B Saff, Ian H Sloan, Yu Guang Wang, and Robert S Womersley · 2018
Later among the works it cites.
Points on manifolds with asymptotically optimal covering radius
Anna Breger, Martin Ehler, and Manuel Gräf · 2018
Later among the works it cites.
Implicit bias of gradient descent on linear convolutional networks
Suriya Gunasekar, Jason D Lee, Daniel Soudry, and Nati Srebro · 2018
Later among the works it cites.
Orthogonal weight normalization: Solution to optimization over multiple dependent stiefel manifolds in deep neural networks
Lei Huang, Xianglong Liu, Bo Lang, Adams Wei Yu, Yongliang Wang, and Bo Li · 2018
Later among the works it cites.
Averaging weights leads to wider optima and better generalization
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Large-margin softmax loss for convolutional neural networks
Weiyang Liu, Yandong Wen, Zhiding Yu, and Meng Yang · 2016
Cited alongside, same era.
No bad local minima: Data independent training error guarantees for multilayer neural networks
Daniel Soudry and Yair Carmon · 2016
Cited alongside, same era.
Matching networks for one shot learning
Oriol Vinyals, Charles Blundell, Timothy Lillicrap, Daan Wierstra, et al · 2016
Cited alongside, same era.
Full-capacity unitary recurrent neural networks
Scott Wisdom, Thomas Powers, John Hershey, Jonathan Le Roux, and Les Atlas · 2016
Cited alongside, same era.
Neural photo editing with introspective adversarial networks
Andrew Brock, Theodore Lim, James M Ritchie, and Nick Weston · 2017
Cited alongside, same era.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina · 2017
Cited alongside, same era.
Deep semi-random features for nonlinear function approximation
Kenji Kawaguchi, Bo Xie, and Le Song · 2018
Later among the works it cites.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein · 2018
Later among the works it cites.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Yuanzhi Li, Tengyu Ma, and Hongyang Zhang · 2018
Later among the works it cites.
Learning towards minimum hyperspherical energy
Weiyang Liu, Rongmei Lin, Zhen Liu, Lixin Liu, Zhiding Yu, Bo Dai, and Le Song · 2018
Later among the works it cites.
Piggyback: Adapting a single network to multiple tasks by learning to mask weights
Arun Mallya, Dillon Davis, and Svetlana Lazebnik · 2018
Later among the works it cites.
Additive margin softmax for face verification
Feng Wang, Weiyang Liu, Haijun Liu, and Jian Cheng · 2018
Later among the works it cites.
Cosface: Large margin cosine loss for deep face recognition
Hao Wang, Yitong Wang, Zheng Zhou, Xing Ji, Dihong Gong, Jingchao Zhou, Zhifeng Li, and Wei Liu · 2018
Later among the works it cites.
A closer look at few-shot classification
Wei-Yu Chen, Yen-Cheng Liu, Zsolt Kira, Yu-Chiang Frank Wang, and Jia-Bin Huang · 2019
Later among the works it cites.
Arcface: Additive angular margin loss for deep face recognition
Jiankang Deng, Jia Guo, Niannan Xue, and Stefanos Zafeiriou · 2019
Later among the works it cites.
Acnet: Strengthening the kernel skeletons for powerful cnn via asymmetric convolution blocks
Xiaohan Ding, Yuchen Guo, Guiguang Ding, and Jungong Han · 2019
Later among the works it cites.
Orthogonal deep neural networks
Kui Jia, Shuai Li, Yuxin Wen, Tongliang Liu, and Dacheng Tao · 2019
Later among the works it cites.
Trivializations for gradient-based optimization on manifolds
Mario Lezcano-Casado · 2019
Later among the works it cites.
Cheap orthogonal constraints in neural networks: A simple parametrization of the orthogonal and unitary group
Mario Lezcano-Casado and David Martínez-Rubio · 2019
Later among the works it cites.
Neural similarity learning
Weiyang Liu, Zhen Liu, James M Rehg, and Le Song · 2019
Later among the works it cites.
On the convergence of adam and beyond
Sashank J Reddi, Satyen Kale, and Sanjiv Kumar · 2019
Later among the works it cites.
Controllable orthogonalization in training dnns
Lei Huang, Li Liu, Fan Zhu, Diwen Wan, Zehuan Yuan, Bo Li, and Ling Shao · 2020
Closest in time.
Efficient riemannian optimization on the stiefel manifold via the cayley transform
Jun Li, Fuxin Li, and Sinisa Todorovic · 2020
Closest in time.
Regularizing neural networks via minimizing hyperspherical energy
Rongmei Lin, Weiyang Liu, Zhen Liu, Chen Feng, Zhiding Yu, James M. Rehg, Li Xiong, and Le Song · 2020
Closest in time.
Deep isometric learning for visual recognition
Haozhi Qi, Chong You, Xiaolong Wang, Yi Ma, and Jitendra Malik · 2020
Closest in time.
Orthogonal relation transforms with graph context modeling for knowledge graph embedding
Yun Tang, Jing Huang, Guangtao Wang, Xiaodong He, and Bowen Zhou · 2020
Closest in time.
Learning with hyperspherical uniformity
Weiyang Liu, Rongmei Lin, Zhen Liu, Li Xiong, Bernhard Schölkopf, and Adrian Weller · 2021
Closest in time.