Fetching the paper…
Reading the bibliography…
The success of deep learning is due in large part to our ability to solve certain massive non-convex optimization problems with relative ease.
Chiyuan Zhang, Samy Bengio, and Yoram Singer · 1902
Earlier work this paper cites.
Differentiable ranks and sorting using optimal transport
Marco Cuturi, Olivier Teboul, and Jean-Philippe Vert · 1905
Earlier work this paper cites.
Johanni Brea, Berfin Simsek, Bernd Illing, and Wulfram Gerstner · 1907
Earlier work this paper cites.
Assignment problems and the location of economic activities
Tjalling C. Koopmans and Martin Beckmann · 1957
Earlier work this paper cites.
P-complete approximation problems
Sartaj Sahni and Teofilo F. Gonzalez · 1976
Earlier work this paper cites.
A shortest augmenting path algorithm for dense and sparse linear assignment problems
Roy Jonker and A. Volgenant · 1987
Earlier work this paper cites.
On the algebraic structure of feedforward network weight spaces
Robert Hecht-Nielsen · 1990
Earlier work this paper cites.
On the geometry of feedforward neural network error surfaces
An Mei Chen, Haw-minn Lu, and Robert Hecht-Nielsen · 1993
Earlier work this paper cites.
Network Optimization: Continuous and Discrete Methods
D.P. Bertsekas · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
The organization of behavior: A neuropsychological theory
Donald Olding Hebb · 2005
Earlier work this paper cites.
Pattern recognition and machine learning, 5th Edition
Christopher M. Bishop · 2007
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
The hungarian method for the assignment problem
Harold W. Kuhn · 2010
Earlier work this paper cites.
Revisiting ”qualitatively characterizing neural network optimization problems”
Jonathan Frankle · 2012
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron C. Courville · 2013
Earlier work this paper cites.
The Quadratic Assignment Problem: Theory and Algorithms
E. Cela · 2013
Earlier work this paper cites.
Maximum quadratic assignment problem: Reduction from maximum label cover and lp-based approximation algorithm
Konstantin Makarychev, Rajsekar Manokaran, and Maxim Sviridenko · 2014
Earlier work this paper cites.
How transferable are features in deep neural networks?
Jason Yosinski, Jeff Clune, Yoshua Bengio, and Hod Lipson · 2014
Earlier work this paper cites.
A method for finding similarity between multi-layer perceptrons by forward bipartite alignment
Stephen C. Ashmore and Michael S. Gashler · 2015
Earlier work this paper cites.
Convex relaxations for permutation problems
Fajwel Fogel, Rodolphe Jenatton, Francis R. Bach, and Alexandre d’Aspremont · 2015
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
Ian J. Goodfellow and Oriol Vinyals · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Lei Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton · 2016
Earlier work this paper cites.
Binarynet: Training deep neural networks with weights and activations constrained to +1 or -1
Matthieu Courbariaux and Yoshua Bengio · 2016
Earlier work this paper cites.
On implementing 2d rectangular assignment algorithms
David Frederic Crouse · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Top-n recommender system via matrix completion
Zhao Kang, Chong Peng, and Qiang Cheng · 2016
Cited alongside, same era.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Cited alongside, same era.
Convergent learning: Do different neural networks learn the same representations?
Yixuan Li, Jason Yosinski, Jeff Clune, Hod Lipson, and John E. Hopcroft · 2016
Cited alongside, same era.
Xnor-net: Imagenet classification using binary convolutional neural networks
Mohammad Rastegari, Vicente Ordonez, Joseph Redmon, and Ali Farhadi · 2016
Cited alongside, same era.
Instance normalization: The missing ingredient for fast stylization
Dmitry Ulyanov, Andrea Vedaldi, and Victor S. Lempitsky · 2016
Cited alongside, same era.
Federated learning with matched averaging
Hongyi Wang, Mikhail Yurochkin, Yuekai Sun, Dimitris S. Papailiopoulos, and Yasaman Khazaeni · 2020
Later among the works it cites.
Group normalization
Yuxin Wu and Kaiming He · 2020
Later among the works it cites.
Faster policy learning with continuous-time gradients
Samuel K. Ainsworth, Kendall Lowrey, John Thickstun, Zaïd Harchaoui, and Siddhartha S. Srinivasa · 2021
Later among the works it cites.
Loss surface simplexes for mode connecting volumes and fast ensembling
Gregory W. Benton, Wesley J. Maddox, Sanae Lotfi, and Andrew Gordon Wilson · 2021
Later among the works it cites.
The role of permutation invariance in linear mode connectivity of neural networks
Rahim Entezari, Hanie Sedghi, Olga Saukh, and Behnam Neyshabur · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Daniel Freeman and Joan Bruna · 2017
Cited alongside, same era.
An introduction to trajectory optimization: How to do your own direct collocation
Matthew Kelly · 2017
Cited alongside, same era.
SGDR: stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2017
Cited alongside, same era.
Communication-efficient learning of deep networks from decentralized data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Agüera y Arcas · 2017
Cited alongside, same era.
oi-vae: Output interpretable vaes for nonlinear group factor analysis
Samuel K. Ainsworth, Nicholas J. Foti, Adrian K. C. Lee, and Emily B. Fox · 2018
Cited alongside, same era.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lénaïc Chizat and Francis R. Bach · 2018
Cited alongside, same era.
Essentially no barriers in neural network energy landscape
Felix Draxler, Kambis Veschgini, Manfred Salmhofer, and Fred A. Hamprecht · 2018
Cited alongside, same era.
Yiding Jiang, Vaishnavh Nagarajan, Christina Baek, and J. Zico Kolter · 2021
Later among the works it cites.
LLC: accurate, multi-purpose learnt low-dimensional binary codes
Aditya Kusupati, Matthew Wallingford, Vivek Ramanujan, Raghav Somani, Jae Sung Park, Krishna Pillutla, Prateek Jain, Sham M. Kakade, and Ali Farhadi · 2021
Later among the works it cites.
On monotonic linear interpolation of neural network parameters
James Lucas, Juhan Bae, Michael R. Zhang, Stanislav Fort, Richard S. Zemel, and Roger B. Grosse · 2021
Later among the works it cites.
Merging models with fisher-weighted averaging
Michael Matena and Colin Raffel · 2021
Later among the works it cites.
Do wide and deep networks learn the same things? uncovering how neural network representations vary with width and depth
Thao Nguyen, Maithra Raghu, and Simon Kornblith · 2021
Later among the works it cites.
Differentiable sorting networks for scalable sorting and ranking supervision
Felix Petersen, Christian Borgelt, Hilde Kuehne, and Oliver Deussen · 2021
Later among the works it cites.
Geometry of the loss landscape in overparameterized neural networks: Symmetries and invariances
Berfin Simsek, François Ged, Arthur Jacot, Francesco Spadaro, Clément Hongler, Wulfram Gerstner, and Johanni Brea · 2021
Later among the works it cites.
Training neural networks with fixed sparse masks
Yi-Lin Sung, Varun Nair, and Colin Raffel · 2021
Later among the works it cites.
What can linear interpolation of neural network loss landscapes tell us?
Tiffany Vlaar and Jonathan Frankle · 2021
Later among the works it cites.
Tent: Fully test-time adaptation by entropy minimization
Dequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen, and Trevor Darrell · 2021
Later among the works it cites.
Learning neural network subspaces
Mitchell Wortsman, Maxwell Horton, Carlos Guestrin, Ali Farhadi, and Mohammad Rastegari · 2021
Later among the works it cites.
Random initialisations performing above chance and how to find them, 2022
Frederik Benzing, Simon Schug, Robert Meier, Johannes von Oswald, Yassir Akram, Nicolas Zucchet, Laurence Aitchison, and Angelika Steger · 2022
Closest in time.
Gradient flow dynamics of shallow relu networks for square loss and orthogonal inputs
Etienne Boursier, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2022
Closest in time.
On the symmetries of deep learning models and their internal representations
Charles Godfrey, Davis Brown, Tegan Emerson, and Henry Kvinge · 2022
Closest in time.
Patching open-vocabulary models by interpolating weights
Gabriel Ilharco, Mitchell Wortsman, Samir Yitzhak Gadre, Shuran Song, Hannaneh Hajishirzi, Simon Kornblith, Ali Farhadi, and Ludwig Schmidt · 2022
Closest in time.
REPAIR: renormalizing permuted activations for interpolation repair
Keller Jordan, Hanie Sedghi, Olga Saukh, Rahim Entezari, and Behnam Neyshabur · 2022
Closest in time.
Linear connectivity reveals generalization strategies
Jeevesh Juneja, Rachit Bansal, Kyunghyun Cho, João Sedoc, and Naomi Saphra · 2022
Closest in time.
Deep neural network fusion via graph matching with applications to model ensemble and federated learning
Chang Liu, Chenfei Lou, Runzhong Wang, Alan Yuhan Xi, Li Shen, and Junchi Yan · 2022
Closest in time.
A convnet for the 2020s
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie · 2022
Closest in time.
Monotonic differentiable sorting networks
Felix Petersen, Christian Borgelt, Hilde Kuehne, and Oliver Deussen · 2022
Closest in time.
Deep networks on toroids: Removing symmetries reveals the structure of flat regions in the landscape geometry
Fabrizio Pittorino, Antonio Ferraro, Gabriele Perugini, Christoph Feinauer, Carlo Baldassi, and Riccardo Zecchina · 2022
Closest in time.
A call to build models like we build open-source software
Colin Raffel · 2022
Closest in time.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Mitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs, Raphael Gontijo-Lopes, Ari S Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, et al · 2022
Closest in time.