Fetching the paper…
Reading the bibliography…
Adversarial training is a widely used strategy for making neural networks resistant to adversarial perturbations.
On a modification of chebyshev’s inequality and of the error formula of laplace
Sergei Bernstein · 1924
Earlier work this paper cites.
A measure of asymptotic efficiency for tests of a hypothesis based on the sum of observations
Herman Chernoff · 1952
Earlier work this paper cites.
Probability inequalities for sums of bounded random variables
Wassily Hoeffding · 1963
Earlier work this paper cites.
Dynamic half-space reporting, geometric optimization, and minimum spanning trees
Pankaj K Agarwal, David Eppstein, and Jiri Matousek · 1992
Earlier work this paper cites.
Adaptive estimation of a quadratic functional by model selection
B. Laurent and P. Massart · 2000
Earlier work this paper cites.
Gaussian processes: inequalities, small ball probabilities and applications
Wenbo V Li and Q-M Shao · 2001
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2013
Earlier work this paper cites.
Faster algorithms via approximation theory
Sushant Sachdeva, Nisheeth K Vishnoi, et al · 2014
Earlier work this paper cites.
Principal component projection without principal component analysis
Roy Frostig, Cameron Musco, Christopher Musco, and Aaron Sidford · 2016
Earlier work this paper cites.
Towards evaluating the robustness of neural networks
Nicholas Carlini and David Wagner · 2017
Earlier work this paper cites.
Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples
Anish Athalye, Nicholas Carlini, and David Wagner · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2018
Earlier work this paper cites.
Learning and generalization in overparameterized neural networks, going beyond two layers
Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang · 2019
Earlier work this paper cites.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Earlier work this paper cites.
On the convergence rate of training recurrent neural networks
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Earlier work this paper cites.
Gradient descent finds global minima of deep neural networks
Simon S Du, Jason D Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2019
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2019
Cited alongside, same era.
Solving empirical risk minimization in the current matrix multiplication time
Yin Tat Lee, Zhao Song, and Qiuyi Zhang · 2019
Cited alongside, same era.
Quadratic suffices for over-parametrization via matrix chernoff bound
Zhao Song and Xin Yang · 2019
Cited alongside, same era.
Slide: In defense of smart algorithms over hardware acceleration for large-scale deep learning systems
Beidi Chen, Tharun Medini, James Farwell, Charlie Tai, Anshumali Shrivastava, et al · 2020
Cited alongside, same era.
A nearly-linear time algorithm for linear programs with small treewidth: A multiscale representation of robust central path
Sally Dong, Yin Tat Lee, and Guanghao Ye · 2021
Later among the works it cites.
Accelerating slide deep learning on modern cpus: Vectorization, quantizations, memory optimizations, and more
Shabnam Daghaghi, Nicholas Meisburger, Mengnan Zhao, and Anshumali Shrivastava · 2021
Later among the works it cites.
Fl-ntk: A neural tangent kernel-based framework for federated learning analysis
Baihe Huang, Xiaoxiao Li, Zhao Song, and Xin Yang · 2021
Later among the works it cites.
A faster algorithm for solving general lps
Shunhua Jiang, Zhao Song, Omri Weinstein, and Hengjie Zhang · 2021
Later among the works it cites.
Breaking the linear iteration cost barrier for some well-known conditional gradient methods using maxip data-structures
Anshumali Shrivastava, Zhao Song, and Zhaozhuo Xu · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A faster interior point method for semidefinite programming
Haotian Jiang, Tarun Kathuria, Yin Tat Lee, Swati Padmanabhan, and Zhao Song · 2020
Cited alongside, same era.
Reformer: The efficient transformer
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya · 2020
Cited alongside, same era.
Generalized leverage score sampling for neural networks
Jason D Lee, Ruoqi Shen, Zhao Song, Mengdi Wang, and Zheng Yu · 2020
Cited alongside, same era.
Toward moderate overparameterization: Global convergence guarantees for training shallow neural networks
Samet Oymak and Mahdi Soltanolkotabi · 2020
Cited alongside, same era.
Fast algorithm for solving structured convex programs
Guanghao Ye · 2020
Cited alongside, same era.
Gradient descent optimizes over-parameterized deep relu networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2020
Cited alongside, same era.
Over-parameterized adversarial training: An analysis overcoming the curse of dimensionality
Yi Zhang, Orestis Plevrakis, Simon S Du, Xingguo Li, Zhao Song, and Sanjeev Arora · 2020
Cited alongside, same era.
Anshumali Shrivastava, Zhao Song, and Zhaozhuo Xu · 2021
Later among the works it cites.
Oblivious sketching-based central path method for linear programming
Zhao Song and Zheng Yu · 2021
Later among the works it cites.
Does preprocessing help training over-parameterized neural networks?
Zhao Song, Shuo Yang, and Ruizhe Zhang · 2021
Later among the works it cites.
Training multi-layer over-parametrized neural network in subquadratic time
Zhao Song, Lichen Zhang, and Ruizhe Zhang · 2021
Later among the works it cites.
Solving sdp faster: A robust ipm framework and efficient implementation
Baihe Huang, Shunhua Jiang, Zhao Song, Runzhou Tao, and Ruizhe Zhang · 2022
Closest in time.
Training overparametrized neural networks in sublinear time
Hang Hu, Zhao Song, Omri Weinstein, and Danyang Zhuo · 2022
Closest in time.
A faster interior-point method for sum-of-squares optimization
Shunhua Jiang, Bento Natura, and Omri Weinstein · 2022
Closest in time.
Bounding the width of neural networks via coupled initialization a worst case analysis
Alexander Munteanu, Simon Omlor, Zhao Song, and David Woodruff · 2022
Closest in time.
Accelerating frank-wolfe algorithm using low-dimensional and adaptive data structures
Zhao Song, Zhaozhuo Xu, Yuanyuan Yang, and Lichen Zhang · 2022
Closest in time.
Speeding up sparsification using inner product search data structures
Zhao Song, Zhaozhuo Xu, and Lichen Zhang · 2022
Closest in time.
Speeding up optimizations via data structures: Faster search, sample and maintenance
Lichen Zhang · 2022
Closest in time.