Fetching the paper…
Reading the bibliography…
Stochastic gradient descent (SGD) now acts as a fundamental part of optimization in current machine learning.
On a modification of chebyshev’s inequality and of the error formula of laplace
Sergei Bernstein · 1924
Earlier work this paper cites.
A measure of asymptotic efficiency for tests of a hypothesis based on the sum of observations
Herman Chernoff · 1952
Earlier work this paper cites.
Probability inequalities for sums of bounded random variables
Wassily Hoeffding · 1963
Earlier work this paper cites.
Quad trees a data structure for retrieval on composite keys
Raphael A Finkel and Jon Louis Bentley · 1974
Earlier work this paper cites.
Multidimensional binary search trees used for associative searching
Jon Louis Bentley · 1975
Earlier work this paper cites.
Dynamic half-space reporting, geometric optimization, and minimum spanning trees
Pankaj K Agarwal, David Eppstein, and Jiri Matousek · 1992
Earlier work this paper cites.
Efficient partition trees
Jiri Matousek · 1992
Earlier work this paper cites.
Reporting points in halfspaces
Jiri Matousek · 1992
Earlier work this paper cites.
Algebraic complexity theory
Peter Bürgisser, Michael Clausen, and Mohammad A Shokrollahi · 1997
Earlier work this paper cites.
Approximate nearest neighbors: towards removing the curse of dimensionality
Piotr Indyk and Rajeev Motwani · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Multiresolution gray-scale and rotation invariant texture classification with local binary patterns
Timo Ojala, Matti Pietikainen, and Topi Maenpaa · 2002
Earlier work this paper cites.
Kernel methods for predicting protein–protein interactions
Asa Ben-Hur and William Stafford Noble · 2005
Earlier work this paper cites.
Histograms of oriented gradients for human detection
Navneet Dalal and Bill Triggs · 2005
Earlier work this paper cites.
Optimal halfspace range reporting in three dimensions
Peyman Afshani and Timothy M Chan · 2009
Earlier work this paper cites.
Feature selection in the tensor product feature space
Aaron Smalter, Jun Huan, and Gerald Lushington · 2009
Earlier work this paper cites.
Robust object tracking with online multiple instance learning
Boris Babenko, Ming-Hsuan Yang, and Serge Belongie · 2010
Earlier work this paper cites.
Natural language processing (almost) from scratch
Ronan Collobert, Jason Weston, Léon Bottou, Michael Karlen, Koray Kavukcuoglu, and Pavel Kuksa · 2011
Earlier work this paper cites.
Optimal partition trees
Timothy M Chan · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Multiplying matrices faster than coppersmith-winograd
Virginia Vassilevska Williams · 2012
Earlier work this paper cites.
Fusion with diffusion for robust visual tracking
Yu Zhou, Xiang Bai, Wenyu Liu, and Longin Latecki · 2012
Earlier work this paper cites.
Fast matrix multiplication
Markus Bläser · 2013
Earlier work this paper cites.
Powers of tensors and fast matrix multiplication
François Le Gall · 2014
Earlier work this paper cites.
Song Han, Huizi Mao, and William J Dally · 2015
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Song Han, Jeff Pool, John Tran, and William Dally · 2015
Earlier work this paper cites.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Field-aware factorization machines for ctr prediction
Yuchin Juan, Yong Zhuang, Wei-Sheng Chin, and Chih-Jen Lin · 2016
Cited alongside, same era.
Sgd learns the conjugate kernel class of the network
Amit Daniely · 2017
Cited alongside, same era.
Deep learning in bioinformatics
Seonwoo Min, Byunghan Lee, and Sungroh Yoon · 2017
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Cited alongside, same era.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Cited alongside, same era.
Discrepancy minimization in input-sparsity time
Yichuan Deng, Zhao Song, and Omri Weinstein · 2022
Later among the works it cites.
A faster small treewidth sdp solver
Yuzhou Gu and Zhao Song · 2022
Later among the works it cites.
Solving sdp faster: A robust ipm framework and efficient implementation
Baihe Huang, Shunhua Jiang, Zhao Song, Runzhou Tao, and Ruizhe Zhang · 2022
Later among the works it cites.
Deep kronecker neural networks: A general framework for neural networks with adaptive activation functions
Ameya D Jagtap, Yeonjong Shin, Kenji Kawaguchi, and George Em Karniadakis · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Cited alongside, same era.
On the convergence rate of training recurrent neural networks
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Cited alongside, same era.
Dynamic matrix inverse: Improved algorithms and matching conditional lower bounds
Jan van den Brand, Danupon Nanongkai, and Thatchaphol Saranurak · 2019
Cited alongside, same era.
Orthogonal range searching in moderate dimensions: kd trees and range trees strike back
Timothy M Chan · 2019
Cited alongside, same era.
Solving linear programs in the current matrix multiplication time
Michael B Cohen, Yin Tat Lee, and Zhao Song · 2019
Cited alongside, same era.
Sparse fourier transform over lattices: A unified approach to signal reconstruction
Zhao Song, Baocheng Sun, Omri Weinstein, and Ruizhe Zhang · 2022
Later among the works it cites.
A theoretical analysis on feature learning in neural networks: Emergence from inputs and advantage over fixed features
Zhenmei Shi, Junyi Wei, and Yingyu Liang · 2022
Later among the works it cites.
Speeding up sparsification using inner product search data structures
Zhao Song, Zhaozhuo Xu, and Lichen Zhang · 2022
Later among the works it cites.
Faster algorithm for structured john ellipsoid computation
Zhao Song, Xin Yang, Yuanyuan Yang, and Tianyi Zhou · 2022
Later among the works it cites.
Bypass exponential time preprocessing: Fast neural network training via weight-data correlation preprocessing
Josh Alman, Jiehao Liang, Zhao Song, Ruizhe Zhang, and Danyang Zhuo · 2023
Closest in time.
Fast attention requires bounded entries
Josh Alman and Zhao Song · 2023
Closest in time.
Attention scheme inspired softmax regression
Yichuan Deng, Zhihang Li, and Zhao Song · 2023
Closest in time.
Randomized and deterministic attention sparsification algorithms for over-parameterized feature dimension
Yichuan Deng, Sridhar Mahadevan, and Zhao Song · 2023
Closest in time.
An over-parameterized exponential regression
Yeqi Gao, Sridhar Mahadevan, and Zhao Song · 2023
Closest in time.
Solving regularized exp, cosh and sinh regression problems
Zhihang Li, Zhao Song, and Tianyi Zhou · 2023
Closest in time.
Application of kronecker convolutions in deep learning technique for automated detection of kidney stones with coronal ct images
Kiran Kumar Patro, Jaya Prakash Allam, Bala Chakravarthy Neelapu, Ryszard Tadeusiewicz, U Rajendra Acharya, Mohamed Hammad, Ozal Yildirim, and Paweł Pławiak · 2023
Closest in time.
An online and unified algorithm for projection matrix vector multiplication with application to empirical risk minimization
Lianke Qin, Zhao Song, Lichen Zhang, and Danyang Zhuo · 2023
Closest in time.
Quartic samples suffice for fourier interpolation
Zhao Song, Baocheng Sun, Omri Weinstein, and Ruizhe Zhang · 2023
Closest in time.
Sketching meets differential privacy: Fast algorithm for dynamic kronecker projection maintenance
Zhao Song, Xin Yang, Yuanyuan Yang, and Lichen Zhang · 2023
Closest in time.
Convergence and generalization of wide neural networks with large bias
Hongru Yang, Ziyu Jiang, Ruizhe Zhang, Zhangyang Wang, and Yingbin Liang · 2023
Closest in time.
Algorithm and hardness for dynamic attention maintenance in large language models
Jan van den Brand, Zhao Song, and Tianyi Zhou · 2024
Closest in time.
Deep kronecker network
Long Feng and Guang Yang · 2024
Closest in time.
Jiuxiang Gu, Chenyang Li, Yingyu Liang, Zhenmei Shi, and Zhao Song · 2024
Closest in time.
A sublinear adversarial training algorithm
Yeqi Gao, Lianke Qin, Zhao Song, and Yitan Wang · 2024
Closest in time.
Efficient sgd neural network training via sublinear activated neuron identification
Lianke Qin, Zhao Song, and Yuanyuan Yang · 2024
Closest in time.
Provable guarantees for neural networks via gradient feature learning
Zhenmei Shi, Junyi Wei, and Yingyu Liang · 2024
Closest in time.
Training multi-layer over-parametrized neural network in subquadratic time
Zhao Song, Lichen Zhang, and Ruizhe Zhang · 2024
Closest in time.
An iterative algorithm for rescaled hyperbolic functions regression
Yeqi Gao, Zhao Song, and Junze Yin · 2025
Closest in time.