Extensions of lipschitz mappings into a hilbert space
William B Johnson and Joram Lindenstrauss · 1984
Earlier work this paper cites.
The volumetric barrier for semidefinite programming
Kurt M Anstreicher · 2000
Earlier work this paper cites.
Approximate nearest neighbors and the fast johnson-lindenstrauss transform
Nir Ailon and Bernard Chazelle · 2006
Earlier work this paper cites.
Multiplying matrices faster than coppersmith-winograd
Virginia Vassilevska Williams · 2012
Earlier work this paper cites.
Osnap: Faster numerical linear algebra algorithms via sparser subspace embeddings
Jelani Nelson and Huy L Nguyên · 2013
Earlier work this paper cites.
Compressed matrix multiplication
Rasmus Pagh · 2013
Earlier work this paper cites.
Fast and scalable polynomial kernels via explicit feature maps
Ninh Pham and Rasmus Pagh · 2013
Earlier work this paper cites.
Subspace embeddings for the polynomial kernel
Haim Avron, Huy Nguyen, and David Woodruff · 2014
Earlier work this paper cites.
Powers of tensors and fast matrix multiplication
François Le Gall · 2014
Earlier work this paper cites.
Large-scale random features for kernel regression
Valero Laparra, Diego Marcos Gonzalez, Devis Tuia, and Gustau Camps-Valls · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Low-rank approximation and regression in input sparsity time
Kenneth L Clarkson and David P Woodruff · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Differentially private identity and equivalence testing of discrete distributions
Maryam Aliakbarpour, Ilias Diakonikolas, and Ronitt Rubinfeld · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Original
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Sketching for kronecker product regression and p-splines
Huaian Diao, Zhao Song, Wen Sun, and David Woodruff · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Private testing of distributions via sample permutations
Maryam Aliakbarpour, Ilias Diakonikolas, Daniel Kane, and Ronitt Rubinfeld · 2019
Earlier work this paper cites.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Earlier work this paper cites.
On the convergence rate of training recurrent neural networks
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Earlier work this paper cites.
Solving linear programs in the current matrix multiplication time
Michael B Cohen, Yin Tat Lee, and Zhao Song · 2019
Earlier work this paper cites.
Solving empirical risk minimization in the current matrix multiplication time
Yin Tat Lee, Zhao Song, and Qiuyi Zhang · 2019
Earlier work this paper cites.
Camembert: a tasty french language model
Original
Louis Martin, Benjamin Muller, Pedro Javier Ortiz Suarez, Yoann Dupont, Laurent Romary, Eric Villemonte de La Clergerie, Djame Seddah, and Benoit Sagot · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Matrix theory: optimization, concentration, and algorithms
Zhao Song · 2019
Earlier work this paper cites.
Towards a zero-one law for column subset selection
Zhao Song, David Woodruff, and Peilin Zhong · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
A deterministic linear program solver in current matrix multiplication time
Jan van den Brand · 2020
Earlier work this paper cites.
An efficient protocol for distributed column subset selection in the entrywise ℓ p \ell_{p} norm
Shuli Jiang, Dongyu Li, Irene Mengze Li, Arvind V Mahankali, and David Woodruff · 2020
Earlier work this paper cites.
Attention-based sentiment analysis using convolutional and recurrent neural network
Mohd Usama, Belal Ahmad, Enmin Song, M Shamim Hossain, Mubarak Alrashoud, and Ghulam Muhammad · 2020
Earlier work this paper cites.
Why are adaptive methods good for attention models?
Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim, Sashank Reddi, Sanjiv Kumar, and Suvrit Sra · 2020
Earlier work this paper cites.
Local differential privacy is equivalent to contraction of an f f -divergence
Shahab Asoodeh, Maryam Aliakbarpour, and Flavio P Calmon · 2021
Earlier work this paper cites.
A refined laser method and faster matrix multiplication
Josh Alman and Virginia Vassilevska Williams · 2021
Earlier work this paper cites.
Attention mechanism for neural machine translation: A survey
Weihua He, Yongyun Wu, and Xiaohua Li · 2021
Earlier work this paper cites.
Streaming and distributed algorithms for robust column subset selection
Shuli Jiang, Dennis Li, Irene Mengze Li, Arvind V Mahankali, and David Woodruff · 2021
Earlier work this paper cites.
Faster dynamic matrix inverse for faster lps
Shunhua Jiang, Zhao Song, Omri Weinstein, and Hengjie Zhang · 2021
Earlier work this paper cites.
Oblivious sketching-based central path method for linear programming
Zhao Song and Zheng Yu · 2021
Earlier work this paper cites.
Approximating how single head attention learns
Original
Charlie Snell, Ruiqi Zhong, Dan Klein, and Jacob Steinhardt · 2021
Earlier work this paper cites.
Training multi-layer over-parametrized neural network in subquadratic time
Original
Zhao Song, Lichen Zhang, and Ruizhe Zhang · 2021
Earlier work this paper cites.
Breaking the linear iteration cost barrier for some well-known conditional gradient methods using maxip data-structures
Zhaozhuo Xu, Zhao Song, and Anshumali Shrivastava · 2021
Earlier work this paper cites.
Off-tanet: A lightweight neural micro-expression recognizer with optical flow features and integrated attention mechanism
Jiahao Zhang, Feng Liu, and Aimin Zhou · 2021
Earlier work this paper cites.
Bypass exponential time preprocessing: Fast neural network training via weight-data correlation preprocessing
Original
Josh Alman, Jiehao Liang, Zhao Song, Ruizhe Zhang, and Danyang Zhuo · 2022
Earlier work this paper cites.
What learning algorithm is in-context learning? investigations with linear models
Original
Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou · 2022
Earlier work this paper cites.