Fetching the paper…
Reading the bibliography…
Large language models have become ubiquitous in modern life, finding applications in various domains such as natural language processing, language translation, and speech recognition.
A model and an hypothesis for language structure
Victor H Yngve · 1960
Earlier work this paper cites.
Trainable grammars for speech recognition
James K Baker · 1979
Earlier work this paper cites.
A stochastic parts program and noun phrase parser for unrestricted text
Kenneth Ward Church · 1989
Earlier work this paper cites.
On the computational complexity and geometry of the first-order theory of the reals. part i: Introduction. preliminaries. the geometry of semi-algebraic sets. the decision problem for the existential theory of the reals
James Renegar · 1992
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Mitchell Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz · 1993
Earlier work this paper cites.
On the combinatorial and algebraic complexity of quantifier elimination
Saugata Basu, Richard Pollack, and Marie-Françoise Roy · 1996
Earlier work this paper cites.
Tree-bank grammars
Eugene Charniak · 1996
Earlier work this paper cites.
Parsing algorithms and metrics
Joshua Goodman · 1996
Earlier work this paper cites.
A maximum entropy model for part-of-speech tagging
Adwait Ratnaparkhi · 1996
Earlier work this paper cites.
Generation that exploits corpus-based statistical knowledge
Irene Langkilde and Kevin Knight · 1998
Earlier work this paper cites.
Foundations of statistical natural language processing
Christopher Manning and Hinrich Schutze · 1999
Earlier work this paper cites.
Accurate unlexicalized parsing
Dan Klein and Christopher D Manning · 2003
Earlier work this paper cites.
Feature-rich part-of-speech tagging with a cyclic dependency network
Kristina Toutanova, Dan Klein, Christopher D Manning, and Yoram Singer · 2003
Earlier work this paper cites.
Improved approximation algorithms for large matrices via random projections
Tamas Sarlos · 2006
Earlier work this paper cites.
Improved inference for unlexicalized parsing
Slav Petrov and Dan Klein · 2007
Earlier work this paper cites.
Low-rank approximation and regression in input sparsity time
Kenneth L Clarkson and David P Woodruff · 2013
Earlier work this paper cites.
Osnap: Faster numerical linear algebra algorithms via sparser subspace embeddings
Jelani Nelson and Huy L Nguyên · 2013
Earlier work this paper cites.
Compressed matrix multiplication
Rasmus Pagh · 2013
Earlier work this paper cites.
Fast and scalable polynomial kernels via explicit feature maps
Ninh Pham and Rasmus Pagh · 2013
Earlier work this paper cites.
Subspace embeddings for the polynomial kernel
Haim Avron, Huy Nguyen, and David Woodruff · 2014
Earlier work this paper cites.
Optimal cur matrix decompositions
Christos Boutsidis and David P Woodruff · 2014
Earlier work this paper cites.
Sketching as a tool for numerical linear algebra
David P Woodruff · 2014
Earlier work this paper cites.
Optimal principal component analysis in distributed and streaming models
Christos Boutsidis, David P Woodruff, and Peilin Zhong · 2016
Earlier work this paper cites.
Weighted low rank approximations with provable guarantees
Ilya Razenshteyn, Zhao Song, and David P Woodruff · 2016
Earlier work this paper cites.
Distributed low rank approximation of implicit functions of a matrix
David P Woodruff and Peilin Zhong · 2016
Cited alongside, same era.
Low rank approximation with entrywise l1-norm error
Zhao Song, David P Woodruff, and Peilin Zhong · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Sketching for kronecker product regression and p-splines
Huaian Diao, Zhao Song, Wen Sun, and David Woodruff · 2018
Cited alongside, same era.
Fast sketching of polynomial kernels of polynomial degree
Zhao Song, David Woodruff, Zheng Yu, and Lichen Zhang · 2021
Later among the works it cites.
Approximating how single head attention learns
Charlie Snell, Ruiqi Zhong, Dan Klein, and Jacob Steinhardt · 2021
Later among the works it cites.
Training multi-layer over-parametrized neural network in subquadratic time
Zhao Song, Lichen Zhang, and Ruizhe Zhang · 2021
Later among the works it cites.
What learning algorithm is in-context learning? investigations with linear models
Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jill Burstein, Christy Doran, and Thamar Solorio · 2019
Cited alongside, same era.
Dynamic matrix inverse: Improved algorithms and matching conditional lower bounds
Jan van den Brand, Danupon Nanongkai, and Thatchaphol Saranurak · 2019
Cited alongside, same era.
A near-optimal algorithm for approximating the john ellipsoid
Michael B Cohen, Ben Cousins, Yin Tat Lee, and Xin Yang · 2019
Cited alongside, same era.
Solving linear programs in the current matrix multiplication time
Michael B Cohen, Yin Tat Lee, and Zhao Song · 2019
Cited alongside, same era.
Optimal sketching for kronecker product regression and low rank approximation
Huaian Diao, Rajesh Jayaram, Zhao Song, Wen Sun, and David Woodruff · 2019
Cited alongside, same era.
Transfer learning in natural language processing
Sebastian Ruder, Matthew E Peters, Swabha Swayamdipta, and Thomas Wolf · 2019
Cited alongside, same era.
Average case column subset selection for entrywise ℓ 1 \ell_{1} -norm loss
Zhao Song, David Woodruff, and Peilin Zhong · 2019
Cited alongside, same era.
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Later among the works it cites.
Discrepancy minimization in input-sparsity time
Yichuan Deng, Zhao Song, and Omri Weinstein · 2022
Later among the works it cites.
What can transformers learn in-context? a case study of simple function classes
Shivam Garg, Dimitris Tsipras, Percy Liang, and Gregory Valiant · 2022
Later among the works it cites.
The Oxford handbook of computational linguistics
Ruslan Mitkov · 2022
Later among the works it cites.
Transformers learn in-context by gradient descent
Johannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov · 2022
Later among the works it cites.
Dynamic tensor product regression
Aravind Reddy, Zhao Song, and Lichen Zhang · 2022
Later among the works it cites.
Faster algorithm for structured john ellipsoid computation
Zhao Song, Xin Yang, Yuanyuan Yang, and Tianyi Zhou · 2022
Later among the works it cites.
Speeding up optimizations via data structures: Faster search, sample and maintenance
Lichen Zhang · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al · 2022
Later among the works it cites.
Fast attention requires bounded entries
Josh Alman and Zhao Song · 2023
Closest in time.
Algorithms and hardness for dynamic attention maintenance in large language models
Jan van den Brand, Zhao Song, and Tianyi Zhou · 2023
Closest in time.
Randomized and deterministic attention sparsification algorithms for over-parameterized feature dimension
Yichuan Deng, Sridhar Mahadevan, and Zhao Song · 2023
Closest in time.
An over-parameterized exponential regression
Yeqi Gao, Sridhar Mahadevan, and Zhao Song · 2023
Closest in time.
Convex minimization with integer minima in O ~ ( n 4 ) \widetilde{O}(n^{4}) time
Haotian Jiang, Yin Tat Lee, Zhao Song, and Lichen Zhang · 2023
Closest in time.
A watermark for large language models
John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein · 2023
Closest in time.
How do transformers learn topic structure: Towards a mechanistic understanding
Yuchen Li, Yuanzhi Li, and Andrej Risteski · 2023
Closest in time.
Solving regularized exp, cosh and sinh regression problems
Zhihang Li, Zhao Song, and Tianyi Zhou · 2023
Closest in time.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig · 2023
Closest in time.
Kdeformer: Accelerating transformers via kernel density estimation
Amir Zandieh, Insu Han, Majid Daliri, and Amin Karbasi · 2023
Closest in time.
Do transformers parse while predicting the masked word?
Haoyu Zhao, Abhishek Panigrahi, Rong Ge, and Sanjeev Arora · 2023
Closest in time.