Fetching the paper…
Reading the bibliography…
The Transformer architecture is widely deployed in many popular and impactful Large Language Models.
Distributional structure
Zellig S. Harris · 1954
Earlier work this paper cites.
A statistical interpretation of term specificity and its application in retrieval , pp. 132–142
Karen Sparck Jones · 1988
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Kurt Hornik, Maxwell Stinchcombe, and Halbert White · 1989
Earlier work this paper cites.
A Comparison of the Computational Power of Sigmoid and Boolean Threshold Circuits , pp. 127–151
W. Maass, G. Schnitger, and E. D. Sontag · 1994
Earlier work this paper cites.
Approximate nearest neighbors: towards removing the curse of dimensionality
Piotr Indyk and Rajeev Motwani · 1998
Earlier work this paper cites.
On the complexity of k-sat
Russell Impagliazzo and Ramamohan Paturi · 2000
Earlier work this paper cites.
Chapter 26 - continuous nearest neighbor search
Yufei Tao, Dimitris Papadias, and Qiongmao Shen · 2002
Earlier work this paper cites.
Latent dirichlet allocation
David M. Blei, Andrew Y. Ng, and Michael I. Jordan · 2003
Earlier work this paper cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E. Peters, and Arman Cohan · 2004
Earlier work this paper cites.
A new algorithm for optimal 2-constraint satisfaction and its implications
Ryan Williams · 2004
Earlier work this paper cites.
Nearest-Neighbor Methods in Learning and Vision: Theory and Practice (Neural Information Processing)
Gregory Shakhnarovich, Trevor Darrell, and Piotr Indyk · 2006
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Z. Li, Madian Khabsa, Han Fang, and Hao Ma · 2006
Earlier work this paper cites.
Near-optimal hashing algorithms for approximate nearest neighbor in high dimensions
Alexandr Andoni and Piotr Indyk · 2008
Earlier work this paper cites.
Contextual document similarity for content-based literature recommender systems
Malte Ostendorff · 2008
Earlier work this paper cites.
Popular conjectures imply strong lower bounds for dynamic problems
Amir Abboud and Virginia Vassilevska Williams · 2014
Earlier work this paper cites.
Distributed representations of sentences and documents
Quoc Le and Tomas Mikolov · 2014
Earlier work this paper cites.
On the number of linear regions of deep neural networks
Guido Montúfar, Razvan Pascanu, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Tight hardness results for lcs and other sequence similarity measures
Amir Abboud, Arturs Backurs, and Virginia Vassilevska Williams · 2015
Earlier work this paper cites.
Probabilistic polynomials and hamming nearest neighbors
Josh Alman and Ryan Williams · 2015
Earlier work this paper cites.
Optimal data-dependent hashing for approximate near neighbors
Alexandr Andoni and Ilya Razenshteyn · 2015
Earlier work this paper cites.
Edit distance cannot be computed in strongly subquadratic time (unless seth is false)
Arturs Backurs and Piotr Indyk · 2015
Earlier work this paper cites.
Hardness of easy problems: Basing hardness on popular conjectures such as the strong exponential time hypothesis (invited talk)
Virginia Vassilevska Williams · 2015
Earlier work this paper cites.
Bag-of-embeddings for text classification
Peng Jin, Yue Zhang, Xingyuan Chen, and Yunqing Xia · 2016
Earlier work this paper cites.
Computer Science Theory Stack Exchange, 2017
Pairwise comparison of bit vectors · 2017
Earlier work this paper cites.
Optimal hashing-based time-space trade-offs for approximate near neighbors
Alexandr Andoni, Thijs Laarhoven, Ilya Razenshteyn, and Erik Waingarten · 2017
Earlier work this paper cites.
Plagiarism detection using document similarity based on distributed representation
Kensuke Baba, Tetsuya Nakatoh, and Toshiro Minami · 2017
Earlier work this paper cites.
On the fine-grained complexity of empirical risk minimization: Kernel methods and neural networks
Arturs Backurs, Piotr Indyk, and Ludwig Schmidt · 2017
Cited alongside, same era.
Spectrally-normalized margin bounds for neural networks
Peter L Bartlett, Dylan J Foster, and Matus J Telgarsky · 2017
Cited alongside, same era.
On the fine-grained complexity of one-dimensional dynamic programming
Marvin Künnemann, Ramamohan Paturi, and Stefan Schneider · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
More consequences of falsifying seth and the orthogonal vectors conjecture
Amir Abboud, Karl Bringmann, Holger Dell, and Jesper Nederlof · 2018
Cited alongside, same era.
Neural tangent kernel: convergence and generalization in neural networks
Rethinking attention with performers
Krzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamás Sarlós, Peter Hawkins, Jared Quincy Davis, Afroz Mohiuddin, Lukasz Kaiser, David Benjamin Belanger, Lucy J. Colwell, and Adrian Weller · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Later among the works it cites.
Bag-of-words technique in natural language processing: A primer for radiologists
Krishna Juluru, Hao-Hsin Shih, Krishna Nand Keshava Murthy, and Pierre Elnajjar · 2021
Later among the works it cites.
Efficient content-based sparse attention with routing transformers
Aurko Roy, Mohammad Saffar, Ashish Vaswani, and David Grangier · 2021
Later among the works it cites.
Synthesizer: Rethinking self-attention for transformer models
Yi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan, Zhe Zhao, and Che Zheng · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Similarity search techniques in exploratory search: A review
Mohammed Najah Mahdi, Abdul Rahim Ahmad, and Roslan Ismail · 2018
Cited alongside, same era.
Hardness of approximate nearest neighbor search
Aviad Rubinstein · 2018
Cited alongside, same era.
On Some Fine-Grained Questions in Algorithms and Complexity , pp. 3447–3487
Virginia Vassilevska Williams · 2018
Cited alongside, same era.
Subcubic equivalences between path, matrix, and triangle problems
Virginia Vassilevska Williams and R. Ryan Williams · 2018
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Self-attention networks can process bounded hierarchical languages
Shunyu Yao, Binghui Peng, Christos H. Papadimitriou, and Karthik Narasimhan · 2021
Later among the works it cites.
Overcoming a theoretical limitation of self-attention
David Chiang and Peter Cholak · 2022
Later among the works it cites.
Formal language recognition by hard attention transformers: Perspectives from circuit complexity
Yiding Hao, Dana Angluin, and Robert Frank · 2022
Later among the works it cites.
Saturated transformers are constant-depth threshold circuits
William Merrill, Ashish Sabharwal, and Noah A. Smith · 2022
Later among the works it cites.
Formal algorithms for transformers, 2022
Mary Phuong and Marcus Hutter · 2022
Later among the works it cites.
Comparative performance analysis of k-nearest neighbour (knn) algorithm and its different variants for disease prediction
Shahadat Uddin, Ibtisham Haque, Haohui Lu, Mohammad Ali Moni, and Ergun Gide · 2022
Later among the works it cites.
On the computational complexity of self-attention
Feyza Duman Keles, Pruthuvi Mahesakya Wijewardena, and Chinmay Hegde · 2023
Later among the works it cites.
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao · 2023
Later among the works it cites.
Transformers learn shortcuts to automata
Bingbin Liu, Jordan T. Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Kdeformer: accelerating transformers via kernel density estimation
Amir Zandieh, Insu Han, Majid Daliri, and Amin Karbasi · 2023
Later among the works it cites.
Finer-grained hardness of kernel density estimation
Josh Alman and Yunfeng Guan · 2024
Closest in time.
Fast attention requires bounded entries
Josh Alman and Zhao Song · 2024
Closest in time.
Tensor ranks and the fine-grained complexity of dynamic programming
Josh Alman, Ethan Turok, Hantao Yu, and Hengzhi Zhang · 2024
Closest in time.
Practical near neighbor search via group testing
Joshua Engels, Benjamin Coleman, and Anshumali Shrivastava · 2024
Closest in time.
Hyperattention: Long-context attention in near-linear time
Insu Han, Rajesh Jayaram, Amin Karbasi, Vahab Mirrokni, David P. Woodruff, and Amir Zandieh · 2024
Closest in time.
On statistical rates and provably efficient criteria of latent diffusion transformers (dits)
Jerry Yao-Chieh Hu, Weimin Wu, Zhuoru Li, Sophia Pi, Zhao Song, and Han Liu · 2024
Closest in time.
The expressive power of transformers with chain of thought
William Merrill and Ashish Sabharwal · 2024
Closest in time.
What Formal Languages Can Transformers Express? A Survey
Lena Strobl, William Merrill, Gail Weiss, David Chiang, and Dana Angluin · 2024
Closest in time.
Statistically meaningful approximation: a case study on approximating turing machines with transformers
Colin Wei, Yining Chen, and Tengyu Ma · 2024
Closest in time.