Fetching the paper…
Reading the bibliography…
Characterizing the express power of the Transformer architecture is critical to understanding its capacity limits and scaling law.
The boolean formula value problem is in alogtime
Samuel R Buss · 1987
Earlier work this paper cites.
An optimal parallel algorithm for formula evaluation
S Buss, S Cook, Arvind Gupta, and Vijaya Ramachandran · 1992
Earlier work this paper cites.
Time, hardware, and uniformity
D Mix Barrington and Neil Immerman · 1994
Earlier work this paper cites.
Introduction to the theory of computation
Michael Sipser · 1996
Earlier work this paper cites.
Descriptive complexity
Neil Immerman · 1998
Earlier work this paper cites.
Introduction to circuit complexity: a uniform approach
Heribert Vollmer · 1999
Earlier work this paper cites.
On the complexity of k-sat
Russell Impagliazzo and Ramamohan Paturi · 2001
Earlier work this paper cites.
Uniform constant-depth threshold circuits for division and iterated multiplication
William Hesse, Eric Allender, and David A Mix Barrington · 2002
Earlier work this paper cites.
Computational complexity: a modern approach
Sanjeev Arora and Boaz Barak · 2009
Earlier work this paper cites.
Petar Velickovi, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Lio, and Yoshua Bengio · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
On some fine-grained questions in algorithms and complexity
Virginia Vassilevska Williams · 2018
Earlier work this paper cites.
Heterogeneous graph attention network
Xiao Wang, Houye Ji, Chuan Shi, Bai Wang, Yanfang Ye, Peng Cui, and Philip S Yu · 2019
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2020
Earlier work this paper cites.
How attentive are graph attention networks?
Shaked Brody, Uri Alon, and Eran Yahav · 2021
Earlier work this paper cites.
Gpt-neox-20b: An open-source autoregressive language model
Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, et al · 2022
Earlier work this paper cites.
What is my math transformer doing?–three results on interpretability and generalization
François Charton · 2022
Earlier work this paper cites.
A nearly optimal size coreset algorithm with nearly linear time
Yichuan Deng, Zhao Song, Yitan Wang, and Yuanyuan Yang · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa · 2022
Earlier work this paper cites.
Transformers learn shortcuts to automata
Bingbin Liu, Jordan T Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang · 2022
Earlier work this paper cites.
Saturated transformers are constant-depth threshold circuits
William Merrill, Ashish Sabharwal, and Noah A Smith · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou · 2022
Earlier work this paper cites.
Long time no see! open-domain conversation with long-term persona memory
Xinchao Xu, Zhibin Gou, Wenquan Wu, Zheng-Yu Niu, Hua Wu, Haifeng Wang, and Shihang Wang · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Earlier work this paper cites.
Fast attention requires bounded entries
Josh Alman and Zhao Song · 2023
Earlier work this paper cites.
Josh Alman and Zhao Song · 2023
Earlier work this paper cites.
Pythia: A suite for analyzing large language models across training and scaling
Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al · 2023
Earlier work this paper cites.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2023
Cited alongside, same era.
Towards revealing the mystery behind chain of thought: a theoretical perspective
Guhao Feng, Bohang Zhang, Yuntian Gu, Haotian Ye, Di He, and Liwei Wang · 2023
Cited alongside, same era.
An over-parameterized exponential regression
Yeqi Gao, Sridhar Mahadevan, and Zhao Song · 2023
Cited alongside, same era.
Retrieval-augmented generation for large language models: A survey
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, and Haofen Wang · 2023
Cited alongside, same era.
On sparse modern hopfield model
Jerry Yao-Chieh Hu, Donglin Yang, Dennis Wu, Chenwei Xu, Bo-Yu Chen, and Han Liu · 2023
Cited alongside, same era.
Wikibench: Community-driven data curation for ai evaluation on wikipedia
Tzu-Sheng Kuo, Aaron Lee Halfaker, Zirui Cheng, Jiwoo Kim, Meng-Hsin Wu, Tongshuang Wu, Kenneth Holstein, and Haiyi Zhu · 2024
Closest in time.
Na Liu, Liangyu Chen, Xiaoyu Tian, Wei Zou, Kaijiang Chen, and Ming Cui · 2024
Closest in time.
Chenyang Li, Yingyu Liang, Zhenmei Shi, Zhao Song, and Tianyi Zhou · 2024
Closest in time.
Fast john ellipsoid computation with differential privacy optimization
Xiaoyu Li, Yingyu Liang, Zhenmei Shi, Zhao Song, and Junwei Yu · 2024
Closest in time.
Fine-grained attention i/o complexity: Comprehensive analysis for backward passes
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A logic for expressing log-precision transformers
William Merrill and Ashish Sabharwal · 2023
Cited alongside, same era.
The parallelism tradeoff: Limitations of log-precision transformers
William Merrill and Ashish Sabharwal · 2023
Cited alongside, same era.
A theoretical analysis of nearest neighbor search on approximate near neighbor graph
Anshumali Shrivastava, Zhao Song, and Zhaozhuo Xu · 2023
Cited alongside, same era.
An automatic learning rate schedule algorithm for achieving faster convergence and steeper descent
Zhao Song and Chiwun Yang · 2023
Cited alongside, same era.
The rise and potential of large language model based agents: A survey
Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou, et al · 2023
Cited alongside, same era.
The claude 3 model family: Opus, sonnet, haiku, 2024
Anthropic · 2024
Cited alongside, same era.
Fast rope attention: Combining the polynomial method and fast fourier transform
Josh Alman and Zhao Song · 2024
Cited alongside, same era.
Xiaoyu Li, Yingyu Liang, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Closest in time.
Yingyu Liang, Heshan Liu, Zhenmei Shi, Zhao Song, Zhuoyan Xu, and Junze Yin · 2024
Closest in time.
Beyond linear approximations: A novel pruning approach for attention matrix
Yingyu Liang, Jiangxuan Long, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Closest in time.
A tighter complexity analysis of sparsegpt
Xiaoyu Li, Yingyu Liang, Zhenmei Shi, and Zhao Song · 2024
Closest in time.
Chain of thought empowers transformers to solve inherently serial problems
Zhiyuan Li, Hong Liu, Denny Zhou, and Tengyu Ma · 2024
Closest in time.
Looped relu mlps may be all you need as practical programmable computers
Yingyu Liang, Zhizhou Sha, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Closest in time.
Looped relu mlps may be all you need as practical programmable computers
Yingyu Liang, Zhizhou Sha, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Closest in time.
Multi-layer transformers gradient can be approximated in almost linear time
Yingyu Liang, Zhizhou Sha, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Closest in time.
Toward infinite-long prefix in transformer
Yingyu Liang, Zhenmei Shi, Zhao Song, and Chiwun Yang · 2024
Closest in time.
Differential privacy of cross-attention with provable guarantee
Yingyu Liang, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Closest in time.
Tensor attention training: Provably efficient learning of higher-order transformers
Yingyu Liang, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Closest in time.
AI @ Meta Llama Team · 2024
Closest in time.
Evaluating very long-term conversational memory of llm agents
Adyasha Maharana, Dong-Ho Lee, Sergey Tulyakov, Mohit Bansal, Francesco Barbieri, and Yuwei Fang · 2024
Closest in time.
Introducing openai o1-preview
OpenAI · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu · 2024
Closest in time.
Zhenmei Shi, Yifei Ming, Xuan-Phi Nguyen, Yingyu Liang, and Shafiq Joty · 2024
Closest in time.
Training multi-layer over-parametrized neural network in subquadratic time
Zhao Song, Lichen Zhang, and Ruizhe Zhang · 2024
Closest in time.
Uniform memory retrieval with larger capacity for modern hopfield models
Dennis Wu, Jerry Yao-Chieh Hu, Teng-Yun Hsiao, and Han Liu · 2024
Closest in time.
STanhop: Sparse tandem hopfield model for memory-enhanced time series prediction
Dennis Wu, Jerry Yao-Chieh Hu, Weijian Li, Bo-Yu Chen, and Han Liu · 2024
Closest in time.
Bishop: Bi-directional cellular learning for tabular data with generalized sparse modern hopfield model
Chenwei Xu, Yu-Chao Huang, Jerry Yao-Chieh Hu, Weijian Li, Ammar Gilani, Hsi-Sheng Goan, and Han Liu · 2024
Closest in time.
Do large language models have compositional ability? an investigation into limitations and scalability
Zhuoyan Xu, Zhenmei Shi, and Yingyu Liang · 2024
Closest in time.
Self-guide: Better task-specific instruction following via self-synthetic finetuning
Chenyang Zhao, Xueying Jia, Vijay Viswanathan, Tongshuang Wu, and Graham Neubig · 2024
Closest in time.