Fetching the paper…
Reading the bibliography…
Attention-based transformers have been remarkably successful at modeling generative processes across various domains and modalities.
Some complexity questions related to distributive computing (preliminary report)
Andrew Chi-Chih Yao · 1979
Earlier work this paper cites.
Markov Chains
J. R. Norris · 1997
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
On the ability and limitations of transformers to recognize formal languages, 2020
Satwik Bhattamishra, Kabir Ahuja, and Navin Goyal · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Are transformers universal approximators of sequence-to-sequence functions?
Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank Reddi, and Sanjiv Kumar · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale, 2021
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Earlier work this paper cites.
Masked autoencoders are scalable vision learners, 2021
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2021
Earlier work this paper cites.
Attention is Turing-complete
Jorge Pérez, Pablo Barceló, and Javier Marinkovic · 2021
Earlier work this paper cites.
Thinking like transformers
Gail Weiss, Yoav Goldberg, and Eran Yahav · 2021
Earlier work this paper cites.
In-context learning and induction heads, 2022
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2022
Cited alongside, same era.
What learning algorithm is in-context learning? investigations with linear models, 2023
Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou · 2023
Cited alongside, same era.
Transformers as statisticians: Provable in-context learning with in-context algorithm selection, 2023
Yu Bai, Fan Chen, Huan Wang, Caiming Xiong, and Song Mei · 2023
Cited alongside, same era.
Birth of a transformer: A memory viewpoint
Alberto Bietti, Vivien Cabannes, Diane Bouchacourt, Herve Jegou, and Leon Bottou · 2023
Cited alongside, same era.
Looped transformers as programmable computers
Angeliki Giannou, Shashank Rajput, Jy-Yong Sohn, Kangwook Lee, Jason D. Lee, and Dimitris Papailiopoulos · 2023
Cited alongside, same era.
GPT-2 modular codebase implementation
Matteo Pagliardini · 2023
Later among the works it cites.
Representational strengths and limitations of transformers
Clayton Sanford, Daniel J Hsu, and Matus Telgarsky · 2023
Later among the works it cites.
The evolution of statistical induction heads: In-context learning markov chains, 2024
Benjamin L. Edelman, Ezra Edelman, Surbhi Goel, Eran Malach, and Nikolaos Tsilivis · 2024
Closest in time.
The developmental landscape of in-context learning
Jesse Hoogland, George Wang, Matthew Farrugia-Roberts, Liam Carroll, Susan Wei, and Daniel Murfet · 2024
Closest in time.
Dual operating modes of in-context learning, 2024
Ziqian Lin and Kangwook Lee · 2024
Closest in time.
Attention with markov: A framework for principled analysis of transformers via markov chains
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Localizing model behavior with path patching
Nicholas Goldowsky-Dill, Chris MacLeod, Lucas Sato, and Aryaman Arora · 2023
Cited alongside, same era.
A theoretical understanding of shallow vision transformers: Learning, generalization, and sample complexity
Hongkang Li, Meng Wang, Sijia Liu, and Pin-Yu Chen · 2023
Cited alongside, same era.
Transformers learn shortcuts to automata, 2023
Bingbin Liu, Jordan T. Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang · 2023
Cited alongside, same era.
On the role of attention in prompt-tuning
Samet Oymak, Ankit Singh Rawat, Mahdi Soltanolkotabi, and Christos Thrampoulidis · 2023
Cited alongside, same era.
Statistically meaningful approximation: a case study on approximating Turing machines with transformers
Colin Wei, Yining Chen, and Tengyu Ma
Cited in the paper.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al
Cited in the paper.
Ashok Vardhan Makkuva, Marco Bondaschi, Adway Girish, Alliot Nagle, Martin Jaggi, Hyeji Kim, and Michael Gastpar · 2024
Closest in time.
How transformers learn causal structure with gradient descent, 2024
Eshaan Nichani, Alex Damian, and Jason D. Lee · 2024
Closest in time.
Toward a theory of tokenization in llms, 2024
Nived Rajaraman, Jiantao Jiao, and Kannan Ramchandran · 2024
Closest in time.
Transformers, parallel computation, and logarithmic depth, 2024
Clayton Sanford, Daniel Hsu, and Matus Telgarsky · 2024
Closest in time.