Fetching the paper…
Reading the bibliography…
Besides natural language processing, transformers exhibit extraordinary performance in solving broader applications, including scientific computing and computer vision.
On the numerical properties of an iterative method for computing the moore–penrose generalized inverse
Torsten Söderström and GW Stewart · 1974
Earlier work this paper cites.
An improved newton iteration for the generalized inverse of a matrix, with applications
Victor Pan and Robert Schreiber · 1991
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting
Shiyang Li, Xiaoyong Jin, Yao Xuan, Xiyou Zhou, Wenhu Chen, Yu-Xiang Wang, and Xifeng Yan · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations
Maziar Raissi, Paris Perdikaris, and George E Karniadakis · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Earlier work this paper cites.
Choose a transformer: Fourier or galerkin
Shuhao Cao · 2021
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Earlier work this paper cites.
What can transformers learn in-context? a case study of simple function classes
Shivam Garg, Dimitris Tsipras, Percy S Liang, and Gregory Valiant · 2022
Earlier work this paper cites.
Making transformers solve compositional tasks
Santiago Ontanon, Joshua Ainslie, Zachary Fisher, and Vaclav Cvicek · 2022
Cited alongside, same era.
Train short, test long: Attention with linear biases enables input length extrapolation
Ofir Press, Noah Smith, and Mike Lewis · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Cited alongside, same era.
Transformers learn to implement preconditioned gradient descent for in-context learning
Kwangjun Ahn, Xiang Cheng, Hadi Daneshmand, and Suvrit Sra · 2023
Cited alongside, same era.
What learning algorithm is in-context learning? investigations with linear models
Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou · 2023
Cited alongside, same era.
In-context convergence of transformers
Yu Huang, Yuan Cheng, and Yingbin Liang · 2023
Later among the works it cites.
MathPrompter: Mathematical reasoning using large language models
Shima Imani, Liang Du, and Harsh Shrivastava · 2023
Later among the works it cites.
The impact of positional encoding on length generalization in transformers
Amirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Payel Das, and Siva Reddy · 2023
Later among the works it cites.
Dissecting chain-of-thought: Compositionality through in-context filtering and learning
Yingcong Li, Kartik Sreenivasan, Angeliki Giannou, Dimitris Papailiopoulos, and Samet Oymak · 2023
Later among the works it cites.
Transformers learn in-context by gradient descent
Johannes Von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yu Bai, Fan Chen, Huan Wang, Caiming Xiong, and Song Mei · 2023
Cited alongside, same era.
Towards revealing the mystery behind chain of thought: a theoretical perspective
Guhao Feng, Yuntian Gu, Bohang Zhang, Haotian Ye, Di He, and Liwei Wang · 2023
Cited alongside, same era.
Sparsegpt: Massive language models can be accurately pruned in one-shot
Elias Frantar and Dan Alistarh · 2023
Cited alongside, same era.
Deqing Fu, Tian-Qi Chen, Robin Jia, and Vatsal Sharan · 2023
Cited alongside, same era.
Looped transformers as programmable computers
Angeliki Giannou, Shashank Rajput, Jy-Yong Sohn, Kangwook Lee, Jason D. Lee, and Dimitris Papailiopoulos · 2023
Cited alongside, same era.
Deep neural networks for solving large linear systems arising from high-dimensional problems
Yiqi Gu and Michael K Ng · 2023
Cited alongside, same era.
Trained transformers learn linear models in-context
Ruiqi Zhang, Spencer Frei, and Peter L Bartlett
Cited in the paper.
Xiaodong Chen, Yuxuan Hu, and Jing Zhang · 2024
Closest in time.
How do transformers learn in-context beyond simple functions? a case study on learning with representations
Tianyu Guo, Wei Hu, Song Mei, Huan Wang, Caiming Xiong, Silvio Savarese, and Yu Bai · 2024
Closest in time.
One step of gradient descent is provably the optimal in-context learner with one layer of linear self-attention
Arvind Mahankali, Tatsunori B Hashimoto, and Tengyu Ma · 2024
Closest in time.
Sheared LLaMA: Accelerating language model pre-training via structured pruning
Mengzhou Xia, Tianyu Gao, Zhiyuan Zeng, and Danqi Chen · 2024
Closest in time.
Looped transformers are better at learning learning algorithms
Liu Yang, Kangwook Lee, Robert D Nowak, and Dimitris Papailiopoulos · 2024
Closest in time.
Metamath: Bootstrap your own mathematical questions for large language models
Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengying Liu, Yu Zhang, James T Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu · 2024
Closest in time.