Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are powerful models that can learn concepts at the inference stage via in-context learning (ICL).
Segmentation-aware convolutional networks using local attention masks
Adam W Harley, Konstantinos G Derpanis, and Iasonas Kokkinos · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Mask-guided attention network for occluded pedestrian detection
Yanwei Pang, Jin Xie, Muhammad Haris Khan, Rao Muhammad Anwer, Fahad Shahbaz Khan, and Ling Shao · 2019
Earlier work this paper cites.
Novel positional encodings to enable tree-based transformers
Vighnesh Shiv and Chris Quirk · 2019
Earlier work this paper cites.
How does bert answer questions? a layer-wise analysis of transformer representations
Betty Van Aken, Benjamin Winter, Alexander Löser, and Felix A Gers · 2019
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret · 2020
Earlier work this paper cites.
Linear attention mechanism: An efficient attention for semantic segmentation
Rui Li, Jianlin Su, Chenxi Duan, and Shunyi Zheng · 2020
Earlier work this paper cites.
Polar relative positional encoding for video-language segmentation
Ke Ning, Lingxi Xie, Fei Wu, and Qi Tian · 2020
Earlier work this paper cites.
Bi-modal progressive mask attention for fine-grained recognition
Kaitao Song, Xiu-Shen Wei, Xiangbo Shu, Ren-Jie Song, and Jianfeng Lu · 2020
Earlier work this paper cites.
Mask attention networks: Rethinking and strengthen transformer
Zhihao Fan, Yeyun Gong, Dayiheng Liu, Zhongyu Wei, Siyuan Wang, Jian Jiao, Nan Duan, Ruofei Zhang, and Xuanjing Huang · 2021
Earlier work this paper cites.
Efficient attention: Attention with linear complexities
Zhuoran Shen, Mingyuan Zhang, Haiyu Zhao, Shuai Yi, and Hongsheng Li · 2021
Earlier work this paper cites.
How many layers and why? an analysis of the model depth in transformers
Antoine Simoulin and Benoit Crabbé · 2021
Cited alongside, same era.
What learning algorithm is in-context learning? investigations with linear models
Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou · 2022
Cited alongside, same era.
Rewiring with positional encodings for graph neural networks
Rickard Brüel-Gabrielsson, Mikhail Yurochkin, and Justin Solomon · 2022
Cited alongside, same era.
Masked-attention mask transformer for universal image segmentation
Bowen Cheng, Ishan Misra, Alexander G Schwing, Alexander Kirillov, and Rohit Girdhar · 2022
Cited alongside, same era.
What can transformers learn in-context? a case study of simple function classes
Shivam Garg, Dimitris Tsipras, Percy S Liang, and Gregory Valiant · 2022
Cited alongside, same era.
Dissecting chain-of-thought: A study on compositional in-context learning of mlps
Yingcong Li, Kartik Sreenivasan, Angeliki Giannou, Dimitris Papailiopoulos, and Samet Oymak · 2023
Later among the works it cites.
Linrec: Linear attention mechanism for long-term sequential recommender systems
Langming Liu, Liu Cai, Chi Zhang, Xiangyu Zhao, Jingtong Gao, Wanyu Wang, Yifu Lv, Wenqi Fan, Yiqi Wang, Ming He, et al · 2023
Later among the works it cites.
Transformers learn in-context by gradient descent
Johannes Von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov · 2023
Later among the works it cites.
On the role of unstructured training data in transformers’ in-context learning capabilities
Kevin Christian Wibisono and Yixin Wang · 2023
Later among the works it cites.
Trained transformers learn linear models in-context
Ruiqi Zhang, Spencer Frei, and Peter L Bartlett · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Generalized classification of satellite image time series with thermal positional encoding
Joachim Nyborg, Charlotte Pelletier, and Ira Assent · 2022
Cited alongside, same era.
cosformer: Rethinking softmax in attention
Zhen Qin, Weixuan Sun, Hui Deng, Dongxu Li, Yunshen Wei, Baohong Lv, Junjie Yan, Lingpeng Kong, and Yiran Zhong · 2022
Cited alongside, same era.
Linear complexity randomized self-attention mechanism
Lin Zheng, Chong Wang, and Lingpeng Kong · 2022
Cited alongside, same era.
Transformers as statisticians: Provable in-context learning with in-context algorithm selection
Yu Bai, Fan Chen, Huan Wang, Caiming Xiong, and Song Mei · 2023
Cited alongside, same era.
In-context learning of large language models explained as kernel regression
Chi Han, Ziqi Wang, Han Zhao, and Heng Ji · 2023
Cited alongside, same era.
In-context convergence of transformers
Yu Huang, Yuan Cheng, and Yingbin Liang · 2023
Cited alongside, same era.
Transformers learn to implement preconditioned gradient descent for in-context learning
Kwangjun Ahn, Xiang Cheng, Hadi Daneshmand, and Suvrit Sra
Cited in the paper.
Later among the works it cites.
Siyu Chen, Heejune Sheen, Tianhao Wang, and Zhuoran Yang · 2024
Closest in time.
Superiority of multi-head attention in in-context linear regression
Yingqian Cui, Jie Ren, Pengfei He, Jiliang Tang, and Yue Xing · 2024
Closest in time.
Juno Kim and Taiji Suzuki · 2024
Closest in time.
Chain of thought empowers transformers to solve inherently serial problems
Zhiyuan Li, Hong Liu, Denny Zhou, and Tengyu Ma · 2024
Closest in time.
How transformers learn causal structure with gradient descent
Eshaan Nichani, Alex Damian, and Jason D Lee · 2024
Closest in time.