Fetching the paper…
Reading the bibliography…
Transformers excel at in-context learning (ICL) -- learning from demonstrations without parameter updates -- but how they do so remains a mystery.
On the reciprocal of the general algebraic matrix
E.H Moore · 1920
Earlier work this paper cites.
Iterative berechung der reziproken matrix
Günther Schulz · 1933
Earlier work this paper cites.
An iterative method for computing the generalized inverse of an arbitrary matrix
Adi Ben-Israel · 1965
Earlier work this paper cites.
On the numerical properties of an iterative method for computing the moore- penrose generalized inverse
Torsten Soderstrom and G. W. Stewart · 1974
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
A.S. Nemirovski and D.B Yudin · 1983
Earlier work this paper cites.
On the limited memory bfgs method for large scale optimization
Dong C. Liu and Jorge Nocedal · 1989
Earlier work this paper cites.
An improved newton iteration for the generalized inverse of a matrix, with applications
Victor Y. Pan and Robert S. Schreiber · 1991
Earlier work this paper cites.
Long Short-Term Memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Numerical optimization
Jorge Nocedal and Stephen J Wright · 1999
Earlier work this paper cites.
Convex optimization
Stephen P Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
On lower and upper bounds in smooth and strongly convex optimization
Yossi Arjevani, Shai Shalev-Shwartz, and Ohad Shamir · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Memory-sample tradeoffs for linear regression with small error
Vatsal Sharan, Aaron Sidford, and Gregory Valiant · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Calibrate before use: Improving few-shot performance of language models
Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh · 2021
Earlier work this paper cites.
What learning algorithm is in-context learning? investigations with linear models
Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou · 2022
Earlier work this paper cites.
Efficientvit: Enhanced linear attention for high-resolution low-computation visual recognition
Han Cai, Chuang Gan, and Song Han · 2022
Earlier work this paper cites.
Palm: Scaling language modeling with pathways, 2022
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel · 2022
Cited alongside, same era.
What can transformers learn in-context? a case study of simple function classes
Shivam Garg, Dimitris Tsipras, Percy Liang, and Gregory Valiant · 2022
Cited alongside, same era.
Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
Mor Geva, Avi Caciularu, Kevin Wang, and Yoav Goldberg · 2022
Cited alongside, same era.
How much does attention actually attend? questioning the importance of attention in pretrained transformers, 2022
Data curation alone can stabilize in-context learning
Ting-Yun Chang and Robin Jia · 2023
Closest in time.
Dola: Decoding by contrasting layers improves factuality in large language models, 2023
Yung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim, James Glass, and Pengcheng He · 2023
Closest in time.
Towards automated circuit discovery for mechanistic interpretability, 2023
Arthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso · 2023
Closest in time.
Why can gpt learn in-context? language models secretly perform gradient descent as meta-optimizers
Damai Dai, Yutao Sun, Li Dong, Yaru Hao, Zhifang Sui, and Furu Wei · 2023
Closest in time.
Looped transformers as programmable computers
Angeliki Giannou, Shashank Rajput, Jy-Yong Sohn, Kangwook Lee, Jason D. Lee, and Dimitris Papailiopoulos · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Michael Hassid, Hao Peng, Daniel Rotem, Jungo Kasai, Ivan Montero, Noah A. Smith, and Roy Schwartz · 2022
Cited alongside, same era.
What makes good in-context examples for GPT-3?
Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, and Weizhu Chen · 2022
Cited alongside, same era.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp · 2022
Cited alongside, same era.
Noisy channel language model prompting for few-shot text classification
Sewon Min, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer · 2022
Cited alongside, same era.
MetaICL: Learning to learn in context
Sewon Min, Mike Lewis, Luke Zettlemoyer, and Hannaneh Hajishirzi · 2022
Cited alongside, same era.
Rethinking the role of demonstrations: What makes in-context learning work?
Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer · 2022
Cited alongside, same era.
Grokking: Generalization beyond overfitting on small algorithmic datasets, 2022
Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra · 2022
Cited alongside, same era.
Scaling language models: Methods, analysis & insights from training gopher, 2022
Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks, Maribeth Rauh, Po-Sen Huang, Amelia Glaese, Johannes Welbl, Sumanth Dathathri, Saffron Huang, Jonathan Uesato, John Mellor, Irina Higgins, Antonia Creswell, Nat McAleese, Amy Wu, Erich Elsen, Siddhant Jayakumar, Elena Buchatskaya, David Budden, Esme Sutherland, Karen Simonyan, Michela Paganini, Laurent Sifre, Lena Martens, Xiang Lorraine Li, Adhiguna Kuncoro, Aida Nematzadeh, Elena Gribovskaya, Domenic Donato, Angeliki Lazaridou, Arthur Mensch, Jean-Baptiste Lespiau, Maria Tsimpoukelli, Nikolai Grigorev, Doug Fritz, Thibault Sottiaux, Mantas Pajarskas, Toby Pohlen, Zhitao Gong, Daniel Toyama, Cyprien de Masson d’Autume, Yujia Li, Tayfun Terzi, Vladimir Mikulik, Igor Babuschkin, Aidan Clark, Diego de Las Casas, Aurelia Guy, Chris Jones, James Bradbury, Matthew Johnson, Blake Hechtman, Laura Weidinger, Iason Gabriel, William Isaac, Ed Lockhart, Simon Osindero, Laura Rimell, Chris Dyer, Oriol Vinyals, Kareem Ayoub, Jeff Stanway, Lorrayne Bennett, Demis Hassabis, Koray Kavukcuoglu, and Geoffrey Irving · 2022
Cited alongside, same era.
Learning to retrieve prompts for in-context learning
Ohad Rubin, Jonathan Herzig, and Jonathan Berant · 2022
Cited alongside, same era.
In-context learning of large language models explained as kernel regression, 2023
Chi Han, Ziqi Wang, Han Zhao, and Heng Ji · 2023
Closest in time.
Exploring the relationship between model architecture and in-context learning ability, 2023
Ivan Lee, Nan Jiang, and Taylor Berg-Kirkpatrick · 2023
Closest in time.
Transformers as algorithms: Generalization and stability in in-context learning
Yingcong Li, Muhammed Emrullah Ildiz, Dimitris Papailiopoulos, and Samet Oymak · 2023
Closest in time.
Progress measures for grokking via mechanistic interpretability, 2023
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt · 2023
Closest in time.
In-context example selection with influences
Tai Nguyen and Eric Wong · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Pretraining task diversity and the emergence of non-bayesian in-context learning for regression, 2023
Allan Raventós, Mansheej Paul, Feng Chen, and Surya Ganguli · 2023
Closest in time.
Selective annotation makes language models better few-shot learners
Hongjin Su, Jungo Kasai, Chen Henry Wu, Weijia Shi, Tianlu Wang, Jiayi Xin, Rui Zhang, Mari Ostendorf, Luke Zettlemoyer, Noah A. Smith, and Tao Yu · 2023
Closest in time.
Uncovering mesa-optimization algorithms in transformers
Johannes von Oswald, Eyvind Niklasson, Maximilian Schlegel, Seijin Kobayashi, Nicolas Zucchet, Nino Scherrer, Nolan Miller, Mark Sandler, Blaise Agüera y Arcas, Max Vladymyrov, Razvan Pascanu, and Joao Sacramento · 2023
Closest in time.
Larger language models do in-context learning differently, 2023
Jerry Wei, Jason Wei, Yi Tay, Dustin Tran, Albert Webson, Yifeng Lu, Xinyun Chen, Hanxiao Liu, Da Huang, Denny Zhou, and Tengyu Ma · 2023
Closest in time.
Replacing softmax with relu in vision transformers, 2023
Mitchell Wortsman, Jaehoon Lee, Justin Gilmer, and Simon Kornblith · 2023
Closest in time.
Trained transformers learn linear models in-context
Ruiqi Zhang, Spencer Frei, and Peter L. Bartlett · 2023
Closest in time.
One step of gradient descent is provably the optimal in-context learner with one layer of linear self-attention
Arvind V. Mahankali, Tatsunori Hashimoto, and Tengyu Ma · 2024
Closest in time.
Linear transformers are versatile in-context learners, 2024
Max Vladymyrov, Johannes von Oswald, Mark Sandler, and Rong Ge · 2024
Closest in time.