Fetching the paper…
Reading the bibliography…
It has recently been demonstrated empirically that in-context learning emerges in transformers when certain distributional properties are present in the training data, but this ability can also diminish upon further training.
Introduction to online convex optimization, 2023
Elad Hazan · 1909
Earlier work this paper cites.
On the generalization ability of on-line learning algorithms
Nicolò Cesa-Bianchi, Alex Conconi, and Claudio Gentile · 2001
Earlier work this paper cites.
Introduction to statistical learning theory
Olivier Bousquet, Stéphane Boucheron, and Gábor Lugosi · 2003
Earlier work this paper cites.
Information theory and statistics: a tutorial
Imre Csiszár and Paul C. Shields · 2004
Earlier work this paper cites.
Prediction, Learning, and Games
Nicolò Cesa-Bianchi and Gábor Lugosi · 2006
Earlier work this paper cites.
Concentration inequalities: A nonasymptotic theory of independence
Stéphane Boucheron, Gábor Lugosi, and Pascal Massart · 2012
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
M. Zinkevich · 2012
Earlier work this paper cites.
Gated feedback recurrent neural networks
Junyoung Chung, Caglar Gulcehre, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Human-level concept learning through probabilistic program induction
Brenden M. Lake, Ruslan Salakhutdinov, and Joshua B. Tenenbaum · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Cited alongside, same era.
What can transformers learn in-context? a case study of simple function classes
Shivam Garg, Dimitris Tsipras, Percy Liang, and Gregory Valiant · 2022
Cited alongside, same era.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2022
Cited alongside, same era.
An explanation of in-context learning as implicit bayesian inference
Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma · 2022
Cited alongside, same era.
Rishabh Agarwal, Avi Singh, Lei M. Zhang, Bernd Bohnet, Luis Rosias, Stephanie Chan, Biao Zhang, Ankesh Anand, Zaheer Abbas, Azade Nova, John D. Co-Reyes, Eric Chu, Feryal Behbahani, Aleksandra Faust, and Hugo Larochelle · 2024
Closest in time.
Unveiling induction heads: Provable training dynamics and feature learning in transformers
Siyu Chen, Heejune Sheen, Tianhao Wang, and Zhuoran Yang · 2024
Closest in time.
The evolution of statistical induction heads: In-context learning markov chains, 2024
Benjamin L. Edelman, Ezra Edelman, Surbhi Goel, Eran Malach, and Nikolaos Tsilivis · 2024
Closest in time.
Gemini: A family of highly capable multimodal models
Gemini Team, Google · 2024
Closest in time.
Dual operating modes of in-context learning
Ziqian Lin and Kangwook Lee · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
What learning algorithm is in-context learning? investigations with linear models
Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou · 2023
Cited alongside, same era.
Transformers as statisticians: Provable in-context learning with in-context algorithm selection
Yu Bai, Fan Chen, Huan Wang, Caiming Xiong, and Song Mei · 2023
Cited alongside, same era.
Context is environment
Sharut Gupta, Stefanie Jegelka, David Lopez-Paz, and Kartik Ahuja · 2023
Cited alongside, same era.
In-context learning creates task vectors
Roee Hendel, Mor Geva, and Amir Globerson · 2023
Cited alongside, same era.
Transformers as algorithms: Generalization and stability in in-context learning
Yingcong Li, Muhammed Emrullah Ildiz, Dimitris Papailiopoulos, and Samet Oymak · 2023
Cited alongside, same era.
The transient nature of emergent in-context learning in transformers
Aaditya Singh, Stephanie Chan, Ted Moskovitz, Erin Grant, Andrew Saxe, and Felix Hill · 2023
Cited alongside, same era.
Transformers learn in-context by gradient descent
Johannes Von Oswald, Eyvind Niklasson, Ettore Randazzo, Joao Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov · 2023
Cited alongside, same era.
Larger language models do in-context learning differently, 2023
Jerry Wei, Jason Wei, Yi Tay, Dustin Tran, Albert Webson, Yifeng Lu, Xinyun Chen, Hanxiao Liu, Da Huang, Denny Zhou, and Tengyu Ma · 2023
Cited alongside, same era.
Closest in time.
How transformers learn causal structure with gradient descent
Eshaan Nichani, Alex Damian, and Jason D. Lee · 2024
Closest in time.
In-context learning through the bayesian prism
Madhur Panwar, Kabir Ahuja, and Navin Goyal · 2024
Closest in time.
The mechanistic basis of data dependence and abrupt learning in an in-context classification task
Gautam Reddy · 2024
Closest in time.
One-layer transformers fail to solve the induction heads task
Clayton Sanford, Daniel Hsu, and Matus Telgarsky · 2024
Closest in time.
Why larger language models do in-context learning differently?
Zhenmei Shi, Junyi Wei, Zhuoyan Xu, and Yingyu Liang · 2024
Closest in time.
What needs to go right for an induction head? a mechanistic study of in-context learning circuits and their formation
Aaditya K Singh, Ted Moskovitz, Felix Hill, Stephanie C.Y. Chan, and Andrew M Saxe · 2024
Closest in time.
How many pretraining tasks are needed for in-context learning of linear regression?
Jingfeng Wu, Difan Zou, Zixiang Chen, Vladimir Braverman, Quanquan Gu, and Peter Bartlett · 2024
Closest in time.
Trained transformers learn linear models in-context
Ruiqi Zhang, Spencer Frei, and Peter Bartlett · 2024
Closest in time.