Fetching the paper…
Reading the bibliography…
Transformers have demonstrated effectiveness in in-context solving data-fitting problems from various (latent) models, as reported by Garg et al.
Regression shrinkage and selection via the lasso
Robert Tibshirani · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Janvin · 2003
Earlier work this paper cites.
A tutorial on training recurrent neural networks , covering bppt , rtrl , ekf and the ” echo state network ” approach - semantic scholar
Herbert Jaeger · 2005
Earlier work this paper cites.
Training recurrent neural networks
Geoffrey E. Hinton and Ilya Sutskever · 2013
Earlier work this paper cites.
Openml: networked science in machine learning
Joaquin Vanschoren, Jan N. van Rijn, Bernd Bischl, and Luis Torgo · 2013
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Guillaume Alain and Yoshua Bengio · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam M. Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Trellis networks for sequence modeling
Shaojie Bai, J. Zico Kolter, and Vladlen Koltun · 2018
Earlier work this paper cites.
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser · 2018
Earlier work this paper cites.
Shaojie Bai, J. Zico Kolter, and Vladlen Koltun · 2019
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G. Carbonell, Quoc V. Le, and Ruslan Salakhutdinov · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
R-transformer: Recurrent neural network enhanced transformer
Z. Wang, Yao Ma, Zitao Liu, and Jiliang Tang · 2019
Earlier work this paper cites.
Multiscale deep equilibrium models
Shaojie Bai, Vladlen Koltun, and J. Zico Kolter · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, T. J. Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeff Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Deep learning techniques for inverse problems in imaging
Greg Ongie, Ajil Jalal, Christopher A. Metzler, Richard Baraniuk, Alexandros G. Dimakis, and Rebecca M. Willett · 2020
Earlier work this paper cites.
Monotone operator equilibrium networks
Ezra Winston and J. Zico Kolter · 2020
Cited alongside, same era.
Stabilizing equilibrium models by jacobian regularization
Shaojie Bai, Vladlen Koltun, and J. Zico Kolter · 2021
Cited alongside, same era.
Linear algebra with transformers
Franccois Charton · 2021
Cited alongside, same era.
Perceiver: General perception with iterative attention
Andrew Jaegle, Felix Gimeno, Andrew Brock, Andrew Zisserman, Oriol Vinyals, and João Carreira · 2021
Cited alongside, same era.
Transformers can do bayesian inference
Samuel Muller, Noah Hollmann, Sebastian Pineda Arango, Josif Grabocka, and Frank Hutter · 2021
Cited alongside, same era.
An explanation of in-context learning as implicit bayesian inference
Rethinking the role of demonstrations: What makes in-context learning work?
Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer · 2022
Later among the works it cites.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, T. J. Henighan, Benjamin Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, John Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom B. Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Christopher Olah · 2022
Later among the works it cites.
Transformers learn in-context by gradient descent
Johannes von Oswald, Eyvind Niklasson, E. Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Huai hsin Chi, F. Xia, Quoc Le, and Denny Zhou · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma · 2021
Cited alongside, same era.
What learning algorithm is in-context learning? investigations with linear models
Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou · 2022
Cited alongside, same era.
Neural deep equilibrium solvers
Shaojie Bai, Vladlen Koltun, and J. Zico Kolter · 2022
Cited alongside, same era.
End-to-end algorithm synthesis with recurrent networks: Logical extrapolation without overthinking
Arpit Bansal, Avi Schwarzschild, Eitan Borgnia, Zeyad Ali Sami Emam, Furong Huang, Micah Goldblum, and Tom Goldstein · 2022
Cited alongside, same era.
Simplicity bias in transformers and their ability to learn sparse boolean functions
S. Bhattamishra, Arkil Patel, Varun Kanade, and Phil Blunsom · 2022
Cited alongside, same era.
What is my math transformer doing? - three results on interpretability and generalization
Franccois Charton · 2022
Cited alongside, same era.
Hardness of noise-free learning for two-hidden-layer neural networks
Sitan Chen, Aravind Gollakota, Adam R. Klivans, and Raghu Meka · 2022
Cited alongside, same era.
Unveiling transformers with lego: a synthetic reasoning task
Yi Zhang, Arturs Backurs, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, and Tal Wagner · 2022
Later among the works it cites.
Teaching algorithmic reasoning via in-context learning
Hattie Zhou, Azade Nova, H. Larochelle, Aaron C. Courville, Behnam Neyshabur, and Hanie Sedghi · 2022
Later among the works it cites.
Transformers learn to implement preconditioned gradient descent for in-context learning
Kwangjun Ahn, Xiang Cheng, Hadi Daneshmand, and Suvrit Sra · 2023
Closest in time.
Transformers as statisticians: Provable in-context learning with in-context algorithm selection
Yu Bai, Fan Chen, Haiquan Wang, Caiming Xiong, and Song Mei · 2023
Closest in time.
Causallm is not optimal for in-context learning
Nan Ding, Tomer Levinboim, Jialin Wu, Sebastian Goodman, and Radu Soricut · 2023
Closest in time.
Faith and fate: Limits of transformers on compositionality
Nouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li, Liwei Jian, Bill Yuchen Lin, Peter West, Chandra Bhagavatula, Ronan Le Bras, Jena D. Hwang, Soumya Sanyal, Sean Welleck, Xiang Ren, Allyson Ettinger, Zaïd Harchaoui, and Yejin Choi · 2023
Closest in time.
Deqing Fu, Tian-Qi Chen, Robin Jia, and Vatsal Sharan · 2023
Closest in time.
Looped transformers as programmable computers
Angeliki Giannou, Shashank Rajput, Jy yong Sohn, Kangwook Lee, Jason D. Lee, and Dimitris Papailiopoulos · 2023
Closest in time.
Teaching arithmetic to small transformers
Nayoung Lee, Kartik Sreenivasan, Jason D Lee, Kangwook Lee, and Dimitris Papailiopoulos · 2023
Closest in time.
Transformers as algorithms: Generalization and stability in in-context learning
Yingcong Li, Muhammed Emrullah Ildiz, Dimitris Papailiopoulos, and Samet Oymak · 2023
Closest in time.
Arvind V. Mahankali, Tatsunori Hashimoto, and Tengyu Ma · 2023
Closest in time.
The effects of pretraining task diversity on in-context learning of ridge regression
Allan Raventos, Mansheej Paul, Feng Chen, and Surya Ganguli · 2023
Closest in time.
How well can transformers emulate in-context newton’s method?
Angeliki Giannou, Liu Yang, Tianhao Wang, Dimitris Papailiopoulos, and Jason D Lee · 2024
Closest in time.
Dual operating modes of in-context learning
Ziqian Lin and Kangwook Lee · 2024
Closest in time.
Ruiqi Zhang, Jingfeng Wu, and Peter L Bartlett · 2024
Closest in time.