Fetching the paper…
Reading the bibliography…
In-context learning (ICL) is a type of prompting where a transformer model operates on a sequence of (input, output) examples and performs inference on-the-fly.
Bayes estimates for the linear model
Dennis V Lindley and Adrian FM Smith · 1972
Earlier work this paper cites.
Rates of convergence for empirical processes of stationary mixing sequences
Bin Yu · 1994
Earlier work this paper cites.
System identification
Lennart Ljung · 1998
Earlier work this paper cites.
Algorithmic stability and generalization performance
Olivier Bousquet and André Elisseeff · 2000
Earlier work this paper cites.
Stability and generalization
Olivier Bousquet and André Elisseeff · 2002
Earlier work this paper cites.
Rademacher complexity bounds for non-iid processes
Mehryar Mohri and Afshin Rostamizadeh · 2008
Earlier work this paper cites.
A theory of learning from different domains
Shai Ben-David, John Blitzer, Koby Crammer, Alex Kulesza, Fernando Pereira, and Jennifer Wortman Vaughan · 2010
Earlier work this paper cites.
Stability bounds for stationary φ \varphi -mixing and β \beta -mixing processes
Mehryar Mohri and Afshin Rostamizadeh · 2010
Earlier work this paper cites.
Generalization bounds for time series prediction with non-stationary processes
Vitaly Kuznetsov and Mehryar Mohri · 2014
Earlier work this paper cites.
Time series prediction and online learning
Vitaly Kuznetsov and Mehryar Mohri · 2016
Earlier work this paper cites.
A vector-contraction inequality for rademacher complexities
Andreas Maurer · 2016
Earlier work this paper cites.
The benefit of multitask representation learning
Andreas Maurer, Massimiliano Pontil, and Bernardino Romera-Paredes · 2016
Earlier work this paper cites.
Regularized linear system identification using atomic, nuclear and kernel-based norms: The role of the stability constraint
Gianluigi Pillonetto, Tianshi Chen, Alessandro Chiuso, Giuseppe De Nicolao, and Lennart Ljung · 2016
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Earlier work this paper cites.
Nonparametric risk bounds for time-series forecasting
Daniel J McDonald, Cosma Rohilla Shalizi, and Mark Schervish · 2017
Earlier work this paper cites.
Geometry of optimization and implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, Ruslan Salakhutdinov, and Nathan Srebro · 2017
Earlier work this paper cites.
Foundations of machine learning
Mehryar Mohri, Afshin Rostamizadeh, and Ameet Talwalkar · 2018
Earlier work this paper cites.
Learning without mixing: Towards a sharp analysis of linear system identification
Max Simchowitz, Horia Mania, Stephen Tu, Michael I Jordan, and Benjamin Recht · 2018
Earlier work this paper cites.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Earlier work this paper cites.
High-dimensional probability: An introduction with applications in data science
Roman Vershynin · 2018
Earlier work this paper cites.
On the value of target data in transfer learning
Steve Hanneke and Samory Kpotufe · 2019
Cited alongside, same era.
A tutorial on concentration bounds for system identification
Nikolai Matni and Stephen Tu · 2019
Cited alongside, same era.
Near optimal finite time identification of arbitrary linear dynamical systems
Tuhin Sarkar and Alexander Rakhlin · 2019
Cited alongside, same era.
High-dimensional statistics: A non-asymptotic viewpoint
Martin J Wainwright · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
On the sample complexity of the linear quadratic regulator
Sarah Dean, Horia Mania, Nikolai Matni, Benjamin Recht, and Stephen Tu · 2020
Cited alongside, same era.
Provable benefit of multitask representation learning in reinforcement learning
Yuan Cheng, Songtao Feng, Jing Yang, Hong Zhang, and Yingbin Liang · 2022
Later among the works it cites.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Later among the works it cites.
Maml and anil provably learn representations
Liam Collins, Aryan Mokhtari, Sewoong Oh, and Sanjay Shakkottai · 2022
Later among the works it cites.
Why can gpt learn in-context? language models secretly perform gradient descent as meta optimizers
Damai Dai, Yutao Sun, Li Dong, Yaru Hao, Zhifang Sui, and Furu Wei · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Few-shot learning via learning the representation, provably
Simon S Du, Wei Hu, Sham M Kakade, Jason D Lee, and Qi Lei · 2020
Cited alongside, same era.
Learning nonlinear dynamical systems from a single trajectory
Dylan Foster, Tuhin Sarkar, and Alexander Rakhlin · 2020
Cited alongside, same era.
Meta-learning for mixed linear regression
Weihao Kong, Raghav Somani, Zhao Song, Sham Kakade, and Sewoong Oh · 2020
Cited alongside, same era.
Active learning for nonlinear system identification with guarantees
Horia Mania, Michael I Jordan, and Benjamin Recht · 2020
Cited alongside, same era.
On the theory of transfer learning: The importance of task diversity
Nilesh Tripuraneni, Michael Jordan, and Chi Jin · 2020
Cited alongside, same era.
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant · 2021
Cited alongside, same era.
Mohamad Kazem Shirani Faradonbeh and Aditya Modi · 2022
Later among the works it cites.
What can transformers learn in-context? a case study of simple function classes
Shivam Garg, Dimitris Tsipras, Percy Liang, and Gregory Valiant · 2022
Later among the works it cites.
Tabpfn: A transformer that solves small tabular classification problems in a second
Noah Hollmann, Samuel Müller, Katharina Eggensperger, and Frank Hutter · 2022
Later among the works it cites.
In-context reinforcement learning with algorithm distillation
Michael Laskin, Luyu Wang, Junhyuk Oh, Emilio Parisotto, Stephen Spencer, Richie Steigerwald, DJ Strouse, Steven Hansen, Angelos Filos, Ethan Brooks, et al · 2022
Later among the works it cites.
Provable and efficient continual representation learning
Yingcong Li, Mingchen Li, M Salman Asif, and Samet Oymak · 2022
Later among the works it cites.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, et al · 2022
Later among the works it cites.
Non-stationary representation learning in sequential linear bandits
Yuzhen Qin, Tommaso Menara, Samet Oymak, Shinung Ching, and Fabio Pasqualetti · 2022
Later among the works it cites.
Non-asymptotic and accurate learning of nonlinear dynamical systems
Yahya Sattar and Samet Oymak · 2022
Later among the works it cites.
Finite sample identification of low-order lti systems via nuclear norm regularization
Yue Sun, Samet Oymak, and Maryam Fazel · 2022
Later among the works it cites.
Statistical learning theory for control: A finite sample perspective
Anastasios Tsiamis, Ingvar Ziemann, Nikolai Matni, and George J Pappas · 2022
Later among the works it cites.
Transformers learn in-context by gradient descent
Johannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov · 2022
Later among the works it cites.
Multi-task imitation learning for linear dynamical systems
Thomas T Zhang, Katie Kang, Bruce D Lee, Claire Tomlin, Sergey Levine, Stephen Tu, and Nikolai Matni · 2022
Later among the works it cites.
Learning with little mixing
Ingvar Ziemann and Stephen Tu · 2022
Later among the works it cites.
Single trajectory nonparametric learning of nonlinear dynamics
Ingvar M Ziemann, Henrik Sandberg, and Nikolai Matni · 2022
Later among the works it cites.
Smoothed online learning for prediction in piecewise affine systems
Adam Block, Max Simchowitz, and Russ Tedrake · 2023
Closest in time.