Fetching the paper…
Reading the bibliography…
State of the art foundation models such as GPT-4 perform surprisingly well at in-context learning (ICL), a variant of meta-learning concerning the learned ability to solve tasks during a neural network forward pass, exploiting contextual information provided as input to the model.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020) · 1901
Earlier work this paper cites.
Generating long sequences with sparse transformers
Child, R., Gray, S., Radford, A., and Sutskever, I. (2019) · 1904
Earlier work this paper cites.
Blockwise self-attention for long document understanding
Qiu, J., Ma, H., Levy, O., Yih, S. W.-t., Wang, S., and Tang, J. (2019) · 1911
Earlier work this paper cites.
A new approach to linear filtering and prediction problems
Kalman, R. E. (1960) · 1960
Earlier work this paper cites.
A maximization technique occurring in the statistical analysis of probabilistic functions of markov chains
Baum, L. E., Petrie, T., Soules, G., and Weiss, N. (1970) · 1970
Earlier work this paper cites.
A tutorial on hidden markov models and selected applications in speech recognition
Rabiner, L. R. (1989) · 1989
Earlier work this paper cites.
Novel approach to nonlinear/non-gaussian bayesian state estimation
Gordon, N. J., Salmond, D. J., and Smith, A. F. (1993) · 1993
Earlier work this paper cites.
Longformer: The long-document transformer
Beltagy, I., Peters, M. E., and Cohan, A. (2020) · 2004
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
Wang, S., Li, B. Z., Khabsa, M., Fang, H., and Ma, H. (2020) · 2006
Earlier work this paper cites.
Rethinking attention with performers
Choromanski, K., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J., Mohiuddin, A., Kaiser, L., et al. (2020) · 2009
Earlier work this paper cites.
Xgboost: A scalable tree boosting system
Chen, T. and Guestrin, C. (2016) · 2016
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I. (2017) · 2017
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. (2019) · 2019
Earlier work this paper cites.
The Pile: An 800GB dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al. (2020) · 2020
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F. (2020) · 2020
Earlier work this paper cites.
Transformer feed-forward layers are key-value memories
Geva, M., Schuster, R., Berant, J., and Levy, O. (2021) · 2021
Earlier work this paper cites.
Meta-learning in neural networks: A survey
Hospedales, T., Antoniou, A., Micaelli, P., and Storkey, A. (2021) · 2021
Earlier work this paper cites.
Finetuning pretrained transformers into rnns
Kasai, J., Peng, H., Zhang, Y., Yogatama, D., Ilharco, G., Pappas, N., Mao, Y., Chen, W., and Smith, N. A. (2021) · 2021
Earlier work this paper cites.
Peng, H., Pappas, N., Yogatama, D., Schwartz, R., Smith, N. A., and Kong, L. (2021) · 2021
Earlier work this paper cites.
Train short, test long: Attention with linear biases enables input length extrapolation
Press, O., Smith, N., and Lewis, M. (2021) · 2021
Cited alongside, same era.
Long Range Arena: A Benchmark for Efficient Transformers
Tay, Y., Dehghani, M., Abnar, S., Shen, Y., Bahri, D., Pham, P., Rao, J., Yang, L., Ruder, S., and Metzler, D. (2021) · 2021
Cited alongside, same era.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Wang, B. and Komatsuzaki, A. (2021) · 2021
Cited alongside, same era.
Data distributional properties drive emergent in-context learning in transformers
Chan, S. C. Y., Santoro, A., Lampinen, A. K., Wang, J. X., Singh, A., Richemond, P. H., McClelland, J., and Hill, F. (2022) · 2022
Cited alongside, same era.
Towards learning universal hyperparameter optimizers with transformers
Chen, Y., Song, X., Lee, C., Wang, Z., Zhang, R., Dohan, D., Kawakami, K., Kochanski, G., Doucet, A., Ranzato, M., et al. (2022) · 2022
Cited alongside, same era.
In-context learning creates task vectors
Hendel, R., Geva, M., and Globerson, A. (2023) · 2023
Later among the works it cites.
TabPFN: A transformer that solves small tabular classification problems in a second
Hollmann, N., Müller, S., Eggensperger, K., and Hutter, F. (2023) · 2023
Later among the works it cites.
Exploring the relationship between model architecture and in-context learning ability
Lee, I., Jiang, N., and Berg-Kirkpatrick, T. (2023) · 2023
Later among the works it cites.
Long range language modeling via gated state spaces
Mehta, H., Gupta, A., Cutkosky, A., and Neyshabur, B. (2023) · 2023
Later among the works it cites.
Pfns4bo: In-context learning for bayesian optimization
Müller, S., Feurer, M., Hollmann, N., and Hutter, F. (2023) · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dai, D., Sun, Y., Dong, L., Hao, Y., Sui, Z., and Wei, F. (2022) · 2022
Cited alongside, same era.
What can transformers learn in-context? A case study of simple function classes
Garg, S., Tsipras, D., Liang, P. S., and Valiant, G. (2022) · 2022
Cited alongside, same era.
It’s raw! audio generation with state-space models
Goel, K., Gu, A., Donahue, C., and Ré, C. (2022) · 2022
Cited alongside, same era.
Diagonal state spaces are as effective as structured state spaces
Gupta, A., Gu, A., and Berant, J. (2022) · 2022
Cited alongside, same era.
Transformers can do bayesian inference
Müller, S., Hollmann, N., Arango, S. P., Grabocka, J., and Hutter, F. (2022) · 2022
Cited alongside, same era.
S4nd: Modeling images and videos as multidimensional signals with state spaces
Nguyen, E., Goel, K., Gu, A., Downs, G., Shah, P., Dao, T., Baccus, S., and Ré, C. (2022) · 2022
Cited alongside, same era.
In-context learning and induction heads
Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., et al. (2022) · 2022
Cited alongside, same era.
Peng, B., Alcaide, E., Anthony, Q., Albalak, A., Arcadinho, S., Cao, H., Cheng, X., Chung, M., Grella, M., GV, K. K., et al. (2023) · 2023
Later among the works it cites.
Do pretrained transformers really learn in-context by gradient descent?
Shen, L., Mishra, A., and Khashabi, D. (2023) · 2023
Later among the works it cites.
Simplified state space layers for sequence modeling
Smith, J. T., Warrington, A., and Linderman, S. (2023) · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al. (2023) · 2023
Later among the works it cites.
Transformers learn in-context by gradient descent
Von Oswald, J., Niklasson, E., Randazzo, E., Sacramento, J., Mordvintsev, A., Zhmoginov, A., and Vladymyrov, M. (2023) · 2023
Later among the works it cites.
Selective structured state-spaces for long-form video understanding
Wang, J., Zhu, W., Wang, P., Yu, X., Liu, L., Omar, M., and Hamid, R. (2023) · 2023
Later among the works it cites.
Efficient bayesian learning curve extrapolation using prior-data fitted networks
Adriaensen, S., Rakotoarison, H., Müller, S., and Hutter, F. (2024) · 2024
Closest in time.
In-context language learning: Architectures and Algorithms
Akyürek, E., Wang, B., Kim, Y., and Andreas, J. (2024) · 2024
Closest in time.
Forecastpfn: Synthetically-trained zero-shot forecasting
Dooley, S., Khurana, G. S., Mohapatra, C., Naidu, S. V., and White, C. (2024) · 2024
Closest in time.
U-Mamba: Enhancing long-range dependency for biomedical image segmentation
Ma, J., Li, F., and Wang, B. (2024) · 2024
Closest in time.
Can mamba learn how to learn? a comparative study on in-context learning tasks
Park, J., Park, J., Xiong, Z., Lee, N., Cho, J., Oymak, S., Lee, K., and Papailiopoulos, D. (2024) · 2024
Closest in time.
The mechanistic basis of data dependence and abrupt learning in an in-context classification task
Reddy, G. (2024) · 2024
Closest in time.
Function vectors in large language models
Todd, E., Li, M. L., Sharma, A. S., Mueller, A., Wallace, B. C., and Bau, D. (2024) · 2024
Closest in time.