Fetching the paper…
Reading the bibliography…
Large language models (LLM) have emerged as a powerful tool for AI, with the key ability of in-context learning (ICL), where they can perform well on unseen tasks based on a brief series of task examples without necessitating any adjustments to the model parameters.
The evaluation of the collision matrix
Wick, G.-C · 1950
Earlier work this paper cites.
Sparse coding with an overcomplete basis set: A strategy employed by v1?
Olshausen, B. and Field, D · 1997
Earlier work this paper cites.
Sparse coding and decorrelation in primary visual cortex during natural vision
Vinje, W. E. and Gallant, J. L · 2000
Earlier work this paper cites.
Latent dirichlet allocation
Blei, D. M., Ng, A. Y., and Jordan, M. I · 2003
Earlier work this paper cites.
An isserlis’ theorem for mixed gaussian variables: Application to the auto-bispectral density
Michalowicz, J., Nichols, J., Bucholtz, F., and Olson, C · 2009
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
SentEval: An evaluation toolkit for universal sentence representations
Conneau, A. and Kiela, D · 2018
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Low-rank bottleneck in multi-head attention models
Bhojanapalli, S., Yun, C., Rawat, A. S., Reddi, S., and Kumar, S · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Learning parities with neural networks
Daniely, A. and Malach, E · 2020
Earlier work this paper cites.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al · 2021
Earlier work this paper cites.
Scatterbrain: Unifying sparse and low-rank attention
Chen, B., Dao, T., Winsor, E., Song, Z., Rudra, A., and Ré, C · 2021
Earlier work this paper cites.
Lighter and better: low-rank decomposed self-attention networks for next-item recommendation
Fan, X., Liu, Z., Lian, J., Zhao, W. X., Xie, X., and Wen, J.-R · 2021
Earlier work this paper cites.
Transformer feed-forward layers are key-value memories
Geva, M., Schuster, R., Berant, J., and Levy, O · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
Lester, B., Al-Rfou, R., and Constant, N · 2021
Earlier work this paper cites.
Prefix-tuning: Optimizing continuous prompts for generation
Li, X. L. and Liang, P · 2021
Earlier work this paper cites.
Metaicl: Learning to learn in context
Min, S., Lewis, M., Zettlemoyer, L., and Hajishirzi, H · 2021
Earlier work this paper cites.
Show your work: Scratchpads for intermediate computation with language models
Nye, M., Andreassen, A. J., Gur-Ari, G., Michalewski, H., Austin, J., Bieber, D., Dohan, D., Lewkowycz, A., Bosma, M., Luan, D., et al · 2021
Earlier work this paper cites.
Linear transformers are secretly fast weight programmers
Schlag, I., Irie, K., and Schmidhuber, J · 2021
Earlier work this paper cites.
Calibrate before use: Improving few-shot performance of language models
Zhao, Z., Wallace, E., Feng, S., Klein, D., and Singh, S · 2021
Earlier work this paper cites.
In-context examples selection for machine translation
Agrawal, S., Zhou, C., Lewis, M., Zettlemoyer, L., and Ghazvininejad, M · 2022
Earlier work this paper cites.
Hidden progress in deep learning: Sgd learns parities near the computational limit
Barak, B., Edelman, B., Goel, S., Kakade, S., Malach, E., and Zhang, C · 2022
Earlier work this paper cites.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2022
Earlier work this paper cites.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., et al · 2022
Cited alongside, same era.
Why can gpt learn in-context? language models secretly perform gradient descent as meta optimizers
Dai, D., Sun, Y., Dong, L., Hao, Y., Sui, Z., and Wei, F · 2022
Cited alongside, same era.
A survey for in-context learning
Dong, Q., Li, L., Dai, D., Zheng, C., Wu, Z., Chang, B., Sun, X., Xu, J., and Sui, Z · 2022
Cited alongside, same era.
What can transformers learn in-context? a case study of simple function classes
Garg, S., Tsipras, D., Liang, P. S., and Valiant, G · 2022
Cited alongside, same era.
LoRA: Low-rank adaptation of large language models
Hu, E. J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Vitality: Unifying low-rank and sparse approximation for vision transformer acceleration with a linear taylor attention
Dass, J., Wu, S., Shi, H., Li, C., Ye, Z., Wang, Z., and Lin, Y · 2023
Later among the works it cites.
Improving language model negotiation with self-play and in-context learning from ai feedback
Fu, Y., Peng, H., Khot, T., and Lapata, M · 2023
Later among the works it cites.
Llama-adapter v2: Parameter-efficient visual instruction model
Gao, P., Han, J., Zhang, R., Lin, Z., Geng, S., Zhou, A., Zhang, W., Lu, P., He, C., Yue, X., et al · 2023
Later among the works it cites.
In-context convergence of transformers
Huang, Y., Cheng, Y., and Liang, Y · 2023
Later among the works it cites.
Transformers as multi-task feature selectors: Generalization analysis of in-context learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Opt-iml: Scaling language model instruction meta learning through the lens of generalization
Iyer, S., Lin, X. V., Pasunuru, R., Mihaylov, T., Simig, D., Yu, P., Shuster, K., Wang, T., Liu, Q., Koura, P. S., et al · 2022
Cited alongside, same era.
Vision transformers provably learn spatial structure
Jelassi, S., Sander, M., and Li, Y · 2022
Cited alongside, same era.
Demonstrate-search-predict: Composing retrieval and language models for knowledge-intensive nlp
Khattab, O., Santhanam, K., Li, X. L., Hall, D., Liang, P., Potts, C., and Zaharia, M · 2022
Cited alongside, same era.
Rethinking the role of demonstrations: What makes in-context learning work?
Min, S., Lyu, X., Holtzman, A., Artetxe, M., Lewis, M., Hajishirzi, H., and Zettlemoyer, L · 2022
Cited alongside, same era.
Cross-task generalization via natural language crowdsourcing instructions
Mishra, S., Khashabi, D., Baral, C., and Hajishirzi, H · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
A theoretical analysis on feature learning in neural networks: Emergence from inputs and advantage over fixed features
Shi, Z., Wei, J., and Liang, Y · 2022
Cited alongside, same era.
Li, H., Wang, M., Lu, S., Wan, H., Cui, X., and Chen, P.-Y · 2023
Later among the works it cites.
Understanding the robustness of self-supervised learning through topic modeling
Luo, Z., Wu, S., Weng, C., Zhou, M., and Ge, R · 2023
Later among the works it cites.
Mahankali, A., Hashimoto, T. B., and Ma, T · 2023
Later among the works it cites.
Fine-tuning language models with just forward passes
Malladi, S., Gao, T., Nichani, E., Damian, A., Lee, J. D., Chen, D., and Arora, S · 2023
Later among the works it cites.
Introducing ChatGPT
OpenAI · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
What in-context learning ’learns’ in-context: Disentangling task recognition and task learning
Pan, J., Gao, T., Chen, H., and Chen, D · 2023
Later among the works it cites.
Trainable transformer in transformer
Panigrahi, A., Malladi, S., Xia, M., and Arora, S · 2023
Later among the works it cites.
Din-sql: Decomposed in-context learning of text-to-sql with self-correction
Pourreza, M. and Rafiei, D · 2023
Later among the works it cites.
Pretraining task diversity and the emergence of non-bayesian in-context learning for regression
Raventos, A., Paul, M., Chen, F., and Ganguli, S · 2023
Later among the works it cites.
Transformers learn in-context by gradient descent
Von Oswald, J., Niklasson, E., Randazzo, E., Sacramento, J., Mordvintsev, A., Zhmoginov, A., and Vladymyrov, M · 2023
Later among the works it cites.
Symbol tuning improves in-context learning in language models
Wei, J., Hou, L., Lampinen, A. K., Chen, X., Huang, D., Tay, Y., Chen, X., Lu, Y., Zhou, D., Ma, T., and Le, Q. V · 2023
Later among the works it cites.
On the role of unstructured training data in transformers’ in-context learning capabilities
Wibisono, K. C. and Wang, Y · 2023
Later among the works it cites.
Improving foundation models for few-shot learning via multitask finetuning
Xu, Z., Shi, Z., Wei, J., Li, Y., and Liang, Y · 2023
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., and Narasimhan, K. R · 2023
Later among the works it cites.
Linear attention is (maybe) all you need (to understand transformer optimization)
Ahn, K., Cheng, X., Song, M., Yun, C., Jadbabaie, A., and Sra, S · 2024
Closest in time.
How do transformers learn in-context beyond simple functions? a case study on learning with representations
Guo, T., Hu, W., Mei, S., Wang, H., Xiong, C., Savarese, S., and Bai, Y · 2024
Closest in time.
The mechanistic basis of data dependence and abrupt learning in an in-context classification task
Reddy, G · 2024
Closest in time.
How many pretraining tasks are needed for in-context learning of linear regression?
Wu, J., Zou, D., Chen, Z., Braverman, V., Gu, Q., and Bartlett, P. L · 2024
Closest in time.
Do large language models have compositional ability? an investigation into limitations and scalability
Xu, Z., Shi, Z., and Liang, Y · 2024
Closest in time.
Step-back prompting enables reasoning via abstraction in large language models
Zheng, H. S., Mishra, S., Chen, X., Cheng, H.-T., Chi, E. H., Le, Q. V., and Zhou, D · 2024
Closest in time.