Fetching the paper…
Reading the bibliography…
In-context learning is a key paradigm in large language models (LLMs) that enables them to generalize to new tasks and domains by simply prompting these models with a few exemplars without explicit parameter updates.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay · 2011
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
R. Socher, A. Perelygin, J. Wu, J. Chuang, C. D. Manning, A. Y. Ng, and C. Potts · 2013
Earlier work this paper cites.
Good debt or bad debt: Detecting semantic orientations in economic texts
P. Malo, A. Sinha, P. Korhonen, J. Wallenius, and P. Takala · 2014
Earlier work this paper cites.
Character-level convolutional networks for text classification
X. Zhang, J. Zhao, and Y. LeCun · 2015
Earlier work this paper cites.
Senteval: An evaluation toolkit for universal sentence representations
A. Conneau and D. Kiela · 2018
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
A. Wang, Y. Pruksachatkun, N. Nangia, A. Singh, J. Michael, F. Hill, O. Levy, and S. Bowman · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Earlier work this paper cites.
Ethos: an online hate speech detection dataset
I. Mollas, Z. Chrysopoulou, S. Karlos, and G. Tsoumakas · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
E. J. Hu, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen, et al · 2021
Earlier work this paper cites.
Transformers can do bayesian inference
S. Müller, N. Hollmann, S. P. Arango, J. Grabocka, and F. Hutter · 2021
Earlier work this paper cites.
Llm.int8(): 8-bit matrix multiplication for transformers at scale, 2022
T. Dettmers, M. Lewis, Y. Belkada, and L. Zettlemoyer · 2022
Earlier work this paper cites.
What can transformers learn in-context? a case study of simple function classes
S. Garg, D. Tsipras, P. S. Liang, and G. Valiant · 2022
Earlier work this paper cites.
Can language models learn from explanations in context?
A. Lampinen, I. Dasgupta, S. Chan, K. Mathewson, M. Tessler, A. Creswell, J. McClelland, J. Wang, and F. Hill · 2022
Cited alongside, same era.
Metaicl: Learning to learn in context
S. Min, M. Lewis, L. Zettlemoyer, and H. Hajishirzi · 2022
Cited alongside, same era.
Rethinking the role of demonstrations: What makes in-context learning work?
S. Min, X. Lyu, A. Holtzman, M. Artetxe, M. Lewis, H. Hajishirzi, and L. Zettlemoyer · 2022
Cited alongside, same era.
Transformer neural processes: Uncertainty-aware meta learning via sequence modeling
T. Nguyen and A. Grover · 2022
Cited alongside, same era.
Do prompt-based models really understand the meaning of their prompts?
A. Webson and E. Pavlick · 2022
Cited alongside, same era.
Emergent abilities of large language models
J. Wei, Y. Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzler, et al · 2022
Larger language models do in-context learning differently
J. Wei, J. Wei, Y. Tay, D. Tran, A. Webson, Y. Lu, X. Chen, H. Liu, D. Huang, D. Zhou, et al · 2023
Later among the works it cites.
How many pretraining tasks are needed for in-context learning of linear regression?
J. Wu, D. Zou, Z. Chen, V. Braverman, Q. Gu, and P. Bartlett · 2023
Later among the works it cites.
Sheared llama: Accelerating language model pre-training via structured pruning
M. Xia, T. Gao, Z. Zeng, and D. Chen · 2023
Later among the works it cites.
Trained transformers learn linear models in-context
R. Zhang, S. Frei, and P. L. Bartlett · 2023
Later among the works it cites.
Group preference optimization: Few-shot alignment of large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Cited alongside, same era.
Why can gpt learn in-context? language models secretly perform gradient descent as meta-optimizers
D. Dai, Y. Sun, L. Dong, Y. Hao, S. Ma, Z. Sui, and F. Wei · 2023
Cited alongside, same era.
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. d. l. Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, et al · 2023
Cited alongside, same era.
Expt: Synthetic pretraining for few-shot experimental design
T. Nguyen, S. Agrawal, and A. Grover · 2023
Cited alongside, same era.
Large language models can be easily distracted by irrelevant context
F. Shi, X. Chen, K. Misra, N. Scales, D. Dohan, E. H. Chi, N. Schärli, and D. Zhou · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Cited alongside, same era.
S. Zhao, J. Dang, and A. Grover · 2023
Later among the works it cites.
R. Agarwal, A. Singh, L. M. Zhang, B. Bohnet, S. Chan, A. Anand, Z. Abbas, A. Nova, J. D. Co-Reyes, E. Chu, et al · 2024
Closest in time.
Transformers learn to implement preconditioned gradient descent for in-context learning
K. Ahn, X. Cheng, H. Daneshmand, and S. Sra · 2024
Closest in time.
Transformers as statisticians: Provable in-context learning with in-context algorithm selection
Y. Bai, F. Chen, H. Wang, C. Xiong, and S. Mei · 2024
Closest in time.
In-context learning with long-context models: An in-depth exploration
A. Bertsch, M. Ivgi, U. Alon, J. Berant, M. R. Gormley, and G. Neubig · 2024
Closest in time.
Premise order matters in reasoning with large language models
X. Chen, R. A. Chi, X. Wang, and D. Zhou · 2024
Closest in time.
Lico: Large language models for in-context molecular optimization
T. Nguyen and A. Grover · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
M. Reid, N. Savinov, D. Teplyashin, D. Lepikhin, T. Lillicrap, J.-b. Alayrac, R. Soricut, A. Lazaridou, O. Firat, J. Schrittwieser, et al · 2024
Closest in time.
The learnability of in-context learning
N. Wies, Y. Levine, and A. Shashua · 2024
Closest in time.