Fetching the paper…
Reading the bibliography…
Large Language Models(LLMs) have been attracting attention due to a ability called in-context learning(ICL).
C. Finn, P. Abbeel, and S. Levine, “Model-agnostic meta-learning for fast adaptation of deep networks,” in International conference on machine learning . PMLR, 2017, pp. 1126–1135
2017
Earlier work this paper cites.
G. Ke, Q. Meng, T. Finley, T. Wang, W. Chen, W. Ma, Q. Ye, and T.-Y. Liu, “Lightgbm: A highly efficient gradient boosting decision tree,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
S. Chan, A. Santoro, A. Lampinen, J. Wang, A. Singh, P. Richemond, J. McClelland, and F. Hill, “Data distributional properties drive emergent in-context learning in transformers,” Advances in Neural Information Processing Systems , vol. 35, pp. 18 878–18 891, 2022
2022
Cited alongside, same era.
D. Chijiwa, S. Yamaguchi, A. Kumagai, and Y. Ida, “Meta-ticket: Finding optimal subnetworks for few-shot learning within randomly initialized neural networks,” Advances in Neural Information Processing Systems , vol. 35, pp. 25 264–25 277, 2022
2022
Cited alongside, same era.
S. Garg, D. Tsipras, P. S. Liang, and G. Valiant, “What can transformers learn in-context? a case study of simple function classes,” Advances in Neural Information Processing Systems , vol. 35, pp. 30 583–30 598, 2022
2022
Cited alongside, same era.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou et al. , “Chain-of-thought prompting elicits reasoning in large language models,” Advances in Neural Information Processing Systems , vol. 35, pp. 24 824–24 837, 2022
2023
Closest in time.
2023
Closest in time.
J. Von Oswald, E. Niklasson, E. Randazzo, J. Sacramento, A. Mordvintsev, A. Zhmoginov, and M. Vladymyrov, “Transformers learn in-context by gradient descent,” in International Conference on Machine Learning . PMLR, 2023, pp. 35 151–35 174
2023
Closest in time.
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.