Fetching the paper…
Reading the bibliography…
Transformer-based large language models have achieved remarkable performance across various natural language processing tasks.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
A multiscale visualization of attention in the transformer model
Vig, J. 2019 · 1906
Earlier work this paper cites.
Notes on the n-person game—ii: The value of an n-person game
Shapley, L. S. 1951 · 1951
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018 · 2018
Earlier work this paper cites.
Rise: Randomized input sampling for explanation of black-box models
Petsiuk, V.; Das, A.; and Saenko, K. 2018 · 2018
Earlier work this paper cites.
Explaining Image Classifiers by Counterfactual Generation
Chang, C.-H.; Creager, E.; Goldenberg, A.; and Duvenaud, D. 2019 · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020 · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Earlier work this paper cites.
Limitations of language models in arithmetic and symbolic induction
Qian, J.; Wang, H.; Li, Z.; Li, S.; and Yan, X. 2022 · 2022
Cited alongside, same era.
Galactica: A large language model for science
Taylor, R.; Kardas, M.; Cucurull, G.; Scialom, T.; Hartshorn, A.; Saravia, E.; Poulton, A.; Kerkez, V.; and Stojnic, R. 2022 · 2022
Cited alongside, same era.
Lamda: Language models for dialog applications
Thoppilan, R.; De Freitas, D.; Hall, J.; Shazeer, N.; Kulshreshtha, A.; Cheng, H.-T.; Jin, A.; Bos, T.; Baker, L.; Du, Y.; et al. 2022 · 2022
Cited alongside, same era.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small
Wang, K.; Variengien, A.; Conmy, A.; Shlegeris, B.; and Steinhardt, J. 2022 · 2022
Cited alongside, same era.
Understanding addition in transformers
Quirke, P.; et al. 2023 · 2023
Later among the works it cites.
Positional description matters for transformers arithmetic
Shen, R.; Bubeck, S.; Eldan, R.; Lee, Y. T.; Li, Y.; and Zhang, Y. 2023 · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Team, G.; Anil, R.; Borgeaud, S.; Wu, Y.; Alayrac, J.-B.; Yu, J.; Soricut, R.; Schalkwyk, J.; Dai, A. M.; Hauth, A.; et al. 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
Cited alongside, same era.
Teaching arithmetic to small transformers
Lee, N.; Sreenivasan, K.; Lee, J. D.; Lee, K.; and Papailiopoulos, D. 2023 · 2023
Cited alongside, same era.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Li, J.; Li, D.; Savarese, S.; and Hoi, S. 2023 · 2023
Cited alongside, same era.
Circuit component reuse across tasks in transformer language models
Merullo, J.; Eickhoff, C.; and Pavlick, E. 2023 · 2023
Cited alongside, same era.
Yang, Z.; Ding, M.; Lv, Q.; Jiang, Z.; He, Z.; Guo, Y.; Bai, J.; and Tang, J. 2023 · 2023
Later among the works it cites.
Scaling instruction-finetuned language models
Chung, H. W.; Hou, L.; Longpre, S.; Zoph, B.; Tay, Y.; Fedus, W.; Li, Y.; Wang, X.; Dehghani, M.; Brahma, S.; et al. 2024 · 2024
Closest in time.
Faith and fate: Limits of transformers on compositionality
Dziri, N.; Lu, X.; Sclar, M.; Li, X. L.; Jiang, L.; Lin, B. Y.; Welleck, S.; West, P.; Bhagavatula, C.; Le Bras, R.; et al. 2024 · 2024
Closest in time.
Mini-gemini: Mining the potential of multi-modality vision language models
Li, Y.; Zhang, Y.; Wang, C.; Zhong, Z.; Chen, Y.; Chu, R.; Liu, S.; and Jia, J. 2024 · 2024
Closest in time.