Fetching the paper…
Reading the bibliography…
The current large language models are mainly based on decode-only structure transformers, which have great in-context learning (ICL) capabilities.
Gestalt principles
Todorovic, D · 2008
Earlier work this paper cites.
Recurrent neural network based language model
Mikolov, T., Karafiát, M., Burget, L., Cernockỳ, J., and Khudanpur, S · 2010
Earlier work this paper cites.
Lstm neural networks for language modeling
Sundermeyer, M., Schlüter, R., and Ney, H · 2012
Earlier work this paper cites.
The lambada dataset: Word prediction requiring a broad discourse context
Paperno, D., Kruszewski, G., Lazaridou, A., Pham, N.-Q., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fernández, R · 2016
Earlier work this paper cites.
Attention is all you need
Vaswani, A · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Shift: A zero flop, zero parameter alternative to spatial convolutions
Wu, B., Wan, A., Yue, X., Jin, P., Zhao, S., Golmant, N., Gholaminejad, A., Gonzalez, J., and Keutzer, K · 2018
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
A mathematical framework for transformer circuits
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Winogrande: An adversarial winograd schema challenge at scale
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y · 2021
Earlier work this paper cites.
Token shift transformer for video classification
Zhang, H., Hao, Y., and Ngo, C.-W · 2021
Earlier work this paper cites.
In-context learning and induction heads
Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Johnston, S., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
Gqa: Training generalized multi-query transformer models from multi-head checkpoints
Ainslie, J., Lee-Thorp, J., de Jong, M., Zemlyanskiy, Y., Lebron, F., and Sanghai, S · 2023
Cited alongside, same era.
Rethinking the role of scale for in-context learning: An interpretability-based case study at 66 billion scale
Bansal, H., Gopalakrishnan, K., Dingliwal, S., Bodapati, S., Kirchhoff, K., and Roth, D · 2023
Cited alongside, same era.
Towards automated circuit discovery for mechanistic interpretability
Conmy, A., Mavor-Parker, A., Lynch, A., Heimersheim, S., and Garriga-Alonso, A · 2023
Cited alongside, same era.
Mamba: Linear-time sequence modeling with selective state spaces
Gu, A. and Dao, T · 2023
Induction heads as an essential mechanism for pattern matching in in-context learning
Crosbie, J. and Shutova, E · 2024
Closest in time.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Closest in time.
Learning transformer programs
Friedman, D., Wettig, A., and Chen, D · 2024
Closest in time.
Selective attention improves transformer
Leviathan, Y., Kalman, M., and Matias, Y · 2024
Closest in time.
Fineweb-edu, May 2024
Lozhkov, A., Ben Allal, L., von Werra, L., and Wolf, T · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Large learning rates improve generalization: But how large are we talking about?
Lobacheva, E., Pokonechny, E., Kodryan, M., and Vetrov, D · 2023
Cited alongside, same era.
Rwkv: Reinventing rnns for the transformer era
Peng, B., Alcaide, E., Anthony, Q., Albalak, A., Arcadinho, S., Biderman, S., Cao, H., Cheng, X., Chung, M., Grella, M., et al · 2023
Cited alongside, same era.
Retentive network: A successor to transformer for large language models
Sun, Y., Dong, L., Huang, S., Ma, S., Xia, Y., Xue, J., Wang, J., and Wei, F · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Cited alongside, same era.
Sub-task decomposition enables learning in sequence to sequence tasks
Wies, N., Levine, Y., and Shashua, A · 2023
Cited alongside, same era.
In-context language learning: Architectures and algorithms
Akyürek, E., Wang, B., Kim, Y., and Andreas, J · 2024
Cited alongside, same era.
Fundamental limitations on subquadratic alternatives to transformers
Alman, J. and Yu, H · 2024
Cited alongside, same era.
Penedo, G., Kydlícek, H., Allal, L. B., and Wolf, T · 2024
Closest in time.
Transformers on markov data: Constant depth suffices
Rajaraman, N., Bondaschi, M., Ramchandran, K., Gastpar, M., and Makkuva, A. V · 2024
Closest in time.
Identifying semantic induction heads to understand in-context learning
Ren, J., Guo, Q., Yan, H., Liu, D., Zhang, Q., Qiu, X., and Lin, D · 2024
Closest in time.
Out-of-distribution generalization via composition: a lens through induction heads in transformers
Song, J., Xu, Z., and Zhong, Y · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y · 2024
Closest in time.
An empirical study of mamba-based language models
Waleffe, R., Byeon, W., Riach, D., Norick, B., Korthikanti, V., Dao, T., Gu, A., Hatamizadeh, A., Singh, S., Narayanan, D., et al · 2024
Closest in time.
How transformers implement induction heads: Approximation and optimization analysis
Wang, M., Yu, R., Wu, L., et al · 2024
Closest in time.
Base of roPE bounds context length
Xu, M., Men, X., Wang, B., Zhang, Q., Lin, H., Han, X., and weipeng chen · 2024
Closest in time.
Yang, A., Yang, B., Hui, B., Zheng, B., Yu, B., Zhou, C., Li, C., Li, C., Liu, D., Huang, F., et al · 2024
Closest in time.