Fetching the paper…
Reading the bibliography…
Recent works attribute the capability of in-context learning (ICL) in large pre-trained language models to implicitly simulating and fine-tuning an internal model (e.g., linear or 2-layer MLP) during inference.
Mining and summarizing customer reviews
Hu, M. and Liu, B · 2004
Earlier work this paper cites.
A sentimental education: Sentiment analysis using subjectivity summarization based on minimum cuts
Pang, B. and Lee, L · 2004
Earlier work this paper cites.
Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales
Pang, B. and Lee, L · 2005
Earlier work this paper cites.
Annotating expressions of opinions and emotions in language
Wiebe, J., Wilson, T., and Cardie, C · 2005
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A. Y., and Potts, C · 2013
Earlier work this paper cites.
Character-level convolutional networks for text classification
Zhang, X., Zhao, J., and LeCun, Y · 2015
Earlier work this paper cites.
Concrete problems in ai safety
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D · 2016
Earlier work this paper cites.
Using fast weights to attend to the recent past, 2016
Ba, J., Hinton, G., Mnih, V., Leibo, J. Z., and Ionescu, C · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Hendrycks, D. and Gimpel, K · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Earlier work this paper cites.
Equilibrium propagation: Bridging the gap between energy-based models and backpropagation
Scellier, B. and Bengio, Y · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Scalable agent alignment via reward modeling: a research direction
Leike, J., Krueger, D., Everitt, T., Martic, M., Maini, V., and Legg, S · 2018
Earlier work this paper cites.
Efficient training of bert by progressively stacking
Gong, L., He, D., Li, Z., Qin, T., Wang, L., and Liu, T · 2019
Earlier work this paper cites.
On the turing completeness of modern neural network architectures
Pérez, J., Marinković, J., and Barceló, P · 2019
Earlier work this paper cites.
Root mean square layer normalization
Zhang, B. and Sennrich, R · 2019
Earlier work this paper cites.
On the ability and limitations of transformers to recognize formal languages
Bhattamishra, S., Ahuja, K., and Goyal, N · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Leveraging passage retrieval with generative models for open domain question answering
Izacard, G. and Grave, E · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
A mathematical exploration of why language models help solve downstream tasks
Saunshi, N., Malladi, S., and Arora, S · 2020
Earlier work this paper cites.
Glu variants improve transformer
Shazeer, N · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., DasSarma, N., et al · 2021
Earlier work this paper cites.
A mathematical framework for transformer circuits
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., et al · 2021
Earlier work this paper cites.
Surface form competition: Why the highest probability answer isn’t always right
Holtzman, A., West, P., Shwartz, V., Choi, Y., and Zettlemoyer, L · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Cited alongside, same era.
Going beyond linear transformers with recurrent fast weight programmers
Irie, K., Schlag, I., Csordás, R., and Schmidhuber, J · 2021
Cited alongside, same era.
Attention is turing-complete
Perez, J., Barcelo, P., and Marinkovic, J · 2021
Cited alongside, same era.
Train short, test long: Attention with linear biases enables input length extrapolation
Press, O., Smith, N. A., and Lewis, M · 2021
Cited alongside, same era.
Ul2: Unifying language learning paradigms
Tay, Y., Dehghani, M., Tran, V. Q., Garcia, X., Wei, J., Wang, X., Chung, H. W., Bahri, D., Schuster, T., Zheng, S., et al · 2022
Later among the works it cites.
Interpretability in the wild: a circuit for indirect object identification in GPT-2 small
Wang, K. R., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J · 2022
Later among the works it cites.
An explanation of in-context learning as implicit bayesian inference
Xie, S. M., Raghunathan, A., Liang, P., and Ma, T · 2022
Later among the works it cites.
Transformers learn to implement preconditioned gradient descent for in-context learning
Ahn, K., Cheng, X., Daneshmand, H., and Sra, S · 2023
Closest in time.
Transformers as statisticians: Provable in-context learning with in-context algorithm selection
Bai, Y., Chen, F., Wang, H., Xiong, C., and Mei, S · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Linear transformers are secretly fast weight memory systems
Schlag, I., Irie, K., and Schmidhuber, J · 2021
Cited alongside, same era.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Lu, Y., Pan, S., Murtadha, A., Wen, B., and Liu, Y · 2021
Cited alongside, same era.
Wei, C., Chen, Y., and Ma, T · 2021
Cited alongside, same era.
Thinking like transformers
Weiss, G., Goldberg, Y., and Yahav, E · 2021
Cited alongside, same era.
Self-attention networks can process bounded hierarchical languages
Yao, S., Peng, B., Papadimitriou, C., and Narasimhan, K · 2021
Cited alongside, same era.
What learning algorithm is in-context learning? investigations with linear models
Akyurek, E., Schuurmans, D., Andreas, J., Ma, T., and Zhou, D · 2022
Cited alongside, same era.
Exploring length generalization in large language models
Anil, C., Wu, Y., Andreassen, A., Lewkowycz, A., Misra, V., Ramasesh, V., Slone, A., Gur-Ari, G., Dyer, E., and Neyshabur, B · 2022
Cited alongside, same era.
Transformers implement functional gradient descent to learn non-linear functions in context
Cheng, X., Chen, Y., and Sra, S · 2023
Closest in time.
A toy model of universality: Reverse engineering how networks learn group operations
Chughtai, B., Chan, L., and Nanda, N · 2023
Closest in time.
Towards automated circuit discovery for mechanistic interpretability
Conmy, A., Mavor-Parker, A. N., Lynch, A., Heimersheim, S., and Garriga-Alonso, A · 2023
Closest in time.
Looped transformers as programmable computers, 2023
Giannou, A., Rajput, S., yong Sohn, J., Lee, K., Lee, J. D., and Papailiopoulos, D · 2023
Closest in time.
A theory of emergent in-context learning as implicit structure induction
Hahn, M. and Goyal, N · 2023
Closest in time.
In-context learning of large language models explained as kernel regression
Han, C., Wang, Z., Zhao, H., and Ji, H · 2023
Closest in time.
Atlas: Few-shot learning with retrieval augmented language models
Izacard, G., Lewis, P., Lomeli, M., Hosseini, L., Petroni, F., Schick, T., Dwivedi-Yu, J., Joulin, A., Riedel, S., and Grave, E · 2023
Closest in time.
A latent space theory for emergent abilities in large language models
Jiang, H · 2023
Closest in time.
Tracr: Compiled transformers as a laboratory for interpretability
Lindner, D., Kramár, J., Rahtz, M., McGrath, T., and Mikulik, V · 2023
Closest in time.
Transformers learn shortcuts to automata
Liu, B., Ash, J. T., Goel, S., Krishnamurthy, A., and Zhang, C · 2023
Closest in time.
Mahankali, A., Hashimoto, T. B., and Ma, T · 2023
Closest in time.
Fine-tuning language models with just forward passes
Malladi, S., Gao, T., Nichani, E., Damian, A., Lee, J. D., Chen, D., and Arora, S · 2023
Closest in time.
Progress measures for grokking via mechanistic interpretability
Nanda, N., Chan, L., Lieberum, T., Smith, J., and Steinhardt, J · 2023
Closest in time.
Efficient training of language models using few-shot learning
Reddi, S. J., Miryoosefi, S., Karp, S., Krishnan, S., Kale, S., Kim, S., and Kumar, S · 2023
Closest in time.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Closest in time.
Uncovering mesa-optimization algorithms in transformers
von Oswald, J., Niklasson, E., Schlegel, M., Kobayashi, S., Zucchet, N., Scherrer, N., Miller, N., Sandler, M., Vladymyrov, M., Pascanu, R., et al · 2023
Closest in time.
Shall we pretrain autoregressive language models with retrieval? a comprehensive study
Wang, B., Ping, W., Xu, P., McAfee, L., Liu, Z., Shoeybi, M., Dong, Y., Kuchaiev, O., Li, B., Xiao, C., Anandkumar, A., and Catanzaro, B · 2023
Closest in time.
The learnability of in-context learning
Wies, N., Levine, Y., and Shashua, A · 2023
Closest in time.
What algorithms can transformers learn? a study in length generalization
Zhou, H., Bradley, A., Littwin, E., Razin, N., Saremi, O., Susskind, J., Bengio, S., and Nakkiran, P · 2023
Closest in time.