Fetching the paper…
Reading the bibliography…
Large-scale neural language models exhibit a remarkable capacity for in-context learning (ICL): they can infer novel functions from datasets provided as input.
Memory-augmented recurrent neural networks can learn generalized Dyck languages
Mirac Suzgun, Sebastian Gehrmann, Yonatan Belinkov, and Stuart M Shieber · 1911
Earlier work this paper cites.
Language identification in the limit
E Mark Gold · 1967
Earlier work this paper cites.
An n log n algorithm for minimizing states in a finite automaton
John Hopcroft · 1971
Earlier work this paper cites.
Maximum likelihood from incomplete data via the EM algorithm
Arthur P Dempster, Nan M Laird, and Donald B Rubin · 1977
Earlier work this paper cites.
A theory of the learnable
Leslie G Valiant · 1984
Earlier work this paper cites.
Identifying languages from stochastic examples
Dana Angluin · 1988
Earlier work this paper cites.
Probabilistic inductive inference
Leonard Pitt · 1989
Earlier work this paper cites.
A tutorial on hidden markov models and selected applications in speech recognition
Lawrence R Rabiner · 1989
Earlier work this paper cites.
Finding structure in time
Jeffrey L Elman · 1990
Earlier work this paper cites.
An empirical study of smoothing techniques for language modeling
Stanley F. Chen and Joshua Goodman · 1996
Earlier work this paper cites.
Long Short-Term Memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
LSTM recurrent networks learn simple context-free and context-sensitive languages
Felix A Gers and E Schmidhuber · 2001
Earlier work this paper cites.
Glu variants improve transformer
Noam Shazeer · 2002
Earlier work this paper cites.
Links between probabilistic automata and hidden Markov models: Probability distributions, learning models and induction algorithms
P. Dupont, F. Denis, and Y. Esposito · 2004
Earlier work this paper cites.
Faster and smaller n-gram language models
Adam Pauls and Dan Klein · 2011
Earlier work this paper cites.
Does string-based neural MT learn source syntax?
Xing Shi, Inkit Padhi, and Kevin Knight · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Sequential neural networks as automata
William Merrill · 2019
Earlier work this paper cites.
Root mean square layer normalization
Biao Zhang and Rico Sennrich · 2019
Earlier work this paper cites.
On the ability and limitations of Transformers to recognize formal languages
Satwik Bhattamishra, Kabir Ahuja, and Navin Goyal · 2020
Earlier work this paper cites.
RNNs can generate bounded hierarchical languages with optimal memory
John Hewitt, Michael Hahn, Surya Ganguli, Percy Liang, and Christopher D. Manning · 2020
Cited alongside, same era.
Transformers are RNNs: Fast autoregressive Transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret · 2020
Cited alongside, same era.
On the linguistic capacity of real-time counter automata
William Merrill · 2020
Cited alongside, same era.
An interpretability illusion for BERT
Tolga Bolukbasi, Adam Pearce, Ann Yuan, Andy Coenen, Emily Reif, Fernanda Viégas, and Martin Wattenberg · 2021
Cited alongside, same era.
A mathematical framework for Transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, et al · 2021
Cited alongside, same era.
Zoology: Measuring and improving recall in efficient language models
Simran Arora, Sabri Eyuboglu, Aman Timalsina, Isys Johnson, Michael Poli, James Zou, Atri Rudra, and Christopher Ré · 2023
Later among the works it cites.
Understanding in-context learning in Transformers and LLMs by learning to learn discrete functions
Satwik Bhattamishra, Arkil Patel, Phil Blunsom, and Varun Kanade · 2023
Later among the works it cites.
Why can GPT learn in-context? Language models secretly perform gradient descent as meta-optimizers
Damai Dai, Yutao Sun, Li Dong, Yaru Hao, Shuming Ma, Zhifang Sui, and Furu Wei · 2023
Later among the works it cites.
Compositional semantic parsing with large language models
Andrew Drozdov, Nathanael Schärli, Ekin Akyürek, Nathan Scales, Xinying Song, Xinyun Chen, Olivier Bousquet, and Denny Zhou · 2023
Later among the works it cites.
Hungry hungry hippos: Towards language modeling with state space models
Daniel Y Fu, Tri Dao, Khaled Kamal Saab, Armin W Thomas, Atri Rudra, and Christopher Re · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shuangfei Zhai, Walter Talbott, Nitish Srivastava, Chen Huang, Hanlin Goh, Ruixiang Zhang, and Josh Susskind · 2021
Cited alongside, same era.
Probing classifiers: Promises, shortcomings, and advances
Yonatan Belinkov · 2022
Cited alongside, same era.
Data distributional properties drive emergent in-context learning in Transformers
Stephanie Chan, Adam Santoro, Andrew Lampinen, Jane Wang, Aaditya Singh, Pierre Richemond, James McClelland, and Felix Hill · 2022
Cited alongside, same era.
What makes instruction learning hard? An investigation and a new challenge in a synthetic environment
Matthew Finlayson, Kyle Richardson, Ashish Sabharwal, and Peter Clark · 2022
Cited alongside, same era.
What can Transformers learn in-context? A case study of simple function classes
Shivam Garg, Dimitris Tsipras, Percy S Liang, and Gregory Valiant · 2022
Cited alongside, same era.
Efficiently modeling long sequences with structured state spaces
Albert Gu, Karan Goel, and Christopher Ré · 2022
Cited alongside, same era.
Transformer quality in linear time
Weizhe Hua, Zihang Dai, Hanxiao Liu, and Quoc V. Le · 2022
Cited alongside, same era.
Later among the works it cites.
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao · 2023
Later among the works it cites.
A theory of emergent in-context learning as implicit structure induction
Michael Hahn and Navin Goyal · 2023
Later among the works it cites.
Exploring the relationship between model architecture and in-context learning ability
Ivan Lee, Nan Jiang, and Taylor Berg-Kirkpatrick · 2023
Later among the works it cites.
The parallelism tradeoff: Limitations of log-precision transformers
William Merrill and Ashish Sabharwal · 2023
Later among the works it cites.
RWKV: Reinventing RNNs for the transformer era
Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Huanqi Cao, Xin Cheng, Michael Chung, Matteo Grella, Kranthi Kiran GV, et al · 2023
Later among the works it cites.
Hyena hierarchy: Towards larger convolutional language models
Michael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y Fu, Tri Dao, Stephen Baccus, Yoshua Bengio, Stefano Ermon, and Christopher Ré · 2023
Later among the works it cites.
SlimPajama: A 627B token cleaned and deduplicated version of RedPajama
Daria Soboleva, Faisal Al-Khateeb, Robert Myers, Jacob R Steeves, Joel Hestness, and Nolan Dey · 2023
Later among the works it cites.
Retentive network: A successor to transformer for large language models
Yutao Sun, Li Dong, Shaohan Huang, Shuming Ma, Yuqing Xia, Jilong Xue, Jianyong Wang, and Furu Wei · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
MLRegTest: A benchmark for the machine learning of regular languages
Sam van der Poel, Dakotah Lambert, Kalina Kostyszyn, Tiantian Gao, Rahul Verma, Derek Andersen, Joanne Chau, Emily Peterson, Cody St Clair, Paul Fodor, et al · 2023
Later among the works it cites.
Larger language models do in-context learning differently
Jerry Wei, Jason Wei, Yi Tay, Dustin Tran, Albert Webson, Yifeng Lu, Xinyun Chen, Hanxiao Liu, Da Huang, Denny Zhou, et al · 2023
Later among the works it cites.
Transformers are uninterpretable with myopic methods: a case study with bounded Dyck grammars
Kaiyue Wen, Yuchen Li, Bingbin Liu, and Andrej Risteski · 2023
Later among the works it cites.
Gated linear attention Transformers with hardware-efficient training
Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda, and Yoon Kim · 2023
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu · 2024
Closest in time.