Fetching the paper…
Reading the bibliography…
In the 1990s, the constant error carousel and gating were introduced as the central ideas of the Long Short-Term Memory (LSTM).
Geometrical and statistical properties of systems of linear inequalities with applications in pattern recognition
T. M. Cover · 1965
Earlier work this paper cites.
A simple neural network generating an interactive memory
J. A. Anderson · 1972
Earlier work this paper cites.
Correlation matrix memories
T. Kohonen · 1972
Earlier work this paper cites.
Associatron – a model of associative memory
K. Nakano · 1972
Earlier work this paper cites.
Distinctive features, categorical perception, and probability learning: Some applications of a neural model
J. Anderson, J. Silverstein, S. Ritz, and R. Jones · 1977
Earlier work this paper cites.
Storing covariance with nonlinearly interacting neurons
T. J. Sejnowski · 1977
Earlier work this paper cites.
Optimising synaptic learning rules in linear associative memories
P. Dayan and D. J. Willshaw · 1991
Earlier work this paper cites.
Untersuchungen zu dynamischen neuronalen Netzen
S. Hochreiter · 1991
Earlier work this paper cites.
Learning to control fast-weight memories: An alternative to recurrent nets
J. Schmidhuber · 1992
Earlier work this paper cites.
The international corpus of English (ICE) project
S. Greenbaum and G. Nelson · 1996
Earlier work this paper cites.
Learning to forget: Continual prediction with LSTM
F. A. Gers, J. Schmidhuber, and F. Cummins · 2000
Earlier work this paper cites.
Gradient flow in recurrent nets: the difficulty of learning long-term dependencies
S. Hochreiter, Y. Bengio, P. Frasconi, and J. Schmidhuber · 2000
Earlier work this paper cites.
Learning to learn using gradient descent
S. Hochreiter, A. Steven Younger, and Peter R. Conwell · 2001
Earlier work this paper cites.
Fast model-based protein homology detection without alignment
S. Hochreiter, M. Heusel, and K. Obermayer · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
The ACL anthology network corpus
D. R. Radev, P. Muthukrishnan, and V. Qazvinian · 2009
Earlier work this paper cites.
Learning word vectors for sentiment analysis
A. L. Maas, R. E. Daly, P. T. Pham, D. Huang, A. Y. Ng, and C. Potts · 2011
Earlier work this paper cites.
Parsing noun phrases in the Penn Treebank
D. Vadas and J. R. Curran · 2011
Earlier work this paper cites.
Generating sequences with recurrent neural networks
A. Graves · 2013
Earlier work this paper cites.
Deep learning in neural networks: An overview
J. Schmidhuber · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. V. Le · 2014
Earlier work this paper cites.
Learning to execute
W. Zaremba and I. Sutskever · 2014
Earlier work this paper cites.
LSTM: A search space odyssey
K. Greff, R. K. Srivastava, J. Koutník, B. R. Steunebrink, and J. Schmidhuber · 2015
Earlier work this paper cites.
The unreasonable effectiveness of recurrent neural networks
A. Karpathy · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Earlier work this paper cites.
Training very deep networks
R. K. Srivastava, K. Greff, and J. Schmidhuber · 2015
Earlier work this paper cites.
Demographic dialectal variation in social media: A case study of African-American English
S. L. Blodgett, L. Green, and B. O’Connor · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Dense associative memory for pattern recognition
D. Krotov and J. J. Hopfield · 2016
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context
D. Paperno, G. Kruszewski, A. Lazaridou, N.-Q. Pham, R. Bernardi, S. Pezzelle, M. Baroni, Gemma G. Boleda, and R. Fernández · 2016
Earlier work this paper cites.
Dense associative memory is robust to adversarial inputs
D. Krotov and J. J. Hopfield · 2017
Earlier work this paper cites.
Pointer sentinel mixture models
S. Merity, C. Xiong, J. Bradbury, and R. Socher · 2017
Earlier work this paper cites.
Learning to generate reviews and discovering sentiment
A. Radford, R. Jozefowicz, and I. Sutskever · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Starcraft II: A new challenge for reinforcement learning
O. Vinyals, T. Ewalds, S. Bartunov, et al · 2017
Earlier work this paper cites.
Think you have solved question answering? Try ARC, the AI2 reasoning challenge
P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord · 2018
Earlier work this paper cites.
Rainfall-runoff modelling using long short-term memory (LSTM) networks
F. Kratzert, D. Klotz, C. Brenner, K. Schulz, and M. Herrnegger · 2018
Earlier work this paper cites.
Learning long-range spatial dependencies with horizontal gated recurrent units
D. Linsley, J. Kim, V. Veerabadran, C. Windolf, and T. Serre · 2018
Earlier work this paper cites.
Online normalizer calculation for softmax
M. Milakov and N. Gimelshein · 2018
Earlier work this paper cites.
Group normalization
Y. Wu and K. He · 2018
Earlier work this paper cites.
What is Gab: A bastion of free speech or an alt-right echo chamber
S. Zannettou, B. Bradlyn, E. DeCristofaro, H. Kwak, M. Sirivianos, G. Stringini, and J. Blackburn · 2018
Earlier work this paper cites.
Identifying and controlling important neurons in neural machine translation
A. Bau, Y. Belinkov, H. Sajjad, N. Durrani, F. Dalvi, and J. Glass · 2019
Cited alongside, same era.
A comprehensive survey of deep learning for image captioning
M. D. Hossain, F. Sohel, M. F. Shiratuddin, and H. Laga · 2019
Cited alongside, same era.
OpenAI Five defeats Dota 2 world champions
A. Karpathy · 2019
Cited alongside, same era.
Benchmarking a catchment-aware long short-term memory network (LSTM) for large-scale hydrological modeling
F. Kratzert, D. Klotz, G. Shalev, G. Klambauer, S. Hochreiter, and G. Nearing · 2019
Cited alongside, same era.
The emergence of number and syntax units in LSTM language models
Y. Lakretz, G. Kruszewski, T. Desbordes, D. Hupkes, S. Dehaene, and M. Baroni · 2019
Cited alongside, same era.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2019
M2D2: A massively multi-domain language modeling dataset
M. Reid, V. Zhong, S. Gururangan, and L. Zettlemoyer · 2022
Later among the works it cites.
BLOOM: A 176B-parameter open-access multilingual language model
T. L. Scao, A. Fan, C. Akiki, et al · 2022
Later among the works it cites.
ChatGPT: Optimizing language models for dialogue
J. Schulman, B. Zoph, C. Kim, J. Hilton, et al · 2022
Later among the works it cites.
Simplified state space layers for sequence modeling
J. T. H. Smith, A. Warrington, and S. W. Linderman · 2022
Later among the works it cites.
AlexaTM 20B: Few-shot learning using a large-scale multilingual Seq2Seq model
S. Soltan, S. Ananthakrishnan, J. FitzGerald, R. Gupta, W. Hamza, H. Khan, C. Peris, S. Rawls, A. Rosenbaum, A. Rumshisky, C. S. Prakash, M. Sridhar, F. Triefenbach, A. Verma, G. Tur, and P. Natarajan · 2022
Later among the works it cites.
LaMDA: Language models for dialog applications
R. Thoppilan, D. deFreitas, J. Hall, et al · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2019
Cited alongside, same era.
Megatron-LM: Training multi-billion parameter language models using model parallelism
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro · 2019
Cited alongside, same era.
HellaSwag: Can a machine really finish your sentence?
R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi · 2019
Cited alongside, same era.
Piqa: Reasoning about physical commonsense in natural language
Y. Bisk, R. Zellers, R. LeBras, J. Gao, and Y. Choi · 2020
Cited alongside, same era.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, et al · 2020
Cited alongside, same era.
Later among the works it cites.
Pretraining without attention
J. Wang, J. N. Yan, A. Gu, and A. M. Rush · 2022
Later among the works it cites.
GLM-130B: An open bilingual pre-trained model
A. Zeng, X. Liu, Z. Du, et al · 2022
Later among the works it cites.
OPT: Open pre-trained transformer language models
S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V. Lin, T. Mihaylov, M. Ott, S. Shleifer, K. Shuster, D. Simig, P. S. Koura, A. Sridhar, T. Wang, and L. Zettlemoyer · 2022
Later among the works it cites.
GPT-4 technical report
J. Achiam, S. Adler, S. Agarwal, et al · 2023
Later among the works it cites.
Zoology: Measuring and improving recall in efficient language models
S. Arora, S. Eyuboglu, A. Timalsina, I. Johnson, M. Poli, J. Zou, A. Rudra, and C. Ré · 2023
Later among the works it cites.
Neural networks and the Chomsky hierarchy
G. Delétang, A. Ruoss, J. Grau-Moya, T. Genewein, L. K. Wenliang, E. Catt, C. Cundy, M. Hutter, S. Legg, J. Veness, and P. A. Ortega · 2023
Later among the works it cites.
Hungry hungry hippos: Towards language modeling with state space models
D. Y. Fu, T. Dao, K. K. Saab, A. W. Thomas, A. Rudra, and C. Re · 2023
Later among the works it cites.
Gemini: A family of highly capable multimodal models
Gemini Team Google · 2023
Later among the works it cites.
Mamba: Linear-time sequence modeling with selective state spaces
A. Gu and T. Dao · 2023
Later among the works it cites.
GateLoop: Fully data-controlled linear recurrence for sequence modeling
T. Katsch · 2023
Later among the works it cites.
Holistic evaluation of language models
P. Liang, R. Bommasani, T. Lee, et al · 2023
Later among the works it cites.
Paloma: A benchmark for evaluating language model fit
I. Magnusson, A. Bhagia, V. Hofmann, et al · 2023
Later among the works it cites.
The parallelism tradeoff: Limitations of log-precision transformers
W. Merrill and A. Sabharwal · 2023
Later among the works it cites.
Resurrecting recurrent neural networks for long sequences
A. Orvieto, S. L. Smith, A. Gu, A. Fernando, C. Gulcehre, R. Pascanu, and S. De · 2023
Later among the works it cites.
The RefinedWeb dataset for Falcon LLM: Outperforming curated corpora with web data, and web data only
G. Penedo, Q. Malartic, D. Hesslow, R. Cojocaru, A. Cappelli, H. Alobeidli, B. Pannier, E. Almazrouei, and J. Launay · 2023
Later among the works it cites.
RWKV: Reinventing RNNs for the transformer era
B. Peng, E. Alcaide, Q. Anthony, et al · 2023
Later among the works it cites.
Hyena hierarchy: Towards larger convolutional language models
M. Poli, S. Massaroli, E. Nguyen, D. Y. Fu, T. Dao, S. Baccus, Y. Bengio, S. Ermon, and C. Ré · 2023
Later among the works it cites.
Hierarchically gated recurrent neural network for sequence modeling
Z. Qin, S. Yang, and Y. Zhong · 2023
Later among the works it cites.
SlimPajama: A 627B token cleaned and deduplicated version of RedPajama
D. Soboleva, F. Al-Khateeb, R. Myers, J. R. Steeves, J. Hestness, and N. Dey · 2023
Later among the works it cites.
Dolma: an open corpus of three trillion tokens for language model pretraining research
L. Soldaini, R. Kinney, A. Bhagia, et al · 2023
Later among the works it cites.
Retentive network: A successor to transformer for large language models
Y. Sun, L. Dong, S. Huang, S. Ma, Y. Xia, J. Xue, J. Wang, and F. Wei · 2023
Later among the works it cites.
EleutherAI/lm-evaluation-harness: Major refactor, 2023
L. Sutawika, L. Gao, H. Schoelkopf, et al · 2023
Later among the works it cites.
Redpajama: an open dataset for training large language models, 2023
TogetherComputer · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample · 2023
Later among the works it cites.
Gated linear attention transformers with hardware-efficient training
S. Yang, B. Wang, Y. Shen, R. Panda, and Y. Kim · 2023
Later among the works it cites.
Flashattention-2: Faster attention with better parallelism and work partitioning
T. Dao · 2024
Closest in time.
Griffin: Mixing gated linear recurrences with local attention for efficient language models
S. De, S. L. Smith, A. Fernando, A. Botev, G. Cristian-Muraru, A. Gu, R. Haroun, L. Berrada, Y. Chen, S. Srinivasan, G. Desjardins, A. Doucet, D. Budden, Y. W. Teh, R. Pascanu, N. DeFreitas, and C. Gulcehre · 2024
Closest in time.
The illusion of state in state-space models
W. Merrill, J. Petty, and A. Sabharwal · 2024
Closest in time.
Global prediction of extreme floods in ungauged watersheds
G. Nearing, D. Cohen, V. Dube, M. Gauch, O. Gilon, S. Harrigan, A. Hassidim, D. Klotz, F. Kratzert, A. Metzger, S. Nevo, F. Pappenberger, C. Prudhomme, G. Shalev, S. Shenzis, T. Y. Tekalign, D. Weitzner, and Y. M. B. Kosko · 2024
Closest in time.
Eagle and Finch: RWKV with matrix-valued states and dynamic recurrence
B. Peng, D. Goldstein, Q. Anthony, et al · 2024
Closest in time.
Mechanistic design and scaling of hybrid architectures
M. Poli, A. W. Thomas, E. Nguyen, P. Ponnusamy, B. Deiseroth, K. Kersting, T. Suzuki, B. Hie, S. Ermon, C. Ré, C. Zhang, and S. Massaroli · 2024
Closest in time.
HGRN2: Gated linear RNNs with state expansion
Z. Qin, S. Yang, W. Sun, X. Shen, D. Li, W. Sun, and Y. Zhong · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
M. Reid, N. Savinov, D. Teplyashin, et al · 2024
Closest in time.
EleutherAI/lm-evaluation-harness, 2024
L. Sutawika, H. Schoelkopf, L. Gao, B. Abbasi, S. Biderman, J. Tow, B. fattori, C. Lovering, farzanehnakhaee70, J. Phang, A. Thite, Fazz, T. Wang, N. Muennighoff, Aflah, sdtblck, nopperl, gakada, tttyuntian, researcher2, Chris, J. Etxaniz, H. A. Lee, Z. Kasner, Khalid, J. Hsu, A. Kanekar, P. S. Ammanamanchi, V. Boykis, and AndyZwei · 2024
Closest in time.
FLA: A Triton-based library for hardware-efficient implementations of linear attention mechanism, 2024
S. Yang and Y. Zhang · 2024
Closest in time.