Fetching the paper…
Reading the bibliography…
The Transformer architecture excels in a variety of language modeling tasks, outperforming traditional neural architectures such as RNN and LSTM.
Lower bounds for multiplication via network coding
Afshani, P.; Freksen, C. B.; Kamma, L.; and Larsen, K. G. 2019 · 1902
Earlier work this paper cites.
SciBERT: A pretrained language model for scientific text
Beltagy, I.; Lo, K.; and Cohan, A. 2019 · 1903
Earlier work this paper cites.
Zhou, S.; and Schoellig, A. P. 2019 · 1912
Earlier work this paper cites.
Large language models in medicine
Thirunavukarasu, A. J.; Ting, D. S. J.; Elangovan, K.; Gutierrez, L.; Tan, T. F.; and Ting, D. S. W. 2023 · 1940
Earlier work this paper cites.
Properties of predictors for autoregressive time series
Fuller, W. A.; and Hasza, D. P. 1981 · 1981
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
Cybenko, G. 1989 · 1989
Earlier work this paper cites.
The vanishing gradient problem during learning recurrent neural nets and problem solutions
Hochreiter, S. 1998 · 1998
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J.; McCandlish, S.; Henighan, T.; Brown, T. B.; Chess, B.; Child, R.; Gray, S.; Radford, A.; Wu, J.; and Amodei, D. 2020 · 2001
Earlier work this paper cites.
Recurrent neural networks
Medsker, L. R.; Jain, L.; et al. 2001 · 2001
Earlier work this paper cites.
Addressing some limitations of transformers with feedback memory
Fan, A.; Lavril, T.; Grave, E.; Joulin, A.; and Sukhbaatar, S. 2020 · 2002
Earlier work this paper cites.
Quantifying attention flow in transformers
Abnar, S.; and Zuidema, W. 2020 · 2005
Earlier work this paper cites.
Multilayer perceptron and neural networks
Popescu, M.-C.; Balas, V. E.; Perescu-Popescu, L.; and Mastorakis, N. 2009 · 2009
Earlier work this paper cites.
Inferring algorithmic patterns with stack-augmented recurrent nets
Joulin, A.; and Mikolov, T. 2015 · 2015
Earlier work this paper cites.
Preventing gradient explosions in gated recurrent units
Kanai, S.; Fujiwara, Y.; and Iwamura, S. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Dehghani, M.; Gouws, S.; Vinyals, O.; Uszkoreit, J.; and Kaiser, Ł. 2018 · 2018
Cited alongside, same era.
A review of weight optimization techniques in recurrent neural networks
Alqushaibi, A.; Abdulkadir, S. J.; Rais, H. M.; and Al-Tashi, Q. 2020 · 2020
Cited alongside, same era.
Transformers are rnns: Fast autoregressive transformers with linear attention
Katharopoulos, A.; Vyas, A.; Pappas, N.; and Fleuret, F. 2020 · 2020
Cited alongside, same era.
Beyond exploding and vanishing gradients: analysing RNN training using attractors and smoothness
Jiang, A. Q.; Sablayrolles, A.; Mensch, A.; Bamford, C.; Chaplot, D. S.; Casas, D. d. l.; Bressand, F.; Lengyel, G.; Lample, G.; Saulnier, L.; et al. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; et al. 2023 · 2023
Later among the works it cites.
Large language models fail on trivial alterations to theory-of-mind tasks
Ullman, T. 2023 · 2023
Later among the works it cites.
Bloomberggpt: A large language model for finance
Wu, S.; Irsoy, O.; Lu, S.; Dabravolski, V.; Dredze, M.; Gehrmann, S.; Kambadur, P.; Rosenberg, D.; and Mann, G. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ribeiro, A. H.; Tiels, K.; Aguirre, L. A.; and Schön, T. 2020 · 2020
Cited alongside, same era.
Large language models still can’t plan (a benchmark for LLMs on planning and reasoning about change)
Valmeekam, K.; Olmo, A.; Sreedharan, S.; and Kambhampati, S. 2022 · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Xia, F.; Chi, E.; Le, Q. V.; Zhou, D.; et al. 2022 · 2022
Cited alongside, same era.
Recurring the transformer for video action recognition
Yang, J.; Dong, X.; Liu, L.; Zhang, C.; Shen, J.; and Yu, D. 2022 · 2022
Cited alongside, same era.
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
Cited alongside, same era.
Tighter Bounds on the Expressivity of Transformer Encoders, May 2023
Chiang, D.; Cholak, P.; and Pillay, A. 2023 · 2023
Cited alongside, same era.
A survey of chain of thought reasoning: Advances, frontiers and future
Chu, Z.; Chen, J.; Chen, Q.; Yu, W.; He, T.; Wang, H.; Peng, W.; Liu, M.; Qin, B.; and Liu, T. 2023 · 2023
Cited alongside, same era.
Zhang, X.; Li, S.; Hauer, B.; Shi, N.; and Kondrak, G. 2023 · 2023
Later among the works it cites.
Physics of language models: Part 3.1, knowledge storage and extraction
Zhu, Z. A.; and Li, Y. 2023 · 2023
Later among the works it cites.
A survey on evaluation of large language models
Chang, Y.; Wang, X.; Wang, J.; Wu, Y.; Yang, L.; Zhu, K.; Chen, H.; Yi, X.; Wang, C.; Wang, Y.; et al. 2024 · 2024
Closest in time.
Faith and fate: Limits of transformers on compositionality
Dziri, N.; Lu, X.; Sclar, M.; Li, X. L.; Jiang, L.; Lin, B. Y.; Welleck, S.; West, P.; Bhagavatula, C.; Le Bras, R.; et al. 2024 · 2024
Closest in time.
Chain of thought empowers transformers to solve inherently serial problems
Li, Z.; Liu, H.; Zhou, D.; and Ma, T. 2024 · 2024
Closest in time.
Prollama: A protein large language model for multi-task protein language processing
Lv, L.; Lin, Z.; Li, H.; Liu, Y.; Cui, J.; Chen, C. Y.-C.; Yuan, L.; and Tian, Y. 2024 · 2024
Closest in time.
Chain of thought utilization in large language models and application in nephrology
Miao, J.; Thongprayoon, C.; Suppadungsuk, S.; Krisanapan, P.; Radhakrishnan, Y.; and Cheungpasitporn, W. 2024 · 2024
Closest in time.
Lower Bounds on the Expressivity of Recurrent Neural Language Models
Svete, A.; Nowak, F.; Sahabdeen, A. M.; and Cotterell, R. 2024 · 2024
Closest in time.
Large language model as attributed training data generator: A tale of diversity and bias
Yu, Y.; Zhuang, Y.; Zhang, J.; Meng, Y.; Ratner, A. J.; Krishna, R.; Shen, J.; and Zhang, C. 2024 · 2024
Closest in time.