Fetching the paper…
Reading the bibliography…
The development of deep learning architectures is a resource-demanding process, due to a vast design space, long prototyping times, and high compute costs associated with at-scale model training and evaluation.
“Attention is all you need”
Ashish Vaswani et al · 2017
Earlier work this paper cites.
“Outrageously large neural networks: The sparsely-gated mixture-of-experts layer”
Noam Shazeer et al · 2017
Earlier work this paper cites.
“A structured self-attentive sentence embedding”
Zhouhan Lin et al · 2017
Earlier work this paper cites.
“On the practical computatifonal power of finite precision RNNs for language recognition”
Gail Weiss, Yoav Goldberg and Eran Yahav · 2018
Earlier work this paper cites.
“Augmented neural odes”
Emilien Dupont, Arnaud Doucet and Yee Teh · 2019
Earlier work this paper cites.
“Root mean square layer normalization”
Biao Zhang and Rico Sennrich · 2019
Earlier work this paper cites.
“Scaling laws for neural language models”
Jared Kaplan et al · 2020
Earlier work this paper cites.
“Transformers are rnns: Fast autoregressive transformers with linear attention”
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas and François Fleuret · 2020
Earlier work this paper cites.
“Glu variants improve transformer”
Noam Shazeer · 2020
Earlier work this paper cites.
“The pile: An 800gb dataset of diverse text for language modeling”
Leo Gao et al · 2020
Earlier work this paper cites.
“RNNs can generate bounded hierarchical languages with optimal memory”
John Hewitt et al · 2020
Earlier work this paper cites.
“Dissecting neural odes”
Stefano Massaroli et al · 2020
Earlier work this paper cites.
“Gshard: Scaling giant models with conditional computation and automatic sharding”
Dmitry Lepikhin et al · 2020
Earlier work this paper cites.
“Primer: Searching for efficient transformers for language modeling”
David So et al · 2021
Earlier work this paper cites.
“Linear transformers are secretly fast weight programmers”
Imanol Schlag, Kazuki Irie and Jürgen Schmidhuber · 2021
Earlier work this paper cites.
“A mathematical framework for transformer circuits”
Nelson Elhage et al · 2021
Cited alongside, same era.
“Efficiently modeling long sequences with structured state spaces”
Albert Gu, Karan Goel and Christopher Ré · 2021
Cited alongside, same era.
“Training compute-optimal large language models”
Jordan Hoffmann et al · 2022
Cited alongside, same era.
“In-context learning and induction heads”
Catherine Olsson et al · 2022
Cited alongside, same era.
“Hungry hungry hippos: Towards language modeling with state space models”
Daniel Fu et al · 2022
Cited alongside, same era.
“Mamba: Linear-time sequence modeling with selective state spaces”
Albert Gu and Tri Dao · 2023
Later among the works it cites.
“Gated Linear Attention Transformers with Hardware-Efficient Training”
Songlin Yang et al · 2023
Later among the works it cites.
“Understanding in-context learning in transformers and llms by learning to learn discrete functions”
Satwik Bhattamishra, Arkil Patel, Phil Blunsom and Varun Kanade · 2023
Later among the works it cites.
“Zoology: Measuring and Improving Recall in Efficient Language Models”
Simran Arora et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Mega: moving average equipped gated attention”
Xuezhe Ma et al · 2022
Cited alongside, same era.
“Transformer quality in linear time”
Weizhe Hua, Zihang Dai, Hanxiao Liu and Quoc Le · 2022
Cited alongside, same era.
“Neural architecture search: Insights from 1000 papers”
Colin White et al · 2023
Cited alongside, same era.
“Llama 2: Open foundation and fine-tuned chat models”
Hugo Touvron et al · 2023
Cited alongside, same era.
Albert Jiang et al · 2023
Cited alongside, same era.
“Dissecting recall of factual associations in auto-regressive language models”
Mor Geva, Jasmijn Bastings, Katja Filippova and Amir Globerson · 2023
Cited alongside, same era.
“RWKV: Reinventing RNNs for the Transformer Era”
Bo Peng et al · 2023
Cited alongside, same era.
Mahan Fathi et al · 2023
Later among the works it cites.
“The Languini Kitchen: Enabling Language Modelling Research at Different Scales of Compute”
Aleksandar Stanić et al · 2023
Later among the works it cites.
“Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws”
Nikhil Sardana and Jonathan Frankle · 2023
Later among the works it cites.
“Laughing Hyena Distillery: Extracting Compact Recurrences From Convolutions”
Stefano Massaroli et al · 2023
Later among the works it cites.
“Fact Finding: Attempting to Reverse-Engineer Factual Recall on the Neuron Level”
Neel Nanda, Senthooran Rajamanoharan, János Kramár and Rohin Shah · 2023
Later among the works it cites.
“Effectively modeling time series with simple discrete state spaces”
Michael Zhang et al · 2023
Later among the works it cites.
“In-Context Language Learning: Architectures and Algorithms”
Ekin Akyürek, Bailin Wang, Yoon Kim and Jacob Andreas · 2024
Closest in time.
“Sparse modular activation for efficient sequence modeling”
Liliang Ren et al · 2024
Closest in time.
“DeepSeek LLM: Scaling Open-Source Language Models with Longtermism”
Xiao Bi et al · 2024
Closest in time.
“Roformer: Enhanced transformer with rotary position embedding”
Jianlin Su et al · 2024
Closest in time.