Fetching the paper…
Reading the bibliography…
Designing efficient and effective architectural backbones has been in the core of research efforts to enhance the capability of foundation models.
“On information-type measure of difference of probability distributions and indirect observations”
Imre Csiszar · 1967
Earlier work this paper cites.
“Ridge regression: applications to nonorthogonal problems”
Arthur Hoerl and Robert Kennard · 1970
Earlier work this paper cites.
“Ridge, a computer program for calculating ridge regression estimates”
Donald Hilt and Donald Seegrist · 1977
Earlier work this paper cites.
“Neural networks and physical systems with emergent collective computational abilities.”
John Hopfield · 1982
Earlier work this paper cites.
“Neural network capacity using delta rule”
DL Prados and SC Kak · 1989
Earlier work this paper cites.
“Local learning algorithms”
Leon Bottou and Vladimir Vapnik · 1992
Earlier work this paper cites.
“Robust estimation of a location parameter”
Peter Huber · 1992
Earlier work this paper cites.
“Learning to control fast-weight memories: An alternative to recurrent nets. Accepted for publication in”
JH Schmidhuber · 1992
Earlier work this paper cites.
“The reality of repressed memories.”
Elizabeth Loftus · 1993
Earlier work this paper cites.
“Reducing the ratio between learning complexity and number of time varying variables in fully recurrent nets”
Jürgen Schmidhuber · 1993
Earlier work this paper cites.
“Regression shrinkage and selection via the lasso”
Robert Tibshirani · 1996
Earlier work this paper cites.
“Long Short-term Memory”
Jürgen Schmidhuber and Sepp Hochreiter · 1997
Earlier work this paper cites.
“Learning and memory”
Hideyuki Okano, Tomoo Hirano and Evan Balaban · 2000
Earlier work this paper cites.
“Memory and the brain”
Lee Robertson · 2002
Earlier work this paper cites.
“The organization of behavior: A neuropsychological theory”
Donald Hebb · 2005
Earlier work this paper cites.
“SVM-KNN: Discriminative nearest neighbor classification for visual category recognition”
Hao Zhang, Alexander Berg, Michael Maire and Jitendra Malik · 2006
Earlier work this paper cites.
“The elements of statistical learning”
Trevor Hastie, Robert Tibshirani and Jerome Friedman · 2009
Earlier work this paper cites.
“Online domain adaptation of a pre-trained cascade of classifiers”
Vidit Jain and Erik Learned-Miller · 2011
Earlier work this paper cites.
“Online learning and online convex optimization”
Shai Shalev-Shwartz · 2012
Earlier work this paper cites.
“A unified convergence analysis of block successive minimization methods for nonsmooth optimization”
Meisam Razaviyayn, Mingyi Hong and Zhi-Quan Luo · 2013
Earlier work this paper cites.
“Incremental majorization-minimization optimization with application to large-scale machine learning”
Julien Mairal · 2015
Earlier work this paper cites.
“LSTM: A search space odyssey”
Klaus Greff, Rupesh Srivastava, Jan Koutnk, Bas Steunebrink and Jürgen Schmidhuber · 2016
Earlier work this paper cites.
“Introduction to online convex optimization”
Elad Hazan · 2016
Earlier work this paper cites.
“Gaussian error linear units (gelus)”
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
“Dense associative memory for pattern recognition”
Dmitry Krotov and John Hopfield · 2016
Earlier work this paper cites.
“The LAMBADA dataset: Word prediction requiring a broad discourse context”
Denis Paperno, German Kruszewski, Angeliki Lazaridou, Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda and Raquel Fernandez · 2016
Earlier work this paper cites.
“Pointer Sentinel Mixture Models”
Stephen Merity, Caiming Xiong, James Bradbury and Richard Socher · 2017
Earlier work this paper cites.
“Neural semantic encoders”
Tsendsuren Munkhdalai and Hong Yu · 2017
Earlier work this paper cites.
“Delta networks for optimized recurrent network computation”
Daniel Neil, Jun Lee, Tobi Delbruck and Shih-Chii Liu · 2017
Earlier work this paper cites.
“Learning and memory: Basic principles, processes, and procedures”
W Terry · 2017
Earlier work this paper cites.
“Attention is All you Need”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, Łukasz Kaiser and Illia Polosukhin · 2017
Earlier work this paper cites.
“Attention is All you Need”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, Łukasz Kaiser and Illia Polosukhin · 2017
Earlier work this paper cites.
“Think you have solved question answering? try arc, the ai2 reasoning challenge”
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick and Oyvind Tafjord · 2018
Earlier work this paper cites.
“BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions”
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins and Kristina Toutanova · 2019
Cited alongside, same era.
“Online model distillation for efficient video inference”
Ravi Mullapudi, Steven Chen, Keyi Zhang, Deva Ramanan and Kayvon Fatahalian · 2019
Cited alongside, same era.
“Metalearned neural memory”
Tsendsuren Munkhdalai, Alessandro Sordoni, Tong Wang and Adam Trischler · 2019
Cited alongside, same era.
“HellaSwag: Can a Machine Really Finish Your Sentence?”
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi and Yejin Choi · 2019
Cited alongside, same era.
“Root mean square layer normalization”
Biao Zhang and Rico Sennrich · 2019
Cited alongside, same era.
“Piqa: Reasoning about physical commonsense in natural language”
Yonatan Bisk, Rowan Zellers, Jianfeng Gao and Yejin Choi · 2020
“Griffin: Mixing gated linear recurrences with local attention for efficient language models”
Soham De, Samuel Smith, Anushan Fernando, Aleksandar Botev, George Cristian-Muraru, Albert Gu, Ruba Haroun, Leonard Berrada, Yutian Chen and Srivatsan Srinivasan · 2024
Later among the works it cites.
“Towards scalable and stable parallelization of nonlinear rnns”
Xavier Gonzalez, Andrew Warrington, Jimmy Smith and Scott Linderman · 2024
Later among the works it cites.
“Unlocking state-tracking in linear rnns through negative eigenvalues”
Riccardo Grazzi, Julien Siems, Jörg Franke, Arber Zela, Frank Hutter and Massimiliano Pontil · 2024
Later among the works it cites.
“Mamba: Linear-Time Sequence Modeling with Selective State Spaces”
Albert Gu and Tri Dao · 2024
Later among the works it cites.
“RULER: What’s the Real Context Size of Your Long-Context Language Models?”
Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia and Boris Ginsburg · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“An image is worth 16x16 words: Transformers for image recognition at scale”
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold and Sylvain Gelly · 2020
Cited alongside, same era.
“Scaling laws for neural language models”
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu and Dario Amodei · 2020
Cited alongside, same era.
“Transformers are rnns: Fast autoregressive transformers with linear attention”
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas and François Fleuret · 2020
Cited alongside, same era.
“Exploring the limits of transfer learning with a unified text-to-text transformer”
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li and Peter Liu · 2020
Cited alongside, same era.
“Going beyond linear transformers with recurrent fast weight programmers”
Kazuki Irie, Imanol Schlag, Robert Csordas and Jurgen Schmidhuber · 2021
Cited alongside, same era.
“Hierarchical associative memory”
Dmitry Krotov · 2021
Cited alongside, same era.
Later among the works it cites.
Jerry-Chieh Hu, Dennis Wu and Han Liu · 2024
Later among the works it cites.
“A survey on long video generation: Challenges, methods, and prospects”
Chengxuan Li, Di Huang, Zeyu Lu, Yang Xiao, Qingqi Pei and Lei Bai · 2024
Later among the works it cites.
“On the expressive power of modern hopfield networks”
Xiaoyu Li, Yuanpeng Li, Yingyu Liang, Zhenmei Shi and Zhao Song · 2024
Later among the works it cites.
“Parallelizing non-linear sequential models over the sequence length”
Yi Lim, Qi Zhu, Joshua Selfridge and Muhammad Kasim · 2024
Later among the works it cites.
“Longhorn: State space models are amortized online learners”
Bo Liu, Rui Wang, Lemeng Wu, Yihao Feng, Peter Stone and Qiang Liu · 2024
Later among the works it cites.
“Lost in the middle: How language models use long contexts”
Nelson Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni and Percy Liang · 2024
Later among the works it cites.
“Exponential capacity of dense associative memories”
Carlo Lucibello and Marc Mézard · 2024
Later among the works it cites.
“The Illusion of State in State-Space Models”
William Merrill, Jackson Petty and Ashish Sabharwal · 2024
Later among the works it cites.
“The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale”
Guilherme Penedo, Hynek Kydlcek, Loubna allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Werra and Thomas Wolf · 2024
Later among the works it cites.
“Eagle and finch: Rwkv with matrix-valued states and dynamic recurrence”
Bo Peng, Daniel Goldstein, Quentin Anthony, Alon Albalak, Eric Alcaide, Stella Biderman, Eugene Cheah, Xingjian Du, Teddy Ferdinan and Haowen Hou · 2024
Later among the works it cites.
“HGRN2: Gated Linear RNNs with State Expansion”
Zhen Qin, Songlin Yang, Weixuan Sun, Xuyang Shen, Dong Li, Weigao Sun and Yiran Zhong · 2024
Later among the works it cites.
“Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling”
Liliang Ren, Yang Liu, Yadong Lu, Yelong Shen, Chen Liang and Weizhu Chen · 2024
Later among the works it cites.
“Roformer: Enhanced transformer with rotary position embedding”
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo and Yunfeng Liu · 2024
Later among the works it cites.
“Learning to (learn at test time): Rnns with expressive hidden states”
Yu Sun, Xinhao Li, Karan Dalal, Jiarui Xu, Arjun Vikram, Genghan Zhang, Yann Dubois, Xinlei Chen, Xiaolong Wang and Sanmi Koyejo · 2024
Later among the works it cites.
Matteo Tiezzi, Michele Casoni, Alessandro Betti, Tommaso Guidi, Marco Gori and Stefano Melacci · 2024
Later among the works it cites.
“Long-context Protein Language Model”
Yingheng Wang, Zichen Wang, Gil Sadeh, Luca Zancato, Alessandro Achille, George Karypis and Huzefa Rangwala · 2024
Later among the works it cites.
“Gated Delta Networks: Improving Mamba2 with Delta Rule”
Songlin Yang, Jan Kautz and Ali Hatamizadeh · 2024
Later among the works it cites.
“Gated Linear Attention Transformers with Hardware-Efficient Training”
Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda and Yoon Kim · 2024
Later among the works it cites.
“Parallelizing linear transformers with the delta rule over sequence length”
Songlin Yang, Bailin Wang, Yu Zhang, Yikang Shen and Yoon Kim · 2024
Later among the works it cites.
“One-Minute Video Generation with Test-Time Training”
Karan Dalal, Daniel Koceja, Gashon Hussein, Jiarui Xu, Yue Zhao, Youjin Song, Shihao Han, Ka Cheung, Jan Kautz and Carlos Guestrin · 2025
Closest in time.
“Lattice: Learning to Efficiently Compress the Memory”
M. Karami and V. Mirrokni · 2025
Closest in time.
“Minimax-01: Scaling foundation models with lightning attention”
Aonian Li, Bangwei Gong, Bo Yang, Boji Shan, Chang Liu, Cheng Zhu, Chunhao Zhang, Congchao Guo, Da Chen and Dong Li · 2025
Closest in time.
“RWKV-7" Goose" with Expressive Dynamic State Evolution”
Bo Peng, Ruichong Zhang, Daniel Goldstein, Eric Alcaide, Haowen Hou, Janna Lu, William Merrill, Guangyu Song, Kaifeng Tan and Saiteja Utpala · 2025
Closest in time.
“Rwkv-7" goose" with expressive dynamic state evolution”
Bo Peng, Ruichong Zhang, Daniel Goldstein, Eric Alcaide, Haowen Hou, Janna Lu, William Merrill, Guangyu Song, Kaifeng Tan and Saiteja Utpala · 2025
Closest in time.
“Information theory: From coding to learning”
Yury Polyanskiy and Yihong Wu · 2025
Closest in time.
“Implicit Language Models are RNNs: Balancing Parallelization and Expressivity”
Mark Schöne, Babak Rahmani, Heiner Kremer, Fabian Falck, Hitesh Ballani and Jannes Gladrow · 2025
Closest in time.
“DeltaProduct: Increasing the Expressivity of DeltaNet Through Products of Householders”
Julien Siems, Timur Carstensen, Arber Zela, Frank Hutter, Massimiliano Pontil and Riccardo Grazzi · 2025
Closest in time.
“Test-time regression: a unifying framework for designing sequence models with associative memory”
Ke Wang, Jiaxin Shi and Emily Fox · 2025
Closest in time.