Training compute-optimal large language models, 2022
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, and Laurent Sifre · 2022
Later among the works it cites.
The dual form of neural networks revisited: Connecting test time predictions to training patterns via spotlights of attention
Kazuki Irie, Róbert Csordás, and Jürgen Schmidhuber · 2022
Later among the works it cites.
Neural differential equations for learning to program neural nets through continuous learning rules
Kazuki Irie, Francesco Faccio, and Jürgen Schmidhuber · 2022
Later among the works it cites.
A modern self-referential weight matrix that learns to modify itself
Kazuki Irie, Imanol Schlag, Róbert Csordás, and Jürgen Schmidhuber · 2022
Later among the works it cites.
Images as weight matrices: Sequential image generation through synaptic learning rules
Kazuki Irie and Jürgen Schmidhuber · 2022
Later among the works it cites.
The devil in linear transformer
Original
Zhen Qin, Xiaodong Han, Weixuan Sun, Dongxu Li, Lingpeng Kong, Nick Barnes, and Yiran Zhong · 2022
Later among the works it cites.
Gpt-4 technical report
Original
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Later among the works it cites.
EasyLM: A Simple And Scalable Training Framework for Large Language Models
Xinyang Geng · 2023
Later among the works it cites.
Mamba: Linear-time sequence modeling with selective state spaces
Original
Albert Gu and Tri Dao · 2023
Later among the works it cites.
Test-time training on nearest neighbors for large language models
Original
Moritz Hardt and Yu Sun · 2023
Later among the works it cites.
Practical computational power of linear transformers and their recurrent and self-referential extensions
Kazuki Irie, Róbert Csordás, and Jürgen Schmidhuber · 2023
Later among the works it cites.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica · 2023
Later among the works it cites.
Rwkv: Reinventing rnns for the transformer era
Original
Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Stella Biderman, Huanqi Cao, Xin Cheng, Michael Chung, Matteo Grella, et al · 2023
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding, 2023
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu · 2023
Later among the works it cites.
Learning to (learn at test time)
Original
Yu Sun, Xinhao Li, Karan Dalal, Chloe Hsu, Sanmi Koyejo, Carlos Guestrin, Xiaolong Wang, Tatsunori Hashimoto, and Xinlei Chen · 2023
Later among the works it cites.
Test-time training on video streams
Original
Renhao Wang, Yu Sun, Yossi Gandelsman, Xinlei Chen, Alexei A Efros, and Xiaolong Wang · 2023
Later among the works it cites.
Effective long-context scaling of foundation models, 2023
Wenhan Xiong, Jingyu Liu, Igor Molybog, Hejia Zhang, Prajjwal Bhargava, Rui Hou, Louis Martin, Rashi Rungta, Karthik Abinav Sankararaman, Barlas Oguz, Madian Khabsa, Han Fang, Yashar Mehdad, Sharan Narang, Kshitiz Malik, Angela Fan, Shruti Bhosale, Sergey Edunov, Mike Lewis, Sinong Wang, and Hao Ma · 2023
Later among the works it cites.
Gated linear attention transformers with hardware-efficient training
Original
Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda, and Yoon Kim · 2023
Later among the works it cites.
You just found out your book was used to train ai. now what?, 2023
Authors Guild · 2024
Closest in time.
xlstm: Extended long short-term memory
Original
Maximilian Beck, Korbinian Pöppel, Markus Spanring, Andreas Auer, Oleksandra Prudnikova, Michael Kopp, Günter Klambauer, Johannes Brandstetter, and Sepp Hochreiter · 2024
Closest in time.
Transformers are ssms: Generalized models and efficient algorithms through structured state space duality
Original
Tri Dao and Albert Gu · 2024
Closest in time.
Griffin: Mixing gated linear recurrences with local attention for efficient language models
Original
Soham De, Samuel L Smith, Anushan Fernando, Aleksandar Botev, George Cristian-Muraru, Albert Gu, Ruba Haroun, Leonard Berrada, Yutian Chen, Srivatsan Srinivasan, et al · 2024
Closest in time.
In the long (context) run, 2023
Harm de Vries · 2024
Closest in time.
Unlocking state-tracking in linear rnns through negative eigenvalues
Riccardo Grazzi, Julien Siems, Arber Zela, Jörg KH Franke, Frank Hutter, and Massimiliano Pontil · 2024
Closest in time.
Strangely, matrix multiplications on gpus run faster when given "predictable" data! [short], 2024
Horace He · 2024
Closest in time.
World model on million-length video and language with blockwise ringattention
Original
Hao Liu, Wilson Yan, Matei Zaharia, and Pieter Abbeel · 2024
Closest in time.
Eagle and finch: Rwkv with matrix-valued states and dynamic recurrence
Original
Bo Peng, Daniel Goldstein, Quentin Anthony, Alon Albalak, Eric Alcaide, Stella Biderman, Eugene Cheah, Teddy Ferdinan, Haowen Hou, Przemysław Kazienko, et al · 2024
Closest in time.
Parallelizing linear transformers with the delta rule over sequence length
Original
Songlin Yang, Bailin Wang, Yu Zhang, Yikang Shen, and Yoon Kim · 2024
Closest in time.