Fetching the paper…
Reading the bibliography…
Transformer architectures have become a dominant paradigm for domains like language modeling but suffer in many inference settings due to their quadratic-time self-attention.
“Multilingual Neural Machine Translation with Knowledge Distillation”, 2019
Xu Tan, Yi Ren, Di He, Tao Qin, Zhou Zhao and Tie-Yan Liu · 1902
Earlier work this paper cites.
“TextKD-GAN: Text Generation using KnowledgeDistillation and Generative Adversarial Networks”, 2019
Md. Haidar and Mehdi Rezagholizadeh · 1905
Earlier work this paper cites.
“Self-Knowledge Distillation in Natural Language Processing”, 2019
Sangchul Hahn and Heeyoul Choi · 1908
Earlier work this paper cites.
“Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer”, 2023
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li and Peter. Liu · 1910
Earlier work this paper cites.
Ze Yang, Linjun Shou, Ming Gong, Wutao Lin and Daxin Jiang · 1910
Earlier work this paper cites.
“Distilling Knowledge Learned in BERT for Text Generation”, 2020
Yen-Chun Chen, Zhe Gan, Yu Cheng, Jingzhou Liu and Jingjing Liu · 1911
Earlier work this paper cites.
“Understanding Knowledge Distillation in Non-autoregressive Machine Translation”, 2021
Chunting Zhou, Graham Neubig and Jiatao Gu · 1911
Earlier work this paper cites.
“On the Hankel-norm approximation of upper-triangular operators and matrices”
P. Dewilde and A.. van Veen · 1993
Earlier work this paper cites.
“On low-complexity approximation of matrices”
A.J. van Veen and P. Dewilde · 1994
Earlier work this paper cites.
“Time-Varying Systems and Computations” Springer Science+Business Media Dordrecht 1998, Springer Book Archive
Patrick Dewilde and Alle-Jan Veen · 1998
Earlier work this paper cites.
“Fast stable solver for sequentially semi-separable linear systems of equations”
Shiv Chandrasekaran, Patrick Dewilde, Ming Gu, T Pals and Alle-Jan van Veen · 2002
Earlier work this paper cites.
“Balanced truncation of linear time-varying systems”
H. Sandberg and A. Rantzer · 2003
Earlier work this paper cites.
“A Survey of Model Reduction by Balanced Truncation and Some New Results”
Serkan Gugercin and Athanasios. Antoulas · 2004
Earlier work this paper cites.
“Language Models are Few-Shot Learners”, 2020
Tom. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever and Dario Amodei · 2005
Earlier work this paper cites.
“Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention”, 2020
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas and François Fleuret · 2006
Earlier work this paper cites.
“Model reduction of linear time-varying systems over finite horizons”
Samuel. Melchior, Paul Van Dooren and Kyle. Gallivan · 2013
Earlier work this paper cites.
“Distilling the Knowledge in a Neural Network”, 2015
Geoffrey Hinton, Oriol Vinyals and Jeff Dean · 2015
Earlier work this paper cites.
“The LAMBADA Dataset: Word Prediction Requiring a Broad Discourse Context”
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc-Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda and Raquel Fernández · 2016
Earlier work this paper cites.
“Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge”
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick and Oyvind Tafjord · 2018
Earlier work this paper cites.
“Attention-Guided Answer Distillation for Machine Reading Comprehension”, 2018
Minghao Hu, Yuxing Peng, Furu Wei, Zhen Huang, Dongsheng Li, Nan Yang and Ming Zhou · 2018
Cited alongside, same era.
“HellaSwag: Can a Machine Really Finish Your Sentence?”
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi and Yejin Choi · 2019
Cited alongside, same era.
“PIQA: Reasoning about Physical Commonsense in Natural Language”
Yonatan Bisk, Rowan Zellers, Jianfeng Gao and Yejin Choi · 2020
Cited alongside, same era.
“Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers”
Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang and Ming Zhou · 2020
Cited alongside, same era.
“Finetuning pretrained transformers into rnns”
Jungo Kasai, Hao Peng, Yizhe Zhang, Dani Yogatama, Gabriel Ilharco, Nikolaos Pappas, Yi Mao, Weizhu Chen and Noah Smith · 2021
“Retentive Network: A Successor to Transformer for Large Language Models”, 2023
Yutao Sun, Li Dong, Shaohan Huang, Shuming Ma, Yuqing Xia, Jilong Xue, Jianyong Wang and Furu Wei · 2023
Later among the works it cites.
“LLaMA: Open and Efficient Foundation Language Models”, 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave and Guillaume Lample · 2023
Later among the works it cites.
“Attention Is All You Need”, 2023
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan. Gomez, Lukasz Kaiser and Illia Polosukhin · 2023
Later among the works it cites.
“Selective Structured State-Spaces for Long-Form Video Understanding”, 2023
Jue Wang, Wentao Zhu, Pichao Wang, Xiang Yu, Linda Liu, Mohamed Omar and Raffay Hamid · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Winogrande: An Adversarial Winograd Schema Challenge at Scale”
Keisuke Sakaguchi, Ronan Bras, Chandra Bhagavatula and Yejin Choi · 2021
Cited alongside, same era.
“FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness”, 2022
Tri Dao, Daniel. Fu, Stefano Ermon, Atri Rudra and Christopher Ré · 2022
Cited alongside, same era.
“Efficiently Modeling Long Sequences with Structured State Spaces”, 2022
Albert Gu, Karan Goel and Christopher Ré · 2022
Cited alongside, same era.
“Long Range Language Modeling via Gated State Spaces”, 2022
Harsh Mehta, Ankit Gupta, Ashok Cutkosky and Behnam Neyshabur · 2022
Cited alongside, same era.
“Hungry Hungry Hippos: Towards Language Modeling with State Space Models”, 2023
Daniel. Fu, Tri Dao, Khaled. Saab, Armin. Thomas, Atri Rudra and Christopher Ré · 2023
Cited alongside, same era.
“Modeling Sequences with Structured State Spaces”, 2023
Albert Gu · 2023
Cited alongside, same era.
“Mamba: Linear-Time Sequence Modeling with Selective State Spaces”, 2023
Albert Gu and Tri Dao · 2023
Cited alongside, same era.
“Sheared llama: Accelerating language model pre-training via structured pruning”
Mengzhou Xia, Tianyu Gao, Zhiyuan Zeng and Danqi Chen · 2023
Later among the works it cites.
“xLSTM: Extended Long Short-Term Memory”, 2024
Maximilian Beck, Korbinian Pöppel, Markus Spanring, Andreas Auer, Oleksandra Prudnikova, Michael Kopp, Günter Klambauer, Johannes Brandstetter and Sepp Hochreiter · 2024
Closest in time.
“Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality”
Tri Dao and Albert Gu · 2024
Closest in time.
“MambaVision: A Hybrid Mamba-Transformer Vision Backbone”, 2024
Ali Hatamizadeh and Jan Kautz · 2024
Closest in time.
“Hydra: Bidirectional State Space Models Through Generalized Matrix Mixers”, 2024
Sukjun Hwang, Aakash Lahoti, Tri Dao and Albert Gu · 2024
Closest in time.
“Jamba: A Hybrid Transformer-Mamba Language Model”, 2024
Opher Lieber, Barak Lenz, Hofit Bata, Gal Cohen, Jhonathan Osin, Itay Dalmedigos, Erez Safahi, Shaked Meirom, Yonatan Belinkov, Shai Shalev-Shwartz, Omri Abend, Raz Alon, Tomer Asida, Amir Bergman, Roman Glozman, Michael Gokhman, Avashalom Manevich, Nir Ratner, Noam Rozen, Erez Shwartz, Mor Zusman and Yoav Shoham · 2024
Closest in time.
“Linearizing Large Language Models”, 2024
Jean Mercat, Igor Vasiljevic, Sedrick Keh, Kushal Arora, Achal Dave, Adrien Gaidon and Thomas Kollar · 2024
Closest in time.
“What does the Knowledge Neuron Thesis Have to do with Knowledge?”, 2024
Jingcheng Niu, Andrew Liu, Zining Zhu and Gerald Penn · 2024
Closest in time.
“HGRN2: Gated Linear RNNs with State Expansion”, 2024
Zhen Qin, Songlin Yang, Weixuan Sun, Xuyang Shen, Dong Li, Weigao Sun and Yiran Zhong · 2024
Closest in time.
“Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling”, 2024
Liliang Ren, Yang Liu, Yadong Lu, Yelong Shen, Chen Liang and Weizhu Chen · 2024
Closest in time.
“The Mamba in the Llama: Distilling and Accelerating Hybrid Models”, 2024
Junxiong Wang, Daniele Paliotta, Avner May, Alexander. Rush and Tri Dao · 2024
Closest in time.
“PoinTramba: A Hybrid Transformer-Mamba Framework for Point Cloud Analysis”, 2024
Zicheng Wang, Zhenghao Chen, Yiming Wu, Zhen Zhao, Luping Zhou and Dong Xu · 2024
Closest in time.
“Gated Linear Attention Transformers with Hardware-Efficient Training”, 2024
Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda and Yoon Kim · 2024
Closest in time.
“The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax Mimicry”, 2024
Michael Zhang, Kush Bhatia, Hermann Kumbong and Christopher Ré · 2024
Closest in time.