Fetching the paper…
Reading the bibliography…
State Space Models (SSMs) have emerged as an appealing alternative to Transformers for large language models, achieving state-of-the-art accuracy with constant memory complexity which allows for holding longer context lengths than attention-based networks.
“Language models are few-shot learners”
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry and Amanda Askell · 1901
Earlier work this paper cites.
“A library of Hadamard matrices”
Neil Sloane · 1999
Earlier work this paper cites.
Song Han, Huizi Mao and William Dally · 2015
Earlier work this paper cites.
“Pointer sentinel mixture models”
Stephen Merity, Caiming Xiong, James Bradbury and Richard Socher · 2016
Earlier work this paper cites.
“Convolutional neural networks using logarithmic data representation”
Daisuke Miyashita, Edward Lee and Boris Murmann · 2016
Earlier work this paper cites.
“The LAMBADA dataset: Word prediction requiring a broad discourse context”
Denis Paperno, Germ\’an Kruszewski, Angeliki Lazaridou, Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda and Raquel Fern\’andez · 2016
Earlier work this paper cites.
“Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge”
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick and Oyvind Tafjord · 2018
Earlier work this paper cites.
“Sigmoid-weighted linear units for neural network function approximation in reinforcement learning”
Stefan Elfwing, Eiji Uchibe and Kenji Doya · 2018
Earlier work this paper cites.
“Quantization and training of neural networks for efficient integer-arithmetic-only inference”
Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam and Dmitry Kalenichenko · 2018
Earlier work this paper cites.
“Fully quantized network for object detection”
Rundong Li, Yan Wang, Feng Liang, Hongwei Qin, Junjie Yan and Rui Fan · 2019
Earlier work this paper cites.
“Haq: Hardware-aware automated quantization with mixed precision”
Kuan Wang, Zhijian Liu, Yujun Lin, Ji Lin and Song Han · 2019
Earlier work this paper cites.
“HellaSwag: Can a Machine Really Finish Your Sentence?”
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi and Yejin Choi · 2019
Earlier work this paper cites.
“Root mean square layer normalization”
Biao Zhang and Rico Sennrich · 2019
Earlier work this paper cites.
“PIQA: Reasoning about Physical Commonsense in Natural Language”
Yonatan Bisk, Rowan Zellers, Ronan Bras, Jianfeng Gao and Yejin Choi · 2020
Earlier work this paper cites.
“Hippo: Recurrent memory with optimal polynomial projections”
Albert Gu, Tri Dao, Stefano Ermon, Atri Rudra and Christopher R\’e · 2020
Earlier work this paper cites.
“WinoGrande: An Adversarial Winograd Schema Challenge at Scale”
Keisuke Sakaguchi, Ronan Bras, Chandra Bhagavatula and Yejin Choi · 2020
Earlier work this paper cites.
“High performance depthwise and pointwise convolutions on mobile devices”
Pengfei Zhang, Eric Lo and Baotong Lu · 2020
Earlier work this paper cites.
“The Pile: An 800GB Dataset of Diverse Text for Language Modeling”
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser and Connor Leahy · 2021
Earlier work this paper cites.
“Efficiently modeling long sequences with structured state spaces”
Albert Gu, Karan Goel and Christopher R\’e · 2021
Earlier work this paper cites.
“Optimizing depthwise separable convolution operations on gpus”
Gangzhao Lu, Weizhe Zhang and Zheng Wang · 2021
Earlier work this paper cites.
“Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale”
Tim Dettmers, Mike Lewis, Younes Belkada and Luke Zettlemoyer · 2022
Earlier work this paper cites.
“Gptq: Accurate post-training quantization for generative pre-trained transformers”
Elias Frantar, Saleh Ashkboos, Torsten Hoefler and Dan Alistarh · 2022
Earlier work this paper cites.
“A survey of quantization methods for efficient neural network inference”
Amir Gholami, Sehoon Kim, Zhen Dong, Zhewei Yao, Michael Mahoney and Kurt Keutzer · 2022
Cited alongside, same era.
“It’s raw! audio generation with state-space models”
Karan Goel, Albert Gu, Chris Donahue and Christopher R\’e · 2022
Cited alongside, same era.
“On the parameterization and initialization of diagonal state space models”
Albert Gu, Karan Goel, Ankit Gupta and Christopher R\’e · 2022
Cited alongside, same era.
“S4nd: Modeling images and videos as multidimensional signals using state spaces”
Eric Nguyen, Karan Goel, Albert Gu, Gordon Downs, Preey Shah, Tri Dao, Stephen Baccus and Christopher R\’e · 2022
Cited alongside, same era.
“Optimal clipping and magnitude-aware differentiation for improved quantization-aware training”
Charbel Sakr, Steve Dai, Rangha Venkatesan, Brian Zimmer, William Dally and Brucek Khailany · 2022
Cited alongside, same era.
“Atom: Low-bit quantization for efficient and accurate llm serving”
Yilong Zhao, Chien-Yu Lin, Kan Zhu, Zihao Ye, Lequn Chen, Size Zheng, Luis Ceze, Arvind Krishnamurthy, Tianqi Chen and Baris Kasikci · 2023
Later among the works it cites.
“A survey on model compression for large language models”
Xunyu Zhu, Jian Li, Yong Liu, Can Ma and Weiping Wang · 2023
Later among the works it cites.
“Slicegpt: Compress large language models by deleting rows and columns”
Saleh Ashkboos, Maximilian Croci, Marcelo Gennari Nascimento, Torsten Hoefler and James Hensman · 2024
Closest in time.
“Quarot: Outlier-free 4-bit inference in rotated llms”
Saleh Ashkboos, Amirkeivan Mohtashami, Maximilian Croci, Bo Li, Martin Jaggi, Dan Alistarh, Torsten Hoefler and James Hensman · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Simplified state space layers for sequence modeling”
Jimmy Smith, Andrew Warrington and Scott Linderman · 2022
Cited alongside, same era.
“Opt: Open pre-trained transformer language models”
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li and Xi Lin · 2022
Cited alongside, same era.
“Analysis of quantization on mlp-based vision models”
Lingran Zhao, Zhen Dong and Kurt Keutzer · 2022
Cited alongside, same era.
“Number Systems for Deep Neural Network Architectures: A Survey”, 2023
Ghada Alsuhli, Vasileios Sakellariou, Hani Saleh, Mahmoud Al-Qutayri, Baker Mohammad and Thanos Stouraitis · 2023
Cited alongside, same era.
“Pythia: A suite for analyzing large language models across training and scaling”
Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Khan, Shivanshu Purohit, USVSN Prashanth and Edward Raff · 2023
Cited alongside, same era.
“Nvidia hopper h100 gpu: Scaling performance”
Jack Choquette · 2023
Cited alongside, same era.
“A framework for few-shot language model evaluation”
Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang and Andy Zou · 2023
Cited alongside, same era.
Maximilian Beck, Korbinian P\"oppel, Markus Spanring, Andreas Auer, Oleksandra Prudnikova, Michael Kopp, G\"unter Klambauer, Johannes Brandstetter and Sepp Hochreiter · 2024
Closest in time.
“Causal depthwise conv1d in CUDA with a PyTorch interface”, 2024
Tri Dao · 2024
Closest in time.
“Fast Hadamard Transform in CUDA, with a PyTorch interface”, 2024
Tri Dao · 2024
Closest in time.
“Qlora: Efficient finetuning of quantized llms”
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman and Luke Zettlemoyer · 2024
Closest in time.
Albert Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Chaplot, Diego de Casas, Emma Hanna and Florian Bressand · 2024
Closest in time.
Tanishq Kumar, Zachary Ankner, Benjamin Spector, Blake Bordelon, Niklas Muennighoff, Mansheej Paul, Cengiz Pehlevan, Christopher R\’e and Aditi Raghunathan · 2024
Closest in time.
“FP8 Quantization: The Power of the Exponent”, 2024
Andrey Kuzmin, Mart Baalen, Yuwei Ren, Markus Nagel, Jorn Peters and Tijmen Blankevoort · 2024
Closest in time.
“Videomamba: State space model for efficient video understanding”
Kunchang Li, Xinhao Li, Yi Wang, Yinan He, Yali Wang, Limin Wang and Yu Qiao · 2024
Closest in time.
“Jamba: A Hybrid Transformer-Mamba Language Model”
Opher Lieber, Barak Lenz, Hofit Bata, Gal Cohen, Jhonathan Osin, Itay Dalmedigos, Erez Safahi, Shaked Meirom, Yonatan Belinkov, Shai Shalev-Shwartz, Omri Abend, Raz Alon, Tomer Asida, Amir Bergman, Roman Glozman, Michael Gokhman, Avashalom Manevich, Nir Ratner, Noam Rozen, Erez Shwartz, Mor Zusman and Yoav Shoham · 2024
Closest in time.
“Jamba: A hybrid transformer-mamba language model”
Opher Lieber, Barak Lenz, Hofit Bata, Gal Cohen, Jhonathan Osin, Itay Dalmedigos, Erez Safahi, Shaked Meirom, Yonatan Belinkov and Shai Shalev-Shwartz · 2024
Closest in time.
“Qserve: W4a8kv4 quantization and system co-design for efficient llm serving”
Yujun Lin, Haotian Tang, Shang Yang, Zhekai Zhang, Guangxuan Xiao, Chuang Gan and Song Han · 2024
Closest in time.
“Mamba-PTQ: Outlier Channels in Recurrent Large Language Models”
Alessandro Pierro and Steven Abreu · 2024
Closest in time.
“Quip#: Even better LLM quantization with hadamard incoherence and lattice codebooks”
Albert Tseng, Jerry Chee, Qingyao Sun, Volodymyr Kuleshov and Christopher De · 2024
Closest in time.
“Mambabyte: Token-free selective state space model”
Junxiong Wang, Tushaar Gangavarapu, Jing Yan and Alexander Rush · 2024
Closest in time.
“Efficient low-rank backpropagation for vision transformer adaptation”
Yuedong Yang, Hung-Yueh Chiang, Guihong Li, Diana Marculescu and Radu Marculescu · 2024
Closest in time.
“Cobra: Extending mamba to multi-modal large language model for efficient inference”
Han Zhao, Min Zhang, Wei Zhao, Pengxiang Ding, Siteng Huang and Donglin Wang · 2024
Closest in time.
“A Survey on Efficient Inference for Large Language Models”
Zixuan Zhou, Xuefei Ning, Ke Hong, Tianyu Fu, Jiaming Xu, Shiyao Li, Yuming Lou, Luning Wang, Zhihang Yuan and Xiuhong Li · 2024
Closest in time.
“Vision mamba: Efficient visual representation learning with bidirectional state space model”
Lianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang, Wenyu Liu and Xinggang Wang · 2024
Closest in time.