Fetching the paper…
Reading the bibliography…
High-dimensional token embeddings underpin Large Language Models (LLMs), as they can capture subtle semantic information and significantly enhance the modelling of complex language patterns.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Tensor-train decomposition
I. V. Oseledets · 2011
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson · 2013
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2017
Earlier work this paper cites.
Groupreduce: Block-wise low-rank approximation for neural language model shrinking
Patrick Chen, Si Si, Yang Li, Ciprian Chelba, and Cho-Jui Hsieh · 2018
Earlier work this paper cites.
Online embedding compression for text classification using low rank matrix factorization
Anish Acharya, Rahul Goel, Angeliki Metallinou, and Inderjit Dhillon · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Distilbert, a distilled version of bert: Smaller, faster, cheaper and lighter
V Sanh · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown · 2020
Earlier work this paper cites.
Tensorized embedding layers
Oleksii Hrinchuk, Valentin Khrulkov, Leyla Mirvakhabova, Elena Orlova, and Ivan Oseledets · 2020
Earlier work this paper cites.
Mcunet: Tiny deep learning on iot devices
Ji Lin, Wei-Ming Chen, Yujun Lin, john cohn, Chuang Gan, and Song Han · 2020
Cited alongside, same era.
Improving word embedding factorization for compression using distilled nonlinear neural decomposition
Vasileios Lioutas, Ahmad Rashid, Krtin Kumar, Md Akmal Haidar, and Mehdi Rezagholizadeh · 2020
Cited alongside, same era.
Ladabert: Lightweight adaptation of bert through hybrid model compression
Yihuan Mao, Yujing Wang, Chufan Wu, Chen Zhang, Yang Wang, Yaming Yang, Quanlu Zhang, Yunhai Tong, and Jing Bai · 2020
Cited alongside, same era.
Drone: Data-aware low-rank compression for large nlp models
Patrick Chen, Hsiang-Fu Yu, Inderjit Dhillon, and Cho-Jui Hsieh · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
J. Edward Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen · 2021
KroneckerBERT: Significant compression of pre-trained language models through kronecker decomposition and knowledge distillation
Marzieh Tahaei, Ella Charlaix, Vahid Nia, Ali Ghodsi, and Mehdi Rezagholizadeh · 2022
Later among the works it cites.
Batude: Budget-aware neural network compression based on tucker decomposition
Miao Yin, Huy Phan, Xiao Zang, Siyu Liao, and Bo Yuan · 2022
Later among the works it cites.
Efficient GPT model pre-training using tensor train matrix representation
Viktoriia Chekalina, Georgiy Novikov, Julia Gusak, Alexander Panchenko, and Ivan Oseledets · 2023
Closest in time.
Cerebras-gpt: Open compute-optimal language models trained on the cerebras wafer-scale cluster
Nolan Dey, Gurpreet Gosal, Hemant Khachane, William Marshall, Ribhu Pathria, Marvin Tom, Joel Hestness, et al · 2023
Closest in time.
Alternating local enumeration (tnale): Solving tensor network structure search with fewer evaluations
Chao Li, Junhua Zeng, Chunmei Li, Cesar F Caiafa, and Qibin Zhao · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Monarch: Expressive structured matrices for efficient and accurate training
Tri Dao, Beidi Chen, Nimit S Sohoni, Arjun Desai, Michael Poli, Jessica Grogan, Alexander Liu, Aniruddh Rao, Atri Rudra, and Christopher Re · 2022
Cited alongside, same era.
Kronecker decomposition for GPT compression
Ali Edalati, Marzieh Tahaei, Ahmad Rashid, Vahid Nia, James Clark, and Mehdi Rezagholizadeh · 2022
Cited alongside, same era.
Language model compression with weighted low-rank factorization
Yen-Chang Hsu, Ting Hua, Sungen Chang, Qian Lou, Yilin Shen, and Hongxia Jin · 2022
Cited alongside, same era.
Permutation search of tensor network structures via local sampling
Chao Li, Junhua Zeng, Zerui Tao, and Qibin Zhao · 2022
Cited alongside, same era.
Compute better spent: Replacing dense layers with structured matrices
Shikai Qiu, Andres Potapczynski, Marc Anton Finzi, Micah Goldblum, and Andrew Gordon Wilson
Cited in the paper.
Closest in time.
Tqcompressor: improving tensor decomposition methods in neural networks via permutations
V Abronin, A Naumov, D Mazur, D Bystrov, K Tsarova, Ar Melnikov, I Oseledets, and R Brasher · 2024
Closest in time.
Melting point: Mobile evaluation of language transformers
Stefanos Laskaridis, Kleomenis Kateveas, Lorenzo Minto, and Hamed Haddadi · 2024
Closest in time.
MobileLLM: Optimizing sub-billion parameter language models for on-device use cases
Zechun Liu, Changsheng Zhao, Forrest Iandola, Chen Lai, Yuandong Tian, Igor Fedorov, Yunyang Xiong, Ernie Chang, Yangyang Shi, Raghuraman Krishnamoorthi, Liangzhen Lai, and Vikas Chandra · 2024
Closest in time.
Openelm: An efficient language model family with open-source training and inference framework
Sachin Mehta, Mohammad Hossein Sekhavat, Qingqing Cao, Maxwell Horton, Yanzi Jin, Chenfan Sun, Iman Mirzadeh, Mahyar Najibi, Dmitry Belenko, Peter Zatloukal, et al · 2024
Closest in time.