Fetching the paper…
Reading the bibliography…
To deploy LLMs on resource-contained platforms such as mobile robots and smartphones, non-transformers LLMs have achieved major breakthroughs.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Recurrent neural network based language model
Mikolov, T., Karafiát, M., Burget, L., Černocký, J., and Khudanpur, S · 2010
Earlier work this paper cites.
Exploring the Limits of Language Modeling, February 2016
Jozefowicz, R., Vinyals, O., Schuster, M., Shazeer, N., and Wu, Y · 2016
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context, June 2016
Paperno, D., Kruszewski, G., Lazaridou, A., Pham, Q. N., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fernández, R · 2016
Earlier work this paper cites.
Online Embedding Compression for Text Classification using Low Rank Matrix Factorization, November 2018
Acharya, A., Goel, R., Metallinou, A., and Dhillon, I · 2018
Earlier work this paper cites.
GroupReduce: Block-Wise Low-Rank Approximation for Neural Language Model Shrinking, June 2018
Chen, P. H., Si, S., Li, Y., Chelba, C., and Hsieh, C.-j · 2018
Earlier work this paper cites.
The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks, March 2019
Frankle, J. and Carbin, M · 2019
Earlier work this paper cites.
Fundamentals of Recurrent Neural Network (RNN) and Long Short-Term Memory (LSTM) network
Sherstinsky, A · 2019
Earlier work this paper cites.
Compressing Pre-trained Language Models by Matrix Decomposition
Ben Noach, M. and Goldberg, Y · 2020
Earlier work this paper cites.
The Pile: An 800gb dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., Presser, S., and Leahy, C · 2020
Earlier work this paper cites.
GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow, March 2021
Black, S., Gao, L., Wang, P., Leahy, C., and Biderman, S · 2021
Cited alongside, same era.
LoRA: Low-Rank Adaptation of Large Language Models, October 2021
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Cited alongside, same era.
LANGUAGE MODEL COMPRESSION WITH WEIGHTED LOW-RANK FACTORIZATION
Hsu, Y.-C., Hua, T., Chang, S.-E., Lou, Q., Shen, Y., and Jin, H · 2022
Cited alongside, same era.
OPT: Open Pre-trained Transformer Language Models, June 2022
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., Mihaylov, T., Ott, M., Shleifer, S., Shuster, K., Simig, D., Koura, P. S., Sridhar, A., Wang, T., and Zettlemoyer, L · 2022
Cited alongside, same era.
Supervised Clustering Loss for Clustering-Friendly Sentence Embeddings: An Application to Intent Clustering
Barnabò, G., Uva, A., Pollastrini, S., Rubagotti, C., and Bernardi, D · 2023
Cited alongside, same era.
An Empirical Analysis and Resource Footprint Study of Deploying Large Language Models on Edge Devices
Dhar, N., Deng, B., Lo, D., Wu, X., Zhao, L., and Suo, K · 2024
Closest in time.
RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation, April 2024
Jin, C., Zhang, Z., Jiang, X., Liu, F., Liu, X., Liu, X., and Jin, X · 2024
Closest in time.
SpecInfer: Accelerating Large Language Model Serving with Tree-based Speculative Inference and Verification
Miao, X., Oliaro, G., Zhang, Z., Cheng, X., Wang, Z., Zhang, Z., Wong, R. Y. Y., Zhu, A., Yang, L., Shi, X., Shi, C., Chen, Z., Arfeen, D., Abhyankar, R., and Jia, Z · 2024
Closest in time.
Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence, April 2024
Peng, B., Goldstein, D., Anthony, Q., Albalak, A., Alcaide, E., Biderman, S., Cheah, E., Du, X., Ferdinan, T., Hou, H., Kazienko, P., GV, K. K., Kocoń, J., Koptyra, B., Krishna, S., McClelland Jr., R., Muennighoff, N., Obeid, F., Saito, A., Song, G., Tu, H., Woźniak, S., Zhang, R., Zhao, B., Zhao, Q., Zhou, P., Zhu, J., and Zhu, R.-J · 2024
Closest in time.
Context-Aware Clustering using Large Language Models, May 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The case for 4-bit precision: K-bit Inference Scaling Laws
Dettmers, T. and Zettlemoyer, L · 2023
Cited alongside, same era.
GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers, March 2023
Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D · 2023
Cited alongside, same era.
Deja Vu: Contextual Sparsity for Efficient LLMs at Inference Time, October 2023
Liu, Z., Wang, J., Dao, T., Zhou, T., Yuan, B., Song, Z., Shrivastava, A., Zhang, C., Tian, Y., Re, C., and Chen, B · 2023
Cited alongside, same era.
RWKV: Reinventing RNNs for the Transformer Era, December 2023
Peng, B., Alcaide, E., Anthony, Q., Albalak, A., Arcadinho, S., Biderman, S., Cao, H., Cheng, X., Chung, M., Grella, M., GV, K. K., He, X., Hou, H., Lin, J., Kazienko, P., Kocon, J., Kong, J., Koptyra, B., Lau, H., Mantri, K. S. I., Mom, F., Saito, A., Song, G., Tang, X., Wang, B., Wind, J. S., Wozniak, S., Zhang, R., Zhang, Z., Zhao, Q., Zhou, P., Zhou, Q., Zhu, J., and Zhu, R.-J · 2023
Cited alongside, same era.
https://github.com/RWKV/rwkv.cpp
rwkv.cpp · 2024
Cited alongside, same era.
https://github.com/keeeeenw/MicroLlama
Microllama
Cited in the paper.
CHAI: Clustered Head Attention for Efficient LLM Inference
Agarwal, S., Acun, B., Hosmer, B., Elhoushi, M., Lee, Y., Venkataraman, S., Papailiopoulos, D., and Wu, C.-J
Cited in the paper.
Tipirneni, S., Adkathimar, R., Choudhary, N., Hiranandani, G., Amjad, R. A., Ioannidis, V. N., Yuan, C., and Reddy, C. K · 2024
Closest in time.
Large Language Models Enable Few-Shot Clustering
Viswanathan, V., Gashteovski, K., Gashteovski, K., Lawrence, C., Wu, T., and Neubig, G · 2024
Closest in time.
PowerInfer-2: Fast Large Language Model Inference on a Smartphone, June 2024
Xue, Z., Song, Y., Mi, Z., Chen, L., Xia, Y., and Chen, H · 2024
Closest in time.
Tinyllama: An open-source small language model, 2024
Zhang, P., Zeng, G., Wang, T., and Lu, W · 2024
Closest in time.
https://blog.rwkv.com/p/rwkvcpp-shipping-to-half-a-billion
Rwkv.cpp - shipping to 1.5 billion systems worldwide · 2025
Closest in time.