Squeezellm: Dense-and-sparse quantization
Original
Kim, S., Hooper, C., Gholami, A., Dong, Z., Li, X., Shen, S., Mahoney, M. W., and Keutzer, K · 2023
Later among the works it cites.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., and Stoica, I · 2023
Later among the works it cites.
Alpacaeval: An automatic evaluator of instruction-following models
Li, X., Zhang, T., Dubois, Y., Taori, R., Gulrajani, I., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Later among the works it cites.
Awq: Activation-aware weight quantization for llm compression and acceleration
Original
Lin, J., Tang, J., Tang, H., Yang, S., Dang, X., and Han, S · 2023
Later among the works it cites.
Locally typical sampling
Meister, C., Pimentel, T., Wiher, G., and Cotterell, R · 2023
Later among the works it cites.
Specinfer: Accelerating generative llm serving with speculative inference and token tree verification
Original
Miao, X., Oliaro, G., Zhang, Z., Cheng, X., Wang, Z., Wong, R. Y. Y., Chen, Z., Arfeen, D., Abhyankar, R., and Jia, Z · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
Tiny vicuna 1b
Pan, J · 2023
Later among the works it cites.
ShareGPT
ShareGPT · 2023
Later among the works it cites.
Accelerating llm inference with staged speculative decoding
Original
Spector, B. and Re, C · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Original
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Zephyr: Direct distillation of lm alignment, 2023
Tunstall, L., Beeching, E., Lambert, N., Rajani, N., Rasul, K., Belkada, Y., Huang, S., von Werra, L., Fourrier, C., Habib, N., Sarrazin, N., Sanseviero, O., Rush, A. M., and Wolf, T · 2023
Later among the works it cites.
Speculative decoding: Lossless speedup of autoregressive translation, 2023
Xia, H., Ge, T., Chen, S.-Q., Wei, F., and Sui, Z · 2023
Later among the works it cites.
H _ 2 \_2 o: Heavy-hitter oracle for efficient generative inference of large language models
Original
Zhang, Z., Sheng, Y., Zhou, T., Chen, T., Zheng, L., Cai, R., Song, Z., Tian, Y., Ré, C., Barrett, C., et al · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena, 2023
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., and Stoica, I · 2023
Later among the works it cites.
Tinyllama: An open-source small language model, 2024
Zhang, P., Zeng, G., Wang, T., and Lu, W · 2024
Closest in time.