Fetching the paper…
Reading the bibliography…
Although quantization for linear layers has been widely used, its application to accelerate the attention process remains limited.
Perplexity—a measure of the difficulty of speech recognition tasks
Jelinek, F., Mercer, R. L., Bahl, L. R., and Baker, J. K · 1977
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Librispeech: An asr corpus based on public domain audio books
Panayotov, V., Chen, G., Povey, D., and Khudanpur, S · 2015
Earlier work this paper cites.
The lambada dataset: Word prediction requiring a broad discourse context
Paperno, D., Kruszewski, G., Lazaridou, A., Pham, N.-Q., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fernández, R · 2016
Earlier work this paper cites.
Improved techniques for training gans
Salimans, T., Goodfellow, I., Zaremba, W., Cheung, V., Radford, A., and Chen, X · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A · 2017
Earlier work this paper cites.
Online normalizer calculation for softmax
Milakov, M. and Gimelshein, N · 2018
Earlier work this paper cites.
Learning robust global representations by penalizing local predictive power
Wang, H., Ge, S., Lipton, Z., and Xing, E. P · 2019
Earlier work this paper cites.
Pytorch image models
Wightman, R · 2019
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2020
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
Wang, S., Li, B. Z., Khabsa, M., Fang, H., and Ma, H · 2020
Earlier work this paper cites.
Rethinking attention with performers
Choromanski, K. M., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J. Q., Mohiuddin, A., Kaiser, L., Belanger, D. B., Colwell, L. J., and Weller, A · 2021
Earlier work this paper cites.
Twins: Revisiting the design of spatial attention in vision transformers
Chu, X., Tian, Z., Wang, Y., Zhang, B., Ren, H., Wei, X., Xia, H., and Shen, C · 2021
Earlier work this paper cites.
Clipscore: A reference-free evaluation metric for image captioning
Hessel, J., Holtzman, A., Forbes, M., Le Bras, R., and Choi, Y · 2021
Earlier work this paper cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B · 2021
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with IO-awareness
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Re, C · 2022
Earlier work this paper cites.
xformers: A modular and hackable transformer modelling library
Lefaudeux, B., Massa, F., Liskovich, D., Xiong, W., Caggiano, V., Naren, S., Xu, M., Hu, J., Tintore, M., Zhang, S., Labatut, P., Haziza, D., Wehrstedt, L., Reizenstein, J., and Sizov, G · 2022
Earlier work this paper cites.
Uniformer: Unified transformer for efficient spatial-temporal representation learning
Li, K., Wang, Y., Peng, G., Song, G., Liu, Y., Li, H., and Qiao, Y · 2022
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2022
Cited alongside, same era.
Metaformer is actually what you need for vision
Yu, W., Luo, M., Zhou, P., Si, C., Zhou, Y., Wang, X., Feng, J., and Yan, S · 2022
Cited alongside, same era.
QLoRA: Efficient finetuning of quantized LLMs
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L · 2023
Cited alongside, same era.
Llmtest needle in a haystack - pressure testing llms
Kamradt, G · 2023
Cited alongside, same era.
CUTLASS: CUDA Templates for Linear Algebra Subroutines and Solvers
NVIDIA · 2023
Cited alongside, same era.
Introducing stable diffusion 3.5
Stability AI · 2023
Cited alongside, same era.
Moa: Mixture of sparse attention for automatic large language model compression
Fu, T., Huang, H., Ning, X., Zhang, G., Chen, B., Wu, T., Wang, H., Huang, Z., Li, S., Yan, S., Dai, G., Yang, H., and Wang, Y · 2024
Closest in time.
Seerattention: Learning intrinsic sparse attention in your llms
Gao, Y., Zeng, Z., Du, D., Cao, S., So, H. K.-H., Cao, T., Yang, F., and Yang, M · 2024
Closest in time.
Chatglm: A family of large language models from glm-130b to glm-4 all tools
GLM, T., Zeng, A., Xu, B., Wang, B., Zhang, C., Yin, D., Rojas, D., Feng, G., Zhao, H., Lai, H., Yu, H., Wang, H., Sun, J., Zhang, J., Cheng, J., Gui, J., Tang, J., Zhang, J., Li, J., Zhao, L., Wu, L., Zhong, L., Liu, M., Huang, M., Zhang, P., Zheng, Q., Lu, R., Duan, S., Zhang, S., Cao, S., Yang, S., Tam, W. L., Zhao, W., Liu, X., Xia, X., Zhang, X., Gu, X., Lv, X., Liu, X., Liu, X., Yang, X., Song, X., Zhang, X., An, Y., Xu, Y., Niu, Y., Yang, Y., Li, Y., Bai, Y., Dong, Y., Qi, Z., Wang, Z., Yang, Z., Du, Z., Hou, Z., and Wang, Z · 2024
Closest in time.
MInference 1.0: Accelerating pre-filling for long-context LLMs via dynamic sparse attention
Jiang, H., LI, Y., Zhang, C., Wu, Q., Luo, X., Ahn, S., Han, Z., Abdi, A. H., Li, D., Lin, C.-Y., Yang, Y., and Qiu, L · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Cited alongside, same era.
Exploring video quality assessment on user generated contents from aesthetic and technical perspectives
Wu, H., Zhang, E., Liao, L., Chen, C., Hou, J., Wang, A., Sun, W., Yan, Q., and Lin, W · 2023
Cited alongside, same era.
Smoothquant: Accurate and efficient post-training quantization for large language models
Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., and Han, S · 2023
Cited alongside, same era.
Imagereward: Learning and evaluating human preferences for text-to-image generation
Xu, J., Liu, X., Wu, Y., Tong, Y., Li, Q., Ding, M., Tang, J., and Dong, Y · 2023
Cited alongside, same era.
Dpm-solver-v3: Improved diffusion ode solver with empirical model statistics
Zheng, K., Lu, C., Chen, J., and Zhu, J · 2023
Cited alongside, same era.
Quarot: Outlier-free 4-bit inference in rotated LLMs
Ashkboos, S., Mohtashami, A., Croci, M. L., Li, B., Cameron, P., Jaggi, M., Alistarh, D., Hoefler, T., and Hensman, J · 2024
Cited alongside, same era.
Closest in time.
Hunyuanvideo: A systematic framework for large video generative models
Kong, W., Tian, Q., Zhang, Z., Min, R., Dai, Z., Zhou, J., Xiong, J., Li, X., Wu, B., Zhang, J., Wu, K., Lin, Q., Wang, A., Wang, A., Li, C., Huang, D., Yang, F., Tan, H., Wang, H., Song, J., Bai, J., Wu, J., Xue, J., Wang, J., Yuan, J., Wang, K., Liu, M., Li, P., Li, S., Wang, W., Yu, W., Deng, X., Li, Y., Long, Y., Chen, Y., Cui, Y., Peng, Y., Yu, Z., He, Z., Xu, Z., Zhou, Z., Xu, Z., Tao, Y., Lu, Q., Liu, S., Zhou, D., Wang, H., Yang, Y., Wang, D., Liu, Y., Jiang, J., and Zhong, C · 2024
Closest in time.
Playground v2.5: Three insights towards enhancing aesthetic quality in text-to-image generation
Li, D., Kamko, A., Akhgari, E., Sabet, A., Xu, L., and Doshi, S · 2024
Closest in time.
Evalcrafter: Benchmarking and evaluating large video generation models
Liu, Y., Cun, X., Liu, X., Wang, X., Zhang, Y., Chen, H., Liu, Y., Zeng, T., Chan, R., and Shan, Y · 2024
Closest in time.
Flashattention-3: Fast and accurate attention with asynchrony and low-precision
Shah, J., Bikshandi, G., Zhang, Y., Thakkar, V., Ramani, P., and Dao, T · 2024
Closest in time.
Skip-attention: Improving vision transformers by paying less attention
Venkataramanan, S., Ghodrati, A., Asano, Y. M., Porikli, F., and Habibian, A · 2024
Closest in time.
Framebridge: Improving image-to-video generation with bridge models
Wang, Y., Chen, Z., Chen, X., Zhu, J., and Chen, J · 2024
Closest in time.
Infllm: Training-free long-context extrapolation for llms with an efficient context memory
Xiao, C., Zhang, P., Han, X., Xiao, G., Lin, Y., Zhang, Z., Liu, Z., and Sun, M · 2024
Closest in time.
∞ \infty Bench: Extending long context evaluation beyond 100K tokens
Zhang, X., Chen, Y., Hu, S., Xu, Z., Chen, J., Hao, M., Han, X., Thai, Z., Wang, S., Liu, Z., and Sun, M · 2024
Closest in time.
Identifying and solving conditional image leakage in image-to-video diffusion model
Zhao, M., Zhu, H., Xiang, C., Zheng, K., Li, C., and Zhu, J · 2024
Closest in time.
Identifying sensitive weights via post-quantization integral
Hu, Y., Huang, W., Liang, Z., Chen, C., Zhang, J., Zhu, J., and Chen, J · 2025
Closest in time.
QServe:w4a8KV4 quantization and system co-design for efficient LLM serving
Lin, Y., Tang, H., Yang, S., Zhang, Z., Xiao, G., Gan, C., and Han, S · 2025
Closest in time.
Sparse videogen: Accelerating video diffusion transformers with spatial-temporal sparsity
Xi, H., Yang, S., Zhao, Y., Xu, C., Li, M., Li, X., Lin, Y., Cai, H., Zhang, J., Li, D., et al · 2025
Closest in time.
Sage: A framework of precise retrieval for rag
Zhang, J., Li, G., and Su, J · 2025
Closest in time.
Zheng, K., Chen, Y., Chen, H., He, G., Liu, M.-Y., Zhu, J., and Zhang, Q · 2025
Closest in time.