Fetching the paper…
Reading the bibliography…
Linear attention is an efficient attention mechanism that has recently emerged as a promising alternative to conventional softmax attention.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering, 2018
Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A · 2018
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language, 2019
Bisk, Y., Zellers, R., Bras, R. L., Gao, J., and Choi, Y · 2019
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions, 2019
Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al · 2019
Earlier work this paper cites.
Winogrande: An adversarial winograd schema challenge at scale, 2019
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y · 2019
Earlier work this paper cites.
Socialiqa: Commonsense reasoning about social interactions, 2019
Sap, M., Rashkin, H., Chen, D., LeBras, R., and Choi, Y · 2019
Earlier work this paper cites.
Triton: an intermediate language and compiler for tiled neural network computations
Tillet, P., Kung, H.-T., and Cox, D. D · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?, 2019
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Rethinking attention with performers
Choromanski, K., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlós, T., Hawkins, P., Davis, J., Mohiuddin, A., Kaiser, L., Belanger, D., Colwell, L. J., and Weller, A · 2020
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding, 2021
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Rethinking positional encoding in language pre-training
Ke, G., He, D., and Liu, T.-Y · 2021
Earlier work this paper cites.
Random feature attention
Peng, H., Pappas, N., Yogatama, D., Schwartz, R., Smith, N. A., and Kong, L · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Cited alongside, same era.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Lu, Y., Pan, S., Wen, B., and Liu, Y · 2021
Cited alongside, same era.
Kerple: Kernelized relative positional embedding for length extrapolation, 2022
Chi, T.-C., Fan, T.-H., Ramadge, P. J., and Rudnicky, A. I · 2022
Cited alongside, same era.
FlashAttention: Fast and memory-efficient exact attention with IO-awareness
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Ré, C · 2022
Cited alongside, same era.
Flashattention-2: Faster attention with better parallelism and work partitioning
Dao, T · 2023
Later among the works it cites.
A framework for few-shot language model evaluation, 12 2023
Gao, L., Tow, J., Abbasi, B., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., Le Noac’h, A., Li, H., McDonell, K., Muennighoff, N., Ociepa, C., Phang, J., Reynolds, L., Schoelkopf, H., Skowron, A., Sutawika, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A · 2023
Later among the works it cites.
C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models, 2023
Huang, Y., Bai, Y., Zhu, Z., Zhang, J., Zhang, J., Su, T., Liu, J., Lv, C., Zhang, Y., Lei, J., Fu, Y., Sun, M., and He, J · 2023
Later among the works it cites.
MAP: Low-data regime multimodal learning with adapter-based pre-training and prompting
Li, W., Li, D., Li, W., Wang, Y., Jie, H., and Zhong, Y · 2023
Later among the works it cites.
Visual instruction tuning
Liu, H., Li, C., Wu, Q., and Lee, Y. J · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hua, W., Dai, Z., Liu, H., and Le, Q. V · 2022
Cited alongside, same era.
BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Li, J., Li, D., Xiong, C., and Hoi, S · 2022
Cited alongside, same era.
Linear video transformer with feature fixation, 2022
Lu, K., Liu, Z., Wang, J., Sun, W., Qin, Z., Li, D., Shen, X., Deng, H., Han, X., Dai, Y., and Zhong, Y · 2022
Cited alongside, same era.
Train short, test long: Attention with linear biases enables input length extrapolation
Press, O., Smith, N., and Lewis, M · 2022
Cited alongside, same era.
The devil in linear transformer
Qin, Z., Han, X., Sun, W., Li, D., Kong, L., Barnes, N., and Zhong, Y · 2022
Cited alongside, same era.
Opt: Open pre-trained transformer language models, 2022
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., Mihaylov, T., Ott, M., Shleifer, S., Shuster, K., Simig, D., Koura, P. S., Sridhar, A., Wang, T., and Zettlemoyer, L · 2022
Cited alongside, same era.
Linear complexity randomized self-attention mechanism
Zheng, L., Wang, C., and Kong, L · 2022
Cited alongside, same era.
Multimodal variational auto-encoder based audio-visual segmentation
Mao, Y., Zhang, J., Xiang, M., Zhong, Y., and Dai, Y · 2023
Later among the works it cites.
RWKV: Reinventing RNNs for the transformer era
Peng, B., Alcaide, E., Anthony, Q., Albalak, A., Arcadinho, S., Biderman, S., Cao, H., Cheng, X., Chung, M., Derczynski, L., Du, X., Grella, M., Gv, K., He, X., Hou, H., Kazienko, P., Kocon, J., Kong, J., Koptyra, B., Lau, H., Lin, J., Mantri, K. S. I., Mom, F., Saito, A., Song, G., Tang, X., Wind, J., Woźniak, S., Zhang, Z., Zhou, Q., Zhu, J., and Zhu, R.-J · 2023
Later among the works it cites.
Fine-grained audible video description
Shen, X., Li, D., Zhou, J., Qin, Z., He, B., Han, X., Li, A., Dai, Y., Kong, L., Wang, M., Qiao, Y., and Zhong, Y · 2023
Later among the works it cites.
Vicinity vision transformer
Sun, W., Qin, Z., Deng, H., Wang, J., Zhang, Y., Zhang, K., Barnes, N., Birchfield, S., Kong, L., and Zhong, Y · 2023
Later among the works it cites.
Efficient streaming language models with attention sinks, 2023
Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M · 2023
Later among the works it cites.
Gated linear attention transformers with hardware-efficient training, 2023
Yang, S., Wang, B., Shen, Y., Panda, R., and Kim, Y · 2023
Later among the works it cites.
Efficient attention via control variates
Zheng, L., Yuan, J., Wang, C., and Kong, L · 2023
Later among the works it cites.
Audio-visual segmentation with semantics, 2023
Zhou, J., Shen, X., Wang, J., Zhang, J., Sun, W., Zhang, J., Birchfield, S., Guo, D., Kong, L., Wang, M., and Zhong, Y · 2023
Later among the works it cites.
Improving audio-visual segmentation with bidirectional generation
Hao, D., Mao, Y., He, B., Han, X., Dai, Y., and Zhong, Y · 2024
Closest in time.
Exploring transformer extrapolation
Qin, Z., Zhong, Y., and Deng, H · 2024
Closest in time.