Fetching the paper…
Reading the bibliography…
Transformer-based language models have found many diverse applications requiring them to process sequences of increasing length.
Generating long sequences with sparse transformers
Child, R., Gray, S., Radford, A., and Sutskever, I · 1904
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Longformer: The long-document transformer
Beltagy, I., Peters, M. E., and Cohan, A · 2004
Earlier work this paper cites.
Rethinking attention with performers
Choromanski, K., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlós, T., Hawkins, P., Davis, J., Mohiuddin, A., Kaiser, L., Belanger, D., Colwell, L. J., and Weller, A · 2009
Earlier work this paper cites.
Efficient transformers: A survey
Tay, Y., Dehghani, M., Bahri, D., and Metzler, D · 2009
Earlier work this paper cites.
The human knowledge compression contest
Hutter, M · 2012
Earlier work this paper cites.
Practical and optimal lsh for angular distance, 2015
Andoni, A., Indyk, P., Laarhoven, T., Razenshteyn, I., and Schmidt, L · 2015
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Earlier work this paper cites.
Are sixteen heads really better than one?
Michel, P., Levy, O., and Neubig, G · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners, 2019
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Earlier work this paper cites.
Triton: An intermediate language and compiler for tiled neural network computations
Tillet, P., Kung, H. T., and Cox, D · 2019
Earlier work this paper cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Voita, E., Talbot, D., Moiseev, F., Sennrich, R., and Titov, I · 2019
Cited alongside, same era.
Losing heads in the lottery: Pruning transformer attention in neural machine translation
Behnke, M. and Heafield, K · 2020
Cited alongside, same era.
OpenWebText2 dataset, as part of ‘the Pile: An 800gb dataset of diverse text for language modeling‘
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., Presser, S., and Leahy, C · 2020
Cited alongside, same era.
Power-bert: Accelerating BERT inference via progressive word-vector elimination
Goyal, S., Choudhury, A. R., Raje, S., Chakaravarthy, V. T., Sabharwal, Y., and Verma, A · 2020
Cited alongside, same era.
Transformers are RNNs: Fast autoregressive transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2020
Cited alongside, same era.
Improving language models by retrieving from trillions of tokens
Borgeaud, S., Mensch, A., Hoffmann, J., Cai, T., Rutherford, E., Millican, K., van den Driessche, G., Lespiau, J., Damoc, B., Clark, A., de Las Casas, D., Guy, A., Menick, J., Ring, R., Hennigan, T., Huang, S., Maggiore, L., Jones, C., Cassirer, A., Brock, A., Paganini, M., Irving, G., Vinyals, O., Osindero, S., Simonyan, K., Rae, J. W., Elsen, E., and Sifre, L · 2021
Later among the works it cites.
Differentiable subset pruning of transformer heads
Li, J., Cotterell, R., and Sachan, M · 2021
Later among the works it cites.
Random feature attention
Peng, H., Pappas, N., Yogatama, D., Schwartz, R., Smith, N. A., and Kong, L · 2021
Later among the works it cites.
Spatten: Efficient sparse attention architecture with cascade token and head pruning
Wang, H., Zhang, Z., and Han, S · 2021
Later among the works it cites.
Ethical and social risks of harm from language models
Weidinger, L., Mellor, J., Rauh, M., Griffin, C., Uesato, J., Huang, P., Cheng, M., Glaese, M., Balle, B., Kasirzadeh, A., Kenton, Z., Brown, S., Hawkins, W., Stepleton, T., Biles, C., Birhane, A., Haas, J., Rimell, L., Hendricks, L. A., Isaac, W., Legassick, S., Irving, G., and Gabriel, I · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reformer: The efficient transformer
Kitaev, N., Kaiser, L., and Levskaya, A · 2020
Cited alongside, same era.
A mixture of h - 1 heads is better than h heads
Peng, H., Schwartz, R., Li, D., and Smith, N. A · 2020
Cited alongside, same era.
Fixed encoder self-attention patterns in transformer-based machine translation
Raganato, A., Scherrer, Y., and Tiedemann, J · 2020
Cited alongside, same era.
Linformer: Self-attention with linear complexity
Wang, S., Li, B. Z., Khabsa, M., Fang, H., and Ma, H · 2020
Cited alongside, same era.
Big bird: Transformers for longer sequences
Zaheer, M., Guruganesh, G., Dubey, K. A., Ainslie, J., Alberti, C., Ontañón, S., Pham, P., Ravula, A., Wang, Q., Yang, L., and Ahmed, A · 2020
Cited alongside, same era.
Scheduled drophead: A regularization method for transformer models
Zhou, W., Ge, T., Wei, F., Zhou, M., and Xu, K · 2020
Cited alongside, same era.
On the dangers of stochastic parrots: Can language models be too big?
Bender, E. M., Gebru, T., McMillan-Major, A., and Shmitchell, S · 2021
Cited alongside, same era.
Later among the works it cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Ré, C · 2022
Later among the works it cites.
Learned token pruning for transformers
Kim, S., Shen, S., Thorsley, D., Gholami, A., Kwon, W., Hassoun, J., and Keutzer, K · 2022
Later among the works it cites.
cosformer: Rethinking softmax in attention
Qin, Z., Sun, W., Deng, H., Li, D., Wei, Y., Lv, B., Yan, J., Kong, L., and Zhong, Y · 2022
Later among the works it cites.
Memorizing transformers, 2022
Wu, Y., Rabe, M. N., Hutchins, D., and Szegedy, C · 2022
Later among the works it cites.
Structured pruning learns compact and accurate models
Xia, M., Zhong, Z., and Chen, D · 2022
Later among the works it cites.
Linear complexity randomized self-attention mechanism
Zheng, L., Wang, C., and Kong, L · 2022
Later among the works it cites.