Fetching the paper…
Reading the bibliography…
The Transformer architecture has revolutionized deep learning on sequential data, becoming ubiquitous in state-of-the-art solutions for a wide variety of applications.
Generating long sequences with sparse transformers
Child, R., Gray, S., Radford, A., and Sutskever, I · 1904
Earlier work this paper cites.
Reducing BERT pre-training time from 3 days to 76 minutes
You, Y., Li, J., Hseu, J., Song, X., Demmel, J., and Hsieh, C · 1904
Earlier work this paper cites.
Parallel prefix computation
Ladner, R. E. and Fischer, M. J · 1980
Earlier work this paper cites.
Achieving logarithmic growth of temporal and spatial complexity in reverse automatic differentiation
Griewank, A · 1992
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn Treebank
Marcus, M. P., Santorini, B., and Marcinkiewicz, M. A · 2004
Earlier work this paper cites.
Evaluating Derivatives: Principles and Techniques of Algorithmic Differentiation, Second Edition
Griewank, A. and Walther, A · 2008
Earlier work this paper cites.
Rethinking attention with Performers
Choromanski, K., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J., Mohiuddin, A., Kaiser, L., Belanger, D., Colwell, L., and Weller, A · 2009
Earlier work this paper cites.
Large text compression benchmark, 2009
Mahoney, M · 2009
Earlier work this paper cites.
Thinking in parallel: Some basic data-parallel algorithms and techniques
Vishkin, U · 2010
Earlier work this paper cites.
Long range arena: A benchmark for efficient transformers
Tay, Y., Dehghani, M., Abnar, S., Shen, Y., Bahri, D., Pham, P., Rao, J., Yang, L., Ruder, S., and Metzler, D · 2011
Earlier work this paper cites.
Learning phrase representations using RNN encoder–decoder for statistical machine translation
Cho, K., van Merriënboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y · 2014
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Mané, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanhoucke, V., Vasudevan, V., Viégas, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., and Zheng, X · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Cited alongside, same era.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Cited alongside, same era.
Training deep nets with sublinear memory cost
Chen, T., Xu, B., Zhang, C., and Guestrin, C · 2016
Cited alongside, same era.
fairseq: A fast, extensible toolkit for sequence modeling, 2019
Ott, M., Edunov, S., Baevski, A., Fan, A., Gross, S., Ng, N., Grangier, D., and Auli, M · 2019
Later among the works it cites.
Stabilizing transformers for reinforcement learning
Parisotto, E., Song, H. F., Rae, J. W., Pascanu, R., Gulcehre, C., Jayakumar, S. M., Jaderberg, M., Kaufman, R. L., Clark, A., Noury, S., et al · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Later among the works it cites.
Language models are few-shot learners, 2020
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Closest in time.
Transformers are RNNs: Fast autoregressive transformers with linear attention
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hendrycks, D. and Gimpel, K · 2016
Cited alongside, same era.
Automatic differentiation in pytorch
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Cited alongside, same era.
Neural ordinary differential equations
Chen, R. T. Q., Rubanova, Y., Bettencourt, J., and Duvenaud, D. K · 2018
Cited alongside, same era.
Scaling neural machine translation
Ott, M., Edunov, S., Grangier, D., and Auli, M · 2018
Cited alongside, same era.
Parmar, N., Vaswani, A., Uszkoreit, J., Kaiser, L., Shazeer, N., and Ku, A · 2018
Cited alongside, same era.
Factorized attention: Self-attention with linear complexities
Shen, Z., Zhang, M., Yi, S., Yan, J., and Zhao, H · 2018
Cited alongside, same era.
Transformer-XL: Language modeling with longer-term dependency, 2019
Dai, Z., Yang, Z., Yang, Y., Cohen, W. W., Carbonell, J., Le, Q. V., and Salakhutdinov, R · 2019
Cited alongside, same era.
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2020
Closest in time.
Reformer: The efficient transformer
Kitaev, N., Kaiser, L., and Levskaya, A · 2020
Closest in time.
ALBERT: A lite BERT for self-supervised learning of language representations
Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., and Soricut, R · 2020
Closest in time.
Linear attention mechanism: An efficient attention for semantic segmentation
Li, R., Duan, C., and Zheng, S · 2020
Closest in time.
Efficient content-based sparse attention with routing transformers
Roy, A., Saffar, M., Vaswani, A., and Grangier, D · 2020
Closest in time.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 2020
Closest in time.
Lite transformer with long-short range attention
Wu*, Z., Liu*, Z., Lin, J., Lin, Y., and Han, S · 2020
Closest in time.
On layer normalization in the transformer architecture, 2020
Xiong, R., Yang, Y., He, D., Zheng, K., Zheng, S., Xing, C., Zhang, H., Lan, Y., Wang, L., and Liu, T.-Y · 2020
Closest in time.