Fetching the paper…
Reading the bibliography…
Transformer is a ubiquitous model for natural language processing and has attracted wide attentions in computer vision.
Roberta: A robustly optimized BERT pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 1907
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
“why should i trust you?”: Explaining the predictions of any classifier
Ribeiro, M. T., Singh, S., and Guestrin, C · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z · 2016
Earlier work this paper cites.
Lite transformer with long-short range attention
Wu, Z., Liu, Z., Lin, J., Lin, Y., and Han, S · 2016
Earlier work this paper cites.
Convolutional sequence to sequence learning
Gehring, J., Auli, M., Grangier, D., Yarats, D., and Dauphin, Y. N · 2017
Earlier work this paper cites.
Densely connected convolutional networks
Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K. Q · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Image transformer
Parmar, N., Vaswani, A., Uszkoreit, J., Kaiser, L., Shazeer, N., Ku, A., and Tran, D · 2018
Earlier work this paper cites.
An analysis of encoder representations in transformer-based machine translation
Raganato, A., Tiedemann, J., et al · 2018
Cited alongside, same era.
Self-attention with relative position representations
Shaw, P., Uszkoreit, J., and Vaswani, A · 2018
Cited alongside, same era.
Why self-attention? a targeted evaluation of neural machine translation architectures
Tang, G., Müller, M., Gonzales, A. R., and Sennrich, R · 2018
Cited alongside, same era.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2018
Cited alongside, same era.
Attention augmented convolutional networks
Bello, I., Zoph, B., Vaswani, A., Shlens, J., and Le, Q. V · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Novel positional encodings to enable tree-based transformers
Shiv, V. and Quirk, C · 2019
Later among the works it cites.
The evolved transformer
So, D., Le, Q., and Liang, C · 2019
Later among the works it cites.
Adaptive attention span in transformers
Sukhbaatar, S., Grave, É., Bojanowski, P., and Joulin, A · 2019
Later among the works it cites.
Attention is not not explanation
Wiegreffe, S. and Pinter, Y · 2019
Later among the works it cites.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Scao, T. L., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. M · 2019
Later among the works it cites.
ERASER: A benchmark to evaluate rationalized NLP models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Attention is not explanation
Jain, S. and Wallace, B. C · 2019
Cited alongside, same era.
Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting
Li, S., Jin, X., Xuan, Y., Zhou, X., Chen, W., Wang, Y.-X., and Yan, X · 2019
Cited alongside, same era.
Stand-alone self-attention in vision models
Parmar, N., Ramachandran, P., Vaswani, A., Bello, I., Levskaya, A., and Shlens, J · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2019
Cited alongside, same era.
DeYoung, J., Jain, S., Rajani, N. F., Lehman, E., Xiong, C., Socher, R., and Wallace, B. C · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Later among the works it cites.
Realformer: Transformer likes residual attention, 2020
He, R., Ravula, A., Kanagal, B., and Ainslie, J · 2020
Later among the works it cites.
Reformer: The efficient transformer
Kitaev, N., Kaiser, L., and Levskaya, A · 2020
Later among the works it cites.
Synthesizer: Rethinking self-attention in transformer models
Tay, Y., Bahri, D., Metzler, D., Juan, D.-C., Zhao, Z., and Zheng, C · 2020
Later among the works it cites.