GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S · 2018
Cited alongside, same era.
Attention augmented convolutional networks
Bello, I., Zoph, B., Vaswani, A., Shlens, J., and Le, Q. V · 2019
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Tensorly: Tensor learning in python
Kossaifi, J., Panagakis, Y., Anandkumar, A., and Pantic, M · 2019
Cited alongside, same era.
Are sixteen heads really better than one?
Michel, P., Levy, O., and Neubig, G · 2019
Cited alongside, same era.
fairseq: A fast, extensible toolkit for sequence modeling
Ott, M., Edunov, S., Baevski, A., Fan, A., Gross, S., Ng, N., Grangier, D., and Auli, M · 2019
Cited alongside, same era.
Stand-alone self-attention in vision models
Ramachandran, P., Parmar, N., Vaswani, A., Bello, I., Levskaya, A., and Shlens, J · 2019
Cited alongside, same era.
Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter
Original
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 2019
Cited alongside, same era.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Voita, E., Talbot, D., Moiseev, F., Sennrich, R., and Titov, I · 2019
Cited alongside, same era.