Fetching the paper…
Reading the bibliography…
Transformer-based models have become ubiquitous in natural language processing thanks to their large capacity, innate parallelism and high performance.
GloVe: Global vectors for word representation
J. Pennington, R. Socher, and C. D. Manning · 2014
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang · 2016
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Understanding back-translation at scale
S. Edunov, M. Ott, M. Auli, and D. Grangier · 2018
Earlier work this paper cites.
The narrativeqa reading comprehension challenge
T. Kočiskỳ, J. Schwarz, P. Blunsom, C. Dyer, K. M. Hermann, G. Melis, and E. Grefenstette · 2018
Earlier work this paper cites.
N. Parmar, A. Vaswani, J. Uszkoreit, L. Kaiser, N. Shazeer, and A. Ku · 2018
Earlier work this paper cites.
Deep contextualized word representations
M. E. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer · 2018
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Z. Yang, P. Qi, S. Zhang, Y. Bengio, W. W. Cohen, R. R. Salakhutdinov, and C. D. Manning · 2018
Earlier work this paper cites.
Generating long sequences with sparse transformers
R. Child, S. Gray, A. Radford, and I. Sutskever · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
A structural probe for finding syntax in word representations
J. Hewitt and C. D. Manning · 2019
Cited alongside, same era.
Music transformer
C.-Z. A. Huang, A. Vaswani, J. Uszkoreit, I. Simon, C. Hawthorne, N. Shazeer, A. M. Dai, M. D. Hoffman, M. Dinculescu, and D. Eck · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov · 2019
Cited alongside, same era.
Compositional questions do not necessitate multi-hop reasoning
S. Min, E. Wallace, S. Singh, M. Gardner, H. Hajishirzi, and L. Zettlemoyer · 2019
Cited alongside, same era.
Stand-alone self-attention in vision models
N. Parmar, P. Ramachandran, A. Vaswani, I. Bello, A. Levskaya, and J. Shlens · 2019
Cited alongside, same era.
Language models as knowledge bases?
Bp-transformer: Modelling long-range context via binary partitioning
Z. Ye, Q. Guo, Q. Gan, X. Qiu, and Z. Zhang · 2019
Later among the works it cites.
Longformer: The long-document transformer
I. Beltagy, M. E. Peters, and A. Cohan · 2020
Closest in time.
Injecting numerical reasoning skills into language models
M. Geva, A. Gupta, and J. Berant · 2020
Closest in time.
Dense passage retrieval for open-domain question answering
V. Karpukhin, B. Oğuz, S. Min, L. Wu, S. Edunov, D. Chen, and W.-t. Yih · 2020
Closest in time.
Reformer: The efficient transformer
N. Kitaev, L. Kaiser, and A. Levskaya · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
F. Petroni, T. Rocktäschel, P. Lewis, A. Bakhtin, Y. Wu, A. Miller, and S. Riedel · 2019
Cited alongside, same era.
Blockwise self-attention for long document understanding
J. Qiu, H. Ma, O. Levy, S. W.-t. Yih, S. Wang, and J. Tang · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2019
Cited alongside, same era.
Huggingface’s transformers: State-of-the-art natural language processing
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, and J. Brew · 2019
Cited alongside, same era.
Xlnet: Generalized autoregressive pretraining for language understanding
Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. R. Salakhutdinov, and Q. V. Le · 2019
Cited alongside, same era.
Deep learning for symbolic mathematics
G. Lample and F. Charton · 2020
Closest in time.
Albert: A lite bert for self-supervised learning of language representations
Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, and R. Soricut · 2020
Closest in time.
Compressive transformers for long-range sequence modelling
J. W. Rae, A. Potapenko, S. M. Jayakumar, C. Hillier, and T. P. Lillicrap · 2020
Closest in time.
How much knowledge can you pack into the parameters of a language model?
A. Roberts, C. Raffel, and N. Shazeer · 2020
Closest in time.
Efficient content-based sparse attention with routing transformers
A. Roy, M. Saffar, A. Vaswani, and D. Grangier · 2020
Closest in time.