Fetching the paper…
Reading the bibliography…
Transformer networks have revolutionized NLP representation learning since they were introduced.
Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R Thomas McCoy, Najoung Kim, Benjamin Van Durme, Samuel R Bowman, Dipanjan Das, et al. 2019 · 1905
Earlier work this paper cites.
Word embedding visualization via dictionary learning
Juexiao Zhang, Yubei Chen, Brian Cheung, and Bruno A Olshausen. 2019 · 1910
Earlier work this paper cites.
Brown corpus manual
W. N. Francis and H. Kucera. 1979 · 1979
Earlier work this paper cites.
A fast iterative shrinkage-thresholding algorithm for linear inverse problems
Amir Beck and Marc Teboulle. 2009 · 2009
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2020 · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer. 2011 · 2011
Earlier work this paper cites.
Transformer interpretability beyond attention visualization
Hila Chefer, Shir Gur, and Lior Wolf. 2020 · 2012
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Matthew D Zeiler and Rob Fergus. 2014 · 2014
Earlier work this paper cites.
Sparse overcomplete word vector representations
Manaal Faruqui, Yulia Tsvetkov, Dani Yogatama, Chris Dyer, and Noah A. Smith. 2015 · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Cited alongside, same era.
"why should I trust you?": Explaining the predictions of any classifier
Marco Túlio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Cited alongside, same era.
Dynamic routing between capsules
Sara Sabour, Nicholas Frosst, and Geoffrey E Hinton. 2017 · 2017
Cited alongside, same era.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D. Manning. 2019 · 2019
Later among the works it cites.
Linguistic knowledge and transferability of contextual representations
Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, and Noah A. Smith. 2019 · 2019
Later among the works it cites.
Visualizing and measuring the geometry of BERT
Emily Reif, Ann Yuan, Martin Wattenberg, Fernanda B. Viégas, Andy Coenen, Adam Pearce, and Been Kim. 2019 · 2019
Later among the works it cites.
How can we know what language models know
Zhengbao Jiang, Frank F. Xu, Jun Araki, and Graham Neubig. 2020 · 2020
Later among the works it cites.
A primer in bertology: What we know about how BERT works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2020 · 2020
Later among the works it cites.
Pretrained bert base model (12 layers)
2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Linear algebraic structure of word senses, with applications to polysemy
Sanjeev Arora, Yuanzhi Li, Yingyu Liang, Tengyu Ma, and Andrej Risteski. 2018 · 2018
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
How contextual are contextualized word representations? comparing the geometry of bert, elmo, and GPT-2 embeddings
Kawin Ethayarajh. 2019 · 2019
Cited alongside, same era.
Generic attention-model explainability for interpreting bi-modal and encoder-decoder transformers
Hila Chefer, Shir Gur, and Lior Wolf. 2021 · 2021
Closest in time.
How to represent part-whole hierarchies in a neural network
Geoffrey Hinton. 2021 · 2021
Closest in time.