Fetching the paper…
Reading the bibliography…
In this work we study the presence of expert units in pre-trained Transformer Models (TM), and how they impact a model's performance.
Products of experts
G. E. Hinton · 1999
Earlier work this paper cites.
A neural probabilistic language model
Y. Bengio, R. Ducharme, P. Vincent, and C. Janvin · 2003
Earlier work this paper cites.
Brain changes in the development of expertise: Neuroanatomical and neurophysiological evidence about skill-based adaptations
N. Hill and W. Schneider · 2006
Earlier work this paper cites.
Visualizing data using t-SNE
L. van der Maaten and G. Hinton · 2008
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang · 2016
Earlier work this paper cites.
Fine-grained analysis of sentence embeddings using auxiliary prediction tasks
Y. Adi, E. Kermany, Y. Belinkov, O. Lavi, and Y. Goldberg · 2017
Earlier work this paper cites.
Network dissection: Quantifying interpretability of deep visual representations
D. Bau, B. Zhou, A. Khosla, A. Oliva, and A. Torralba · 2017
Earlier work this paper cites.
Toward controlled generation of text
Z. Hu, Z. Yang, X. Liang, R. Salakhutdinov, and E. P. Xing · 2017
Earlier work this paper cites.
Plug & play generative networks: Conditional iterative generation of images in latent space
A. Nguyen, J. Yosinski, Y. Bengio, A. Dosovitskiy, and J. Clune · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Łukasz Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Sanity checks for saliency maps
J. Adebayo, J. Gilmer, M. Muelly, I. Goodfellow, M. Hardt, and B. Kim · 2018
Earlier work this paper cites.
What you can cram into a single vector: Probing sentence embeddings for linguistic properties
A. Conneau, G. Kruszewski, G. Lample, L. Barrault, and M. Baroni · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Net2vec: Quantifying and explaining how concepts are encoded by filters in deep neural networks
R. Fong and A. Vedaldi · 2018
Cited alongside, same era.
Interpreting recurrent and attention-based neural models: a case study on natural language inference
R. Ghaeini, X. Fern, and P. Tadepalli · 2018
Cited alongside, same era.
Universal language model fine-tuning for text classification
J. Howard and S. Ruder · 2018
Cited alongside, same era.
Know what you don’t know: Unanswerable questions for squad
P. Rajpurkar, R. Jia, and P. Liang · 2018
Cited alongside, same era.
Gan dissection: Visualizing and understanding generative adversarial networks
D. Bau, J.-Y. Zhu, H. Strobelt, Z. Bolei, J. B. Tenenbaum, W. T. Freeman, and A. Torralba · 2019
Cited alongside, same era.
A multi-task approach for disentangling syntax and semantics in sentence representations
Roberta: A robustly optimized bert pretraining approach
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov · 2019
Later among the works it cites.
Excitatory and inhibitory subnetworks are equally selective during decision-making and emerge simultaneously during learning
F. Najafi, G. F. Elsayed, R. Cao, E. Pnevmatikakis, P. E. Latham, J. P. Cunningham, and A. K. Churchland · 2019
Later among the works it cites.
Probing neural network comprehension of natural language arguments
T. Niven and H.-Y. Kao · 2019
Later among the works it cites.
Megatron-lm
Nvidia · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever · 2019
Later among the works it cites.
Adversarial decomposition of text representation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Chen, Q. Tang, S. Wiseman, and K. Gimpel · 2019
Cited alongside, same era.
What does bert look at? an analysis of bert’s attention
K. Clark, U. Khandelwal, O. Levy, and C. D. Manning · 2019
Cited alongside, same era.
Explanations can be manipulated and geometry is to blame
A.-K. Dombrowski, M. Alber, C. J. Anders, M. Ackermann, K.-R. Müller, and P. Kessel · 2019
Cited alongside, same era.
The curious case of neural text degeneration
A. Holtzman, J. Buys, M. Forbes, and Y. Choi · 2019
Cited alongside, same era.
CTRL - A Conditional Transformer Language Model for Controllable Generation
N. S. Keskar, B. McCann, L. Varshney, C. Xiong, and R. Socher · 2019
Cited alongside, same era.
Cross-lingual language model pretraining
G. Lample and A. Conneau · 2019
Cited alongside, same era.
Linguistic knowledge and transferability of contextual representations
N. F. Liu, M. Gardner, Y. Belinkov, M. E. Peters, and N. A. Smith · 2019
Cited alongside, same era.
A. Romanov, A. Rumshisky, A. Rogers, and D. Donahue · 2019
Later among the works it cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
V. Sanh, L. Debut, J. Chaumond, and T. Wolf · 2019
Later among the works it cites.
Just “onesec” for producing multilingual sense-annotated data
B. Scarlini, T. Pasini, and R. Navigli · 2019
Later among the works it cites.
What do you learn from context? probing for sentence structure in contextualized word representations
I. Tenney, P. Xia, B. Chen, A. Wang, A. Poliak, R. T. McCoy, N. Kim, B. V. Durme, S. R. Bowman, D. Das, and E. Pavlick · 2019
Later among the works it cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman · 2019
Later among the works it cites.
Huggingface’s transformers: State-of-the-art natural language processing
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, and J. Brew · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. Salakhutdinov, and Q. V. Le · 2019
Later among the works it cites.