Fetching the paper…
Reading the bibliography…
Transformer-based NLP models are trained using hundreds of millions or even billions of parameters, limiting their applicability in computationally constrained environments.
Distilling task-specific knowledge from BERT into simple neural networks,
R. Tang, Y. Lu, L. Liu, L. Mou, O. Vechtomova, J. Lin, · 1903
Earlier work this paper cites.
Are sixteen heads really better than one?,
P. Michel, O. Levy, G. Neubig, · 1905
Earlier work this paper cites.
1905
Earlier work this paper cites.
1906
Earlier work this paper cites.
Roberta: A robustly optimized BERT pretraining approach,
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, V. Stoyanov, · 1907
Earlier work this paper cites.
1908
Earlier work this paper cites.
1909
Earlier work this paper cites.
1909
Earlier work this paper cites.
1909
Earlier work this paper cites.
1909
Earlier work this paper cites.
1909
Earlier work this paper cites.
J. S. McCarley, Pruning a bert-based question answering model, 2019. arXiv:1910.06360
1910
Earlier work this paper cites.
1910
Earlier work this paper cites.
O. Zafrir, G. Boudoukh, P. Izsak, M. Wasserblat, Q8bert: Quantized 8bit bert, 2019. arXiv:1910.06188
1910
Earlier work this paper cites.
SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation,
D. Cer, M. Diab, E. Agirre, I. Lopez-Gazpio, L. Specia, · 2001
Earlier work this paper cites.
An introduction to variable and feature selection,
I. Guyon, A. Elisseeff, · 2003
Earlier work this paper cites.
2004
Earlier work this paper cites.
2004
Earlier work this paper cites.
2004
Earlier work this paper cites.
2005
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases,
W. B. Dolan, C. Brockett, · 2005
Cited alongside, same era.
A fast learning algorithm for deep belief nets,
G. E. Hinton, S. Osindero, Y. W. Teh, · 2006
Cited alongside, same era.
2006
Cited alongside, same era.
2006
Cited alongside, same era.
The fifth pascal recognizing textual entailment challenge,
L. Bentivogli, I. Dagan, H. T. Dang, D. Giampiccolo, B. Magnini, · 2009
Cited alongside, same era.
Linguistic knowledge and transferability of contextual representations,
N. F. Liu, M. Gardner, Y. Belinkov, M. E. Peters, N. A. Smith, · 2019
Later among the works it cites.
One size does not fit all: Comparing NMT representations of different granularities,
N. Durrani, F. Dalvi, H. Sajjad, Y. Belinkov, P. Nakov, · 2019
Later among the works it cites.
What is one grain of sand in the desert? analyzing individual neurons in deep nlp models,
F. Dalvi, N. Durrani, H. Sajjad, Y. Belinkov, D. A. Bau, J. Glass, · 2019
Later among the works it cites.
Huggingface’s transformers: State-of-the-art natural language processing,
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, J. Brew, · 2019
Later among the works it cites.
Analyzing redundancy in pretrained transformer models,
F. Dalvi, H. Sajjad, N. Durrani, Y. Belinkov, · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Recursive deep models for semantic compositionality over a sentiment treebank,
R. Socher, A. Perelygin, J. Wu, J. Chuang, C. D. Manning, A. Ng, C. Potts, · 2013
Cited alongside, same era.
2014
Cited alongside, same era.
SQuAD: 100,000+ questions for machine comprehension of text,
P. Rajpurkar, J. Zhang, K. Lopyrev, P. Liang, · 2016
Cited alongside, same era.
What do Neural Machine Translation Models Learn about Morphology?,
Y. Belinkov, N. Durrani, F. Dalvi, H. Sajjad, J. Glass, · 2017
Cited alongside, same era.
Understanding and improving morphological learning in the neural machine translation decoder,
F. Dalvi, N. Durrani, H. Sajjad, Y. Belinkov, S. Vogel, · 2017
Cited alongside, same era.
Evaluating layers of representation in neural machine translation on part-of-speech and semantic tagging tasks,
Y. Belinkov, L. Màrquez, H. Sajjad, N. Durrani, F. Dalvi, J. Glass, · 2017
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding,
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, S. Bowman, · 2018
Cited alongside, same era.
Faster and just as accurate: A simple decomposition for transformer models,
Q. Cao, H. Trivedi, A. Balasubramanian, et al., · 2020
Closest in time.
"mobilebert: Task-agnostic compression of bert by progressive knowledge transfer",
Z. Sun, H. Yu, X. Song, R. Liu, Y. Yang, D. Zhou, · 2020
Closest in time.
Comparing rewinding and fine-tuning in neural network pruning,
A. Renda, J. Frankle, M. Carbin, · 2020
Closest in time.
Bert-of-theseus: Compressing bert by progressive module replacing,
C. Xu, W. Zhou, T. Ge, F. Wei, M. Zhou, · 2020
Closest in time.
On the linguistic representational power of neural machine translation models,
Y. Belinkov, N. Durrani, F. Dalvi, H. Sajjad, J. Glass, · 2020
Closest in time.
Interpretability and analysis in neural NLP,
Y. Belinkov, S. Gehrmann, E. Pavlick, · 2020
Closest in time.
Investigating transferability in pretrained language models,
A. Tamkin, T. Singh, D. Giovanardi, N. D. Goodman, · 2020
Closest in time.
What happens to bert embeddings during fine-tuning?,
A. Merchant, E. Rahimtoroghi, E. Pavlick, I. Tenney, · 2020
Closest in time.
Analyzing individual neurons in pre-trained language models,
N. Durrani, H. Sajjad, F. Dalvi, Y. Belinkov, · 2020
Closest in time.
Greedy layer pruning: Decreasing inference time of transformer models,
D. Peer, S. Stabinger, S. Engl, A. J. Rodríguez-Sánchez, · 2021
Closest in time.
Neuron-level Interpretation of Deep NLP Models: A Survey,
H. Sajjad, N. Durrani, F. Dalvi, · 2021
Closest in time.
How transfer learning impacts linguistic knowledge in deep NLP models?,
N. Durrani, H. Sajjad, F. Dalvi, · 2021
Closest in time.
D. Arps, Y. Samih, L. Kallmeyer, H. Sajjad, Probing for constituency structure in neural language models, 2022. doi: 10.48550/ARXIV.2204.06201
2022
Closest in time.
Discovering latent concepts learned in BERT,
F. Dalvi, A. R. Khan, F. Alam, N. Durrani, J. Xu, H. Sajjad, · 2022
Closest in time.
Analyzing encoded concepts in transformer language models,
H. Sajjad, N. Durrani, F. Dalvi, F. Alam, A. R. Khan, J. Xu, · 2022
Closest in time.