Fetching the paper…
Reading the bibliography…
Large Transformer-based models have exhibited superior performance in various natural language processing and computer vision tasks.
Roberta: A robustly optimized bert pretraining approach
Liu, Y · 1907
Earlier work this paper cites.
Structured pruning of a bert-based question answering model
McCarley, J · 1910
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Lai, T. L · 1985
Earlier work this paper cites.
Optimal brain damage
LeCun, Y · 1990
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
Qian, N · 1999
Earlier work this paper cites.
On iterative neural network pruning, reinitialization, and the similarity of masks
Paganini, M · 2001
Earlier work this paper cites.
Poor man’s bert: Smaller and faster transformer models
Sajjad, H · 2004
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Dolan, W. B · 2005
Earlier work this paper cites.
The second PASCAL recognising textual entailment challenge
Bar-Haim, R · 2006
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Dagan, I · 2006
Earlier work this paper cites.
The third PASCAL recognizing textual entailment challenge
Giampiccolo, D · 2007
Earlier work this paper cites.
The fifth pascal recognizing textual entailment challenge
Bentivogli, L · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
The winograd schema challenge
Levesque, H · 2012
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R · 2013
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Han, S · 2015
Cited alongside, same era.
Tensorflow: A system for large-scale machine learning
Abadi, M · 2016
Cited alongside, same era.
Eie: Efficient inference engine on compressed deep neural network
Han, S · 2016
Cited alongside, same era.
SQuAD: 100,000+ questions for machine comprehension of text
Rajpurkar, P · 2016
Cited alongside, same era.
SQuAD: 100,000+ questions for machine comprehension of text
Rajpurkar, P · 2016
Cited alongside, same era.
SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation
Cer, D · 2017
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A · 2019
Later among the works it cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A · 2019
Later among the works it cites.
Neural network acceptability judgments
Warstadt, A · 2019
Later among the works it cites.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T · 2019
Later among the works it cites.
Language models are few-shot learners
Brown, T. B · 2020
Later among the works it cites.
The lottery ticket hypothesis for pre-trained BERT networks
Chen, T · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pruning convolutional neural networks for resource efficient inference
Molchanov, P · 2017
Cited alongside, same era.
Learning sparse neural networks through l_0 regularization
Louizos, C · 2018
Cited alongside, same era.
Piggyback: Adding multiple tasks to a single, fixed network by learning to mask
Mallya, A · 2018
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Williams, A · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J · 2019
Cited alongside, same era.
Global sparse momentum SGD for pruning very deep neural networks
Ding, X · 2019
Cited alongside, same era.
Dosovitskiy, A · 2020
Later among the works it cites.
Reducing transformer depth on demand with structured dropout
Fan, A · 2020
Later among the works it cites.
The Microsoft toolkit of multi-task deep neural networks for natural language understanding
Liu, X · 2020
Later among the works it cites.
Comparing rewinding and fine-tuning in neural network pruning
Renda, A · 2020
Later among the works it cites.
Movement pruning: Adaptive sparsity by fine-tuning
Sanh, V · 2020
Later among the works it cites.
Pruning neural machine translation for speed using group lasso
Behnke, M · 2021
Later among the works it cites.
Block pruning for faster transformers
Lagunas, F · 2021
Later among the works it cites.
Super tickets in pre-trained language models: From model compression to improving generalization
Liang, C · 2021
Later among the works it cites.
Prune once for all: Sparse pre-trained language models
Zafrir, O · 2021
Later among the works it cites.
A biased graph neural network sampler with near-optimal regret
Zhang, Q · 2021
Later among the works it cites.
To prune, or not to prune: Exploring the efficacy of pruning for model compression
Zhu, M · 2021
Later among the works it cites.