Fetching the paper…
Reading the bibliography…
With the rapid adoption of machine learning (ML), a number of domains now use the approach of fine tuning models which were pre-trained on a large corpus of data.
Patient knowledge distillation for bert model compression
S. Sun, Y. Cheng, Z. Gan, and J. Liu · 1908
Earlier work this paper cites.
Connection-level analysis and modeling of network traffic
S. Sarvotham, R. Riedi, and R. Baraniuk · 2001
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
W. B. Dolan and C. Brockett · 2005
Earlier work this paper cites.
Optimization of collective communication operations in mpich
R. Thakur, R. Rabenseifner, and W. Gropp · 2005
Earlier work this paper cites.
Numerical Optimization
J. Nocedal and S. J. Wright · 2006
Earlier work this paper cites.
The lottery ticket hypothesis for pre-trained bert networks
T. Chen, J. Frankle, S. Chang, S. Liu, Y. Zhang, Z. Wang, and M. Carbin · 2007
Earlier work this paper cites.
Learning word vectors for sentiment analysis
A. L. Maas, R. E. Daly, P. T. Pham, D. Huang, A. Y. Ng, and C. Potts · 2011
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
R. Socher, A. Perelygin, J. Wu, J. Chuang, C. D. Manning, A. Ng, and C. Potts · 2013
Earlier work this paper cites.
Cnn features off-the-shelf: an astounding baseline for recognition
A. Sharif Razavian, H. Azizpour, J. Sullivan, and S. Carlsson · 2014
Earlier work this paper cites.
Teaching machines to read and comprehend, 2015
K. M. Hermann, T. Kočiský, E. Grefenstette, L. Espeholt, W. Kay, M. Suleyman, and P. Blunsom · 2015
Earlier work this paper cites.
Autograd: Effortless gradients in numpy
D. Maclaurin, D. Duvenaud, and R. P. Adams · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
X. Zhang, J. Zhao, and Y. LeCun · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Zero-shot transfer learning for event extraction
L. Huang, H. Ji, K. Cho, and C. R. Voss · 2017
Earlier work this paper cites.
Svcca: Singular vector canonical correlation analysis for deep understanding and improvement
M. Raghu, J. Gilmer, J. Yosinski, and J. Sohl-Dickstein · 2017
Earlier work this paper cites.
Zero-shot object detection
A. Bansal, K. Sikka, G. Sharma, R. Chellappa, and A. Divakaran · 2018
Earlier work this paper cites.
Lstd: A low-shot transfer detector for object detection
H. Chen, Y. Wang, G. Wang, and Y. Qiao · 2018
Earlier work this paper cites.
Cinic-10 is not imagenet or cifar-10
L. N. Darlow, E. J. Crowley, A. Antoniou, and A. J. Storkey · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Universal language model fine-tuning for text classification
J. Howard and S. Ruder · 2018
Earlier work this paper cites.
Deep contextualized word representations
M. E. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer · 2018
Earlier work this paper cites.
A new method of region embedding for text classification
C. Qiao, B. Huang, G. Niu, D. Li, D. Dong, W. He, D. Yu, and H. Wu · 2018
Earlier work this paper cites.
Know what you don’t know: Unanswerable questions for squad
P. Rajpurkar, R. Jia, and P. Liang · 2018
Earlier work this paper cites.
Transfer learning via learning to transfer
W. Ying, Y. Zhang, J. Huang, and Q. Yang · 2018
Earlier work this paper cites.
Swag: A large-scale adversarial dataset for grounded commonsense inference
R. Zellers, Y. Bisk, R. Schwartz, and Y. Choi · 2018
Cited alongside, same era.
Fully quantizing a simplified transformer for end-to-end speech recognition
A. Bie, B. Venkitesh, J. Monteiro, M. Haidar, M. Rezagholizadeh, et al · 2019
Cited alongside, same era.
Fine-tune bert with sparse self-attention mechanism
B. Cui, Y. Li, M. Chen, and Z. Zhang · 2019
Cited alongside, same era.
Reducing transformer depth on demand with structured dropout
A. Fan, E. Grave, and A. Joulin · 2019
Cited alongside, same era.
Efficient training of bert by progressively stacking
L. Gong, D. He, Z. Li, T. Qin, L. Wang, and T. Liu · 2019
Cited alongside, same era.
Large-batch training for lstm and beyond
Y. You, J. Hseu, C. Ying, J. Demmel, K. Keutzer, and C.-J. Hsieh · 2019
Later among the works it cites.
O. Zafrir, G. Boudoukh, P. Izsak, and M. Wasserblat · 2019
Later among the works it cites.
https://github.com/magnific0/wondershaper
The wonder shaper 1.4.1 · 2020
Later among the works it cites.
Accordion: Adaptive gradient communication via critical learning regime identification
S. Agarwal, H. Wang, K. Lee, S. Venkataraman, and D. Papailiopoulos · 2020
Later among the works it cites.
Don’t stop pretraining: Adapt language models to domains and tasks
S. Gururangan, A. Marasović, S. Swayamdipta, K. Lo, I. Beltagy, D. Downey, and N. A. Smith · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Parameter-efficient transfer learning for nlp
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. De Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly · 2019
Cited alongside, same era.
H. Jiang, P. He, W. Chen, X. Liu, J. Gao, and T. Zhao · 2019
Cited alongside, same era.
Tinybert: Distilling bert for natural language understanding
X. Jiao, Y. Yin, L. Shang, X. Jiang, X. Chen, L. Li, F. Wang, and Q. Liu · 2019
Cited alongside, same era.
Albert: A lite bert for self-supervised learning of language representations
Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, and R. Soricut · 2019
Cited alongside, same era.
What would elsa do? freezing layers during transformer fine-tuning
J. Lee, R. Tang, and J. Lin · 2019
Cited alongside, same era.
Text summarization with pretrained encoders
Y. Liu and M. Lapata · 2019
Cited alongside, same era.
Rethinking the value of network pruning
Z. Liu, M. Sun, T. Zhou, G. Huang, and T. Darrell · 2019
Cited alongside, same era.
A unified architecture for accelerating distributed DNN training in heterogeneous gpu/cpu clusters
Y. Jiang, Y. Zhu, C. Lan, B. Yi, Y. Cui, and C. Guo · 2020
Later among the works it cites.
Pytorch distributed: Experiences on accelerating data parallel training
S. Li, Y. Zhao, R. Varma, O. Salpekar, P. Noordhuis, T. Li, A. Paszke, J. Smith, B. Vaughan, P. Damania, et al · 2020
Later among the works it cites.
Enhancing the reliability of out-of-distribution image detection in neural networks, 2020
S. Liang, Y. Li, and R. Srikant · 2020
Later among the works it cites.
Fastbert: a self-distilling bert with adaptive inference time
W. Liu, P. Zhou, Z. Zhao, Z. Wang, H. Deng, and Q. Ju · 2020
Later among the works it cites.
Do you have the right scissors? tailoring pre-trained language models via monte-carlo methods
N. Miao, Y. Song, H. Zhou, and L. Li · 2020
Later among the works it cites.
Expbert: Representation engineering with natural language explanations
S. Murty, P. W. Koh, and P. Liang · 2020
Later among the works it cites.
Cerebro: A data system for optimized deep learning model selection
S. Nakandala, Y. Zhang, and A. Kumar · 2020
Later among the works it cites.
Mlperf inference benchmark
V. J. Reddi, C. Cheng, D. Kanter, P. Mattson, G. Schmuelling, C.-J. Wu, B. Anderson, M. Breughe, M. Charlebois, W. Chou, et al · 2020
Later among the works it cites.
The right tool for the job: Matching model and instance complexities
R. Schwartz, G. Stanovsky, S. Swayamdipta, J. Dodge, and N. A. Smith · 2020
Later among the works it cites.
Odin: automated drift detection and recovery in video analytics
A. Suprem, J. Arulraj, C. Pu, and J. Ferreira · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, J. Davison, S. Shleifer, P. von Platen, C. Ma, Y. Jernite, J. Plu, C. Xu, T. L. Scao, S. Gugger, M. Drame, Q. Lhoest, and A. M. Rush · 2020
Later among the works it cites.
Deebert: Dynamic early exiting for accelerating bert inference
J. Xin, R. Tang, J. Lee, Y. Yu, and J. Lin · 2020
Later among the works it cites.
Drawing early-bird tickets: Toward more efficient training of deep networks
H. You, C. Li, P. Xu, Y. Fu, Y. Wang, X. Chen, R. G. Baraniuk, Z. Wang, and Y. Lin · 2020
Later among the works it cites.
Accelerating training of transformer-based language models with progressive layer dropping
M. Zhang and Y. He · 2020
Later among the works it cites.
Incorporating bert into neural machine translation
J. Zhu, Y. Xia, L. Wu, D. He, T. Qin, W. Zhou, H. Li, and T.-Y. Liu · 2020
Later among the works it cites.
A comprehensive survey on transfer learning
F. Zhuang, Z. Qi, K. Duan, D. Xi, Y. Zhu, H. Zhu, H. Xiong, and Q. He · 2020
Later among the works it cites.
Deep residual flow for out of distribution detection
E. Zisselman and A. Tamar · 2020
Later among the works it cites.
Autofreeze: Automatically freezing model blocks to accelerate fine-tuning
Y. Liu, S. Agarwal, and S. Venkataraman · 2021
Closest in time.