Fetching the paper…
Reading the bibliography…
While pre-trained language models (e.g., BERT) have achieved impressive results on different natural language processing tasks, they have large numbers of parameters and suffer from big computational and memory costs, which make them difficult for real-world deployment.
Balanced One-shot Neural Architecture Optimization
Renqian Luo, Tao Qin, and Enhong Chen. 2019 · 1909
Earlier work this paper cites.
Auto-keras: An efficient neural architecture search system. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining . 1946–1956
Haifeng Jin, Qingquan Song, and Xia Hu. 2019 · 1956
Earlier work this paper cites.
Block-wisely Supervised Neural Architecture Search with Knowledge Distillation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 1989–1998
Changlin Li, Jiefeng Peng, Liuchun Yuan, Guangrun Wang, Xiaodan Liang, Liang Lin, and Xiaojun Chang. 2020b · 1998
Earlier work this paper cites.
Automatically Constructing a Corpus of Sentential Paraphrases. In Proceedings of the Third International Workshop on Paraphrasing (IWP2005)
William B. Dolan and Chris Brockett. 2005 · 2005
Earlier work this paper cites.
The PASCAL Recognising Textual Entailment Challenge. In Machine Learning Challenges. Evaluating Predictive Uncertainty, Visual Object Classification, and Recognising Tectual Entailment . 177–190
Ido Dagan, Oren Glickman, and Bernardo Magnini. 2006 · 2006
Earlier work this paper cites.
Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank. In EMNLP . 1631–1642
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization. In ICLR (Poster)
Diederik P Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
SQuAD: 100,000+ Questions for Machine Comprehension of Text. In EMNLP . 2383–2392
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
SemEval-2017 Task 1: Semantic Textual Similarity Multilingual and Crosslingual Focused Evaluation. In Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017) . 1–14
Daniel Cer, Mona Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia. 2017 · 2017
Earlier work this paper cites.
Attention is all you need. In Advances in neural information processing systems . 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Understanding and simplifying one-shot architecture search. In International Conference on Machine Learning . 550–559
Gabriel Bender, Pieter-Jan Kindermans, Barret Zoph, Vijay Vasudevan, and Quoc Le. 2018 · 2018
Earlier work this paper cites.
ProxylessNAS: Direct Neural Architecture Search on Target Task and Hardware. In International Conference on Learning Representations
Han Cai, Ligeng Zhu, and Song Han. 2018 · 2018
Earlier work this paper cites.
Quora question pairs
Zihan Chen, Hongbo Zhang, Xiaoji Zhang, and Leqi Zhao. 2018 · 2018
Earlier work this paper cites.
Neural architecture search: A survey
Thomas Elsken, Jan Hendrik Metzen, and Frank Hutter. 2018 · 2018
Earlier work this paper cites.
Depthwise Separable Convolutions for Neural Machine Translation. In International Conference on Learning Representations
Lukasz Kaiser, Aidan N Gomez, and Francois Chollet. 2018 · 2018
Earlier work this paper cites.
DARTS: Differentiable Architecture Search. In International Conference on Learning Representations
Hanxiao Liu, Karen Simonyan, and Yiming Yang. 2018 · 2018
Earlier work this paper cites.
Know What You Don’t Know: Unanswerable Questions for SQuAD. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) . 784–789
Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018 · 2018
Earlier work this paper cites.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. In International Conference on Learning Representations
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2018 · 2018
Earlier work this paper cites.
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference. In NAACL . 1112–1122
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018 · 2018
Cited alongside, same era.
Pay Less Attention with Lightweight and Dynamic Convolutions. In International Conference on Learning Representations
Felix Wu, Angela Fan, Alexei Baevski, Yann Dauphin, and Michael Auli. 2018 · 2018
Cited alongside, same era.
Once-for-All: Train One Network and Specialize it for Efficient Deployment. In International Conference on Learning Representations
Han Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang, and Song Han. 2019 · 2019
Cited alongside, same era.
Fairnas: Rethinking evaluation fairness of weight sharing neural architecture search
Xiangxiang Chu, Bo Zhang, Ruijun Xu, and Jixiang Li. 2019 · 2019
Cited alongside, same era.
ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators. In International Conference on Learning Representations
Ofir Zafrir, Guy Boudoukh, Peter Izsak, and Moshe Wasserblat. 2019 · 2019
Later among the works it cites.
AdaBERT: Task-Adaptive BERT Compression with Differentiable Neural Architecture Search. In Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence, IJCAI-20 , Christian Bessiere (Ed.). International Joint Conferences on Artificial Intelligence Organization, 2463–2469
Daoyuan Chen, Yaliang Li, Minghui Qiu, Zhen Wang, Bofang Li, Bolin Ding, Hongbo Deng, Jun Huang, Wei Lin, and Jingren Zhou. 2020 · 2020
Later among the works it cites.
Compressing BERT: Studying the Effects of Weight Pruning on Transfer Learning. In Proceedings of the 5th Workshop on Representation Learning for NLP . 143–155
Mitchell Gordon, Kevin Duh, and Nicholas Andrews. 2020 · 2020
Later among the works it cites.
Single path one-shot neural architecture search with uniform sampling. In European Conference on Computer Vision . Springer, 544–560
Zichao Guo, Xiangyu Zhang, Haoyuan Mu, Wen Heng, Zechun Liu, Yichen Wei, and Jian Sun. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kevin Clark, Minh-Thang Luong, Quoc V Le, and Christopher D Manning. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) . Association for Computational Linguistics, Minneapolis, Minnesota, 4171–4186
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Reducing Transformer Depth on Demand with Structured Dropout. In International Conference on Learning Representations
Angela Fan, Edouard Grave, and Armand Joulin. 2019 · 2019
Cited alongside, same era.
Searching for mobilenetv3. In Proceedings of the IEEE International Conference on Computer Vision . 1314–1324
Andrew Howard, Mark Sandler, Grace Chu, Liang-Chieh Chen, Bo Chen, Mingxing Tan, Weijun Wang, Yukun Zhu, Ruoming Pang, Vijay Vasudevan, et al · 2019
Cited alongside, same era.
Large memory layers with product keys. In Advances in Neural Information Processing Systems . 8548–8559
Guillaume Lample, Alexandre Sablayrolles, Marc’Aurelio Ranzato, Ludovic Denoyer, and Hervé Jégou. 2019 · 2019
Cited alongside, same era.
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations. In International Conference on Learning Representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019 · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Cited alongside, same era.
Pruning a bert-based question answering model
J Scott McCarley. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
DynaBERT: Dynamic BERT with Adaptive Width and Depth
Lu Hou, Zhiqi Huang, Lifeng Shang, Xin Jiang, Xiao Chen, and Qun Liu. 2020 · 2020
Later among the works it cites.
TinyBERT: Distilling BERT for Natural Language Understanding. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: Findings . 4163–4174
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2020 · 2020
Later among the works it cites.
Applying depthwise separable and multi-channel convolutional neural networks of varied kernel size on semantic trajectories
Antonios Karatzoglou, Nikolai Schnell, and Michael Beigl. 2020 · 2020
Later among the works it cites.
FreeDOM: A Transferable Neural Architecture for Structured Information Extraction on Web Documents. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining . 1092–1102
Bill Yuchen Lin, Ying Sheng, Nguyen Vo, and Sandeep Tata. 2020 · 2020
Later among the works it cites.
Neural architecture search with gbdt
Renqian Luo, Xu Tan, Rui Wang, Tao Qin, Enhong Chen, and Tie-Yan Liu. 2020 · 2020
Later among the works it cites.
Designing network design spaces. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 10428–10436
Ilija Radosavovic, Raj Prateek Kosaraju, Ross Girshick, Kaiming He, and Piotr Dollár. 2020 · 2020
Later among the works it cites.
Q-BERT: Hessian Based Ultra Low Precision Quantization of BERT.. In AAAI . 8815–8821
Sheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma, Zhewei Yao, Amir Gholami, Michael W Mahoney, and Kurt Keutzer. 2020 · 2020
Later among the works it cites.
LightPAFF: A Two-Stage Distillation Framework for Pre-training and Fine-tuning
Kaitao Song, Hao Sun, Xu Tan, Tao Qin, Jianfeng Lu, Hongzhi Liu, and Tie-Yan Liu. 2020 · 2020
Later among the works it cites.
MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . 2158–2170
Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou. 2020 · 2020
Later among the works it cites.
Finding Fast Transformers: One-Shot Neural Architecture Search by Component Composition
Henry Tsai, Jayden Ooi, Chun-Sung Ferng, Hyung Won Chung, and Jason Riesa. 2020 · 2020
Later among the works it cites.
MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual , Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin (Eds.)
Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou. 2020a · 2020
Later among the works it cites.
BERT-of-Theseus: Compressing BERT by Progressive Module Replacing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) . 7859–7869
Canwen Xu, Wangchunshu Zhou, Tao Ge, Furu Wei, and Ming Zhou. 2020 · 2020
Later among the works it cites.
Bignas: Scaling up neural architecture search with big single-stage models. In European Conference on Computer Vision . Springer, 702–717
Jiahui Yu, Pengchong Jin, Hanxiao Liu, Gabriel Bender, Pieter-Jan Kindermans, Mingxing Tan, Thomas Huang, Xiaodan Song, Ruoming Pang, and Quoc Le. 2020 · 2020
Later among the works it cites.