Fetching the paper…
Reading the bibliography…
Transformer based Very Large Language Models (VLLMs) like BERT, XLNet and RoBERTa, have recently shown tremendous performance on a large variety of Natural Language Understanding (NLU) tasks.
Optimal brain damage. In Advances in neural information processing systems . 598–605
Yann LeCun, John S Denker, and Sara A Solla. 1990 · 1990
Earlier work this paper cites.
A new algorithm for data compression
Philip Gage. 1994 · 1994
Earlier work this paper cites.
Model compression. In Proceedings of the 12th ACM SIGKDD international conference on Knowledge discovery and data mining . ACM, 535–541
Cristian Buciluǎ, Rich Caruana, and Alexandru Niculescu-Mizil. 2006 · 2006
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Alex Graves, Greg Wayne, and Ivo Danihelka. 2014 · 2014
Earlier work this paper cites.
Song Han, Huizi Mao, and William J Dally. 2015a · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015 · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy. 2015 · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Minh-Thang Luong, Hieu Pham, and Christopher D Manning. 2015 · 2015
Earlier work this paper cites.
Long short-term memory-networks for machine reading
Jianpeng Cheng, Li Dong, and Mirella Lapata. 2016 · 2016
Earlier work this paper cites.
Fixed point quantization of deep convolutional networks. In International Conference on Machine Learning . 2849–2858
Darryl Lin, Sachin Talathi, and Sreekanth Annapureddy. 2016 · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Massive exploration of neural machine translation architectures
Denny Britz, Anna Goldie, Minh-Thang Luong, and Quoc Le. 2017 · 2017
Earlier work this paper cites.
A survey of model compression and acceleration for deep neural networks
Yu Cheng, Duo Wang, Pan Zhou, and Tao Zhang. 2017 · 2017
Earlier work this paper cites.
Attention is all you need. In Advances in neural information processing systems . 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin. 2018 · 2018
Earlier work this paper cites.
Retraining-based iterative weight quantization for deep neural networks
Dongsoo Lee and Byeongwook Kim. 2018 · 2018
Earlier work this paper cites.
Rethinking the value of network pruning
Zhuang Liu, Mingjie Sun, Tinghui Zhou, Gao Huang, and Trevor Darrell. 2018 · 2018
Earlier work this paper cites.
The natural language decathlon: Multitask learning as question answering
Bryan McCann, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher. 2018 · 2018
Cited alongside, same era.
Deep contextualized word representations
Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Know What You Don’t Know: Unanswerable Questions for SQuAD
Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018 · 2018
Cited alongside, same era.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2018 · 2018
Cited alongside, same era.
Fully Quantized Transformer for Improved Translation
Gabriele Prato, Ella Charlaix, and Mehdi Rezagholizadeh. 2019 · 2019
Closest in time.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Closest in time.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2019 · 2019
Closest in time.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 2019
Closest in time.
Megatron-lm: Training multi-billion parameter language models using gpu model parallelism
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ekaterina Arkhangelskaia and Sourav Dutta. 2019 · 2019
Cited alongside, same era.
transformers. zip: Compressing Transformers with Pruning and Quantization
Robin Cheong and Robel Daniel. 2019 · 2019
Cited alongside, same era.
What Does BERT Look At? An Analysis of BERT’s Attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D Manning. 2019 · 2019
Cited alongside, same era.
Cross-lingual Language Model Pretraining. In Advances in Neural Information Processing Systems . 7057–7067
Alexis Conneau and Guillaume Lample. 2019 · 2019
Cited alongside, same era.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, William W Cohen, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. 2019 · 2019
Cited alongside, same era.
A structural probe for finding syntax in word representations. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) . 4129–4138
John Hewitt and Christopher D Manning. 2019 · 2019
Cited alongside, same era.
TinyBERT: Distilling BERT for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2019 · 2019
Cited alongside, same era.
Spanbert: Improving pre-training by representing and predicting spans
Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S Weld, Luke Zettlemoyer, and Omer Levy. 2019 · 2019
Cited alongside, same era.
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019 · 2019
Closest in time.
What does BERT Learn from Multiple-Choice Reading Comprehension Datasets?
Chenglei Si, Shuohang Wang, Min-Yen Kan, and Jing Jiang. 2019 · 2019
Closest in time.
Patient knowledge distillation for bert model compression
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu. 2019a · 2019
Closest in time.
Ernie 2.0: A continual pre-training framework for language understanding
Yu Sun, Shuohuan Wang, Yukun Li, Shikun Feng, Hao Tian, Hua Wu, and Haifeng Wang. 2019b · 2019
Closest in time.
Distilling Task-Specific Knowledge from BERT into Simple Neural Networks
Raphael Tang, Yao Lu, Linqing Liu, Lili Mou, Olga Vechtomova, and Jimmy Lin. 2019 · 2019
Closest in time.
Bert rediscovers the classical nlp pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Closest in time.
Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 2019
Closest in time.
Universal adversarial triggers for nlp
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh. 2019a · 2019
Closest in time.
AllenNLP Interpret: A Framework for Explaining Predictions of NLP Models
Eric Wallace, Jens Tuyls, Junlin Wang, Sanjay Subramanian, Matt Gardner, and Sameer Singh. 2019b · 2019
Closest in time.
Do NLP Models Know Numbers? Probing Numeracy in Embeddings
Eric Wallace, Yizhong Wang, Sujian Li, Sameer Singh, and Matt Gardner. 2019c · 2019
Closest in time.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2019 · 2019
Closest in time.
XLNet: Generalized Autoregressive Pretraining for Language Understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V Le. 2019 · 2019
Closest in time.
Improving Neural Network Quantization without Retraining using Outlier Channel Splitting. In International Conference on Machine Learning . 7543–7552
Ritchie Zhao, Yuwei Hu, Jordan Dotzel, Chris De Sa, and Zhiru Zhang. 2019 · 2019
Closest in time.
Deconstructing lottery tickets: Zeros, signs, and the supermask
Hattie Zhou, Janice Lan, Rosanne Liu, and Jason Yosinski. 2019 · 2019
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention. In International conference on machine learning . 2048–2057
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. 2015 · 2057
Closest in time.