Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Original
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, et al · 1909
Earlier work this paper cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Original
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
Q8BERT: Quantized 8Bit BERT
Original
Ofir Zafrir, Guy Boudoukh, Peter Izsak, and Moshe Wasserblat. 2019 · 1910
Earlier work this paper cites.
TwinBERT: Distilling Knowledge to Twin-Structured BERT Models for Efficient Retrieval
Original
Wenhao Lu, Jian Jiao, and Ruofei Zhang. 2020 · 2002
Earlier work this paper cites.
Deep Neural Networks for YouTube Recommendations. In RecSys 2016
Paul Covington, Jay Adams, and Emre Sargin. 2016 · 2016
Earlier work this paper cites.
BranchyNet: Fast Inference via Early Exiting from Deep Neural Networks. In ICPR 2016
Surat Teerapittayanon, Bradley McDanel, and H. T. Kung. 2016 · 2016
Earlier work this paper cites.
Learning to Generate Product Reviews from Attributes. In EACL 2017
Li Dong, Shaohan Huang, Furu Wei, et al · 2017
Earlier work this paper cites.
Visually-aware fashion recommendation and design with generative image models. In ICDM 2017
Wang-Cheng Kang, Chen Fang, Zhaowen Wang, and Julian McAuley. 2017 · 2017
Earlier work this paper cites.
Neural Rating Regression with Abstractive Tips Generation for Recommendation. In SIGIR 2017
Piji Li, Zihao Wang, Zhaochun Ren, Lidong Bing, and Wai Lam. 2017 · 2017
Earlier work this paper cites.
Attention is All you Need. In NeurIPS 2017
Ashish Vaswani, Noam Shazeer, Niki Parmar, et al · 2017
Earlier work this paper cites.
To prune, or not to prune: exploring the efficacy of pruning for model compression
Original
Michael Zhu and Suyog Gupta. 2017 · 2017
Earlier work this paper cites.
Multi-Scale Dense Networks for Resource Efficient Image Classification. In ICLR 2018
Gao Huang, Danlu Chen, Tianhong Li, et al · 2018
Earlier work this paper cites.
Deep Interest Network for Click-Through Rate Prediction. In KDD 2018
Guorui Zhou, Chengru Song, Xiaoqiang Zhu, et al · 2018
Earlier work this paper cites.
Towards Knowledge-Based Recommender Dialog System. In EMNLP 2019
Qibin Chen, Junyang Lin, Yichang Zhang, et al · 2019
Earlier work this paper cites.
Towards knowledge-based personalized product description generation in e-commerce. In KDD 2019
Qibin Chen, Junyang Lin, Yichang Zhang, et al · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL 2019
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Unified Language Model Pre-training for Natural Language Understanding and Generation. In NeurIPS 2019
Li Dong, Nan Yang, Wenhui Wang, et al · 2019
Earlier work this paper cites.
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks. In NeurIPS 2019
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019 · 2019
Earlier work this paper cites.
Justifying Recommendations using Distantly-Labeled Reviews and Fine-Grained Aspects. In EMNLP 2019
Jianmo Ni, Jiacheng Li, and Julian McAuley. 2019a · 2019
Earlier work this paper cites.
Justifying Recommendations using Distantly-Labeled Reviews and Fine-Grained Aspects. In EMNLP 2019
Jianmo Ni, Jiacheng Li, and Julian McAuley. 2019b · 2019
Earlier work this paper cites.
Language Models are Unsupervised Multitask Learners
Alec Radford, Jeff Wu, Rewon Child, et al · 2019
Earlier work this paper cites.
Language Models are Few-Shot Learners. In NeurIPS 2020
Tom B. Brown, Benjamin Mann, Nick Ryder, et al · 2020
Earlier work this paper cites.