Improving the Interpretability of Deep Neural Networks with Knowledge Distillation
Xuan Liu, Xiaoguang Wang, and Stan Matwin. 2018a · 2018
Later among the works it cites.
Multi-Label Image Classification via Knowledge Distillation from Weakly-Supervised Detection
Yongcheng Liu, Lu Sheng, Jing Shao, Junjie Yan, Shiming Xiang, and Chunhong Pan. 2018b · 2018
Later among the works it cites.
Wasserstein Auto-Encoders
Ilya Tolstikhin, Sylvain Gelly, Olivier Bousquet, and Bernhard Scholkopf. 2018 · 2018
Later among the works it cites.
Distilled Wasserstein Learning for Word Embedding and Topic Modeling
Hongteng Xu, Wenlin Wang, Wei Liu, and Lawrence Carin. 2018 · 2018
Later among the works it cites.
Decoupling Sparsity and Smoothness in the Dirichlet Variational Autoencoder Topic Model
Sophie Burkhardt and Stefan Kramer. 2019 · 2019
Later among the works it cites.
What does BERT look at? an analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Linguistic knowledge and transferability of contextual representations
Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, and Noah A. Smith. 2019 · 2019
Later among the works it cites.
Topic modeling with Wasserstein autoencoders
Feng Nan, Ran Ding, Ramesh Nallapati, and Bing Xiang. 2019 · 2019
Later among the works it cites.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Original
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019 · 2019
Later among the works it cites.
DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter
Original
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 2019
Later among the works it cites.
Natural Language Generation for Effective Knowledge Distillation
Raphael Tang, Yao Lu, and Jimmy Lin. 2019a · 2019
Later among the works it cites.
ATM:Adversarial-neural Topic Model
Rui Wang, Deyu Zhou, and Yulan He. 2019 · 2019
Later among the works it cites.
Topic Modeling in Embedding Spaces
Adji B Dieng, Francisco JR Ruiz, and David M Blei. 2020 · 2020
Closest in time.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020 · 2020
Closest in time.
Creating Something from Nothing: Unsupervised Knowledge Distillation for Cross-Modal Hashing
Hengtong Hu, Lingxi Xie, Richang Hong, and Qi Tian. 2020 · 2020
Closest in time.
Neural Topic Modeling with Bidirectional Adversarial Training
Rui Wang, Xuemeng Hu, Deyu Zhou, Yulan He, Yuxuan Xiong, Chenchen Ye, and Haiyang Xu. 2020 · 2020
Closest in time.
TextBrewer: An Open-Source Knowledge Distillation Toolkit for Natural Language Processing
Original
Ziqing Yang, Yiming Cui, Zhipeng Chen, Wanxiang Che, Ting Liu, Shijin Wang, and Guoping Hu. 2020 · 2020
Closest in time.
Neural models for documents with metadata
Dallas Card, Chenhao Tan, and Noah A. Smith. 2018 · 2040
Closest in time.