Fetching the paper…
Reading the bibliography…
Pre-trained language models have been applied to various NLP tasks with considerable performance gains.
Biographies, bollywood, boom-boxes and blenders: Domain adaptation for sentiment classification
John Blitzer, Mark Dredze, and Fernando Pereira. 2007 · 2007
Earlier work this paper cites.
A survey on transfer learning
Sinno Jialin Pan and Qiang Yang. 2010 · 2010
Earlier work this paper cites.
Compressing deep convolutional networks using vector quantization
Yunchao Gong, Liu Liu, Ming Yang, and Lubomir D. Bourdev. 2014 · 2014
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Song Han, Jeff Pool, John Tran, and William J. Dally. 2015 · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeffrey Dean. 2015 · 2015
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
How transferable are neural networks in NLP applications?
Lili Mou, Zhao Meng, Rui Yan, Ge Li, Yan Xu, Lu Zhang, and Zhi Jin. 2016 · 2016
Earlier work this paper cites.
Meta-learning with memory-augmented neural networks
Adam Santoro, Sergey Bartunov, Matthew Botvinick, Daan Wierstra, and Timothy P. Lillicrap. 2016 · 2016
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine. 2017 · 2017
Earlier work this paper cites.
Adversarial multi-task learning for text classification
Pengfei Liu, Xipeng Qiu, and Xuanjing Huang. 2017 · 2017
Earlier work this paper cites.
Data-free knowledge distillationfor deep neural networks
Raphael Gontijo Lopes, Stefano Fenu, and Thad Starner. 2017 · 2017
Earlier work this paper cites.
Knowledge adaptation: Teaching to adapt
Sebastian Ruder, Parsa Ghaffari, and John G. Breslin. 2017 · 2017
Earlier work this paper cites.
Prototypical networks for few-shot learning
Jake Snell, Kevin Swersky, and Richard S. Zemel. 2017 · 2017
Earlier work this paper cites.
Distant domain transfer learning
Ben Tan, Yu Zhang, Sinno Jialin Pan, and Qiang Yang. 2017 · 2017
Earlier work this paper cites.
Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results
Antti Tarvainen and Harri Valpola. 2017 · 2017
Earlier work this paper cites.
Soft weight-sharing for neural network compression
Karen Ullrich, Edward Meeds, and Max Welling. 2017 · 2017
Earlier work this paper cites.
Transfer learning for sequence tagging with hierarchical recurrent networks
Zhilin Yang, Ruslan Salakhutdinov, and William W. Cohen. 2017 · 2017
Earlier work this paper cites.
A gift from knowledge distillation: Fast optimization, network minimization and transfer learning
Junho Yim, Donggyu Joo, Ji-Hoon Bae, and Junmo Kim. 2017 · 2017
Earlier work this paper cites.
Learning from multiple teacher networks
Shan You, Chang Xu, Chao Xu, and Dacheng Tao. 2017 · 2017
Cited alongside, same era.
Cross-domain review helpfulness prediction based on convolutional neural networks with auxiliary domain discriminators
Cen Chen, Yinfei Yang, Jun Zhou, Xiaolong Li, and Forrest Sheng Bao. 2018 · 2018
Cited alongside, same era.
Probabilistic model-agnostic meta-learning
Chelsea Finn, Kelvin Xu, and Sergey Levine. 2018 · 2018
Cited alongside, same era.
Born-again neural networks
Tommaso Furlanello, Zachary Chase Lipton, Michael Tschannen, Laurent Itti, and Anima Anandkumar. 2018 · 2018
Cited alongside, same era.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Meta-learning: A survey
Joaquin Vanschoren. 2018 · 2018
Cited alongside, same era.
Well-read students learn better: On the importance of pre-training compact models
Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Multimodal model-agnostic meta-learning via task-aware modulation
Risto Vuorio, Shao-Hua Sun, Hexiang Hu, and Joseph J. Lim. 2019 · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019 · 2019
Later among the works it cites.
Extreme language model compression with optimal subwords and shared projections
Sanqiang Zhao, Raghav Gupta, Yang Song, and Denny Zhou. 2019 · 2019
Later among the works it cites.
Few shot network compression via cross distillation
Haoli Bai, Jiaxiang Wu, Irwin King, and Michael R. Lyu. 2020 · 2020
Closest in time.
Meta-learning deep energy-based memory models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018 · 2018
Cited alongside, same era.
Multi-domain gated CNN for review helpfulness prediction
Cen Chen, Minghui Qiu, Yinfei Yang, Jun Zhou, Jun Huang, Xiaolong Li, and Forrest Sheng Bao. 2019 · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Domain-invariant feature distillation for cross-domain sentiment classification
Mengting Hu, Yike Wu, Shiwan Zhao, Honglei Guo, Renhong Cheng, and Zhong Su. 2019 · 2019
Cited alongside, same era.
Metadistiller: Network self-boosting via meta-learned top-down distillation
Yunhun Jang, Hankook Lee, Sung Ju Hwang, and Jinwoo Shin. 2019 · 2019
Cited alongside, same era.
Meta-learning representations for continual learning
Khurram Javed and Martha White. 2019 · 2019
Cited alongside, same era.
Sergey Bartunov, Jack W. Rae, Simon Osindero, and Timothy P. Lillicrap. 2020 · 2020
Closest in time.
Squeezebert: What can computer vision teach NLP about efficient neural networks?
Forrest N. Iandola, Albert E. Shaw, Ravi Krishna, and Kurt Keutzer. 2020 · 2020
Closest in time.
Few sample knowledge distillation for efficient network compression
Tianhong Li, Jianguo Li, Zhuang Liu, and Changshui Zhang. 2020 · 2020
Closest in time.
Metadistiller: Network self-boosting via meta-learned top-down distillation
Benlin Liu, Yongming Rao, Jiwen Lu, Jie Zhou, and Cho jui Hsieh. 2020 · 2020
Closest in time.
MTSS: learn from multiple domain teachers and become a multi-domain dialogue expert
Shuke Peng, Feng Ji, Zehao Lin, Shaobo Cui, Haiqing Chen, and Yin Zhang. 2020 · 2020
Closest in time.
Mobilebert: a compact task-agnostic BERT for resource-limited devices
Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou. 2020 · 2020
Closest in time.
Meta fine-tuning neural language models for multi-domain text mining
Chengyu Wang, Minghui Qiu, Jun Huang, and Xiaofeng He. 2020 · 2020
Closest in time.
Model compression with two-stage multi-teacher knowledge distillation for web question answering system
Ze Yang, Linjun Shou, Ming Gong, Wutao Lin, and Daxin Jiang. 2020 · 2020
Closest in time.
Zero-shot text classification via reinforced self-training
Zhiquan Ye, Yuxia Geng, Jiaoyan Chen, Jingmin Chen, Xiaoxiao Xu, Suhang Zheng, Feng Wang, Jun Zhang, and Huajun Chen. 2020 · 2020
Closest in time.
Meta-learning for few-shot natural language processing: A survey
Wenpeng Yin. 2020 · 2020
Closest in time.
Cross-domain knowledge distillation for retrieval-based question answering systems
Cen Chen, Chengyu Wang, Minghui Qiu, Dehong Gao, Linbo Jin, and Wang Li. 2021 · 2021
Closest in time.
Learning to augment for data-scarce domain bert knowledge distillation
Lingyun Feng, Minghui Qiu, Yaliang Li, Hai-Tao Zheng, and Ying Shen. 2021 · 2021
Closest in time.