Fetching the paper…
Reading the bibliography…
Adapting pre-trained language models (PrLMs) (e.g., BERT) to new domains has gained much attention recently.
Roberta: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019c · 1907
Earlier work this paper cites.
Alexander Rietzler, Sebastian Stabinger, Paul Opitz, and Stefan Engl. 2019 · 1908
Earlier work this paper cites.
Unsupervised domain adaptation on reading comprehension
Yu Cao, Meng Fang, Baosheng Yu, and Joey Tianyi Zhou. 2019 · 1911
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1911
Earlier work this paper cites.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020 · 2004
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 2005
Earlier work this paper cites.
Semi-supervised classification by low density separation
Olivier Chapelle and Alexander Zien. 2005 · 2005
Earlier work this paper cites.
Biographies, Bollywood, boom-boxes and blenders: Domain adaptation for sentiment classification
John Blitzer, Mark Dredze, and Fernando Pereira. 2007 · 2007
Earlier work this paper cites.
Dataset Shift in Machine Learning
Joaquin Quionero-Candela, Masashi Sugiyama, Anton Schwaighofer, and Neil D Lawrence. 2009 · 2009
Earlier work this paper cites.
A theory of learning from different domains
Shai Ben-David, John Blitzer, Koby Crammer, Alex Kulesza, Fernando Pereira, and Jennifer Wortman Vaughan. 2010 · 2010
Earlier work this paper cites.
Cross-domain sentiment classification via spectral feature alignment
Sinno Jialin Pan, Xiaochuan Ni, Jian-Tao Sun, Qiang Yang, and Zheng Chen. 2010 · 2010
Earlier work this paper cites.
Cross-language text classification using structural correspondence learning
Peter Prettenhofer and Benno Stein. 2010 · 2010
Earlier work this paper cites.
Automatically extracting polarity-bearing topics for cross-domain sentiment classification
Yulan He, Chenghua Lin, and Harith Alani. 2011 · 2011
Earlier work this paper cites.
A kernel two-sample test
Arthur Gretton, Karsten M. Borgwardt, Malte J. Rasch, Bernhard Schölkopf, and Alexander J. Smola. 2012 · 2012
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Learning transferable features with deep adaptation networks
Mingsheng Long, Yue Cao, Jianmin Wang, and Michael I. Jordan. 2015 · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Thang Luong, Hieu Pham, and Christopher D. Manning. 2015 · 2015
Cited alongside, same era.
Supervised representation learning: Transfer learning with deep autoencoders
Fuzhen Zhuang, Xiaohu Cheng, Ping Luo, Sinno Jialin Pan, and Qing He. 2015 · 2015
Cited alongside, same era.
Domain-adversarial training of neural networks
Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, François Laviolette, Mario Marchand, and Victor S. Lempitsky. 2016 · 2016
Cited alongside, same era.
Deep reconstruction-classification networks for unsupervised domain adaptation
Muhammad Ghifary, W. Bastiaan Kleijn, Mengjie Zhang, David Balduzzi, and Wen Li. 2016 · 2016
Cited alongside, same era.
A closer look at memorization in deep networks
Devansh Arpit, Stanislaw Jastrzebski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S. Kanwal, Tegan Maharaj, Asja Fischer, Aaron C. Courville, Yoshua Bengio, and Simon Lacoste-Julien. 2017 · 2017
Cited alongside, same era.
Publicly available clinical BERT embeddings
Emily Alsentzer, John R. Murphy, Willie Boag, Wei-Hung Weng, Di Jin, Tristan Naumann, and Matthew B. A. McDermott. 2019 · 2019
Later among the works it cites.
Mixmatch: A holistic approach to semi-supervised learning
David Berthelot, Nicholas Carlini, Ian J. Goodfellow, Nicolas Papernot, Avital Oliver, and Colin Raffel. 2019 · 2019
Later among the works it cites.
Multi-source cross-lingual model transfer: Learning what to share
Xilun Chen, Ahmed Hassan Awadallah, Hany Hassan, Wei Wang, and Claire Cardie. 2019 · 2019
Later among the works it cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Unsupervised domain adaptation of contextualized embeddings for sequence labeling
Xiaochuang Han and Jacob Eisenstein. 2019 · 2019
Later among the works it cites.
Visualizing and understanding the effectiveness of BERT
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Temporal ensembling for semi-supervised learning
Samuli Laine and Timo Aila. 2017 · 2017
Cited alongside, same era.
Asymmetric tri-training for unsupervised domain adaptation
Kuniaki Saito, Yoshitaka Ushiku, and Tatsuya Harada. 2017 · 2017
Cited alongside, same era.
Cross-lingual distillation for text classification
Ruochen Xu and Yiming Yang. 2017 · 2017
Cited alongside, same era.
Jointly extracting relations with class ties via effective deep ranking
Hai Ye, Wenhan Chao, Zhunchen Luo, and Zhoujun Li. 2017 · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals. 2017 · 2017
Cited alongside, same era.
Adversarial deep averaging networks for cross-lingual sentiment classification
Xilun Chen, Yu Sun, Ben Athiwaratkun, Claire Cardie, and Kilian Weinberger. 2018 · 2018
Cited alongside, same era.
Adaptive semi-supervised learning for cross-domain sentiment classification
Ruidan He, Wee Sun Lee, Hwee Tou Ng, and Daniel Dahlmeier. 2018 · 2018
Cited alongside, same era.
Yaru Hao, Li Dong, Furu Wei, and Ke Xu. 2019 · 2019
Later among the works it cites.
Learning deep representations by mutual information estimation and maximization
R. Devon Hjelm, Alex Fedorov, Samuel Lavoie-Marchildon, Karan Grewal, Philip Bachman, Adam Trischler, and Yoshua Bengio. 2019 · 2019
Later among the works it cites.
Drop to adapt: Learning discriminative features for unsupervised domain adaptation
Seungmin Lee, Dongwan Kim, Namil Kim, and Seong-Gyun Jeong. 2019 · 2019
Later among the works it cites.
Zero-shot entity linking by reading entity descriptions
Lajanugen Logeswaran, Ming-Wei Chang, Kenton Lee, Kristina Toutanova, Jacob Devlin, and Honglak Lee. 2019 · 2019
Later among the works it cites.
Domain adaptation with BERT-based domain classification and data selection
Xiaofei Ma, Peng Xu, Zhiguo Wang, Ramesh Nallapati, and Bing Xiang. 2019 · 2019
Later among the works it cites.
To tune or not to tune? Adapting pretrained representations to diverse tasks
Matthew E Peters, Sebastian Ruder, and Noah A Smith. 2019 · 2019
Later among the works it cites.
XLNet: generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019 · 2019
Later among the works it cites.
Jointly learning semantic parser and natural language generator via dual information maximization
Hai Ye, Wenjie Li, and Lu Wang. 2019 · 2019
Later among the works it cites.
Mutual mean-teaching: Pseudo label refinery for unsupervised domain adaptation on person re-identification
Yixiao Ge, Dapeng Chen, and Hongsheng Li. 2020 · 2020
Closest in time.
BioBERT: a pre-trained biomedical language representation model for biomedical text mining
Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. 2020 · 2020
Closest in time.
Unsupervised domain adaptation of a pretrained cross-lingual language model
Juntao Li, Ruidan He, Hai Ye, Hwee Tou Ng, Lidong Bing, and Rui Yan. 2020 · 2020
Closest in time.
Contrastive representation distillation
Yonglong Tian, Dilip Krishnan, and Phillip Isola. 2020 · 2020
Closest in time.