Fetching the paper…
Reading the bibliography…
Pretrained language models (PTLMs) are typically learned over a large, static corpus and further fine-tuned for various downstream tasks.
On tiny episodic memories in continual learning
Arslan Chaudhry, Marcus Rohrbach, Mohamed Elhoseiny, Thalaiyasingam Ajanthan, Puneet K Dokania, Philip HS Torr, and Marc’Aurelio Ranzato. 2019 · 1902
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019b · 1907
Earlier work this paper cites.
Catastrophic forgetting, rehearsal and pseudorehearsal
Anthony V. Robins. 1995 · 1995
Earlier work this paper cites.
Covid-twitter-bert: A natural language processing model to analyse covid-19 content on twitter
Martin Müller, Marcel Salathé, and Per Egil Kummervold. 2020 · 2005
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey E. Hinton, Oriol Vinyals, and J. Dean. 2015 · 2015
Earlier work this paper cites.
Recurrent neural network language model adaptation with curriculum learning
Yangyang Shi, Martha Larson, and Catholijn M. Jonker. 2015 · 2015
Earlier work this paper cites.
Hashtag recommendation using attention-based convolutional neural network
Yuyun Gong and Qi Zhang. 2016 · 2016
Earlier work this paper cites.
Chemprot-3.0: a global chemical biology diseases mapping
Jens Vindahl. 2016 · 2016
Earlier work this paper cites.
SemEval 2017 task 10: ScienceIE - extracting keyphrases and relations from scientific publications
Isabelle Augenstein, Mrinal Das, Sebastian Riedel, Lakshmi Vikraman, and Andrew McCallum. 2017 · 2017
Earlier work this paper cites.
PubMed 200k RCT: a dataset for sequential sentence classification in medical abstracts
Franck Dernoncourt and Ji Young Lee. 2017 · 2017
Earlier work this paper cites.
icarl: Incremental classifier and representation learning
Sylvestre-Alvise Rebuffi, Alexander Kolesnikov, Georg Sperl, and Christoph H. Lampert. 2017 · 2017
Earlier work this paper cites.
SemEval 2018 task 2: Multilingual emoji prediction
Francesco Barbieri, Jose Camacho-Collados, Francesco Ronzano, Luis Espinosa-Anke, Miguel Ballesteros, Valerio Basile, Viviana Patti, and Horacio Saggion. 2018 · 2018
Earlier work this paper cites.
Lifelong learning via progressive distillation and retrospection
Saihui Hou, Xinyu Pan, Chen Change Loy, Zilei Wang, and Dahua Lin. 2018 · 2018
Earlier work this paper cites.
Examining temporality in document classification
Xiaolei Huang and Michael J. Paul. 2018 · 2018
Earlier work this paper cites.
Measuring the evolution of a scientific field through citation frames
David Jurgens, Srijan Kumar, Raine Hoover, Dan McFarland, and Dan Jurafsky. 2018 · 2018
Earlier work this paper cites.
Learning without forgetting
Zhizhong Li and Derek Hoiem. 2018 · 2018
Earlier work this paper cites.
Multi-task identification of entities, relations, and coreference for scientific knowledge graph construction
Yi Luan, Luheng He, Mari Ostendorf, and Hannaneh Hajishirzi. 2018 · 2018
Earlier work this paper cites.
Progress & compress: A scalable framework for continual learning
Jonathan Schwarz, Wojciech Czarnecki, Jelena Luketina, Agnieszka Grabska-Barwinska, Yee Whye Teh, Razvan Pascanu, and Raia Hadsell. 2018 · 2018
Earlier work this paper cites.
SciBERT: A pretrained language model for scientific text
Iz Beltagy, Kyle Lo, and Arman Cohan. 2019 · 2019
Earlier work this paper cites.
Episodic memory in lifelong language learning
Cyprien de Masson d’Autume, Sebastian Ruder, Lingpeng Kong, and Dani Yogatama. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Parameter-efficient transfer learning for NLP
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019 · 2019
Cited alongside, same era.
Continual learning for sentence representations using conceptors
Tianlin Liu, Lyle Ungar, and João Sedoc. 2019a · 2019
Cited alongside, same era.
The materials science procedural text corpus: Annotating materials synthesis procedures with shallow semantic structures
Sheshera Mysore, Zachary Jensen, Edward Kim, Kevin Huang, Haw-Shiuan Chang, Emma Strubell, Jeffrey Flanigan, Andrew McCallum, and Elsa Olivetti. 2019 · 2019
Cited alongside, same era.
Covid-19 tweets analysis through transformer language models
Abdul Hameed Azeemi and Adeel Waheed. 2021 · 2021
Closest in time.
Co2l: Contrastive continual learning
Hyuntak Cha, Jaeho Lee, and Jinwoo Shin. 2021 · 2021
Closest in time.
Time-aware language models as temporal knowledge bases
Bhuwan Dhingra, Jeremy R. Cole, Julian Martin Eisenschlos, D. Gillick, Jacob Eisenstein, and William W. Cohen. 2021 · 2021
Closest in time.
Seed: Self-supervised distillation for visual representation
Zhiyuan Fang, Jianfeng Wang, Lijuan Wang, L. Zhang, Yezhou Yang, and Zicheng Liu. 2021 · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Patient knowledge distillation for BERT model compression
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu. 2019 · 2019
Cited alongside, same era.
Sentence embedding alignment for lifelong relation extraction
Hong Wang, Wenhan Xiong, Mo Yu, Xiaoxiao Guo, Shiyu Chang, and William Yang Wang. 2019 · 2019
Cited alongside, same era.
An empirical investigation towards efficient multi-domain language model pre-training
Kristjan Arumae, Qing Sun, and Parminder Bhatia. 2020 · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Cited alongside, same era.
Lifelong language knowledge distillation
Yung-Sung Chuang, Shang-Yu Su, and Yun-Nung Chen. 2020 · 2020
Cited alongside, same era.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020 · 2020
Cited alongside, same era.
TinyBERT: Distilling BERT for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2020 · 2020
Cited alongside, same era.
Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021 · 2021
Closest in time.
Demix layers: Disentangling domains for modular language modeling
Suchin Gururangan, Michael Lewis, Ari Holtzman, Noah A. Smith, and Luke Zettlemoyer. 2021 · 2021
Closest in time.
Dynamic language models for continuously evolving content
Spurthi Amba Hombaiah, Tao Chen, Mingyang Zhang, Michael Bendersky, and Marc-Alexander Najork. 2021 · 2021
Closest in time.
Continual learning for text classification with information disentanglement based regularization
Yufan Huang, Yanzhe Zhang, Jiaao Chen, Xuezhi Wang, and Diyi Yang. 2021 · 2021
Closest in time.
Towards continual knowledge learning of language models
Joel Jang, Seonghyeon Ye, Sohee Yang, Joongbo Shin, Janghoon Han, Gyeonghun Kim, Stanley Jungkyu Choi, and Minjoon Seo. 2021 · 2021
Closest in time.
Rational LAMOL: A rationale-based lifelong learning framework
Kasidis Kanwatchara, Thanapapas Horsuwan, Piyawat Lertvittayakumjorn, Boonserm Kijsirikul, and Peerapon Vateekul. 2021 · 2021
Closest in time.
Assessing temporal generalization in neural language models
Angeliki Lazaridou, A. Kuncoro, E. Gribovskaya, Devang Agrawal, Adam Liska, Tayfun Terzi, Mai Gimenez, Cyprien de Masson d’Autume, Sebastian Ruder, Dani Yogatama, Kris Cao, Tomás Kociský, Susannah Young, and P. Blunsom. 2021 · 2021
Closest in time.
Time waits for no one! analysis and challenges of temporal misalignment
Kelvin Luu, Daniel Khashabi, Suchin Gururangan, Karishma Mandyam, and Noah A. Smith. 2021 · 2021
Closest in time.
Multidomain pretrained language models for green NLP
Antonis Maronikolakis and Hinrich Schütze. 2021 · 2021
Closest in time.
Merging models with fisher-weighted averaging
Michael Matena and Colin Raffel. 2021 · 2021
Closest in time.
AdapterFusion: Non-destructive task composition for transfer learning
Jonas Pfeiffer, Aishwarya Kamath, Andreas Rücklé, Kyunghyun Cho, and Iryna Gurevych. 2021 · 2021
Closest in time.
Temporal adaptation of bert and performance on downstream document classification: Insights from social media
Paul Röttger and J. Pierrehumbert. 2021 · 2021
Closest in time.
Adapt-and-distill: Developing small, fast and effective pretrained language models for domains
Yunzhi Yao, Shaohan Huang, Wenhui Wang, Li Dong, and Furu Wei. 2021 · 2021
Closest in time.
Pretrained language model in continual learning: A comparative study
Tongtong Wu, Massimo Caccia, Zhuang Li, Yuan-Fang Li, Guilin Qi, and Gholamreza Haffari. 2022 · 2022
Closest in time.