Fetching the paper…
Reading the bibliography…
The mainstream machine learning paradigms for NLP often work with two underlying presumptions.
Universal grammar
Montague Richard. 1970 · 1970
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael McCloskey and Neal J. Cohen. 1989 · 1989
Earlier work this paper cites.
A system for incremental learning based on algorithmic probability
Ray J. Solomonoff. 1989 · 1989
Earlier work this paper cites.
Lifelong robot learning
Sebastian Thrun and Tom M. Mitchell. 1995 · 1995
Earlier work this paper cites.
Syntactic structures
Noam Chomsky. 2002 · 2002
Earlier work this paper cites.
The task rehearsal method of life-long learning: Overcoming impoverished data
Daniel L. Silver and Robert E. Mercer. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Toward an architecture for never-ending language learning
Andrew Carlson, Justin Betteridge, Bryan Kisiel, Burr Settles, Estevam R. Hruschka Jr., and Tom M. Mitchell. 2010 · 2010
Earlier work this paper cites.
The turking test: Can language models understand instructions?
Avia Efrat and Omer Levy. 2020 · 2010
Earlier work this paper cites.
Few-shot text generation with pattern-exploiting training
Timo Schick and Hinrich Schütze. 2020 · 2012
Earlier work this paper cites.
Learning from natural instructions
Dan Goldwasser and Dan Roth. 2014 · 2014
Earlier work this paper cites.
Lifelong learning for sentiment classification
Zhiyuan Chen, Nianzu Ma, and Bing Liu. 2015 · 2015
Earlier work this paper cites.
Toward continual learning for conversational agents
Sungjin Lee. 2017 · 2017
Cited alongside, same era.
Distantly supervised lifelong learning for large-scale social media sentiment analysis
Rui Xia, Jie Jiang, and Huihui He. 2017 · 2017
Cited alongside, same era.
The natural language decathlon: Multitask learning as question answering
Bryan McCann, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher. 2018 · 2018
Cited alongside, same era.
Overcoming catastrophic forgetting with hard attention to the task
Joan Serrà, Didac Suris, Marius Miron, and Alexandros Karatzoglou. 2018 · 2018
Cited alongside, same era.
Quoref: A reading comprehension dataset with questions requiring coreferential reasoning
Pradeep Dasigi, Nelson F. Liu, Ana Marasovic, Noah A. Smith, and Matt Gardner. 2019 · 2019
Cited alongside, same era.
BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Later among the works it cites.
Using task descriptions in lifelong machine learning for improved performance and zero-shot transfer
Mohammad Rostami, David Isele, and Eric Eaton. 2020 · 2020
Later among the works it cites.
Winogrande: An adversarial winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2020 · 2020
Later among the works it cites.
Learn continually, generalize rapidly: Lifelong knowledge accumulation for few-shot learning
Xisen Jin, Bill Yuchen Lin, Mohammad Rostami, and Xiang Ren. 2021 · 2021
Later among the works it cites.
Adapting BERT for continual learning of a sequence of aspect sentiment classification tasks
Zixuan Ke, Hu Xu, and Bing Liu. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Episodic memory in lifelong language learning
Cyprien de Masson d’Autume, Sebastian Ruder, Lingpeng Kong, and Dani Yogatama. 2019 · 2019
Cited alongside, same era.
Cosmos QA: machine reading comprehension with contextual commonsense reasoning
Lifu Huang, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Sentence embedding alignment for lifelong relation extraction
Hong Wang, Wenhan Xiong, Mo Yu, Xiaoxiao Guo, Shiyu Chang, and William Yang Wang. 2019 · 2019
Cited alongside, same era.
Continual lifelong learning in natural language processing: A survey
Magdalena Biesialska, Katarzyna Biesialska, and Marta R. Costa-jussà. 2020 · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Cited alongside, same era.
Dynamic memory to alleviate catastrophic forgetting in continuous learning settings
Johannes Hofmanninger, Matthias Perkonigg, James A. Brink, Oleg Pianykh, Christian Herold, and Georg Langs. 2020 · 2020
Cited alongside, same era.
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2021 · 2021
Later among the works it cites.
Cross-task generalization via natural language crowdsourcing instructions
Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2021 · 2021
Later among the works it cites.
Exploiting cloze-questions for few-shot text classification and natural language inference
Timo Schick and Hinrich Schütze. 2021 · 2021
Later among the works it cites.
Incremental few-shot text classification with multi-round new classes: Formulation, dataset and system
Congying Xia, Wenpeng Yin, Yihao Feng, and Philip S. Yu. 2021 · 2021
Later among the works it cites.
DER: dynamically expandable representation for class incremental learning
Shipeng Yan, Jiangwei Xie, and Xuming He. 2021 · 2021
Later among the works it cites.
Negative training for neural dialogue response generation
Tianxing He and James R. Glass. 2020 · 2058
Closest in time.