Fetching the paper…
Reading the bibliography…
Recent work on applying large language models (LMs) achieves impressive performance in many NLP applications.
Continual learning for sentence representations using conceptors
Tianlin Liu, Lyle Ungar, and Joao Sedoc. 2019a · 1904
Earlier work this paper cites.
Three scenarios for continual learning
Gido M. van de Ven and Andreas S. Tolias. 2019 · 1904
Earlier work this paper cites.
Claudio Greco, Barbara Plank, Raquel Fernández, and Raffaella Bernardi. 2019 · 1906
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael McCloskey and Neal J Cohen. 1989 · 1989
Earlier work this paper cites.
Dark experience for general continual learning: a strong, simple baseline
Pietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati, and Simone Calderara. 2020 · 2004
Earlier work this paper cites.
Don’t stop pretraining: adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A Smith. 2020a · 2004
Earlier work this paper cites.
Exploring fine-tuning techniques for pre-trained cross-lingual models via continual learning
Zihan Liu, Genta Indra Winata, Andrea Madotto, and Pascale Fung. 2020b · 2004
Earlier work this paper cites.
Lifelong language knowledge distillation
Yung-Sung Chuang, Shang-Yu Su, and Yun-Nung Chen. 2020 · 2010
Earlier work this paper cites.
Continual learning in task-oriented dialogue systems
Andrea Madotto, Zhaojiang Lin, Zhenpeng Zhou, Seungwhan Moon, Paul Crook, Bing Liu, Zhou Yu, Eunjoon Cho, and Zhiguang Wang. 2020 · 2012
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, Jeff Dean, et al. 2015 · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Jake Zhao, and Yann LeCun. 2015 · 2015
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil C. Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell. 2016 · 2016
Earlier work this paper cites.
Pathnet: Evolution channels gradient descent in super neural networks
Chrisantha Fernando, Dylan Banarse, Charles Blundell, Yori Zwols, David Ha, Andrei A. Rusu, Alexander Pritzel, and Daan Wierstra. 2017 · 2017
Earlier work this paper cites.
Gradient episodic memory for continual learning
David Lopez-Paz and Marc’Aurelio Ranzato. 2017 · 2017
Earlier work this paper cites.
icarl: Incremental classifier and representation learning
Sylvestre-Alvise Rebuffi, Alexander Kolesnikov, Georg Sperl, and Christoph H. Lampert. 2017 · 2017
Earlier work this paper cites.
Continual learning in generative adversarial nets
Ari Seff, Alex Beatson, Daniel Suo, and Han Liu. 2017 · 2017
Earlier work this paper cites.
Continual learning with deep generative replay
Hanul Shin, Jung Kwon Lee, Jaehong Kim, and Jiwon Kim. 2017 · 2017
Earlier work this paper cites.
Few-shot learning through an information retrieval lens
Eleni Triantafillou, Richard S. Zemel, and Raquel Urtasun. 2017 · 2017
Earlier work this paper cites.
One-shot unsupervised cross domain translation
Sagie Benaim and Lior Wolf. 2018 · 2018
Cited alongside, same era.
Lifelong machine learning
Zhiyuan Chen and Bing Liu. 2018 · 2018
Cited alongside, same era.
Overcoming catastrophic interference using conceptor-aided backpropagation
Xu He and Herbert Jaeger. 2018 · 2018
Cited alongside, same era.
Few-shot charge prediction with discriminative legal attributes
Zikun Hu, Xiang Li, Cunchao Tu, Zhiyuan Liu, and Maosong Sun. 2018 · 2018
Cited alongside, same era.
Measuring the evolution of a scientific field through citation frames
David Jurgens, Srijan Kumar, Raine Hoover, Daniel A. McFarland, and Dan Jurafsky. 2018 · 2018
Cited alongside, same era.
Regularized training objective for continued training for domain adaptation in neural machine translation
Huda Khayrallah, Brian Thompson, Kevin Duh, and Philipp Koehn. 2018 · 2018
Lamol: Language modeling is all you need for lifelong language learning
Fan-Keng Sun, Cheng-Hao Ho, and Hung-Yi Lee. 2020 · 2020
Later among the works it cites.
Efficient meta lifelong-learning with limited memory
Zirui Wang, Sanket Vaibhav Mehta, Barnabás Póczos, and Jaime Carbonell. 2020 · 2020
Later among the works it cites.
Making pre-trained language models better few-shot learners
Tianyu Gao, Adam Fisch, and Danqi Chen. 2021 · 2021
Later among the works it cites.
Ppt: Pre-trained prompt tuning for few-shot learning
Yuxian Gu, Xu Han, Zhiyuan Liu, and Minlie Huang. 2021 · 2021
Later among the works it cites.
Demix layers: Disentangling domains for modular language modeling
Suchin Gururangan, Mike Lewis, Ari Holtzman, Noah A. Smith, and Luke Zettlemoyer. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Multi-task identification of entities, relations, and coreference for scientific knowledge graph construction
Yi Luan, Luheng He, Mari Ostendorf, and Hannaneh Hajishirzi. 2018 · 2018
Cited alongside, same era.
Overcoming catastrophic forgetting with hard attention to the task
Joan Serrà, Didac Suris, Marius Miron, and Alexandros Karatzoglou. 2018 · 2018
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Parameter-efficient transfer learning for NLP
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019 · 2019
Cited alongside, same era.
Compositional language continual learning
Yuanpeng Li, Liang Zhao, Kenneth Church, and Mohamed Elhoseiny. 2019 · 2019
Cited alongside, same era.
A progressive model to enable continual learning for semantic slot filling
Yilin Shen, Xiangyu Zeng, and Hongxia Jin. 2019 · 2019
Cited alongside, same era.
Junxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick, and Graham Neubig. 2021 · 2021
Later among the works it cites.
Continual learning for text classification with information disentanglement based regularization
Yufan Huang, Yanzhe Zhang, Jiaao Chen, Xuezhi Wang, and Diyi Yang. 2021 · 2021
Later among the works it cites.
Learn continually, generalize rapidly: Lifelong knowledge accumulation for few-shot learning
Xisen Jin, Bill Yuchen Lin, Mohammad Rostami, and Xiang Ren. 2021 · 2021
Later among the works it cites.
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant. 2021 · 2021
Later among the works it cites.
An empirical investigation of the role of pre-training in lifelong learning
Sanket Vaibhav Mehta, Darshan Patil, Sarath Chandar, and Emma Strubell. 2021 · 2021
Later among the works it cites.
LFPT5: A unified framework for lifelong few-shot language learning based on prompt tuning of T5
Chengwei Qin and Shafiq Joty. 2021 · 2021
Later among the works it cites.
Scaling language models: Methods, analysis & insights from training gopher
Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, H. Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks, Maribeth Rauh, Po-Sen Huang, Amelia Glaese, Johannes Welbl, Sumanth Dathathri, Saffron Huang, Jonathan Uesato, John Mellor, Irina Higgins, Antonia Creswell, Nat McAleese, Amy Wu, Erich Elsen, Siddhant M. Jayakumar, Elena Buchatskaya, David Budden, Esme Sutherland, Karen Simonyan, Michela Paganini, Laurent Sifre, Lena Martens, Xiang Lorraine Li, Adhiguna Kuncoro, Aida Nematzadeh, Elena Gribovskaya, Domenic Donato, Angeliki Lazaridou, Arthur Mensch, Jean-Baptiste Lespiau, Maria Tsimpoukelli, Nikolai Grigorev, Doug Fritz, Thibault Sottiaux, Mantas Pajarskas, Toby Pohlen, Zhitao Gong, Daniel Toyama, Cyprien de Masson d’Autume, Yujia Li, Tayfun Terzi, Vladimir Mikulik, Igor Babuschkin, Aidan Clark, Diego de Las Casas, Aurelia Guy, Chris Jones, James Bradbury, Matthew Johnson, Blake A. Hechtman, Laura Weidinger, Iason Gabriel, William S. Isaac, Edward Lockhart, Simon Osindero, Laura Rimell, Chris Dyer, Oriol Vinyals, Kareem Ayoub, Jeff Stanway, Lorrayne Bennett, Demis Hassabis, Koray Kavukcuoglu, and Geoffrey Irving. 2021 · 2021
Later among the works it cites.
Exploiting cloze-questions for few-shot text classification and natural language inference
Timo Schick and Hinrich Schütze. 2021 · 2021
Later among the works it cites.
Incremental few-shot text classification with multi-round new classes: Formulation, dataset and system
Congying Xia, Wenpeng Yin, Yihao Feng, and Philip S. Yu. 2021 · 2021
Later among the works it cites.
Online continual learning through mutual information maximization
Yiduo Guo, Bing Liu, and Dongyan Zhao. 2022 · 2022
Closest in time.
ELLE: efficient lifelong pre-training for emerging data
Yujia Qin, Jiajie Zhang, Yankai Lin, Zhiyuan Liu, Peng Li, Maosong Sun, and Jie Zhou. 2022 · 2022
Closest in time.
Shaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, Elton Zheng, Rewon Child, Reza Yazdani Aminabadi, Julie Bernauer, Xia Song, Mohammad Shoeybi, Yuxiong He, Michael Houston, Saurabh Tiwary, and Bryan Catanzaro. 2022 · 2022
Closest in time.
Contintin: Continual learning from task instructions
Wenpeng Yin, Jia Li, and Caiming Xiong. 2022 · 2022
Closest in time.