Fetching the paper…
Reading the bibliography…
Continual learning necessitates the continual adaptation of models to newly emerging tasks while minimizing the catastrophic forgetting of old ones.
1907
Earlier work this paper cites.
1911
Earlier work this paper cites.
R. Weischedel and A. Brunstein, “BBN pronoun coreference and entity type corpus,” Linguistic Data Consortium, Philadelphia , vol. 112, 2005. [Online]. Available: https://catalog.ldc.upenn.edu/LDC2005T33
2005
Earlier work this paper cites.
2006
Earlier work this paper cites.
C. Walker, S. Strassel, J. Medero, and K. Maeda, “ACE 2005 multilingual training corpus,” Linguistic Data Consortium, Philadelphia , vol. 57, p. 45, 2006. [Online]. Available: https://catalog.ldc.upenn.edu/LDC2006T06
2006
Earlier work this paper cites.
R. Weischedel, M. Palmer, M. Marcus, E. Hovy, S. Pradhan, L. Ramshaw, N. Xue, A. Taylor, J. Kaufman, M. Franchini et al. , “OntoNotes release 5.0 ldc2013t19,” Linguistic Data Consortium, Philadelphia, PA , vol. 23, 2013. [Online]. Available: https://catalog.ldc.upenn.edu/LDC2013T19
2013
Earlier work this paper cites.
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska et al. , “Overcoming catastrophic forgetting in neural networks,” Proceedings of the national academy of sciences , vol. 114, no. 13, pp. 3521–3526, 2017. [Online]. Available: https://www.pnas.org/doi/epdf/10.1073/pnas.1611835114
2017
Earlier work this paper cites.
S.-W. Lee, J.-H. Kim, J. Jun, J.-W. Ha, and B.-T. Zhang, “Overcoming catastrophic forgetting by incremental moment matching,” in Proceedings of the 31st International Conference on Neural Information Processing Systems , 2017, pp. 4655–4665. [Online]. Available: https://proceedings.neurips.cc/paper/2017/file/f708f064faaf32a43e4d3c784e6af9ea-Paper.pdf
2017
Earlier work this paper cites.
Z. Li and D. Hoiem, “Learning without forgetting,” IEEE transactions on pattern analysis and machine intelligence , vol. 40, no. 12, pp. 2935–2947, 2017. [Online]. Available: https://ieeexplore.ieee.org/abstract/document/8107520
2017
Earlier work this paper cites.
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean, “Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,” in International Conference on Learning Representations , 2017. [Online]. Available: https://openreview.net/pdf?id=B1ckMDqlg
2017
Earlier work this paper cites.
Y. Zhang, V. Zhong, D. Chen, G. Angeli, and C. D. Manning, “Position-aware attention and supervised data improve slot filling,” in Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing , 2017, pp. 35–45. [Online]. Available: https://nlp.stanford.edu/pubs/zhang2017tacred.pdf
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Chaudhry, P. K. Dokania, T. Ajanthan, and P. H. Torr, “Riemannian walk for incremental learning: Understanding forgetting and intransigence,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 532–547. [Online]. Available: https://openaccess.thecvf.com/content_ECCV_2018/papers/Arslan_Chaudhry__Riemannian_Walk_ECCV_2018_paper.pdf
2018
Earlier work this paper cites.
D. Isele and A. Cosgun, “Selective experience replay for lifelong learning,” in Proceedings of the AAAI Conference on Artificial Intelligence , 2018. [Online]. Available: https://ojs.aaai.org/index.php/AAAI/article/view/11595/11454
2018
Earlier work this paper cites.
X. Han, H. Zhu, P. Yu, Z. Wang, Y. Yao, Z. Liu, and M. Sun, “FewRel: A large-scale supervised few-shot relation classification dataset with state-of-the-art evaluation,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing , 2018, pp. 4803–4809. [Online]. Available: https://aclanthology.org/D18-1514.pdf
2018
Earlier work this paper cites.
G. I. Parisi, R. Kemker, J. L. Part, C. Kanan, and S. Wermter, “Continual lifelong learning with neural networks: A review,” Neural Networks , vol. 113, pp. 54–71, 2019. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0893608019300231
2019
Earlier work this paper cites.
D. Rolnick, A. Ahuja, J. Schwarz, T. Lillicrap, and G. Wayne, “Experience replay for continual learning,” Advances in Neural Information Processing Systems , vol. 32, 2019. [Online]. Available: https://proceedings.neurips.cc/paper/2019/file/fa7cdfad1a5aaf8370ebeda47a1ff1c3-Paper.pdf
2019
Earlier work this paper cites.
H. Wang, W. Xiong, M. Yu, X. Guo, S. Chang, and W. Y. Wang, “Sentence embedding alignment for lifelong relation extraction,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) , 2019, pp. 796–806. [Online]. Available: https://aclanthology.org/N19-1086.pdf
2019
Earlier work this paper cites.
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. De Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly, “Parameter-efficient transfer learning for NLP,” in International Conference on Machine Learning . PMLR, 2019, pp. 2790–2799. [Online]. Available: http://proceedings.mlr.press/v97/houlsby19a/houlsby19a.pdf
2019
Cited alongside, same era.
L. B. Soares, N. Fitzgerald, J. Ling, and T. Kwiatkowski, “Matching the blanks: Distributional similarity for relation learning,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , 2019, pp. 2895–2905. [Online]. Available: https://aclanthology.org/P19-1279.pdf
2019
Cited alongside, same era.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf
2020
Cited alongside, same era.
B. Lester, R. Al-Rfou, and N. Constant, “The power of scale for parameter-efficient prompt tuning,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , 2021, pp. 3045–3059. [Online]. Available: https://aclanthology.org/2021.emnlp-main.243.pdf
2021
Later among the works it cites.
N. Ding, G. Xu, Y. Chen, X. Wang, X. Han, P. Xie, H. Zheng, and Z. Liu, “Few-NERD: A few-shot named entity recognition dataset,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) , 2021, pp. 3198–3213. [Online]. Available: https://aclanthology.org/2021.acl-long.248.pdf
2021
Later among the works it cites.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al. , “Training language models to follow instructions with human feedback,” Advances in Neural Information Processing Systems , vol. 35, pp. 27 730–27 744, 2022. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2022/file/b1efde53be364a73914f58805a001731-Paper-Conference.pdf
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
X. Han, Y. Dai, T. Gao, Y. Lin, Z. Liu, P. Li, M. Sun, and J. Zhou, “Continual relation learning via episodic memory activation and reconsolidation,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , 2020, pp. 6429–6440. [Online]. Available: https://aclanthology.org/2020.acl-main.573.pdf
2020
Cited alongside, same era.
M. Zhao, T. Lin, F. Mi, M. Jaggi, and H. Schütze, “Masking as an efficient alternative to finetuning for pretrained language models,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing , 2020, pp. 2226–2241. [Online]. Available: https://aclanthology.org/2020.emnlp-main.174.pdf
2020
Cited alongside, same era.
J. Zhang, J. Zhang, S. Ghosh, D. Li, S. Tasci, L. Heck, H. Zhang, and C.-C. J. Kuo, “Class-incremental learning via deep model consolidation,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2020, pp. 1131–1140. [Online]. Available: https://openaccess.thecvf.com/content_WACV_2020/papers/Zhang_Class-incremental_Learning_via_Deep_Model_Consolidation_WACV_2020_paper.pdf
2020
Cited alongside, same era.
2021
Cited alongside, same era.
A. Daruna, M. Gupta, M. Sridharan, and S. Chernova, “Continual learning of knowledge graph embeddings,” IEEE Robotics and Automation Letters , vol. 6, no. 2, pp. 1128–1135, 2021. [Online]. Available: https://ieeexplore.ieee.org/abstract/document/9343669
2021
Cited alongside, same era.
N. Monaikul, G. Castellucci, S. Filice, and O. Rokhlenko, “Continual learning for named entity recognition,” in Proceedings of the AAAI Conference on Artificial Intelligence , 2021, pp. 13 570–13 577. [Online]. Available: https://ojs.aaai.org/index.php/AAAI/article/view/17600/17407
2021
Cited alongside, same era.
X. Gu, L. Liu, H. Yu, J. Li, C. Chen, and J. Han, “On the transformer growth for progressive BERT training,” in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , 2021, pp. 5174–5180. [Online]. Available: https://aclanthology.org/2021.naacl-main.406.pdf
2021
Cited alongside, same era.
X. L. Li and P. Liang, “Prefix-Tuning: Optimizing continuous prompts for generation,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) , 2021, pp. 4582–4597. [Online]. Available: https://aclanthology.org/2021.acl-long.353.pdf
2021
Cited alongside, same era.
T. Gao, A. Fisch, and D. Chen, “Making pre-trained language models better few-shot learners,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) , 2021, pp. 3816–3830. [Online]. Available: https://aclanthology.org/2021.acl-long.295.pdf
2021
Cited alongside, same era.
2022
Later among the works it cites.
2022
Later among the works it cites.
X. Jin, D. Zhang, H. Zhu, W. Xiao, S.-W. Li, X. Wei, A. Arnold, and X. Ren, “Lifelong pretraining: Continually adapting language models to emerging corpora,” in Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , 2022, pp. 4764–4780. [Online]. Available: https://aclanthology.org/2022.naacl-main.351.pdf
2022
Later among the works it cites.
K. Zhao, H. Xu, J. Yang, and K. Gao, “Consistent representation learning for continual relation extraction,” in Findings of the Association for Computational Linguistics: ACL 2022 , 2022, pp. 3402–3411. [Online]. Available: https://aclanthology.org/2022.findings-acl.268.pdf
2022
Later among the works it cites.
E. B. Zaken, Y. Goldberg, and S. Ravfogel, “Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , 2022, pp. 1–9. [Online]. Available: https://aclanthology.org/2022.acl-short.1.pdf
2022
Later among the works it cites.
A. Razdaibiedina, Y. Mao, R. Hou, M. Khabsa, M. Lewis, and A. Almahairi, “Progressive prompts: Continual learning for language models,” in The Eleventh International Conference on Learning Representations , 2022. [Online]. Available: https://openreview.net/pdf?id=UJTgQBc91_
2022
Later among the works it cites.
Y. Qin, J. Zhang, Y. Lin, Z. Liu, P. Li, M. Sun, and J. Zhou, “Elle: Efficient lifelong pre-training for emerging data,” in Findings of the Association for Computational Linguistics: ACL 2022 , 2022, pp. 2789–2810. [Online]. Available: https://aclanthology.org/2022.findings-acl.220.pdf
2022
Later among the works it cites.
W. Fedus, B. Zoph, and N. Shazeer, “Switch Transformers: Scaling to trillion parameter models with simple and efficient sparsity,” The Journal of Machine Learning Research , vol. 23, no. 1, pp. 5232–5270, 2022. [Online]. Available: https://jmlr.org/papers/volume23/21-0998/21-0998.pdf
2022
Later among the works it cites.
Z. Zhang, Y. Lin, Z. Liu, P. Li, M. Sun, and J. Zhou, “MoEfication: Transformer feed-forward layers are mixtures of experts,” in Findings of the Association for Computational Linguistics: ACL 2022 , 2022, pp. 877–890. [Online]. Available: https://aclanthology.org/2022.findings-acl.71.pdf
2022
Later among the works it cites.
S. Gururangan, M. Lewis, A. Holtzman, N. A. Smith, and L. Zettlemoyer, “DEMix layers: Disentangling domains for modular language modeling,” in Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , 2022, pp. 5557–5576. [Online]. Available: https://aclanthology.org/2022.naacl-main.407.pdf
2022
Later among the works it cites.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Z. Ke, Y. Shao, H. Lin, T. Konishi, G. Kim, and B. Liu, “Continual pre-training of language models,” in The Eleventh International Conference on Learning Representations , 2023. [Online]. Available: https://openreview.net/pdf?id=m_GDIItaI3o
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.