Fetching the paper…
Reading the bibliography…
Continual learning (CL) aims to enable information systems to learn from a continuous data stream across time.
Curriculum learning for domain adaptation in neural machine translation
Xuan Zhang, Pamela Shapiro, Gaurav Kumar, Paul McNamee, Marine Carpuat, and Kevin Duh. 2019 · 1915
Earlier work this paper cites.
Online distilling from checkpoints for neural machine translation
Hao-Ran Wei, Shujian Huang, Ran Wang, Xin-yu Dai, and Jiajun Chen. 2019 · 1941
Earlier work this paper cites.
Incremental learning from noisy data
Jeffrey C Schlimmer and Richard H Granger. 1986 · 1986
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael McCloskey and Neal J Cohen. 1989 · 1989
Earlier work this paper cites.
A system for incremental learning based on algorithmic probability
Ray J Solomonoff. 1989 · 1989
Earlier work this paper cites.
Learning and development in neural networks: The importance of starting small
Jeffrey L Elman. 1993 · 1993
Earlier work this paper cites.
Effective learning in dynamic environments by explicit context tracking
Gerhard Widmer and Miroslav Kubat. 1993 · 1993
Earlier work this paper cites.
Continual learning in reinforcement environments
Mark B. Ring. 1994 · 1994
Earlier work this paper cites.
Catastrophic forgetting, rehearsal and pseudorehearsal
Anthony Robins. 1995 · 1995
Earlier work this paper cites.
Is learning the n-th thing any easier than learning the first?
Sebastian Thrun. 1996 · 1996
Earlier work this paper cites.
Multitask learning
Rich Caruana. 1997 · 1997
Earlier work this paper cites.
Learning to Learn: Introduction and Overview , page 3–17. Kluwer Academic Publishers, USA
Sebastian Thrun and Lorien Pratt. 1998 · 1998
Earlier work this paper cites.
On-Line Learning and Stochastic Approximations , page 9–42. Cambridge University Press, USA
Léon Bottou. 1999 · 1999
Earlier work this paper cites.
The task rehearsal method of life-long learning: Overcoming impoverished data
Daniel L Silver and Robert E Mercer. 2002 · 2002
Earlier work this paper cites.
Large scale online learning
Léon Bottou and Yann LeCun. 2004 · 2004
Earlier work this paper cites.
Prediction, Learning, and Games
Nicolo Cesa-Bianchi and Gabor Lugosi. 2006 · 2006
Earlier work this paper cites.
Strategies for lifelong knowledge extraction from the web
Michele Banko and Oren Etzioni. 2007 · 2007
Earlier work this paper cites.
Curriculum learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston. 2009 · 2009
Earlier work this paper cites.
Causality
Judea Pearl. 2009 · 2009
Earlier work this paper cites.
Toward an architecture for never-ending language learning
Andrew Carlson, Justin Betteridge, Bryan Kisiel, Burr Settles, Estevam R Hruschka, and Tom M Mitchell. 2010 · 2010
Earlier work this paper cites.
A survey on transfer learning
S. J. Pan and Q. Yang. 2010 · 2010
Earlier work this paper cites.
icarl: Incremental classifier and representation learning
Sylvestre-Alvise Rebuffi, Alexander Kolesnikov, Georg Sperl, and Christoph H Lampert. 2017b · 2010
Earlier work this paper cites.
Caltech-UCSD Birds 200
P. Welinder, S. Branson, T. Mita, C. Wah, F. Schroff, S. Belongie, and P. Perona. 2010 · 2010
Earlier work this paper cites.
The Caltech-UCSD Birds-200-2011 Dataset
C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie. 2011 · 2011
Earlier work this paper cites.
Online learning and online convex optimization
Shai Shalev-Shwartz. 2012 · 2012
Earlier work this paper cites.
Lifelong machine learning systems: Beyond learning algorithms
Daniel L Silver, Qiang Yang, and Lianghao Li. 2013 · 2013
Earlier work this paper cites.
Incremental on-line adaptation of pomdp-based dialogue managers to extended domains
Milica Gasic, Dongho Kim, Pirros Tsiakoulis, Catherine Breslin, Matthew Henderson, Martin Szummer, Blaise Thomson, and Steve J. Young. 2014 · 2014
Earlier work this paper cites.
Lifelong learning for sentiment classification
Zhiyuan Chen, Nianzu Ma, and Bing Liu. 2015 · 2015
Earlier work this paper cites.
Multi-task learning for multiple language translation
Daxiang Dong, Hua Wu, Wei He, Dianhai Yu, and Haifeng Wang. 2015 · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeffrey Dean. 2015 · 2015
Earlier work this paper cites.
Human-level concept learning through probabilistic program induction
Brenden M Lake, Ruslan Salakhutdinov, and Joshua B Tenenbaum. 2015 · 2015
Earlier work this paper cites.
Stanford neural machine translation systems for spoken language domain
Minh-Thang Luong and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Online multitask learning for machine translation quality estimation
José G. C. de Souza, Matteo Negri, Elisa Ricci, and Marco Turchi. 2015 · 2015
Earlier work this paper cites.
Multi-way, multilingual neural machine translation with a shared attention mechanism
Orhan Firat, Kyunghyun Cho, and Yoshua Bengio. 2016 · 2016
Earlier work this paper cites.
Fast domain adaptation for neural machine translation
Markus Freitag and Yaser Al-Onaizan. 2016 · 2016
Earlier work this paper cites.
Toward multilingual neural machine translation with universal encoder and decoder
Thanh-Le Ha, Jan Niehues, and Alexander H. Waibel. 2016 · 2016
Earlier work this paper cites.
Less-forgetting learning in deep neural networks
Heechul Jung, Jeongwoo Ju, Minju Jung, and Junmo Kim. 2016 · 2016
Earlier work this paper cites.
Sequence-level knowledge distillation
Yoon Kim and Alexander M. Rush. 2016 · 2016
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al. 2016 · 2016
Earlier work this paper cites.
Ask me anything: Dynamic memory networks for natural language processing
Ankit Kumar, Ozan Irsoy, Peter Ondruska, Mohit Iyyer, James Bradbury, Ishaan Gulrajani, Victor Zhong, Romain Paulus, and Richard Socher. 2016 · 2016
Earlier work this paper cites.
Learning without forgetting
Zhizhong Li and Derek Hoiem. 2016 · 2016
Earlier work this paper cites.
Andrei A Rusu, Neil C Rabinowitz, Guillaume Desjardins, Hubert Soyer, James Kirkpatrick, Koray Kavukcuoglu, Razvan Pascanu, and Raia Hadsell. 2016 · 2016
Earlier work this paper cites.
Lifelong-RL: Lifelong relaxation labeling for separating entities and aspects in opinion targets
Lei Shu, Bing Liu, Hu Xu, and Annice Kim. 2016 · 2016
Earlier work this paper cites.
Continuously learning neural dialogue management
Pei-Hao Su, Milica Gasic, Nikola Mrksic, Lina Rojas-Barahona, Stefan Ultes, David Vandyke, Tsung-Hsien Wen, and Steve Young. 2016 · 2016
Earlier work this paper cites.
Expert gate: Lifelong learning with a network of experts
Rahaf Aljundi, Punarjay Chakravarty, and Tinne Tuytelaars. 2017 · 2017
Earlier work this paper cites.
An empirical comparison of domain adaptation methods for neural machine translation
Chenhui Chu, Raj Dabre, and Sadao Kurohashi. 2017 · 2017
Earlier work this paper cites.
Multi-domain neural machine translation through unsupervised adaptation
M. Amin Farajian, Marco Turchi, Matteo Negri, and Marcello Federico. 2017 · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine. 2017 · 2017
Cited alongside, same era.
Audio set: An ontology and human-labeled dataset for audio events
Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter. 2017 · 2017
Cited alongside, same era.
Google’s multilingual neural machine translation system: Enabling zero-shot translation
Melvin Johnson, Mike Schuster, Quoc V. Le, Maxim Krikun, Yonghui Wu, Zhifeng Chen, Nikhil Thorat, Fernanda Viégas, Martin Wattenberg, Greg Corrado, Macduff Hughes, and Jeffrey Dean. 2017 · 2017
Cited alongside, same era.
Learning to remember rare events
Łukasz Kaiser, Ofir Nachum, Aurko Roy, and Samy Bengio. 2017 · 2017
Cited alongside, same era.
Systematic generalization: What is required and can it be learned?
Dzmitry Bahdanau, Shikhar Murty, Michael Noukhovitch, Thien Huu Nguyen, Harm de Vries, and Aaron Courville. 2019 · 2019
Later among the works it cites.
Simple, scalable adaptation for neural machine translation
Ankur Bapna and Orhan Firat. 2019 · 2019
Later among the works it cites.
From system 1 deep learning to system 2 deep learning
Yoshua Bengio. 2019 · 2019
Later among the works it cites.
Episodic memory in lifelong language learning
Cyprien de Masson d’Autume, Sebastian Ruder, Lingpeng Kong, and Dani Yogatama. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Continual learning via neural pruning
Siavash Golkar, Micheal Kagan, and Kyunghyun Cho. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Measuring catastrophic forgetting in neural networks
Ronald Kemker, Marc McClure, Angelina Abitino, Tyler L. Hayes, and Christopher Kanan. 2017 · 2017
Cited alongside, same era.
Core50: a new dataset and benchmark for continuous object recognition
Vincenzo Lomonaco and Davide Maltoni. 2017 · 2017
Cited alongside, same era.
Gradient episodic memory for continual learning
David Lopez-Paz and Marc’Aurelio Ranzato. 2017 · 2017
Cited alongside, same era.
Regularization techniques for fine-tuning in neural machine translation
Antonio Valerio Miceli Barone, Barry Haddow, Ulrich Germann, and Rico Sennrich. 2017 · 2017
Cited alongside, same era.
An overview of multi-task learning in deep neural networks
Sebastian Ruder. 2017 · 2017
Cited alongside, same era.
Continual learning with deep generative replay
Hanul Shin, Jung Kwon Lee, Jaehong Kim, and Jiwon Kim. 2017 · 2017
Cited alongside, same era.
Lifelong learning CRF for supervised aspect extraction
Lei Shu, Hu Xu, and Bing Liu. 2017 · 2017
Cited alongside, same era.
Later among the works it cites.
Psycholinguistics meets continual learning: Measuring catastrophic forgetting in visual question answering
Claudio Greco, Barbara Plank, Raquel Fernández, and Raffaella Bernardi. 2019 · 2019
Later among the works it cites.
Parameter-efficient transfer learning for NLP
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019 · 2019
Later among the works it cites.
Learn to grow: A continual structure learning framework for overcoming catastrophic forgetting
Xilai Li, Yingbo Zhou, Tianfu Wu, Richard Socher, and Caiming Xiong. 2019 · 2019
Later among the works it cites.
Continual learning for sentence representations using conceptors
Tianlin Liu, Lyle Ungar, and João Sedoc. 2019 · 2019
Later among the works it cites.
Lifelong and interactive learning of factual knowledge in dialogues
Sahisnu Mazumder, Bing Liu, Shuai Wang, and Nianzu Ma. 2019 · 2019
Later among the works it cites.
Meta-learning improves lifelong relation extraction
Abiola Obamuyide and Andreas Vlachos. 2019 · 2019
Later among the works it cites.
Practical deep learning with bayesian principles
Kazuki Osawa, Siddharth Swaroop, Mohammad Emtiyaz E Khan, Anirudh Jain, Runa Eschenhagen, Richard E Turner, and Rio Yokota. 2019 · 2019
Later among the works it cites.
Continual lifelong learning with neural networks: A review
German I Parisi, Ronald Kemker, Jose L Part, Christopher Kanan, and Stefan Wermter. 2019 · 2019
Later among the works it cites.
Latent replay for real-time continual learning
Lorenzo Pellegrini, Gabrile Graffieti, Vincenzo Lomonaco, and Davide Maltoni. 2019 · 2019
Later among the works it cites.
A comprehensive, application-oriented study of catastrophic forgetting in DNNs
B. Pfülb and A. Gepperth. 2019 · 2019
Later among the works it cites.
Competence-based curriculum learning for neural machine translation
Emmanouil Antonios Platanios, Otilia Stretcu, Graham Neubig, Barnabas Poczos, and Tom Mitchell. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
Room to Glo: A systematic comparison of semantic change detection approaches with word embeddings
Philippa Shoemark, Farhana Ferdousi Liza, Dong Nguyen, Scott Hale, and Barbara McGillivray. 2019 · 2019
Later among the works it cites.
BERT and PALs: Projected attention layers for efficient adaptation in multi-task learning
Asa Cooper Stickland and Iain Murray. 2019 · 2019
Later among the works it cites.
Multilingual neural machine translation with knowledge distillation
Xu Tan, Yi Ren, Di He, Tao Qin, and Tie-Yan Liu. 2019 · 2019
Later among the works it cites.
Unsupervised pretraining for neural machine translation using elastic weight consolidation
Dušan Variš and Ondřej Bojar. 2019 · 2019
Later among the works it cites.
Sentence embedding alignment for lifelong relation extraction
Hong Wang, Wenhan Xiong, Mo Yu, Xiaoxiao Guo, Shiyu Chang, and William Yang Wang. 2019b · 2019
Later among the works it cites.
Open-world learning and application to product classification
Hu Xu, Bing Liu, Lei Shu, and P. Yu. 2019 · 2019
Later among the works it cites.
Learning and evaluating general linguistic intelligence
Dani Yogatama, Cyprien de Masson d’Autume, Jerome Connor, Tomas Kocisky, Mike Chrzanowski, Lingpeng Kong, Angeliki Lazaridou, Wang Ling, Lei Yu, Chris Dyer, and Phil Blunsom. 2019 · 2019
Later among the works it cites.
Learning to continually learn
Shawn Beaulieu, Lapo Frati, Thomas Miconi, Joel Lehman, Kenneth O. Stanley, Jeff Clune, and Nick Cheney. 2020 · 2020
Closest in time.
Abductive commonsense reasoning
Chandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi, Ari Holtzman, Hannah Rashkin, Doug Downey, Wen tau Yih, and Yejin Choi. 2020 · 2020
Closest in time.
Uncertainty-guided continual learning with bayesian neural networks
Sayna Ebrahimi, Mohamed Elhoseiny, Trevor Darrell, and Marcus Rohrbach. 2020 · 2020
Closest in time.
Carlos Escolano, Marta R. Costa-jussà, José A. R. Fonollosa, and Mikel Artetxe. 2020 · 2020
Closest in time.
Causalm: Causal model explanation through counterfactual language models
Amir Feder, Nadav Oved, Uri Shalit, and Roi Reichart. 2020 · 2020
Closest in time.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020 · 2020
Closest in time.
Meta-learning in neural networks: A survey
Timothy Hospedales, Antreas Antoniou, Paul Micaelli, and Amos Storkey. 2020 · 2020
Closest in time.
Xtreme: A massively multilingual multi-task benchmark for evaluating cross-lingual generalization
Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020 · 2020
Closest in time.
Compositional language continual learning
Yuanpeng Li, Liang Zhao, Kenneth Church, and Mohamed Elhoseiny. 2020 · 2020
Closest in time.
Learning on the job: Online lifelong and continual learning
Bing Liu. 2020 · 2020
Closest in time.
Adapterhub: A framework for adapting transformers
Jonas Pfeiffer, Andreas Rücklé, Clifton Poth, Aishwarya Kamath, Ivan Vulić, Sebastian Ruder, Kyunghyun Cho, and Iryna Gurevych. 2020 · 2020
Closest in time.
English intermediate-task training improves zero-shot cross-lingual transfer too
Jason Phang, Iacer Calixto, Phu Mon Htut, Yada Pruksachatkun, Haokun Liu, Clara Vania, Katharina Kann, and Samuel R. Bowman. 2020 · 2020
Closest in time.
XCOPA: A multilingual dataset for causal commonsense reasoning
Edoardo Maria Ponti, Goran Glavaš, Olga Majewska, Qianchu Liu, Ivan Vulić, and Anna Korhonen. 2020 · 2020
Closest in time.
Evaluation of lifelong learning systems
Yevhenii Prokopalo, Sylvain Meignier, Olivier Galibert, Loic Barrault, and Anthony Larcher. 2020 · 2020
Closest in time.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Closest in time.
Self-induced curriculum learning in self-supervised neural machine translation
Dana Ruiter, Josef van Genabith, and Cristina España-Bonet. 2020 · 2020
Closest in time.
Toward training recurrent neural networks for lifelong learning
Shagun Sodhani, Sarath Chandar, and Yoshua Bengio. 2020 · 2020
Closest in time.
LAMOL: LAnguage MOdeling for Lifelong Language Learning
Fan-Keng Sun, Cheng-Hao Ho, and Hung-Yi Lee. 2020 · 2020
Closest in time.
Batchensemble: an alternative approach to efficient ensemble and lifelong learning
Yeming Wen, Dustin Tran, and Jimmy Ba. 2020 · 2020
Closest in time.
Overcoming catastrophic forgetting during domain adaptation of neural machine translation
Brian Thompson, Jeremy Gwinnup, Huda Khayrallah, Kevin Duh, and Philipp Koehn. 2019 · 2068
Closest in time.