Fetching the paper…
Reading the bibliography…
We introduce a lifelong language learning setup where a model needs to learn from a stream of text examples without any dataset identifier.
Catastrophic interference in connectionist networks: The sequential learning problem
McCloskey, M. and Cohen, N. J · 1989
Earlier work this paper cites.
Connectionist models of recognition memory: constraints imposed by learning and forgetting functions
Ratcliff, R · 1990
Earlier work this paper cites.
Memory–a century of consolidation
McGaugh, J. L · 2000
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Adam: a method for stochastic optimization
Kingma, D. P. and Ba, J. L · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Zhang, X., Zhao, J., and LeCun, Y · 2015
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Joshi, M., Choi, E., Weld, D., and Zettlemoyer, L · 2017
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
Kirkpatrick, J., Pascanu, R., Rabinowitz, N., Veness, J., Desjardins, G., Rusu, A. A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., Hassabis, D., Clopath, C., Kumaran, D., and Hadsell, R · 2017
Earlier work this paper cites.
Gradient episodic memory for continuum learning
Lopez-Paz, D. and Ranzato, M · 2017
Earlier work this paper cites.
An overview of multi-task learning in deep neural networks
Ruder, S · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Cited alongside, same era.
Continual learning through synaptic intelligence
Zenke, F., Poole, B., and Ganguli, S · 2017
Cited alongside, same era.
Riemannian walk for incremental learning: Understanding forgetting and intransigence
Chaudhry, A., Dokania, P. K., Ajanthan, T., and Torr, P. H · 2018
Cited alongside, same era.
QuAC: Question answering in context
Choi, E., He, H., Iyyer, M., Yatskar, M., tau Yih, W., Choi, Y., Liang, P., and Zettlemoyer, L · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Cited alongside, same era.
Differentiable plasticity: training plastic neural networks with backpropagation
Miconi, T., Clune, J., and Stanley, K. O · 2018
Later among the works it cites.
Deep contextualized word representations
Peters, M. E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L · 2018
Later among the works it cites.
Progress and compress: A scalable framework for continual learning
Schwarz, J., Luketina, J., Czarnecki, W. M., Grabska-Barwinska, A., Teh, Y. W., Pascanu, R., and Hadsell, R · 2018
Later among the works it cites.
Memory-based parameter adaptation
Sprechmann, P., Jayakumar, S. M., Rae, J. W., Pritzel, A., Uria, B., and Vinyals, O · 2018
Later among the works it cites.
Efficient lifelong learning with A-GEM
Chaudhry, A., Ranzato, M., Rohrbach, M., and Elhoseiny, M · 2019
Closest in time.
Adaptive posterior learning: few-shot learning with a surprise-based memory module
Ramalho, T. and Garnelo, M · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Universal language model fine-tuning for text classification
Howard, J. and Ruder, S · 2018
Cited alongside, same era.
Constituency parsing with a self-attentive encoder
Kitaev, N. and Klein, D · 2018
Cited alongside, same era.
Higher-order coreference resolution with coarse-to-fine inference
Lee, K., He, L., and Zettlemoyer, L · 2018
Cited alongside, same era.
The natural language decathlon: Multitask learning as question answering
McCann, B., Keskar, N. S., Xiong, C., and Socher, R · 2018
Cited alongside, same era.
Closest in time.
An empirical study of example forgetting during deep neural network learning
Toneva, M., Sordoni, A., des Combes, R. T., Trischler, A., Bengio, Y., and Gordon, G. J · 2019
Closest in time.
Sentence embedding alignment for lifelong relation extraction
Wang, H., Xiong, W., Yu, M., Guo, X., Chang, S., and Wang, W. Y · 2019
Closest in time.
Learning and evaluating general linguistic intelligence
Yogatama, D., de Masson d’Autume, C., Connor, J., Kocisky, T., Chrzanowski, M., Kong, L., Lazaridou, A., Ling, W., Yu, L., Dyer, C., and Blunsom, P · 2019
Closest in time.