Fetching the paper…
Reading the bibliography…
Pretraining on a large-scale corpus has become a standard method to build general language models (LMs).
Catastrophic forgetting, rehearsal and pseudorehearsal
Robins, A · 1995
Earlier work this paper cites.
Towards a human-like open-domain chatbot
Adiwardana, D., Luong, M., So, D. R., Hall, J., Fiedel, N., Thoppilan, R., Yang, Z., Kulshreshtha, A., Nemade, G., Lu, Y., and Le, Q. V · 2001
Earlier work this paper cites.
Recurrent neural network based language model
Mikolov, T., Karafiát, M., Burget, L., Cernocký, J. H., and Khudanpur, S · 2010
Earlier work this paper cites.
Generating text with recurrent neural networks
Sutskever, I., Martens, J., and Hinton, G · 2011
Earlier work this paper cites.
SemEval-2012 task 7: Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Gordon, A., Kozareva, Z., and Roemmele, M · 2012
Earlier work this paper cites.
Online incremental feature learning with denoising autoencoders
Zhou, G., Sohn, K., and Lee, H · 2012
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Mikolov, T., Chen, K., Corrado, G., and Dean, J · 2013
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., and Le, Q. V · 2014
Earlier work this paper cites.
Net2net: Accelerating learning via knowledge transfer
Chen, T., Goodfellow, I., and Shlens, J · 2015
Earlier work this paper cites.
Semi-supervised sequence learning
Dai, A. M. and Le, Q. V · 2015
Earlier work this paper cites.
Skip-thought vectors
Kiros, R., Zhu, Y., Salakhutdinov, R. R., Zemel, R., Urtasun, R., Torralba, A., and Fidler, S · 2015
Earlier work this paper cites.
Findings of the 2016 conference on machine translation
Bojar, O., Chatterjee, R., Federmann, C., Graham, Y., Haddow, B., Huck, M., Yepes, A. J., Koehn, P., Logacheva, V., Monz, C., et al · 2016
Earlier work this paper cites.
Rusu, A. A., Rabinowitz, N. C., Desjardins, G., Soyer, H., Kirkpatrick, J., Kavukcuoglu, K., Pascanu, R., and Hadsell, R · 2016
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
Hestness, J., Narang, S., Ardalani, N., Diamos, G. F., Jun, H., Kianinejad, H., Patwary, M. M. A., Yang, Y., and Zhou, Y · 2017
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
Joshi, M., Choi, E., Weld, D., and Zettlemoyer, L · 2017
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
Kirkpatrick, J., Pascanu, R., Rabinowitz, N. C., Veness, J., Desjardins, G., Rusu, A. A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., Hassabis, D., Clopath, C., Kumaran, D., and Hadsell, R · 2017
Earlier work this paper cites.
Learning without forgetting
Li, Z. and Hoiem, D · 2017
Earlier work this paper cites.
Gradient episodic memory for continual learning
Lopez-Paz, D. and Ranzato, M · 2017
Earlier work this paper cites.
icarl: Incremental classifier and representation learning
Rebuffi, S., Kolesnikov, A., Sperl, G., and Lampert, C. H · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q. V., Hinton, G. E., and Dean, J · 2017
Earlier work this paper cites.
Continual learning with deep generative replay
Shin, H., Lee, J. K., Kim, J., and Kim, J · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Cited alongside, same era.
Efficient lifelong learning with a-gem
Chaudhry, A., Ranzato, M., Rohrbach, M., and Elhoseiny, M · 2018
Cited alongside, same era.
Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Kudo, T. and Richardson, J · 2018
Cited alongside, same era.
Learning without forgetting
Li, Z. and Hoiem, D · 2018
Cited alongside, same era.
Packnet: Adding multiple tasks to a single network by iterative pruning
Mallya, A. and Lazebnik, S · 2018
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2018
Xlnet: Generalized autoregressive pretraining for language understanding
Yang, Z., Dai, Z., Yang, Y., Carbonell, J., Salakhutdinov, R. R., and Le, Q. V · 2019
Later among the works it cites.
Continual lifelong learning in natural language processing: A survey
Biesialska, M., Biesialska, K., and Costa-jussà, M. R · 2020
Later among the works it cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Later among the works it cites.
Electra: Pre-training text encoders as discriminators rather than generators
Clark, K., Luong, M.-T., Le, Q. V., and Manning, C. D · 2020
Later among the works it cites.
Meta-learning with sparse experience replay for lifelong language learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Adafactor: Adaptive learning rates with sublinear memory cost
Shazeer, N. and Stern, M · 2018
Cited alongside, same era.
Mesh-tensorflow: Deep learning for supercomputers
Shazeer, N., Cheng, Y., Parmar, N., Tran, D., Vaswani, A., Koanantakool, P., Hawkins, P., Lee, H., Hong, M., Young, C., Sepassi, R., and Hechtman, B · 2018
Cited alongside, same era.
Lifelong learning with dynamically expandable networks
Yoon, J., Yang, E., Lee, J., and Hwang, S. J · 2018
Cited alongside, same era.
Record: Bridging the gap between human and machine commonsense reading comprehension
Zhang, S., Liu, X., Liu, J., Gao, J., Duh, K., and Durme, B. V · 2018
Cited alongside, same era.
Online continual learning with maximally interfered retrieval
Aljundi, R., Caccia, L., Belilovsky, E., Caccia, M., Lin, M., Charlin, L., and Tuytelaars, T · 2019
Cited alongside, same era.
Episodic memory in lifelong language learning
d’Autume, C. d. M., Ruder, S., Kong, L., and Yogatama, D · 2019
Cited alongside, same era.
Holla, N., Mishra, P., Yannakoudakis, H., and Shutova, E · 2020
Later among the works it cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Later among the works it cites.
LAMOL: language modeling for lifelong language learning
Sun, F., Ho, C., and Lee, H · 2020
Later among the works it cites.
LAMOL: language modeling for lifelong language learning
Sun, F., Ho, C., and Lee, H · 2020
Later among the works it cites.
Batchensemble: an alternative approach to efficient ensemble and lifelong learning
Wen, Y., Tran, D., and Ba, J · 2020
Later among the works it cites.
Drill: Dynamic representations for imbalanced lifelong learning
Ahrens, K., Abawi, F., and Wermter, S · 2021
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B., and Shazeer, N · 2021
Later among the works it cites.
Demix layers: Disentangling domains for modular language modeling
Gururangan, S., Lewis, M., Holtzman, A., Smith, N. A., and Zettlemoyer, L · 2021
Later among the works it cites.
Continual learning for text classification with information disentanglement based regularization
Huang, Y., Zhang, Y., Chen, J., Wang, X., and Yang, D · 2021
Later among the works it cites.
Towards a robust experimental framework and benchmark for lifelong language learning
Hussain, A., Holla, N., Mishra, P., Yannakoudakis, H., and Shutova, E · 2021
Later among the works it cites.
Learn continually, generalize rapidly: Lifelong knowledge accumulation for few-shot learning
Jin, X., Lin, B. Y., Rostami, M., and Ren, X · 2021
Later among the works it cites.
Beyond distillation: Task-level mixture-of-experts for efficient inference
Kudugunta, S., Huang, Y., Bapna, A., Krikun, M., Lepikhin, D., Luong, M.-T., and Firat, O · 2021
Later among the works it cites.
GShard: Scaling giant models with conditional computation and automatic sharding
Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z · 2021
Later among the works it cites.
Glam: Efficient scaling of language models with mixture-of-experts
Du, N., Huang, Y., Dai, A. M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A. W., Firat, O., et al · 2022
Later among the works it cites.
On continual model refinement in out-of-distribution data streams
Lin, B. Y., Wang, S., Lin, X. V., Jia, R., Xiao, L., Ren, X., and Yih, W.-t · 2022
Later among the works it cites.