Fetching the paper…
Reading the bibliography…
Pretrained language models (PLMs) are today the primary model for natural language processing.
Roberta: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 1907
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael McCloskey and Neal J. Cohen · 1989
Earlier work this paper cites.
Connectionist models of recognition memory: constraints imposed by learning and forgetting functions
Roger Ratcliff · 1990
Earlier work this paper cites.
Inhibiting your native language: The role of retrieval-induced forgetting during second-language acquisition
Benjamin J Levy, Nathan D McVeigh, Alejandra Marful, and Michael C Anderson · 2007
Earlier work this paper cites.
Oscillatory brain activity before and after an internal context change—evidence for a reset of encoding processes
Bernhard Pastötter, Karl-Heinz Bäuml, and Simon Hanslmayr · 2008
Earlier work this paper cites.
The role of forgetting in the evolution and learning of language
Jeffrey Barrett and Kevin JS Zollman · 2009
Earlier work this paper cites.
Metalearning
Tom Schaul and Jürgen Schmidhuber · 2010
Earlier work this paper cites.
Learning to learn
Sebastian Thrun and Lorien Pratt · 2012
Earlier work this paper cites.
Powerplay: Training an increasingly general problem solver by continually searching for the simplest still unsolvable problem
Jürgen Schmidhuber · 2013
Earlier work this paper cites.
Why forget? on the adaptive value of memory loss
Simon Nørby · 2015
Earlier work this paper cites.
Learning to learn by gradient descent by gradient descent
Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W Hoffman, David Pfau, Tom Schaul, Brendan Shillingford, and Nando De Freitas · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al · 2017
Earlier work this paper cites.
Gradient episodic memory for continual learning
David Lopez-Paz and Marc’Aurelio Ranzato · 2017
Earlier work this paper cites.
Continual learning with deep generative replay
Hanul Shin, Jung Kwon Lee, Jaehong Kim, and Jiwon Kim · 2017
Earlier work this paper cites.
XNLI: Evaluating cross-lingual sentence representations
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov · 2018
Earlier work this paper cites.
Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson · 2018
Earlier work this paper cites.
Packnet: Adding multiple tasks to a single network by iterative pruning
Arun Mallya and Svetlana Lazebnik · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford and Karthik Narasimhan · 2018
Earlier work this paper cites.
Progress & compress: A scalable framework for continual learning
Jonathan Schwarz, Wojciech Czarnecki, Jelena Luketina, Agnieszka Grabska-Barwinska, Yee Whye Teh, Razvan Pascanu, and Raia Hadsell · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin · 2019
Cited alongside, same era.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli · 2019
Cited alongside, same era.
Continual lifelong learning with neural networks: A review
German I Parisi, Ronald Kemker, Jose L Part, Christopher Kanan, and Stefan Wermter · 2019
Cited alongside, same era.
Experience replay for continual learning
David Rolnick, Arun Ahuja, Jonathan Schwarz, Timothy Lillicrap, and Gregory Wayne · 2019
Cited alongside, same era.
On the cross-lingual transferability of monolingual representations
Unks everywhere: Adapting multilingual language models to new scripts
Jonas Pfeiffer, Ivan Vulić, Iryna Gurevych, and Sebastian Ruder · 2021
Later among the works it cites.
Knowledge evolution in neural networks
Ahmed Taha, Abhinav Shrivastava, and Larry S Davis · 2021
Later among the works it cites.
Ethical and social risks of harm from language models
Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, Zac Kenton, Sasha Brown, Will Hawkins, Tom Stepleton, Courtney Biles, Abeba Birhane, Julia Haas, Laura Rimell, Lisa Anne Hendricks, William Isaac, Sean Legassick, Geoffrey Irving, and Iason Gabriel · 2021
Later among the works it cites.
Masakhaner 2.0: Africa-centric transfer learning for named entity recognition
David Ifeoluwa Adelani, Graham Neubig, Sebastian Ruder, Shruti Rijhwani, Michael Beukman, Chester Palen-Michel, Constantine Lignos, Jesujoba O Alabi, Shamsuddeen H Muhammad, Peter Nabende, et al · 2022
Later among the works it cites.
Adapting pre-trained language models to African languages via multilingual adaptive fine-tuning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mikel Artetxe, Sebastian Ruder, and Dani Yogatama · 2020
Cited alongside, same era.
Shawn Beaulieu, Lapo Frati, Thomas Miconi, Joel Lehman, Kenneth O Stanley, Jeff Clune, and Nick Cheney · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Cited alongside, same era.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov · 2020
Cited alongside, same era.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A Smith · 2020
Cited alongside, same era.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Cited alongside, same era.
Mlqa: Evaluating cross-lingual extractive question answering
Patrick Lewis, Barlas Oguz, Ruty Rinott, Sebastian Riedel, and Holger Schwenk · 2020
Cited alongside, same era.
Jesujoba O. Alabi, David Ifeoluwa Adelani, Marius Mosbach, and Dietrich Klakow · 2022
Later among the works it cites.
Composable sparse fine-tuning for cross-lingual transfer
Alan Ansell, Edoardo Ponti, Anna Korhonen, and Ivan Vulić · 2022
Later among the works it cites.
Refactor gnns: Revisiting factorisation-based models from a message-passing perspective
Yihong Chen, Pushkar Mishra, Luca Franceschi, Pasquale Minervini, Pontus Stenetorp, and Sebastian Riedel · 2022
Later among the works it cites.
Americasnli: Evaluating zero-shot natural language understanding of pretrained multilingual models in truly low-resource languages
Abteen Ebrahimi, Manuel Mager, Arturo Oncevay, Vishrav Chaudhary, Luis Chiruzzo, Angela Fan, John Ortega, Ricardo Ramos, Annette Rios Gonzales, Ivan Meza-Ruiz, et al · 2022
Later among the works it cites.
Towards continual reinforcement learning: A review and perspectives
Khimya Khetarpal, Matthew Riemer, Irina Rish, and Doina Precup · 2022
Later among the works it cites.
Kelly Marchisio, Patrick Lewis, Yihong Chen, and Mikel Artetxe · 2022
Later among the works it cites.
Memory-based model editing at scale
Eric Mitchell, Charles Lin, Antoine Bosselut, Christopher D Manning, and Chelsea Finn · 2022
Later among the works it cites.
The primacy bias in deep reinforcement learning
Evgenii Nikishin, Max Schwarzer, Pierluca D’Oro, Pierre-Luc Bacon, and Aaron Courville · 2022
Later among the works it cites.
Lifting the curse of multilinguality by pre-training modular transformers
Jonas Pfeiffer, Naman Goyal, Xi Lin, Xian Li, James Cross, Sebastian Riedel, and Mikel Artetxe · 2022
Later among the works it cites.
Taxonomy of risks posed by language models
Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, et al · 2022
Later among the works it cites.
Fortuitous forgetting in connectionist networks
Hattie Zhou, Ankit Vani, Hugo Larochelle, and Aaron Courville · 2022
Later among the works it cites.
Sample-efficient reinforcement learning by breaking the replay ratio barrier
Pierluca D’Oro, Max Schwarzer, Evgenii Nikishin, Pierre-Luc Bacon, Marc G Bellemare, and Aaron Courville · 2023
Closest in time.
Building a subspace of policies for scalable continual learning
Jean-Baptiste Gaya, Thang Doan, Lucas Caccia, Laure Soulier, Ludovic Denoyer, and Roberta Raileanu · 2023
Closest in time.
Delving deeper into cross-lingual visual question answering
Chen Liu, Jonas Pfeiffer, Anna Korhonen, Ivan Vulić, and Iryna Gurevych · 2023
Closest in time.
Understanding plasticity in neural networks
Clare Lyle, Zeyu Zheng, Evgenii Nikishin, Bernardo Avila Pires, Razvan Pascanu, and Will Dabney · 2023
Closest in time.
Deep reinforcement learning with plasticity injection
Evgenii Nikishin, Junhyuk Oh, Georg Ostrovski, Clare Lyle, Razvan Pascanu, Will Dabney, and Andre Barreto · 2023
Closest in time.
Learn, unlearn and relearn: An online learning paradigm for deep neural networks
Vijaya Raghavan T Ramkumar, Elahe Arani, and Bahram Zonooz · 2023
Closest in time.
Llama: Open and efficient foundation language models, 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Closest in time.