Fetching the paper…
Reading the bibliography…
Despite their wide adoption, the underlying training and memorization dynamics of very large language models is not well understood.
Does Learning Require Memorization? A Short Tale about a Long Tail
Vitaly Feldman · 1906
Earlier work this paper cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 1907
Earlier work this paper cites.
Evaluating forgetting curves
Geoffrey R Loftus · 1985
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael McCloskey and Neal J Cohen · 1989
Earlier work this paper cites.
Connectionist models of recognition memory: Constraints imposed by learning and forgetting functions
Roger Ratcliff · 1990
Earlier work this paper cites.
Scaling Laws for Neural Language Models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2001
Earlier work this paper cites.
Why verbs are harder to learn than nouns: Initial insights from a computational model of intention recognition in situated word learning
Michael Fleischman and Deb Roy · 2005
Earlier work this paper cites.
The elements of statistical learning: data mining, inference and prediction
James Franklin · 2005
Earlier work this paper cites.
Calibrating noise to sensitivity in private data analysis
Cynthia Dwork, Frank McSherry, Kobbi Nissim, and Adam Smith · 2006
Earlier work this paper cites.
Expanding retrieval practice promotes short-term retention, but equally spaced retrieval enhances long-term retention
Jeffrey D Karpicke and Henry L Roediger III · 2007
Earlier work this paper cites.
The eu proposal for a general data protection regulation and the roots of the ‘right to be forgotten’
Alessandro Mantelero · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Effects of spaced retrieval training on semantic memory in alzheimer’s disease: A systematic review
Shiri Oren, Charlene Willerton, and Jeff Small · 2014
Earlier work this paper cites.
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
The right time to learn: mechanisms and optimization of spaced learning
Paul Smolen, Yili Zhang, and John H Byrne · 2016
Earlier work this paper cites.
Repeat before forgetting: Spaced repetition for efficient and effective training of neural networks
Hadi Amiri, Timothy Miller, and Guergana Savova · 2017
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al · 2017
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2017
Earlier work this paper cites.
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, et al · 2017
Earlier work this paper cites.
Svcca: Singular vector canonical correlation analysis for deep learning dynamics and interpretability
Maithra Raghu, Justin Gilmer, Jason Yosinski, and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
The eu general data protection regulation (gdpr)
Paul Voigt and Axel Von dem Bussche · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Earlier work this paper cites.
Critical learning periods in deep networks
Alessandro Achille, Matteo Rovere, and Stefano Soatto · 2018
Earlier work this paper cites.
Lifelong machine learning
Zhiyuan Chen and Bing Liu · 2018
Earlier work this paper cites.
Gradient descent happens in a tiny subspace
Guy Gur-Ari, Daniel A Roberts, and Ethan Dyer · 2018
Earlier work this paper cites.
Insights on representational similarity in neural networks with canonical correlation
Ari Morcos, Maithra Raghu, and Samy Bengio · 2018
Earlier work this paper cites.
Leveraging random label memorization for unsupervised pre-training
Vinaychandran Pondenkandath, Michele Alberti, Sammer Puran, Rolf Ingold, and Marcus Liwicki · 2018
Earlier work this paper cites.
General data protection regulation (gdpr)
General Data Protection Regulation · 2018
Earlier work this paper cites.
Understanding learning dynamics of language models with svcca
Naomi Saphra and Adam Lopez · 2018
Earlier work this paper cites.
An empirical study of example forgetting during deep neural network learning
Mariya Toneva, Alessandro Sordoni, Remi Tachet des Combes, Adam Trischler, Yoshua Bengio, and Geoffrey J Gordon · 2018
Earlier work this paper cites.
The secret sharer: Evaluating and testing unintended memorization in neural networks
Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos, and Dawn Song · 2019
Earlier work this paper cites.
Does learning require memorization
Vitaly Feldman · 2019
Earlier work this paper cites.
Understanding the scope and impact of the california consumer privacy act of 2018
Elizabeth Liz Harding, Jarno J Vanto, Reece Clark, L Hannah Ji, and Sara C Ainsworth · 2019
Cited alongside, same era.
Membership inference attacks on sequence-to-sequence models
Sorami Hisamoto, Matt Post, and Kevin Duh · 2019
Cited alongside, same era.
Generalization through memorization: Nearest neighbor language models
Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis · 2019
Cited alongside, same era.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Training dynamics for text summarization models
Tanya Goyal, Jiacheng Xu, Junyi Jessy Li, and Greg Durrett · 2021
Later among the works it cites.
Danny Hernandez, Jared Kaplan, Tom Henighan, and Sam McCandlish · 2021
Later among the works it cites.
How bpe affects memorization in transformers
Eugene Kharitonov, Marco Baroni, and Dieuwke Hupkes · 2021
Later among the works it cites.
Deduplicating training data makes language models better
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini · 2021
Later among the works it cites.
Large language models can be strong differentially private learners
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H Miller, and Sebastian Riedel · 2019
Cited alongside, same era.
A constructive prediction of the generalization error across scales
Jonathan S Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, and Nir Shavit · 2019
Cited alongside, same era.
Auditing data provenance in text-generation models
Congzheng Song and Vitaly Shmatikov · 2019
Cited alongside, same era.
Better fine-tuning by reducing representational collapse
Armen Aghajanyan, Akshat Shrivastava, Anchit Gupta, Naman Goyal, Luke Zettlemoyer, and Sonal Gupta · 2020
Cited alongside, same era.
Aim, 6 2020
Gor Arakelyan, Gevorg Soghomonyan, and The Aim team · 2020
Cited alongside, same era.
Recall and learn: Fine-tuning deep pretrained language models with less forgetting
Sanyuan Chen, Yutai Hou, Yiming Cui, Wanxiang Che, Ting Liu, and Xiangzhan Yu · 2020
Cited alongside, same era.
Pretrained language model embryology: The birth of albert
Cheng-Han Chiang, Sung-Feng Huang, and Hung-yi Lee · 2020
Cited alongside, same era.
Xuechen Li, Florian Tramer, Percy Liang, and Tatsunori Hashimoto · 2021
Later among the works it cites.
Probing across time: What does roberta know and when?
Leo Z Liu, Yizhong Wang, Jungo Kasai, Hannaneh Hajishirzi, and Noah A Smith · 2021
Later among the works it cites.
On robustness of generative representations against catastrophic forgetting
Wojciech Masarczyk, Kamil Deja, and Tomasz Trzcinski · 2021
Later among the works it cites.
R Thomas McCoy, Paul Smolensky, Tal Linzen, Jianfeng Gao, and Asli Celikyilmaz · 2021
Later among the works it cites.
Wide neural networks forget less catastrophically
Seyed Iman Mirzadeh, Arslan Chaudhry, Huiyi Hu, Razvan Pascanu, Dilan Gorur, and Mehrdad Farajtabar · 2021
Later among the works it cites.
Deep double descent: Where bigger models and more data hurt
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever · 2021
Later among the works it cites.
Scaling language models: Methods, analysis & insights from training gopher
Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al · 2021
Later among the works it cites.
Effect of scale on catastrophic forgetting in neural networks
Vinay Venkatesh Ramasesh, Aitor Lewkowycz, and Ethan Dyer · 2021
Later among the works it cites.
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
Later among the works it cites.
Observing the learning curve of nmt systems with regard to linguistic phenomena
Patrick Stadler, Vivien Macketanz, and Eleftherios Avramidis · 2021
Later among the works it cites.
Analyzing the source and target contributions to predictions in neural machine translation
Elena Voita, Rico Sennrich, and Ivan Titov · 2021
Later among the works it cites.
Counterfactual Memorization in Neural Language Models
Chiyuan Zhang, Daphne Ippolito, Katherine Lee, Matthew Jagielski, Florian Tramèr, and Nicholas Carlini · 2021
Later among the works it cites.
Cm3: A causal masked multimodal model of the internet
Armen Aghajanyan, Bernie Huang, Candace Ross, Vladimir Karpukhin, Hu Xu, Naman Goyal, Dmytro Okhonko, Mandar Joshi, Gargi Ghosh, Mike Lewis, et al · 2022
Closest in time.
A review on language models as knowledge bases
Badr AlKhamissi, Millicent Li, Asli Celikyilmaz, Mona Diab, and Marjan Ghazvininejad · 2022
Closest in time.
Analyzing the mono-and cross-lingual pretraining dynamics of multilingual language models
Terra Blevins, Hila Gonen, and Luke Zettlemoyer · 2022
Closest in time.
Quantifying memorization across neural language models
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang · 2022
Closest in time.
Word acquisition in neural language models
Tyler A. Chang and Benjamin K. Bergen · 2022
Closest in time.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Closest in time.
Unified scaling laws for routed language models
Aidan Clark, Diego de las Casas, Aurelia Guy, Arthur Mensch, Michela Paganini, Jordan Hoffmann, Bogdan Damoc, Blake Hechtman, Trevor Cai, Sebastian Borgeaud, et al · 2022
Closest in time.
spaCy 3: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing
Matthew Honnibal and Ines Montani · 2022
Closest in time.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Yejin Bang, Andrea Madotto, and Pascale Fung · 2022
Closest in time.
Quantifying privacy risks of masked language models using membership inference attacks
Fatemehsadat Mireshghallah, Kartik Goyal, Archit Uniyal, Taylor Berg-Kirkpatrick, and Reza Shokri · 2022
Closest in time.
On the dynamics of gender learning in speech translation
Beatrice Savoldi, Marco Gaido, Luisa Bentivogli, Matteo Negri, and Marco Turchi · 2022
Closest in time.
Chenze Shao and Yang Feng · 2022
Closest in time.
Shaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, et al · 2022
Closest in time.
Transformer memory as a differentiable search index
Yi Tay, Vinh Q Tran, Mostafa Dehghani, Jianmo Ni, Dara Bahri, Harsh Mehta, Zhen Qin, Kai Hui, Zhe Zhao, Jai Gupta, et al · 2022
Closest in time.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou · 2022
Closest in time.
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al · 2022
Closest in time.