Fetching the paper…
Reading the bibliography…
To produce accurate predictions, language models (LMs) must balance between generalization and memorization.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Bert is not a knowledge base (yet): Factual knowledge vs. name-based reasoning in unsupervised qa
Nina Poerner, Ulli Waltinger, and Hinrich Schütze. 2019 · 1911
Earlier work this paper cites.
GloVe: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014 · 2014
Earlier work this paper cites.
Idiom token classification using sentential distributed semantics
Giancarlo Salton, Robert Ross, and John Kelleher. 2016 · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017 · 2017
Earlier work this paper cites.
LIdioms: A multilingual linked idioms data set
Diego Moussallem, Mohamed Ahmed Sherif, Diego Esteves, Marcos Zampieri, and Axel-Cyrille Ngonga Ngomo. 2018 · 2018
Earlier work this paper cites.
The secret sharer: Evaluating and testing unintended memorization in neural networks
Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos, and Dawn Song. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Earlier work this paper cites.
Still a pain in the neck: Evaluating text representations on lexical composition
Vered Shwartz and Ido Dagan. 2019 · 2019
Earlier work this paper cites.
Auditing data provenance in text-generation models
Congzheng Song and Vitaly Shmatikov. 2019 · 2019
Earlier work this paper cites.
Electra: Pre-training text encoders as discriminators rather than generators
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020 · 2020
Earlier work this paper cites.
Does learning require memorization? a short tale about a long tail
Vitaly Feldman. 2020 · 2020
Earlier work this paper cites.
What neural networks memorize and why: Discovering the long tail via influence estimation
Vitaly Feldman and Chiyuan Zhang. 2020 · 2020
Earlier work this paper cites.
MAGPIE: A large corpus of potentially idiomatic expressions
Hessel Haagsma, Johan Bos, and Malvina Nissim. 2020 · 2020
Cited alongside, same era.
How can we know what language models know?
Zhengbao Jiang, Frank F Xu, Jun Araki, and Graham Neubig. 2020 · 2020
Cited alongside, same era.
Text-to-text pre-training for data-to-text tasks
Mihir Kale and Abhinav Rastogi. 2020 · 2020
Cited alongside, same era.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020 · 2020
Cited alongside, same era.
E-BERT: Efficient-yet-effective entity embeddings for BERT
Nina Poerner, Ulli Waltinger, and Hinrich Schütze. 2020 · 2020
Cited alongside, same era.
Epie dataset: A corpus for possible idiomatic expressions
Prateek Saxena and Soma Paul. 2020 · 2020
The curious case of hallucinations in neural machine translation
Vikas Raunak, Arul Menezes, and Marcin Junczys-Dowmunt. 2021 · 2021
Later among the works it cites.
Quantifying memorization across neural language models
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang. 2022 · 2022
Closest in time.
It’s not rocket science: Interpreting figurative language in narratives
Tuhin Chakrabarty, Yejin Choi, and Vered Shwartz. 2022 · 2022
Closest in time.
Knowledge neurons in pretrained transformers
Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei. 2022 · 2022
Closest in time.
Can transformer be too compositional? analysing idiom processing in neural machine translation
Verna Dankers, Christopher Lucas, and Ivan Titov. 2022 · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
BERTnesia: Investigating the capture and forgetting of knowledge in BERT
Jonas Wallat, Jaspreet Singh, and Avishek Anand. 2020 · 2020
Cited alongside, same era.
When is memorization of irrelevant training data necessary for high-accuracy learning?
Gavin Brown, Mark Bun, Vitaly Feldman, Adam Smith, and Kunal Talwar. 2021 · 2021
Cited alongside, same era.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al. 2021 · 2021
Cited alongside, same era.
Measuring and improving consistency in pretrained language models
Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard Hovy, Hinrich Schütze, and Yoav Goldberg. 2021 · 2021
Cited alongside, same era.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. 2021 · 2021
Cited alongside, same era.
Contextualized embeddings encode monolingual and cross-lingual knowledge of idiomaticity
Samin Fakharian and Paul Cook. 2021 · 2021
Cited alongside, same era.
Mor Geva, Avi Caciularu, Kevin Ro Wang, and Yoav Goldberg. 2022 · 2022
Closest in time.
Deduplicating training data makes language models better
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. 2022 · 2022
Closest in time.
Data contamination: From memorization to exploitation
Inbal Magar and Roy Schwartz. 2022 · 2022
Closest in time.
Locating and editing factual knowledge in gpt
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022 · 2022
Closest in time.
How to dissect a Muppet: The structure of transformer embedding spaces
Timothee Mickus, Denis Paperno, and Mathieu Constant. 2022 · 2022
Closest in time.
Finding memo: Extractive memorization in constrained sequence generation tasks
Vikas Raunak and Arul Menezes. 2022 · 2022
Closest in time.
What makes reading comprehension questions difficult?
Saku Sugawara, Nikita Nangia, Alex Warstadt, and Samuel Bowman. 2022 · 2022
Closest in time.
Memorisation versus generalisation in pre-trained language models
Michael Tänzer, Sebastian Ruder, and Marek Rei. 2022 · 2022
Closest in time.
Memorization without overfitting: Analyzing the training dynamics of large language models
Kushal Tirumala, Aram H Markosyan, Luke Zettlemoyer, and Armen Aghajanyan. 2022 · 2022
Closest in time.