Fetching the paper…
Reading the bibliography…
Training data memorization in NLP can both be beneficial (e.g., closed-book QA) and undesirable (personal data extraction).
Mathijs Mul and Willem Zuidema · 1906
Earlier work this paper cites.
Human Behavior and the Principle of Least Effort
George Zipf · 1949
Earlier work this paper cites.
On Learning the Past Tenses of English Verbs
D E Rumelhart and J McClelland · 1986
Earlier work this paper cites.
On language and connectionism: Analysis of a parallel distributed processing model of language acquisition
Steven Pinker and Alan Prince · 1988
Earlier work this paper cites.
Lexical entries and rules of language: A multidisciplinary study of German inflection
Harald Clahsen · 1999
Earlier work this paper cites.
Words and rules: The ingredients of language
Steven Pinker · 1999
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
A parallel architecture perspective on language processing
Ray Jackendoff · 2007
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning · 2015
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2015
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Earlier work this paper cites.
The case for learned index structures
Tim Kraska, Alex Beutel, Ed H Chi, Jeffrey Dean, and Neoklis Polyzotis · 2017
Earlier work this paper cites.
Membership inference attacks against machine learning models
Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov · 2017
Earlier work this paper cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Improving text-to-SQL evaluation methodology
Catherine Finegan-Dollak, Jonathan K Kummerfeld, Li Zhang, Karthik Ramanathan, Sesh Sadasivam, Rui Zhang, and Dragomir Radev · 2018
Earlier work this paper cites.
Visualisation and ‘diagnostic classifiers’ reveal how recurrent and recursive neural networks process hierarchical structure
Dieuwke Hupkes, Sara Veldhoen, and Willem Zuidema · 2018
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson · 2018
Earlier work this paper cites.
Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks
Brenden Lake and Marco Baroni · 2018
Earlier work this paper cites.
Deja vu: an empirical evaluation of the memorization properties of convnets
Alexandre Sablayrolles, Matthijs Douze, Cordelia Schmid, and Hervé Jégou · 2018
Earlier work this paper cites.
Privacy risk in machine learning: Analyzing the connection to overfitting
Samuel Yeom, Irene Giacomelli, Matt Fredrikson, and Somesh Jha · 2018
Cited alongside, same era.
What does bert look at? an analysis of bert’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D Manning · 2019
Cited alongside, same era.
Lossless data compression with transformer
Gautier Izacard, Armand Joulin, and Edouard Grave · 2019
Cited alongside, same era.
Measuring compositional generalization: A comprehensive method on realistic data
Daniel Keysers, Nathanael Schärli, Nathan Scales, Hylke Buisman, Daniel Furrer, Sergii Kashubin, Nikola Momchev, Danila Sinopalnikov, Lukasz Stafiniak, Tibor Tihon, et al · 2019
Cited alongside, same era.
Transcoding compositionally: Using attention to find more generalizable solutions
Kris Korrel, Dieuwke Hupkes, Verna Dankers, and Elia Bruni · 2019
Cited alongside, same era.
Compositionality decomposed: how do neural networks generalise?
Dieuwke Hupkes, Verna Dankers, Mathijs Mul, and Elia Bruni · 2020
Later among the works it cites.
Generalization through memorization: Nearest neighbor language models
Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis · 2020
Later among the works it cites.
What they do when in doubt: a study of inductive biases in seq2seq learners
Eugene Kharitonov and Rahma Chaabouni · 2020
Later among the works it cites.
COGS: a compositional generalization challenge based on semantic interpretation
Najoung Kim and Tal Linzen · 2020
Later among the works it cites.
Question and answer test-train overlap in open-domain question answering datasets
Patrick Lewis, Pontus Stenetorp, and Sebastian Riedel · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Cited alongside, same era.
Are sixteen heads really better than one?
Paul Michel, Omer Levy, and Graham Neubig · 2019
Cited alongside, same era.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Cited alongside, same era.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H Miller, and Sebastian Riedel · 2019
Cited alongside, same era.
Bpe-dropout: Simple and effective subword regularization
Ivan Provilkov, Dmitrii Emelianenko, and Elena Voita · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2019
Cited alongside, same era.
Rajiv Movva and Jason Y Zhao · 2020
Later among the works it cites.
How much knowledge can you pack into the parameters of a language model?
Adam Roberts, Colin Raffel, and Noam Shazeer · 2020
Later among the works it cites.
The tatoeba translation challenge–realistic data sets for low resource and multilingual mt
Jörg Tiedemann · 2020
Later among the works it cites.
OPUS-MT – building open translation services for the world
Jörg Tiedemann and Santhosh Thottingal · 2020
Later among the works it cites.
Nncp v2: Lossless data compression with transformer
Fabrice Bellard · 2021
Closest in time.
Can transformers jump around right in natural language? assessing performance transfer from scan
Rahma Chaabouni, Roberto Dessì, and Eugene Kharitonov · 2021
Closest in time.
The paradox of the compositionality of natural language: a neural machine translation case study
Verna Dankers, Elia Bruni, and Dieuwke Hupkes · 2021
Closest in time.
The flores-101 evaluation benchmark for low-resource and multilingual machine translation
Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc’Aurelio Ranzato, Francisco Guzman, and Angela Fan · 2021
Closest in time.
Efficient nearest neighbor language models
Junxian He, Graham Neubig, and Taylor Berg-Kirkpatrick · 2021
Closest in time.
Nearest neighbor machine translation
Urvashi Khandelwal, Angela Fan, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis · 2021
Closest in time.
Text-free prosody-aware generative spoken language modeling
Eugene Kharitonov, Ann Lee, Adam Polyak, Yossi Adi, Jade Copet, Kushal Lakhotia, Tu-Anh Nguyen, Morgane Rivière, Abdelrahman Mohamed, Emmanuel Dupoux, et al · 2021
Closest in time.
Generative spoken language modeling from raw audio
Kushal Lakhotia, Evgeny Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu-Anh Nguyen, Jade Copet, Alexei Baevski, Adelrahman Mohamed, and Emmanuel Dupoux · 2021
Closest in time.
PAQ: 65 million probably-asked questions and what you can do with them
Patrick Lewis, Yuxiang Wu, Linqing Liu, Pasquale Minervini, Heinrich Küttler, Aleksandra Piktus, Pontus Stenetorp, and Sebastian Riedel · 2021
Closest in time.
Effects of parameter norm growth during transformer training: Inductive bias from gradient descent, 2021
William Merrill, Vivek Ramanujan, Yoav Goldberg, Roy Schwartz, and Noah Smith · 2021
Closest in time.
Koustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau, Adina Williams, and Douwe Kiela · 2021
Closest in time.
Understanding unintended memorization in language models under federated learning
Om Dipakbhai Thakkar, Swaroop Ramaswamy, Rajiv Mathews, and Francoise Beaufays · 2021
Closest in time.