Fetching the paper…
Reading the bibliography…
When training a neural network, it will quickly memorise some source-target mappings from your dataset but never learn some others.
Overfitting and undercomputing in machine learning
Tom Dietterich. 1995 · 1995
Earlier work this paper cites.
A lightweight evaluation framework for machine translation reordering
David Talbot, Hideto Kazawa, Hiroshi Ichikawa, Jason Katz-Brown, Masakazu Seno, and Franz J Och. 2011 · 2011
Earlier work this paper cites.
Efficient word alignment with Markov Chain Monte Carlo
Robert Östling and Jörg Tiedemann. 2016 · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals. 2017 · 2017
Earlier work this paper cites.
Hallucinations in neural machine translation
Katherine Lee, Orhan Firat, Ashish Agarwal, Clara Fannjiang, and David Sussillo. 2018 · 2018
Earlier work this paper cites.
The secret sharer: Evaluating and testing unintended memorization in neural networks
Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos, and Dawn Song. 2019 · 2019
Earlier work this paper cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Earlier work this paper cites.
An empirical study of example forgetting during deep neural network learning
Mariya Toneva, Alessandro Sordoni, Remi Tachet des Combes, Adam Trischler, Yoshua Bengio, and Geoffrey J. Gordon. 2019 · 2019
Earlier work this paper cites.
Benign overfitting in linear regression
Peter L Bartlett, Philip M Long, Gábor Lugosi, and Alexander Tsigler. 2020 · 2020
Earlier work this paper cites.
Does learning require memorization? A short tale about a long tail
Vitaly Feldman. 2020 · 2020
Earlier work this paper cites.
What neural networks memorize and why: Discovering the long tail via influence estimation
Vitaly Feldman and Chiyuan Zhang. 2020 · 2020
Cited alongside, same era.
Dataset cartography: Mapping and diagnosing datasets with training dynamics
Swabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A Smith, and Yejin Choi. 2020 · 2020
Cited alongside, same era.
OPUS-MT – building open translation services for the world
Jörg Tiedemann and Santhosh Thottingal. 2020 · 2020
Cited alongside, same era.
Improving massively multilingual neural machine translation and zero-shot translation
Biao Zhang, Philip Williams, Ivan Titov, and Rico Sennrich. 2020 · 2020
Cited alongside, same era.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al. 2021 · 2021
Cited alongside, same era.
Quantifying memorization across neural language models
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang. 2022 · 2022
Later among the works it cites.
No language left behind: Scaling human-centered machine translation
Marta R Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, et al. 2022 · 2022
Later among the works it cites.
Can transformer be too compositional? Analysing idiom processing in neural machine translation
Verna Dankers, Christopher Lucas, and Ivan Titov. 2022 · 2022
Later among the works it cites.
The flores-101 evaluation benchmark for low-resource and multilingual machine translation
Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc’Aurelio Ranzato, Francisco Guzman, and Angela Fan. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Daniel D’souza, Zach Nussbaum, Chirag Agarwal, and Sara Hooker. 2021 · 2021
Cited alongside, same era.
How BPE affects memorization in transformers
Eugene Kharitonov, Marco Baroni, and Dieuwke Hupkes. 2021 · 2021
Cited alongside, same era.
The curious case of hallucinations in neural machine translation
Vikas Raunak, Arul Menezes, and Marcin Junczys-Dowmunt. 2021 · 2021
Cited alongside, same era.
How benign is benign overfitting?
Amartya Sanyal, Puneet K. Dokania, Varun Kanade, and Philip H. S. Torr. 2021 · 2021
Cited alongside, same era.
Language modeling, lexical translation, reordering: The training process of NMT through the lens of classical SMT
Elena Voita, Rico Sennrich, and Ivan Titov. 2021 · 2021
Cited alongside, same era.
Counterfactual memorization in neural language models
Chiyuan Zhang, Daphne Ippolito, Katherine Lee, Matthew Jagielski, Florian Tramèr, and Nicholas Carlini. 2021 · 2021
Cited alongside, same era.
Measures of information reflect memorization patterns
Rachit Bansal, Danish Pruthi, and Yonatan Belinkov. 2022 · 2022
Cited alongside, same era.
Fatemehsadat Mireshghallah, Archit Uniyal, Tianhao Wang, David Evans, and Taylor Berg-Kirkpatrick. 2022 · 2022
Later among the works it cites.
Finding memo: Extractive memorization in constrained sequence generation tasks
Vikas Raunak and Arul Menezes. 2022 · 2022
Later among the works it cites.
Memorization without overfitting: Analyzing the training dynamics of large language models
Kushal Tirumala, Aram H Markosyan, Luke Zettlemoyer, and Armen Aghajanyan. 2022 · 2022
Later among the works it cites.
An empirical study of memorization in NLP
Xiaosen Zheng and Jing Jiang. 2022 · 2022
Later among the works it cites.
Speak, memory: An archaeology of books known to chatgpt/gpt-4
Kent K Chang, Mackenzie Cramer, Sandeep Soni, and David Bamman. 2023 · 2023
Closest in time.
Nuno M Guerreiro, Elena Voita, and André FT Martins. 2023 · 2023
Closest in time.
Understanding transformer memorization recall through idioms
Adi Haviv, Ido Cohen, Jacob Gidron, Roei Schuster, Yoav Goldberg, and Mor Geva. 2023 · 2023
Closest in time.