Fetching the paper…
Reading the bibliography…
The impressive capabilities of large language models (LLMs) have sparked debate over whether these models genuinely generalize to unseen tasks or predominantly rely on memorizing vast amounts of pretraining data.
Europarl: A parallel corpus for statistical machine translation
P. Koehn · 2005
Earlier work this paper cites.
Proceedings of the Fourth Workshop on Statistical Machine Translation , Athens, Greece, Mar. 2009. Association for Computational Linguistics
C. Callison-Burch, P. Koehn, C. Monz, and J. Schroeder, editors · 2009
Earlier work this paper cites.
Enriching phrase tables for statistical machine translation using mixed embeddings
P. Passban, Q. Liu, and A. Way · 2016
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
M. Joshi, E. Choi, D. Weld, and L. Zettlemoyer · 2017
Earlier work this paper cites.
Learning joint multilingual sentence representations with neural machine translation
H. Schwenk and M. Douze · 2017
Earlier work this paper cites.
How large are lions? inducing distributions over quantitative attributes
Y. Elazar, A. Mahabal, D. Ramachandran, T. Bedrax-Weiss, and D. Roth · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. B. Brown · 2020
Earlier work this paper cites.
Does learning require memorization? a short tale about a long tail
V. Feldman · 2020
Earlier work this paper cites.
What neural networks memorize and why: Discovering the long tail via influence estimation
V. Feldman and C. Zhang · 2020
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima, et al · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2020
Earlier work this paper cites.
Estimating training data influence by tracing gradient descent
G. Pruthi, F. Liu, S. Kale, and M. Sundararajan · 2020
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
T. Zhang*, V. Kishore*, F. Wu*, K. Q. Weinberger, and Y. Artzi · 2020
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?
E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell · 2021
Cited alongside, same era.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al · 2021
Cited alongside, same era.
Quantifying memorization across neural language models
N. Carlini, D. Ippolito, M. Jagielski, K. Lee, F. Tramer, and C. Zhang · 2022
Cited alongside, same era.
Data distributional properties drive emergent in-context learning in transformers
S. Chan, A. Santoro, A. Lampinen, J. Wang, A. Singh, P. Richemond, J. McClelland, and F. Hill · 2022
Cited alongside, same era.
Measuring causal effects of data statistics on language model’sfactual’predictions
Y. Elazar, N. Kassner, S. Ravfogel, A. Feder, A. Ravichander, M. Mosbach, Y. Belinkov, H. Schütze, and Y. Goldberg · 2022
Backtracking mathematical reasoning of language models to the pretraining data
Y. Razeghi, H. Ivison, S. Singh, and Y. Elazar · 2023
Later among the works it cites.
Large language models are latent variable models: Explaining and finding good demonstrations for in-context learning
X. Wang, W. Zhu, M. Saxon, M. Steyvers, and W. Y. Wang · 2023
Later among the works it cites.
Counterfactual memorization in neural language models
C. Zhang, D. Ippolito, K. Lee, M. Jagielski, F. Tramèr, and N. Carlini · 2023
Later among the works it cites.
Parallel structures in pre-training data yield in-context learning
Y. Chen, C. Zhao, Z. Yu, K. McKeown, and H. He · 2024
Closest in time.
What’s in my big data?, 2024
Y. Elazar, A. Bhagia, I. Magnusson, A. Ravichander, D. Schwenk, A. Suhr, P. Walsh, D. Groeneveld, L. Soldaini, S. Singh, H. Hajishirzi, N. A. Smith, and J. Dodge · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Data contamination: From memorization to exploitation
I. Magar and R. Schwartz · 2022
Cited alongside, same era.
Text embeddings by weakly-supervised contrastive pre-training
L. Wang, N. Yang, X. Huang, B. Jiao, L. Yang, D. Jiang, R. Majumder, and F. Wei · 2022
Cited alongside, same era.
An explanation of in-context learning as implicit bayesian inference
S. M. Xie, A. Raghunathan, P. Liang, and T. Ma · 2022
Cited alongside, same era.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Cited alongside, same era.
A theory for emergence of complex skills in language models, 2023
S. Arora and A. Goyal · 2023
Cited alongside, same era.
Pythia: A suite for analyzing large language models across training and scaling
S. Biderman, H. Schoelkopf, Q. G. Anthony, H. Bradley, K. O’Brien, E. Hallahan, M. A. Khan, S. Purohit, U. S. Prashanth, E. Raff, et al · 2023
Cited alongside, same era.
Sok: Memorization in general-purpose large language models
V. Hartmann, A. Suri, V. Bindschaedler, D. Evans, S. Tople, and R. West · 2023
Cited alongside, same era.
OLMo: Accelerating the science of language models
D. Groeneveld, I. Beltagy, E. Walsh, A. Bhagia, R. Kinney, O. Tafjord, A. Jha, H. Ivison, I. Magnusson, Y. Wang, S. Arora, D. Atkinson, R. Authur, K. Chandu, A. Cohan, J. Dumas, Y. Elazar, Y. Gu, J. Hessel, T. Khot, W. Merrill, J. Morrison, N. Muennighoff, A. Naik, C. Nam, M. Peters, V. Pyatkin, A. Ravichander, D. Schwenk, S. Shah, W. Smith, E. Strubell, N. Subramani, M. Wortsman, P. Dasigi, N. Lambert, K. Richardson, L. Zettlemoyer, J. Dodge, K. Lo, L. Soldaini, N. Smith, and H. Hajishirzi · 2024
Closest in time.
Investigating data contamination for pre-training language models
M. Jiang, K. Z. Liu, M. Zhong, R. Schaeffer, S. Ouyang, J. Han, and S. Koyejo · 2024
Closest in time.
Lmd3: Language model data density dependence
J. Kirchenbauer, G. Honke, G. Somepalli, J. Geiping, D. Ippolito, K. Lee, T. Goldstein, and D. Andre · 2024
Closest in time.
Infini-gram: Scaling unbounded n-gram language models to a trillion tokens
J. Liu, S. Min, L. Zettlemoyer, Y. Choi, and H. Hajishirzi · 2024
Closest in time.
Evaluating n n -gram novelty of language models using rusty-dawg
W. Merrill, N. A. Smith, and Y. Elazar · 2024
Closest in time.
Detection and measurement of syntactic templates in generated text
C. Shaib, Y. Elazar, J. J. Li, and B. C. Wallace · 2024
Closest in time.
Functional benchmarks for robust evaluation of reasoning performance, and the reasoning gap
S. Srivastava, A. PV, S. Menon, A. Sukumar, A. Philipose, S. Prince, S. Thomas, et al · 2024
Closest in time.
X. Wang, A. Amayuelas, K. Zhang, L. Pan, W. Chen, and W. Y. Wang · 2024
Closest in time.