Fetching the paper…
Reading the bibliography…
Despite the remarkable capabilities of modern large language models (LLMs), the mechanisms behind their problem-solving abilities remain elusive.
Algorithmic stability and generalization performance
O. Bousquet and A. Elisseeff · 2000
Earlier work this paper cites.
Norm-based capacity control in neural networks
B. Neyshabur, R. Tomioka, and N. Srebro · 2015
Earlier work this paper cites.
Estimating accuracy from unlabeled data: A bayesian approach
E. A. Platanios, A. Dubey, and T. Mitchell · 2016
Earlier work this paper cites.
A closer look at memorization in deep networks
D. Arpit, S. Jastrzebski, N. Ballas, D. Krueger, E. Bengio, M. S. Kanwal, T. Maharaj, A. Fischer, A. Courville, Y. Bengio, et al · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization, 2017
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2017
Earlier work this paper cites.
Nearly-tight vc-dimension and pseudodimension bounds for piecewise linear neural networks
P. L. Bartlett, N. Harvey, C. Liaw, and A. Mehrabian · 2019
Earlier work this paper cites.
Fantastic generalization measures and where to find them
Y. Jiang, B. Neyshabur, H. Mobahi, D. Krishnan, and S. Bengio · 2019
Earlier work this paper cites.
Generalization in deep networks: The role of distance from initialization
V. Nagarajan and J. Z. Kolter · 2019
Earlier work this paper cites.
Does learning require memorization? a short tale about a long tail
V. Feldman · 2020
Earlier work this paper cites.
What neural networks memorize and why: Discovering the long tail via influence estimation
V. Feldman and C. Zhang · 2020
Earlier work this paper cites.
Extracting training data from large language models
N. Carlini, F. Tramer, E. Wallace, M. Jagielski, A. Herbert-Voss, K. Lee, A. Roberts, T. Brown, D. Song, U. Erlingsson, et al · 2021
Cited alongside, same era.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al · 2021
Cited alongside, same era.
Measuring mathematical problem solving with the math dataset
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt · 2021
Cited alongside, same era.
Training data leakage analysis in language models
H. A. Inan, O. Ramadan, L. Wutschitz, D. Jones, V. Rühle, J. Withers, and R. Sim · 2021
Cited alongside, same era.
# instag: Instruction tagging for analyzing supervised fine-tuning of large language models
K. Lu, H. Yuan, Z. Yuan, R. Lin, J. Lin, C. Tan, C. Zhou, and J. Zhou · 2023
Later among the works it cites.
A preliminary study of the intrinsic relationship between complexity and alignment
Y. Zhao, B. Yu, B. Hui, H. Yu, F. Huang, Y. Li, and N. L. Zhang · 2023
Later among the works it cites.
A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Yang, A. Fan, et al · 2024
Closest in time.
Dsdm: Model-aware dataset selection with datamodels
L. Engstrom, A. Feldmann, and A. Madry · 2024
Closest in time.
Be like a goldfish, don’t memorize! mitigating memorization in generative llms
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Jiang, V. Nagarajan, C. Baek, and J. Z. Kolter · 2021
Cited alongside, same era.
Leveraging unlabeled data to predict out-of-distribution performance
S. Garg, S. Balakrishnan, Z. C. Lipton, B. Neyshabur, and H. Sedghi · 2022
Cited alongside, same era.
Memorization without overfitting: Analyzing the training dynamics of large language models
K. Tirumala, A. Markosyan, L. Zettlemoyer, and A. Aghajanyan · 2022
Cited alongside, same era.
Alpagasus: Training a better alpaca with fewer data
L. Chen, S. Li, J. Yan, H. Wang, K. Gunaratna, V. Yadav, Z. Tang, V. Srinivasan, T. Zhou, H. Huang, et al · 2023
Cited alongside, same era.
Studying large language model generalization with influence functions
R. Grosse, J. Bae, C. Anil, N. Elhage, A. Tamkin, A. Tajdini, B. Steiner, D. Li, E. Durmus, E. Perez, et al · 2023
Cited alongside, same era.
M. Li, Y. Zhang, Z. Li, J. Chen, L. Chen, N. Cheng, J. Wang, T. Zhou, and J. Xiao · 2023
Cited alongside, same era.
A. Hans, Y. Wen, N. Jain, J. Kirchenbauer, H. Kazemi, P. Singhania, S. Singh, G. Somepalli, J. Geiping, A. Bhatele, et al · 2024
Closest in time.
W. Liu, W. Zeng, K. He, Y. Jiang, and J. He · 2024
Closest in time.
D. Mekala, A. Nguyen, and J. Shang · 2024
Closest in time.
Rl on incorrect synthetic data scales the efficiency of llm math reasoning by eight-fold
A. Setlur, S. Garg, X. Geng, N. Garg, V. Smith, and A. Kumar · 2024
Closest in time.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, M. Zhang, Y. Li, Y. Wu, and D. Guo · 2024
Closest in time.
Gemma 2: Improving open language models at a practical size
G. Team, M. Riviere, S. Pathak, P. G. Sessa, C. Hardin, S. Bhupatiraju, L. Hussenot, T. Mesnard, B. Shahriari, A. Ramé, et al · 2024
Closest in time.