Fetching the paper…
Reading the bibliography…
We prove theoretically that generalization improves not only through data scaling but also by compressing internal representations.
Shannon, C.E. A Mathematical Theory of Communication
1948
Earlier work this paper cites.
Naftali Tishby, Fernando C. Pereira, and William Bialek. The information bottleneck method. arXiv preprint physics/0004057, 2000
2000
Earlier work this paper cites.
Cover, T.M. and Thomas, J.A. Elements of Information Theory
2006
Earlier work this paper cites.
Ohad Shamir, Sivan Sabato, and Naftali Tishby. Learning and generalization with the information bottleneck. Theoretical Computer Science , 411(29):2696–2711, 2010
2010
Earlier work this paper cites.
Giraldo, L.G.S., Rao, M., and Principe, J.C. Measures of entropy from data using infinitely divisible kernels
2014
Earlier work this paper cites.
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. Language Models are Unsupervised Multitask Learners
2019
Earlier work this paper cites.
2020
Cited alongside, same era.
2022
Cited alongside, same era.
Tadros, T., Krishnan, G.P., Ramyaa, R., Bazhenov, M Sleep-like unsupervised replay reduces catastrophic forgetting in artificial neural networks
2022
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Later among the works it cites.
modded-nanogpt: Speedrunning the NanoGPT baseline
Keller Jordan, Jeremy Bernstein, Brendan Rappazzo, @fernbear.bsky.social, Vlado Boza, Jiacheng You, Franz Cesista, Braden Koszarsky, and @Grad62304977 · 2024
Later among the works it cites.
DeepSeek AI. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
2025
Closest in time.
2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
Yu, F., Arora, H.S., and Johnson, M. Iterative Graph Alignment
2024
Cited alongside, same era.
Tishby, N. and Zaslavsky, N. Deep Learning and the Information Bottleneck Principle
Cited in the paper.
Oscar Skean and Md Rifat Arefin and Dan Zhao and Niket Patel and Jalal Naghiyev and Yann LeCun and Ravid Shwartz-Ziv Layer by Layer: Uncovering Hidden Representations in Language Models
Cited in the paper.
Ravid Shwartz-Ziv and Naftali Tishby Opening the Black Box of Deep Neural Networks via Information
Cited in the paper.
Cited in the paper.
Yuzhen Huang and Jinghan Zhang and Zifei Shan and Junxian He Compression Represents Intelligence Linearly
Cited in the paper.
Oscar C González, Yury Sokolov, Giri P Krishnan, Jean Erik Delanois, Maxim Bazhenov Can sleep protect memorise from catastrophic forgetting?
Cited in the paper.
2025
Closest in time.