Fetching the paper…
Reading the bibliography…
We present a universal theoretical framework for understanding long-context language modeling based on a bipartite mutual information scaling law that we rigorously verify in natural language.
Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context (2019)
Dai, Z. et al · 1901
Earlier work this paper cites.
Generating Long Sequences with Sparse Transformers (2019)
Child, R., Gray, S., Radford, A. & Sutskever, I · 1904
Earlier work this paper cites.
On Variational Bounds of Mutual Information (2019)
Poole, B., Ozair, S., van den Oord, A., Alemi, A. A. & Tucker, G · 1905
Earlier work this paper cites.
Adaptive Attention Span in Transformers (2019)
Sukhbaatar, S., Grave, E., Bojanowski, P. & Joulin, A · 1905
Earlier work this paper cites.
Mutual Information Scaling and Expressive Power of Sequence Models (2019)
Shen, H · 1905
Earlier work this paper cites.
HuggingFace’s Transformers: State-of-the-art Natural Language Processing (2020)
Wolf, T. et al · 1910
Earlier work this paper cites.
Compressive Transformers for Long-Range Sequence Modelling (2019)
Rae, J. W., Potapenko, A., Jayakumar, S. M. & Lillicrap, T. P · 1911
Earlier work this paper cites.
PyTorch: An Imperative Style, High-Performance Deep Learning Library (2019)
Paszke, A. et al · 1912
Earlier work this paper cites.
Asymptotic evaluation of certain markov process expectations for large time, i
Donsker, M. D. & Varadhan, S. S · 1975
Earlier work this paper cites.
On bounds for packings on a sphere and in space
Kabatiansky, G. A. & Levenshtein, V. I · 1978
Earlier work this paper cites.
Der bekannte Grenzwert der redundanzfreien Information in Texten - eine Fehlinterpretation der Shannonschen Experimente?
Hilberg, W · 1990
Earlier work this paper cites.
Entropy and Long-Range Correlations in Literary English (1994)
Ebeling, W. & Pöschel, T · 1994
Earlier work this paper cites.
Entropy and Long-Range Correlations in Literary English
Ebeling, W. & Pöschel, T · 1994
Earlier work this paper cites.
The rate-distortion dimension of sets and measures
Kawabata, T. & Dembo, A · 1994
Earlier work this paper cites.
Long-range correlations between letters and sentences in texts
Ebeling, W. & Neiman, A · 1995
Earlier work this paper cites.
Power laws and universality
Stanley, H. E · 1995
Earlier work this paper cites.
Predictive Information (1999)
Bialek, W. & Tishby, N · 1999
Earlier work this paper cites.
Zipf and Heaps Laws’ Coefficients Depend on Language
Gelbukh, A. & Sidorov, G · 2001
Earlier work this paper cites.
Scaling Laws for Neural Language Models (2020)
Kaplan, J. et al · 2001
Earlier work this paper cites.
LONG-RANGE FRACTAL CORRELATIONS IN LITERARY CORPORA
MONTEMURRO, M. A. & PURY, P. A · 2002
Earlier work this paper cites.
Efficient Estimation of Mutual Information for Strongly Dependent Variables (2015)
Gao, S., Steeg, G. V. & Galstyan, A · 2003
Earlier work this paper cites.
Longformer: The Long-Document Transformer (2020)
Beltagy, I., Peters, M. E. & Cohan, A · 2004
Earlier work this paper cites.
Estimating mutual information
Kraskov, A., Stögbauer, H. & Grassberger, P · 2004
Earlier work this paper cites.
Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention (2020)
Katharopoulos, A., Vyas, A., Pappas, N. & Fleuret, F · 2006
Earlier work this paper cites.
CLUB: A Contrastive Log-ratio Upper Bound of Mutual Information (2020)
Cheng, P. et al · 2006
Earlier work this paper cites.
Entropy Estimates from Insufficient Samplings (2008)
Grassberger, P · 2008
Earlier work this paper cites.
Excess entropy in natural language: present state and perspectives (2011)
Debowski, L · 2011
Earlier work this paper cites.
Conditional likelihood maximisation: A unifying framework for information theoretic feature selection
Brown, G., Pocock, A., Zhao, M.-J. & Luján, M · 2012
Earlier work this paper cites.
Sphere packing bounds via spherical codes
Cohn, H. & Zhao, Y · 2014
Earlier work this paper cites.
Deep learning and the information bottleneck principle
Tishby, N. & Zaslavsky, N · 2015
Earlier work this paper cites.
The Relaxed Hilberg Conjecture: A Review and New Experimental Support
Łukasz Debowski · 2015
Earlier work this paper cites.
InfoGAN: Interpretable Representation Learning by Information Maximizing Generative Adversarial Nets
Chen, X. et al · 2016
Earlier work this paper cites.
The Entropy of Words - Learnability and Expressivity across More than 1000 Languages (2017)
Bentz, C., Alikaniotis, D., Cysouw, M. & i Cancho, R. F · 2017
Earlier work this paper cites.
Critical Behavior in Physics and Probabilistic Formal Languages
Lin, H. W. & Tegmark, M · 2017
Earlier work this paper cites.
Syntactic dependencies correspond to word pairs with high mutual information
Futrell, R., Qian, P., Gibson, E., Fedorenko, E. & Blank, I · 2019
Earlier work this paper cites.
Learning deep representations by mutual information estimation and maximization (2019)
Hjelm, R. D. et al · 2019
Earlier work this paper cites.
Backflow Transformations via Neural Networks for Quantum Many-Body Wave Functions
Luo, D. & Clark, B. K · 2019
Earlier work this paper cites.
Machine learning and the physical sciences
Carleo, G. et al · 2019
Earlier work this paper cites.
Representation learning with contrastive predictive coding (2019)
van den Oord, A., Li, Y. & Vinyals, O · 2019
Cited alongside, same era.
Language Models are Few-Shot Learners
Brown, T. et al · 2020
Cited alongside, same era.
Learning to summarize from human feedback (2020)
Stiennon, N. et al · 2020
Cited alongside, same era.
On Mutual Information Maximization for Representation Learning
Tschannen, M., Djolonga, J., Rubenstein, P. K., Gelly, S. & Lucic, M · 2020
Cited alongside, same era.
Big Bird: Transformers for Longer Sequences
Zaheer, M. et al · 2020
Cited alongside, same era.
The Information Bottleneck Problem and its Applications in Machine Learning
Goldfeld, Z. & Polyanskiy, Y · 2020
Cited alongside, same era.
Efficient Memory Management for Large Language Model Serving with PagedAttention (2023)
Kwon, W. et al · 2023
Later among the works it cites.
Gauge-invariant and anyonic-symmetric autoregressive neural network for quantum lattice models
Luo, D. et al · 2023
Later among the works it cites.
From tensor-network quantum states to tensorial recurrent neural networks
Wu, D., Rossi, R., Vicentini, F. & Carleo, G · 2023
Later among the works it cites.
ANTN: Bridging autoregressive neural networks and tensor networks for quantum many-body simulation
Chen, Z., Newhouse, L., Chen, E., Luo, D. & Soljacic, M · 2023
Later among the works it cites.
The Linear Representation Hypothesis and the Geometry of Large Language Models (2023)
Park, K., Choe, Y. J. & Veitch, V · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Quantum Natural Gradient
Stokes, J., Izaac, J., Killoran, N. & Carleo, G · 2020
Cited alongside, same era.
Deep Learning Enabled Strain Mapping of Single-Atom Defects in Two-Dimensional Transition Metal Dichalcogenides with Sub-Picometer Precision
Lee, C.-H. et al · 2020
Cited alongside, same era.
Compressive Transformers for Long-Range Sequence Modelling
Rae, J. W., Potapenko, A., Jayakumar, S. M., Hillier, C. & Lillicrap, T. P · 2020
Cited alongside, same era.
Long-Short Transformer: Efficient Transformers for Language and Vision (2021)
Zhu, C. et al · 2021
Cited alongside, same era.
MINE: Mutual Information Neural Estimation (2021)
Belghazi, M. I. et al · 2021
Cited alongside, same era.
Show Your Work: Scratchpads for Intermediate Computation with Language Models (2021)
Nye, M. et al · 2021
Cited alongside, same era.
OpenAI et al · 2024
Later among the works it cites.
The Llama 3 Herd of Models (2024)
Grattafiori, A. et al · 2024
Later among the works it cites.
DeepSeek-V3 Technical Report (2024)
DeepSeek-AI et al · 2024
Later among the works it cites.
DRT-o1: Optimized Deep Reasoning Translation via Long Chain-of-Thought (2024)
Wang, J., Meng, F., Liang, Y. & Zhou, J · 2024
Later among the works it cites.
Mamba: Linear-Time Sequence Modeling with Selective State Spaces (2024)
Gu, A. & Dao, T · 2024
Later among the works it cites.
Dao, T. & Gu, A · 2024
Later among the works it cites.
Learning to (Learn at Test Time): RNNs with Expressive Hidden States (2024)
Sun, Y. et al · 2024
Later among the works it cites.
Explaining neural scaling laws
Bahri, Y., Dyer, E., Kaplan, J., Lee, J. & Sharma, U · 2024
Later among the works it cites.
A Dynamical Model of Neural Scaling Laws (2024)
Bordelon, B., Atanasov, A. & Pehlevan, C · 2024
Later among the works it cites.
Nayak, A. K. & Varshney, L. R · 2024
Later among the works it cites.
Transformers learn variable-order markov chains in-context (2024)
Zhou, R., Tian, C. & Diggavi, S · 2024
Later among the works it cites.
HMT: Hierarchical Memory Transformer for Long Context Language Processing (2024)
He, Z., Qin, Z., Prakriya, N., Sun, Y. & Cong, J · 2024
Later among the works it cites.
xLSTM: Extended Long Short-Term Memory (2024)
Beck, M. et al · 2024
Later among the works it cites.
Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models (2024)
De, S. et al · 2024
Later among the works it cites.
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision (2024)
Shah, J. et al · 2024
Later among the works it cites.
Qin, Z. et al · 2024
Later among the works it cites.
Entangling Intelligence: AI-Quantum Crossovers and Perspectives
Chen, Z. & Luo, D · 2024
Later among the works it cites.
OccamLLM: Fast and Exact Language Model Arithmetic in a Single Step
Dugan, O. et al · 2024
Later among the works it cites.
TENG: Time-Evolving Natural Gradient for Solving PDEs With Deep Neural Nets Toward Machine Precision
Chen, Z., Mccarran, J., Vizcaino, E., Soljacic, M. & Luo, D · 2024
Later among the works it cites.
QuanTA: Efficient high-rank fine-tuning of llms with quantum-informed tensor adaptation
Chen, Z. et al · 2024
Later among the works it cites.
Photonic probabilistic machine learning using quantum vacuum noise
Choi, S. et al · 2024
Later among the works it cites.
Tensor networks and efficient descriptions of classical data (2024)
Lu, S., Kanász-Nagy, M., Kukuljan, I. & Cirac, J. I · 2024
Later among the works it cites.
Mathematica, Version 14.2
Inc., W. R · 2024
Later among the works it cites.
On the Origins of Linear Representations in Large Language Models (2024)
Jiang, Y., Rajendran, G., Ravikumar, P., Aragam, B. & Veitch, V · 2024
Later among the works it cites.
FlexAttention for Efficient High-Resolution Vision-Language Models (2024)
Li, J. et al · 2024
Later among the works it cites.
Gemini: A family of highly capable multimodal models (2025)
Team, G. et al · 2025
Closest in time.
Comanici, G. et al · 2025
Closest in time.
Deepseek-r1 incentivizes reasoning in llms through reinforcement learning
Guo, D. et al · 2025
Closest in time.
Qwen2.5 technical report (2025)
Qwen et al · 2025
Closest in time.
Yang, A. et al · 2025
Closest in time.
Guo, H. et al · 2025
Closest in time.
Multimodal foundation models for material property prediction and discovery
Moro, V. et al · 2025
Closest in time.