Fetching the paper…
Reading the bibliography…
The rapid advancement of large language models (LLMs) calls for a rigorous theoretical framework to explain their empirical success.
On computable numbers, with an application to the entscheidungsproblem
Alan Mathison Turing et al · 1936
Earlier work this paper cites.
A preliminary report on a general theory of inductive inference
Ray J Solomonoff · 1960
Earlier work this paper cites.
Three approaches to the quantitative definition ofinformation’
Andrei N Kolmogorov · 1965
Earlier work this paper cites.
On the length of programs for computing finite binary sequences
Gregory J Chaitin · 1966
Earlier work this paper cites.
Algorithmic information theory
Gregory J Chaitin · 1977
Earlier work this paper cites.
Kolmogorov’s contributions to information theory and algorithmic complexity
Thomas M Cover, Peter Gacs, and Robert M Gray · 1989
Earlier work this paper cites.
Language models are few-shot learners, 2020
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2005
Earlier work this paper cites.
Universal artificial intelligence: Sequential decisions based on algorithmic probability
Marcus Hutter · 2005
Earlier work this paper cites.
An introduction to Kolmogorov complexity and its applications , volume 3
Ming Li, Paul Vitányi, et al · 2008
Earlier work this paper cites.
Algorithmic randomness and complexity
Rodney G Downey and Denis R Hirschfeldt · 2010
Earlier work this paper cites.
SMS Spam Collection
Tiago Almeida and Jos Hidalgo · 2011
Earlier work this paper cites.
Solomonoff induction: A solution to the problem of the priors?
Aron Vallinder · 2012
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Jake Zhao, and Yann LeCun · 2015
Earlier work this paper cites.
The description length of deep learning models
Léonard Blier and Yann Ollivier · 2018
Earlier work this paper cites.
Universal artificial intelligence: Practical agents and fundamental challenges
Tom Everitt and Marcus Hutter · 2018
Earlier work this paper cites.
CARER: Contextualized affect representations for emotion recognition
Elvis Saravia, Hsien-Chi Toby Liu, Yen-Hao Huang, Junlin Wu, and Yi-Shin Chen · 2018
Earlier work this paper cites.
Transformer-based image compression
Ming Lu, Peiyao Guo, Huiqing Shi, Chuntong Cao, and Zhan Ma · 2021
Earlier work this paper cites.
Language models as zero-shot planners: Extracting actionable knowledge for embodied agents, 2022
Wenlong Huang, Pieter Abbeel, Deepak Pathak, and Igor Mordatch · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa · 2022
Earlier work this paper cites.
Transformerlens
Neel Nanda and Joseph Bloom · 2022
Cited alongside, same era.
Emergent abilities of large language models, 2022
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
Language modeling is compression
Grégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt, Tim Genewein, Christopher Mattern, Jordi Grau-Moya, Li Kevin Wenliang, Matthew Aitchison, Laurent Orseau, et al · 2023
Cited alongside, same era.
Metagpt: Meta programming for multi-agent collaborative framework
Introducing openai o1-preview
OpenAI · 2024
Later among the works it cites.
Scaling llm test-time compute optimally can be more effective than scaling model parameters, 2024
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar · 2024
Later among the works it cites.
Qwq: Reflect deeply on the boundaries of the unknown, November 2024
Qwen Team · 2024
Later among the works it cites.
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Mei Li, Mingfeng Xue, Pei Zhang, Qin Zhu, Rui Men, Runji Lin, Tianhao Li, Tianyi Tang, Tingyu Xia, Xingzhang Ren, Xuancheng Ren, Yang Fan, Yang Su, Yichang Zhang, Yu Wan, Yuqiong Liu, Zeyu Cui, Zhenru Zhang, and Zihan Qiu · 2024
Later among the works it cites.
What matters in training a gpt4-style language model with multimodal inputs?
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sirui Hong, Xiawu Zheng, Jonathan Chen, Yuheng Cheng, Ceyao Zhang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, Liyang Zhou, Chenyu Ran, et al · 2023
Cited alongside, same era.
Lans: A layout-aware neural solver for plane geometry problem
Zhong-Zhi Li, Ming-Liang Zhang, Fei Yin, and Cheng-Lin Liu · 2023
Cited alongside, same era.
Locating and editing factual associations in gpt, 2023
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2023
Cited alongside, same era.
Baolin Peng, Chunyuan Li, Pengcheng He, Michel Galley, and Jianfeng Gao · 2023
Cited alongside, same era.
Label words are anchors: An information flow perspective for understanding in-context learning, 2023
Lean Wang, Lei Li, Damai Dai, Deli Chen, Hao Zhou, Fandong Meng, Jie Zhou, and Xu Sun · 2023
Cited alongside, same era.
Eva-kellm: A new benchmark for evaluating knowledge editing of llms, 2023
Suhang Wu, Minlong Peng, Yue Chen, Jinsong Su, and Mingming Sun · 2023
Cited alongside, same era.
Language modeling is compression, 2024
Grégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt, Tim Genewein, Christopher Mattern, Jordi Grau-Moya, Li Kevin Wenliang, Matthew Aitchison, Laurent Orseau, Marcus Hutter, and Joel Veness · 2024
Cited alongside, same era.
Chao Deng, Jiale Yuan, Pi Bu, Peijie Wang, Zhong-Zhi Li, Jian Xu, Xiao-Hui Li, Yuan Gao, Jun Song, Bo Zheng, et al · 2024
Cited alongside, same era.
Yan Zeng, Hanbo Zhang, Jiani Zheng, Jiangnan Xia, Guoqiang Wei, Yang Wei, Yuchen Zhang, Tao Kong, and Ruihua Song · 2024
Later among the works it cites.
Distributed rule vectors is a key mechanism in large language models’ in-context learning, 2024
Bowen Zheng, Ming Ma, Zhongqiao Lin, and Tianming Yang · 2024
Later among the works it cites.
Ultraif: Advancing instruction following from the wild
Kaikai An, Li Sheng, Ganqu Cui, Shuzheng Si, Ning Ding, Yu Cheng, and Baobao Chang · 2025
Closest in time.
Parameters vs. context: Fine-grained control of knowledge reliance in language models
Baolong Bi, Shenghua Liu, Yiwei Wang, Yilong Xu, Junfeng Fang, Lingrui Mei, and Xueqi Cheng · 2025
Closest in time.
Yuyao Ge, Shenghua Liu, Yiwei Wang, Lingrui Mei, Lizhe Chen, Baolong Bi, and Xueqi Cheng · 2025
Closest in time.
Axis: Efficient human-agent-computer interaction with api-first llm-based agents, 2025
Junting Lu, Zhiyang Zhang, Fangkai Yang, Jue Zhang, Lu Wang, Chao Du, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang, and Qi Zhang · 2025
Closest in time.
a1: Steep test-time scaling law via environment augmented generation, 2025
Lingrui Mei, Shenghua Liu, Yiwei Wang, Baolong Bi, Yuyao Ge, Jun Wan, Yurong Wu, and Xueqi Cheng · 2025
Closest in time.
Qwen2.5 technical report, 2025
Qwen, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Mei Li, Mingfeng Xue, Pei Zhang, Qin Zhu, Rui Men, Runji Lin, Tianhao Li, Tianyi Tang, Tingyu Xia, Xingzhang Ren, Xuancheng Ren, Yang Fan, Yang Su, Yichang Zhang, Yu Wan, Yuqiong Liu, Zeyu Cui, Zhenru Zhang, and Zihan Qiu · 2025
Closest in time.
A practical review of mechanistic interpretability for transformer-based language models, 2025
Daking Rai, Yilun Zhou, Shi Feng, Abulhair Saparov, and Ziyu Yao · 2025
Closest in time.
Kimi k1.5: Scaling reinforcement learning with llms, 2025
Kimi Team, Angang Du, Bofei Gao, Bowei Xing, Changjiu Jiang, Cheng Chen, Cheng Li, Chenjun Xiao, Chenzhuang Du, Chonghua Liao, Chuning Tang, Congcong Wang, Dehao Zhang, Enming Yuan, Enzhe Lu, Fengxiang Tang, Flood Sung, and et al · 2025
Closest in time.
Unifying two types of scaling laws from the perspective of conditional kolmogorov complexity
Jun Wan · 2025
Closest in time.
Vulnerability of text-to-image models to prompt template stealing: A differential evolution approach
Yurong Wu, Fangwen Mu, Qiuhong Zhang, Jinjing Zhao, Xinrun Xu, Lingrui Mei, Yang Wu, Lin Shi, Junjie Wang, Zhiming Ding, et al · 2025
Closest in time.
Redstar: Does scaling long-cot data unlock better slow-reasoning systems?
Haotian Xu, Xing Wu, Weinong Wang, Zhongzhi Li, Da Zheng, Boyuan Chen, Yi Hu, Shijia Kang, Jiaming Ji, Yingying Zhang, et al · 2025
Closest in time.
Vem: Environment-free exploration for training gui agent with value environment model
Jiani Zheng, Lu Wang, Fangkai Yang, Chaoyun Zhang, Lingrui Mei, Wenjie Yin, Qingwei Lin, Dongmei Zhang, Saravan Rajmohan, and Qi Zhang · 2025
Closest in time.
Trustrag: Enhancing robustness and trustworthiness in rag, 2025
Huichi Zhou, Kin-Hei Lee, Zhonghao Zhan, Yue Chen, Zhenhao Li, Zhaoyang Wang, Hamed Haddadi, and Emine Yilmaz · 2025
Closest in time.