Fetching the paper…
Reading the bibliography…
Since the success of GPT, large language models (LLMs) have been revolutionizing machine learning and have initiated the so-called LLM prompting paradigm.
An unsolvable problem of elementary number theory
Alonzo Church · 1936
Earlier work this paper cites.
λ \lambda -definability and recursiveness
Stephen Cole Kleene · 1936
Earlier work this paper cites.
A universal Turing machine with two internal states
Claude E. Shannon · 1956
Earlier work this paper cites.
A variant to Turing’s theory of computing machines
Hao Wang · 1957
Earlier work this paper cites.
On context-free languages and push-down automata
Marcel Paul Schützenberger · 1963
Earlier work this paper cites.
On the computational complexity of algorithms
Juris Hartmanis and Richard E. Stearns · 1965
Earlier work this paper cites.
Two-tape simulation of multitape Turing machines
Fred C. Hennie and Richard Edwin Stearns · 1966
Earlier work this paper cites.
Visual feature extraction by a multilayered network of analog threshold elements
Kunihiko Fukushima · 1969
Earlier work this paper cites.
Effect of guard digits and normalization options on floating point multiplication
R. Goodman and Alan Feldstein · 1977
Earlier work this paper cites.
Computability, complexity, and languages: fundamentals of theoretical computer science
Martin Davis, Ron Sigal, and Elaine J Weyuker · 1994
Earlier work this paper cites.
Computational Complexity: A Modern Approach
Sanjeev Arora and Boaz Barak · 2009
Earlier work this paper cites.
Wang’s B machines are efficiently universal, as is Hasenjaeger’s small universal electromechanical toy
Turlough Neary, Damien Woods, Niall Murphy, and Rainer Glaschick · 2014
Earlier work this paper cites.
Layer normalization
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Improving language understanding by generative pre-training, 2018
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Earlier work this paper cites.
On the Turing completeness of modern neural network architectures
Jorge Pérez, Javier Marinković, and Pablo Barceló · 2019
Earlier work this paper cites.
On the computational power of Transformers and its implications in sequence modeling
Satwik Bhattamishra, Arkil Patel, and Navin Goyal · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Theoretical limitations of self-attention in neural sequence models
Michael Hahn · 2020
Cited alongside, same era.
Attention is Turing-complete
Jorge Pérez, Pablo Barceló, and Javier Marinkovic · 2021
Cited alongside, same era.
What learning algorithm is in-context learning? investigations with linear models
Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou · 2022
Cited alongside, same era.
Formal language recognition by hard attention Transformers: Perspectives from circuit complexity
Yiding Hao, Dana Angluin, and Robert Frank · 2022
Cited alongside, same era.
DIMES: A differentiable meta solver for combinatorial optimization problems
Ruizhong Qiu, Zhiqing Sun, and Yiming Yang · 2022
Cited alongside, same era.
Tighter bounds on the expressivity of Transformer encoders
David Chiang, Peter Cholak, and Anand Pillay · 2023
Learning universal predictors
Jordi Grau-Moya, Tim Genewein, Marcus Hutter, Laurent Orseau, Gregoire Deletang, Elliot Catt, Anian Ruoss, Li Kevin Wenliang, Christopher Mattern, Matthew Aitchison, and Joel Veness · 2024
Closest in time.
On the sensitivity of individual fairness: Measures and robust algorithms
Xinyu He, Jian Kang, Ruizhong Qiu, Fei Wang, Jose Sepulveda, and Hanghang Tong · 2024
Closest in time.
Universal length generalization with Turing programs
Kaiying Hou, David Brandfonbrener, Sham Kakade, Samy Jelassi, and Eran Malach · 2024
Closest in time.
Chain of thought empowers Transformers to solve inherently serial problems
Zhiyuan Li, Hong Liu, Denny Zhou, and Tengyu Ma · 2024
Closest in time.
BackTime: Backdoor attacks on multivariate time series forecasting
Xiao Lin, Zhining Liu, Dongqi Fu, Ruizhong Qiu, and Hanghang Tong · 2024
Closest in time.
Introducing Llama 3.1: Our most capable models to date, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Auto-regressive next-token predictors are universal learners
Eran Malach · 2023
Cited alongside, same era.
The parallelism tradeoff: Limitations of log-precision Transformers
William Merrill and Ashish Sabharwal · 2023
Cited alongside, same era.
Reconstructing graph diffusion history from a single snapshot
Ruizhong Qiu, Dingsu Wang, Lei Ying, H Vincent Poor, Yifang Zhang, and Hanghang Tong · 2023
Cited alongside, same era.
How powerful are decoder-only Transformer neural models?
Jesse Roberts · 2023
Cited alongside, same era.
Transformers learn in-context by gradient descent
Johannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov · 2023
Cited alongside, same era.
Networked time series imputation via position-aware graph enhanced variational autoencoders
Dingsu Wang, Yuchen Yan, Ruizhong Qiu, Yada Zhu, Kaiyu Guan, Andrew Margenot, and Hanghang Tong · 2023
Cited alongside, same era.
Meta · 2024
Closest in time.
Hello gpt-4o, 2024
OpenAI · 2024
Closest in time.
Gradient compressed sensing: A query-efficient gradient estimator for high-dimensional zeroth-order optimization
Ruizhong Qiu and Hanghang Tong · 2024
Closest in time.
Linear Transformers are versatile in-context learners
Max Vladymyrov, Johannes von Oswald, Mark Sandler, and Rong Ge · 2024
Closest in time.
Robust watermarking for diffusion models: A unified multi-dimensional recipe, 2024
Tianxin Wei, Ruizhong Qiu, Yifan Chen, Yunzhe Qi, Jiacheng Lin, Wenju Xu, Sreyashi Nag, Ruirui Li, Hanqing Lu, Zhengyang Wang, Chen Luo, Hui Liu, Suhang Wang, Jingrui He, Qi He, and Xianfeng Tang · 2024
Closest in time.
Fair anomaly detection for imbalanced groups
Ziwei Wu, Lecheng Zheng, Yuancheng Yu, Ruizhong Qiu, John Birge, and Jingrui He · 2024
Closest in time.
Discrete-state continuous-time diffusion for graph generation
Zhe Xu, Ruizhong Qiu, Yuzhong Chen, Huiyuan Chen, Xiran Fan, Menghai Pan, Zhichen Zeng, Mahashweta Das, and Hanghang Tong · 2024
Closest in time.
Ensuring user-side fairness in dynamic recommender systems
Hyunsik Yoo, Zhichen Zeng, Jian Kang, Ruizhong Qiu, David Zhou, Zhining Liu, Fei Wang, Charlie Xu, Eunice Chan, and Hanghang Tong · 2024
Closest in time.
Graph mixup on approximate Gromov–Wasserstein geodesics
Zhichen Zeng, Ruizhong Qiu, Zhe Xu, Zhining Liu, Yuchen Yan, Tianxin Wei, Lei Ying, Jingrui He, and Hanghang Tong · 2024
Closest in time.
Trained Transformers learn linear models in-context
Ruiqi Zhang, Spencer Frei, and Peter L. Bartlett · 2024
Closest in time.
AIM: Attributing, interpreting, mitigating data unfairness
Zhining Liu, Ruizhong Qiu, Zhichen Zeng, Yada Zhu, Hendrik Hamann, and Hanghang Tong · 2025
Closest in time.
Generalizable recommender system during temporal popularity distribution shifts
Hyunsik Yoo, Ruizhong Qiu, Charlie Xu, Fei Wang, and Hanghang Tong · 2025
Closest in time.