Fetching the paper…
Reading the bibliography…
Chain-of-Thought (CoT) reasoning is known to improve Large Language Models both empirically and in terms of theoretical approximation power.
ε \varepsilon -entropy and ε \varepsilon -capacity of sets in functional spaces
Andrey Kolmogorov and Vladimir Tikhomirov · 1959
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Kurt Hornik, Maxwell Stinchcombe, and Halbert White · 1989
Earlier work this paper cites.
The Nature of Statistical Learning Theory
Vladimir Vapnik · 1995
Earlier work this paper cites.
Matplotlib: A 2D graphics environment
J. D. Hunter · 2007
Earlier work this paper cites.
Learning theory estimates via integral operators and their approximations
Steve Smale and Ding-Xuan Zhou · 2007
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2017
Diederik P. Kingma and Jimmy Ba · 2017
Earlier work this paper cites.
Failures of gradient-based deep learning, 2017
Shai Shalev-Shwartz, Ohad Shamir, and Shaked Shammah · 2017
Earlier work this paper cites.
PyTorch: An imperative style, high-performance deep learning library, 2019
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Earlier work this paper cites.
On the Turing completeness of modern neural network architectures, 2019
Jorge Pérez, Javier Marinković, and Pablo Barceló · 2019
Earlier work this paper cites.
Understanding the role of individual units in a deep neural network
David Bau, Jun-Yan Zhu, Hendrik Strobelt, Agata Lapedriza, Bolei Zhou, and Antonio Torralba · 2020
Earlier work this paper cites.
Language models are few-shot learners, 2020
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Theoretical limitations of self-attention in neural sequence models
Michael Hahn · 2020
Earlier work this paper cites.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Earlier work this paper cites.
An interpretability illusion for BERT, 2021
Tolga Bolukbasi, Adam Pearce, Ann Yuan, Andy Coenen, Emily Reif, Fernanda Viégas, and Martin Wattenberg · 2021
Earlier work this paper cites.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the MATH dataset, 2021
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive NLP tasks, 2021
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela · 2021
Earlier work this paper cites.
A survey of transformers, 2021
Tianyang Lin, Yuxin Wang, Xiangyang Liu, and Xipeng Qiu · 2021
Earlier work this paper cites.
Show your work: Scratchpads for intermediate computation with language models, 2021
Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, Charles Sutton, and Augustus Odena · 2021
Earlier work this paper cites.
The merged-staircase property: a necessary and nearly sufficient condition for SGD learning of sparse functions on two-layer neural networks, 2022
Emmanuel Abbe, Enric Boix-Adsera, and Theodor Misiakiewicz · 2022
Earlier work this paper cites.
A path towards autonomous machine intelligence, 2022
Yann LeCun · 2022
Cited alongside, same era.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2022
Cited alongside, same era.
Interpretability in the wild: a circuit for indirect object identification in GPT-2 small, 2022
Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt · 2022
Cited alongside, same era.
An explanation of in-context learning as implicit bayesian inference, 2022
Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma · 2022
Cited alongside, same era.
What learning algorithm is in-context learning? Investigations with linear models, 2023
Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou · 2023
Cited alongside, same era.
Progress measures for grokking via mechanistic interpretability, 2023
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt · 2023
Later among the works it cites.
RWKV: Reinventing RNNs for the transformer era, 2023
Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Stella Biderman, Huanqi Cao, Xin Cheng, Michael Chung, Matteo Grella, Kranthi Kiran GV, Xuzheng He, Haowen Hou, Jiaju Lin, Przemyslaw Kazienko, Jan Kocon, Jiaming Kong, Bartlomiej Koptyra, Hayden Lau, Krishna Sri Ipsit Mantri, Ferdinand Mom, Atsushi Saito, Guangyu Song, Xiangru Tang, Bolun Wang, Johan S. Wind, Stanislaw Wozniak, Ruichong Zhang, Zhenyuan Zhang, Qihang Zhao, Peng Zhou, Qinghua Zhou, Jian Zhu, and Rui-Jie Zhu · 2023
Later among the works it cites.
The mechanistic basis of data dependence and abrupt learning in an in-context classification task, 2023
Gautam Reddy · 2023
Later among the works it cites.
Outliers with opposing signals have an outsized effect on neural network optimization, 2023
Elan Rosenfeld and Andrej Risteski · 2023
Later among the works it cites.
Transformers as recognizers of formal languages: A survey on expressivity, 2023
Lena Strobl, William Merrill, Gail Weiss, David Chiang, and Dana Angluin · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
SGD with large step sizes learns sparse features, 2023
Maksym Andriushchenko, Aditya Varre, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2023
Cited alongside, same era.
A theory for emergence of complex skills in language models, 2023
Sanjeev Arora and Anirudh Goyal · 2023
Cited alongside, same era.
Hidden progress in deep learning: SGD learns parities near the computational limit, 2023
Boaz Barak, Benjamin L. Edelman, Surbhi Goel, Sham Kakade, Eran Malach, and Cyril Zhang · 2023
Cited alongside, same era.
Birth of a transformer: A memory viewpoint, 2023
Alberto Bietti, Vivien Cabannes, Diane Bouchacourt, Herve Jegou, and Leon Bottou · 2023
Cited alongside, same era.
Neural networks and the Chomsky hierarchy, 2023
Grégoire Delétang, Anian Ruoss, Jordi Grau-Moya, Tim Genewein, Li Kevin Wenliang, Elliot Catt, Chris Cundy, Marcus Hutter, Shane Legg, Joel Veness, and Pedro A. Ortega · 2023
Cited alongside, same era.
Towards revealing the mystery behind chain of thought: A theoretical perspective, 2023
Guhao Feng, Bohang Zhang, Yuntian Gu, Haotian Ye, Di He, and Liwei Wang · 2023
Cited alongside, same era.
What can transformers learn in-context? A case study of simple function classes, 2023
Shivam Garg, Dimitris Tsipras, Percy Liang, and Gregory Valiant · 2023
Cited alongside, same era.
Later among the works it cites.
Attention is all you need, 2023
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2023
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models, 2023
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou · 2023
Later among the works it cites.
The clock and the pizza: Two stories in mechanistic explanation of neural networks, 2023
Ziqian Zhong, Ziming Liu, Max Tegmark, and Jacob Andreas · 2023
Later among the works it cites.
Revisiting feature prediction for learning visual representations from video, 2024
Adrien Bardes, Quentin Garrido, Jean Ponce, Xinlei Chen, Michael Rabbat, Yann LeCun, Mahmoud Assran, and Nicolas Ballas · 2024
Closest in time.
Sudden drops in the loss: Syntax acquisition, phase transitions, and simplicity bias in MLMs, 2024
Angelica Chen, Ravid Shwartz-Ziv, Kyunghyun Cho, Matthew L. Leavitt, and Naomi Saphra · 2024
Closest in time.
Vision transformers need registers, 2024
Timothée Darcet, Maxime Oquab, Julien Mairal, and Piotr Bojanowski · 2024
Closest in time.
The evolution of statistical induction heads: In-context learning Markov chains, 2024
Benjamin L. Edelman, Ezra Edelman, Surbhi Goel, Eran Malach, and Nikolaos Tsilivis · 2024
Closest in time.
Think before you speak: Training language models with pause tokens, 2024
Sachin Goyal, Ziwei Ji, Ankit Singh Rawat, Aditya Krishna Menon, Sanjiv Kumar, and Vaishnavh Nagarajan · 2024
Closest in time.
Chain of thought empowers transformers to solve inherently serial problems, 2024
Zhiyuan Li, Hong Liu, Denny Zhou, and Tengyu Ma · 2024
Closest in time.
The expressive power of transformers with chain of thought, 2024
William Merrill and Ashish Sabharwal · 2024
Closest in time.
How transformers learn causal structure with gradient descent, 2024
Eshaan Nichani, Alex Damian, and Jason D. Lee · 2024
Closest in time.
GPT-4 technical report, 2024
OpenAI et al · 2024
Closest in time.
Transformers, parallel computation, and logarithmic depth, 2024
Clayton Sanford, Daniel Hsu, and Matus Telgarsky · 2024
Closest in time.
Memory mosaics, 2024
Jianyu Zhang, Niklas Nolte, Ranajoy Sadhukhan, Beidi Chen, and Léon Bottou · 2024
Closest in time.