Fetching the paper…
Reading the bibliography…
Artificial intelligence is making spectacular progress, and one of the best examples is the development of large language models (LLMs) such as OpenAI's GPT series.
Generating Long Sequences with Sparse Transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 1904
Earlier work this paper cites.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 1910
Earlier work this paper cites.
Xxii. programming a computer for playing chess
Claude E Shannon · 1950
Earlier work this paper cites.
Studies in linguistic analysis
John Rupert Firth · 1957
Earlier work this paper cites.
Empirical explorations of the logic theory machine: a case study in heuristic
Allen Newell, John Clifford Shaw, and Herbert A Simon · 1957
Earlier work this paper cites.
The perceptron: a probabilistic model for information storage and organization in the brain
Frank Rosenblatt · 1958
Earlier work this paper cites.
Procedural reflection in programming languages volume i
Brian Cantwell Smith · 1982
Earlier work this paper cites.
The modularity of mind
Jerry A Fodor · 1983
Earlier work this paper cites.
A general framework for parallel distributed processing
David E. Rumelhart, Geoffrey E. Hinton, and James L. McClelland · 1986
Earlier work this paper cites.
Society of mind
Marvin Minsky · 1988
Earlier work this paper cites.
The recent excitement about neural networks
Francis Crick · 1989
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
George V. Cybenko · 1989
Earlier work this paper cites.
On the computational power of neural nets
Hava T Siegelmann and Eduardo D Sontag · 1992
Earlier work this paper cites.
The mathematics of statistical machine translation: Parameter estimation
Peter F Brown, Stephen A Della Pietra, Vincent J Della Pietra, Robert L Mercer, et al · 1993
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Mitchell Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz · 1993
Earlier work this paper cites.
An introduction to computational learning theory
Michael J Kearns and Umesh Vazirani · 1994
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Foundations of statistical natural language processing
Christopher Manning and Hinrich Schutze · 1999
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, and Pascal Vincent · 2000
Earlier work this paper cites.
Scaling Laws for Neural Language Models, January 2020
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2001
Earlier work this paper cites.
Behind Deep Blue: Building the computer that defeated the world chess champion
Feng-Hsiung Hsu · 2002
Earlier work this paper cites.
Pattern theory: the mathematics of perception
David Mumford · 2002
Earlier work this paper cites.
Information theory, inference and learning algorithms
David JC MacKay · 2003
Earlier work this paper cites.
Machines who think: A personal inquiry into the history and prospects of artificial intelligence
Pamela McCorduck and Cli Cfe · 2004
Earlier work this paper cites.
On the Linguistic Capacity of Real-Time Counter Automata
William Merrill · 2004
Earlier work this paper cites.
Dimensions of Neural-symbolic Integration - A Structured Survey, November 2005
Sebastian Bader and Pascal Hitzler · 2005
Earlier work this paper cites.
Language Models are Few-Shot Learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2005
Earlier work this paper cites.
Finding Universal Grammatical Relations in Multilingual BERT
Ethan A. Chi, John Hewitt, and Christopher D. Manning · 2005
Earlier work this paper cites.
Pattern recognition and machine learning
Christopher M Bishop and Nasser M Nasrabadi · 2006
Earlier work this paper cites.
Efficient selectivity and backup operators in monte-carlo tree search
Rémi Coulom · 2006
Earlier work this paper cites.
High Dimensional Statistical Inference and Random Matrices, November 2006
Iain M. Johnstone · 2006
Earlier work this paper cites.
Hopfield Networks is All You Need, April 2021
Hubert Ramsauer, Bernhard Schäfl, Johannes Lehner, Philipp Seidl, Michael Widrich, Thomas Adler, Lukas Gruber, Markus Holzleitner, Milena Pavlović, Geir Kjetil Sandve, Victor Greiff, David Kreil, Michael Kopp, Günter Klambauer, Johannes Brandstetter, and Sepp Hochreiter · 2008
Earlier work this paper cites.
Computational complexity: a modern approach
Sanjeev Arora and Boaz Barak · 2009
Earlier work this paper cites.
Speech and language processing, 2009
Dan Jurafsky and James H Martin · 2009
Earlier work this paper cites.
Probabilistic graphical models: principles and techniques
Daphne Koller and Nir Friedman · 2009
Earlier work this paper cites.
Information, physics, and computation
Marc Mezard and Andrea Montanari · 2009
Earlier work this paper cites.
The quest for artificial intelligence
Nils J Nilsson · 2009
Earlier work this paper cites.
Mathematical Foundations for a Compositional Distributional Model of Meaning, March 2010
Bob Coecke, Mehrnoosh Sadrzadeh, and Stephen Clark · 2010
Earlier work this paper cites.
Vision: A computational investigation into the human representation and processing of visual information
David Marr · 2010
Earlier work this paper cites.
Pattern theory: the stochastic analysis of real-world signals
David Mumford and Agnès Desolneux · 2010
Earlier work this paper cites.
Artificial intelligence a modern approach
Stuart J Russell · 2010
Earlier work this paper cites.
Inductive Biases for Deep Learning of Higher-Level Cognition, August 2022
Anirudh Goyal and Yoshua Bengio · 2011
Earlier work this paper cites.
Fast and slow thinking
Daniel Kahneman · 2011
Earlier work this paper cites.
A survey of monte carlo tree search methods
Cameron B. Browne, Edward Powley, Daniel Whitehouse, Simon M. Lucas, Peter I. Cowling, Philipp Rohlfshagen, Stephen Tavener, Diego Perez, Spyridon Samothrakis, and Simon Colton · 2012
Earlier work this paper cites.
Ronald DeVore, Boris Hanin, and Guergana Petrova · 2012
Earlier work this paper cites.
Neurosymbolic AI: The 3rd Wave, December 2020
Artur d’Avila Garcez and Luis C. Lamb · 2012
Earlier work this paper cites.
Lecture notes on metric embeddings
Jirǐ Matoušek · 2013
Earlier work this paper cites.
Efficient Estimation of Word Representations in Vector Space, September 2013
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
GloVe: Global Vectors for Word Representation
Jeffrey Pennington, Richard Socher, and Christopher Manning · 2014
Cited alongside, same era.
The Loss Surfaces of Multilayer Networks, January 2015
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2015
Cited alongside, same era.
Human-level concept learning through probabilistic program induction
B. M. Lake, R. Salakhutdinov, and J. B. Tenenbaum · 2015
Cited alongside, same era.
Popular talks and private discussion, 2015
Yann LeCun · 2015
Cited alongside, same era.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Cited alongside, same era.
Deep unsupervised learning using nonequilibrium thermodynamics
Language Models (Mostly) Know What They Know, July 2022
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield Dodds, Nova DasSarma, Eli Tran-Johnson, Scott Johnston, Sheer El-Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bowman, Stanislav Fort, Deep Ganguli, Danny Hernandez, Josh Jacobson, Jackson Kernion, Shauna Kravec, Liane Lovitt, Kamal Ndousse, Catherine Olsson, Sam Ringer, Dario Amodei, Tom Brown, Jack Clark, Nicholas Joseph, Ben Mann, Sam McCandlish, Chris Olah, and Jared Kaplan · 2022
Later among the works it cites.
A path towards autonomous machine intelligence, 2022
Yann LeCun · 2022
Later among the works it cites.
Solving Quantitative Reasoning Problems with Language Models, June 2022
Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra · 2022
Later among the works it cites.
Holistic Evaluation of Language Models, November 2022
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christopher Ré, Diana Acosta-Navas, Drew A. Hudson, Eric Zelikman, Esin Durmus, Faisal Ladhak, Frieda Rong, Hongyu Ren, Huaxiu Yao, Jue Wang, Keshav Santhanam, Laurel Orr, Lucia Zheng, Mert Yuksekgonul, Mirac Suzgun, Nathan Kim, Neel Guha, Niladri Chatterji, Omar Khattab, Peter Henderson, Qian Huang, Ryan Chi, Sang Michael Xie, Shibani Santurkar, Surya Ganguli, Tatsunori Hashimoto, Thomas Icard, Tianyi Zhang, Vishrav Chaudhary, William Wang, Xuechen Li, Yifan Mai, Yuhui Zhang, and Yuta Koreeda · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli · 2015
Cited alongside, same era.
Group equivariant convolutional networks
Taco Cohen and Max Welling · 2016
Cited alongside, same era.
Deep learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Cited alongside, same era.
Statistical Physics, Optimization, Inference, and Message-Passing Algorithms: Lecture Notes of the Les Houches School of Physics: Special Issue, October 2013
Florent Krzakala, Federico Ricci-Tersenghi, Lenka Zdeborova, Eric W Tramel, Riccardo Zecchina, and Leticia F Cugliandolo · 2016
Cited alongside, same era.
Building Machines That Learn and Think Like People
Brenden M Lake, Tomer D Ullman, Joshua B Tenenbaum, and Samuel J Gershman · 2016
Cited alongside, same era.
Yuri Manin and Matilde Marcolli · 2016
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2017
Cited alongside, same era.
Later among the works it cites.
Transformers Learn Shortcuts to Automata, October 2022
Bingbin Liu, Jordan T. Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang · 2022
Later among the works it cites.
A Solvable Model of Neural Scaling Laws, October 2022
Alexander Maloney, Daniel A. Roberts, and James Sully · 2022
Later among the works it cites.
Saturated Transformers are Constant-Depth Threshold Circuits
William Merrill, Ashish Sabharwal, and Noah A. Smith · 2022
Later among the works it cites.
Mechanistic interpretability, variables, and the importance of interpretable bases, 2022
Chris Olah · 2022
Later among the works it cites.
In-context Learning and Induction Heads, September 2022
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2022
Later among the works it cites.
Formal Algorithms for Transformers, July 2022
Mary Phuong and Marcus Hutter · 2022
Later among the works it cites.
Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets
Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra · 2022
Later among the works it cites.
Measuring and Narrowing the Compositionality Gap in Language Models, October 2022
Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah A. Smith, and Mike Lewis · 2022
Later among the works it cites.
Beyond neural scaling laws: beating power law scaling via data pruning, June 2022
Ben Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli, and Ari S. Morcos · 2022
Later among the works it cites.
Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, et al · 2022
Later among the works it cites.
Chess as a Testbed for Language Model State Tracking, May 2022
Shubham Toshniwal, Sam Wiseman, Karen Livescu, and Kevin Gimpel · 2022
Later among the works it cites.
Emergent Abilities of Large Language Models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus · 2022
Later among the works it cites.
An Explanation of In-context Learning as Implicit Bayesian Inference, July 2022
Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma · 2022
Later among the works it cites.
Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer, March 2022
Greg Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen, and Jianfeng Gao · 2022
Later among the works it cites.
MiniF2F: a cross-system benchmark for formal Olympiad-level mathematics, February 2022
Kunhao Zheng, Jesse Michael Han, and Stanislas Polu · 2022
Later among the works it cites.
ProofNet: Autoformalizing and Formally Proving Undergraduate-Level Mathematics, February 2023
Zhangir Azerbayev, Bartosz Piotrowski, Hailey Schoelkopf, Edward W. Ayers, Dragomir Radev, and Jeremy Avigad · 2023
Closest in time.
Sparks of Artificial General Intelligence: Early experiments with GPT-4, March 2023
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang · 2023
Closest in time.
Tighter Bounds on the Expressivity of Transformer Encoders, May 2023
David Chiang, Peter Cholak, and Anand Pillay · 2023
Closest in time.
Chatgpt is a blurry jpeg of the web
Ted Chiang · 2023
Closest in time.
A Toy Model of Universality: Reverse Engineering How Networks Learn Group Operations, May 2023
Bilal Chughtai, Lawrence Chan, and Neel Nanda · 2023
Closest in time.
Trapping LLM Hallucinations Using Tagged Context Prompts, June 2023
Philip Feldman, James R. Foulds, and Shimei Pan · 2023
Closest in time.
Learning Transformer Programs, June 2023
Dan Friedman, Alexander Wettig, and Danqi Chen · 2023
Closest in time.
What Can Transformers Learn In-Context? A Case Study of Simple Function Classes, January 2023
Shivam Garg, Dimitris Tsipras, Percy Liang, and Gregory Valiant · 2023
Closest in time.
Grokking modular arithmetic, January 2023
Andrey Gromov · 2023
Closest in time.
Thilo Hagendorff · 2023
Closest in time.
A Theory of Emergent In-Context Learning as Implicit Structure Induction, March 2023
Michael Hahn and Navin Goyal · 2023
Closest in time.
Energy Transformer, February 2023
Benjamin Hoover, Yuchen Liang, Bao Pham, Rameswar Panda, Hendrik Strobelt, Duen Horng Chau, Mohammed J. Zaki, and Dmitry Krotov · 2023
Closest in time.
Consistency Analysis of ChatGPT, March 2023
Myeongjun Jang and Thomas Lukasiewicz · 2023
Closest in time.
Kenneth Li, Aspen K. Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg · 2023
Closest in time.
Inference-Time Intervention: Eliciting Truthful Answers from a Language Model, June 2023
Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg · 2023
Closest in time.
Let’s Verify Step by Step, May 2023
Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe · 2023
Closest in time.
Dissociating language and thought in large language models: a cognitive perspective, January 2023
Kyle Mahowald, Anna A. Ivanova, Idan A. Blank, Nancy Kanwisher, Joshua B. Tenenbaum, and Evelina Fedorenko · 2023
Closest in time.
Mathematical Structure of Syntactic Merge, May 2023
Matilde Marcolli, Noam Chomsky, and Robert Berwick · 2023
Closest in time.
The Parallelism Tradeoff: Limitations of Log-Precision Transformers, April 2023
William Merrill and Ashish Sabharwal · 2023
Closest in time.
The Quantization Model of Neural Scaling, March 2023
Eric J. Michaud, Ziming Liu, Uzay Girit, and Max Tegmark · 2023
Closest in time.
Progress measures for grokking via mechanistic interpretability, January 2023
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt · 2023
Closest in time.
Are emergent abilities of large language models a mirage?
Rylan Schaeffer, Brando Miranda, and Oluwasanmi Koyejo · 2023
Closest in time.
An Introduction to Transformers, July 2023
Richard E. Turner · 2023
Closest in time.
The Learnability of In-Context Learning, March 2023
Noam Wies, Yoav Levine, and Amnon Shashua · 2023
Closest in time.
What Is ChatGPT Doing… and Why Does It Work?
Stephen Wolfram · 2023
Closest in time.
Tree of Thoughts: Deliberate Problem Solving with Large Language Models, May 2023
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan · 2023
Closest in time.
Beyond Positive Scaling: How Negation Impacts Scaling Trends of Language Models, May 2023
Yuhui Zhang, Michihiro Yasunaga, Zhengping Zhou, Jeff Z. HaoChen, James Zou, Percy Liang, and Serena Yeung · 2023
Closest in time.
Do Transformers Parse while Predicting the Masked Word?, March 2023
Haoyu Zhao, Abhishek Panigrahi, Rong Ge, and Sanjeev Arora · 2023
Closest in time.
A Survey of Large Language Models, September 2023
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie, and Ji-Rong Wen · 2023
Closest in time.