Fetching the paper…
Reading the bibliography…
The widespread adoption of large language models (LLMs) makes it important to recognize their strengths and limitations.
Language Models are Few-Shot Learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; Agarwal, S.; Herbert-Voss, A.; Krueger, G.; Henighan, T.; Child, R.; Ramesh, A.; Ziegler, D.; Wu, J.; Winter, C.; Hesse, C.; Chen, M.; Sigler, E.; Litwin, M.; Gray, S.; Chess, B.; Clark, J.; Berner, C.; McCandlish, S.; Radford, A.; Sutskever, I.; and Amodei, D. 2020 · 1901
Earlier work this paper cites.
RoBERTa: A robustly optimized BERT pretraining approach
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019 · 1907
Earlier work this paper cites.
Mechanization in problem solving: The effect of Einstellung
Luchins, A. S. 1942 · 1942
Earlier work this paper cites.
Computing machinery and intelligence
Turing, A. M. 1950 · 1950
Earlier work this paper cites.
The scars of human evolution
Krogman, W. M. 1951 · 1951
Earlier work this paper cites.
Binary codes capable of correcting deletions, insertions, and reversals
Levenshtein, V. I. 1966 · 1966
Earlier work this paper cites.
The Codebreakers
Kahn, D. 1967 · 1967
Earlier work this paper cites.
On the genesis of abstract ideas
Posner, M. I.; and Keele, S. W. 1968 · 1968
Earlier work this paper cites.
The spandrels of San Marco and the Panglossian paradigm: A critique of the adaptationist programme
Gould, S.; and Lewontin, R. 1979 · 1979
Earlier work this paper cites.
Contextual effects on word perception and eye movements during reading
Ehrlich, S. F.; and Rayner, K. 1981 · 1981
Earlier work this paper cites.
The Appeal of Parallel Distributed Processing
McClelland, J. L.; Rumelhart, D. E.; and Hinton, G. E. 1986 · 1986
Earlier work this paper cites.
On learning the past tenses of English verbs
Rumelhart, D. E.; and McClelland, J. L. 1986 · 1986
Earlier work this paper cites.
Toward a universal law of generalization for psychological science
Shepard, R. N. 1987 · 1987
Earlier work this paper cites.
Connectionism and cognitive architecture: A critical analysis
Fodor, J. A.; and Pylyshyn, Z. W. 1988 · 1988
Earlier work this paper cites.
On language and connectionism: Analysis of a parallel distributed processing model of language acquisition
Pinker, S.; and Prince, A. 1988 · 1988
Earlier work this paper cites.
On the proper treatment of connectionism
Smolensky, P. 1988 · 1988
Earlier work this paper cites.
A comparative study of cache recovery by three corvid species
Balda, R. P.; and Kamil, A. C. 1989 · 1989
Earlier work this paper cites.
The Adaptive Character of Thought
Anderson, J. R. 1990 · 1990
Earlier work this paper cites.
Finding structure in time
Elman, J. L. 1990 · 1990
Earlier work this paper cites.
Tensor product variable binding and the representation of symbolic structures in connectionist systems
Smolensky, P. 1990 · 1990
Earlier work this paper cites.
Distributed representations, simple recurrent networks, and grammatical structure
Elman, J. L. 1991 · 1991
Earlier work this paper cites.
Foundations of Vision
Wandell, B. A. 1995 · 1995
Earlier work this paper cites.
Long short-term memory
Hochreiter, S.; and Schmidhuber, J. 1997 · 1997
Earlier work this paper cites.
Adaptations, exaptations, and spandrels
Buss, D. M.; Haselton, M. G.; Shackelford, T. K.; Bleske, A. L.; and Wakefield, J. C. 1998 · 1998
Earlier work this paper cites.
Rethinking eliminative connectionism
Marcus, G. F. 1998 · 1998
Earlier work this paper cites.
The Code Book
Singh, S. 1999 · 1999
Earlier work this paper cites.
A neural probabilistic language model
Bengio, Y.; Ducharme, R.; and Vincent, P. 2000 · 2000
Earlier work this paper cites.
Three kinds of adaptationism
Godfrey-Smith, P. 2001 · 2001
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J.; McCandlish, S.; Henighan, T.; Brown, T. B.; Chess, B.; Child, R.; Gray, S.; Radford, A.; Wu, J.; and Amodei, D. 2020 · 2001
Earlier work this paper cites.
Language and number: A bilingual training study
Spelke, E. S.; and Tsivkin, S. 2001 · 2001
Earlier work this paper cites.
Empirical assessment of stimulus poverty arguments
Pullum, G. K.; and Scholz, B. C. 2002 · 2002
Earlier work this paper cites.
Underdetermination in language games: Survey & analysis of Pig Latin dialects
Vaux, B.; and Nevins, A. I. 2003 · 2003
Earlier work this paper cites.
Longformer: The Long-Document Transformer
Beltagy, I.; Peters, M. E.; and Cohan, A. 2020 · 2004
Earlier work this paper cites.
The perils of being bipedal
Latimer, B. 2005 · 2005
Earlier work this paper cites.
Transferring inductive biases through knowledge distillation
Abnar, S.; Dehghani, M.; and Zuidema, W. 2020 · 2006
Earlier work this paper cites.
Can machine learning be secure?
Barreno, M.; Nelson, B.; Sears, R.; Joseph, A. D.; and Tygar, J. D. 2006 · 2006
Earlier work this paper cites.
The PASCAL Recognising Textual Entailment Challenge
Dagan, I.; Glickman, O.; and Magnini, B. 2006 · 2006
Earlier work this paper cites.
Functional explanation and the function of explanation
Lombrozo, T.; and Carey, S. 2006 · 2006
Earlier work this paper cites.
Harmony optimization and the computational architecture of the mind/brain
Smolensky, P.; and Legendre, G. 2006 · 2006
Earlier work this paper cites.
A weakly informative default prior distribution for logistic and other regression models
Gelman, A.; Jakulin, A.; Pittau, M. G.; and Su, Y.-S. 2008 · 2008
Earlier work this paper cites.
Simultaneous Inference in General Parametric Models
Hothorn, T.; Bretz, F.; and Westfall, P. 2008 · 2008
Earlier work this paper cites.
Natural Language Processing with Python
Bird, S.; Loper, E.; and Klein, E. 2009 · 2009
Earlier work this paper cites.
Quirks of Human Anatomy: An Evo-Devo Look at the Human Body
Held Jr, L. I. 2009 · 2009
Earlier work this paper cites.
Regularization paths for generalized linear models via coordinate descent
Friedman, J.; Hastie, T.; and Tibshirani, R. 2010 · 2010
Earlier work this paper cites.
Cognition, Evolution, and Behavior
Shettleworth, S. J. 2010 · 2010
Earlier work this paper cites.
Language models are open knowledge graphs
Wang, C.; Liu, X.; and Song, D. 2020 · 2010
Earlier work this paper cites.
Rational integration of noisy evidence and prior semantic expectations in sentence interpretation
Gibson, E.; Bergen, L.; and Piantadosi, S. T. 2013 · 2013
Earlier work this paper cites.
The effect of word predictability on reading time is logarithmic
Smith, N. J.; and Levy, R. 2013 · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
Szegedy, C.; Zaremba, W.; Sutskever, I.; Bruna, J.; Erhan, D.; Goodfellow, I.; and Fergus, R. 2014 · 2014
Earlier work this paper cites.
Explaining and Harnessing Adversarial Examples
Goodfellow, I.; Shlens, J.; and Szegedy, C. 2015 · 2015
Earlier work this paper cites.
Deep learning
LeCun, Y.; Bengio, Y.; and Hinton, G. 2015 · 2015
Earlier work this paper cites.
The ancestral shape hypothesis: an evolutionary explanation for the occurrence of intervertebral disc herniation in humans
Plomp, K. A.; Viðarsdóttir, U. S.; Weston, D. A.; Dobney, K.; and Collard, M. 2015 · 2015
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? Debiasing word embeddings
Bolukbasi, T.; Chang, K.-W.; Zou, J. Y.; Saligrama, V.; and Kalai, A. T. 2016 · 2016
Earlier work this paper cites.
Are we smart enough to know how smart animals are?
De Waal, F. 2016 · 2016
Earlier work this paper cites.
Neural Text Generation from Structured Data with Application to the Biography Domain
Lebret, R.; Grangier, D.; and Auli, M. 2016 · 2016
Earlier work this paper cites.
Neural Machine Translation of Rare Words with Subword Units
Sennrich, R.; Haddow, B.; and Birch, A. 2016 · 2016
Earlier work this paper cites.
ggplot2: Elegant Graphics for Data Analysis
Wickham, H. 2016 · 2016
Earlier work this paper cites.
Semantics derived automatically from language corpora contain human-like biases
Caliskan, A.; Bryson, J. J.; and Narayanan, A. 2017 · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C.; Abbeel, P.; and Levine, S. 2017 · 2017
Cited alongside, same era.
Pointer sentinel mixture models
Merity, S.; Xiong, C.; Bradbury, J.; and Socher, R. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Think you have solved question answering? Try ARC, the AI2 reasoning challenge
Clark, P.; Cowhey, I.; Etzioni, O.; Khot, T.; Sabharwal, A.; Schoenick, C.; and Tafjord, O. 2018 · 2018
Cited alongside, same era.
Recasting Gradient-Based Meta-Learning as Hierarchical Bayes
Language models are general-purpose interfaces
Hao, Y.; Song, H.; Dong, L.; Huang, S.; Chi, Z.; Wang, W.; Ma, S.; and Wei, F. 2022 · 2022
Later among the works it cites.
State-of-the-art generalisation research in NLP: a taxonomy and review
Hupkes, D.; Giulianelli, M.; Dankers, V.; Artetxe, M.; Elazar, Y.; Pimentel, T.; Christodoulopoulos, C.; Lasri, K.; Saphra, N.; Sinclair, A.; et al. 2022 · 2022
Later among the works it cites.
Evaluation gaps in machine learning practice
Hutchinson, B.; Rostamzadeh, N.; Greer, C.; Heller, K.; and Prabhakaran, V. 2022 · 2022
Later among the works it cites.
Models In a Spelling Bee: Language Models Implicitly Learn the Character Composition of Tokens
Itzhak, I.; and Levy, O. 2022 · 2022
Later among the works it cites.
Language models (mostly) know what they know
Kadavath, S.; Conerly, T.; Askell, A.; Henighan, T.; Drain, D.; Perez, E.; Schiefer, N.; Hatfield-Dodds, Z.; DasSarma, N.; Tran-Johnson, E.; Johnston, S.; El-Showk, S.; Jones, A.; Elhage, N.; Hume, T.; Chen, A.; Bai, Y.; Bowman, S.; Fort, S.; Ganguli, D.; Hernandez, D.; Jacobson, J.; Kernion, J.; Kravec, S.; Lovitt, L.; Ndousse, K.; Olsson, C.; Ringer, S.; Amodei, D.; Brown, T.; Clark, J.; Joseph, N.; Mann, B.; McCandlish, S.; Olah, C.; and Kaplan, J. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Grant, E.; Finn, C.; Levine, S.; Darrell, T.; and Griffiths, T. 2018 · 2018
Cited alongside, same era.
Generalization without Systematicity: On the Compositional Skills of Sequence-to-Sequence Recurrent Networks
Lake, B. M.; and Baroni, M. 2018 · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Radford, A.; Narasimhan, K.; Salimans, T.; and Sutskever, I. 2018 · 2018
Cited alongside, same era.
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference
Williams, A.; Nangia, N.; and Bowman, S. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Cited alongside, same era.
Troubling trends in machine learning scholarship
Lipton, Z. C.; and Steinhardt, J. 2019 · 2019
Cited alongside, same era.
Mechanistic versus functional understanding
Lombrozo, T.; and Wilkenfeld, D. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
What do tokens know about their characters and how do they know it?
Kaushal, A.; and Mahowald, K. 2022 · 2022
Later among the works it cites.
Kim, N.; Linzen, T.; and Smolensky, P. 2022 · 2022
Later among the works it cites.
Large language models are zero-shot reasoners
Kojima, T.; Gu, S. S.; Reid, M.; Matsuo, Y.; and Iwasawa, Y. 2022 · 2022
Later among the works it cites.
Lampinen, A. K. 2022 · 2022
Later among the works it cites.
Deduplicating Training Data Makes Language Models Better
Lee, K.; Ippolito, D.; Nystrom, A.; Zhang, C.; Eck, D.; Callison-Burch, C.; and Carlini, N. 2022 · 2022
Later among the works it cites.
Can you hear me now? Sensitive comparisons of human and machine perception
Lepori, M. A.; and Firestone, C. 2022 · 2022
Later among the works it cites.
Holistic evaluation of language models
Liang, P.; Bommasani, R.; Lee, T.; Tsipras, D.; Soylu, D.; Yasunaga, M.; Zhang, Y.; Narayanan, D.; Wu, Y.; Kumar, A.; Newman, B.; Yuan, B.; Yan, B.; Zhang, C.; Cosgrove, C.; Manning, C. D.; Ré, C.; Acosta-Navas, D.; Hudson, D. A.; Zelikman, E.; Durmus, E.; Ladhak, F.; Rong, F.; Ren, H.; Yao, H.; Wang, J.; Santhanam, K.; Orr, L.; Zheng, L.; Yuksekgonul, M.; Suzgun, M.; Kim, N.; Guha, N.; Chatterji, N.; Khattab, O.; Henderson, P.; Huang, Q.; Chi, R.; Xie, S. M.; Santurkar, S.; Ganguli, S.; Hashimoto, T.; Icard, T.; Zhang, T.; Chaudhary, V.; Wang, W.; Li, X.; Mai, Y.; Zhang, Y.; and Koreeda, Y. 2022 · 2022
Later among the works it cites.
Entailment Semantics Can Be Extracted from an Ideal Language Model
Merrill, W.; Warstadt, A.; and Linzen, T. 2022 · 2022
Later among the works it cites.
Reducing Conversational Agents’ Overconfidence Through Linguistic Calibration
Mielke, S. J.; Szlam, A.; Dinan, E.; and Boureau, Y.-L. 2022 · 2022
Later among the works it cites.
Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?
Min, S.; Lyu, X.; Holtzman, A.; Artetxe, M.; Lewis, M.; Hajishirzi, H.; and Zettlemoyer, L. 2022 · 2022
Later among the works it cites.
Transformers Can Do Bayesian Inference
Müller, S.; Hollmann, N.; Arango, S. P.; Grabocka, J.; and Hutter, F. 2022 · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C. L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; Schulman, J.; Hilton, J.; Kelton, F.; Miller, L.; Simens, M.; Askell, A.; Welinder, P.; Christiano, P.; Leike, J.; and Lowe, R. 2022 · 2022
Later among the works it cites.
Meaning without reference in large language models
Piantadosi, S.; and Hill, F. 2022 · 2022
Later among the works it cites.
R: A Language and Environment for Statistical Computing
R Core Team. 2022 · 2022
Later among the works it cites.
Scaling Language Models: Methods, Analysis & Insights from Training Gopher
Rae, J. W.; Borgeaud, S.; Cai, T.; Millican, K.; Hoffmann, J.; Song, F.; Aslanides, J.; Henderson, S.; Ring, R.; Young, S.; Rutherford, E.; Hennigan, T.; Menick, J.; Cassirer, A.; Powell, R.; van den Driessche, G.; Hendricks, L. A.; Rauh, M.; Huang, P.-S.; Glaese, A.; Welbl, J.; Dathathri, S.; Huang, S.; Uesato, J.; Mellor, J.; Higgins, I.; Creswell, A.; McAleese, N.; Wu, A.; Elsen, E.; Jayakumar, S.; Buchatskaya, E.; Budden, D.; Sutherland, E.; Simonyan, K.; Paganini, M.; Sifre, L.; Martens, L.; Li, X. L.; Kuncoro, A.; Nematzadeh, A.; Gribovskaya, E.; Donato, D.; Lazaridou, A.; Mensch, A.; Lespiau, J.-B.; Tsimpoukelli, M.; Grigorev, N.; Fritz, D.; Sottiaux, T.; Pajarskas, M.; Pohlen, T.; Gong, Z.; Toyama, D.; de Masson d’Autume, C.; Li, Y.; Terzi, T.; Mikulik, V.; Babuschkin, I.; Clark, A.; de Las Casas, D.; Guy, A.; Jones, C.; Bradbury, J.; Johnson, M.; Hechtman, B.; Weidinger, L.; Gabriel, I.; Isaac, W.; Lockhart, E.; Osindero, S.; Rimell, L.; Dyer, C.; Vinyals, O.; Ayoub, K.; Stanway, J.; Bennett, L.; Hassabis, D.; Kavukcuoglu, K.; and Irving, G. 2022 · 2022
Later among the works it cites.
Impact of pretraining term frequencies on few-shot numerical reasoning
Razeghi, Y.; Logan IV, R. L.; Gardner, M.; and Singh, S. 2022 · 2022
Later among the works it cites.
Talking about large language models
Shanahan, M. 2022 · 2022
Later among the works it cites.
Neurocompositional computing: From the Central Paradox of Cognition to a new generation of AI systems
Smolensky, P.; McCoy, R. T.; Fernandez, R.; Goldrick, M.; and Gao, J. 2022 · 2022
Later among the works it cites.
Do Prompt-Based Models Really Understand the Meaning of Their Prompts?
Webson, A.; and Pavlick, E. 2022 · 2022
Later among the works it cites.
Do All Languages Cost the Same? Tokenization in the Era of Commercial Language Models
Ahia, O.; Kumar, S.; Gonen, H.; Kasai, J.; Mortensen, D. R.; Smith, N. A.; and Tsvetkov, Y. 2023 · 2023
Closest in time.
CoLT5: Faster long-range Transformers with conditional computation
Ainslie, J.; Lei, T.; de Jong, M.; Ontañón, S.; Brahma, S.; Zemlyanskiy, Y.; Uthus, D.; Guo, M.; Lee-Thorp, J.; Tay, Y.; et al. 2023 · 2023
Closest in time.
Arkoudas, K. 2023 · 2023
Closest in time.
Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
BIG-bench authors. 2023 · 2023
Closest in time.
Using cognitive psychology to understand GPT-3
Binz, M.; and Schulz, E. 2023 · 2023
Closest in time.
Sparks of artificial general intelligence: Early experiments with GPT-4
Bubeck, S.; Chandrasekaran, V.; Eldan, R.; Gehrke, J.; Horvitz, E.; Kamar, E.; Lee, P.; Lee, Y. T.; Li, Y.; Lundberg, S.; et al. 2023 · 2023
Closest in time.
Mathematics, word problems, common sense, and artificial intelligence
Davis, E. 2023 · 2023
Closest in time.
Toxicity in ChatGPT: Analyzing persona-assigned language models
Deshpande, A.; Murahari, V.; Rajpurohit, T.; Kalyan, A.; and Narasimhan, K. 2023 · 2023
Closest in time.
Faith and Fate: Limits of Transformers on Compositionality
Dziri, N.; Lu, X.; Sclar, M.; Li, X. L.; Jian, L.; Lin, B. Y.; West, P.; Bhagavatula, C.; Bras, R. L.; Hwang, J. D.; et al. 2023 · 2023
Closest in time.
LMentry: A Language Model Benchmark of Elementary Language Tasks
Efrat, A.; Honovich, O.; and Levy, O. 2023 · 2023
Closest in time.
Understanding In-Context Learning via Supportive Pretraining Data
Han, X.; Simig, D.; Mihaylov, T.; Tsvetkov, Y.; Celikyilmaz, A.; and Wang, T. 2023 · 2023
Closest in time.
Holtzman, A.; West, P.; and Zettlemoyer, L. 2023 · 2023
Closest in time.
Entity Tracking in Language Models
Kim, N.; and Schuster, S. 2023 · 2023
Closest in time.
Kosoy, E.; Reagan, E. R.; Lai, L.; Gopnik, A.; and Cobb, D. K. 2023 · 2023
Closest in time.
Lost in the middle: How language models use long contexts
Liu, N. F.; Lin, K.; Hewitt, J.; Paranjape, A.; Bevilacqua, M.; Petroni, F.; and Liang, P. 2023 · 2023
Closest in time.
Dissociating language and thought in large language models: A cognitive perspective
Mahowald, K.; Ivanova, A. A.; Blank, I. A.; Kanwisher, N.; Tenenbaum, J. B.; and Fedorenko, E. 2023 · 2023
Closest in time.
Auto-Regressive Next-Token Predictors are Universal Learners
Malach, E. 2023 · 2023
Closest in time.
Modeling rapid language learning by distilling Bayesian priors into artificial neural networks
McCoy, R. T.; and Griffiths, T. L. 2023 · 2023
Closest in time.
How Much Do Language Models Copy From Their Training Data? Evaluating Linguistic Novelty in Text Generation Using RAVEN
McCoy, R. T.; Smolensky, P.; Linzen, T.; Gao, J.; and Celikyilmaz, A. 2023 · 2023
Closest in time.
Inverse Scaling: When Bigger Isn’t Better
McKenzie, I. R.; Lyzhov, A.; Pieler, M.; Parrish, A.; Mueller, A.; Prabhu, A.; McLean, E.; Kirtland, A.; Ross, A.; Liu, A.; et al. 2023 · 2023
Closest in time.
How do we know how smart AI systems are?
Mitchell, M. 2023 · 2023
Closest in time.
The debate over understanding in AI’s large language models
Mitchell, M.; and Krakauer, D. C. 2023 · 2023
Closest in time.
Mollo, D. C.; and Millière, R. 2023 · 2023
Closest in time.
OpenAI. 2023 · 2023
Closest in time.
The ROOTS Search Tool: Data Transparency for LLMs
Piktus, A.; Akiki, C.; Villegas, P.; Laurençon, H.; Dupont, G.; Luccioni, S.; Jernite, Y.; and Rogers, A. 2023 · 2023
Closest in time.
Elastic net regularization paths for all generalized linear models
Tay, J. K.; Narasimhan, B.; and Hastie, T. 2023 · 2023
Closest in time.
Llama 2: Open Foundation and Fine-Tuned Chat Models
Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; Bikel, D.; Blecher, L.; Ferrer, C. C.; Chen, M.; Cucurull, G.; Esiobu, D.; Fernandes, J.; Fu, J.; Fu, W.; Fuller, B.; Gao, C.; Goswami, V.; Goyal, N.; Hartshorn, A.; Hosseini, S.; Hou, R.; Inan, H.; Kardas, M.; Kerkez, V.; Khabsa, M.; Kloumann, I.; Korenev, A.; Koura, P. S.; Lachaux, M.-A.; Lavril, T.; Lee, J.; Liskovich, D.; Lu, Y.; Mao, Y.; Martinet, X.; Mihaylov, T.; Mishra, P.; Molybog, I.; Nie, Y.; Poulton, A.; Reizenstein, J.; Rungta, R.; Saladi, K.; Schelten, A.; Silva, R.; Smith, E. M.; Subramanian, R.; Tan, X. E.; Tang, B.; Taylor, R.; Williams, A.; Kuan, J. X.; Xu, P.; Yan, Z.; Zarov, I.; Zhang, Y.; Fan, A.; Kambadur, M.; Narang, S.; Rodriguez, A.; Stojnic, R.; Edunov, S.; and Scialom, T. 2023 · 2023
Closest in time.
Adversarial Policies Beat Superhuman Go AIs
Wang, T. T.; Gleave, A.; Tseng, T.; Pelrine, K.; Belrose, N.; Miller, J.; Dennis, M. D.; Duan, Y.; Pogrebniak, V.; Levine, S.; and Russell, S. 2023 · 2023
Closest in time.
Wu, Z.; Qiu, L.; Ross, A.; Akyürek, E.; Chen, B.; Wang, B.; Kim, N.; Andreas, J.; and Kim, Y. 2023 · 2023
Closest in time.
Tree of Thoughts: Deliberate Problem Solving with Large Language Models
Yao, S.; Yu, D.; Zhao, J.; Shafran, I.; Griffiths, T. L.; Cao, Y.; and Narasimhan, K. 2023 · 2023
Closest in time.
Yiu, E.; Kosoy, E.; and Gopnik, A. 2023 · 2023
Closest in time.
Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
Zhou, D.; Schärli, N.; Hou, L.; Wei, J.; Scales, N.; Wang, X.; Schuurmans, D.; Cui, C.; Bousquet, O.; Le, Q. V.; and Chi, E. H. 2023 · 2023
Closest in time.