Fetching the paper…
Reading the bibliography…
Large language models like GPT-4 have achieved remarkable proficiency in a broad spectrum of language-based tasks, some of which are traditionally associated with hallmarks of human intelligence.
Thirunavukarasu, A. J., Ting, D. S. J., Elangovan, K., Gutierrez, L., Tan, T. F. & Ting, D. S. W. (2023), ‘Large language models in medicine’, Nature Medicine
1940
Earlier work this paper cites.
Turing, A. M. (1950), ‘Computing Machinery and Intelligence’, Mind
1950
Earlier work this paper cites.
Osgood, C. E. (1952), ‘The nature and measurement of meaning’, Psychological bulletin
1952
Earlier work this paper cites.
Wittgenstein, L. (1953), Philosophical Investigations
1953
Earlier work this paper cites.
Harris, Z. S. (1954), ‘Distributional structure’, Word
1954
Earlier work this paper cites.
Firth, J. R. (1957), ‘A synopsis of linguistic theory, 1930-1955’, Studies in linguistic analysis
1955
Earlier work this paper cites.
Weaver, W. (1955), Translation, in
1955
Earlier work this paper cites.
Chomsky, N. (1957), Syntactic Structures
1957
Earlier work this paper cites.
Winograd, T. (1971), ‘Procedures as a Representation for Data in a Computer Program for Understanding Natural Language’
1971
Earlier work this paper cites.
Fodor, J. A. (1975), The Language of Thought
1975
Earlier work this paper cites.
Putnam, H. (1975), ‘The Meaning of ’Meaning”, Minnesota Studies in the Philosophy of Science
1975
Earlier work this paper cites.
Salton, G., Wong, A. & Yang, C. S. (1975), ‘A vector space model for automatic indexing’, Communications of the ACM
1975
Earlier work this paper cites.
Hume, D. (1978), A Treatise of Human Nature
1978
Earlier work this paper cites.
Kripke, S. (1980), Naming and Necessity
1980
Earlier work this paper cites.
Searle, J. R. (1980), ‘Minds, Brains, and Programs’, Behavioral and Brain Sciences
1980
Earlier work this paper cites.
Block, N. (1981), ‘Psychologism and Behaviorism’, The Philosophical Review
1981
Earlier work this paper cites.
Block, N. (1986), ‘Advertisement for a Semantics for Psychology’, Midwest Studies in Philosophy
1986
Earlier work this paper cites.
Fodor, J. A. & Pylyshyn, Z. W. (1988), ‘Connectionism and cognitive architecture: A critical analysis’, Cognition
1988
Earlier work this paper cites.
Pinker, S. & Prince, A. (1988), ‘On language and connectionism: Analysis of a parallel distributed processing model of language acquisition’, Cognition
1988
Earlier work this paper cites.
Smolensky, P. (1988), ‘On the proper treatment of connectionism’, Behavioral and Brain Sciences
1988
Earlier work this paper cites.
Smolensky, P. (1989), Connectionism and Constituent Structure, in
1989
Earlier work this paper cites.
Harnad, S. (1990), ‘The symbol grounding problem’, Physica D: Nonlinear Phenomena
1990
Earlier work this paper cites.
Schmidhuber, J. (1990), Towards Compositional Learning with Dynamic Neural Networks
1990
Earlier work this paper cites.
Macdonald, C. (1995), Classicism Vs. Connectionism, in
1995
Earlier work this paper cites.
Hochreiter, S. & Schmidhuber, J. (1997), ‘Long Short-Term Memory’, Neural Computation
1997
Earlier work this paper cites.
Marconi, D. (1997), Lexical Competence
1997
Earlier work this paper cites.
Jelinek, F. (1998), Statistical Methods for Speech Recognition
1998
Earlier work this paper cites.
Sober, E. (1998), Morgan’s canon, in
1998
Earlier work this paper cites.
Bengio, Y., Ducharme, R. & Vincent, P. (2000), A Neural Probabilistic Language Model, in
2000
Earlier work this paper cites.
Chomsky, N. (2000), Knowledqe of Lanquaqe: Its Nature, Oriqin and Use, in
2000
Earlier work this paper cites.
Baier, A. C. (2002), Hume: The Reflective Women’s Epistemologist?, in
2002
Earlier work this paper cites.
2005
Earlier work this paper cites.
Tomasello, M. (2009), Constructing a Language
2009
Earlier work this paper cites.
Lasnik, H. & Lohndal, T. (2010), ‘Government–binding/principles and parameters theory’, WIREs Cognitive Science
2010
Earlier work this paper cites.
2013
Earlier work this paper cites.
Cho, K., van Merrienboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H. & Bengio, Y. (2014), ‘Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation’
2014
Earlier work this paper cites.
Dąbrowska, E. (2015), ‘What exactly is Universal Grammar, and has anyone seen it?’, Frontiers in Psychology
2015
Earlier work this paper cites.
Buckner, C. (2017), Understanding Associative and Cognitive Explanations in Comparative Psychology, in
2017
Earlier work this paper cites.
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S. & Amodei, D. (2017), Deep Reinforcement Learning from Human Preferences, in
2017
Earlier work this paper cites.
Lake, B. M., Ullman, T. D., Tenenbaum, J. B. & Gershman, S. J. (2017), ‘Building machines that learn and think like people’, Behavioral and Brain Sciences
2017
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł. & Polosukhin, I. (2017), Attention is All you Need, in
2017
Cited alongside, same era.
Ha, D. & Schmidhuber, J. (2018), ‘World Models’
2018
Cited alongside, same era.
Lake, B. & Baroni, M. (2018), Generalization without Systematicity: On the Compositional Skills of Sequence-to-Sequence Recurrent Networks, in
2018
Cited alongside, same era.
Auersperg, A. M. I. & von Bayern, A. M. P. (2019), ‘Who’s a clever bird — now? A brief history of parrot cognition’, Behaviour
2019
Cited alongside, same era.
Chollet, F. (2019), ‘On the Measure of Intelligence’
2019
Cited alongside, same era.
Aiyappa, R., An, J., Kwak, H. & Ahn, Y.-Y. (2023), ‘Can we trust the evaluation on ChatGPT?’
2023
Later among the works it cites.
Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., Chu, E., Clark, J. H., Shafey, L. E., Huang, Y., Meier-Hellstern, K., Mishra, G., Moreira, E., Omernick, M., Robinson, K., Ruder, S., Tay, Y., Xiao, K., Xu, Y., Zhang, Y., Abrego, G. H., Ahn, J., Austin, J., Barham, P., Botha, J., Bradbury, J., Brahma, S., Brooks, K., Catasta, M., Cheng, Y., Cherry, C., Choquette-Choo, C. A., Chowdhery, A., Crepy, C., Dave, S., Dehghani, M., Dev, S., Devlin, J., Díaz, M., Du, N., Dyer, E., Feinberg, V., Feng, F., Fienber, V., Freitag, M., Garcia, X., Gehrmann, S., Gonzalez, L., Gur-Ari, G., Hand, S., Hashemi, H., Hou, L., Howland, J., Hu, A., Hui, J., Hurwitz, J., Isard, M., Ittycheriah, A., Jagielski, M., Jia, W., Kenealy, K., Krikun, M., Kudugunta, S., Lan, C., Lee, K., Lee, B., Li, E., Li, M., Li, W., Li, Y., Li, J., Lim, H., Lin, H., Liu, Z., Liu, F., Maggioni, M., Mahendru, A., Maynez, J., Misra, V., Moussalem, M., Nado, Z., Nham, J., Ni, E., Nystrom, A., Parrish, A., Pellat, M., Polacek, M., Polozov, A., Pope, R., Qiao, S., Reif, E., Richter, B., Riley, P., Ros, A. C., Roy, A., Saeta, B., Samuel, R., Shelby, R., Slone, A., Smilkov, D., So, D. R., Sohn, D., Tokumine, S., Valter, D., Vasudevan, V., Vodrahalli, K., Wang, X., Wang, P., Wang, Z., Wang, T., Wieting, J., Wu, Y., Xu, K., Xu, Y., Xue, L., Yin, P., Yu, J., Zhang, Q., Zheng, S., Zheng, C., Zhou, W., Zhou, D., Petrov, S. & Wu, Y. (2023), ‘PaLM 2 Technical Report’
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Keysers, D., Schärli, N., Scales, N., Buisman, H., Furrer, D., Kashubin, S., Momchev, N., Sinopalnikov, D., Stafiniak, L., Tihon, T., Tsarkov, D., Wang, X., van Zee, M. & Bousquet, O. (2019), Measuring Compositional Generalization: A Comprehensive Method on Realistic Data, in
2019
Cited alongside, same era.
Tshitoyan, V., Dagdelen, J., Weston, L., Dunn, A., Rong, Z., Kononova, O., Persson, K. A., Ceder, G. & Jain, A. (2019), ‘Unsupervised word embeddings capture latent knowledge from materials science literature’, Nature
2019
Cited alongside, same era.
Wallace, E., Wang, Y., Li, S., Singh, S. & Gardner, M. (2019), Do NLP Models Know Numbers? Probing Numeracy in Embeddings, in
2019
Cited alongside, same era.
Akyürek, E., Akyürek, A. F. & Andreas, J. (2020), Learning to Recombine and Resample Data For Compositional Generalization, in
2020
Cited alongside, same era.
Andreas, J. (2020), Good-Enough Compositional Data Augmentation, in
2020
Cited alongside, same era.
Bender, E. M. & Koller, A. (2020), Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data, in
2020
Cited alongside, same era.
Boleda, G. (2020), ‘Distributional Semantics and Linguistic Theory’, Annual Review of Linguistics
2020
Cited alongside, same era.
2023
Later among the works it cites.
Betker, J., Goh, G., Jing, L., Brooks, T., Wang, J., Li, L., Ouyang, L., Zhuang, J., Lee, J., Guo, Y. et al. (2023), ‘Improving image generation with better captions’, Computer Science. https://cdn. openai. com/papers/dall-e-3. pdf
2023
Later among the works it cites.
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., Nori, H., Palangi, H., Ribeiro, M. T. & Zhang, Y. (2023), ‘Sparks of Artificial General Intelligence: Early experiments with GPT-4’
2023
Later among the works it cites.
Buckner, C. J. (2023), From Deep Learning to Rational Machines: What the History of Philosophy Can Teach Us about the Future of Artificial Intelligence
2023
Later among the works it cites.
Grynbaum, M. M. & Mac, R. (2023), ‘The Times Sues OpenAI and Microsoft Over A.I. Use of Copyrighted Work’, The New York Times
2023
Later among the works it cites.
He, Z., Xie, Z., Jha, R., Steck, H., Liang, D., Feng, Y., Majumder, B. P., Kallus, N. & Mcauley, J. (2023), Large Language Models as Zero-Shot Conversational Recommenders, in
2023
Later among the works it cites.
Herbold, S., Hautli-Janisz, A., Heuer, U., Kikteva, Z. & Trautsch, A. (2023), ‘A large-scale comparison of human-written versus ChatGPT-generated essays’, Scientific Reports
2023
Later among the works it cites.
Hupkes, D., Giulianelli, M., Dankers, V., Artetxe, M., Elazar, Y., Pimentel, T., Christodoulopoulos, C., Lasri, K., Saphra, N., Sinclair, A., Ulmer, D., Schottmann, F., Batsuren, K., Sun, K., Sinha, K., Khalatbari, L., Ryskina, M., Frieske, R., Cotterell, R. & Jin, Z. (2023), ‘A taxonomy and review of generalization research in NLP’, Nature Machine Intelligence
2023
Later among the works it cites.
Jones, C. & Bergen, B. (2023), ‘Does GPT-4 Pass the Turing Test?’
2023
Later among the works it cites.
Karhade, M. (2023), ‘GPT-4: 8 Models in One ; The Secret is Out’
2023
Later among the works it cites.
Kasneci, E., Sessler, K., Küchemann, S., Bannert, M., Dementieva, D., Fischer, F., Gasser, U., Groh, G., Günnemann, S., Hüllermeier, E., Krusche, S., Kutyniok, G., Michaeli, T., Nerdel, C., Pfeffer, J., Poquet, O., Sailer, M., Schmidt, A., Seidel, T., Stadler, M., Weller, J., Kuhn, J. & Kasneci, G. (2023), ‘ChatGPT for good? On opportunities and challenges of large language models for education’, Learning and Individual Differences
2023
Later among the works it cites.
Kheiri, K. & Karimi, H. (2023), ‘SentimentGPT: Exploiting GPT for Advanced Sentiment Analysis and its Departure from Current Machine Learning’
2023
Later among the works it cites.
Lake, B. M. & Baroni, M. (2023), ‘Human-like systematic generalization through a meta-learning neural network’, Nature
2023
Later among the works it cites.
Lavechin, M., Sy, Y., Titeux, H., Blandón, M. A. C., Räsänen, O., Bredin, H., Dupoux, E. & Cristia, A. (2023), ‘BabySLM: Language-acquisition-friendly benchmark of self-supervised spoken language models’
2023
Later among the works it cites.
Lee, N., Sreenivasan, K., Lee, J., Lee, K. & Papailiopoulos, D. (2023), Teaching Arithmetic to Small Transformers, in
2023
Later among the works it cites.
Liang, W., Zhang, Y., Cao, H., Wang, B., Ding, D., Yang, X., Vodrahalli, K., He, S., Smith, D., Yin, Y., McFarland, D. & Zou, J. (2023), ‘Can large language models provide useful feedback on research papers? A large-scale empirical analysis’
2023
Later among the works it cites.
Long, B., Goodin, S., Kachergis, G., Marchman, V. A., Radwan, S. F., Sparks, R. Z., Xiang, V., Zhuang, C., Hsu, O., Newman, B., Yamins, D. L. K. & Frank, M. C. (2023), ‘The BabyView camera: Designing a new head-mounted camera to capture children’s early social and visual environments’, Behavior Research Methods
2023
Later among the works it cites.
Mandelkern, M. & Linzen, T. (2023), ‘Do Language Models Refer?’
2023
Later among the works it cites.
McCoy, R. T., Yao, S., Friedman, D., Hardy, M. & Griffiths, T. L. (2023), ‘Embers of Autoregression: Understanding Large Language Models Through the Problem They are Trained to Solve’
2023
Later among the works it cites.
McGrath, S., Russin, J., Pavlick, E. & Feiman, R. (2023), ‘Properties of LoTs: The footprints or the bear itself?’
2023
Later among the works it cites.
Mirchandani, S., Xia, F., Florence, P., Ichter, B., Driess, D., Arenas, M. G., Rao, K., Sadigh, D. & Zeng, A. (2023), ‘Large Language Models as General Pattern Machines’
2023
Later among the works it cites.
Mirowski, P., Mathewson, K. W., Pittman, J. & Evans, R. (2023), Co-Writing Screenplays and Theatre Scripts with Language Models: Evaluation by Industry Professionals, in
2023
Later among the works it cites.
Mollo, D. C. & Millière, R. (2023), ‘The Vector Grounding Problem’
2023
Later among the works it cites.
Murty, S., Sharma, P., Andreas, J. & Manning, C. D. (2023), ‘Grokking of Hierarchical Structure in Vanilla Transformers’
2023
Later among the works it cites.
OpenAI (2023 a
2023
Later among the works it cites.
OpenAI (2023 b
2023
Later among the works it cites.
Pavlick, E. (2023), ‘Symbols and grounding in large language models’, Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences
2023
Later among the works it cites.
Piantadosi, S. (2023), ‘Modern language models refute Chomsky’s approach to language’
2023
Later among the works it cites.
Portelance, E. & Jasbi, M. (2023), ‘The roles of neural networks in language acquisitio’
2023
Later among the works it cites.
Savelka, J., Agarwal, A., An, M., Bogart, C. & Sakr, M. (2023), Thrilled by Your Progress! Large Language Models (GPT-4) No Longer Struggle to Pass Assessments in Higher Education Programming Courses, in
2023
Later among the works it cites.
Savelka, J., Ashley, K. D., Gray, M. A., Westermann, H. & Xu, H. (2023), Can GPT-4 Support Analysis of Textual Data in Tasks Requiring Highly Specialized Domain Expertise?, in
2023
Later among the works it cites.
Schut, L., Tomasev, N., McGrath, T., Hassabis, D., Paquet, U. & Kim, B. (2023), ‘Bridging the Human-AI Knowledge Gap: Concept Discovery and Transfer in AlphaZero’
2023
Later among the works it cites.
Shinn, N., Cassano, F., Berman, E., Gopinath, A., Narasimhan, K. & Yao, S. (2023), ‘Reflexion: Language Agents with Verbal Reinforcement Learning’
2023
Later among the works it cites.
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S. & Scialom, T. (2023), ‘Llama 2: Open Foundation and Fine-Tuned Chat Models’
2023
Later among the works it cites.
Wang, L., Lyu, C., Ji, T., Zhang, Z., Yu, D., Shi, S. & Tu, Z. (2023), ‘Document-Level Machine Translation with Large Language Models’
2023
Later among the works it cites.
Wang, R., Todd, G., Yuan, E., Xiao, Z., Côté, M.-A. & Jansen, P. (2023), ‘ByteSized32: A Corpus and Challenge Task for Generating Task-Specific World Models Expressed as Text Games’
2023
Later among the works it cites.
Warstadt, A., Mueller, A., Choshen, L., Wilcox, E., Zhuang, C., Ciro, J., Mosquera, R., Paranjabe, B., Williams, A., Linzen, T. & Cotterell, R. (2023), Findings of the BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora, in
2023
Later among the works it cites.
Zhang, T., Ladhak, F., Durmus, E., Liang, P., McKeown, K. & Hashimoto, T. B. (2023), ‘Benchmarking Large Language Models for News Summarization’
2023
Later among the works it cites.
Zhou, A., Wang, K., Lu, Z., Shi, W., Luo, S., Qin, Z., Lu, S., Jia, A., Song, L., Zhan, M. & Li, H. (2023), ‘Solving Challenging Math Word Problems Using GPT-4 Code Interpreter with Code-based Self-Verification’
2023
Later among the works it cites.