Fetching the paper…
Reading the bibliography…
We developed a benchmark set to assess the generalization of state-of-the-art large language models on problems beyond linguistic tasks and evaluate it on a systematic progression of GPT models (GPT-3.5, GPT-4, GPT-4o, GPT-4o-mini).
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. (2020) · 1901
Earlier work this paper cites.
Am. J. Psychol. 15 , 201–292
Spearman, C. (1904). ”general intelligence,” objectively determined and measured · 1904
Earlier work this paper cites.
The Measurement of Adult Intelligence
Wechsler, D · 1944
Earlier work this paper cites.
Mind LIX , 433–460. URL: https://doi.org/10.1093/mind/LIX.236.433
Turing, A. M. (1950). Computing machinery and intelligence · 1950
Earlier work this paper cites.
Syntactic Structures
Chomsky, N · 1957
Earlier work this paper cites.
The Development of Intelligence in Children ( 81–111)
Binet, A., and Simon, T · 1961
Earlier work this paper cites.
Journal of Educational Psychology 54 , 1–22. doi: 10.1037/h0046743
Cattell, R. B. (1963). Theory of fluid and crystallized intelligence: A critical experiment · 1963
Earlier work this paper cites.
J. Chem. Inf. Comput. Sci. 28 , 31–36. URL: https://doi.org/10.1021/ci00057a005
Weininger, D. (1988). Smiles, a chemical language and information system. 1. introduction to methodology and encoding rules · 1988
Earlier work this paper cites.
On Language: The Diversity of Human Language-Structure and its Influence on the Mental Development of Mankind
Humboldt, W · 1988
Earlier work this paper cites.
Human Cognitive Abilities: A Survey of Factor-Analytic Studies
Carroll, J. B · 1993
Earlier work this paper cites.
Intelligence 24 , 79–132. URL: https://www.sciencedirect.com/science/article/pii/S0160289697900143 . doi: https://doi.org/10.1016/S0160-2896(97)90014-3
Gottfredson, L. S. (1997). Why g matters: The complexity of everyday life · 1997
Earlier work this paper cites.
The g factor: The science of mental ability
Jensen, A · 1998
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P. (2016) · 2016
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I. (2017) · 2017
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. (2018) · 2018
Earlier work this paper cites.
HellaSwag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y. (2019) · 2019
Earlier work this paper cites.
Adversarial NLI: A new benchmark for natural language understanding
Nie, Y., Williams, A., Dinan, E., Bansal, M., Weston, J., and Kiela, D. (2020) · 2020
Earlier work this paper cites.
Minds and Machines 30 , 681 – 694. URL: https://api.semanticscholar.org/CorpusID:228954221
Floridi, L., and Chiriatti, M. (2020). Gpt-3: Its nature, scope, limits, and consequences · 2020
Earlier work this paper cites.
Language (technology) is power: A critical survey of “bias” in NLP
Blodgett, S. L., Barocas, S., Daumé III, H., and Wallach, H. (2020) · 2020
Earlier work this paper cites.
Analysis of minimax algorithm using tic-tac-toe
Swaminathan, B., Vaishali, R., and Subashri, T. (2020) · 2020
Earlier work this paper cites.
J. Game Theory 9 , 1–7. doi: 10.5923/j.jgt.20200901.01
Alkaraz, S. H., El-Seidy, E., and Morcos, N. S. (2020). Tic-tac-toe: Understanding the minimax algorithm · 2020
Cited alongside, same era.
On the dangers of stochastic parrots: Can language models be too big?
Bender, E. M., Gebru, T., McMillan-Major, A., and Shmitchell, S. (2021a) · 2021
Cited alongside, same era.
On the dangers of stochastic parrots: Can language models be too big?
Bender, E. M., Gebru, T., McMillan-Major, A., and Shmitchell, S. (2021b) · 2021
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P. F., Leike, J., and Lowe, R. (2022) · 2022
Cited alongside, same era.
TruthfulQA: Measuring how models mimic human falsehoods
Lin, S., Hilton, J., and Evans, O. (2022) · 2022
Cited alongside, same era.
Formal languages and neural models for learning on sequences
Merrill, W. (2023) · 2023
Later among the works it cites.
Nature 624 , 570–578
Boiko, D. A., MacKnight, R., Kline, B., and Gomes, G. (2023). Autonomous chemical research with large language models · 2023
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T., Cao, Y., and Narasimhan, K. (2023) · 2023
Later among the works it cites.
IEEE Access 12 , 6518–6531. doi: 10.1109/ACCESS.2024.3349952
Fields, J., Chovanec, K., and Madiraju, P. (2024). A survey of text classification with transformers: How wide? how large? how long? how accurate? how expensive? how safe? · 2024
Closest in time.
Electronics 13 , 1532. doi: 10.3390/electronics13081532
Topsakal, O., and Harper, J. (2024). Benchmarking large language model (llm) performance for game playing via tic-tac-toe · 2024
Closest in time.
ARC Prize 2024: ARC-AGI Competition
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., ichter, b., Xia, F., Chi, E., Le, Q. V., and Zhou, D. (2022) · 2022
Cited alongside, same era.
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B. (2023) · 2023
Cited alongside, same era.
arXiv. URL: https://doi.org/10.48550/arXiv.2302.13971
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G. (2023). Llama: Open and efficient foundation language models · 2023
Cited alongside, same era.
J. Mach. Learn. Res. 24
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvyashchenko, S., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N., Prabhakaran, V., Reif, E., Du, N., Hutchinson, B., Pope, R., Bradbury, J., Austin, J., Isard, M., Gur-Ari, G., Yin, P., Duke, T., Levskaya, A., Ghemawat, S., Dev, S., Michalewski, H., Garcia, X., Misra, V., Robinson, K., Fedus, L., Zhou, D., Ippolito, D., Luan, D., Lim, H., Zoph, B., Spiridonov, A., Sepassi, R., Dohan, D., Agrawal, S., Omernick, M., Dai, A. M., Pillai, T. S., Pellat, M., Lewkowycz, A., Moreira, E., Child, R., Polozov, O., Lee, K., Zhou, Z., Wang, X., Saeta, B., Diaz, M., Firat, O., Catasta, M., Wei, J., Meier-Hellstern, K., Eck, D., Dean, J., Petrov, S., and Fiedel, N. (2023). Palm: scaling language modeling with pathways · 2023
Cited alongside, same era.
arXiv. URL: https://doi.org/10.48550/arXiv.2303.12712
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S. M., Nori, H., Palangi, H., Ribeiro, M. T., and Zhang, Y. (2023). Sparks of artificial general intelligence: Early experiments with GPT-4 · 2023
Cited alongside, same era.
A framework for few-shot language model evaluation
Gao, L., Tow, J., Abbasi, B., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., Le Noac’h, A., Li, H., McDonell, K., Muennighoff, N., Ociepa, C., Phang, J., Reynolds, L., Schoelkopf, H., Skowron, A., Sutawika, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A. (2023) · 2023
Cited alongside, same era.
Testing spatial reasoning of large language models: the case of tic-tac-toe
Liga, D., and Pasetto, L. (2023) · 2023
Cited alongside, same era.
Infinite MonkeyLab42 (2024) · 2024
Closest in time.
Nature Machine Intelligence 6 , 161–169
Jablonka, K. M., Schwaller, P., Ortega-Guerrero, A., and Smit, B. (2024). Leveraging large language models for predictive chemistry · 2024
Closest in time.
microsoft/phi-2
Microsoft (2024a) · 2024
Closest in time.
Jackfram/llama-68m
JackFram (2024) · 2024
Closest in time.
openai-community/gpt2-medium
OpenAI (2024a) · 2024
Closest in time.
sshleifer/tiny-gpt2
Shleifer, S. (2024) · 2024
Closest in time.
Tinyllama/tinyllama-1.1b-chat-v1.0
TinyLlama (2024) · 2024
Closest in time.
mistralai/mixtral-8x7b-instruct-v0.1
Mistralai (2024) · 2024
Closest in time.
microsoft/dialogpt-medium
Microsoft (2024b) · 2024
Closest in time.
microsoft/phi-3-mini-4k-instruct
Microsoft (2024c) · 2024
Closest in time.
distilbert/distilgpt2
Face, H. (2024) · 2024
Closest in time.
openai-community/gpt2
OpenAI (2024b) · 2024
Closest in time.
Falcon-7b-instruct
UAE, T. (2024) · 2024
Closest in time.
Childplay github repository
(2025) · 2025
Closest in time.