Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have demonstrated remarkable performance on various quantitative reasoning and knowledge benchmarks.
Problems and solutions on thermodynamics and Statistical Mechanics
Yung-kuo Lim · 1996
Earlier work this paper cites.
Problems and solutions in quantum mechanics: Major, American universities ph. D. qualifying questions and, solutions
Yung-kuo Lim · 1998
Earlier work this paper cites.
Problems and solutions on Mechanics
Yung-kuo Lim and Yuan-qi Qiang · 2001
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2005
Earlier work this paper cites.
Barbri Practice Questions: Multistate Testing Practice Questions
Barbri · 2007
Earlier work this paper cites.
Problems and solutions on electromagnetism
Yung-kuo Lim · 2007
Earlier work this paper cites.
Berkeley problems in Mathematics
Paulo N de Souza and Jorge N. Silva · 2008
Earlier work this paper cites.
Measuring massive multitask language understanding, 2020
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2009
Earlier work this paper cites.
MAWPS: A math word problem repository
Rik Koncel-Kedziorski, Subhro Roy, Aida Amini, Nate Kushman, and Hannaneh Hajishirzi · 2016
Earlier work this paper cites.
Solving general arithmetic word problems, 2016
Subhro Roy and Dan Roth · 2016
Earlier work this paper cites.
McGraw-Hill Education 3 MCAT Practice Tests, Third Edition
Candice McCloskey Campbell, Shaun Murphree, Jennifer M. Warner, Amy B. Wachholz, Kathy A. Zahler, and George J. Hademenos · 2017
Earlier work this paper cites.
Putnam and beyond
Răzvan Gelca and Titu Andreescu · 2017
Earlier work this paper cites.
Program induction by rationale generation: Learning, to solve and explain algebraic word problems
Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom · 2017
Earlier work this paper cites.
Sympy: Symbolic computing in python
Aaron Meurer, Christopher P. Smith, Mateusz Paprocki, Ondřej Čertík, Sergey B. Kirpichev, Matthew Rocklin, AMiT Kumar, Sergiu Ivanov, Jason K. Moore, Sartaj Singh, Thilina Rathnayake, Sean Vig, Brian E. Granger, Richard P. Muller, Francesco Bonazzi, Harsh Gupta, Shivam Vats, Fredrik Johansson, Fabian Pedregosa, Matthew J. Curry, Andy R. Terrel, Štěpán Roučka, Ashutosh Saboo, Isuru Fernando, Sumith Kulal, Robert Cimrman, and Anthony Scopatz · 2017
Earlier work this paper cites.
Undergraduate Mathematics Competitions (1995-2016): Taras Shevchenko National University of Kyiv
Volodymyr Brayman and A. G. Kukush · 2018
Earlier work this paper cites.
CommonsenseQA: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant · 2018
Earlier work this paper cites.
HotpotQA: A dataset for diverse, explainable multi-hop question answering, 2018
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning · 2018
Earlier work this paper cites.
On the measure of intelligence
François Chollet · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Problems and solutions on optics
Swee Cheng Lim, Choy Heng Lai, and Leong Chuan Kwek · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Scaling laws for neural language models, 2020
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
A diverse corpus for evaluating and developing English math word problem solvers
Shen-yun Miao, Chao-Chun Liang, and Keh-Yih Su · 2020
Earlier work this paper cites.
The dangers of underclaiming: Reasons for caution when reporting how NLP systems fail, 2021
Samuel R. Bowman · 2021
Cited alongside, same era.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Cited alongside, same era.
Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies, 2021
Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant · 2021
Cited alongside, same era.
Qualifying examination for fall 2021, Aug 2021
Department of Mathematics Harvard University · 2021
Cited alongside, same era.
Measuring mathematical problem solving with the MATH dataset
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt · 2021
Broken neural scaling laws, 2023
Ethan Caballero, Kshitij Gupta, Irina Rish, and David Krueger · 2023
Closest in time.
GPT retakes my midterm and gets an A, 2023
Bryan Caplan · 2023
Closest in time.
Can Large Language Models be an alternative to human evaluations?
Cheng-Han Chiang and Hung-yi Lee · 2023
Closest in time.
GPTScore: Evaluate as you desire
Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu · 2023
Closest in time.
Chain-of-thought hub: A continuous effort to measure large language models’ reasoning performance, 2023
Yao Fu, Litu Ou, Mingyu Chen, Yuhao Wan, Hao Peng, and Tushar Khot · 2023
Closest in time.
Large language models are not abstract reasoners, 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Show your work: Scratchpads for intermediate computation with language models, 2021
Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, Charles Sutton, and Augustus Odena · 2021
Cited alongside, same era.
Are NLP models really able to solve simple math word problems?, 2021
Arkil Patel, Satwik Bhattamishra, and Navin Goyal · 2021
Cited alongside, same era.
Revisiting neural scaling laws in language and vision
Ibrahim M Alabdulmohsin, Behnam Neyshabur, and Xiaohua Zhai · 2022
Cited alongside, same era.
Constitutional AI: Harmlessness from AI, feedback, 2022
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, Kamile Lukosuite, Liane Lovitt, Michael Sellitto, Nelson Elhage, Nicholas Schiefer, Noemi Mercado, Nova DasSarma, Robert Lasenby, Robin Larson, Sam Ringer, Scott Johnston, Shauna Kravec, Sheer El Showk, Stanislav Fort, Tamera Lanham, Timothy Telleen-Lawton, Tom Conerly, Tom Henighan, Tristan Hume, Samuel R. Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan · 2022
Cited alongside, same era.
Michael Bommarito II and Daniel Martin Katz · 2022
Cited alongside, same era.
PaLM: Scaling language modeling with Pathways, 2022
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel · 2022
Cited alongside, same era.
Training compute-optimal large language models, 2022
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, and Laurent Sifre · 2022
Cited alongside, same era.
Gaël Gendron, Qiming Bao, Michael Witbrock, and Gillian Dobbie · 2023
Closest in time.
Introducing PaLM 2, 2023
Zoubin Ghahramani · 2023
Closest in time.
GPT-4 passes the bar exam
Daniel Martin Katz, Michael James Bommarito, Shang Gao, and Pablo Arredondo · 2023
Closest in time.
Large language models are state-of-the-art evaluators of translation quality
Tom Kocmi and Christian Federmann · 2023
Closest in time.
Large language models are zero-shot reasoners, 2023
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa · 2023
Closest in time.
Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models
Tiffany H Kung, Morgan Cheatham, Arielle Medenilla, Czarina Sillos, Lorie De Leon, Camille Elepaño, Maria Madriaga, Rimel Aggabao, Giezel Diaz-Candido, James Maningo, et al · 2023
Closest in time.
A systematic study and comprehensive evaluation of ChatGPT on benchmark datasets, 2023
Md Tahmid Rahman Laskar, M Saiful Bari, Mizanur Rahman, Md Amran Hossen Bhuiyan, Shafiq Joty, and Jimmy Xiangji Huang · 2023
Closest in time.
Making large language models better reasoners with step-aware verifier, 2023
Yifei Li, Zeqi Lin, Shizhuo Zhang, Qiang Fu, Bei Chen, Jian-Guang Lou, and Weizhu Chen · 2023
Closest in time.
Let’s verify step by step, 2023
Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe · 2023
Closest in time.
G-eval: NLG evaluation using GPT-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu · 2023
Closest in time.
Experimental evidence on the productivity effects of generative artificial intelligence
Shakked Noy and Whitney Zhang · 2023
Closest in time.
GPT-4 technical report, 2023
OpenAI · 2023
Closest in time.
An independent evaluation of ChatGPT on mathematical word problems (MWP)
Paulo Shakarian, Abhinav Koyyalamudi, Noel Ngu, and Lakshmivihari Mareedu · 2023
Closest in time.
Clever Hans or Neural Theory of Mind? Stress testing social reasoning in large language models
Natalie Shapira, Mosh Levy, Seyed Hossein Alavi, Xuhui Zhou, Yejin Choi, Yoav Goldberg, Maarten Sap, and Vered Shwartz · 2023
Closest in time.
Large language models still can’t plan (a benchmark for LLMs on planning and reasoning about change), 2023
Karthik Valmeekam, Alberto Olmo, Sarath Sreedharan, and Subbarao Kambhampati · 2023
Closest in time.
Self-consistency improves chain of thought reasoning in language models, 2023
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou · 2023
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models, 2023
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan · 2023
Closest in time.