Fetching the paper…
Reading the bibliography…
Language models (LMs) can hallucinate when performing complex mathematical reasoning.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Analysing mathematical reasoning abilities of neural models
David Saxton, Edward Grefenstette, Felix Hill, and Pushmeet Kohli. 2019 · 1904
Earlier work this paper cites.
The use of deep learning for symbolic integration: A review of (lample and charton, 2019)
Ernest Davis. 2019 · 1912
Earlier work this paper cites.
Deep learning for symbolic mathematics
Guillaume Lample and François Charton. 2019 · 1912
Earlier work this paper cites.
Theory of the contribution of excitons to the complex dielectric constant of crystals
JJ Hopfield. 1958 · 1958
Earlier work this paper cites.
Transformers as soft reasoners over language
Peter Clark, Oyvind Tafjord, and Kyle Richardson. 2020 · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Gleu: Automatic evaluation of sentence-level fluency
Andrew Mutton, Mark Dras, Stephen Wan, and Robert Dale. 2007 · 2007
Earlier work this paper cites.
Causal inference in statistics: An overview
Judea Pearl. 2009 · 2009
Earlier work this paper cites.
Generative language modeling for automated theorem proving
Stanislas Polu and Ilya Sutskever. 2020 · 2009
Earlier work this paper cites.
Premise selection for mathematics by corpus analysis and kernel methods
Jesse Alama, Tom Heskes, Daniel Kühlwein, Evgeni Tsivtsivadze, and Josef Urban. 2014 · 2014
Earlier work this paper cites.
Formalizing physics: automation, presentation and foundation issues
Cezary Kaliszyk, Josef Urban, Umair Siddique, Sanaz Khan-Afshar, Cvetan Dunchev, and Sofiene Tahar. 2015 · 2015
Earlier work this paper cites.
Reasoning about quantities in natural language
Subhro Roy, Tim Vieira, and Dan Roth. 2015 · 2015
Earlier work this paper cites.
Premise selection for theorem proving by deep graph embedding
Mingzhe Wang, Yihe Tang, Jian Wang, and Jia Deng. 2017 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Automatic derivation of formulas using reforcement learning
MinZhong Luo and Li Liu. 2018 · 2018
Earlier work this paper cites.
Manipulating type-i and type-ii dirac polaritons in cavity-embedded honeycomb metasurfaces
Charlie-Ray Mann, Thomas J Sturges, Guillaume Weick, William L Barnes, and Eros Mariani. 2018 · 2018
Earlier work this paper cites.
Toward an artificial intelligence physicist for unsupervised learning
Tailin Wu and Max Tegmark. 2019 · 2019
Earlier work this paper cites.
Natural language premise selection: Finding supporting statements for mathematical text
Deborah Ferreira and André Freitas. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Earlier work this paper cites.
A promising path towards autoformalization and general artificial intelligence
Christian Szegedy. 2020 · 2020
Earlier work this paper cites.
Wei Zhao, Goran Glavaš, Maxime Peyrard, Yang Gao, Robert West, and Steffen Eger. 2020 · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021 · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021 · 2021
Earlier work this paper cites.
Language models of isabelle proofs
Albert Qiaochu Jiang, Wenda Li, Jesse Michael Han, and Yuhuai Wu Lisa. 2021 · 2021
Earlier work this paper cites.
Similarity-based equational inference in physics
Jordan Meadows and André Freitas. 2021 · 2021
Cited alongside, same era.
Naturalproofs: Mathematical theorem proving in natural language
Sean Welleck, Jiacheng Liu, Ronan Le Bras, Hannaneh Hajishirzi, Yejin Choi, and Kyunghyun Cho. 2021 · 2021
Cited alongside, same era.
Few-shot training llms for project-specific code-summarization
Toufique Ahmed and Premkumar Devanbu. 2022 · 2022
Cited alongside, same era.
A qubo formulation for top- τ \tau eigencentrality nodes
Prosper D Akrobotu, Tamsin E James, Christian FA Negre, and Susan M Mniszewski. 2022 · 2022
Cited alongside, same era.
Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W Cohen. 2022 · 2022
Cited alongside, same era.
Will we run out of data? an analysis of the limits of scaling datasets in machine learning
Pablo Villalobos, Jaime Sevilla, Lennart Heim, Tamay Besiroglu, Marius Hobbhahn, and Anson Ho. 2022 · 2022
Later among the works it cites.
Autoformalization with large language models
Yuhuai Wu, Albert Q. Jiang, Wenda Li, Markus N. Rabe, Charles Staats, Mateja Jamnik, and Christian Szegedy. 2022 · 2022
Later among the works it cites.
Llemma: An open language model for mathematics
Zhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos, Stephen McAleer, Albert Q Jiang, Jia Deng, Stella Biderman, and Sean Welleck. 2023 · 2023
Later among the works it cites.
Baldur: Whole-proof generation and repair with large language models
Emily First, Markus N Rabe, Talia Ringer, and Yuriy Brun. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022 · 2022
Cited alongside, same era.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022 · 2022
Cited alongside, same era.
On the limitations of reference-free evaluations of generated text
Daniel Deutsch, Rotem Dror, and Dan Roth. 2022 · 2022
Cited alongside, same era.
A neural network solves, explains, and generates university math problems by program synthesis and few-shot learning at human level
Iddo Drori, Sarah Zhang, Reece Shuttleworth, Leonard Tang, Albert Lu, Elizabeth Ke, Kevin Liu, Linda Chen, Sunny Tran, Newman Cheng, et al. 2022 · 2022
Cited alongside, same era.
Physics-informed neural networks for solving reynolds-averaged navier–stokes equations
Hamidreza Eivazi, Mojtaba Tahani, Philipp Schlatter, and Ricardo Vinuesa. 2022 · 2022
Cited alongside, same era.
To be or not to be an integer? encoding variables for mathematical text
Deborah Ferreira, Mokanarangan Thayaparan, Marco Valentino, Julia Rozanova, and Andre Freitas. 2022 · 2022
Cited alongside, same era.
Enhancing neural mathematical reasoning by abductive combination with symbolic library
Yangyang Hu and Yang Yu. 2022 · 2022
Cited alongside, same era.
Simon Frieder, Luca Pinchetti, Ryan-Rhys Griffiths, Tommaso Salvatori, Thomas Lukasiewicz, Philipp Christian Petersen, Alexis Chevalier, and Julius Berner. 2023 · 2023
Later among the works it cites.
Openagi: When llm meets domain experts
Yingqiang Ge, Wenyue Hua, Jianchao Ji, Juntao Tan, Shuyuan Xu, and Yongfeng Zhang. 2023 · 2023
Later among the works it cites.
Solving math word problems by combining language models with symbolic solvers
Joy He-Yueya, Gabriel Poesia, Rose E. Wang, and Noah D. Goodman. 2023 · 2023
Later among the works it cites.
Llm-adapters: An adapter family for parameter-efficient fine-tuning of large language models
Zhiqiang Hu, Yihuai Lan, Lei Wang, Wanyu Xu, Ee-Peng Lim, Roy Ka-Wei Lee, Lidong Bing, and Soujanya Poria. 2023 · 2023
Later among the works it cites.
Evaluating the logical reasoning ability of chatgpt and gpt-4
Hanmeng Liu, Ruoxi Ning, Zhiyang Teng, Jian Liu, Qiji Zhou, and Yue Zhang. 2023 · 2023
Later among the works it cites.
Introduction to mathematical language processing: Informal proofs, word problems, and supporting tasks
Jordan Meadows and André Freitas. 2023 · 2023
Later among the works it cites.
An independent evaluation of chatgpt on mathematical word problems (mwp)
Paulo Shakarian, Abhinav Koyyalamudi, Noel Ngu, and Lakshmivihari Mareedu. 2023 · 2023
Later among the works it cites.
Multi-operational mathematical derivations in latent space
Marco Valentino, Jordan Meadows, Lan Zhang, and André Freitas. 2023 · 2023
Later among the works it cites.
Learning multi-step reasoning by solving arithmetic tasks
Tianduo Wang and Wei Lu. 2023 · 2023
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. 2023 · 2023
Later among the works it cites.
Harnessing the power of llms in practice: A survey on chatgpt and beyond
Jingfeng Yang, Hongye Jin, Ruixiang Tang, Xiaotian Han, Qizhang Feng, Haoming Jiang, Bing Yin, and Xia Hu. 2023 · 2023
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L Griffiths, Yuan Cao, and Karthik Narasimhan. 2023 · 2023
Later among the works it cites.
How well do large language models perform in arithmetic tasks?
Zheng Yuan, Hongyi Yuan, Chuanqi Tan, Wei Wang, and Songfang Huang. 2023 · 2023
Later among the works it cites.
Premise order matters in reasoning with large language models
Xinyun Chen, Ryan A Chi, Xuezhi Wang, and Denny Zhou. 2024 · 2024
Closest in time.
Mathematics, word problems, common sense, and artificial intelligence
Ernest Davis. 2024 · 2024
Closest in time.
Augmenting math word problems via iterative question composing
Haoxiong Liu and Andrew Chi-Chih Yao. 2024 · 2024
Closest in time.
Quantum many-body physics calculations with large language models
Haining Pan, Nayantara Mudur, Will Taranto, Maria Tikhanovskaya, Subhashini Venugopalan, Yasaman Bahri, Michael P Brenner, and Eun-Ah Kim. 2024 · 2024
Closest in time.
Verification and refinement of natural language explanations through llm-symbolic theorem proving
Xin Quan, Marco Valentino, Louise A. Dennis, and André Freitas. 2024 · 2024
Closest in time.
Openmathinstruct-1: A 1.8 million math instruction tuning dataset
Shubham Toshniwal, Ivan Moshkov, Sean Narenthiran, Daria Gitman, Fei Jia, and Igor Gitman. 2024 · 2024
Closest in time.
Solving olympiad geometry without human demonstrations
Trieu H Trinh, Yuhuai Wu, Quoc V Le, He He, and Thang Luong. 2024 · 2024
Closest in time.
Leandojo: Theorem proving with retrieval-augmented language models
Kaiyu Yang, Aidan Swope, Alex Gu, Rahul Chalamala, Peiyang Song, Shixing Yu, Saad Godil, Ryan J Prenger, and Animashree Anandkumar. 2024 · 2024
Closest in time.