Fetching the paper…
Reading the bibliography…
This paper investigates how hallucination rates in Large Language Models (LLMs) may be controlled via a symbolic data generation framework, exploring a fundamental relationship between the rate of certain mathematical errors and types of input intervention.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020 · 1910
Earlier work this paper cites.
Equational reasoning and term rewriting systems
Plaisted, D. A. 1993 · 1993
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K.; Roukos, S.; Ward, T.; and Zhu, W.-J. 2002 · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y. 2004 · 2004
Earlier work this paper cites.
BLEURT: Learning robust metrics for text generation
Sellam, T.; Das, D.; and Parikh, A. P. 2020 · 2004
Earlier work this paper cites.
GLEU: Automatic evaluation of sentence-level fluency
Mutton, A.; Dras, M.; Wan, S.; and Dale, R. 2007 · 2007
Earlier work this paper cites.
Generative language modeling for automated theorem proving
Polu, S.; and Sutskever, I. 2020 · 2009
Earlier work this paper cites.
NTCIR-12 MathIR Task Overview
Zanibbi, R.; Aizawa, A.; Kohlhase, M.; Ounis, I.; Topic, G.; and Davila, K. 2016 · 2016
Earlier work this paper cites.
Manipulating type-I and type-II Dirac polaritons in cavity-embedded honeycomb metasurfaces
Mann, C.-R.; Sturges, T. J.; Weick, G.; Barnes, W. L.; and Mariani, E. 2018 · 2018
Earlier work this paper cites.
Semantic code search via equational reasoning
Premtoon, V.; Koppel, J.; and Solar-Lezama, A. 2020 · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M.; Tworek, J.; Jun, H.; Yuan, Q.; Pinto, H. P. d. O.; Kaplan, J.; Edwards, H.; Burda, Y.; Joseph, N.; Brockman, G.; et al. 2021 · 2021
Earlier work this paper cites.
Language models of isabelle proofs
Jiang, A. Q.; Li, W.; Han, J. M.; and Lisa, Y. W. 2021 · 2021
Earlier work this paper cites.
Retrieval augmentation reduces hallucination in conversation
Shuster, K.; Poff, S.; Chen, M.; Kiela, D.; and Weston, J. 2021 · 2021
Earlier work this paper cites.
Naturalproofs: Mathematical theorem proving in natural language
Welleck, S.; Liu, J.; Bras, R. L.; Hajishirzi, H.; Choi, Y.; and Cho, K. 2021 · 2021
Earlier work this paper cites.
Few-shot training LLMs for project-specific code-summarization
Ahmed, T.; and Devanbu, P. 2022 · 2022
Earlier work this paper cites.
A QUBO formulation for top- τ \tau eigencentrality nodes
Akrobotu, P. D.; James, T. E.; Negre, C. F.; and Mniszewski, S. M. 2022 · 2022
Earlier work this paper cites.
ATTEMPT: Parameter-Efficient Multi-task Tuning via Attentional Mixtures of Soft Prompts
Asai, A.; Salehi, M.; Peters, M.; and Hajishirzi, H. 2022 · 2022
Earlier work this paper cites.
Palm: Scaling language modeling with pathways
Chowdhery, A.; Narang, S.; Devlin, J.; Bosma, M.; Mishra, G.; Roberts, A.; Barham, P.; Chung, H. W.; Sutton, C.; Gehrmann, S.; et al. 2022 · 2022
Earlier work this paper cites.
A neural network solves, explains, and generates university math problems by program synthesis and few-shot learning at human level
Drori, I.; Zhang, S.; Shuttleworth, R.; Tang, L.; Lu, A.; Ke, E.; Liu, K.; Chen, L.; Tran, S.; Cheng, N.; et al. 2022 · 2022
Cited alongside, same era.
Draft, sketch, and prove: Guiding formal theorem provers with informal proofs
Jiang, A. Q.; Welleck, S.; Zhou, J. P.; Li, W.; Liu, J.; Jamnik, M.; Lacroix, T.; Wu, Y.; and Lample, G. 2022 · 2022
Cited alongside, same era.
Solving quantitative reasoning problems with language models
Lewkowycz, A.; Andreassen, A.; Dohan, D.; Dyer, E.; Michalewski, H.; Ramasesh, V.; Slone, A.; Anil, C.; Schlag, I.; Gutman-Solo, T.; et al. 2022 · 2022
Cited alongside, same era.
A survey of deep learning for mathematical reasoning
Lu, P.; Qiu, L.; Yu, W.; Welleck, S.; and Chang, K.-W. 2022 · 2022
Cited alongside, same era.
Llama: Open and efficient foundation language models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023 · 2023
Closest in time.
Wysocka, M.; Wysocki, O.; Delmas, M.; Mutel, V.; and Freitas, A. 2023 · 2023
Closest in time.
Harnessing the power of llms in practice: A survey on chatgpt and beyond
Yang, J.; Jin, H.; Tang, R.; Han, X.; Feng, Q.; Jiang, H.; Yin, B.; and Hu, X. 2023 · 2023
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S.; Yu, D.; Zhao, J.; Shafran, I.; Griffiths, T. L.; Cao, Y.; and Narasimhan, K. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Polu, S.; Han, J. M.; Zheng, K.; Baksys, M.; Babuschkin, I.; and Sutskever, I. 2022 · 2022
Cited alongside, same era.
Llm-planner: Few-shot grounded planning for embodied agents with large language models
Song, C. H.; Wu, J.; Washington, C.; Sadler, B. M.; Chao, W.-L.; and Su, Y. 2022 · 2022
Cited alongside, same era.
Galactica: A large language model for science
Taylor, R.; Kardas, M.; Cucurull, G.; Scialom, T.; Hartshorn, A.; Saravia, E.; Poulton, A.; Kerkez, V.; and Stojnic, R. 2022 · 2022
Cited alongside, same era.
Will we run out of data? An analysis of the limits of scaling datasets in Machine Learning
Villalobos, P.; Sevilla, J.; Heim, L.; Besiroglu, T.; Hobbhahn, M.; and Ho, A. 2022 · 2022
Cited alongside, same era.
Naturalprover: Grounded mathematical proof generation with language models
Welleck, S.; Liu, J.; Lu, X.; Hajishirzi, H.; and Choi, Y. 2022 · 2022
Cited alongside, same era.
Evaluating Token-Level and Passage-Level Dense Retrieval Models for Math Information Retrieval
Zhong, W.; Yang, J.-H.; and Lin, J. 2022 · 2022
Cited alongside, same era.
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
Cited alongside, same era.
Llemma: An open language model for mathematics
Azerbayev, Z.; Schoelkopf, H.; Paster, K.; Santos, M. D.; McAleer, S.; Jiang, A. Q.; Deng, J.; Biderman, S.; and Welleck, S. 2023 · 2023
Cited alongside, same era.
Yuan, Z.; Yuan, H.; Tan, C.; Wang, W.; and Huang, S. 2023 · 2023
Closest in time.
Premise Order Matters in Reasoning with Large Language Models
Chen, X.; Chi, R. A.; Wang, X.; and Zhou, D. 2024 · 2024
Closest in time.
Scaling instruction-finetuned language models
Chung, H. W.; Hou, L.; Longpre, S.; Zoph, B.; Tay, Y.; Fedus, W.; Li, Y.; Wang, X.; Dehghani, M.; Brahma, S.; et al. 2024 · 2024
Closest in time.
Dubey, A.; Jauhri, A.; Pandey, A.; Kadian, A.; Al-Dahle, A.; Letman, A.; Mathur, A.; Schelten, A.; Yang, A.; Fan, A.; et al. 2024 · 2024
Closest in time.
Augmenting math word problems via iterative question composing
Liu, H.; and Yao, A. C.-C. 2024 · 2024
Closest in time.
Exploring the Limits of Fine-grained LLM-based Physics Inference via Premise Removal Interventions
Meadows, J.; James, T.; and Freitas, A. 2024 · 2024
Closest in time.
A Symbolic Framework for Evaluating Mathematical Reasoning and Generalisation with Transformers
Meadows, J.; Valentino, M.; Teney, D.; and Freitas, A. 2024 · 2024
Closest in time.
Introducing meta llama 3: The most capable openly available llm to date
Meta, A. 2024 · 2024
Closest in time.
Gsm-symbolic: Understanding the limitations of mathematical reasoning in large language models
Mirzadeh, I.; Alizadeh, K.; Shahrokhi, H.; Tuzel, O.; Bengio, S.; and Farajtabar, M. 2024 · 2024
Closest in time.
Verification and Refinement of Natural Language Explanations through LLM-Symbolic Theorem Proving
Quan, X.; Valentino, M.; Dennis, L. A.; and Freitas, A. 2024 · 2024
Closest in time.
OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset
Toshniwal, S.; Moshkov, I.; Narenthiran, S.; Gitman, D.; Jia, F.; and Gitman, I. 2024 · 2024
Closest in time.
Solving olympiad geometry without human demonstrations
Trinh, T. H.; Wu, Y.; Le, Q. V.; He, H.; and Luong, T. 2024 · 2024
Closest in time.
On the Nature of Explanation: An Epistemological-Linguistic Perspective for Explanation-Based Natural Language Inference
Valentino, M.; and Freitas, A. 2024 · 2024
Closest in time.
Multi-Operational Mathematical Derivations in Latent Space
Valentino, M.; Meadows, J.; Zhang, L.; and Freitas, A. 2024 · 2024
Closest in time.