Fetching the paper…
Reading the bibliography…
This paper proposes a methodology for generating and perturbing detailed derivations of equations at scale, aided by a symbolic engine, to evaluate the generalisability of Transformers to out-of-distribution mathematical reasoning problems.
Scibert: A pretrained language model for scientific text
Iz Beltagy, Kyle Lo, and Arman Cohan. 2019 · 1903
Earlier work this paper cites.
Analysing mathematical reasoning abilities of neural models
David Saxton, Edward Grefenstette, Felix Hill, and Pushmeet Kohli. 2019 · 1904
Earlier work this paper cites.
Learning the difference that makes a difference with counterfactually-augmented data
Divyansh Kaushik, Eduard Hovy, and Zachary C Lipton. 2019 · 1909
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al. 2019 · 1910
Earlier work this paper cites.
The use of deep learning for symbolic integration: A review of (lample and charton, 2019)
Ernest Davis. 2019 · 1912
Earlier work this paper cites.
Deep learning for symbolic mathematics
Guillaume Lample and François Charton. 2019 · 1912
Earlier work this paper cites.
The MATHEMATICA® book, version 4
Stephen Wolfram. 1999 · 1999
Earlier work this paper cites.
Transformers as soft reasoners over language
Peter Clark, Oyvind Tafjord, and Kyle Richardson. 2020 · 2002
Earlier work this paper cites.
Isarstep: a benchmark for high-level mathematical reasoning
Wenda Li, Lei Yu, Yuhuai Wu, and Lawrence C Paulson. 2020 · 2006
Earlier work this paper cites.
On the stability of fine-tuning bert: Misconceptions, explanations, and strong baselines
Marius Mosbach, Maksym Andriushchenko, and Dietrich Klakow. 2020 · 2006
Earlier work this paper cites.
Mathematical reasoning via self-supervised skip-tree training
Markus N Rabe, Dennis Lee, Kshitij Bansal, and Christian Szegedy. 2020 · 2006
Earlier work this paper cites.
Int: An inequality benchmark for evaluating generalization in theorem proving
Yuhuai Wu, Albert Qiaochu Jiang, Jimmy Ba, and Roger Grosse. 2020 · 2007
Earlier work this paper cites.
Causal inference in statistics: An overview
Judea Pearl. 2009 · 2009
Earlier work this paper cites.
Generative language modeling for automated theorem proving
Stanislas Polu and Ilya Sutskever. 2020 · 2009
Earlier work this paper cites.
The lean theorem prover (system description)
Leonardo de Moura, Soonho Kong, Jeremy Avigad, Floris van Doorn, and Jakob von Raumer. 2015 · 2015
Earlier work this paper cites.
Sympy: symbolic computing in python
Aaron Meurer, Christopher P Smith, Mateusz Paprocki, Ondřej Čertík, Sergey B Kirpichev, Matthew Rocklin, AMiT Kumar, Sergiu Ivanov, Jason K Moore, Sartaj Singh, et al. 2017 · 2017
Earlier work this paper cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Tangent-cft: An embedding model for mathematical formulas
Behrooz Mansouri, Shaurya Rohatgi, Douglas W Oard, Jian Wu, C Lee Giles, and Richard Zanibbi. 2019 · 2019
Earlier work this paper cites.
String correction using the damerau-levenshtein distance
Chunchun Zhao and Sartaj Sahni. 2019 · 2019
Earlier work this paper cites.
Evaluating token-level and passage-level dense retrieval models for math information retrieval
Wei Zhong, Jheng-Hong Yang, and Jimmy Lin. 2022 · 2019
Earlier work this paper cites.
Adversarial machine learning-industry perspectives
Ram Shankar Siva Kumar, Magnus Nyström, John Lambert, Andrew Marshall, Mario Goertzel, Andi Comissoneru, Matt Swann, and Sharon Xia. 2020 · 2020
Earlier work this paper cites.
Adversarial nli: A new benchmark for natural language understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2020 · 2020
Earlier work this paper cites.
Beyond accuracy: Behavioral testing of NLP models with CheckList
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020 · 2020
Cited alongside, same era.
On the value of out-of-distribution testing: An example of goodhart’s law
Damien Teney, Ehsan Abbasnejad, Kushal Kafle, Robik Shrestha, Christopher Kanan, and Anton Van Den Hengel. 2020 · 2020
Cited alongside, same era.
Adapting text embeddings for causal inference
Victor Veitch, Dhanya Sridhar, and David Blei. 2020 · 2020
Cited alongside, same era.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021 · 2021
Cited alongside, same era.
A causal framework for distribution generalization
Rune Christiansen, Niklas Pfister, Martin Emil Jakobsen, Nicola Gnecco, and Jonas Peters. 2021 · 2021
Cited alongside, same era.
Hybrid tokenization and datasets for solving mathematics and science problems using transformers
Pratik Mandlecha, Snehith Kumar Chatakonda, Neeraj Kollepara, and Pawan Kumar. 2022 · 2022
Later among the works it cites.
Jordan Meadows, Zili Zhou, and Andre Freitas. 2022 · 2022
Later among the works it cites.
Combining sparse and dense information retrieval
Vít Novotnỳ and Michal Štefánik. 2022 · 2022
Later among the works it cites.
Direct and indirect effects
Judea Pearl. 2022 · 2022
Later among the works it cites.
Transformer-encoder and decoder models for questions on math
Anja Reusch, Maik Thiele, and Wolfgang Lehner. 2022 · 2022
Later among the works it cites.
A survey of adversarial defences and robustness in nlp
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Amnesic probing: Behavioral explanation with amnesic counterfactuals
Yanai Elazar, Shauli Ravfogel, Alon Jacovi, and Yoav Goldberg. 2021 · 2021
Cited alongside, same era.
Similarity-based equational inference in physics
Jordan Meadows and André Freitas. 2021 · 2021
Cited alongside, same era.
Are nlp models really able to solve simple math word problems?
Arkil Patel, Satwik Bhattamishra, and Navin Goyal. 2021 · 2021
Cited alongside, same era.
Mathbert: A pre-trained model for mathematical formula understanding
Shuai Peng, Ke Yuan, Liangcai Gao, and Zhi Tang. 2021 · 2021
Cited alongside, same era.
Probing the probing paradigm: Does probing accuracy entail task relevance?
Abhilasha Ravichander, Yonatan Belinkov, and Eduard Hovy. 2021 · 2021
Cited alongside, same era.
Mathbert: A pre-trained language model for general nlp tasks in mathematics education
Jia Tracy Shen, Michiharu Yamashita, Ethan Prihar, Neil Heffernan, Xintao Wu, Ben Graff, and Dongwon Lee. 2021 · 2021
Cited alongside, same era.
Towards grounded natural language proof generation
Sean Welleck, Jiacheng Liu, Jesse Michael Han, and Yejin Choi. 2021 · 2021
Cited alongside, same era.
Goyal Shreya, Sumanth Doddapaneni, Mitesh M Khapra, and Balaraman Ravindran. 2022 · 2022
Later among the works it cites.
A causal framework to quantify the robustness of mathematical reasoning with language models
Alessandro Stolfo, Zhijing Jin, Kumar Shridhar, Bernhard Schölkopf, and Mrinmaya Sachan. 2022 · 2022
Later among the works it cites.
Ijs at textgraphs-16 natural language premise selection task: Will contextual information improve natural language premise selection?
Thi Hong Hanh Tran, Matej Martinc, Antoine Doucet, and Senja Pollak. 2022 · 2022
Later among the works it cites.
TextGraphs 2022 shared task on natural language premise selection
Marco Valentino, Deborah Ferreira, Mokanarangan Thayaparan, André Freitas, and Dmitry Ustalov. 2022a · 2022
Later among the works it cites.
Will we run out of data? an analysis of the limits of scaling datasets in machine learning
Pablo Villalobos, Jaime Sevilla, Lennart Heim, Tamay Besiroglu, Marius Hobbhahn, and Anson Ho. 2022 · 2022
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022 · 2022
Later among the works it cites.
Symbolic brittleness in sequence models: on systematic generalization in symbolic mathematics
Sean Welleck, Peter West, Jize Cao, and Yejin Choi. 2022 · 2022
Later among the works it cites.
Llemma: An open language model for mathematics
Zhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos, Stephen McAleer, Albert Q Jiang, Jia Deng, Stella Biderman, and Sean Welleck. 2023 · 2023
Closest in time.
Mathematical capabilities of chatgpt
Simon Frieder, Luca Pinchetti, Ryan-Rhys Griffiths, Tommaso Salvatori, Thomas Lukasiewicz, Philipp Christian Petersen, Alexis Chevalier, and Julius Berner. 2023 · 2023
Closest in time.
Bert is not the count: Learning to match mathematical statements with proofs
Weixian Waylon Li, Yftah Ziser, Maximin Coavoux, and Shay B Cohen. 2023 · 2023
Closest in time.
Algebra error classification with large language models
Hunter McNichols, Mengxue Zhang, and Andrew Lan. 2023 · 2023
Closest in time.
Introduction to mathematical language processing: Informal proofs, word problems, and supporting tasks
Jordan Meadows and André Freitas. 2023 · 2023
Closest in time.
Generating mathematical derivations with large language models
Jordan Meadows, Marco Valentino, and Andre Freitas. 2023 · 2023
Closest in time.
Interventional probing in high dimensions: An nli case study
Julia Rozanova, Marco Valentino, Lucas Cordeiro, and André Freitas. 2023a · 2023
Closest in time.
A survey of methods for revealing and overcoming weaknesses of data-driven natural language understanding
Viktor Schlegel, Goran Nenadic, and Riza Batista-Navarro. 2023 · 2023
Closest in time.
Multi-operational mathematical derivations in latent space
Marco Valentino, Jordan Meadows, Lan Zhang, and André Freitas. 2023 · 2023
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L Griffiths, Yuan Cao, and Karthik Narasimhan. 2023 · 2023
Closest in time.