Fetching the paper…
Reading the bibliography…
How can we trust the correctness of a learned model on a particular input of interest? Model accuracy is typically measured on average over a distribution of inputs, giving no guarantee for any fixed input.
The Art of Computer Programming, Volume II: Seminumerical Algorithms
Donald E. Knuth · 1969
Earlier work this paper cites.
A theory of the learnable
Leslie G. Valiant · 1972
Earlier work this paper cites.
The knowledge complexity of interactive proof-systems (extended abstract)
Shafi Goldwasser, Silvio Micali, and Charles Rackoff · 1985
Earlier work this paper cites.
On the complexity of space bounded interactive proofs (extended abstract)
Anne Condon and Richard J. Lipton · 1989
Earlier work this paper cites.
IP = PSPACE
Adi Shamir · 1992
Earlier work this paper cites.
Optimal depth neural networks for multiplication and related problems
Kai-Yeung Siu and Vwani P. Roychowdhury · 1992
Earlier work this paper cites.
Probabilistically checkable debate systems and nonapproximability of pspace-hard functions
Anne Condon, Joan Feigenbaum, Carsten Lund, and Peter W. Shor · 1995
Earlier work this paper cites.
On the complexity of interactive proofs with bounded communication
Oded Goldreich and Johan Håstad · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S. Sutton, David A. McAllester, Satinder Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
On interactive proofs with a laconic prover
Oded Goldreich, Salil P. Vadhan, and Avi Wigderson · 2002
Earlier work this paper cites.
Verifying and decoding in constant depth
Shafi Goldwasser, Dan Gutfreund, Alexander Healy, Tali Kaufman, and Guy N. Rothblum · 2007
Earlier work this paper cites.
Probabilistic proof systems: A primer
Oded Goldreich · 2008
Earlier work this paper cites.
Generative language modeling for automated theorem proving
Stanislas Polu and Ilya Sutskever · 2009
Earlier work this paper cites.
Interactive proofs of proximity: delegating computation in sublinear time
Guy N. Rothblum, Salil P. Vadhan, and Avi Wigderson · 2013
Earlier work this paper cites.
Understanding Machine Learning - From Theory to Algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Earlier work this paper cites.
SEPIA: search for proofs using inferred automata
Thomas Gransden, Neil Walkinshaw, and Rajeev Raman · 2015
Earlier work this paper cites.
Delegating computation: Interactive proofs for muggles
Shafi Goldwasser, Yael Tauman Kalai, and Guy N. Rothblum · 2015
Earlier work this paper cites.
Investigating the ability of neural networks to learn simple modular arithmetic, 2017
Theodoros Palamas · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F. Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Geoffrey Irving, Paul F. Christiano, and Dario Amodei · 2018
Cited alongside, same era.
Simple doubly-efficient interactive proof systems for locally-characterizable sets
Oded Goldreich and Guy N. Rothblum · 2018
Cited alongside, same era.
Overcoming exploration in reinforcement learning with demonstrations
Ashvin Nair, Bob McGrew, Marcin Andrychowicz, Wojciech Zaremba, and Pieter Abbeel · 2018
Cited alongside, same era.
Proofwriter: Generating implications, proofs, and abductive statements over natural language
Oyvind Tafjord, Bhavana Dalvi, and Peter Clark · 2020
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe · 2022
Later among the works it cites.
Exploration in deep reinforcement learning: A survey
Pawel Ladosz, Lilian Weng, Minwoo Kim, and Hyondong Oh · 2022
Later among the works it cites.
Mathematical capabilities of chatgpt
Simon Frieder, Luca Pinchetti, Alexis Chevalier, Ryan-Rhys Griffiths, Tommaso Salvatori, Thomas Lukasiewicz, Philipp Petersen, and Julius Berner · 2023
Later among the works it cites.
Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting
Miles Turpin, Julian Michael, Ethan Perez, and Samuel R. Bowman · 2023
Later among the works it cites.
Can chatgpt defend its belief in truth? evaluating LLM reasoning via debate
Boshi Wang, Xiang Yue, and Huan Sun · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cem Anil, Guodong Zhang, Yuhuai Wu, and Roger B. Grosse · 2021
Cited alongside, same era.
Interactive proofs for verifying machine learning
Shafi Goldwasser, Guy N. Rothblum, Jonathan Shafer, and Amir Yehudayoff · 2021
Cited alongside, same era.
Measuring mathematical problem solving with the MATH dataset
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt · 2021
Cited alongside, same era.
Investigating the limitations of the transformers with simple arithmetic tasks
Rodrigo Frassetto Nogueira, Zhiying Jiang, and Jimmy Lin · 2021
Cited alongside, same era.
Smooth and strong pcps
Orr Paradise · 2021
Cited alongside, same era.
Doubly efficient interactive proofs for general arithmetic circuits with linear prover time
Jiaheng Zhang, Tianyi Liu, Weijie Wang, Yinuo Zhang, Dawn Song, Xiang Xie, and Yupeng Zhang · 2021
Cited alongside, same era.
Constant-round interactive proofs for delegating computation
Omer Reingold, Guy N. Rothblum, and Ron D. Rothblum · 2021
Cited alongside, same era.
Scalable AI safety via doubly-efficient debate
Jonah Brown-Cohen, Geoffrey Irving, and Georgios Piliouras · 2023
Later among the works it cites.
Pseudointelligence: A unifying lens on language model evaluation
Shikhar Murty, Orr Paradise, and Pratyusha Sharma · 2023
Later among the works it cites.
Leandojo: Theorem proving with retrieval-augmented language models
Kaiyu Yang, Aidan M. Swope, Alex Gu, Rahul Chalamala, Peiyang Song, Shixing Yu, Saad Godil, Ryan J. Prenger, and Animashree Anandkumar · 2023
Later among the works it cites.
Auto-regressive next-token predictors are universal learners
Eran Malach · 2023
Later among the works it cites.
Kaleidoscope: Semantically-grounded, context-specific ML model evaluation
Harini Suresh, Divya Shanmugam, Tiffany Chen, Annie G. Bryan, Alexander D’Amour, John V. Guttag, and Arvind Satyanarayan · 2023
Later among the works it cites.
Can transformers learn the greatest common divisor?
François Charton · 2024
Closest in time.
Teaching arithmetic to small transformers
Nayoung Lee, Kartik Sreenivasan, Jason D. Lee, Kangwook Lee, and Dimitris Papailiopoulos · 2024
Closest in time.
Neural interactive proofs
Lewis Hammond and Sam Adam-Day · 2024
Closest in time.
Prover-verifier games improve legibility of LLM outputs
Jan Hendrik Kirchner, Yining Chen, Harri Edwards, Jan Leike, Nat McAleese, and Yuri Burda · 2024
Closest in time.
Interpretability guarantees with Merlin-Arthur classifiers
Stephan Wäldchen, Kartikey Sharma, Berkant Turan, Max Zimmer, and Sebastian Pokutta · 2024
Closest in time.
Solving olympiad geometry without human demonstrations
Trieu H. Trinh, Yuhuai Wu, Quoc V. Le, He He, and Thang Luong · 2024
Closest in time.
Mathematical discoveries from program search with large language models
Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Matej Balog, M Pawan Kumar, Emilien Dupont, Francisco JR Ruiz, Jordan S Ellenberg, Pengming Wang, Omar Fawzi, et al · 2024
Closest in time.
Let’s verify step by step
Hunter Lightman, Vineet Kosaraju, Yura Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe · 2024
Closest in time.