Fetching the paper…
Reading the bibliography…
If machine learning models were to achieve superhuman abilities at various reasoning or decision-making tasks, how would we go about evaluating such models, given that humans would necessarily be poor proxies for ground truth? In this paper, we propose a framework for evaluating superhuman models via consistency checks.
Parallel search of strongly ordered game trees
T Anthony Marsland and Murray Campbell · 1982
Earlier work this paper cites.
Encyclopaedia of chess openings, volume B (2nd ed.)
Lev Abramov, Vladimir Bagirov, Mikhail Botvinnik, Srdan Cvetkovic, Miroslav Filip, Efim Geller, Aivars Gipslis, Eduard Gufeld, Vlastimil Hort, Garry Kasparov, Viktor Korchnoi, Zdenko Krnic, Bent Larsen, Aleksandar Matanović, Nikolay Minev, John Nunn, Bruno Parma, Lev Polugaevsky, Alexey Suetin, Evgeny Sveshnikov, Mark Taimanov, Dragan Ugrinovic, and Wolfgang Uhlmann · 1984
Earlier work this paper cites.
Metamorphic testing: a new approach for generating next test cases
Tsong Y Chen, Shing C Cheung, and Shiu Ming Yiu · 1998
Earlier work this paper cites.
Deep blue
Murray Campbell, A Joseph Hoane Jr, and Feng-hsiung Hsu · 2002
Earlier work this paper cites.
Semi-supervised learning
Olivier Chapelle, Bernhard Scholkopf, and Alexander Zien · 2009
Earlier work this paper cites.
Testing and validating machine learning classifiers by metamorphic testing
Xiaoyuan Xie, Joshua WK Ho, Christian Murphy, Gail Kaiser, Baowen Xu, and Tsong Yueh Chen · 2011
Earlier work this paper cites.
Fairness through awareness
Cynthia Dwork, Moritz Hardt, Toniann Pitassi, Omer Reingold, and Richard Zemel · 2012
Earlier work this paper cites.
Playing Atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2013
Earlier work this paper cites.
Concrete problems in AI safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Machine bias
Julia Angwin, Jeff Larson, Surya Mattu, and Lauren Kirchner · 2016
Earlier work this paper cites.
Big data’s disparate impact
Solon Barocas and Andrew D Selbst · 2016
Earlier work this paper cites.
Scott Garrabrant, Tsvi Benson-Tilsen, Andrew Critch, Nate Soares, and Jessica Taylor · 2016
Earlier work this paper cites.
Equality of opportunity in supervised learning
Moritz Hardt, Eric Price, and Nati Srebro · 2016
Earlier work this paper cites.
Good and safe uses of ai oracles
Stuart Armstrong and Xavier O’Rorke · 2017
Earlier work this paper cites.
Adversarial examples for evaluating reading comprehension systems
Robin Jia and Percy Liang · 2017
Earlier work this paper cites.
Reluplex: An efficient SMT solver for verifying deep neural networks
Guy Katz, Clark Barrett, David L Dill, Kyle Julian, and Mykel J Kochenderfer · 2017
Earlier work this paper cites.
Alpha Zero’s “alien” chess shows the power, and the peculiarity, of AI, 2017
Will Knight · 2017
Earlier work this paper cites.
DeepXplore: Automated whitebox testing of deep learning systems
Kexin Pei, Yinzhi Cao, Junfeng Yang, and Suman Jana · 2017
Earlier work this paper cites.
Mastering the game of Go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy Lillicrap, Fan Hui, Laurent Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis · 2017
Earlier work this paper cites.
Leela Chess Zero
Lc0 developers · 2018
Earlier work this paper cites.
The accuracy, fairness, and limits of predicting recidivism
Julia Dressel and Hany Farid · 2018
Earlier work this paper cites.
Geoffrey Irving, Paul Christiano, and Dario Amodei · 2018
Earlier work this paper cites.
Human decisions and machine predictions
Jon Kleinberg, Himabindu Lakkaraju, Jure Leskovec, Jens Ludwig, and Sendhil Mullainathan · 2018
Earlier work this paper cites.
Virtual adversarial training: A regularization method for supervised and semi-supervised learning
Takeru Miyato, Shin-ichi Maeda, Masanori Koyama, and Shin Ishii · 2018
Earlier work this paper cites.
A general reinforcement learning algorithm that masters chess, Shogi, and Go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2018
Earlier work this paper cites.
DeepTest: Automated testing of deep-neural-network-driven autonomous cars
Yuchi Tian, Kexin Pei, Suman Jana, and Baishakhi Ray · 2018
Earlier work this paper cites.
Fairness definitions explained
Sahil Verma and Julia Rubin · 2018
Cited alongside, same era.
DeepRoad: GAN-based metamorphic testing and input validation framework for autonomous driving systems
Mengshi Zhang, Yuqun Zhang, Lingming Zhang, Cong Liu, and Sarfraz Khurshid · 2018
Cited alongside, same era.
Neural legal judgment prediction in English
Ilias Chalkidis, Ion Androutsopoulos, and Nikolaos Aletras · 2019
Cited alongside, same era.
Benchmarking neural network robustness to common corruptions and perturbations
Dan Hendrycks and Thomas Dietterich · 2019
Cited alongside, same era.
Writeup: Progress on AI safety via debate, 2020, 2020
Beth Barnes, Paul Christiano, L Ouyang, and G Irving · 2020
Cited alongside, same era.
Predictability and surprise in large generative models
Deep Ganguli, Danny Hernandez, Liane Lovitt, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova Dassarma, Dawn Drain, Nelson Elhage, Sheer El Showk, Stanislav Fort, Zac Hatfield-Dodds, Tom Henighan, Scott Johnston, Andy Jones, Nicholas Joseph, Jackson Kernian, Shauna Kravec, Ben Mann, Neel Nanda, Kamal Ndousse, Catherine Olsson, Daniela Amodei, Tom Brown, Jared Kaplan, Sam McCandlish, Christopher Olah, Dario Amodei, and Jack Clark · 2022
Later among the works it cites.
X-risk analysis for AI research
Dan Hendrycks and Mantas Mazeika · 2022
Later among the works it cites.
BECEL: Benchmark for consistency evaluation of language models
Myeongjun Jang, Deuk Sin Kwon, and Thomas Lukasiewicz · 2022
Later among the works it cites.
Machine bias
Surya Mattu Julia Angwin, Jeff Larson and Lauren Kirchner · 2022
Later among the works it cites.
Are AlphaZero-like agents robust to adversarial perturbations?
Li-Cheng Lan, Huan Zhang, Ti-Rong Wu, Meng-Yu Tsai, I Wu, Cho-Jui Hsieh, et al · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ilias Chalkidis, Manos Fergadiotis, Prodromos Malakasiotis, Nikolaos Aletras, and Ion Androutsopoulos · 2020
Cited alongside, same era.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Cited alongside, same era.
Evaluating models’ local decision boundaries via contrast sets
Matt Gardner, Yoav Artzi, Victoria Basmova, Jonathan Berant, Ben Bogin, Sihao Chen, Pradeep Dasigi, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, et al · 2020
Cited alongside, same era.
Performative prediction, 2020
Juan C. Perdomo, Tijana Zrnic, Celestine Mendler-Dünner, and Moritz Hardt · 2020
Cited alongside, same era.
Beyond accuracy: Behavioral testing of NLP models with CheckList
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh · 2020
Cited alongside, same era.
Testing monotonicity of machine learning models, 2020
Arnab Sharma and Heike Wehrheim · 2020
Cited alongside, same era.
Evolutionary algorithms and their applications to engineering problems
Adam Slowik and Halina Kwasnicka · 2020
Cited alongside, same era.
Later among the works it cites.
Holistic evaluation of language models
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christopher Ré, Diana Acosta-Navas, Drew A. Hudson, Eric Zelikman, Esin Durmus, Faisal Ladhak, Frieda Rong, Hongyu Ren, Huaxiu Yao, Jue Wang, Keshav Santhanam, Laurel Orr, Lucia Zheng, Mert Yuksekgonul, Mirac Suzgun, Nathan Kim, Neel Guha, Niladri Chatterji, Omar Khattab, Peter Henderson, Qian Huang, Ryan Chi, Sang Michael Xie, Shibani Santurkar, Surya Ganguli, Tatsunori Hashimoto, Thomas Icard, Tianyi Zhang, Vishrav Chaudhary, William Wang, Xuechen Li, Yifan Mai, Yuhui Zhang, and Yuta Koreeda · 2022
Later among the works it cites.
Forecasting future world events with neural networks
Andy Zou, Tristan Xiao, Ryan Jia, Joe Kwon, Mantas Mazeika, Richard Li, Dawn Song, Jacob Steinhardt, Owain Evans, and Dan Hendrycks · 2022
Later among the works it cites.
What is Lc0?, 2018
Lc0 authors · 2023
Closest in time.
A cookbook of self-supervised learning
Randall Balestriero, Mark Ibrahim, Vlad Sobal, Ari Morcos, Shashank Shekhar, Tom Goldstein, Florian Bordes, Adrien Bardes, Gregoire Mialon, Yuandong Tian, et al · 2023
Closest in time.
AI scientists: Safe and useful AI?, 2023
Yoshua Bengio · 2023
Closest in time.
Sparks of artificial general intelligence: Early experiments with GPT-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang · 2023
Closest in time.
Artificial influence: An analysis of AI-driven persuasion
Matthew Burtell and Thomas Woodside · 2023
Closest in time.
URL http://caissabase.co.uk/
Caissabase, 2023 · 2023
Closest in time.
Nondeterminism in Non-determinism in GPT-4 is caused by Sparse MoE, 2023
Sam Chann · 2023
Closest in time.
Eliciting latent knowledge: How to tell if your eyes deceive you, 2022
Paul Christiano, Ajeya Cotra, and Mark Xu · 2023
Closest in time.
LM vs LM: Detecting factual errors via cross examination
Roi Cohen, May Hamri, Mor Geva, and Amir Globerson · 2023
Closest in time.
Syzygy endgame tablebases, 2023
Niklas Fiekas · 2023
Closest in time.
Auto-GPT: An autonomous GPT-4 experiment, 2023
Significant Gravitas · 2023
Closest in time.
Consistency analysis of ChatGPT
Myeongjun Jang and Thomas Lukasiewicz · 2023
Closest in time.
Stockfish and Lc0, test at different number of nodes, Nov 2022
Marco Meloni · 2023
Closest in time.
Incentivizing honest performative predictions with proper scoring rules
Caspar Oesterheld, Johannes Treutlein, Emery Cooper, and Rubi Hudson · 2023
Closest in time.
Alexander Pan, Chan Jun Shern, Andy Zou, Nathaniel Li, Steven Basart, Thomas Woodside, Jonathan Ng, Hanlin Zhang, Scott Emmons, and Dan Hendrycks · 2023
Closest in time.
Manifold Markets: User GPT-4 (Bot), 2023
Markus Sobkowski · 2023
Closest in time.
Stockfish 15.1, 2023
Stockfish 15.1 · 2023
Closest in time.
Stockfish official repository
Stockfish developers · 2023
Closest in time.
Miles Turpin, Julian Michael, Ethan Perez, and Samuel R Bowman · 2023
Closest in time.