Fetching the paper…
Reading the bibliography…
As machine learning models become more general, we need to characterise them in richer, more meaningful ways.
The Origins of Intelligence In The Child
Jean Piaget · 1923
Earlier work this paper cites.
Verification of Forecasts Expressed in Terms of Probability
Glenn W. Brier · 1950
Earlier work this paper cites.
A new vector partition of the probability score
Allan H. Murphy · 1973
Earlier work this paper cites.
Assessment in infancy: Ordinal scales of psychological development
Ina C. Užgiris and J. McV. Hunt · 1975
Earlier work this paper cites.
Individual differences in cognitive functions
Jan-Eric Gustafsson and Johan Olav Undheim · 1996
Earlier work this paper cites.
Human-like social skills in dogs?
Brian Hare and Michael Tomasello · 2005
Earlier work this paper cites.
A temporal same-object advantage in the tunnel effect: facilitated change detection for persisting objects
Jonathan I Flombaum and Brian J Scholl · 2006
Earlier work this paper cites.
Exceeding our grasp: Science, history, and the problem of unconceived alternatives , volume 1
P Kyle Stanford · 2006
Earlier work this paper cites.
Object persistence in philosophy and psychology
Brian J Scholl · 2007
Earlier work this paper cites.
Intuitive physical reasoning about occluded objects by inexperienced chicks
Cinzia Chiandetti and Giorgio Vallortigara · 2010
Earlier work this paper cites.
Cattell–Horn–Carroll abilities and cognitive tests: What we’ve learned from 20 years of research
Timothy Keith and Matthew R Reynolds · 2010
Earlier work this paper cites.
How cognitive modeling can benefit from hierarchical bayesian models
Michael D Lee · 2011
Earlier work this paper cites.
The No-U-Turn sampler: Adaptively setting path lengths in Hamiltonian Monte Carlo
Matthew D. Hoffman and Andrew Gelman · 2014
Earlier work this paper cites.
Adversarial examples for evaluating reading comprehension systems
Robin Jia and Percy Liang · 2017
Cited alongside, same era.
Building machines that learn and think like people
Brenden M Lake, Tomer D Ullman, Joshua B Tenenbaum, and Samuel J Gershman · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Unity: A general platform for intelligent agents
Arthur Juliani, Vincent-Pierre Berges, Esh Vckay, Yuan Gao, Hunter Henry, Marwan Mattar, and Danny Lange · 2018
Cited alongside, same era.
The animal-ai environment: Training and testing animal-like artificial cognition
Holistic evaluation of language models
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al · 2022
Later among the works it cites.
Evaluating object permanence in embodied agents using the animal-ai environment
Konstantinos Voudouris, Niall Donnelly, Danaja Rutar, Ryan Burnell, John Burden, José Hernández-Orallo, and Lucy G Cheke · 2022
Later among the works it cites.
Language, common sense, and the winograd schema challenge
Jacob Browning and Yann LeCun · 2023
Closest in time.
Google colaboratory
Google · 2023
Closest in time.
Running cognitive evaluations on large language models: The do’s and the don’ts, 2023
Anna A. Ivanova · 2023
Closest in time.
How do we know how smart ai systems are?
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Benjamin Beyret, José Hernández-Orallo, Lucy Cheke, Marta Halina, Murray Shanahan, and Matthew Crosby · 2019
Cited alongside, same era.
Vindicating methodological triangulation
Remco Heesen, Liam Kofi Bright, and Andrew Zucker · 2019
Cited alongside, same era.
Hellaswag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Cited alongside, same era.
The animal-AI testbed and competition
Matthew Crosby, Benjamin Beyret, Murray Shanahan, Jose Hernandez-Orallo, Lucy Cheke, and Marta Halina · 2020
Cited alongside, same era.
Artificial intelligence and the common sense of animals
Murray Shanahan, Matthew Crosby, Benjamin Beyret, and Lucy Cheke · 2020
Cited alongside, same era.
Dynabench: Rethinking benchmarking in NLP
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, et al · 2021
Cited alongside, same era.
Reduced, reused and recycled: The life of a dataset in machine learning research
Bernard Koch, Emily Denton, Alex Hanna, and Jacob G Foster · 2021
Cited alongside, same era.
Ai and the everything in the whole wide world benchmark
Deborah Raji, Emily Denton, Emily M. Bender, Alex Hanna, and Amandalynne Paullada · 2021
Cited alongside, same era.
Melanie Mitchell · 2023
Closest in time.
Probabilistic Machine Learning: Advanced Topics
Kevin P. Murphy · 2023
Closest in time.
Investigating object permanence in deep reinforcement learning agents
Konstantinos Voudouris, Jason Darwin Liu, Natasza Siwinska, Wout Schellaert, and Lucy G Cheke · 2024
Closest in time.
Mastering diverse control tasks through world models
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap · 2025
Closest in time.
Hibayes: A hierarchical bayesian modeling framework for ai evaluation statistics, 2025
Lennart Luettgau, Harry Coppock, Magda Dubois, Christopher Summerfield, and Cozmin Ududec · 2025
Closest in time.
The animal-ai environment: A virtual laboratory for comparative cognition and artificial intelligence research
Konstantinos Voudouris, Ben Slater, Lucy G Cheke, Wout Schellaert, José Hernández-Orallo, Marta Halina, Matishalin Patel, Ibrahim Alhas, Matteo G Mecattaf, John Burden, et al · 2025
Closest in time.
General Scales Unlock AI Evaluation with Explanatory and Predictive Power, March 2025
Lexin Zhou, Lorenzo Pacchiardi, Fernando Martínez-Plumed, Katherine M. Collins, Yael Moros-Daval, Seraphina Zhang, Qinlin Zhao, Yitian Huang, Luning Sun, Jonathan E. Prunty, Zongqian Li, Pablo Sánchez-García, Kexin Jiang Chen, Pablo A. M. Casares, Jiyun Zu, John Burden, Behzad Mehrbakhsh, David Stillwell, Manuel Cebrian, Jindong Wang, Peter Henderson, Sherry Tongshuang Wu, Patrick C. Kyllonen, Lucy Cheke, Xing Xie, and José Hernández-Orallo · 2025
Closest in time.