Fetching the paper…
Reading the bibliography…
More than one hundred benchmarks have been developed to test the commonsense knowledge and commonsense reasoning abilities of artificial intelligence (AI) systems.
1904
Earlier work this paper cites.
1904
Earlier work this paper cites.
1905
Earlier work this paper cites.
1908
Earlier work this paper cites.
1908
Earlier work this paper cites.
1909
Earlier work this paper cites.
1909
Earlier work this paper cites.
1909
Earlier work this paper cites.
1909
Earlier work this paper cites.
1910
Earlier work this paper cites.
1910
Earlier work this paper cites.
1910
Earlier work this paper cites.
1911
Earlier work this paper cites.
1911
Earlier work this paper cites.
Allen, James F. “Maintaining knowledge about temporal intervals.” Communications of the ACM 26, no. 11 (1983): 832-843. https://dl.acm.org/doi/pdf/10.1145/182.358434
1983
Earlier work this paper cites.
Wimmer, Heinz and Joseph Perner. “Beliefs about beliefs: Representation and constraining function of wrong beliefs in young children’s understanding of deception”. Cognition,
1983
Earlier work this paper cites.
Lenat, Douglas B., Mayank Prakash, and Mary Shepherd. “CYC: Using common sense knowledge to overcome brittleness and knowledge acquisition bottlenecks.” AI magazine 6, no. 4 (1985): 65-65. https://ojs.aaai.org/index.php/aimagazine/article/view/510
1985
Earlier work this paper cites.
Davis, Ernest. Representations of Commonsense Knowledge
1990
Earlier work this paper cites.
Kearns, Michael. “Efficient noise-tolerant learning from statistical queries.” Journal of the ACM (JACM) 45, no. 6 (1998): 983-1006
1998
Earlier work this paper cites.
Charness, Neil, Eyal M. Reingold, Marc Pomplun, and Dave M. Stampe. “The perceptual aspect of skilled performance in chess: Evidence from eye movements.” Memory & cognition 29, no. 8 (2001): 1146-1152. https://link.springer.com/article/10.3758/BF03206384
2001
Earlier work this paper cites.
Brown, T.L., H.E. LeMay. B. Bursten, Chemistry: The Central Science
2003
Earlier work this paper cites.
2003
Earlier work this paper cites.
2004
Earlier work this paper cites.
2004
Earlier work this paper cites.
2005
Earlier work this paper cites.
2005
Earlier work this paper cites.
2005
Earlier work this paper cites.
2005
Earlier work this paper cites.
2005
Earlier work this paper cites.
2005
Earlier work this paper cites.
2007
Earlier work this paper cites.
2007
Earlier work this paper cites.
Vogt, Stine, and Svein Magnussen. “Expertise in pictorial perception: Eye-movement patterns and visual memory in artists and laymen.” Perception 36, no. 1 (2007): 91-100. https://journals.sagepub.com/doi/abs/10.1068/p5262
2007
Earlier work this paper cites.
2008
Earlier work this paper cites.
van Harmelen, Frank, Vladimir Lifschitz, and Bruce Porter (eds). Handbook of Knowledge Representation
2008
Earlier work this paper cites.
Deng, Jia, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. “Imagenet: A large-scale hierarchical image database.” CVPR
2009
Earlier work this paper cites.
Dagan, Ido, Bill Dolan, Bernardo Magnini, and Dan Roth. 2010. “Recognizing textual entailment: Rational, evaluation and approaches–erratum.” Natural Language Engineering 16, no. 1: 105-105
2010
Earlier work this paper cites.
Deutscher, Guy, Through the Language Glass: Why the World Looks Different in Different Languages,
2010
Earlier work this paper cites.
2010
Earlier work this paper cites.
2010
Earlier work this paper cites.
2010
Earlier work this paper cites.
Kahneman, Daniel, Thinking Fast and Slow,
2011
Earlier work this paper cites.
Roemmele, Melissa, Cosmin Adrian Bejan, and Andrew S. Gordon. “Choice of Plausible Alternatives: An Evaluation of Commonsense Causal Reasoning.” In AAAI spring symposium: logical formalizations of commonsense reasoning, pp. 90-95. 2011. https://people.ict.usc.edu/gordon/public_html/publications/AAAI-SPRING11A.PDF
2011
Earlier work this paper cites.
Gervais, Will M., and Ara Norenzayan. “Analytic thinking promotes religious disbelief.” Science (2012)
2012
Earlier work this paper cites.
Levesque, Hector, Ernest Davis, and Leora Morgenstern. “The Winograd Schema Challenge”. Principles of Knowledge Representation and Reasoning
2012
Earlier work this paper cites.
Davis, Ernest. “Qualitative Spatial Reasoning in Interpreting Text and Narrative.” Spatial Cognition and Computation,
2013
Earlier work this paper cites.
Lin, Tsung-Yi, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. “Microsoft COCO: Common objects in context.” In European conference on computer vision, pp. 740-755. Springer, Cham, 2014. https://link.springer.com/chapter/10.1007/978-3-319-10602-1_48
2014
Earlier work this paper cites.
McWhorter, John. The Language Hoax: Why The World Looks the Same in Any Language,
2014
Earlier work this paper cites.
Mueller, Erik T. Commonsense reasoning: an event calculus based approach
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
Caba Heilbron, Fabian, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. “Activitynet: A large-scale video benchmark for human activity understanding.” In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 961-970. 2015 https://svl.stanford.edu/assets/papers/Heilbron_ActivityNet_A_Large-Scale_2015_CVPR_paper.pdf
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
Davis, Ernest. “How to Write Science Questions that are Easy for People and Hard for Computers,” AI Magazine,
2016
Cited alongside, same era.
Gordon, Andrew. “Commonsense interpretation of triangle behavior.” In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 30, no. 1. 2016. https://ojs.aaai.org/index.php/AAAI/article/view/9881
2016
Cited alongside, same era.
Mostafazadeh, Nasrin, Nathanael Chambers, Xiaodong He, Devi Parikh, Dhruv Batra, Lucy Vanderwende, Pushmeet Kohli, and James Allen. “A corpus and cloze evaluation for deeper understanding of commonsense stories.” In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 839-849. 2016. https://aclanthology.org/N16-1098.pdf
2016
Cited alongside, same era.
Hong, Yining, Li Yi, Josh Tenenbaum, Antonio Torralba, and Chuang Gan. “Ptr: A benchmark for part-based conceptual, relational, and physical reasoning.” Advances in Neural Information Processing Systems 34 (2021): 17427-17440. https://proceedings.neurips.cc/paper/2021/hash/918f5cd5a5c0d48671d4d4fc54bab2e9-Abstract.html
2021
Later among the works it cites.
Kayser, Maxime, Oana-Maria Camburu, Leonard Salewski, Cornelius Emde, Virginie Do, Zeynep Akata, and Thomas Lukasiewicz. “E-ViL: A dataset and benchmark for natural language explanations in vision-language tasks.” In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 1244-1254. 2021. https://openaccess.thecvf.com/content/ICCV2021/html/Kayser_E-ViL_A_Dataset_and_Benchmark_for_Natural_Language_Explanations_in_ICCV_2021_paper.html
2021
Later among the works it cites.
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
Thomee, Bart, David A. Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li. “YFCC100M: The new data in multimedia research.” Communications of the ACM 59, no. 2 (2016): 64-73. https://dl.acm.org/doi/abs/10.1145/2812802
2016
Cited alongside, same era.
Davis, Ernest. “The Logical Depth of Reasoning about Other Minds.” Advances in Cognitive Systems,
2017
Cited alongside, same era.
Evans, Jonathan St BT. “Dual process theory: Perspectives and problems.” In Wim de Neys (ed.) Dual process theory 2.0
2017
Cited alongside, same era.
Gordon, Andrew S., and Jerry R. Hobbs. A formal theory of commonsense psychology: How people think people think
2017
Cited alongside, same era.
Goyal, Raghav, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel et al. “The ‘something something’ video database for learning and evaluating visual common sense.” In Proceedings of the IEEE international conference on computer vision, pp. 5842-5850. 2017. https://openaccess.thecvf.com/content_ICCV_2017/papers/Goyal_The_Something_Something_ICCV_2017_paper.pdf
2017
Cited alongside, same era.
Krishna, Ranjay, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen et al. “Visual genome: Connecting language and vision using crowdsourced dense image annotations.” International journal of computer vision 123, no. 1 (2017): 32-73. https://link.springer.com/article/10.1007/S11263-016-0981-7
2017
Cited alongside, same era.
Rohrbach, Anna, Atousa Torabi, Marcus Rohrbach, Niket Tandon, Christopher Pal, Hugo Larochelle, Aaron Courville, and Bernt Schiele. “Movie description.” International Journal of Computer Vision 123, no. 1 (2017): 94-120. https://link.springer.com/article/10.1007/s11263-016-0987-1
2017
Cited alongside, same era.
Speer, Robyn, Joshua Chin, and Catherine Havasi. “ConceptNet 5.5: An open multilingual graph of general knowledge.” In Thirty-first AAAI conference on artificial intelligence. 2017. https://ojs.aaai.org/index.php/AAAI/article/view/11164
2017
Cited alongside, same era.
2021
Later among the works it cites.
Li, Linjie, Jie Lei, Zhe Gan, and Jingjing Liu. “Adversarial VQA: A new benchmark for evaluating the robustness of vqa models.” In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 2042-2051. 2021. https://openaccess.thecvf.com/content/ICCV2021/html/Li_Adversarial_VQA_A_New_Benchmark_for_Evaluating_the_Robustness_of_ICCV_2021_paper.html
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
Paullada, Amandalynne, Inioluwa Deborah Raji, Emily M. Bender, Emily Denton, and Alex Hanna. “Data and its (dis) contents: A survey of dataset development and use in machine learning research.” Patterns 2, no. 11 (2021): 100336. https://www.sciencedirect.com/science/article/pii/S2666389921001847
2021
Later among the works it cites.
2021
Later among the works it cites.
Riochet, Ronan, Mario Ynocente Castro, Mathieu Bernard, Adam Lerer, Rob Fergus, Véronique Izard, and Emmanuel Dupoux. “IntPhys 2019: A Benchmark for Visual Intuitive Physics Understanding.” IEEE Transactions on Pattern Analysis and Machine Intelligence 44, no. 9 (2021): 5016-5025. https://ieeexplore.ieee.org/abstract/document/9442261
2021
Later among the works it cites.
Sakaguchi, Keisuke, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. “Winogrande: An adversarial Winograd schema challenge at scale.” Communications of the ACM 64, no. 9 (2021): 99-106. https://dl.acm.org/doi/abs/10.1145/3474381
2021
Later among the works it cites.
Seo, Jaehyung, Chanjun Park, Hyeonseok Moon, Sugyeong Eo, Myunghoon Kang, Seounghoon Lee, and Heuiseok Lim. “KommonGen: A Dataset for Korean Generative Commonsense Reasoning Evaluation.” In Annual Conference on Human and Language Technology, pp. 55-60. Human and Language Technology, 2021. https://koreascience.kr/article/CFKO202130060697830.pdf
2021
Later among the works it cites.
Shu, Tianmin, Abhishek Bhandwaldar, Chuang Gan, Kevin Smith, Shari Liu, Dan Gutfreund, Elizabeth Spelke, Joshua Tenenbaum, and Tomer Ullman. “Agent: A benchmark for core psychological reasoning.” In International Conference on Machine Learning, pp. 9614-9625. PMLR, 2021. https://research.ibm.com/publications/agent-a-benchmark-for-core-psychological-reasoning
2021
Later among the works it cites.
2021
Later among the works it cites.
KVL-BERT: Knowledge Enhanced Visual-and-Linguistic BERT for visual commonsense reasoning,
Song, Dandan, Siyi Ma, Zhanchen Sun, Sicheng Yang, Lejian Lia (2021) · 2021
Later among the works it cites.
2021
Later among the works it cites.
Weihs, Luca, Matt Deitke, Aniruddha Kembhavi, and Roozbeh Mottaghi. “Visual room rearrangement.” In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 5922-5931. 2021. https://openaccess.thecvf.com/content/CVPR2021/html/Weihs_Visual_Room_Rearrangement_CVPR_2021_paper.html
2021
Later among the works it cites.
2021
Later among the works it cites.
Aghahadi, Zeinab, and Alireza Talebpour. “Avicenna: a challenge dataset for natural language generation toward commonsense syllogistic reasoning.” Journal of Applied Non-Classical Logics
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
Over-reliance on English hinders cognitive science
Blasi, Damián, Joseph Hencirh, Evangelia Adamou, David Kemmerer, Asifa Majid (2022) · 2022
Later among the works it cites.
Davis, Ernest and Gary Marcus. “Experiments in Commonsense Reasoning in GPT-3: Status Report from June 2022.” Unpublished. https://cs.nyu.edu/~davise/papers/GPT-3-6-22.html
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
Guan, Jian, Zhuoer Feng, Yamei Chen, Ruilin He, Xiaoxi Mao, Changjie Fan, and Minlie Huang. “LOT: A Story-Centric Benchmark for Evaluating Chinese Long Text Understanding and Generation.” Transactions of the Association for Computational Linguistics 10 (2022): 434-451
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capacities of language models
Srivastaa, Aarohi et al. 2022 · 2022
Later among the works it cites.
2022
Later among the works it cites.
Not another Negation Benchmark: The NaN-NLI Test Suite for Sub-clausal Negation
Truong, Hung Thinh, Yulia Otmakhova, Timothy Baldwin, Trevor Cohn, Karin Verspoor, and Jey Han Lau. (2022) · 2022
Later among the works it cites.
2022
Later among the works it cites.
Weihs, Luca, Amanda Rose Yuile, Renée Baillargeon, Cynthia L. Fisher, Gary Marcus, Roozbeh Mottaghi, and Aniruddha Kembhavi. “Benchmarking progress to infant-Level physical reasoning in AI.” Transactions on Machine Learning, 2022. https://openreview.net/pdf?id=9NjqD9i48M
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
Yu, Samuel, Peter Wu, Paul Pu Liang, Ruslan Salakhutdinov, and Louis-Philippe Morency. “PACS: A Dataset for Physical Audiovisual CommonSense Reasoning.” In European Conference on Computer Vision, pp. 292-309. Springer, Cham, 2022. https://ui.adsabs.harvard.edu/abs/2022arXiv220311130Y/abstract
2022
Later among the works it cites.
2023
Closest in time.
2023
Closest in time.
Marcus, Gary and Ernest Davis. “How Not to Test GPT.” Comm. ACM blog
2023
Closest in time.
LoBue, Peter, and Alexander Yates. “Types of common-sense knowledge needed for recognizing textual entailment.” In Proceedings of the 49th annual meeting of the association for computational linguistics: human language technologies, pp. 329-334. 2011. https://aclanthology.org/P11-2057.pdf
2057
Closest in time.