Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have demonstrated their ability to replicate human behaviors across a wide range of scenarios.
Hellaswag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019 · 1905
Earlier work this paper cites.
dungeons & dragons , volume 19
Gary Gygax and Dave Arneson. 1974 · 1974
Earlier work this paper cites.
How social an animal? the human capacity for caring
C Daniel Batson. 1990 · 1990
Earlier work this paper cites.
Social intelligence and interaction: Expressions and implications of the social bias in human intelligence
Esther N Goody. 1995 · 1995
Earlier work this paper cites.
A survey of deep learning techniques for neural machine translation
Shuoheng Yang, Yuxin Wang, and Xiaowen Chu. 2020 · 2002
Earlier work this paper cites.
Emotional intelligence and social interaction
Paulo N Lopes, Marc A Brackett, John B Nezlek, Astrid Schütz, Ina Sellin, and Peter Salovey. 2004 · 2004
Earlier work this paper cites.
Why we are social animals: The high road to imitation as social glue
Ap Dijksterhuis. 2005 · 2005
Earlier work this paper cites.
Social intelligence, human intelligence and niche construction
Kim Sterelny. 2007 · 2007
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020 · 2009
Earlier work this paper cites.
A unified model of human semantic knowledge and its disorders
Lang Chen, Matthew A Lambon Ralph, and Timothy T Rogers. 2017 · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018 · 2018
Earlier work this paper cites.
Deep dungeons and dragons: Learning character-action interactions from role-playing game transcripts
Annie Louis and Charles Sutton. 2018 · 2018
Earlier work this paper cites.
From place to site: Negotiating narrative complexity
Robert A Beauregard. 2020 · 2020
Earlier work this paper cites.
Program synthesis with large language models
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. 2021 · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021 · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021 · 2021
Earlier work this paper cites.
Automatic text summarization: A comprehensive survey
Wafaa S El-Kassas, Cherif R Salama, Ahmed A Rafea, and Hoda K Mohamed. 2021 · 2021
Earlier work this paper cites.
Beyond english-centric multilingual machine translation
Angela Fan, Shruti Bhosale, Holger Schwenk, Zhiyi Ma, Ahmed El-Kishky, Siddharth Goyal, Mandeep Baines, Onur Celebi, Guillaume Wenzek, Vishrav Chaudhary, et al. 2021 · 2021
Cited alongside, same era.
Work–family practices and complexity of their usage: a discourse analysis towards socially responsible human resource management
Suvi Heikkinen, Anna-Maija Lämsä, and Charlotta Niemistö. 2021 · 2021
Cited alongside, same era.
Winogrande: An adversarial winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021 · 2021
Cited alongside, same era.
Telling stories through multi-user dialogue by modeling character relations
Wai Man Si, Prithviraj Ammanabrolu, and Mark Riedl. 2021 · 2021
Cited alongside, same era.
Retrieving and reading: A comprehensive survey on open-domain question answering
Fengbin Zhu, Wenqiang Lei, Chao Wang, Jianming Zheng, Soujanya Poria, and Tat-Seng Chua. 2021 · 2021
Pei Zhou, Andrew Zhu, Jennifer Hu, Jay Pujara, Xiang Ren, Chris Callison-Burch, Yejin Choi, and Prithviraj Ammanabrolu. 2022 · 2022
Later among the works it cites.
Chatgpt vs. bard: a comparative study
Imtiaz Ahmed, Ayon Roy, Mashrafi Kajol, Uzma Hasan, Partha Protim Datta, and Md Rokonuzzaman Reza. 2023 · 2023
Later among the works it cites.
Studenteval: A benchmark of student-written prompts for large language models of code
Hannah McLean Babe, Sydney Nguyen, Yangtian Zi, Arjun Guha, Molly Q Feldman, and Carolyn Jane Anderson. 2023 · 2023
Later among the works it cites.
Agentverse: Facilitating multi-agent collaboration and exploring emergent behaviors in agents
Weize Chen, Yusheng Su, Jingwei Zuo, Cheng Yang, Chenfei Yuan, Chen Qian, Chi-Min Chan, Yujia Qin, Yaxi Lu, Ruobing Xie, et al. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Do as i can, not as i say: Grounding language in robotic affordances
Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, et al. 2022 · 2022
Cited alongside, same era.
Michael Bommarito II and Daniel Martin Katz. 2022 · 2022
Cited alongside, same era.
Dungeons and dragons as a dialog challenge for artificial intelligence
Chris Callison-Burch, Gaurav Singh Tomar, Lara Martin, Daphne Ippolito, Suma Bailis, and David Reitter. 2022 · 2022
Cited alongside, same era.
Towards reasoning in large language models: A survey
Jie Huang and Kevin Chen-Chuan Chang. 2022 · 2022
Cited alongside, same era.
The process of question answering: A computer simulation of cognition
Wendy G Lehnert. 2022 · 2022
Cited alongside, same era.
Holistic evaluation of language models
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al. 2022 · 2022
Cited alongside, same era.
Unified-io: A unified model for vision, language, and multi-modal tasks
Jiasen Lu, Christopher Clark, Rowan Zellers, Roozbeh Mottaghi, and Aniruddha Kembhavi. 2022 · 2022
Cited alongside, same era.
Yuzhen Huang, Yuzhuo Bai, Zhihao Zhu, Junlei Zhang, Jinghan Zhang, Tangjun Su, Junteng Liu, Chuancheng Lv, Yikai Zhang, Jiayi Lei, et al. 2023 · 2023
Later among the works it cites.
Mathprompter: Mathematical reasoning using large language models
Shima Imani, Liang Du, and Harsh Shrivastava. 2023 · 2023
Later among the works it cites.
Api-bank: A benchmark for tool-augmented llms
Minghao Li, Feifan Song, Bowen Yu, Haiyang Yu, Zhoujun Li, Fei Huang, and Yongbin Li. 2023 · 2023
Later among the works it cites.
Yuanzhi Liang, Linchao Zhu, and Yi Yang. 2023 · 2023
Later among the works it cites.
Overcoming the limitations of large language models how to enhance llms with human-like cognitive skills
Janna Lipenkova. 2023 · 2023
Later among the works it cites.
Generative agents: Interactive simulacra of human behavior
Joon Sung Park, Joseph C O’Brien, Carrie J Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2023 · 2023
Later among the works it cites.
A survey of hallucination in large foundation models
Vipula Rawte, Amit Sheth, and Amitava Das. 2023 · 2023
Later among the works it cites.
Evaluation and analysis of hallucination in large vision-language models
Junyang Wang, Yiyang Zhou, Guohai Xu, Pengcheng Shi, Chenlin Zhao, Haiyang Xu, Qinghao Ye, Ming Yan, Ji Zhang, Jihua Zhu, et al. 2023 · 2023
Later among the works it cites.
Superclue: A comprehensive chinese large language model benchmark
Liang Xu, Anqi Li, Lei Zhu, Hang Xue, Changtai Zhu, Kangkang Zhao, Haonan He, Xuanwei Zhang, Qiyue Kang, and Zhenzhong Lan. 2023 · 2023
Later among the works it cites.
Satlm: Satisfiability-aided language models using declarative prompting
Xi Ye, Qiaochu Chen, Isil Dillig, and Greg Durrett. 2023 · 2023
Later among the works it cites.
Agieval: A human-centric benchmark for evaluating foundation models
Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang, Shuai Lu, Yanlin Wang, Amin Saied, Weizhu Chen, and Nan Duan. 2023 · 2023
Later among the works it cites.
Fireball: A dataset of dungeons and dragons actual-play with structured game state information
Andrew Zhu, Karmanya Aggarwal, Alexander Feng, Lara J Martin, and Chris Callison-Burch. 2023 · 2023
Later among the works it cites.