Fetching the paper…
Reading the bibliography…
Benchmarks play a significant role in the current evaluation of Large Language Models (LLMs), yet they often overlook the models' abilities to capture the nuances of human language, primarily focusing on evaluating embedded knowledge and technical skills.
Logic and conversation
Herbert Paul Grice. 1975 · 1975
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown. 2020 · 2005
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020 · 2009
Earlier work this paper cites.
Sprachwissenschaft: ein Reader
Ludger Hoffmann. 2010 · 2010
Earlier work this paper cites.
A test for the assessment of pragmatic abilities and cognitive substrates (apacs): Normative data and psychometric properties
Giorgio Arcara and Valentina Bambini. 2016 · 2016
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018 · 2018
Earlier work this paper cites.
Computational investigations of pragmatic effects in natural language
Jad Kabbara. 2019 · 2019
Earlier work this paper cites.
A survey on in-context learning
Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Zhiyong Wu, Baobao Chang, Xu Sun, Jingjing Xu, and Zhifang Sui. 2022 · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022 · 2022
Earlier work this paper cites.
Rethinking the role of demonstrations: What makes in-context learning work?
Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022 · 2022
Earlier work this paper cites.
Pragmatic analysis in natural language processing
R.S. Satpute and A. Agrawal. 2022 · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022 · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Earlier work this paper cites.
Gpt-4 surpassing human performance in linguistic pragmatics
Ljubisa Bojic, Predrag Kovacevic, and Milan Cabarkapa. 2023 · 2023
Cited alongside, same era.
Holistic evaluation of language models
Rishi Bommasani, Percy Liang, and Tony Lee. 2023 · 2023
Cited alongside, same era.
The pragmatic profile of chatgpt: assessing the pragmatic skills of a conversational agent
Chiara Barattieri di San Pietro, Federico Frau, Veronica Mangiaterra, and Valentina Bambini. 2023 · 2023
Cited alongside, same era.
Evaluating large language models: A comprehensive survey
Zishan Guo, Renren Jin, Chuang Liu, Yufei Huang, Dan Shi, Linhao Yu, Yan Liu, Jiaxuan Li, Bojian Xiong, Deyi Xiong, et al. 2023 · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023 · 2023
Later among the works it cites.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. 2023 · 2023
Later among the works it cites.
A survey on evaluation of large language models
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al. 2024 · 2024
Closest in time.
Open llm leaderboard v2
Clémentine Fourrier, Nathan Habib, Alina Lozovskaya, Konrad Szafer, and Thomas Wolf. 2024 · 2024
Closest in time.
Openagi: When llm meets domain experts
Yingqiang Ge, Wenyue Hua, Kai Mei, Juntao Tan, Shuyuan Xu, Zelong Li, Yongfeng Zhang, et al. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Xing Wang, and Zhaopeng Tu. 2023 · 2023
Cited alongside, same era.
An overview of bard: an early experiment with generative ai
James Manyika and Sissie Hsiao. 2023 · 2023
Cited alongside, same era.
Language artificial intelligences’ communicative performance quantified through the gricean conversation theory
Yunju Nam, Hyenyeong Chung, and Upyong Hong. 2023 · 2023
Cited alongside, same era.
Human-like problem-solving abilities in large language models using chatgpt
Graziella Orrù, Andrea Piarulli, Ciro Conversano, and Angelo Gemignani. 2023 · 2023
Cited alongside, same era.
The use of artificial intelligence-based chat-gpt and its challenges for the world of education; from the viewpoint of the development of creative writing skills
Muhammad Shidiq. 2023 · 2023
Cited alongside, same era.
CodeFusion: A pre-trained diffusion model for code generation
Mukul Singh, José Cambronero, Sumit Gulwani, Vu Le, Carina Negreanu, and Gust Verbruggen. 2023 · 2023
Cited alongside, same era.
Application of chatgpt in improving customer sentiment analysis for businesses
Frans Sudirjo, Karno Diantoro, Jassim Ahmad Al-Gasawneh, Hizbul Khootimah Azzaakiyyah, and Abu Muna Almaududi Ausat. 2023 · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. 2023 · 2023
Cited alongside, same era.
A study on large language models’ limitations in multiple-choice question answering
Aisha Khatun and Daniel G Brown. 2024 · 2024
Closest in time.
Ldcc-solar-10.7b
Wonchul Kim. 2024 · 2024
Closest in time.
Assessing the strengths and weaknesses of large language models
Shalom Lappin. 2024 · 2024
Closest in time.
Open Ko-LLM leaderboard: Evaluating large language models in Korean with Ko-h5 benchmark
Chanjun Park, Hyeonwoo Kim, Dahyun Kim, SeongHwan Cho, Sanghoon Kim, Sukyung Lee, Yungi Kim, and Hwalsuk Lee. 2024 · 2024
Closest in time.
Kmmlu: Measuring massive multitask language understanding in korean
Guijin Son, Hanwool Lee, Sungdong Kim, Seungone Kim, Niklas Muennighoff, Taekyoon Choi, Cheonbok Park, Kang Min Yoo, and Stella Biderman. 2024 · 2024
Closest in time.
Harnessing the power of llms in practice: A survey on chatgpt and beyond
Jingfeng Yang, Hongye Jin, Ruixiang Tang, Xiaotian Han, Qizhang Feng, Haoming Jiang, Shaochen Zhong, Bing Yin, and Xia Hu. 2024 · 2024
Closest in time.
Kang Min Yoo, Jaegeun Han, Sookyo In, Heewon Jeon, Jisu Jeong, Jaewook Kang, Hyunwook Kim, Kyung-Min Kim, Munhyong Kim, Sungju Kim, et al. 2024 · 2024
Closest in time.