Fetching the paper…
Reading the bibliography…
Despite the rapid development of large language models (LLMs) for the Korean language, there remains an obvious lack of benchmark datasets that test the requisite Korean cultural and linguistic knowledge.
A mathematical theory of communication
Claude Elwood Shannon. 1948 · 1948
Earlier work this paper cites.
Language Testing in Practice: Designing and Developing Useful Language Tests
L.F. Bachman and A.S. Palmer. 1996 · 1996
Earlier work this paper cites.
Conceptnet—a practical commonsense reasoning tool-kit
Hugo Liu and Push Singh. 2004 · 2004
Earlier work this paper cites.
Neural question generation from text: A preliminary study
Qingyu Zhou, Nan Yang, Furu Wei, Chuanqi Tan, Hangbo Bao, and M. Zhou. 2017 · 2017
Earlier work this paper cites.
Commonsense knowledge in machine intelligence
Niket Tandon, Aparna S. Varde, and Gerard de Melo. 2018 · 2018
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018 · 2018
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018 · 2018
Earlier work this paper cites.
A systematic review of automatic question generation for educational purposes
Ghader Kurdi, Jared Leo, Bijan Parsia, Uli Sattler, and Salam Al-Emari. 2019 · 2019
Earlier work this paper cites.
Cosmos QA: Machine reading comprehension with contextual commonsense reasoning
Lifu Huang, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2019 · 2019
Earlier work this paper cites.
Korquad1.0: Korean qa dataset for machine reading comprehension
Seungyoung Lim, Myungji Kim, and Jooyoul Lee. 2019 · 2019
Earlier work this paper cites.
CommonsenseQA: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019 · 2019
Earlier work this paper cites.
PAWS-X: A cross-lingual adversarial dataset for paraphrase identification
Yinfei Yang, Yuan Zhang, Chris Tar, and Jason Baldridge. 2019 · 2019
Earlier work this paper cites.
HellaSwag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019 · 2019
Earlier work this paper cites.
KorNLI and KorSTS: New benchmark datasets for Korean natural language understanding
Jiyeon Ham, Yo Joong Choe, Kyubyong Park, Ilji Choi, and Hyungjoon Soh. 2020a · 2020
Earlier work this paper cites.
Xtreme: A massively multilingual multi-task benchmark for evaluating cross-lingual generalization
Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020 · 2020
Cited alongside, same era.
XGLUE: A new benchmark dataset for cross-lingual pre-training, understanding and generation
Yaobo Liang, Nan Duan, Yeyun Gong, Ning Wu, Fenfei Guo, Weizhen Qi, Ming Gong, Linjun Shou, Daxin Jiang, Guihong Cao, Xiaodong Fan, Ruofei Zhang, Rahul Agrawal, Edward Cui, Sining Wei, Taroon Bharti, Ying Qiao, Jiun-Hung Chen, Winnie Wu, Shuguang Liu, Fan Yang, Daniel Campos, Rangan Majumder, and Ming Zhou. 2020 · 2020
Cited alongside, same era.
BEEP! Korean corpus of online news comments for toxic speech detection
Jihyung Moon, Won Ik Cho, and Junbum Lee. 2020 · 2020
Cited alongside, same era.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2020 · 2020
Cited alongside, same era.
The falcon series of open language models
Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra Cojocaru, Mérouane Debbah, Étienne Goffinet, Daniel Hesslow, Julien Launay, Quentin Malartic, Daniele Mazzotta, Badreddine Noune, Baptiste Pannier, and Guilherme Penedo. 2023 · 2023
Later among the works it cites.
Atlas: Few-shot learning with retrieval augmented language models
Gautier Izacard, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave. 2023 · 2023
Later among the works it cites.
A technical report for polyglot-ko: Open-source large-scale korean language models
Hyunwoong Ko, Kichang Yang, Minho Ryu, Taekyoon Choi, Seungmu Yang, Jiwung Hyun, Sungho Park, and Kyubyong Park. 2023 · 2023
Later among the works it cites.
Flask: Fine-grained language model evaluation based on alignment skill sets
Seonghyeon Ye, Doyoung Kim, Sungdong Kim, Hyeonbin Hwang, Seungone Kim, Yongrae Jo, James Thorne, Juho Kim, and Minjoon Seo. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Research community dynamics behind popular ai benchmarks
Fernando Martínez-Plumed, Pablo Barredo, Seán Ó hÉigeartaigh, and José Hernández-Orallo. 2021 · 2021
Cited alongside, same era.
Understanding the capabilities, limitations, and societal impact of large language models
Alex Tamkin, Miles Brundage, Jack Clark, and Deep Ganguli. 2021 · 2021
Cited alongside, same era.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021 · 2021
Cited alongside, same era.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021 · 2021
Cited alongside, same era.
Klue: Korean language understanding evaluation
Sungjoon Park, Jihyung Moon, Sungdong Kim, Won Ik Cho, Ji Yoon Han, Jangwon Park, Chisung Song, Junseong Kim, Youngsook Song, Taehwan Oh, Joohong Lee, Juhyun Oh, Sungwon Lyu, Younghoon Jeong, Inkwon Lee, Sangwoo Seo, Dongjun Lee, Hyunwoo Kim, Myeonghwa Lee, Seongbo Jang, Seungwon Do, Sunkyoung Kim, Kyungtae Lim, Jongwon Lee, Kyumin Park, Jamin Shin, Seonghyun Kim, Lucy Park, Lucy Park, Alice Oh, Jung-Woo Ha (NAVER AI Lab), and Kyunghyun Cho. 2021 · 2021
Cited alongside, same era.
EnCBP: A new benchmark dataset for finer-grained cultural background prediction in English
Weicheng Ma, Samiha Datta, Lili Wang, and Soroush Vosoughi. 2022 · 2022
Cited alongside, same era.
KoBEST: Korean balanced evaluation of significant tasks
Myeongjun Jang, Dohyung Kim, Deuk Sin Kwon, and Eric Davis. 2022 · 2022
Cited alongside, same era.
KOLD: Korean offensive language dataset
Younghoon Jeong, Juhyun Oh, Jongwon Lee, Jaimeen Ahn, Jihyung Moon, Sungjoon Park, and Alice Oh. 2022 · 2022
Cited alongside, same era.
Serengeti: Massively multilingual language models for africa
Ife Adebara, AbdelRahim Elmadany, Muhammad Abdul-Mageed, and Alcides Alcoba Inciarte. 2023 · 2023
Later among the works it cites.
Mega: Multilingual evaluation of generative ai
Kabir Ahuja, Harshita Diddee, Rishav Hada, Millicent Ochieng, Krithika Ramesh, Prachi Jain, Akshay Nambi, Tanuja Ganu, Sameer Segal, Maxamed Axmed, Kalika Bali, and Sunayana Sitaram. 2023 · 2023
Later among the works it cites.
Sumanth Doddapaneni, Rahul Aralikatte, Gowtham Ramesh, Shreya Goyal, Mitesh M. Khapra, Anoop Kunchukuttan, and Pratyush Kumar. 2023 · 2023
Later among the works it cites.
C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models
Yuzhen Huang, Yuzhuo Bai, Zhihao Zhu, Junlei Zhang, Jinghan Zhang, Tangjun Su, Junteng Liu, Chuancheng Lv, Yikai Zhang, Jiayi Lei, Yao Fu, Maosong Sun, and Junxian He. 2023 · 2023
Later among the works it cites.
Kobbq: Korean bias benchmark for question answering
Jiho Jin, Jiseon Kim, Nayeon Lee, Haneul Yoo, Alice Oh, and Hwaran Lee. 2023 · 2023
Later among the works it cites.
Hae-rae bench: Evaluation of korean knowledge in language models
Guijin Son, Hanwool Lee, Suwan Kim, Huiseo Kim, Jaecheol Lee, Je Won Yeom, Jihyu Jung, Jung Woo Kim, and Songseong Kim. 2023 · 2023
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R. Brown, and et al. 2023 · 2023
Later among the works it cites.
Agieval: A human-centric benchmark for evaluating foundation models
Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang, Shuai Lu, Yanlin Wang, Amin Saied, Weizhu Chen, and Nan Duan. 2023 · 2023
Later among the works it cites.
BBQ: A hand-built bias benchmark for question answering
Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel Bowman. 2022 · 2086
Closest in time.