Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have made progress in various real-world tasks, which stimulates requirements for the evaluation of LLMs.
Sequential Equilibria
David M. Kreps and Robert Wilson. 1982 · 1982
Earlier work this paper cites.
Volunteering leads to rock–paper–scissors dynamics in a public goods game
Dirk Semmann, Hans-Jürgen Krambeck, and Manfred Milinski. 2003 · 2003
Earlier work this paper cites.
Does the Turing Test Demonstrate Intelligence or Not?. In Proceedings, The Twenty-First National Conference on Artificial Intelligence and the Eighteenth Innovative Applications of Artificial Intelligence Conference, July 16-20, 2006, Boston, Massachusetts, USA . AAAI Press, 1539–1542
Stuart M. Shieber. 2006 · 2006
Earlier work this paper cites.
Idioms: Motivation and etymology
Dmitrij Dobrovol’skij and Elisabeth Piirainen. 2010 · 2010
Earlier work this paper cites.
Interactive analysis of Likert scale data using a multichart visualization tool. In IHC+CLIHC . Brazilian Computer Society / ACM, 358–365
Fábio Petrillo, Andre Suslik Spritzer, Carla Maria Dal Sasso Freitas, and Marcelo Soares Pimenta. 2011 · 2011
Earlier work this paper cites.
167Mei’s Story: “Idioms Solitaire” between Sports Fans
Huatong Sun. 2012 · 2012
Earlier work this paper cites.
Overview of the IWSLT 2017 Evaluation Campaign. In Proceedings of the 14th International Conference on Spoken Language Translation . International Workshop on Spoken Language Translation, Tokyo, Japan, 2–14
Mauro Cettolo, Marcello Federico, Luisa Bentivogli, Jan Niehues, Sebastian Stüker, Katsuhito Sudoh, Koichiro Yoshino, and Christian Federmann. 2017 · 2017
Earlier work this paper cites.
Public goods games and psychological utility: Theory and evidence
Sanjit Dhami, Mengxing Wei, and Ali al Nowaihi. 2019 · 2017
Earlier work this paper cites.
Attention is All you Need. In NIPS . 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Well-Founded Extensive Games with Perfect Information. In Proceedings Eighteenth Conference on Theoretical Aspects of Rationality and Knowledge, TARK 2021, Beijing, China, June 25-27, 2021 (EPTCS, Vol. 335) , Joseph Y. Halpern and Andrés Perea (Eds.). 7–21
Krzysztof R. Apt and Sunil Simon. 2021 · 2021
Earlier work this paper cites.
Program Synthesis with Large Language Models
Jacob Austin, Augustus Odena, Maxwell I. Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie J. Cai, Michael Terry, Quoc V. Le, and Charles Sutton. 2021 · 2021
Earlier work this paper cites.
Measuring Coding Challenge Competence With APPS. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks 2021, December 2021, virtual , Joaquin Vanschoren and Sai-Kit Yeung (Eds.)
Dan Hendrycks, Steven Basart, Saurav Kadavath, Mantas Mazeika, Akul Arora, Ethan Guo, Collin Burns, Samir Puranik, Horace He, Dawn Song, and Jacob Steinhardt. 2021 · 2021
Cited alongside, same era.
PaLM: Scaling Language Modeling with Pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022 · 2022
Cited alongside, same era.
A Survey on Document-level Neural Machine Translation: Methods and Evaluation
Sameen Maruf, Fahimeh Saleh, and Gholamreza Haffari. 2022 · 2022
Cited alongside, same era.
OpenAI. 2023 · 2023
Closest in time.
Human-like problem-solving abilities in large language models using ChatGPT
Graziella Orrù, Andrea Piarulli, Ciro Conversano, and Angelo Gemignani. 2023 · 2023
Closest in time.
GameEval: Evaluating LLMs on Conversational Games
Dan Qiao, Chenfei Wu, Yaobo Liang, Juntao Li, and Nan Duan. 2023 · 2023
Closest in time.
Neural Machine Translation for Low-resource Languages: A Survey
Surangika Ranathunga, En-Shiun Annie Lee, Marjana Prifti Skenduli, Ravi Shekhar, Mehreen Alam, and Rishemjit Kaur. 2023 · 2023
Closest in time.
Testing the General Deductive Reasoning Capacity of Large Language Models Using OOD Examples
Abulhair Saparov, Richard Yuanzhe Pang, Vishakh Padmakumar, Nitish Joshi, Seyed Mehran Kazemi, Najoung Kim, and He He. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Synchromesh: Reliable Code Generation from Pre-trained Language Models. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022 . OpenReview.net
Gabriel Poesia, Alex Polozov, Vu Le, Ashish Tiwari, Gustavo Soares, Christopher Meek, and Sumit Gulwani. 2022 · 2022
Cited alongside, same era.
Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, Quyet V. Do, Yan Xu, and Pascale Fung. 2023 · 2023
Cited alongside, same era.
A Survey on Evaluation of Large Language Models
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, Wei Ye, Yue Zhang, Yi Chang, Philip S. Yu, Qiang Yang, and Xing Xie. 2023 · 2023
Cited alongside, same era.
AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback
Yann Dubois, Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023 · 2023
Cited alongside, same era.
Mathematical Capabilities of ChatGPT
Simon Frieder, Luca Pinchetti, Alexis Chevalier, Ryan-Rhys Griffiths, Tommaso Salvatori, Thomas Lukasiewicz, Philipp Christian Petersen, and Julius Berner. 2023 · 2023
Cited alongside, same era.
ChatGPT for Programming Numerical Methods
Ali Kashefi and Tapan Mukerji. 2023 · 2023
Cited alongside, same era.
New Trends in Machine Translation using Large Language Models: Case Examples with ChatGPT
Chenyang Lyu, Jitao Xu, and Longyue Wang. 2023 · 2023
Cited alongside, same era.
Document-Level Machine Translation with Large Language Models
Longyue Wang, Chenyang Lyu, Tianbo Ji, Zhirui Zhang, Dian Yu, Shuming Shi, and Zhaopeng Tu. 2023a
Cited in the paper.
PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Yidong Wang, Zhuohao Yu, Zhengran Zeng, Linyi Yang, Cunxiang Wang, Hao Chen, Chaoya Jiang, Rui Xie, Jindong Wang, Xing Xie, Wei Ye, Shikun Zhang, and Yue Zhang. 2023b
Cited in the paper.
Closest in time.
Natural Language to Code Generation in Interactive Data Science Notebooks. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023 , Anna Rogers, Jordan L. Boyd-Graber, and Naoaki Okazaki (Eds.). Association for Computational Linguistics, 126–173
Pengcheng Yin, Wen-Ding Li, Kefan Xiao, Abhishek Rao, Yeming Wen, Kensen Shi, Joshua Howland, Paige Bailey, Michele Catasta, Henryk Michalewski, Oleksandr Polozov, and Charles Sutton. 2023 · 2023
Closest in time.
Planning with Large Language Models for Code Generation. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023 . OpenReview.net
Shun Zhang, Zhenfang Chen, Yikang Shen, Mingyu Ding, Joshua B. Tenenbaum, and Chuang Gan. 2023 · 2023
Closest in time.
Judging LLM-as-a-judge with MT-Bench and Chatbot Arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023 · 2023
Closest in time.
PromptBench: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts
Kaijie Zhu, Jindong Wang, Jiaheng Zhou, Zichen Wang, Hao Chen, Yidong Wang, Linyi Yang, Wei Ye, Neil Zhenqiang Gong, Yue Zhang, and Xing Xie. 2023 · 2023
Closest in time.
Efficiently Measuring the Cognitive Ability of LLMs: An Adaptive Testing Perspective
Yan Zhuang, Qi Liu, Yuting Ning, Weizhe Huang, Rui Lv, Zhenya Huang, Guanhao Zhao, Zheng Zhang, Qingyang Mao, Shijin Wang, and Enhong Chen. 2023 · 2023
Closest in time.