Fetching the paper…
Reading the bibliography…
Large language models have demonstrated remarkable few-shot performance on many natural language understanding tasks.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E. Terry · 1952
Earlier work this paper cites.
The proposed uscf rating system, its development, theory, and applications
Arpad E Elo · 1967
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Christopher Maddison, Arthur Guez, Laurent Sifre, George Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Earlier work this paper cites.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm, 2017
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, and Demis Hassabis · 2017
Earlier work this paper cites.
Human-level play in the game of diplomacy by combining language models with strategic reasoning
Anton Bakhtin, Noam Brown, Emily Dinan, Gabriele Farina, Colin Flaherty, Daniel Fried, Andrew Goff, Jonathan Gray, Hengyuan Hu, Athul Paul Jacob, Mojtaba Komeili, Karthik Konath, Minae Kwon, Adam Lerer, Mike Lewis, Alexander H. Miller, Sandra Mitts, Adithya Renduchintala, Stephen Roller, Dirk Rowe, Weiyan Shi, Joe Spisak, Alexander Wei, David J. Wu, Hugh Zhang, and Markus Zijlstra · 2022
Earlier work this paper cites.
Palm: Scaling language modeling with pathways, 2022
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel · 2022
Earlier work this paper cites.
Star: Bootstrapping reasoning with reasoning, 2022
Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah D. Goodman · 2022
Earlier work this paper cites.
Llm-deliberation: Evaluating llms with interactive multi-agent negotiation games
Sahar Abdelnabi, Amr Gomaa, Sarath Sivaprasad, Lea Schonherr, and Mario Fritz · 2023
Earlier work this paper cites.
Lmrl gym: Benchmarks for multi-turn reinforcement learning with language models
Marwa Abdulhai, Isadora White, Charles Burton Snell, Charles Sun, Joey Hong, Yuexiang Zhai, Kelvin Xu, and Sergey Levine · 2023
Earlier work this paper cites.
Playing repeated games with large language models
Elif Akata, Lion Schulz, Julian Coda-Forno, Seong Joon Oh, Matthias Bethge, and Eric Schulz · 2023
Earlier work this paper cites.
clembench: Using game play to evaluate chat-optimized language models as conversational agents
Kranti Chalamalasetti, Jana Gotze, Sherzod Hakimov, Brielen Madureira, P. Sadler, and David Schlangen · 2023
Earlier work this paper cites.
Jiangjie Chen, Siyu Yuan, Rong Ye, Bodhisattwa Prasad Majumder, and Kyle Richardson · 2023
Earlier work this paper cites.
Chatbot arena: An open platform for evaluating llms by human preference
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E. Gonzalez, and Ion Stoica · 2023
Cited alongside, same era.
Reasoning with language model is planning with world model
Shibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong, Zhen Wang, Daisy Zhe Wang, and Zhiting Hu · 2023
Cited alongside, same era.
Unveiling the general intelligence factor in language models: A psychometric approach, 2023
David Ilić · 2023
Cited alongside, same era.
How novices use llm-based code generators to solve cs1 coding tasks in a self-paced learning environment, 2023
Majeed Kazemitabaar, Xinying Hou, Austin Henley, Barbara J. Ericson, David Weintrop, and Tovi Grossman · 2023
Cited alongside, same era.
Api-bank: A comprehensive benchmark for tool-augmented llms, 2023
Arb: Advanced reasoning benchmark for large language models, 2023
Tomohiro Sawada, Daniel Paleka, Alexander Havrilla, Pranav Tadepalli, Paula Vidas, Alexander Kranias, John J. Nay, Kshitij Gupta, and Aran Komatsuzaki · 2023
Later among the works it cites.
Toolformer: Language models can teach themselves to use tools, 2023
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom · 2023
Later among the works it cites.
Zihao Wang, Shaofei Cai, Anji Liu, Xiaojian Ma, and Yitao Liang · 2023
Later among the works it cites.
Agentgpt
Adam Watkins, Srijan Subedi, and Asim Shrestha · 2023
Later among the works it cites.
Xizhou Zhu, Yuntao Chen, Hao Tian, Chenxin Tao, Weijie Su, Chenyu Yang, Gao Huang, Bin Li, Lewei Lu, Xiaogang Wang, Y. Qiao, Zhaoxiang Zhang, and Jifeng Dai · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Minghao Li, Yingxiu Zhao, Bowen Yu, Feifan Song, Hangyu Li, Haiyang Yu, Zhoujun Li, Fei Huang, and Yongbin Li · 2023
Cited alongside, same era.
Avalonbench: Evaluating llms playing the game of avalon, 2023
Jonathan Light, Min Cai, Sheng Shen, and Ziniu Hu · 2023
Cited alongside, same era.
Agentsims: An open-source sandbox for large language model evaluation, 2023
Jiaju Lin, Haoran Zhao, Aochi Zhang, Yiting Wu, Huqiuyue Ping, and Qin Chen · 2023
Cited alongside, same era.
Alympics: Llm agents meet game theory – exploring strategic decision-making with ai agents, 2023
Shaoguang Mao, Yuzhe Cai, Yan Xia, Wenshan Wu, Xun Wang, Fengyi Wang, Tao Ge, and Furu Wei · 2023
Cited alongside, same era.
Evaluating language-model agents on realistic autonomous tasks
METR · 2023
Cited alongside, same era.
Gaia: a benchmark for general ai assistants, 2023
Grégoire Mialon, Clémentine Fourrier, Craig Swift, Thomas Wolf, Yann LeCun, and Thomas Scialom · 2023
Cited alongside, same era.
Gameeval: Evaluating llms on conversational games
Dan Qiao, Chenfei Wu, Yaobo Liang, Juntao Li, and Nan Duan · 2023
Cited alongside, same era.
Autogpt: An autonomous gpt-4 experiment
Toran Bruce Richards · 2023
Cited alongside, same era.
Later among the works it cites.
Llm-coordination: Evaluating and analyzing multi-agent coordination abilities in large language models, 2024
Saaket Agashe, Yue Fan, Anthony Reyna, and Xin Eric Wang · 2024
Closest in time.
Llmarena: Assessing capabilities of large language models in dynamic multi-agent environments
Junzhe Chen, Xuming Hu, Shuodi Liu, Shiyu Huang, Weijuan Tu, Zhaofeng He, and Lijie Wen · 2024
Closest in time.
Gtbench: Uncovering the strategic reasoning limitations of llms via game-theoretic evaluations
Jinhao Duan, Renming Zhang, James Diffenderfer, Bhavya Kailkhura, Lichao Sun, Elias Stengel-Eskin, Mohit Bansal, Tianlong Chen, and Kaidi Xu · 2024
Closest in time.
llm-reasoners: A library for advanced large language model reasoning
maitrix org · 2024
Closest in time.
choix: Inference algorithms for models based on luce’s choice axiom
Lucas Maystre · 2024
Closest in time.
Mmlu-pro: A more robust and challenging multi-task language understanding benchmark, 2024
Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, Tianle Li, Max Ku, Kai Wang, Alex Zhuang, Rongqi Fan, Xiang Yue, and Wenhu Chen · 2024
Closest in time.