Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have demonstrated exceptional performance across a wide range of natural language tasks.
Estimation of item response models using the em algorithm for finite mixtures
David J Woodruff and Bradley A Hanson. 1996 · 1996
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021a · 2009
Earlier work this paper cites.
Multidimensional item response theory models
Mark D Reckase. 2009 · 2009
Earlier work this paper cites.
hdbscan: Hierarchical density based clustering
Leland McInnes, John Healy, Steve Astels, et al. 2017 · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018 · 2018
Earlier work this paper cites.
Umap: Uniform manifold approximation and projection for dimension reduction
Leland McInnes, John Healy, and James Melville. 2018 · 2018
Earlier work this paper cites.
Know what you don’t know: Unanswerable questions for squad
Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018 · 2018
Earlier work this paper cites.
Commonsenseqa: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2018 · 2018
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W Cohen, Ruslan Salakhutdinov, and Christopher D Manning. 2018 · 2018
Earlier work this paper cites.
Ekt: Exercise-aware knowledge tracing for student performance prediction
Qi Liu, Zhenya Huang, Yu Yin, Enhong Chen, Hui Xiong, Yu Su, and Guoping Hu. 2019 · 2019
Earlier work this paper cites.
Item response theory in ai: Analysing machine learning classifiers at the instance level
Fernando Martínez-Plumed, Ricardo BC Prudêncio, Adolfo Martínez-Usó, and José Hernández-Orallo. 2019 · 2019
Earlier work this paper cites.
Neural cognitive diagnosis for intelligent education systems
Fei Wang, Qi Liu, Enhong Chen, Zhenya Huang, Yuying Chen, Yu Yin, Zai Huang, and Shijin Wang. 2020 · 2020
Earlier work this paper cites.
Program synthesis with large language models
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. 2021 · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021 · 2021
Earlier work this paper cites.
Rcd: Relation map driven cognitive diagnosis for intelligent education systems
Weibo Gao, Qi Liu, Zhenya Huang, Yu Yin, Haoyang Bi, Mu-Chun Wang, Jianhui Ma, Shijin Wang, and Yu Su. 2021 · 2021
Earlier work this paper cites.
Evaluation examples are not equally informative: How should that change nlp leaderboards?
Pedro Rodriguez, Joe Barrow, Alexander Miserlis Hoyle, John P Lalor, Robin Jia, and Jordan Boyd-Graber. 2021 · 2021
Earlier work this paper cites.
When ai difficulty is easy: The explanatory power of predicting irt difficulty
Fernando Martínez-Plumed, David Castellano, Carlos Monserrat-Aranda, and José Hernández-Orallo. 2022 · 2022
Earlier work this paper cites.
Automix: Automatically mixing language models
Pranjal Aggarwal, Aman Madaan, Ankit Anand, Srividya Pranavi Potharaju, Swaroop Mishra, Pei Zhou, Aditya Gupta, Dheeraj Rajagopal, Karthik Kappaganthu, Yiming Yang, et al. 2023 · 2023
Cited alongside, same era.
Frugalgpt: How to use large language models while reducing cost and improving performance
Lingjiao Chen, Matei Zaharia, and James Zou. 2023 · 2023
Cited alongside, same era.
Leveraging transferable knowledge concept graph embedding for cold-start cognitive diagnosis
Weibo Gao, Hao Wang, Qi Liu, Fei Wang, Xin Lin, Linan Yue, Zheng Zhang, Rui Lv, and Shijin Wang. 2023 · 2023
Cited alongside, same era.
Tryage: Real-time, intelligent routing of user prompts to large language model
Surya Narayanan Hari and Matt Thomson. 2023 · 2023
Cited alongside, same era.
Automated evaluation of retrieval-augmented language models with task-specific exam generation
Gauthier Guinet, Behrooz Omidvar-Tehrani, Anoop Deoras, and Laurent Callot. 2024 · 2024
Later among the works it cites.
Routerbench: A benchmark for multi-llm routing system
Qitian Jason Hu, Jacob Bieker, Xiuyu Li, Nan Jiang, Benjamin Keigwin, Gaurav Ranganath, Kurt Keutzer, and Shriyash Kaustubh Upadhyay. 2024 · 2024
Later among the works it cites.
C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models
Yuzhen Huang, Yuzhuo Bai, Zhihao Zhu, Junlei Zhang, Jinghan Zhang, Tangjun Su, Junteng Liu, Chuancheng Lv, Yikai Zhang, Yao Fu, et al. 2024 · 2024
Later among the works it cites.
Soyeong Jeong, Jinheon Baek, Sukmin Cho, Sung Ju Hwang, and Jong C Park. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pengcheng He, Jianfeng Gao, and Weizhu Chen. 2023 · 2023
Cited alongside, same era.
What we evaluate when we evaluate recommender systems: Understanding recommender systems’ performance using item response theory
Yang Liu, Alan Medlar, and Dorota Glowacka. 2023 · 2023
Cited alongside, same era.
Routing to the expert: Efficient reward-guided ensemble of large language models
Keming Lu, Hongyi Yuan, Runji Lin, Junyang Lin, Zheng Yuan, Chang Zhou, and Jingren Zhou. 2023 · 2023
Cited alongside, same era.
Large language model routing with benchmark datasets
Tal Shnitzer, Anthony Ou, Mírian Silva, Kate Soule, Yuekai Sun, Justin Solomon, Neil Thompson, and Mikhail Yurochkin. 2023 · 2023
Cited alongside, same era.
Can large language model comprehend ancient chinese? a preliminary test on aclue
Yixuan Zhang and Haonan Li. 2023 · 2023
Cited alongside, same era.
Fairlisa: fair user modeling with limited sensitive attributes information
Zheng Zhang, Qi Liu, Hao Jiang, Fei Wang, Yan Zhuang, Le Wu, Weibo Gao, and Enhong Chen. 2023 · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023 · 2023
Cited alongside, same era.
A survey on mixture of experts
Weilin Cai, Juyong Jiang, Fan Wang, Jing Tang, Sunghun Kim, and Jiayi Huang. 2024 · 2024
Cited alongside, same era.
Haonan Li, Yixuan Zhang, Fajri Koto, Yifei Yang, Hai Zhao, Yeyun Gong, Nan Duan, and Timothy Baldwin. 2024 · 2024
Later among the works it cites.
Routoo: Learning to route to large language models effectively
Alireza Mohammadshahi, Arshad Rafiq Shaikh, and Majid Yazdani. 2024 · 2024
Later among the works it cites.
Routellm: Learning to route llms with preference data
Isaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang, Tianhao Wu, Joseph E Gonzalez, M Waleed Kadous, and Ion Stoica. 2024 · 2024
Later among the works it cites.
Optimising calls to large language models with uncertainty-based two-tier selection
Guillem Ramírez, Alexandra Birch, and Ivan Titov. 2024 · 2024
Later among the works it cites.
Fly-swat or cannon? cost-effective language model choice via meta-modeling
Marija Šakota, Maxime Peyrard, and Robert West. 2024 · 2024
Later among the works it cites.
Dynamollm: Designing llm inference clusters for performance and energy efficiency
Jovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Josep Torrellas, and Esha Choukse. 2024 · 2024
Later among the works it cites.
Adaption-of-thought: Learning question difficulty improves large language models for reasoning
Mayi Xu, Yongqi Li, Ke Sun, and Tieyun Qian. 2024 · 2024
Later among the works it cites.
Decompose, analyze and rethink: Solving intricate problems with human-like reasoning cycle
Shangzi Xue, Zhenya Huang, Jiayu Liu, Xin Lin, Yuting Ning, Binbin Jin, Xin Li, and Qi Liu. 2024 · 2024
Later among the works it cites.
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al. 2024 · 2024
Later among the works it cites.
Towards accurate and fair cognitive diagnosis via monotonic data augmentation
Zheng Zhang, Wei Song, Qi Liu, Qingyang Mao, Yiyan Wang, Weibo Gao, Zhenya Huang, Shijin Wang, and Enhong Chen. 2024 · 2024
Later among the works it cites.
A comprehensive survey of large language models in management: Applications, challenges, and opportunities
Hongke Zhao, Likang Wu, Yuqing Shan, Zonghan Jin, Yuanpei Sui, Zipeng Liu, Nan Feng, Minqiang Li, and Wei Zhang. 2024 · 2024
Later among the works it cites.
Foundation model enhanced derivative-free cognitive diagnosis
Mingjia Li, Hong Qian, Jinglan Lv, Mengliang He, Wei Zhang, and Aimin Zhou. 2025 · 2025
Closest in time.
Debate on graph: a flexible and reliable reasoning framework for large language models
Jie Ma, Zhitao Gao, Qi Chai, Wangchun Sun, Pinghui Wang, Hongbin Pei, Jing Tao, Lingyun Song, Jun Liu, Chen Zhang, et al. 2025 · 2025
Closest in time.