Fetching the paper…
Reading the bibliography…
LLM ensembles are widely used for LLM judges.
Information aggregation, rationality, and the condorcet jury theorem
Austen-Smith, D. and Banks, J. S · 1996
Earlier work this paper cites.
A tutorial on conformal prediction
Shafer, G. and Vovk, V · 2008
Earlier work this paper cites.
Essai sur l’application de l’analyse à la probabilité des décisions rendues à la pluralité des voix
De Condorcet, N. et al · 2014
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
Lin, S., Hilton, J., and Evans, O · 2021
Earlier work this paper cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al · 2022
Earlier work this paper cites.
Angelopoulos, A. N., Bates, S., Fisch, A., Lei, L., and Schuster, T · 2023
Earlier work this paper cites.
Conformal prediction: a unified review of theory and new challenges
Fontana, M., Zeni, G., and Vantini, S · 2023
Earlier work this paper cites.
Encouraging divergent thinking in large language models through multi-agent debate
Liang, T., He, Z., Jiao, W., Wang, X., Wang, Y., Wang, R., Yang, Y., Shi, S., and Tu, Z · 2023
Earlier work this paper cites.
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K · 2023
Earlier work this paper cites.
G-eval: NLG evaluation using gpt-4 with better human alignment
Liu, Y., Iter, D., Xu, Y., Wang, S., Xu, R., and Zhu, C · 2023
Earlier work this paper cites.
FActScore: Fine-grained atomic evaluation of factual precision in long form text generation
Min, S., Krishna, K., Lyu, X., Lewis, M., Yih, W.-t., Koh, P., Iyyer, M., Zettlemoyer, L., and Hajishirzi, H · 2023
Earlier work this paper cites.
Large language models are not fair evaluators, 2023
Wang, P., Li, L., Chen, L., Cai, Z., Zhu, D., Lin, B., Cao, Y., Liu, Q., Liu, T., and Sui, Z · 2023
Earlier work this paper cites.
Tram: Benchmarking temporal reasoning for large language models
Wang, Y. and Zhao, Y · 2023
Earlier work this paper cites.
Evaluating large language models at evaluating instruction following
Zeng, Z., Yu, J., Gao, T., Meng, Y., Goyal, T., and Chen, D · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al · 2023
Cited alongside, same era.
Ice-score: Instructing large language models to evaluate code
Zhuo, T. Y · 2023
Cited alongside, same era.
Cai, Z., Cao, M., Chen, H., Chen, K., Chen, K., Chen, X., Chen, X., Chen, Z., Chen, Z., Chu, P., et al · 2024
Cited alongside, same era.
Llm evaluators recognize and favor their own generations, 2024
Panickssery, A., Bowman, S. R., and Feng, S · 2024
Later among the works it cites.
Offsetbias: Leveraging debiased data for tuning evaluators, 2024
Park, J., Jwa, S., Ren, M., Kim, D., and Choi, S · 2024
Later among the works it cites.
Wisdom of the silicon crowd: Llm ensemble prediction capabilities rival human crowd accuracy, 2024
Schoenegger, P., Tuminauskaite, I., Park, P. S., and Tetlock, P. E · 2024
Later among the works it cites.
Judgebench: A benchmark for evaluating llm-based judges
Tan, S., Zhuang, S., Montgomery, K., Tang, W. Y., Cuadron, A., Wang, C., Popa, R. A., and Stoica, I · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chen, B., Wang, X., Peng, S., Litschko, R., Korhonen, A., and Plank, B · 2024
Cited alongside, same era.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Cited alongside, same era.
Length-controlled alpacaeval: A simple way to debias automatic evaluators, 2024
Dubois, Y., Galambosi, B., Liang, P., and Hashimoto, T. B · 2024
Cited alongside, same era.
Benchmarking cognitive biases in large language models as evaluators, 2024
Koo, R., Lee, M., Raheja, V., Park, J. I., Kim, Z. M., and Kang, D · 2024
Cited alongside, same era.
Rewardbench: Evaluating reward models for language modeling
Lambert, N., Pyatkin, V., Morrison, J., Miranda, L., Lin, B. Y., Chandu, K., Dziri, N., Kumar, S., Zick, T., Choi, Y., et al · 2024
Cited alongside, same era.
Rlaif vs. rlhf: Scaling reinforcement learning from human feedback with ai feedback, 2024
Lee, H., Phatale, S., Mansoor, H., Mesnard, T., Ferret, J., Lu, K., Bishop, C., Hall, E., Carbune, V., Rastogi, A., and Prakash, S · 2024
Cited alongside, same era.
Halludial: A large-scale benchmark for automatic dialogue-level hallucination evaluation
Luo, W., Shen, T., Li, W., Peng, G., Xuan, R., Wang, H., and Yang, X · 2024
Cited alongside, same era.
Language models with conformal factuality guarantees, 2024
Mohri, C. and Hashimoto, T · 2024
Cited alongside, same era.
Wang, T., Kulikov, I., Golovneva, O., Yu, P., Yuan, W., Dwivedi-Yu, J., Pang, R. Y., Fazel-Zarandi, M., Weston, J., and Li, X · 2024
Later among the works it cites.
Meta-rewarding language models: Self-improving alignment with llm-as-a-meta-judge, 2024
Wu, T., Yuan, W., Golovneva, O., Xu, J., Tian, Y., Jiao, J., Weston, J., and Sukhbaatar, S · 2024
Later among the works it cites.
Mitigating llm hallucinations via conformal abstention, 2024
Yadkori, Y. A., Kuzborskij, I., Stutz, D., György, A., Fisch, A., Doucet, A., Beloshapka, I., Weng, W.-H., Yang, Y.-Y., Szepesvári, C., Cemgil, A. T., and Tomasev, N · 2024
Later among the works it cites.
Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., et al · 2024
Later among the works it cites.
Kieval: A knowledge-grounded interactive evaluation framework for large language models, 2024
Yu, Z., Gao, C., Yao, W., Wang, Y., Ye, W., Wang, J., Xie, X., Zhang, Y., and Zhang, S · 2024
Later among the works it cites.
A comprehensive analysis of the effectiveness of large language models as automatic dialogue evaluators
Zhang, C., D’Haro, L. F., Chen, Y., Zhang, M., and Li, H · 2024
Later among the works it cites.
Calderon, N., Reichart, R., and Dror, R · 2025
Closest in time.
Nv-embed: Improved techniques for training llms as generalist embedding models, 2025
Lee, C., Roy, R., Xu, M., Raiman, J., Shoeybi, M., Catanzaro, B., and Ping, W · 2025
Closest in time.
Ensemble of large language models for curated labeling and rating of free-text data, 2025
Qiu, J., Guo, D., Natalie, P., Noelle, P., Cheri, L., and Henry, T. R · 2025
Closest in time.
The lessons of developing process reward models in mathematical reasoning, 2025
Zhang, Z., Zheng, C., Wu, Y., Zhang, B., Lin, R., Yu, B., Liu, D., Zhou, J., and Lin, J · 2025
Closest in time.