Fetching the paper…
Reading the bibliography…
Accurate and consistent evaluation is crucial for decision-making across numerous fields, yet it remains a challenging task due to inherent subjectivity, variability, and scale.
Are Large Language Models Good at Utility Judgments?. In
Hengran Zhang, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke, Yixing Fan, and Xueqi Cheng. 2024c · 1951
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation. In
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries. In
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
The limits of automatic summarisation according to ROUGE. In
Natalie Schluter. 2017 · 2007
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Earlier work this paper cites.
Language Models are Few-Shot Learners. In
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Earlier work this paper cites.
GLM: General Language Model Pretraining with Autoregressive Blank Infilling. In
Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback. In
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe. 2022 · 2022
Earlier work this paper cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al · 2022
Earlier work this paper cites.
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. In
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022 · 2022
Earlier work this paper cites.
Fengshenbang 1.0: Being the foundation of chinese cognitive intelligence
Jiaxing Zhang, Ruyi Gan, Junjie Wang, Yuxiang Zhang, Lin Zhang, Ping Yang, Xinyu Gao, Ziwei Wu, Xiaoqun Dong, Junqing He, et al · 2022
Earlier work this paper cites.
Chatyuan: A large language model for dialogue in chinese and english
Liang Xu Xuanwei Zhang and Kangkang Zhao. 2022 · 2022
Earlier work this paper cites.
Do Language Models Know When They’re Hallucinating References?
Ayush Agrawal, Mirac Suzgun, Lester Mackey, and Adam Tauman Kalai. 2023 · 2023
Earlier work this paper cites.
L-eval: Instituting standardized evaluation for long context language models
Chenxin An, Shansan Gong, Ming Zhong, Xingjian Zhao, Mukai Li, Jun Zhang, Lingpeng Kong, and Xipeng Qiu. 2023 · 2023
Earlier work this paper cites.
Benchmarking Foundation Models with Language-Model-as-an-Examiner. In
Yushi Bai, Jiahao Ying, Yixin Cao, Xin Lv, Yuze He, Xiaozhi Wang, Jifan Yu, Kaisheng Zeng, Yijia Xiao, Haozhe Lyu, Jiayin Zhang, Juanzi Li, and Lei Hou. 2023 · 2023
Earlier work this paper cites.
CLAIR: Evaluating Image Captions with Large Language Models. In
David Chan, Suzanne Petryk, Joseph Gonzalez, Trevor Darrell, and John Canny. 2023b · 2023
Earlier work this paper cites.
Evaluating hallucinations in chinese large language models
Qinyuan Cheng, Tianxiang Sun, Wenwei Zhang, Siyin Wang, Xiangyang Liu, Mozhi Zhang, Junliang He, Mianqiu Huang, Zhangyue Yin, Kai Chen, et al · 2023
Earlier work this paper cites.
Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning. In
Antonia Creswell, Murray Shanahan, and Irina Higgins. 2023 · 2023
Earlier work this paper cites.
Holistic analysis of hallucination in gpt-4v (ision): Bias and interference challenges
Chenhang Cui, Yiyang Zhou, Xinyu Yang, Shirley Wu, Linjun Zhang, James Zou, and Huaxiu Yao. 2023 · 2023
Earlier work this paper cites.
RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment
Hanze Dong, Wei Xiong, Deepanshu Goyal, Yihan Zhang, Winnie Chow, Rui Pan, Shizhe Diao, Jipeng Zhang, Kashun Shum, and Tong Zhang. 2023 · 2023
Earlier work this paper cites.
AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback. In
Yann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023 · 2023
Earlier work this paper cites.
Chang Gao, Haiyun Jiang, Deng Cai, Shuming Shi, and Wai Lam. 2023a · 2023
Earlier work this paper cites.
Human-like summarization evaluation with chatgpt
Mingqi Gao, Jie Ruan, Renliang Sun, Xunjian Yin, Shiping Yang, and Xiaojun Wan. 2023b · 2023
Earlier work this paper cites.
Gemini: a family of highly capable multimodal models
Google. 2023 · 2023
Earlier work this paper cites.
LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models. In
Neel Guha, Julian Nyarko, Daniel E. Ho, Christopher Ré, Adam Chilton, Aditya K, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel N. Rockmore, Diego Zambrano, Dmitry Talisman, Enam Hoque, Faiz Surani, Frank Fagan, Galit Sarfaty, Gregory M. Dickinson, Haggai Porat, Jason Hegland, Jessica Wu, Joe Nudell, Joel Niklaus, John J. Nay, Jonathan H. Choi, Kevin Tobia, Margaret Hagan, Megan Ma, Michael A. Livermore, Nikon Rasumov-Rahe, Nils Holzenberger, Noam Kolt, Peter Henderson, Sean Rehaag, Sharad Goel, Shang Gao, Spencer Williams, Sunny Gandhi, Tom Zur, Varun Iyer, and Zehua Li. 2023 · 2023
Earlier work this paper cites.
Reasoning with Language Model is Planning with World Model. In
Shibo Hao, Yi Gu, Haodi Ma, Joshua Hong, Zhen Wang, Daisy Wang, and Zhiting Hu. 2023 · 2023
Earlier work this paper cites.
ChatGPT for scientific paper writing—promises and perils
Shijun He, Fan Yang, Jian-ping Zuo, and Ze-min Lin. 2023 · 2023
Earlier work this paper cites.
Towards Reasoning in Large Language Models: A Survey. In
Jie Huang and Kevin Chen-Chuan Chang. 2023 · 2023
Earlier work this paper cites.
On the humanity of conversational ai: Evaluating the psychological portrayal of llms. In
Jen-tse Huang, Wenxuan Wang, Eric John Li, Man Ho Lam, Shujie Ren, Youliang Yuan, Wenxiang Jiao, Zhaopeng Tu, and Michael Lyu. 2023 · 2023
Earlier work this paper cites.
Baseline defenses for adversarial attacks against aligned language models
Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping-yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein. 2023 · 2023
Earlier work this paper cites.
Deciphering “the language of nature”: A transformer-based language model for deleterious mutations in proteins
Theodore T. Jiang, Li Fang, and Kai Wang. 2023 · 2023
Earlier work this paper cites.
Maple: Multi-modal prompt learning. In
Muhammad Uzair Khattak, Hanoona Rasheed, Muhammad Maaz, Salman Khan, and Fahad Shahbaz Khan. 2023 · 2023
Earlier work this paper cites.
Prometheus: Inducing Fine-grained Evaluation Capability in Language Models
Seungone Kim, Jamin Shin, Yejin Cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, et al · 2023
Earlier work this paper cites.
Benchmarking cognitive biases in large language models as evaluators
Ryan Koo, Minhwa Lee, Vipul Raheja, Jong Inn Park, Zae Myung Kim, and Dongyeop Kang. 2023 · 2023
Earlier work this paper cites.
Improving Diversity of Demographic Representation in Large Language Models via Collective-Critiques and Self-Voting. In
Preethi Lahoti, Nicholas Blumm, Xiao Ma, Raghavendra Kotikalapudi, Sahitya Potluri, Qijun Tan, Hansa Srinivasan, Ben Packer, Ahmad Beirami, Alex Beutel, and Jilin Chen. 2023 · 2023
Earlier work this paper cites.
Examining query sentiment bias effects on search results in large language models. In
Alice Li and Luanne Sinnamon. 2023 · 2023
Earlier work this paper cites.
Generative judge for evaluating alignment
Junlong Li, Shichao Sun, Weizhe Yuan, Run-Ze Fan, Hai Zhao, and Pengfei Liu. 2023b · 2023
Earlier work this paper cites.
Evaluating Object Hallucination in Large Vision-Language Models. In
Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Xin Zhao, and Ji-Rong Wen. 2023a · 2023
Earlier work this paper cites.
Encouraging divergent thinking in large language models through multi-agent debate
Tian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang, Yan Wang, Rui Wang, Yujiu Yang, Zhaopeng Tu, and Shuming Shi. 2023 · 2023
Earlier work this paper cites.
Let’s verify step by step. In
Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. 2023 · 2023
Earlier work this paper cites.
LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models. In
Yen-Ting Lin and Yun-Nung Chen. 2023 · 2023
Earlier work this paper cites.
ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation. In
Zi Lin, Zihan Wang, Yongqi Tong, Yangkun Wang, Yuxin Guo, Yujia Wang, and Jingbo Shang. 2023 · 2023
Earlier work this paper cites.
Goal-Oriented Prompt Attack and Safety Evaluation for LLMs
Chengyuan Liu, Fubang Zhao, Lizhi Qing, Yangyang Kang, Changlong Sun, Kun Kuang, and Fei Wu. 2023e · 2023
Earlier work this paper cites.
Visual Instruction Tuning. In
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023c · 2023
Earlier work this paper cites.
Languages are rewards: Hindsight finetuning using human feedback
Hao Liu, Carmelo Sferrazza, and Pieter Abbeel. 2023d · 2023
Earlier work this paper cites.
Bioinformatics: Advancing biomedical discovery and innovation in the era of big data and artificial intelligence
Yuan Liu, Yamei Chen, and Leng Han. 2023a · 2023
Earlier work this paper cites.
G-Eval: NLG Evaluation using Gpt-4 with Better Human Alignment. In
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023b · 2023
Earlier work this paper cites.
MathVista: Evaluating Mathematical Reasoning in Visual Contexts
Yuxuan Lu, Xiaoyi Ding, Lingfeng Wang, Zheng Zhu, and Junxian He. 2023 · 2023
Earlier work this paper cites.
Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct
Haipeng Luo, Qingfeng Sun, Can Xu, Pu Zhao, Jianguang Lou, Chongyang Tao, Xiubo Geng, Qingwei Lin, Shifeng Chen, and Dongmei Zhang. 2023a · 2023
Earlier work this paper cites.
WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct
Haipeng Luo, Qingfeng Sun, Can Xu, Pu Zhao, Jianguang Lou, Chongyang Tao, Xiubo Geng, Qingwei Lin, Shifeng Chen, and Dongmei Zhang. 2023b · 2023
Earlier work this paper cites.
Zero-shot listwise document reranking with a large language model
Xueguang Ma, Xinyu Zhang, Ronak Pradeep, and Jimmy Lin. 2023 · 2023
Earlier work this paper cites.
Self-Refine: Iterative Refinement with Self-Feedback. In
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark. 2023 · 2023
Earlier work this paper cites.
FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation. In
Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023 · 2023
Earlier work this paper cites.
OpenAI. 2023 · 2023
Earlier work this paper cites.
Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization. In
Rajkumar Ramamurthy, Prithviraj Ammanabrolu, Kianté Brantley, Jack Hessel, Rafet Sifa, Christian Bauckhage, Hannaneh Hajishirzi, and Yejin Choi. 2023 · 2023
Earlier work this paper cites.
Retrieval-based Evaluation for LLMs: A Case Study in Korean Legal QA. In
Cheol Ryu, Seolhwa Lee, Subeen Pang, Chanyeol Choi, Hojun Choi, Myeonggee Min, and Jy-Yong Sohn. 2023 · 2023
Earlier work this paper cites.
Verbosity bias in preference labeling by large language models
Keita Saito, Akifumi Wachi, Koki Wataoka, and Youhei Akimoto. 2023 · 2023
Earlier work this paper cites.
Languagempc: Large language models as decision makers for autonomous driving
Hao Sha, Yao Mu, Yuxuan Jiang, Li Chen, Chenfeng Xu, Ping Luo, Shengbo Eben Li, Masayoshi Tomizuka, Wei Zhan, and Mingyu Ding. 2023 · 2023
Earlier work this paper cites.
Reflexion: language agents with verbal reinforcement learning. In
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023 · 2023
Earlier work this paper cites.
Evaluation Metrics in the Era of GPT-4: Reliably Evaluating Large Language Models on Sequence to Sequence Tasks. In
Andrea Sottana, Bin Liang, Kai Zou, and Zheng Yuan. 2023 · 2023
Earlier work this paper cites.
Think-on-graph: Deep and responsible reasoning of large language model with knowledge graph
Jiashuo Sun, Chengjin Xu, Lumingyuan Tang, Saizhuo Wang, Chen Lin, Yeyun Gong, Heung-Yeung Shum, and Jian Guo. 2023c · 2023
Earlier work this paper cites.
Aligning large multimodal models with factually augmented rlhf
Zhiqing Sun, Sheng Shen, Shengcao Cao, Haotian Liu, Chunyuan Li, Yikang Shen, Chuang Gan, Liang-Yan Gui, Yu-Xiong Wang, Yiming Yang, et al · 2023
Earlier work this paper cites.
Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human Supervision. In
Zhiqing Sun, Yikang Shen, Qinhong Zhou, Hongxin Zhang, Zhenfang Chen, David D. Cox, Yiming Yang, and Chuang Gan. 2023b · 2023
Earlier work this paper cites.
Stanford Alpaca: An Instruction-following LLaMA model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023 · 2023
Earlier work this paper cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Earlier work this paper cites.
Vicuna: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Earlier work this paper cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Earlier work this paper cites.
Evaluation and analysis of hallucination in large vision-language models
Junyang Wang, Yiyang Zhou, Guohai Xu, Pengcheng Shi, Chenlin Zhao, Haiyang Xu, Qinghao Ye, Ming Yan, Ji Zhang, Jihua Zhu, et al · 2023
Earlier work this paper cites.
Math-Shepherd: Verify and Reinforce LLMs Step-by-Step without Human Annotations
Peiyi Wang, Lei Li, Zhihong Shao, Rui Xu, Damai Dai, Yifei Li, Deli Chen, Yu Wu, and Zhifang Sui. 2023b · 2023
Earlier work this paper cites.
Shepherd: A Critic for Language Model Generation
Tianlu Wang, Ping Yu, Xiaoqing Ellen Tan, Sean O’Brien, Ramakanth Pasunuru, Jane Dwivedi-Yu, Olga Golovneva, Luke Zettlemoyer, Maryam Fazel-Zarandi, and Asli Celikyilmaz. 2023c · 2023
Earlier work this paper cites.
Self-Instruct: Aligning Language Models with Self-Generated Instructions. In
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2023a · 2023
Earlier work this paper cites.
PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Yidong Wang, Zhuohao Yu, Zhengran Zeng, Linyi Yang, Cunxiang Wang, Hao Chen, Chaoya Jiang, Rui Xie, Jindong Wang, Xing Xie, et al · 2023
Cited alongside, same era.
Filling in missing pieces in the co-development of artificial intelligence and environmental science
Zhenyu Wang, Jin Zhang, Pei Hua, Yuanzheng Cui, Chunhui Lu, Xiaojun Wang, Qiuwen Chen, and Peter Krebs. 2023e · 2023
Cited alongside, same era.
Large Language Models are Better Reasoners with Self-Verification. In
Yixuan Weng, Minjun Zhu, Fei Xia, Bin Li, Shizhu He, Shengping Liu, Bin Sun, Kang Liu, and Jun Zhao. 2023 · 2023
Cited alongside, same era.
Large language models are diverse role-players for summarization evaluation. In
Ning Wu, Ming Gong, Linjun Shou, Shining Liang, and Daxin Jiang. 2023 · 2023
Cited alongside, same era.
INSTRUCTSCORE: Towards Explainable Text Generation Evaluation with Automatic Feedback. In
Wenda Xu, Danqing Wang, Liangming Pan, Zhenqiao Song, Markus Freitag, William Wang, and Lei Li. 2023b · 2023
HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition. In
Yuxuan Liu, Tianchi Yang, Shaohan Huang, Zihan Zhang, Haizhen Huang, Furu Wei, Weiwei Deng, Feng Sun, and Qi Zhang. 2024a · 2024
Closest in time.
LLM Comparative Assessment: Zero-shot NLG Evaluation through Pairwise Comparisons using Large Language Models. In
Adian Liusie, Potsawee Manakul, and Mark Gales. 2024 · 2024
Closest in time.
Artificial intelligence for life sciences: A comprehensive guide and future trends
Ming Luo, Wenyu Yang, Long Bai, Lin Zhang, Jia-Wei Huang, Yinhong Cao, Yuhua Xie, Liping Tong, Haibo Zhang, Lei Yu, et al · 2024
Closest in time.
Leveraging Large Language Models for Relevance Judgments in Legal Case Retrieval
Shengjie Ma, Chong Chen, Qi Chu, and Jiaxin Mao. 2024 · 2024
Closest in time.
Evaluating the Performance of Large Language Models via Debates
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Artificial intelligence for science—bridging data to wisdom
Yongjun Xu, Fei Wang, Zhulin An, Qi Wang, and Zhao Zhang. 2023a · 2023
Cited alongside, same era.
Auto-gpt for online decision making: Benchmarks and additional opinions
Hui Yang, Sifu Yue, and Yunzhong He. 2023 · 2023
Cited alongside, same era.
Tree of Thoughts: Deliberate Problem Solving with Large Language Models. In
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. 2023a · 2023
Cited alongside, same era.
ReAct: Synergizing Reasoning and Acting in Language Models. In
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R. Narasimhan, and Yuan Cao. 2023b · 2023
Cited alongside, same era.
SelFee: Iterative Self-Revising LLM Empowered by Self-Feedback Generation
Seonghyeon Ye, Yongrae Jo, Doyoung Kim, Sungdong Kim, Hyeonbin Hwang, and Minjoon Seo. 2023 · 2023
Cited alongside, same era.
Design and Implementation of an LLM system to Improve Response Time for SMEs Technology Credit Evaluation
Sungwook Yoon. 2023 · 2023
Cited alongside, same era.
Advanced prompting as a catalyst: Empowering large language models in the management of gastrointestinal cancers
J Yuan, P Bao, Z Chen, M Yuan, J Zhao, J Pan, Y Xie, Y Cao, Y Wang, Z Wang, et al · 2023
Cited alongside, same era.
Behrad Moniri, Hamed Hassani, and Edgar Dobriban. 2024 · 2024
Closest in time.
Offsetbias: Leveraging debiased data for tuning evaluators
Junsoo Park, Seungyeon Jwa, Meiying Ren, Daeyoung Kim, and Sanghyuk Choi. 2024 · 2024
Closest in time.
Agent q: Advanced reasoning and learning for autonomous ai agents
Pranav Putta, Edmund Mills, Naman Garg, Sumeet Motwani, Chelsea Finn, Divyansh Garg, and Rafael Rafailov. 2024 · 2024
Closest in time.
Large Language Models are Effective Text Rankers with Pairwise Ranking Prompting. In
Zhen Qin, Rolf Jagerman, Kai Hui, Honglei Zhuang, Junru Wu, Le Yan, Jiaming Shen, Tianqi Liu, Jialu Liu, Donald Metzler, Xuanhui Wang, and Michael Bendersky. 2024 · 2024
Closest in time.
Promoting interactions between cognitive science and large language models
Youzhi Qu, Penghui Du, Wenxin Che, Chen Wei, Chi Zhang, Wanli Ouyang, Yatao Bian, Feiyang Xu, Bin Hu, Kai Du, et al · 2024
Closest in time.
Evaluating rag-fusion with ragelo: an automated elo-based framework
Zackary Rackauckas, Arthur Câmara, and Jakub Zavrel. 2024 · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2024 · 2024
Closest in time.
Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment
Vyas Raina, Adian Liusie, and Mark Gales. 2024 · 2024
Closest in time.
Constructing domain-specific evaluation sets for llm-as-a-judge
Ravi Raju, Swayambhoo Jain, Bo Li, Jonathan Li, and Urmish Thakkar. 2024 · 2024
Closest in time.
Branch-Solve-Merge Improves Large Language Model Evaluation and Generation. In
Swarnadeep Saha, Omer Levy, Asli Celikyilmaz, Mohit Bansal, Jason Weston, and Xian Li. 2024 · 2024
Closest in time.
Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
Amrith Setlur, Chirag Nagpal, Adam Fisch, Xinyang Geng, Jacob Eisenstein, Rishabh Agarwal, Alekh Agarwal, Jonathan Berant, and Aviral Kumar. 2024 · 2024
Closest in time.
Lin Shi, Weicheng Ma, and Soroush Vosoughi. 2024 · 2024
Closest in time.
Drug development in the AI era: AlphaFold 3 is coming!
Yi Shi. 2024 · 2024
Closest in time.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. 2024 · 2024
Closest in time.
Llm-as-a-judge & reward model: What they can and cannot do
Guijin Son, Hyunwoo Ko, Hoyoung Lee, Yewon Kim, and Seunghyeok Hong. 2024 · 2024
Closest in time.
Preference Ranking Optimization for Human Alignment. In
Feifan Song, Bowen Yu, Minghao Li, Haiyang Yu, Fei Huang, Yongbin Li, and Houfeng Wang. 2024a · 2024
Closest in time.
Automated Essay Scoring and Revising Based on Open-Source Large Language Models
Yishen Song, Qianta Zhu, Huaibo Wang, and Qinhua Zheng. 2024b · 2024
Closest in time.
Blinded by Generated Contexts: How Language Models Merge Generated and Retrieved Contexts When Knowledge Conflicts?. In
Hexiang Tan, Fei Sun, Wanli Yang, Yuanzhuo Wang, Qi Cao, and Xueqi Cheng. 2024a · 2024
Closest in time.
Judgebench: A benchmark for evaluating llm-based judges
Sijun Tan, Siyuan Zhuang, Kyle Montgomery, William Y Tang, Alejandro Cuadron, Chenguang Wang, Raluca Ada Popa, and Ion Stoica. 2024b · 2024
Closest in time.
Self-Retrieval: Building an Information Retrieval System with One Large Language Model
Qiaoyu Tang, Jiawei Chen, Bowen Yu, Yaojie Lu, Cheng Fu, Haiyang Yu, Hongyu Lin, Fei Huang, Ben He, Xianpei Han, et al · 2024
Closest in time.
Found in the Middle: Permutation Self-Consistency Improves Listwise Ranking in Large Language Models. In
Raphael Tang, Crystina Zhang, Xueguang Ma, Jimmy Lin, and Ferhan Ture. 2024c · 2024
Closest in time.
LLMs in medicine: The need for advanced evaluation systems for disruptive technologies
Yi-Da Tang, Er-Dan Dong, and Wen Gao. 2024b · 2024
Closest in time.
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Aman Singh Thakur, Kartik Choudhary, Venkat Srinik Ramayapally, Sankaran Vaidyanathan, and Dieuwke Hupkes. 2024 · 2024
Closest in time.
MacGyver: Are Large Language Models Creative Problem Solvers?. In
Yufei Tian, Abhilasha Ravichander, Lianhui Qin, Ronan Le Bras, Raja Marjieh, Nanyun Peng, Yejin Choi, Thomas Griffiths, and Faeze Brahman. 2024 · 2024
Closest in time.
DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving. In
Yuxuan Tong, Xiwen Zhang, Rui Wang, Ruidong Wu, and Junxian He. 2024 · 2024
Closest in time.
Reft: Reasoning with reinforced fine-tuning. In
Luong Trung, Xinbo Zhang, Zhanming Jie, Peng Sun, Xiaoran Jin, and Hang Li. 2024 · 2024
Closest in time.
Current and Future State of Evaluation of Large Language Models for Clinical Summarization
Daan Van Veen, Akshay Kolluri, and et al. 2024 · 2024
Closest in time.
Halu-J: Critique-Based Hallucination Judge
Binjie Wang, Steffi Chern, Ethan Chern, and Pengfei Liu. 2024a · 2024
Closest in time.
BioRAG: A RAG-LLM Framework for Biological Question Reasoning
Chengrui Wang, Qingqing Long, Xiao Meng, Xunxin Cai, Chengjun Wu, Zhen Meng, Xuezhi Wang, and Yuanchun Zhou. 2024d · 2024
Closest in time.
Tianlu Wang, Ilia Kulikov, Olga Golovneva, Ping Yu, Weizhe Yuan, Jane Dwivedi-Yu, Richard Yuanzhe Pang, Maryam Fazel-Zarandi, Jason Weston, and Xian Li. 2024b · 2024
Closest in time.
DHP Benchmark: Are LLMs Good NLG Evaluators?
Yicheng Wang, Jiayi Yuan, Yu-Neng Chuang, Zhuoer Wang, Yingchi Liu, Mark Cusick, Param Kulkarni, Zhengping Ji, Yasser Ibrahim, and Xia Hu. 2024f · 2024
Closest in time.
Speculative rag: Enhancing retrieval augmented generation through drafting
Zilong Wang, Zifeng Wang, Long Le, Huaixiu Steven Zheng, Swaroop Mishra, Vincent Perot, Yuwei Zhang, Anush Mattapalli, Ankur Taly, Jingbo Shang, et al · 2024
Closest in time.
AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models
Yuhang Wu, Wenmeng Yu, Yean Cheng, Yan Wang, Xiaohan Zhang, Jiazheng Xu, Ming Ding, and Yuxiao Dong. 2024 · 2024
Closest in time.
Beyond Accuracy: Evaluating Logical Coherence of Mathematical Reasoning in Large Language Models
Zheng Xia, Li Qian, and Hao Wang. 2024 · 2024
Closest in time.
Sorry-bench: Systematically evaluating large language model safety refusal behaviors
Tinghao Xie, Xiangyu Qi, Yi Zeng, Yangsibo Huang, Udari Madhushani Sehwag, Kaixuan Huang, Luxi He, Boyi Wei, Dacheng Li, Ying Sheng, et al · 2024
Closest in time.
Improving Model Factuality with Fine-grained Critique-based Evaluator
Yiqing Xie, Wenxuan Zhou, Pradyot Prakash, Di Jin, Yuning Mao, Quintin Fettes, Arya Talebzadeh, Sinong Wang, Han Fang, Carolyn Rose, et al · 2024
Closest in time.
LLaVA-Critic: Learning to Evaluate Multimodal Models
Tianyi Xiong, Xiyao Wang, Dong Guo, Qinghao Ye, Haoqi Fan, Quanquan Gu, Heng Huang, and Chunyuan Li. 2024 · 2024
Closest in time.
Academically intelligent LLMs are not necessarily socially intelligent
Ruoxi Xu, Hongyu Lin, Xianpei Han, Le Sun, and Yingfei Sun. 2024a · 2024
Closest in time.
Artificial intelligence is restructuring a new world
Yongjun Xu, Fei Wang, and Tangtang Zhang. 2024b · 2024
Closest in time.
Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
An Yang, Beichen Zhang, Binyuan Hui, Bofei Gao, Bowen Yu, Chengpeng Li, Dayiheng Liu, Jianhong Tu, Jingren Zhou, Junyang Lin, Keming Lu, Mingfeng Xue, Runji Lin, Tianyu Liu, Xingzhang Ren, and Zhenru Zhang. 2024b · 2024
Closest in time.
UCFE: A User-Centric Financial Expertise Benchmark for Large Language Models
Yuzhe Yang, Yifei Zhang, Yan Hu, Yilin Guo, Ruoli Gan, Yueru He, Mingcong Lei, Xiao Zhang, Haining Wang, Qianqian Xie, et al · 2024
Closest in time.
Zonghai Yao, Aditya Parashar, Huixue Zhou, Won Seok Jang, Feiyun Ouyang, Zhichao Yang, and Hong Yu. 2024 · 2024
Closest in time.
Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
Jiayi Ye, Yanbo Wang, Yue Huang, Dongping Chen, Qihui Zhang, Nuno Moniz, Tian Gao, Werner Geyer, Chao Huang, Pin-Yu Chen, et al · 2024
Closest in time.
xFinder: Robust and Pinpoint Answer Extraction for Large Language Models
Qingchen Yu, Zifan Zheng, Shichao Song, Zhiyu Li, Feiyu Xiong, Bo Tang, and Ding Chen. 2024b · 2024
Closest in time.
Yangyang Yu, Zhiyuan Yao, Haohang Li, Zhiyang Deng, Yupeng Cao, Zhi Chen, Jordan W Suchow, Rong Liu, Zhenyu Cui, Denghui Zhang, et al · 2024
Closest in time.
Self-rewarding language models
Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Sainbayar Sukhbaatar, Jing Xu, and Jason Weston. 2024 · 2024
Closest in time.
Kamer Ali Yuksel and Hassan Sawaf. 2024 · 2024
Closest in time.
Evaluating and improving tool-augmented computation-intensive math reasoning
Beichen Zhang, Kun Zhou, Xilin Wei, Xin Zhao, Jing Sha, Shijin Wang, and Ji-Rong Wen. 2024d · 2024
Closest in time.
Hengyuan Zhang, Yanru Wu, Dawei Li, Sak Yang, Rui Zhao, Yong Jiang, and Fei Tan. 2024b · 2024
Closest in time.
Evaluation Ethics of LLMs in Legal Domain
Ruizhe Zhang, Haitao Li, Yueyue Wu, Qingyao Ai, Yiqun Liu, Min Zhang, and Shaoping Ma. 2024a · 2024
Closest in time.
Revolutionizing finance with llms: An overview of applications and insights
Huaqin Zhao, Zhengliang Liu, Zihao Wu, Yiwei Li, Tianze Yang, Peng Shu, Shaochen Xu, Haixing Dai, Lin Zhao, Gengchen Mai, et al · 2024
Closest in time.
CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?
Yuwei Zhao, Ziyang Luo, Yuchen Tian, Hongzhan Lin, Weixiang Yan, Annan Li, and Jing Ma. 2024b · 2024
Closest in time.
Cheating automatic llm benchmarks: Null models achieve high win rates
Xiaosen Zheng, Tianyu Pang, Chao Du, Qian Liu, Jing Jiang, and Min Lin. 2024a · 2024
Closest in time.
Harnessing the power of artificial intelligence to combat infectious diseases: Progress, challenges, and future outlook
Hang-Yu Zhou, Yaling Li, Jia-Ying Li, Jing Meng, and Aiping Wu. 2024a · 2024
Closest in time.
Self-discover: Large language models self-compose reasoning structures
Pei Zhou, Jay Pujara, Xiang Ren, Xinyun Chen, Heng-Tze Cheng, Quoc V Le, Ed H Chi, Denny Zhou, Swaroop Mishra, and Huaixiu Steven Zheng. 2024b · 2024
Closest in time.
Beyond Yes and No: Improving Zero-Shot LLM Rankers via Scoring Fine-Grained Relevance Labels. In
Honglei Zhuang, Zhen Qin, Kai Hui, Junru Wu, Le Yan, Xuanhui Wang, and Michael Bendersky. 2024a · 2024
Closest in time.
Agent-as-a-Judge: Evaluate Agents with Agents
Mingchen Zhuge, Changsheng Zhao, Dylan Ashley, Wenyi Wang, Dmitrii Khizbullin, Yunyang Xiong, Zechun Liu, Ernie Chang, Raghuraman Krishnamoorthi, Yuandong Tian, et al · 2024
Closest in time.
Think-j: Learning to think for generative llm-as-a-judge
Hui Huang, Yancheng He, Hongli Zhou, Rui Zhang, Wei Liu, Weixun Wang, Wenbo Su, Bo Zheng, and Jiaheng Liu. 2025 · 2025
Closest in time.
Hadi Mohammadi, Anastasia Giachanou, and Ayoub Bagheri. 2025 · 2025
Closest in time.
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
Yidong Wang, Yunze Song, Tingyuan Zhu, Xuanwang Zhang, Zhuohao Yu, Hao Chen, Chiyu Song, Qiufeng Wang, Cunxiang Wang, Zhen Wu, et al · 2025
Closest in time.
Evaluating the ability of large language models to emulate personality
Yilei Wang, Jiabao Zhao, Deniz S Ones, Liang He, and Xin Xu. 2025b · 2025
Closest in time.
Improve llm-as-a-judge ability as a general ability
Jiachen Yu, Shaoning Sun, Xiaohui Hu, Jiaxu Yan, Kaidong Yu, and Xuelong Li. 2025 · 2025
Closest in time.
Sentient Agent as a Judge: Evaluating Higher-Order Social Cognition in Large Language Models
Bang Zhang, Ruotian Ma, Qingxuan Jiang, Peisong Wang, Jiaqi Chen, Zheng Xie, Xingyu Chen, Yue Wang, Fanghua Ye, Jian Li, et al · 2025
Closest in time.
Crowd comparative reasoning: Unlocking comprehensive evaluations for llm-as-a-judge
Qiyuan Zhang, Yufei Wang, Yuxin Jiang, Liangyou Li, Chuhan Wu, Yasheng Wang, Xin Jiang, Lifeng Shang, Ruiming Tang, Fuyuan Lyu, et al · 2025
Closest in time.
One Token to Fool LLM-as-a-Judge
Yulai Zhao, Haolin Liu, Dian Yu, SY Kung, Haitao Mi, and Dong Yu. 2025 · 2025
Closest in time.
TrueTeacher: Learning Factual Consistency Evaluation with Large Language Models. In
Zorik Gekhman, Jonathan Herzig, Roee Aharoni, Chen Elkind, and Idan Szpektor. 2023 · 2070
Closest in time.