Fetching the paper…
Reading the bibliography…
As Large Language Models (LLMs) are increasingly deployed in highly specialized vertical domains, the evaluation of their domain-specific performance becomes critical.
The probabilistic relevance framework: Bm25 and beyond
Stephen Robertson, Hugo Zaragoza, and 1 others. 2009 · 2009
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Earlier work this paper cites.
Crimekgassitant
Huanyong Liu. 2018 · 2018
Earlier work this paper cites.
Cail2018: A large-scale legal dataset for judgment prediction
Chaojun Xiao, Haoxi Zhong, Zhipeng Guo, Cunchao Tu, Zhiyuan Liu, Maosong Sun, Yansong Feng, Xianpei Han, Zhen Hu, and Heng Wang. 2018 · 2018
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, and Reiichiro Nakano. 2021 · 2021
Earlier work this paper cites.
Diseasekg
Chen Peng, Jun Zhang, Yanchao Xu, and Jianan Yang. 2021 · 2021
Earlier work this paper cites.
Rlprompt: Optimizing discrete text prompts with reinforcement learning
Mingkai Deng, Jianyu Wang, Cheng-Ping Hsieh, Yihan Wang, Han Guo, Tianmin Shu, Meng Song, Eric Xing, and Zhiting Hu. 2022 · 2022
Earlier work this paper cites.
On the opportunities and risks of foundation models for natural language processing in radiology
Walter F Wiggins and Ali S Tejani. 2022 · 2022
Earlier work this paper cites.
Leven: A large-scale chinese legal event detection dataset
Feng Yao, Chaojun Xiao, Xiaozhi Wang, Zhiyuan Liu, Lei Hou, Cunchao Tu, Juanzi Li, Yun Liu, Weixing Shen, and Maosong Sun. 2022 · 2022
Earlier work this paper cites.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang. 2023 · 2023
Earlier work this paper cites.
Yirong Chen, Zhenyu Wang, Xiaofen Xing, Zhipei Xu, Kai Fang, Junhong Wang, Sihang Li, Jieling Wu, Qi Liu, and Xiangmin Xu. 2023 · 2023
Earlier work this paper cites.
Legalbench: A collaboratively built benchmark for measuring legal reasoning in large language models
Neel Guha, Julian Nyarko, Daniel Ho, Christopher Ré, Adam Chilton, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel Rockmore, and Diego Zambrano. 2023 · 2023
Earlier work this paper cites.
Evaluating large language models: A comprehensive survey
Zishan Guo, Renren Jin, Chuang Liu, Yufei Huang, Dan Shi, Linhao Yu, Yan Liu, Jiaxuan Li, Bojian Xiong, and Deyi Xiong. 2023 · 2023
Cited alongside, same era.
Cmmlu: Measuring massive multitask language understanding in chinese
Haonan Li, Yixuan Zhang, Fajri Koto, Yifei Yang, Hai Zhao, Yeyun Gong, Nan Duan, and Timothy Baldwin. 2023 · 2023
Cited alongside, same era.
Bcembedding: Bilingual and crosslingual embedding for rag
Inc. NetEase Youdao. 2023 · 2023
Cited alongside, same era.
Pillow: Enhancing efficient instruction fine-tuning via prompt matching
Zhenting Qi, Xiaoyu Tan, Shaojie Shi, Chao Qu, Yinghui Xu, and Yuan Qi. 2023 · 2023
Cited alongside, same era.
Query-dependent prompt evaluation and optimization with offline inverse rl
Hao Sun, Alihan Hüyük, and Mihaela van der Schaar. 2023 · 2023
Cited alongside, same era.
M3-embedding: Multi-linguality, multi-functionality, multi-granularity text embeddings through self-knowledge distillation
Jianlyu Chen, Shitao Xiao, Peitian Zhang, Kun Luo, Defu Lian, and Zheng Liu. 2024 · 2024
Closest in time.
A comprehensive survey on evaluating large language model applications in the medical industry
Yining Huang, Keke Tang, and Meilian Chen. 2024 · 2024
Closest in time.
Aaron Hurst, Adam Lerer, Adam P. Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, and 1 others. 2024 · 2024
Closest in time.
Prewrite: Prompt rewriting with reinforcement learning
Weize Kong, Spurthi Hombaiah, Mingyang Zhang, Qiaozhu Mei, and Michael Bendersky. 2024 · 2024
Closest in time.
Mt-eval: A multi-turn capabilities evaluation benchmark for large language models
Wai-Chung Kwan, Xingshan Zeng, Yuxin Jiang, Yufei Wang, Liangyou Li, Lifeng Shang, Xin Jiang, Qun Liu, and Kam-Fai Wong. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Caregpt: Medical llm, open source driven for a healthy future
Rongsheng Wang, Ruizhe Zhou, Haoming Chen, Yapeng Wang, and Tao Tan. 2023 · 2023
Cited alongside, same era.
Bloomberggpt: A large language model for finance
Shijie Wu, Ozan Irsoy, Steven Lu, Vadim Dabravolski, Mark Dredze, Sebastian Gehrmann, Prabhanjan Kambadur, David Rosenberg, and Gideon Mann. 2023 · 2023
Cited alongside, same era.
Doctorglm: Fine-tuning your chinese doctor is not a herculean task
Honglin Xiong, Sheng Wang, Yitao Zhu, Zihao Zhao, Yuxiao Liu, Linlin Huang, Qian Wang, and Dinggang Shen. 2023 · 2023
Cited alongside, same era.
Tempera: Test-time prompt editing via reinforcement learning
Tianjun Zhang, Xuezhi Wang, Denny Zhou, Dale Schuurmans, and Joseph E Gonzalez. 2023 · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, and Eric Xing. 2023 · 2023
Cited alongside, same era.
Mt-bench-101: A fine-grained benchmark for evaluating large language models in multi-turn dialogues
Ge Bai, Jie Liu, Xingyuan Bu, Yancheng He, Jiaheng Liu, Zhanhui Zhou, Zhuoran Lin, Wenbo Su, Tiezheng Ge, and Bo Zheng. 2024 · 2024
Cited alongside, same era.
Shanghai housing provident fund
Shanghai Provident Fund Management Center. 2024 · 2024
Cited alongside, same era.
Closest in time.
Introducing llama 3.1: Our most capable models to date
Meta. 2024 · 2024
Closest in time.
Mistral.ai news: Ministraux
Mistral. 2024 · 2024
Closest in time.
A survey of useful llm evaluation
Jilun Peng, Sijia Cheng, Egil Diau, Yungyu Shih, Poheng Chen, Yenting Lin, and Yunnung Chen. 2024 · 2024
Closest in time.
Qwen2.5: A party of foundation models
Qwen. 2024 · 2024
Closest in time.
Zhongjing: Enhancing the chinese medical capabilities of large language model through expert feedback and real-world multi-turn dialogue
Songhua Yang, Hanjie Zhao, Senbin Zhu, Guangyu Zhou, Hongfei Xu, Yuxiang Jia, and Hongying Zan. 2024 · 2024
Closest in time.
KIEval: A knowledge-grounded interactive evaluation framework for large language models
Zhuohao Yu, Chang Gao, Wenjin Yao, Yidong Wang, Wei Ye, Jindong Wang, Xing Xie, Yue Zhang, and Shikun Zhang. 2024 · 2024
Closest in time.
Dyval: Dynamic evaluation of large language models for reasoning tasks
Kaijie Zhu, Jiaao Chen, Jindong Wang, Neil Zhenqiang Gong, Diyi Yang, and Xing Xie. 2024 · 2024
Closest in time.