Fetching the paper…
Reading the bibliography…
Evaluation of Large Language Models (LLMs) is challenging because instruction-following necessitates alignment with human values and the required set of skills varies depending on the instruction.
Criteria for alignment of expectations and assessments in mathematics and science education. research monograph no. 6
Norman Lott Webb · 1997
Earlier work this paper cites.
Alignment of science and mathematics standards and assessments in four states. research monograph no. 18
Norman Lott Webb · 1999
Earlier work this paper cites.
Compositional semantic parsing on semi-structured tables
Panupong Pasupat and Percy Liang · 2015
Earlier work this paper cites.
An analysis of prerequisite skills for reading comprehension
Saku Sugawara and Akiko Aizawa · 2016
Earlier work this paper cites.
Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension
Aniruddha Kembhavi, Minjoon Seo, Dustin Schwenk, Jonghyun Choi, Ali Farhadi, and Hannaneh Hajishirzi · 2017
Earlier work this paper cites.
The e2e dataset: New challenges for end-to-end generation, 2017
Jekaterina Novikova, Ondřej Dušek, and Verena Rieser · 2017
Earlier work this paper cites.
Evaluating quality of chatbots and intelligent conversational agents, 2017
Nicole M. Radziwill and Morgan C. Benton · 2017
Earlier work this paper cites.
Evaluation metrics for machine reading comprehension: Prerequisite skills and readability
Saku Sugawara, Yusuke Kido, Hikaru Yokono, and Akiko Aizawa · 2017
Earlier work this paper cites.
Hierarchical neural story generation, 2018
Angela Fan, Mike Lewis, and Yann Dauphin · 2018
Earlier work this paper cites.
Mapping language to code in programmatic context, 2018
Srinivasan Iyer, Ioannis Konstas, Alvin Cheung, and Luke Zettlemoyer · 2018
Earlier work this paper cites.
Gender bias in coreference resolution, 2018
Rachel Rudinger, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme · 2018
Earlier work this paper cites.
FEVER: a large-scale dataset for fact extraction and VERification
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal · 2018
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering, 2018
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning · 2018
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language, 2019
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi · 2019
Earlier work this paper cites.
Can you unpack that? learning to rewrite questions-in-context
Ahmed Elgohary, Denis Peskov, and Jordan Boyd-Graber · 2019
Earlier work this paper cites.
Counterfactual story reasoning and generation, 2019
Lianhui Qin, Antoine Bosselut, Ari Holtzman, Chandra Bhagavatula, Elizabeth Clark, and Yejin Choi · 2019
Earlier work this paper cites.
Neural network acceptability judgments, 2019
Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman · 2019
Earlier work this paper cites.
Abductive commonsense reasoning, 2020
Chandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi, Ari Holtzman, Hannah Rashkin, Doug Downey, Scott Wen tau Yih, and Yejin Choi · 2020
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models, 2020
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith · 2020
Earlier work this paper cites.
Content planning for neural story generation with aristotelian rescoring
Seraphina Goldfarb-Tarrant, Tuhin Chakrabarty, Ralph Weischedel, and Nanyun Peng · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2020
Earlier work this paper cites.
Commongen: A constrained text generation challenge for generative commonsense reasoning, 2020
Bill Yuchen Lin, Wangchunshu Zhou, Ming Shen, Pei Zhou, Chandra Bhagavatula, Yejin Choi, and Xiang Ren · 2020
Earlier work this paper cites.
Thinking like a skeptic: Defeasible inference in natural language
Rachel Rudinger, Vered Shwartz, Jena D. Hwang, Chandra Bhagavatula, Maxwell Forbes, Ronan Le Bras, Noah A. Smith, and Yejin Choi · 2020
Earlier work this paper cites.
A framework for evaluation of machine reading comprehension gold standards, 2020
Viktor Schlegel, Marco Valentino, André Freitas, Goran Nenadic, and Riza Batista-Navarro · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment, 2021
Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Jackson Kernion, Kamal Ndousse, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, and Jared Kaplan · 2021
Earlier work this paper cites.
Program synthesis with large language models, 2021
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, and Charles Sutton · 2021
Earlier work this paper cites.
Evaluating large language models trained on code, 2021
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Josh Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba · 2021
Earlier work this paper cites.
Decontextualization: Making sentences stand-alone, 2021
Eunsol Choi, Jennimaria Palomaki, Matthew Lamm, Tom Kwiatkowski, Dipanjan Das, and Michael Collins · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Earlier work this paper cites.
A framework for few-shot language model evaluation, September 2021
Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, Jason Phang, Laria Reynolds, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou · 2021
Earlier work this paper cites.
Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies, 2021
Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant · 2021
Earlier work this paper cites.
The disagreement deconvolution: Bringing machine learning performance metrics in line with reality
Mitchell L Gordon, Kaitlyn Zhou, Kayur Patel, Tatsunori Hashimoto, and Michael S Bernstein · 2021
Earlier work this paper cites.
Cosqa: 20,000+ web queries for code search and question answering, 2021
Junjie Huang, Duyu Tang, Linjun Shou, Ming Gong, Ke Xu, Daxin Jiang, Ming Zhou, and Nan Duan · 2021
Earlier work this paper cites.
krippendorffsalpha: An r package for measuring agreement using krippendorff’s alpha coefficient
John Hughes · 2021
Earlier work this paper cites.
Contractnli: A dataset for document-level natural language inference for contracts, 2021
Yuta Koreeda and Christopher D. Manning · 2021
Earlier work this paper cites.
Hurdles to progress in long-form question answering
Kalpesh Krishna, Aurko Roy, and Mohit Iyyer · 2021
Earlier work this paper cites.
Fetaqa: Free-form table question answering, 2021
Linyong Nan, Chiachun Hsieh, Ziming Mao, Xi Victoria Lin, Neha Verma, Rui Zhang, Wojciech Kryściński, Nick Schoelkopf, Riley Kong, Xiangru Tang, Murori Mutuma, Ben Rosand, Isabel Trindade, Renusree Bandaru, Jacob Cunningham, Caiming Xiong, and Dragomir Radev · 2021
Earlier work this paper cites.
Timedial: Temporal commonsense reasoning in dialog, 2021
Lianhui Qin, Aditya Gupta, Shyam Upadhyay, Luheng He, Yejin Choi, and Manaal Faruqui · 2021
Earlier work this paper cites.
QA dataset explosion: A taxonomy of NLP resources for question answering and reading comprehension
Anna Rogers, Matt Gardner, and Isabelle Augenstein · 2021
Earlier work this paper cites.
proscript: Partially ordered scripts generation via pre-trained language models, 2021
Keisuke Sakaguchi, Chandra Bhagavatula, Ronan Le Bras, Niket Tandon, Peter Clark, and Yejin Choi · 2021
Earlier work this paper cites.
Get your vitamin c! robust fact verification with contrastive evidence, 2021
Tal Schuster, Adam Fisch, and Regina Barzilay · 2021
Cited alongside, same era.
self_awareness: a benchmark task to measure self-awareness of language models. In: The Beyond the Imitation Game Benchmark (BIG-bench). GitHub repository: https://github.com/google/BIG-bench , 2021
Roman Sitelew, Jascha Sohl-Dickstein, and Josh Rule · 2021
Cited alongside, same era.
Conditionalqa: A complex reading comprehension dataset with conditional answers, 2021
Haitian Sun, William W. Cohen, and Ruslan Salakhutdinov · 2021
Cited alongside, same era.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le · 2021
Cited alongside, same era.
Topiocqa: Open-domain conversational question answering with topic switching, 2022
Vaibhav Adlakha, Shehzaad Dhuliawala, Kaheer Suleman, Harm de Vries, and Siva Reddy · 2022
Code alpaca: An instruction-following llama model for code generation
Sahil Chaudhary · 2023
Closest in time.
Theoremqa: A theorem-driven question answering dataset
Wenhu Chen, Ming Yin, Max Ku, Elaine Wan, Xueguang Ma, Jianyu Xu, Tony Xia, Xinyi Wang, and Pan Lu · 2023
Closest in time.
Instructeval: Towards holistic evaluation of instruction-tuned large language models
Yew Ken Chia, Pengfei Hong, Lidong Bing, and Soujanya Poria · 2023
Closest in time.
Can large language models be an alternative to human evaluations?, 2023
Cheng-Han Chiang and Hung yi Lee · 2023
Closest in time.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Do as i can, not as i say: Grounding language in robotic affordances, 2022
Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Daniel Ho, Jasmine Hsu, Julian Ibarz, Brian Ichter, Alex Irpan, Eric Jang, Rosario Jauregui Ruano, Kyle Jeffrey, Sally Jesmonth, Nikhil J Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, Kuang-Huei Lee, Sergey Levine, Yao Lu, Linda Luu, Carolina Parada, Peter Pastor, Jornell Quiambao, Kanishka Rao, Jarek Rettinghouse, Diego Reyes, Pierre Sermanet, Nicolas Sievers, Clayton Tan, Alexander Toshev, Vincent Vanhoucke, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, Mengyuan Yan, and Andy Zeng · 2022
Cited alongside, same era.
Measuring progress on scalable oversight for large language models, 2022
Samuel R. Bowman, Jeeyoon Hyun, Ethan Perez, Edwin Chen, Craig Pettit, Scott Heiner, Kamilė Lukošiūtė, Amanda Askell, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Christopher Olah, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Jackson Kernion, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, Liane Lovitt, Nelson Elhage, Nicholas Schiefer, Nicholas Joseph, Noemí Mercado, Nova DasSarma, Robin Larson, Sam McCandlish, Sandipan Kundu, Scott Johnston, Shauna Kravec, Sheer El Showk, Stanislav Fort, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Ben Mann, and Jared Kaplan · 2022
Cited alongside, same era.
Finqa: A dataset of numerical reasoning over financial data, 2022
Zhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan Routledge, and William Yang Wang · 2022
Cited alongside, same era.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al · 2022
Cited alongside, same era.
e-care: a new dataset for exploring explainable causal reasoning, 2022
Li Du, Xiao Ding, Kai Xiong, Ting Liu, and Bing Qin · 2022
Cited alongside, same era.
Understanding dataset difficulty with 𝒱 \mathcal{V} -usable information
Kawin Ethayarajh, Yejin Choi, and Swabha Swayamdipta · 2022
Cited alongside, same era.
Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text, 2022
Sebastian Gehrmann, Elizabeth Clark, and Thibault Sellam · 2022
Cited alongside, same era.
Barda: A belief and reasoning datasetthat separates factual accuracy and reasoning ability, 2023
Peter Clark, Bhavana Dalvi Mishra, and Oyvind Tafjor · 2023
Closest in time.
Diffqg: Generating questions to summarize factual changes, 2023
Jeremy R. Cole, Palak Jain, Julian Martin Eisenschlos, Michael J. Q. Zhang, Eunsol Choi, and Bhuwan Dhingra · 2023
Closest in time.
Qlora: Efficient finetuning of quantized llms
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer · 2023
Closest in time.
Alpacafarm: A simulation framework for methods that learn from human feedback
Yann Dubois, Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto · 2023
Closest in time.
Gptscore: Evaluate as you desire, 2023
Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu · 2023
Closest in time.
Llama-adapter v2: Parameter-efficient visual instruction model
Peng Gao, Jiaming Han, Renrui Zhang, Ziyi Lin, Shijie Geng, Aojun Zhou, Wei Zhang, Pan Lu, Conghui He, Xiangyu Yue, et al · 2023
Closest in time.
Koala: A dialogue model for academic research
Xinyang Geng, Arnav Gudibande, Hao Liu, Eric Wallace, Pieter Abbeel, Sergey Levine, and Dawn Song · 2023
Closest in time.
The false promise of imitating proprietary llms
Arnav Gudibande, Eric Wallace, Charlie Snell, Xinyang Geng, Hao Liu, Pieter Abbeel, Sergey Levine, and Dawn Song · 2023
Closest in time.
Aligning ai with shared human values, 2023
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt · 2023
Closest in time.
Bring your own data! self-supervised evaluation for large language models
Neel Jain, Khalid Saifullah, Yuxin Wen, John Kirchenbauer, Manli Shu, Aniruddha Saha, Micah Goldblum, Jonas Geiping, and Tom Goldstein · 2023
Closest in time.
Exploring the benefits of training expert language models over instruction tuning
Joel Jang, Seungone Kim, Seonghyeon Ye, Doyoung Kim, Lajanugen Logeswaran, Moontae Lee, Kyungjae Lee, and Minjoon Seo · 2023
Closest in time.
Llm-blender: Ensembling large language models with pairwise ranking and generative fusion
Dongfu Jiang, Xiang Ren, and Bill Yuchen Lin · 2023
Closest in time.
Pretraining language models with human preferences
Tomasz Korbak, Kejian Shi, Angelica Chen, Rasika Bhalerao, Christopher L Buckley, Jason Phang, Samuel R Bowman, and Ethan Perez · 2023
Closest in time.
LongEval: Guidelines for human evaluation of faithfulness in long-form summarization
Kalpesh Krishna, Erin Bransom, Bailey Kuehl, Mohit Iyyer, Pradeep Dasigi, Arman Cohan, and Kyle Lo · 2023
Closest in time.
Openassistant conversations – democratizing large language model alignment, 2023
Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi-Rui Tam, Keith Stevens, Abdullah Barhoum, Nguyen Minh Duc, Oliver Stanley, Richárd Nagyfi, Shahul ES, Sameer Suri, David Glushkov, Arnav Dantuluri, Andrew Maguire, Christoph Schuhmann, Huu Nguyen, and Alexander Mattick · 2023
Closest in time.
A systematic study and comprehensive evaluation of chatgpt on benchmark datasets
Md Tahmid Rahman Laskar, M Saiful Bari, Mizanur Rahman, Md Amran Hossen Bhuiyan, Shafiq Joty, and Jimmy Xiangji Huang · 2023
Closest in time.
Can large language models infer and disagree like humans?, 2023
Noah Lee, Na Min An, and James Thorne · 2023
Closest in time.
Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe · 2023
Closest in time.
G-eval: Nlg evaluation using gpt-4 with better human alignment, 2023
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu · 2023
Closest in time.
Wizardcoder: Empowering code large language models with evol-instruct
Ziyang Luo, Can Xu, Pu Zhao, Qingfeng Sun, Xiubo Geng, Wenxiang Hu, Chongyang Tao, Jing Ma, Qingwei Lin, and Daxin Jiang · 2023
Closest in time.
When not to trust language models: Investigating effectiveness of parametric and non-parametric memories, 2023
Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, and Hannaneh Hajishirzi · 2023
Closest in time.
Factscore: Fine-grained atomic evaluation of factual precision in long form text generation
Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Concise answers to complex questions: Summarization of long-form answers, 2023
Abhilash Potluri, Fangyuan Xu, and Eunsol Choi · 2023
Closest in time.
Lamp: When large language models meet personalization, 2023
Alireza Salemi, Sheshera Mysore, Michael Bendersky, and Hamed Zamani · 2023
Closest in time.
Scone: Benchmarking negation reasoning in language models with fine-tuning and in-context learning, 2023
Jingyuan Selena She, Christopher Potts, Samuel R. Bowman, and Atticus Geiger · 2023
Closest in time.
Model evaluation for extreme risks
Toby Shevlane, Sebastian Farquhar, Ben Garfinkel, Mary Phuong, Jess Whittlestone, Jade Leung, Daniel Kokotajlo, Nahema Marchal, Markus Anderljung, Noam Kolt, et al · 2023
Closest in time.
Asqa: Factoid questions meet long-form answers, 2023
Ivan Stelmakh, Yi Luan, Bhuwan Dhingra, and Ming-Wei Chang · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Closest in time.
Doremi: Optimizing data mixtures speeds up language model pretraining
Sang Michael Xie, Hieu Pham, Xuanyi Dong, Nan Du, Hanxiao Liu, Yifeng Lu, Percy Liang, Quoc V Le, Tengyu Ma, and Adams Wei Yu · 2023
Closest in time.
Wizardlm: Empowering large language models to follow complex instructions, 2023
Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang · 2023
Closest in time.
Llama-adapter: Efficient fine-tuning of language models with zero-init attention
Renrui Zhang, Jiaming Han, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, Peng Gao, and Yu Qiao · 2023
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2023
Closest in time.
Agieval: A human-centric benchmark for evaluating foundation models
Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang, Shuai Lu, Yanlin Wang, Amin Saied, Weizhu Chen, and Nan Duan · 2023
Closest in time.