Teaching language models to support answers with verified quotes, 2022
Jacob Menick, Maja Trebacz, Vladimir Mikulik, John Aslanides, Francis Song, Martin Chadwick, Mia Glaese, Susannah Young, Lucy Campbell-Gillingham, Geoffrey Irving, and Nat McAleese · 2022
Later among the works it cites.
Webgpt: Browser-assisted question-answering with human feedback, 2022
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schulman · 2022
Later among the works it cites.
Benchclamp: A benchmark for evaluating language models on semantic parsing, 2022
Subhro Roy, Sam Thomson, Tongfei Chen, Richard Shin, Adam Pauls, Jason Eisner, and Benjamin Van Durme · 2022
Later among the works it cites.
On the evaluation metrics for paraphrase generation
Lingfeng Shen, Lemao Liu, Haiyun Jiang, and Shuming Shi · 2022
Later among the works it cites.
The sensitivity of annotator bias to task definitions in argument mining
Terne Sasha Thorn Jakobsen, Maria Barrett, Anders Søgaard, and David Lassen · 2022
Later among the works it cites.
SaFeRDialogues: Taking feedback gracefully after conversational safety failures
Megan Ung, Jing Xu, and Y-Lan Boureau · 2022
Later among the works it cites.
Adversarial glue: A multi-task benchmark for robustness evaluation of language models, 2022
Boxin Wang, Chejian Xu, Shuohang Wang, Zhe Gan, Yu Cheng, Jianfeng Gao, Ahmed Hassan Awadallah, and Bo Li · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models, 2022
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer · 2022
Later among the works it cites.
A multitask, multilingual, multimodal evaluation of chatgpt on reasoning, hallucination, and interactivity
Original
Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, et al · 2023
Closest in time.
Pythia: A suite for analyzing large language models across training and scaling, 2023
Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar van der Wal · 2023
Closest in time.
Sparks of artificial general intelligence: Early experiments with gpt-4
Original
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al · 2023
Closest in time.
A close look into the calibration of pre-trained language models, 2023
Yangyi Chen, Lifan Yuan, Ganqu Cui, Zhiyuan Liu, and Heng Ji · 2023
Closest in time.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al · 2023
Closest in time.
Free dolly: Introducing the world’s first truly open instruction-tuned llm, 2023
Mike Conover, Matt Hayes, Ankit Mathur, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia, and Reynold Xin · 2023
Closest in time.
Gptscore: Evaluate as you desire
Original
Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu · 2023
Closest in time.
Xiezhi: An ever-updating benchmark for holistic domain knowledge evaluation
Original
Zhouhong Gu, Xiaoxuan Zhu, Haoning Ye, Lin Zhang, Jianchen Wang, Sihang Jiang, Zhuozhi Xiong, Zihan Li, Qianyu He, Rui Xu, et al · 2023
Closest in time.
LLM-Eval: Unified multi-dimensional automatic evaluation for open-domain conversations with large language models, 2023
Yen-Ting Lin and Yun-Nung Chen · 2023
Closest in time.
Benchmarking large language model capabilities for conditional generation
Joshua Maynez, Priyanka Agrawal, and Sebastian Gehrmann · 2023
Closest in time.
The RefinedWeb dataset for Falcon LLM: outperforming curated corpora with web data, and web data only
Original
Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay · 2023
Closest in time.
On the challenges of using black-box apis for toxicity evaluation in research, 2023
Luiza Pozzobon, Beyza Ermis, Patrick Lewis, and Sara Hooker · 2023
Closest in time.
Direct preference optimization: Your language model is secretly a reward model, 2023
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn · 2023
Closest in time.
Alpaca: A strong, replicable instruction-following model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto · 2023
Closest in time.
Introducing mpt-7b: A new standard for open-source, commercially usable llms, 2023
MosaicML NLP Team · 2023
Closest in time.
Llama: Open and efficient foundation language models, 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Closest in time.
Koala: An index for quantifying overlaps with pre-training corpora
Original
Thuy-Trang Vu, Xuanli He, Gholamreza Haffari, and Ehsan Shareghi · 2023
Closest in time.
Slic-hf: Sequence likelihood calibration with human feedback, 2023
Yao Zhao, Rishabh Joshi, Tianqi Liu, Misha Khalman, Mohammad Saleh, and Peter J. Liu · 2023
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Original
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2023
Closest in time.
Multilingual machine translation with large language models: Empirical results and analysis, 2023
Wenhao Zhu, Hongyi Liu, Qingxiu Dong, Jingjing Xu, Shujian Huang, Lingpeng Kong, Jiajun Chen, and Lei Li · 2023
Closest in time.
Can large language models transform computational social science?
Original
Caleb Ziems, William Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang, and Diyi Yang · 2023
Closest in time.