Fetching the paper…
Reading the bibliography…
The paper introduces a framework for the evaluation of the encoding of factual scientific knowledge, designed to streamline the manual evaluation process typically conducted by domain experts.
The Curious Case of Neural Text Degeneration, February 2020
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi · 1904
Earlier work this paper cites.
Language Models as Knowledge Bases?, September 2019
Fabio Petroni, Tim Rocktäschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H. Miller, and Sebastian Riedel · 1909
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
On Faithfulness and Factuality in Abstractive Summarization, May 2020b
Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald · 2005
Earlier work this paper cites.
Gleu: Automatic evaluation of sentence-level fluency
Andrew Mutton, Mark Dras, Stephen Wan, and Robert Dale · 2007
Earlier work this paper cites.
Transformers and the Representation of Biomedical Background Knowledge
Oskar Wysocki, Zili Zhou, Paul O’Regan, Deborah Ferreira, Magdalena Wysocka, Dónal Landers, and André Freitas · 2017
Earlier work this paper cites.
Transformers and the Representation of Biomedical Background Knowledge
Oskar Wysocki, Zili Zhou, Paul O’Regan, Deborah Ferreira, Magdalena Wysocka, Dónal Landers, and André Freitas · 2017
Earlier work this paper cites.
What defines the “kingdom” fungi?
Thomas A. Richards, Guy Leonard, and Jeremy G. Wideman · 2017
Earlier work this paper cites.
Sentence-level fluency evaluation: References help, but can be spared!
Katharina Kann, Sascha Rothe, and Katja Filippova · 2018
Earlier work this paper cites.
International Code of Nomenclature for algae, fungi, and plants (Shenzhen Code) adopted by the Nineteenth International Botanical Congress Shenzhen, China, July 2017
Nick J Turland, John Harry Wiersema, Fred R Barrie, Werner Greuter, David L Hawksworth, Patrick Stephen Herendeen, Sandra Knapp, Wolf-Henning Kusber, De-Zhu Li, Karol Marhold, et al · 2018
Earlier work this paper cites.
Toward computer-made artificial antibiotics
Marcelo Der Torossian Torres and Cesar de la Fuente-Nunez · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Language models are few-shot learners, 2020
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Systematic evaluation of research progress on natural language processing in medicine over the past 20 years: Bibliometric study on pubmed
Jing Wang, Huan Deng, Bangtao Liu, Anbin Hu, Jun Liang, Lingye Fan, Xu Zheng, Tong Wang, and Jianbo Lei · 2020
Earlier work this paper cites.
Are Pretrained Language Models Symbolic Reasoners over Knowledge?
Nora Kassner, Benno Krojer, and Hinrich Schütze · 2020
Earlier work this paper cites.
On faithfulness and factuality in abstractive summarization
Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2020
Earlier work this paper cites.
How context affects language models’ factual predictions, 2020
Fabio Petroni, Patrick Lewis, Aleksandra Piktus, Tim Rocktäschel, Yuxiang Wu, Alexander H. Miller, and Sebastian Riedel · 2020
Earlier work this paper cites.
Unambiguous identification of fungi: where do we stand and how accurate and precise is fungal dna barcoding?
Robert Lücking, M Catherine Aime, Barbara Robbertse, Andrew N Miller, Hiran A Ariyawansa, Takayuki Aoki, Gianluigi Cardinali, Pedro W Crous, Irina S Druzhinina, David M Geiser, et al · 2020
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al · 2020
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell · 2021
Earlier work this paper cites.
Artificial intelligence and antibiotic discovery
Liliana David, Anca Monica Brata, Cristina Mogosan, Cristina Pop, Zoltan Czako, Lucian Muresan, Abdulrahman Ismaiel, Dinu Iuliu Dumitrascu, Daniel Corneliu Leucuta, Mihaela Fadygas Stanculete, et al · 2021
Earlier work this paper cites.
Accelerating antibiotic discovery through artificial intelligence
Marcelo CR Melo, Jacqueline RMA Maasch, and Cesar de la Fuente-Nunez · 2021
Earlier work this paper cites.
Calibrate Before Use: Improving Few-Shot Performance of Language Models, June 2021
Tony Z. Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh · 2021
Earlier work this paper cites.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al · 2021
Earlier work this paper cites.
q 2 q^{2} : Evaluating factual consistency in knowledge-grounded dialogues via question generation and question answering
Or Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman, Idan Szpektor, and Omri Abend · 2021
Earlier work this paper cites.
Understanding factuality in abstractive summarization with FRANK: A benchmark for factuality metrics
Artidoro Pagnoni, Vidhisha Balachandran, and Yulia Tsvetkov · 2021
Earlier work this paper cites.
Measuring attribution in natural language generation models
Hannah Rashkin, Vitaly Nikolaev, Matthew Lamm, Lora Aroyo, Michael Collins, Dipanjan Das, Slav Petrov, Gaurav Singh Tomar, Iulia Turc, and David Reitter · 2021
Earlier work this paper cites.
Do prompt-based models really understand the meaning of their prompts?
Albert Webson and Ellie Pavlick · 2021
Earlier work this paper cites.
How to publish a new fungal species, or name, version 3.0
M Catherine Aime, Andrew N Miller, Takayuki Aoki, Konstanze Bensch, Lei Cai, Pedro W Crous, David L Hawksworth, Kevin D Hyde, Paul M Kirk, Robert Lücking, et al · 2021
Cited alongside, same era.
Taxonomy of Risks posed by Language Models
Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, Courtney Biles, Sasha Brown, Zac Kenton, Will Hawkins, Tom Stepleton, Abeba Birhane, Lisa Anne Hendricks, Laura Rimell, William Isaac, Julia Haas, Sean Legassick, Geoffrey Irving, and Iason Gabriel · 2022
Cited alongside, same era.
Do transformers encode a foundational ontology? probing abstract classes in natural language
Mael Jullien, Marco Valentino, and Andre Freitas · 2022
Cited alongside, same era.
Rational discovery of antimicrobial peptides by means of artificial intelligence
Paola Ruiz Puentes, Maria C Henao, Javier Cifuentes, Carolina Muñoz-Camargo, Luis H Reyes, Juan C Cruz, and Pablo Arbeláez · 2022
Cited alongside, same era.
Impact of Pretraining Term Frequencies on Few-Shot Numerical Reasoning
Large language models are state-of-the-art evaluators of translation quality
Tom Kocmi and Christian Federmann · 2023
Closest in time.
Is ChatGPT a good NLG evaluator? a preliminary study
Jiaan Wang, Yunlong Liang, Fandong Meng, Zengkui Sun, Haoxiang Shi, Zhixu Li, Jinan Xu, Jianfeng Qu, and Jie Zhou · 2023
Closest in time.
Style over substance: Evaluation biases for large language models, 2023
Minghao Wu and Alham Fikri Aji · 2023
Closest in time.
Benchmarking cognitive biases in large language models as evaluators, 2023
Ryan Koo, Minhwa Lee, Vipul Raheja, Jong Inn Park, Zae Myung Kim, and Dongyeop Kang · 2023
Closest in time.
LLM-blender: Ensembling large language models with pairwise ranking and generative fusion
Dongfu Jiang, Xiang Ren, and Bill Yuchen Lin · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yasaman Razeghi, Robert L Logan Iv, Matt Gardner, and Sameer Singh · 2022
Cited alongside, same era.
Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets, January 2022
Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra · 2022
Cited alongside, same era.
Memorization Without Overfitting: Analyzing the Training Dynamics of Large Language Models, November 2022
Kushal Tirumala, Aram H. Markosyan, Luke Zettlemoyer, and Armen Aghajanyan · 2022
Cited alongside, same era.
Measuring reliability of large language models through semantic consistency
Harsh Raj, Domenic Rosati, and Subhabrata Majumdar · 2022
Cited alongside, same era.
Pei Ke, Hao Zhou, Yankai Lin, Peng Li, Jie Zhou, Xiaoyan Zhu, and Minlie Huang · 2022
Cited alongside, same era.
Sophie Henning, William Beluch, Alexander Fraser, and Annemarie Friedrich · 2022
Cited alongside, same era.
Fungal names: a comprehensive nomenclatural repository and knowledge base for fungal taxonomy
Fang Wang, Ke Wang, Lei Cai, Mingjun Zhao, Paul M Kirk, Guomei Fan, Qinglan Sun, Bo Li, Shuai Wang, Zhengfei Yu, Dong Han, Juncai Ma, Linhuan Wu, and Yijian Yao · 2022
Cited alongside, same era.
BigScience Language Open-science Open-access Multilingual (BLOOM) Language Model
BigScience · 2022
Cited alongside, same era.
Memories for virtual AI characters
Fabian Landwehr, Erika Varis Doggett, and Romann M. Weber · 2023
Closest in time.
Systematic assessment of factual knowledge in large language models
Linhao Luo, Thuy-Trang Vu, Dinh Phung, and Gholamreza Haffari · 2023
Closest in time.
Many bioinformatics programming tasks can be automated with chatgpt
Stephen R Piccolo, Paul Denny, Andrew Luxton-Reilly, Samuel Payne, and Perry G Ridge · 2023
Closest in time.
Gilchan Park, Byung-Jun Yoon, Xihaier Luo, Vanessa López-Marrero, Patrick Johnstone, Shinjae Yoo, and Francis J Alexander · 2023
Closest in time.
Bioinfo-bench: A simple benchmark framework for llm bioinformatics skills evaluation
Qiyuan Chen and Cheng Deng · 2023
Closest in time.
Large language models encode clinical knowledge
Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al · 2023
Closest in time.
Neeraj Varshney, Wenlin Yao, Hongming Zhang, Jianshu Chen, and Dong Yu · 2023
Closest in time.
Hallucination is the last thing you need
Shawn Curran, Sam Lansley, and Oliver Bethell · 2023
Closest in time.
Large language models are human-level prompt engineers, 2023
Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models, 2023
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom · 2023
Closest in time.
Leveraging large language models for nlg evaluation: A survey
Zhen Li, Xiaohan Xu, Tao Shen, Can Xu, Jia-Chen Gu, and Chongyang Tao · 2024
Closest in time.
Llms instead of human judges? a large scale empirical study across 20 nlp evaluation tasks
Anna Bavaresco, Raffaella Bernardi, Leonardo Bertolazzi, Desmond Elliott, Raquel Fernández, Albert Gatt, Esam Ghaleb, Mario Giulianelli, Michael Hanna, Alexander Koller, André F. T. Martins, Philipp Mondorf, Vera Neplenbroek, Sandro Pezzelle, Barbara Plank, David Schlangen, Alessandro Suglia, Aditya K Surikuchi, Ece Takmaz, and Alberto Testoni · 2024
Closest in time.
Leveraging large language models for predictive chemistry
Kevin Maik Jablonka, Philippe Schwaller, Andres Ortega-Guerrero, and Berend Smit · 2024
Closest in time.
ChatGPT rates natural language explanation quality like humans: But on which scales?
Fan Huang, Haewoon Kwak, Kunwoo Park, and Jisun An · 2024
Closest in time.
Replacing judges with juries: Evaluating llm generations with a panel of diverse models
Pat Verga, Sebastian Hofstatter, Sophia Althammer, Yixuan Su, Aleksandra Piktus, Arkady Arkhangorodsky, Minjie Xu, Naomi White, and Patrick Lewis · 2024
Closest in time.
Are large language model-based evaluators the solution to scaling up multilingual evaluation?
Rishav Hada, Varun Gumma, Adrian Wynter, Harshita Diddee, Mohamed Ahmed, Monojit Choudhury, Kalika Bali, and Sunayana Sitaram · 2024
Closest in time.
The effectiveness of LLMs as annotators: A comparative overview and empirical analysis of direct representation
Maja Pavlovic and Massimo Poesio · 2024
Closest in time.
Evaluating large language models at evaluating instruction following
Zhiyuan Zeng, Jiatong Yu, Tianyu Gao, Yu Meng, Tanya Goyal, and Danqi Chen · 2024
Closest in time.
Pitfalls of conversational LLMs on news debiasing
Ipek Baris Schlicht, Defne Altiok, Maryanne Taouk, and Lucie Flek · 2024
Closest in time.
Trustllm: Trustworthiness in large language models
Lichao Sun, Yue Huang, Haoran Wang, Siyuan Wu, Qihui Zhang, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, Xiner Li, et al · 2024
Closest in time.
An evaluation of large language models in bioinformatics research
Hengchuang Yin, Zhonghui Gu, Fanhao Wang, Yiparemu Abuduhaibaier, Yanqiao Zhu, Xinming Tu, Xian-Sheng Hua, Xiao Luo, and Yizhou Sun · 2024
Closest in time.
Aparna Elangovan, Ling Liu, Lei Xu, Sravan Bodapati, and Dan Roth · 2024
Closest in time.
Understanding the effects of language-specific class imbalance in multilingual fine-tuning
Vincent Jung and Lonneke van der Plas · 2024
Closest in time.
An llm-based knowledge synthesis and scientific reasoning framework for biomedical discovery
Oskar Wysocki, Magdalena Wysocka, Danilo Carvalho, Alex Teodor Bogatu, Danilo Miranda Gusicuma, Maxime Delmas, Harriet Unsworth, and Andre Freitas · 2024
Closest in time.