Fetching the paper…
Reading the bibliography…
We present SciRIFF (Scientific Resource for Instruction-Following and Finetuning), a dataset of 137K instruction-following instances for training and evaluation, covering 54 tasks.
Medmentions: A large biomedical corpus annotated with umls concepts
Sunil Mohan and Donghui Li. 2019 · 1902
Earlier work this paper cites.
Introduction to the bio-entity recognition task at JNLPBA
Nigel Collier, Tomoko Ohta, Yoshimasa Tsuruoka, Yuka Tateisi, and Jin-Dong Kim. 2004 · 2004
Earlier work this paper cites.
The ACL Anthology reference corpus: A reference dataset for bibliographic research in computational linguistics
Steven Bird, Robert Dale, Bonnie Dorr, Bryan Gibson, Mark Joseph, Min-Yen Kan, Dongwon Lee, Brett Powley, Dragomir Radev, and Yee Fan Tan. 2008 · 2008
Earlier work this paper cites.
Linnaeus: A species name identification system for biomedical literature
Martin Gerner, Goran Nenadic, and Casey M. Bergman. 2010 · 2010
Earlier work this paper cites.
The ddi corpus: An annotated corpus with pharmacological substances and drug–drug interactions
María Herrero-Zazo, Isabel Segura-Bedmar, Paloma Martínez, and Thierry Declerck. 2013 · 2013
Earlier work this paper cites.
Ncbi disease corpus: a resource for disease name recognition and concept normalization
Rezarta Islamaj Doğan, Robert Leaman, and Zhiyong Lu. 2014 · 2014
Earlier work this paper cites.
Anatomical entity mention recognition at literature scale
Sampo Pyysalo and Sophia Ananiadou. 2014 · 2014
Earlier work this paper cites.
The chemdner corpus of chemicals and drugs and its annotation principles
Martin Krallinger, Obdulia Rabal, Florian Leitner, Miguel Vazquez, David Salgado, Zhiyong Lu, Robert Leaman, Yanan Lu, Donghong Ji, Daniel M. Lowe, Roger A. Sayle, Riza Theresa Batista-Navarro, Rafal Rak, Torsten Huber, Tim Rocktäschel, Sérgio Matos, David Campos, Buzhou Tang, Hua Xu, Tsendsuren Munkhdalai, Keun Ho Ryu, S. V. Ramanan, Senthil Nathan, Slavko Žitnik, Marko Bajec, Lutz Weber, Matthias Irmer, Saber A. Akhondi, Jan A. Kors, Shuo Xu, Xin An, Utpal Kumar Sikdar, Asif Ekbal, Masaharu Yoshioka, Thaer M. Dieb, Miji Choi, Karin Verspoor, Madian Khabsa, C. Lee Giles, Hongfang Liu, Komandur Elayavilli Ravikumar, Andre Lamurias, Francisco M. Couto, Hong-Jie Dai, Richard Tzong-Han Tsai, Caglar Ata, Tolga Can, Anabel Usié, Rui Alves, Isabel Segura-Bedmar, Paloma Martínez, Julen Oyarzabal, and Alfonso Valencia. 2015 · 2015
Earlier work this paper cites.
An overview of the bioasq large-scale biomedical semantic indexing and question answering competition
George Tsatsaronis, Georgios Balikas, Prodromos Malakasiotis, Ioannis Partalas, Matthias Zschunke, Michael R Alvers, Dirk Weissenborn, Anastasia Krithara, Sergios Petridis, Dimitris Polychronopoulos, et al. 2015 · 2015
Earlier work this paper cites.
Gnormplus: An integrative approach for tagging genes, gene families, and protein domains
Chih-Hsuan Wei, Hung-Yu Kao, and Zhiyong Lu. 2015 · 2015
Earlier work this paper cites.
Biocreative v cdr task corpus: a resource for chemical disease relation extraction
Jiao Li, Yueping Sun, Robin J. Johnson, Daniela Sciaky, Chih-Hsuan Wei, Robert Leaman, Allan Peter Davis, Carolyn J. Mattingly, Thomas C. Wiegers, and Zhiyong Lu. 2016 · 2016
Earlier work this paper cites.
The colorado richly annotated full text (craft) corpus: Multi-model annotation in the biomedical domain
Kevin Bretonnel Cohen, Karin Verspoor, Karën Fort, Christopher Funk, Michael Bada, Martha Palmer, and Lawrence E. Hunter. 2017 · 2017
Earlier work this paper cites.
Overview of the biocreative vi chemical-protein interaction track
Martin Krallinger, Obdulia Rabal, Saber Ahmad Akhondi, Martín Pérez Pérez, Jesus Santamaría, Gael Pérez Rodríguez, Georgios Tsatsaronis, Ander Intxaurrondo, José Antonio Baso López, Umesh K. Nandal, Erin M. van Buel, Ambika Chandrasekhar, Marleen Rodenburg, Astrid Lægreid, Marius A. Doornenbal, Julen Oyarzábal, Anália Lourenço, and Alfonso Valencia. 2017 · 2017
Earlier work this paper cites.
A discourse-aware attention model for abstractive summarization of long documents
Arman Cohan, Franck Dernoncourt, Doo Soon Kim, Trung Bui, Seokhwan Kim, Walter Chang, and Nazli Goharian. 2018 · 2018
Earlier work this paper cites.
Multi-task identification of entities, relations, and coreference for scientific knowledge graph construction
Yi Luan, Luheng He, Mari Ostendorf, and Hannaneh Hajishirzi. 2018 · 2018
Earlier work this paper cites.
A corpus with multi-level annotations of patients, interventions and outcomes to support language processing for medical literature
Benjamin Nye, Junyi Jessy Li, Roma Patel, Yinfei Yang, Iain Marshall, Ani Nenkova, and Byron Wallace. 2018 · 2018
Earlier work this paper cites.
Structural scaffolds for citation intent classification in scientific publications
Arman Cohan, Waleed Ammar, Madeleine van Zuylen, and Field Cady. 2019 · 2019
Earlier work this paper cites.
PubMedQA: A dataset for biomedical research question answering
Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William Cohen, and Xinghua Lu. 2019 · 2019
Earlier work this paper cites.
Toward systematic review automation: a practical guide to using machine learning tools in research synthesis
Iain James Marshall and Byron C. Wallace. 2019 · 2019
Earlier work this paper cites.
The materials science procedural text corpus: Annotating materials synthesis procedures with shallow semantic structures
Sheshera Mysore, Zachary Jensen, Edward Kim, Kevin Huang, Haw-Shiuan Chang, Emma Strubell, Jeffrey Flanigan, Andrew McCallum, and Elsa Olivetti. 2019 · 2019
Earlier work this paper cites.
Transfer learning in biomedical natural language processing: An evaluation of bert and elmo on ten benchmarking datasets
Yifan Peng, Shankai Yan, and Zhiyong Lu. 2019 · 2019
Earlier work this paper cites.
TLDR: Extreme summarization of scientific documents
Isabel Cachola, Kyle Lo, Arman Cohan, and Daniel Weld. 2020 · 2020
Earlier work this paper cites.
Evidence inference 2.0: More data, better models
Jay DeYoung, Eric Lehman, Benjamin Nye, Iain Marshall, and Byron C. Wallace. 2020 · 2020
Earlier work this paper cites.
AxCell: Automatic extraction of results from machine learning papers
Marcin Kardas, Piotr Czapla, Pontus Stenetorp, Sebastian Ruder, Sebastian Riedel, Ross Taylor, and Robert Stojnic. 2020 · 2020
Earlier work this paper cites.
Chia, a large annotated corpus of clinical trial eligibility criteria
Fabrício Kury, Alex Butler, Chi Yuan, Li-heng Fu, Yingcheng Sun, Hao Liu, Ida Sim, Simona Carini, and Chunhua Weng. 2020 · 2020
Earlier work this paper cites.
Multi-XScience: A large-scale dataset for extreme multi-document summarization of scientific articles
Yao Lu, Yue Dong, and Laurent Charlin. 2020 · 2020
Earlier work this paper cites.
COVID-QA: A question answering dataset for COVID-19
Timo Möller, Anthony Reina, Raghavan Jayakumar, and Malte Pietsch. 2020 · 2020
Earlier work this paper cites.
Fact or fiction: Verifying scientific claims
David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi. 2020 · 2020
Earlier work this paper cites.
A dataset of information-seeking questions and answers anchored in research papers
Pradeep Dasigi, Kyle Lo, Iz Beltagy, Arman Cohan, Noah A. Smith, and Matt Gardner. 2021 · 2021
Earlier work this paper cites.
MS^2: Multi-document summarization of medical studies
Jay DeYoung, Iz Beltagy, Madeleine van Zuylen, Bailey Kuehl, and Lucy Lu Wang. 2021 · 2021
Earlier work this paper cites.
Domain-specific language model pretraining for biomedical natural language processing
Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, and Hoifung Poon. 2021 · 2021
Earlier work this paper cites.
Cross-Task Generalization via Natural Language Crowdsourcing Instructions
Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2021 · 2021
Earlier work this paper cites.
COVID-fact: Fact extraction and verification of real-world claims on COVID-19 pandemic
Arkadiy Saakyan, Tuhin Chakrabarty, and Smaranda Muresan. 2021 · 2021
Earlier work this paper cites.
Evidence-based fact-checking of health-related claims
Mourad Sarrouti, Asma Ben Abacha, Yassine Mrabet, and Dina Demner-Fushman. 2021 · 2021
Earlier work this paper cites.
Byron C. Wallace, Sayantan Saha, Frank Soboczenski, and Iain J. Marshall. 2021 · 2021
Cited alongside, same era.
Multi-label classification for biomedical literature: an overview of the biocreative vii litcovid track for covid-19 literature topic annotations
Qingyu Chen, Alexis Allot, Robert Leaman, Rezarta Islamaj, Jingcheng Du, Li Fang, Kai Wang, Shuo Xu, Yuefu Zhang, Parsa Bagherzadeh, et al. 2022 · 2022
Cited alongside, same era.
Overview of the first shared task on multi perspective scientific document summarization (MuP)
Arman Cohan, Guy Feigenblat, Tirthankar Ghosal, and Michal Shmueli-Scheuer. 2022 · 2022
Cited alongside, same era.
Bigbio: A framework for data-centric biomedical natural language processing
Jason Fries, Leon Weber, Natasha Seelam, Gabriel Altay, Debajyoti Datta, Samuele Garda, Sunny Kang, Rosaline Su, Wojciech Kusa, Samuel Cahyawijaya, Fabio Barth, Simon Ott, Matthias Samwald, Stephen Bach, Stella Biderman, Mario Sänger, Bo Wang, Alison Callahan, Daniel León Periñán, Théo Gigant, Patrick Haller, Jenny Chim, Jose Posada, John Giorgi, Karthik Rangasai Sivaraman, Marc Pàmies, Marianna Nezhurina, Robert Martin, Michael Cullan, Moritz Freidank, Nathan Dahlberg, Shubhanshu Mishra, Shamik Bose, Nicholas Broad, Yanis Labrak, Shlok Deshmukh, Sid Kiblawi, Ayush Singh, Minh Chien Vu, Trishala Neeraj, Jonas Golde, Albert Villanova del Moral, and Benjamin Beilharz. 2022 · 2022
HoneyBee: Progressive instruction finetuning of large language models for materials science
Yu Song, Santiago Miret, Huan Zhang, and Bang Liu. 2023 · 2023
Later among the works it cites.
Augustin Toma, Patrick R. Lawler, Jimmy Ba, Rahul G. Krishnan, Barry Rubin, and Bo Wang. 2023 · 2023
Later among the works it cites.
Bioinstruct: Instruction tuning of large language models for biomedical natural language processing
Hieu Tran, Zhichao Yang, Zonghai Yao, and Hong Yu. 2023 · 2023
Later among the works it cites.
DataFinder: Scientific dataset recommendation from natural language descriptions
Vijay Viswanathan, Luyu Gao, Tongshuang Wu, Pengfei Liu, and Graham Neubig. 2023 · 2023
Later among the works it cites.
Academicgpt: Empowering academic research
Shufa Wei, Xiaolong Xu, Xianbiao Qi, Xi Yin, Jun Xia, Jingyi Ren, Peijun Tang, Yuxiang Zhong, Yihao Chen, Xiaoqin Ren, Yuxin Liang, Liankai Huang, Kai Xie, Weikang Gui, Wei Tan, Shuanglong Sun, Yongquan Hu, Qinxian Liu, Nanjin Li, Chihao Dai, Lihua Wang, Xiaohui Liu, Lei Zhang, and Yutao Xie. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Making science simple: Corpora for the lay summarisation of scientific literature
Tomas Goldsack, Zhihao Zhang, Chenghua Lin, and Carolina Scarton. 2022 · 2022
Cited alongside, same era.
MultiCite: Modeling realistic citations requires moving beyond the single-sentence single-label setting
Anne Lauscher, Brandon Ko, Bailey Kuehl, Sophie Johnson, Arman Cohan, David Jurgens, and Kyle Lo. 2022 · 2022
Cited alongside, same era.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Pan Lu, Swaroop Mishra, Tony Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. 2022 · 2022
Cited alongside, same era.
Biored: a rich biomedical relation extraction dataset
Ling Luo, Po-Ting Lai, Chih-Hsuan Wei, Cecilia N Arighi, and Zhiyong Lu. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022 · 2022
Cited alongside, same era.
In-BoXBART: Get instructions into biomedical multi-task learning
Mihir Parmar, Swaroop Mishra, Mirali Purohit, Man Luo, Murad Mohammad, and Chitta Baral. 2022 · 2022
Cited alongside, same era.
M2d2: A massively multi-domain language modeling dataset
Machel Reid, Victor Zhong, Suchin Gururangan, and Luke Zettlemoyer. 2022 · 2022
Cited alongside, same era.
Galactica: A large language model for science
Ross Taylor, Marcin Kardas, Guillem Cucurull, Thomas Scialom, Anthony Hartshorn, Elvis Saravia, Andrew Poulton, Viktor Kerkez, and Robert Stojnic. 2022 · 2022
Cited alongside, same era.
Later among the works it cites.
Darwin series: Domain specific large language models for natural science
Tong Xie, Yuwei Wan, Wei Huang, Zhenyu Yin, Yixuan Liu, Shaozhou Wang, Qingyuan Linghu, Chunyu Kit, Clara Grazian, Wenjie Zhang, Imran Razzak, and Bram Hoex. 2023 · 2023
Later among the works it cites.
Dynosaur: A dynamic growth paradigm for instruction-tuning data curation
Da Yin, Xiao Liu, Fan Yin, Ming Zhong, Hritik Bansal, Jiawei Han, and Kai-Wei Chang. 2023 · 2023
Later among the works it cites.
Lima: Less is more for alignment
Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, L. Yu, Susan Zhang, Gargi Ghosh, M. Lewis, Luke Zettlemoyer, and Omer Levy. 2023 · 2023
Later among the works it cites.
Schema-driven information extraction from heterogeneous tables
Fan Bai, Junmo Kang, Gabriel Stanovsky, Dayne Freitag, Mark Dredze, and Alan Ritter. 2024 · 2024
Closest in time.
Biomedlm: A 2.7b parameter language model trained on biomedical text
Elliot Bolton, Abhinav Venigalla, Michihiro Yasunaga, David Hall, Betty Xiong, Tony Lee, Roxana Daneshjou, Jonathan Frankle, Percy Liang, Michael Carbin, and Christopher D. Manning. 2024 · 2024
Closest in time.
Sciassess: Benchmarking llm proficiency in scientific literature analysis
Hengxing Cai, Xiaochen Cai, Junhan Chang, Sihang Li, Lin Yao, Changxin Wang, Zhifeng Gao, Hongshuai Wang, Yongge Li, Mujie Lin, Shuwen Yang, Jiankun Wang, Mingjun Xu, Jin Huang, Fang Xi, Jiaxi Zhuang, Yuqi Yin, Yaqi Li, Changhong Chen, Zheng Cheng, Zifeng Zhao, Linfeng Zhang, and Guolin Ke. 2024 · 2024
Closest in time.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinson, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Zhao, Yanping Huang, Andrew Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. 2024 · 2024
Closest in time.
How abilities in large language models are affected by supervised fine-tuning data composition
Guanting Dong, Hongyi Yuan, Keming Lu, Chengpeng Li, Mingfeng Xue, Dayiheng Liu, Wei Wang, Zheng Yuan, Chang Zhou, and Jingren Zhou. 2024 · 2024
Closest in time.
The llama 3 herd of models
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, et al. 2024 · 2024
Closest in time.
Mol-instructions: A large-scale biomolecular instruction dataset for large language models
Yin Fang, Xiaozhuan Liang, Ningyu Zhang, Kangwei Liu, Rui Huang, Zhuo Chen, Xiaohui Fan, and Huajun Chen. 2024 · 2024
Closest in time.
Sciknoweval: Evaluating multi-level scientific knowledge of large language models
Kehua Feng, Keyan Ding, Weijie Wang, Xiang Zhuang, Zeyuan Wang, Ming Qin, Yu Zhao, Jianhua Yao, Qiang Zhang, and Huajun Chen. 2024 · 2024
Closest in time.
Prometheus: Inducing fine-grained evaluation capability in language models
Seungone Kim, Jamin Shin, Yejin Cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, and Minjoon Seo. 2024 · 2024
Closest in time.
Biomistral: A collection of open-source pretrained large language models for medical domains
Yanis Labrak, Adrien Bazoge, Emmanuel Morin, P. Gourraud, Mickael Rouvier, and Richard Dufour. 2024 · 2024
Closest in time.
Scilitllm: How to adapt llms for scientific literature understanding
Sihang Li, Jian Huang, Jiaxi Zhuang, Yaorui Shi, Xiaochen Cai, Mingjun Xu, Xiang Wang, Linfeng Zhang, Guolin Ke, and Hengxing Cai. 2024 · 2024
Closest in time.
MUFFIN: Curating multi-faceted instructions for improving instruction following
Renze Lou, Kai Zhang, Jian Xie, Yuxuan Sun, Janice Ahn, Hanzi Xu, Yu su, and Wenpeng Yin. 2024 · 2024
Closest in time.
Learning to generate instruction tuning datasets for zero-shot task adaptation
Nihal V. Nayak, Yiyang Nan, Avi Trost, and Stephen H. Bach. 2024 · 2024
Closest in time.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Mingchuan Zhang, Y.K. Li, Y. Wu, and Daya Guo. 2024 · 2024
Closest in time.
Mathscale: Scaling instruction tuning for mathematical reasoning
Zhengyang Tang, Xingxing Zhang, Benyou Wang, and Furu Wei. 2024 · 2024
Closest in time.
Openmathinstruct-1: A 1.8 million math instruction tuning dataset
Shubham Toshniwal, Ivan Moshkov, Sean Narenthiran, Daria Gitman, Fei Jia, and Igor Gitman. 2024 · 2024
Closest in time.
Mind your format: Towards consistent evaluation of in-context learning improvements
Anton Voronov, Lena Wolf, and Max Ryabinin. 2024 · 2024
Closest in time.
PMC-LLaMA: toward building open-source language models for medicine
Chaoyi Wu, Weixiong Lin, Xiaoman Zhang, Ya Zhang, Weidi Xie, and Yanfeng Wang. 2024 · 2024
Closest in time.
WizardLM: Empowering large pre-trained language models to follow complex instructions
Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, Qingwei Lin, and Daxin Jiang. 2024 · 2024
Closest in time.
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, et al. 2024 · 2024
Closest in time.
Llasmol: Advancing large language models for chemistry with a large-scale, comprehensive, high-quality instruction tuning dataset
Botao Yu, Frazier N. Baker, Ziqi Chen, Xia Ning, and Huan Sun. 2024 · 2024
Closest in time.
Team Gemini-2.5. 2025 · 2025
Closest in time.
Kimi k2: Open agentic intelligence
Team Kimi. 2025 · 2025
Closest in time.
Demystifying scientific problem-solving in llms by probing knowledge and reasoning
Alan Li, Yixin Liu, Arpan Sarkar, Doug Downey, and Arman Cohan. 2025 · 2025
Closest in time.
Sciarena: An open evaluation platform for foundation models in scientific literature tasks
Yilun Zhao, Kaiyan Zhang, Tiansheng Hu, Sihong Wu, Ronan Le Bras, Taira Anderson, Jonathan Bragg, Joseph Chee Chang, Jesse Dodge, Matt Latzke, Yixin Liu, Charles McGrady, Xiangru Tang, Zihang Wang, Chen Zhao, Hannaneh Hajishirzi, Doug Downey, and Arman Cohan. 2025 · 2025
Closest in time.