Fetching the paper…
Reading the bibliography…
We present MatSci-NLP, a natural language benchmark for evaluating the performance of natural language processing (NLP) models on materials science text.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Scibert: A pretrained language model for scientific text
Iz Beltagy, Kyle Lo, and Arman Cohan. 2019 · 1903
Earlier work this paper cites.
Sheshera Mysore, Zach Jensen, Edward Kim, Kevin Huang, Haw-Shiuan Chang, Emma Strubell, Jeffrey Flanigan, Andrew McCallum, and Elsa Olivetti. 2019 · 1905
Earlier work this paper cites.
Pubmedqa: A dataset for biomedical research question answering
Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William W Cohen, and Xinghua Lu. 2019 · 1909
Earlier work this paper cites.
Modeling task relationships in multi-task learning with multi-gate mixture-of-experts
Jiaqi Ma, Zhe Zhao, Xinyang Yi, Jilin Chen, Lichan Hong, and Ed H Chi. 2018 · 1939
Earlier work this paper cites.
Annotating and extracting synthesis process of all-solid-state batteries from scientific literature
Fusataka Kuniyoshi, Kohei Makino, Jun Ozawa, and Makoto Miwa. 2020 · 2002
Earlier work this paper cites.
The sofc-exp corpus and neural approaches to information extraction in the materials science domain
Annemarie Friedrich, Heike Adel, Federico Tomazic, Johannes Hingerl, Renou Benteau, Anika Maruscyk, and Lukas Lange. 2020 · 2006
Earlier work this paper cites.
Biomegatron: Larger biomedical domain language model
Hoo-Chang Shin, Yang Zhang, Evelina Bakhturina, Raul Puri, Mostofa Patwary, Mohammad Shoeybi, and Raghav Mani. 2020 · 2010
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Sc-comics: a superconductivity corpus for materials informatics
Kyosuke Yamaguchi, Ryoji Asahi, and Yutaka Sasaki. 2020 · 2014
Earlier work this paper cites.
Bioasq: A challenge on large-scale biomedical semantic indexing and question answering
Georgios Balikas, Anastasia Krithara, Ioannis Partalas, and George Paliouras. 2015 · 2015
Earlier work this paper cites.
Multi-task sequence to sequence learning
Minh-Thang Luong, Quoc V Le, Ilya Sutskever, Oriol Vinyals, and Lukasz Kaiser. 2015 · 2015
Earlier work this paper cites.
Focal loss for dense object detection
Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár. 2017 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Biomedical named entity recognition with multilingual bert
Kai Hakala and Sampo Pyysalo. 2019 · 2019
Earlier work this paper cites.
Transfer learning in biomedical natural language processing: An evaluation of bert and elmo on ten benchmarking datasets
Yifan Peng, Shankai Yan, and Zhiyong Lu. 2019 · 2019
Cited alongside, same era.
Bert with history answer embedding for conversational question answering
Chen Qu, Liu Yang, Minghui Qiu, W Bruce Croft, Yongfeng Zhang, and Mohit Iyyer. 2019 · 2019
Cited alongside, same era.
Named entity recognition and normalization applied to large-scale information extraction from the materials science literature
Leigh Weston, Vahe Tshitoyan, John Dagdelen, Olga Kononova, Amalie Trewartha, Kristin A Persson, Gerbrand Ceder, and Anubhav Jain. 2019 · 2019
Cited alongside, same era.
Enriching pre-trained language model with entity information for relation classification
Shanchan Wu and Yifan He. 2019 · 2019
Cited alongside, same era.
Inorganic materials synthesis planning with literature-trained neural networks
Edward Kim, Zach Jensen, Alexander van Grootel, Kevin Huang, Matthew Staib, Sheshera Mysore, Haw-Shiuan Chang, Emma Strubell, Andrew McCallum, Stefanie Jegelka, et al. 2020 · 2020
Opportunities and challenges of text mining in materials research
Olga Kononova, Tanjin He, Haoyan Huo, Amalie Trewartha, Elsa A Olivetti, and Gerbrand Ceder. 2021 · 2021
Later among the works it cites.
Text2event: Controllable sequence-to-structure generation for end-to-end event extraction
Yaojie Lu, Hongyu Lin, Jin Xu, Xianpei Han, Jialong Tang, Annan Li, Le Sun, Meng Liao, and Shaoyi Chen. 2021 · 2021
Later among the works it cites.
Scifive: a text-to-text transformer model for biomedical literature
Long N Phan, James T Anibal, Hieu Tran, Shaurya Chanana, Erol Bahadroglu, Alec Peltekian, and Grégoire Altan-Bonnet. 2021 · 2021
Later among the works it cites.
Machine learning in materials science: From explainable predictions to autonomous design
Ghanshyam Pilania. 2021 · 2021
Later among the works it cites.
Looking through glass: Knowledge discovery from materials science literature using natural language processing
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Biobert: a pre-trained biomedical language representation model for biomedical text mining
Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. 2020 · 2020
Cited alongside, same era.
Text mining for processing conditions of solid-state battery electrolytes
Rubayyat Mahbub, Kevin Huang, Zach Jensen, Zachary D Hood, Jennifer LM Rupp, and Elsa A Olivetti. 2020 · 2020
Cited alongside, same era.
Data-driven materials research enabled by natural language processing and information extraction
Elsa A Olivetti, Jacqueline M Cole, Edward Kim, Olga Kononova, Gerbrand Ceder, Thomas Yong-Jin Han, and Anna M Hiszpanski. 2020 · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al. 2020 · 2020
Cited alongside, same era.
A pre-training technique to localize medical bert and enhance biobert
Shoya Wada, Toshihiro Takeda, Shiro Manabe, Shozo Konishi, Jun Kamohara, and Yasushi Matsumura. 2020 · 2020
Cited alongside, same era.
A dataset of information-seeking questions and answers anchored in research papers
Pradeep Dasigi, Kyle Lo, Iz Beltagy, Arman Cohan, Noah A Smith, and Matt Gardner. 2021 · 2021
Cited alongside, same era.
Database, features, and machine learning model to identify thermally driven metal–insulator transition compounds
Alexandru B Georgescu, Peiwen Ren, Aubrey R Toland, Shengtong Zhang, Kyle D Miller, Daniel W Apley, Elsa A Olivetti, Nicholas Wagner, and James M Rondinelli. 2021 · 2021
Cited alongside, same era.
Vineeth Venugopal, Sourav Sahoo, Mohd Zaki, Manish Agarwal, Nitya Nand Gosvami, and NM Anoop Krishnan. 2021 · 2021
Later among the works it cites.
The impact of domain-specific pre-training on named entity recognition tasks in materials science
Nicholas Walker, Amalie Trewartha, Haoyan Huo, Sanghoon Lee, Kevin Cruse, John Dagdelen, Alexander Dunn, Kristin Persson, Gerbrand Ceder, and Anubhav Jain. 2021 · 2021
Later among the works it cites.
Recent advances and applications of deep learning methods in materials science
Kamal Choudhary, Brian DeCost, Chi Chen, Anubhav Jain, Francesca Tavazza, Ryan Cohn, Cheol Woo Park, Alok Choudhary, Ankit Agrawal, Simon JL Billinge, et al. 2022 · 2022
Later among the works it cites.
Matscibert: A materials domain language model for text mining and information extraction
Tanishq Gupta, Mohd Zaki, NM Krishnan, et al. 2022 · 2022
Later among the works it cites.
Scholarbert: Bigger is not always better
Zhi Hong, Aswathy Ajith, Gregory Pauloski, Eamon Duede, Carl Malamud, Roger Magoulas, Kyle Chard, and Ian Foster. 2022 · 2022
Later among the works it cites.
Batterybert: A pretrained language model for battery database enhancement
Shu Huang and Jacqueline M Cole. 2022 · 2022
Later among the works it cites.
Material science relation extraction (matscire)
MatSciRE. 2022 · 2022
Later among the works it cites.
Ai4mat - neurips 2022
Santiago Miret, Marta Skreta, Benjamin Sanchez-Lengelin, Shyue Ping Ong, Zamyla Morgan-Chan, and Alan Aspuru-Guzik · 2022
Later among the works it cites.
Bloom: A 176b-parameter open-access multilingual language model
Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2022 · 2022
Later among the works it cites.
Galactica: A large language model for science
Ross Taylor, Marcin Kardas, Guillem Cucurull, Thomas Scialom, Anthony Hartshorn, Elvis Saravia, Andrew Poulton, Viktor Kerkez, and Robert Stojnic. 2022 · 2022
Later among the works it cites.
Joint extraction of entities, relations, and events via modeling inter-instance and inter-label dependencies
Minh Van Nguyen, Bonan Min, Franck Dernoncourt, and Thien Nguyen. 2022 · 2022
Later among the works it cites.