Fetching the paper…
Reading the bibliography…
The vast majority of materials science knowledge exists in unstructured natural language, yet structured data is crucial for innovative and systematic materials design.
“Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context”
Zihang Dai et al · 1901
Earlier work this paper cites.
“SciBERT: A Pretrained Language Model for Scientific Text”
Iz Beltagy, Kyle Lo and Arman Cohan · 1903
Earlier work this paper cites.
“ChemDataExtractor: A Toolkit for Automated Extraction of Chemical Information from the Scientific Literature”
Matthew. Swain and Jacqueline. Cole · 1904
Earlier work this paper cites.
“MASS: Masked Sequence to Sequence Pre-training for Language Generation”
Kaitao Song et al · 1905
Earlier work this paper cites.
“The Linear Sum Assignment Problem”
Rainer. Burkard and Ulrich Derigs · 1980
Earlier work this paper cites.
“OpticalBERT and OpticalTable-SQA: Text- and Table-Based Language Models for the Optical-Materials Domain”
Jiuyang Zhao, Shu Huang and Jacqueline. Cole · 1981
Earlier work this paper cites.
“Materials Selection in Mechanical Design”
Michael Ashby · 1999
Earlier work this paper cites.
“Scaling Laws for Neural Language Models”
Jared Kaplan et al · 2001
Earlier work this paper cites.
“Longformer: The Long-Document Transformer”
Iz Beltagy, Matthew. Peters and Arman Cohan · 2004
Earlier work this paper cites.
“Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks”
Patrick Lewis et al · 2005
Earlier work this paper cites.
“Language Models are Few-Shot Learners”
Tom. Brown et al · 2005
Earlier work this paper cites.
“Scaling Properties of Adsorption Energies for Hydrogen-Containing Molecules on Transition-Metal Surfaces”
Frank Abild-Pedersen et al · 2007
Earlier work this paper cites.
“An overview of the Tesseract OCR engine”
Ray Smith · 2007
Earlier work this paper cites.
“Internet resources integrating many small-molecule databases1”
M. Sitzmann, I.V. Filippov and M.C. Nicklaus · 2008
Earlier work this paper cites.
“OSCAR4: a flexible architecture for chemical text-mining”
David Jessop et al · 2011
Earlier work this paper cites.
“ChemicalTagger: A tool for semantic text-mining in chemistry”
Lezan Hawizy, David Jessop, Nico Adams and Peter Murray-Rust · 2011
Earlier work this paper cites.
“Data-Driven High-Throughput Prediction of the 3-D Structure of Small Molecules: Review and Progress. A Response to the Letter by the Cambridge Crystallographic Data Centre”
Pierre Baldi · 2011
Earlier work this paper cites.
“ChemSpot: a hybrid system for chemical named entity recognition”
Tim Rocktäschel, Michael Weidlich and Ulf Leser · 2012
Earlier work this paper cites.
“Authors Guild, Inc. v. HathiTrust” United States District Court for the Southern District of New York, 902 F. Supp. 2d 445, 2012
2012
Earlier work this paper cites.
“matextract”, https://github.com/lamalab-org/matextract-book , 2014
Mara Schilling-Wilhelmi et al · 2014
Earlier work this paper cites.
“Pint: a Python Units Library”, https://github.com/hgrecco/pint , 2014
Hernan. Grecco · 2014
Earlier work this paper cites.
“LeadMine: a grammar and dictionary driven approach to entity recognition”
Daniel Lowe and Roger Sayle · 2015
Earlier work this paper cites.
“CrossRef text and data mining services”
Rachael Lammey · 2015
Earlier work this paper cites.
“PubChem Substance and Compound databases”
Sunghwan Kim et al · 2015
Earlier work this paper cites.
“Machine-learning-assisted materials discovery using failed experiments”
Paul Raccuglia et al · 2016
Earlier work this paper cites.
“Data management in the long tail: Science, software and service”
Christine Borgman et al · 2016
Earlier work this paper cites.
“Responsible Content Mining”
J. Molloy, M. Haeussler, P. Murray-Rust and C. Oppenheim · 2016
Earlier work this paper cites.
“Machine learning in materials informatics: recent applications and prospects”
Rampi Ramprasad et al · 2017
Earlier work this paper cites.
“Information Retrieval and Text Mining Technologies for Chemistry”
Martin Krallinger et al · 2017
Earlier work this paper cites.
“Machine learning for molecular and materials science”
Keith. Butler et al · 2018
Earlier work this paper cites.
“Inverse molecular design using machine learning: Generative models for matter engineering”
Benjamin Sanchez-Lengeling and Alán Aspuru-Guzik · 2018
Earlier work this paper cites.
“Improving language understanding by generative pre-training”
Alec Radford, Karthik Narasimhan, Tim Salimans and Ilya Sutskever · 2018
Earlier work this paper cites.
“doccano: Text Annotation Tool for Human” Software available from https://github.com/doccano/doccano, 2018
Hiroki Nakayama et al · 2018
Earlier work this paper cites.
“Know What You Don’t Know: Unanswerable Questions for SQuAD”
Pranav Rajpurkar, Robin Jia and Percy Liang · 2018
Earlier work this paper cites.
“unyt: Handle, manipulate, and convert data with units in Python”
Nathan. Goldbaum et al · 2018
Earlier work this paper cites.
“Text-mined dataset of inorganic materials synthesis recipes”
Olga Kononova et al · 2019
Earlier work this paper cites.
“Language models are unsupervised multitask learners”
Alec Radford et al · 2019
Earlier work this paper cites.
“The Semantic Scholar Open Research Corpus (S2ORC)”, https://allenai.org/data/s2orc , 2019
Allen Institute for AI · 2019
Earlier work this paper cites.
“ImageDataExtractor: A Tool To Extract and Quantify Data from Microscopy Images”
Karim. Mukaddem, Edward. Beard, Batuhan Yildirim and Jacqueline. Cole · 2019
Earlier work this paper cites.
“A General-Purpose Algorithm for Constrained Sequential Inference”
Daniel Deutsch, Shyam Upadhyay and Dan Roth · 2019
Earlier work this paper cites.
“Anthropogenic biases in chemical reaction data hinder exploratory inorganic synthesis”
Xiwen Jia et al · 2019
Earlier work this paper cites.
“Big-data science in porous materials: materials genomics and machine learning”
Kevin Jablonka, Daniele Ongari, Seyed Moosavi and Berend Smit · 2020
Earlier work this paper cites.
“A universal system for digitization and automatic execution of the chemical synthesis literature”
S.. Mehr et al · 2020
Earlier work this paper cites.
“Jupyter Book”, https://zenodo.org/record/4539666 , 2020
Executable Community · 2020
Earlier work this paper cites.
“Elsevier OA CC-BY Corpus”, https://researchcollaborations.elsevier.com/en/datasets/elsevier-oa-cc-by-corpus , 2020
Elsevier · 2020
Earlier work this paper cites.
“A review of optical chemical structure recognition tools”
Kohulan Rajan, Henning Brinkhaus, Achim Zielesny and Christoph Steinbeck · 2020
Earlier work this paper cites.
“Opportunities and challenges of text mining in materials research”
Olga Kononova et al · 2021
Earlier work this paper cites.
“ChemDataExtractor 2.0: Autopopulated Ontologies for Materials Science”
Juraj Mavračić et al · 2021
Earlier work this paper cites.
“AI trends that I unironically love”, https://cs.stanford.edu/people/chrismre/papers/SIGMOD-Chris-Re-DataCentric-Foundation-Models-KeyNote.pdf" , 2021
Chris Ré · 2021
Earlier work this paper cites.
“Open Reaction Database”, https://open-reaction-database.org , 2021
Open Reaction Database Project Authors · 2021
Earlier work this paper cites.
“The Open Reaction Database”
Steven. Kearnes et al · 2021
Earlier work this paper cites.
“What Makes Good In-Context Examples for GPT- 3 3 ?”
Jiachang Liu et al · 2021
Earlier work this paper cites.
“LoRA: Low-Rank Adaptation of Large Language Models”
Edward. Hu et al · 2021
Earlier work this paper cites.
“Table Transformer”, https://github.com/microsoft/table-transformer, 2021
Brandon Smock and Rohith Pesala · 2021
Earlier work this paper cites.
“Extracting processing and testing parameters from materials science literature for improved property prediction of glasses”
Mohd Zaki, Jayadeva and N.M. Krishnan · 2021
Earlier work this paper cites.
“Democratising deep learning for microscopy with ZeroCostDL4Mic”
Lucas von Chamier et al · 2021
Earlier work this paper cites.
“Recent advances and applications of deep learning methods in materials science”
Kamal Choudhary et al · 2022
Earlier work this paper cites.
“BatteryDataExtractor: battery-aware text-mining software embedded with BERT models”
Shu Huang and Jacqueline. Cole · 2022
Earlier work this paper cites.
“Training language models to follow instructions with human feedback”
Long Ouyang et al · 2022
Earlier work this paper cites.
“Legal reform to enhance global text and data mining research”
Sean. Fiil-Flynn et al · 2022
Earlier work this paper cites.
“Argilla” Software available from https://github.com/argilla-io/argilla, 2022
Argilla Team · 2022
Earlier work this paper cites.
“PDFDataExtractor: A Tool for Reading Scientific Text and Interpreting Metadata from the Typeset Literature in the Portable Document Format”
Miao Zhu and Jacqueline. Cole · 2022
Earlier work this paper cites.
“Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity”
Yao Lu et al · 2022
Earlier work this paper cites.
“Large Language Models are Few-Shot Clinical Information Extractors”
Monica Agrawal et al · 2022
Earlier work this paper cites.
“LangChain”, https://github.com/langchain-ai/langchain , 2022
Harrison Chase · 2022
Earlier work this paper cites.
“LlamaIndex”, https://github.com/jerryjliu/llama_index , 2022
Jerry Liu · 2022
Earlier work this paper cites.
“Unleashing the Power of Knowledge Extraction from Scientific Literature in Catalysis”
Yue Zhang et al · 2022
Earlier work this paper cites.
“MatSciBERT: A materials domain language model for text mining and information extraction”
Tanishq Gupta, Mohd Zaki, N.. Krishnan and Mausam · 2022
Earlier work this paper cites.
“Generalized Visual Language Models”
Lilian Weng · 2022
Earlier work this paper cites.
“Microstructure segmentation with deep learning encoders pre-trained on a large microscopy dataset”
Joshua Stuckner, Bryan Harder and Timothy. Smith · 2022
Earlier work this paper cites.
“Language Models as Agent Models”
Jacob Andreas · 2022
Earlier work this paper cites.
“Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents”
Wenlong Huang, Pieter Abbeel, Deepak Pathak and Igor Mordatch · 2022
Earlier work this paper cites.
“Data-Driven Matching of Experimental Crystal Structures and Gas Adsorption Isotherms of Metal–Organic Frameworks”
Daniele Ongari et al · 2022
Earlier work this paper cites.
“A general-purpose material property data extraction pipeline from large polymer corpora using natural language processing”
Pranav Shetty et al · 2023
Cited alongside, same era.
“Attention Is All You Need”
Ashish Vaswani et al · 2023
Cited alongside, same era.
“Generative Pre-trained Transformer: A Comprehensive Review on Enabling Technologies, Potential Applications, Emerging Challenges, and Future Directions”
Gokul Yenduri et al · 2023
Cited alongside, same era.
“14 examples of how LLMs can transform materials science and chemistry: a reflection on a large language model hackathon”
Kevin Jablonka et al · 2023
Cited alongside, same era.
“Language models and protocol standardization guidelines for accelerating synthesis planning in heterogeneous catalysis”
Manu Suvarna et al · 2023
Cited alongside, same era.
“Flexible, model-agnostic method for materials data extraction from text using general purpose language models”
Maciej. Polak et al · 2024
Closest in time.
“Extracting Polymer Nanocomposite Samples from Full-Length Documents”
Ghazal Khalighinejad, Defne Circi, L.. Brinson and Bhuwan Dhingra · 2024
Closest in time.
“Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs”
Aaditya. Singh and DJ Strouse · 2024
Closest in time.
“Calibrating Large Language Models with Sample Consistency”
Qing Lyu et al · 2024
Closest in time.
“About Europe PMC”, https://europepmc.org/About , 2024
EMBL’s European Bioinformatics Institute · 2024
Closest in time.
“arXiv”, https://arxiv.org/ , 2024
Cornell University · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“LLaMA: Open and Efficient Foundation Language Models”
Hugo Touvron et al · 2023
Cited alongside, same era.
“Faith and Fate: Limits of Transformers on Compositionality”
Nouha Dziri et al · 2023
Cited alongside, same era.
“LIMA: Less Is More for Alignment”
Chunting Zhou et al · 2023
Cited alongside, same era.
“SciCrawler GitHub Repository”, https://github.com/MasterAI-EAM/SciCrawler , 2023
MasterAI-EAM · 2023
Cited alongside, same era.
“SciDownl GitHub Repository”, https://github.com/Tishacy/SciDownl , 2023
Tishacy · 2023
Cited alongside, same era.
“Pygetpapers GitHub Repository”, https://github.com/petermr/pygetpapers , 2023
Peter Murray · 2023
Cited alongside, same era.
“Leakage and the reproducibility crisis in machine-learning-based science”
Sayash Kapoor and Arvind Narayanan · 2023
Cited alongside, same era.
“ChemRxiv”, https://chemrxiv.org/engage/chemrxiv/public-dashboard , 2024
American Chemical Society (ACS) et al · 2024
Closest in time.
“About Sci-Hub”, https://sci-hub.ru/about , 2024
Alexandra Elbakyan · 2024
Closest in time.
“Elsevier Developer Portal”, https://dev.elsevier.com , 2024
Elsevier B.V · 2024
Closest in time.
“Knowledge Graph Extraction from Total Synthesis Documents”
Andres M, Zlatko Jončev and Philippe Schwaller · 2024
Closest in time.
“Accelerating Scientific Discovery with Generative Knowledge Extraction, Graph-Based Representation, and Multimodal Intelligent Graph Reasoning”
Markus. Buehler · 2024
Closest in time.
“A Comprehensive Overview of Large Language Models”
Humza Naveed et al · 2024
Closest in time.
“Language agents achieve superhuman synthesis of scientific knowledge”
Michael. Skarlinski et al · 2024
Closest in time.
“Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference”
Wei-Lin Chiang et al · 2024
Closest in time.
“Are large language models superhuman chemists?”
Adrian Mirza et al · 2024
Closest in time.
“Creation of a structured solar cell material dataset and performance prediction using large language models”
Tong Xie et al · 2024
Closest in time.
“LAB-Bench: Measuring Capabilities of Language Models for Biology Research”
Jon. Laurent et al · 2024
Closest in time.
“No ”Zero-Shot” Without Exponential Data: Pretraining Concept Frequency Determines Multimodal Model Performance”
Vishaal Udandarao et al · 2024
Closest in time.
“The Llama 3 Herd of Models”
AI Llama · 2024
Closest in time.
“Improving Large Language Models for Clinical Named Entity Recognition via Prompt Engineering”
Yan Hu et al · 2024
Closest in time.
Vamsi Kommineni, Birgitta König-Ries and Sheeba Samuel · 2024
Closest in time.
“Table Meets LLM: Can Large Language Models Understand Structured Table Data? A Benchmark and Empirical Study”
Yuan Sui et al · 2024
Closest in time.
“Many-Shot In-Context Learning”
Rishabh Agarwal et al · 2024
Closest in time.
“Chain of Thoughtlessness? An Analysis of CoT in Planning”
Kaya Stechly, Karthik Valmeekam and Subbarao Kambhampati · 2024
Closest in time.
“Code Generation with AlphaCodium: From Prompt Engineering to Flow Engineering”
Tal Ridnik, Dedy Kredo and Itamar Friedman · 2024
Closest in time.
“FOFO: A Benchmark to Evaluate LLMs’ Format-Following Capability”
Congying Xia et al · 2024
Closest in time.
“Large Language Models for Scientific Information Extraction: An Empirical Study for Virology”
Mahsa Shamsabadi, Jennifer D’Souza and Sören Auer · 2024
Closest in time.
“Assessment of Fine-Tuned Large Language Models for Real-World Chemistry and Material Science Applications”
Joren van Herck et al · 2024
Closest in time.
“Large Language Models for Inorganic Synthesis Predictions”
Seongmin Kim, Yousung Jung and Joshua Schrier · 2024
Closest in time.
“Leveraging large language models for predictive chemistry”
Kevin Jablonka, Philippe Schwaller, Andres Ortega-Guerrero and Berend Smit · 2024
Closest in time.
“Data-driven analysis of text-mined seed-mediated growth AuNP syntheses”
Sanghoon Lee et al · 2024
Closest in time.
“GoLLIE: Annotation Guidelines improve Zero-Shot Information-Extraction”
Oscar Sainz et al · 2024
Closest in time.
“Extracting Structured Data from Organic Synthesis Procedures Using a Fine-Tuned Large Language Model”
Qianxiang Ai et al · 2024
Closest in time.
“LoRA Learns Less and Forgets Less”
Dan Biderman et al · 2024
Closest in time.
“How Beneficial Is Pretraining on a Narrow Domain-Specific Corpus for Information Extraction about Photocatalytic Water Splitting?”
Taketomo Isazawa and Jacqueline. Cole · 2024
Closest in time.
“Using Large Language Models for Data Extraction from Tables in Materials Literature”
Defne Circi, Ghazal Khalighinejad, Bhuwan Dhingra and L. Brinson · 2024
Closest in time.
“Image and data mining in reticular chemistry powered by GPT-4V”
Zhiling Zheng et al · 2024
Closest in time.
“Using machine-learning and large-language-model extracted data to predict copolymerizations”
Mara Schilling-Wilhelmi and Kevin Jablonka · 2024
Closest in time.
“Automated electrosynthesis reaction mining with multimodal large language models (MLLMs)”
Shi Leong, Sergio Pablo-García, Zijian Zhang and Alán Aspuru-Guzik · 2024
Closest in time.
“DeepSeek-VL: Towards Real-World Vision-Language Understanding”
Haoyu Lu et al · 2024
Closest in time.
“On the Hidden Mystery of OCR in Large Multimodal Models”
Yuliang Liu et al · 2024
Closest in time.
“Probing the limitations of multimodal language models for chemistry and materials research”
Nawaf Alampara et al · 2024
Closest in time.
“Autonomous data extraction from peer reviewed literature for training machine learning models of oxidation potentials”
Siwoo Lee, Stefan Heinen, Danish Khan and O Anatole · 2024
Closest in time.
“DiSCoMaT: Distantly Supervised Composition Extraction from Tables in Materials Science Articles”
Tanishq Gupta et al · 2024
Closest in time.
“OpenChemIE: An Information Extraction Toolkit For Chemistry Literature”
Vincent Fan et al · 2024
Closest in time.
“Empowering Biomedical Discovery with AI Agents”
Shanghua Gao et al · 2024
Closest in time.
“Toward a Team of AI-made Scientists for Scientific Discovery from Gene Expression Data”
Haoyang Liu et al · 2024
Closest in time.
“ProtAgents: Protein discovery via large language model multi-agent collaborations combining physics and machine learning”
A. Ghafarollahi and M.. Buehler · 2024
Closest in time.
“ACEGEN: Reinforcement learning of generative chemical agents for drug discovery”
Albert Bou et al · 2024
Closest in time.
“Augmenting large language models with chemistry tools”
Andres M. et al · 2024
Closest in time.
“The Landscape of Emerging AI Agent Architectures for Reasoning, Planning, and Tool Calling: A Survey”
Tula Masterman, Sandi Besen, Mason Sawtell and Alex Chao · 2024
Closest in time.
“A Review of Large Language Models and Autonomous Agents in Chemistry”
Mayk Ramos, Christopher. Collison and Andrew. White · 2024
Closest in time.
“A survey on large language model based autonomous agents”
Lei Wang et al · 2024
Closest in time.
“Cognitive Architectures for Language Agents”
Theodore. Sumers, Shunyu Yao, Karthik Narasimhan and Thomas. Griffiths · 2024
Closest in time.
“CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing”
Zhibin Gou et al · 2024
Closest in time.
“ChatDev: Communicative Agents for Software Development”
Chen Qian et al · 2024
Closest in time.
“Understanding the planning of LLM agents: A survey”
Xu Huang et al · 2024
Closest in time.
“Large Language Models as Tool Makers”
Tianle Cai et al · 2024
Closest in time.
“CREATOR: Tool Creation for Disentangling Abstract and Concrete Reasoning of Large Language Models”
Cheng Qian et al · 2024
Closest in time.
“CRAFT: Customizing LLMs by Creating and Retrieving from Specialized Toolsets”
Lifan Yuan et al · 2024
Closest in time.
“Identifying the Risks of LM Agents with an LM-Emulated Sandbox”
Yangjun Ruan et al · 2024
Closest in time.
“Prioritizing Safeguarding Over Autonomy: Risks of LLM Agents for Science”
Xiangru Tang et al · 2024
Closest in time.
“AI Agents That Matter”
Sayash Kapoor et al · 2024
Closest in time.
“Grammar-Constrained Decoding for Structured NLP Tasks without Finetuning”
Saibo Geng, Martin Josifoski, Maxime Peyrard and Robert West · 2024
Closest in time.
“Annotating Materials Science Text: A Semi-automated Approach for Crafting Outputs with Gemini Pro”
Hasan. Sayeed, Trupti Mohanty and Taylor. Sparks · 2024
Closest in time.
“Structured Prompt Interrogation and Recursive Extraction of Semantics (SPIRES): a method for populating knowledge bases using zero-shot learning”
J Caufield et al · 2024
Closest in time.
“MatText: Do Language Models Need More than Text & Scale for Materials Modeling?”
Nawaf Alampara, Santiago Miret and Kevin Jablonka · 2024
Closest in time.
“Are LLMs Ready for Real-World Materials Discovery?”
Santiago Miret and N Krishnan · 2024
Closest in time.
“MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation”
Qian Huang, Jian Vora, Percy Liang and Jure Leskovec · 2024
Closest in time.
“SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models”
Xiaoxuan Wang et al · 2024
Closest in time.
“Large Language Models: A Survey”
Shervin Minaee et al · 2024
Closest in time.
“Automated Chemical Reaction Extraction from Scientific Literature”
Jiang Guo et al · 2045
Closest in time.