Fetching the paper…
Reading the bibliography…
Generating novel and creative scientific hypotheses is a cornerstone in achieving Artificial General Intelligence.
Evaluation metrics for language models
Stanley F Chen, Douglas Beeferman, and Roni Rosenfeld · 1998
Earlier work this paper cites.
Creativity in science: Chance, logic, genius, and zeitgeist
Dean Keith Simonton · 2004
Earlier work this paper cites.
The standard definition of creativity
Mark A Runco and Garrett J Jaeger · 2012
Earlier work this paper cites.
New measures for evaluating creativity in scientific publications
Simona Doboli, Fanshu Zhao, and Alex Doboli · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2015
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Taku Kudo and John Richardson · 2018
Earlier work this paper cites.
Creativity in science-scientific essay
Heidi Angell Strøm · 2018
Earlier work this paper cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Can gpt-3 pass a writer’s turing test?
Katherine Elkins and Jon Chun · 2020
Earlier work this paper cites.
Keep calm and explore: Language models for action generation in text-based games
Shunyu Yao, Rohan Rao, Matthew Hausknecht, and Karthik Narasimhan · 2020
Earlier work this paper cites.
How faithful is your synthetic data? sample-level metrics for evaluating and auditing generative models
Ahmed M. Alaa, Boris van Breugel, Evgeny S. Saveliev, and Mihaela van der Schaar · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Earlier work this paper cites.
Can a transformer pass the wug test? tuning copying bias in neural morphological inflection models
Ling Liu and Mans Hulden · 2021
Earlier work this paper cites.
How much do language models copy from their training data? evaluating linguistic novelty in text generation using raven
R. Thomas McCoy, Paul Smolensky, Tal Linzen, Jianfeng Gao, and Asli Celikyilmaz · 2021
Earlier work this paper cites.
Do as i can, not as i say: Grounding language in robotic affordances
Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, et al · 2022
Cited alongside, same era.
Contrastive decoding: Open-ended text generation as optimization
Xiang Lisa Li, Ari Holtzman, Daniel Fried, Percy Liang, Jason Eisner, Tatsunori Hashimoto, Luke Zettlemoyer, and Mike Lewis · 2022
Cited alongside, same era.
Typical decoding for natural language generation
Clara Meister, Tiago Pimentel, Gian Wiher, and Ryan Cotterell · 2022
Cited alongside, same era.
A contrastive framework for neural text generation
Yixuan Su, Tian Lan, Yan Wang, Dani Yogatama, Lingpeng Kong, and Nigel Collier · 2022
Cited alongside, same era.
Self-consistency improves chain of thought reasoning in language models
User-controlled knowledge fusion in large language models: Balancing creativity and hallucination
Chen Zhang · 2023
Later among the works it cites.
Can large language models transform computational social science?
Caleb Ziems, William B. Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang, and Diyi Yang · 2023
Later among the works it cites.
Mle-bench: Evaluating machine learning agents on machine learning engineering
Jun Shern Chan, Neil Chowdhury, Oliver Jaffe, James Aung, Dane Sherburn, Evan Mays, Giulio Starace, Kevin Liu, Leon Maksin, Tejal Patwardhan, et al · 2024
Later among the works it cites.
Ziru Chen, Shijie Chen, Yuting Ning, Qianheng Zhang, Boshi Wang, Botao Yu, Yifei Li, Zeyi Liao, Chen Wei, Zitong Lu, et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou · 2022
Cited alongside, same era.
Science in the age of large language models
Abeba Birhane, Atoosa Kasirzadeh, David Leslie, and Sandra Wachter · 2023
Cited alongside, same era.
Reasoning with language model is planning with world model
Shibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong, Zhen Wang, Daisy Zhe Wang, and Zhiting Hu · 2023
Cited alongside, same era.
Halueval: A large-scale hallucination evaluation benchmark for large language models
Junyi Li, Xiaoxue Cheng, Wayne Xin Zhao, Jianyun Nie, and Ji rong Wen · 2023
Cited alongside, same era.
Improving knowledge extraction from llms for robotic task learning through agent analysis
James R Lindes and Wray Peter · 2023
Cited alongside, same era.
Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models
Potsawee Manakul, Adian Liusie, and Mark John Francis Gales · 2023
Cited alongside, same era.
Sources of hallucination by large language models on inference tasks
Nick McKenna, Tianyi Li, Liang Cheng, Mohammad Javad Hosseini, Mark Johnson, and Mark Steedman · 2023
Cited alongside, same era.
Numeracy from literacy: Data science as an emergent skill from large language models
David Noever and Forrest McKee · 2023
Cited alongside, same era.
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al · 2024
Later among the works it cites.
Peter Alexander Jansen, Marc-Alexandre Cot’e, Tushar Khot, Erin Bransom, Bhavana Dalvi, Bodhisattwa Prasad Majumder, Oyvind Tafjord, and Peter Clark · 2024
Later among the works it cites.
Can large language models unlock novel scientific research ideas?
Sandeep Kumar, Tirthankar Ghosal, Vinayak Goyal, and Asif Ekbal · 2024
Later among the works it cites.
Discoverybench: Towards data-driven discovery with large language models
Bodhisattwa Prasad Majumder, Harshit Surana, Dhruv Agarwal, Bhavana Dalvi, Abhijeetsingh Meena, Aryan Prakhar, Tirth Vora, Tushar Khot, Ashish Sabharwal, and Peter Clark · 2024
Later among the works it cites.
Collaborative gym: A framework for enabling and evaluating human-agent collaboration
Yijia Shao, Vinay Samuel, Yucheng Jiang, John Yang, and Diyi Yang · 2024
Later among the works it cites.
Can llms generate novel research ideas? a large-scale human study with 100+ nlp researchers
Chenglei Si, Diyi Yang, and Tatsunori Hashimoto · 2024
Later among the works it cites.
The virtual lab: Ai agents design new sars-cov-2 nanobodies with experimental validation
Kyle Swanson, Wesley Wu, Nash L. Bulaong, John E. Pak, and James Zou · 2024
Later among the works it cites.
A comprehensive survey of hallucination mitigation techniques in large language models
S. M. Towhidul Islam Tonmoy, S. M. Mehedi Zaman, Vinija Jain, Anku Rani, Vipula Rawte, Aman Chadha, and Amitava Das · 2024
Later among the works it cites.
Large language models for causal hypothesis generation in science
Kai-Hendrik Cohrs, Emiliano Diaz, Vasileios Sitokonstantinou, Gherardo Varando, and Gustau Camps-Valls · 2025
Closest in time.
Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Anil Palepu, Petar Sirkovic, Artiom Myaskovsky, Felix Weissenberger, Keran Rong, Ryutaro Tanno, et al · 2025
Closest in time.
Llm4sr: A survey on large language models for scientific research
Ziming Luo, Zonglin Yang, Zexin Xu, Wei Yang, and Xinya Du · 2025
Closest in time.
Mlgym: A new framework and benchmark for advancing ai research agents
Deepak Nathani, Lovish Madaan, Nicholas Roberts, Nikolay Bashlykov, Ajay Menon, Vincent Moens, Amar Budhiraja, Despoina Magka, Vladislav Vorotilov, Gaurav Chaurasia, et al · 2025
Closest in time.
Agentrxiv: Towards collaborative autonomous research
Samuel Schmidgall and Michael Moor · 2025
Closest in time.
Agent laboratory: Using llm agents as research assistants
Samuel Schmidgall, Yusheng Su, Ze Wang, Ximeng Sun, Jialian Wu, Xiaodong Yu, Jiang Liu, Zicheng Liu, and Emad Barsoum · 2025
Closest in time.