Fetching the paper…
Reading the bibliography…
While Large Language Models (LLMs) demonstrate remarkable capabilities in scientific tasks such as literature analysis and experimental design (e.g., accurately extracting key findings from papers or generating coherent experimental procedures), existing evaluation benchmarks primarily assess performance using rich contextual inputs.
Creativity
Guilford, J. P · 1950
Earlier work this paper cites.
An analysis of creativity
Rhodes, M · 1961
Earlier work this paper cites.
Creativity and Intelligence: Exploration with Gifted Students (John Wiley & Sons, New York, 1962)
Getzels, J. W. & Jackson, P. W · 1962
Earlier work this paper cites.
The nature of human intelligence (McGraw-Hill, 1967)
Guilford, J. P · 1967
Earlier work this paper cites.
The social psychology of creativity: A componential conceptualization
Amabile, T. M · 1983
Earlier work this paper cites.
The threshold theory regarding creativity and intelligence: An empirical test with gifted and nongifted children
Runco, M. A. & Albert, R. S · 1986
Earlier work this paper cites.
The creative mind: Myths and mechanisms (Routledge, 2004)
Boden, M. A · 2004
Earlier work this paper cites.
Can only intelligent people be creative? A meta-analysis
Kim, K. H · 2005
Earlier work this paper cites.
Relationship of intelligence and creativity in gifted and non-gifted students: An investigation of threshold theory
Preckel, F., Holling, H. & Wiese, M · 2006
Earlier work this paper cites.
Ability differences among people who have commensurate degrees matter for scientific creativity
Park, G., Lubinski, D. & Benbow, C. P · 2008
Earlier work this paper cites.
Intelligence and creativity of polish middle-school students: Looking for the threshold hypothesis
Gralewski, J., Weremczuk, E. & Karwowski, M · 2012
Earlier work this paper cites.
The relationship between intelligence and creativity: New support for the threshold hypothesis by means of empirical breakpoint detection
Jauk, E., Benedek, M., Dunst, B. & Neubauer, A. C · 2013
Earlier work this paper cites.
Re-examining prominent measures of divergent and convergent creativity
Cortes, R. A., Weinberger, A. B., Daker, R. J. & Green, A. E · 2019
Earlier work this paper cites.
Hybrid intelligence
Dellermann, D., Ebel, P., Söllner, M. & Leimeister, J. M · 2019
Earlier work this paper cites.
Quantifying the carbon emissions of machine learning
Lacoste, A., Luccioni, A., Schmidt, V. & Dandres, T · 2019
Earlier work this paper cites.
SciBERT: A pretrained language model for scientific text
Beltagy, I., Lo, K. & Cohan, A · 2019
Earlier work this paper cites.
A reappraisal of the threshold hypothesis of creativity and intelligence
Weiss, S., Steger, D., Schroeders, U. & Wilhelm, O · 2020
Earlier work this paper cites.
A research agenda for hybrid intelligence: augmenting human intellect with collaborative, adaptive, responsible, and explainable artificial intelligence
Akata, Z. et al · 2020
Earlier work this paper cites.
The boundary lens: Theorising academic activity
Cohen, E · 2021
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J. et al · 2022
Earlier work this paper cites.
Mapping citizen science through the lens of human-centered AI
Rafner, J. et al · 2022
Earlier work this paper cites.
Scientific discovery in the age of artificial intelligence
Wang, H. et al · 2023
Earlier work this paper cites.
Surprising combinations of research contents and contexts are related to impact and emerge with scientific outsiders from distant disciplines
Shi, F. & Evans, J · 2023
Earlier work this paper cites.
The impact of large language models on scientific discovery: a preliminary study using GPT-4
AI4Science, M. R. & Quantum, M. A · 2023
Earlier work this paper cites.
Forecasting the future of artificial intelligence with machine learning-based link prediction in an exponentially growing knowledge network
Krenn, M. et al · 2023
Earlier work this paper cites.
Creativity in the age of generative AI
Rafner, J., Beaty, R. E., Kaufman, J. C., Lubart, T. & Sherson, J · 2023
Earlier work this paper cites.
Towards game-based assessment of creative thinking
Rafner, J. et al · 2023
Earlier work this paper cites.
Think outside the code: Brainstorming boosts large language models in code generation
Li, X.-Y., Xue, J.-T., Xie, Z. & Li, M · 2023
Earlier work this paper cites.
Is artificial intelligence more creative than humans?: ChatGPT and the divergent association task
Cropley, D · 2023
Cited alongside, same era.
Generative AI enhances individual creativity but reduces the collective diversity of novel content
Doshi, A. R. & Hauser, O. P · 2023
Cited alongside, same era.
Judging LLM-as-a-judge with MT-Bench and Chatbot Arena
Zheng, L. et al · 2023
Cited alongside, same era.
Can large language models be an alternative to human evaluations?
Chiang, C.-H. & Lee, H.-y · 2023
Cited alongside, same era.
Exploring the use of large language models for reference-free text quality evaluation: An empirical study
Chen, Y., Wang, R., Jiang, H., Shi, S. & Xu, R · 2023
Cited alongside, same era.
ChatGPT outperforms crowd workers for text-annotation tasks
Gilardi, F., Alizadeh, M. & Kubli, M · 2023
Assessing and understanding creativity in large language models
Zhao, Y. et al · 2024
Closest in time.
Length-controlled AlpacaEval: A simple way to debias automatic evaluators
Dubois, Y., Galambosi, B., Liang, P. & Hashimoto, T. B · 2024
Closest in time.
From crowdsourced data to high-quality benchmarks: Arena-Hard and BenchBuilder pipeline
Li, T. et al · 2024
Closest in time.
ChatGPT rates natural language explanation quality like humans: But on which scales?
Huang, F., Kwak, H., Park, K. & An, J · 2024
Closest in time.
Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form Text
Badshah, S. & Sajjad, H · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Accelerating science with human-aware artificial intelligence
Sourati, J. & Evans, J. A · 2023
Cited alongside, same era.
Achiam, J. et al · 2023
Cited alongside, same era.
Bai, J. et al · 2023
Cited alongside, same era.
Jiang, A. Q. et al · 2023
Cited alongside, same era.
Solving olympiad geometry without human demonstrations
Trinh, T. H., Wu, Y., Le, Q. V., He, H. & Luong, T · 2024
Cited alongside, same era.
On the relationship between creative potential and creative achievement: Challenges and future directions
Benedek, M · 2024
Cited alongside, same era.
Closest in time.
Replacing judges with juries: Evaluating LLM generations with a panel of diverse models
Verga, P. et al · 2024
Closest in time.
Prometheus 2: An open source language model specialized in evaluating other language models
Kim, S. et al · 2024
Closest in time.
LiveCodeBench: Holistic and contamination free evaluation of large language models for code
Jain, N. et al · 2024
Closest in time.
Simple synthetic data reduces sycophancy in large language models
Wei, J., Huang, D., Lu, Y., Zhou, D. & Le, Q. V · 2024
Closest in time.
The Claude 3 model family: Opus, Sonnet, Haiku
AI Anthropic · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Team, G. et al · 2024
Closest in time.
Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model
Liu, A. et al · 2024
Closest in time.
Dubey, A. et al · 2024
Closest in time.
The Amazon Nova family of models: Technical report and model card (2024)
Intelligence, A. A. G · 2024
Closest in time.
Abdin, M. et al · 2024
Closest in time.
Gottweis, J. et al · 2025
Closest in time.
We’re different, we’re the same: Creative homogeneity across LLMs
Wenger, E. & Kenett, Y · 2025
Closest in time.
LiveBench: A challenging, contamination-limited LLM benchmark
White, C. et al · 2025
Closest in time.
Gu, J. et al · 2025
Closest in time.
SycEval: Evaluating LLM sycophancy
Fanous, A. et al · 2025
Closest in time.
A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions
Huang, L. et al · 2025
Closest in time.
How do hackathons foster creativity? towards ai collaborative evaluation of creativity at scale
Falk, J. et al · 2025
Closest in time.
Evaluating ai’s ideas: The role of individual creativity and expertise in human-ai co-creativity (2025)
DiStefano, P. V. et al · 2025
Closest in time.
EcoLogits: Evaluate the environmental impacts of generative AI (2025)
Rince, S. & Banse, A · 2025
Closest in time.
Qwen: et al · 2025
Closest in time.
DeepSeek-AI · 2025
Closest in time.
DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning
DeepSeek-AI · 2025
Closest in time.