Fetching the paper…
Reading the bibliography…
We design a suite of minimal algorithmic tasks that are a loose abstraction of open-ended real-world tasks.
A mathematical theory of communication
Shannon, C. E · 1948
Earlier work this paper cites.
Prediction and entropy of printed english
Shannon, C. E · 1951
Earlier work this paper cites.
A review of mental leaps: Analogy in creative thought
Hofstadter, D · 1995
Earlier work this paper cites.
Mental leaps: analogy in creative thought
Holyoak, K. J. and Thagard, P · 1995
Earlier work this paper cites.
Creativity: Flow and the Psychology of Discovery and Invention
Csikszentmihalyi, M · 1996
Earlier work this paper cites.
The Creative Mind - Myths and Mechanisms (2. ed.)
Boden, M. A · 2003
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
Lower bounds for reductions
Kääriäinen, M · 2006
Earlier work this paper cites.
Driven by compression progress: A simple principle explains essential aspects of subjective beauty, novelty, surprise, interestingness, attention, curiosity, creativity, art, science, music, jokes
Schmidhuber, J · 2009
Earlier work this paper cites.
Efficient reductions for imitation learning
Ross, S. and Bagnell, D · 2010
Earlier work this paper cites.
The standard definition of creativity
Runco, M. A. and Jaeger, G. J · 2012
Earlier work this paper cites.
Cognitive science: Leap of thought
Callaway, E · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2014
Earlier work this paper cites.
Generating sentences from a continuous space
Bowman, S. R., Vilnis, L., Vinyals, O., Dai, A. M., Józefowicz, R., and Bengio, S · 2016
Earlier work this paper cites.
Abstractive text summarization using sequence-to-sequence rnns and beyond, 2016
Nallapati, R., Zhou, B., dos santos, C. N., Gulcehre, C., and Xiang, B · 2016
Earlier work this paper cites.
Mode regularized generative adversarial networks
Che, T., Li, Y., Jacob, A. P., Bengio, Y., and Li, W · 2017
Earlier work this paper cites.
Noisy networks for exploration
Fortunato, M., Azar, M. G., Piot, B., Menick, J., Osband, I., Graves, A., Mnih, V., Munos, R., Hassabis, D., Pietquin, O., Blundell, C., and Legg, S · 2017
Earlier work this paper cites.
Z-forcing: Training stochastic recurrent networks
Goyal, A., Sordoni, A., Côté, M.-A., Ke, N. R., and Bengio, Y · 2017
Earlier work this paper cites.
Parameter space noise for exploration
Plappert, M., Houthooft, R., Dhariwal, P., Sidor, S., Chen, R. Y., Chen, X., Asfour, T., Abbeel, P., and Andrychowicz, M · 2017
Earlier work this paper cites.
Improved variational autoencoders for text modeling using dilated convolutions
Yang, Z., Hu, Z., Salakhutdinov, R., and Berg-Kirkpatrick, T · 2017
Earlier work this paper cites.
Non-autoregressive neural machine translation
Gu, J., Bradbury, J., Xiong, C., Li, V. O. K., and Socher, R · 2018
Earlier work this paper cites.
Theoretical insights into memorization in gans
Nagarajan, V., Raffel, C., and Goodfellow, I. J · 2018
Earlier work this paper cites.
Narayan, S., Cohen, S. B., and Lapata, M · 2018
Earlier work this paper cites.
Texygen: A benchmarking platform for text generation models
Zhu, Y., Lu, S., Zheng, L., Guo, J., Zhang, W., Wang, J., and Yu, Y · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Earlier work this paper cites.
A big data approach to computational creativity: The curious case of chef watson
Varshney, L. R., Pinel, F., Varshney, K. R., Bhattacharjya, D., Schörgendorfer, A., and Chee, Y · 2019
Earlier work this paper cites.
Extracting training data from large language models
Carlini, N., Tramèr, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T. B., Song, D. X., Erlingsson, Ú., Oprea, A., and Raffel, C · 2020
Earlier work this paper cites.
Generative adversarial networks
Goodfellow, I. J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A. C., and Bengio, Y · 2020
Earlier work this paper cites.
The pitfalls of simplicity bias in neural networks
Shah, H., Tamuly, K., Raghunathan, A., Jain, P., and Netrapalli, P · 2020
Earlier work this paper cites.
Leap-of-thought: Teaching pre-trained models to systematically reason over implicit knowledge
Talmor, A., Tafjord, O., Clark, P., Goldberg, Y., and Berant, J · 2020
Earlier work this paper cites.
Structured denoising diffusion models in discrete state-spaces
Austin, J., Johnson, D. D., Ho, J., Tarlow, D., and van den Berg, R · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Earlier work this paper cites.
Argmax flows and multinomial diffusion: Learning categorical distributions
Hoogeboom, E., Nielsen, D., Jaini, P., Forr’e, P., and Welling, M · 2021
Earlier work this paper cites.
Why machine reading comprehension models learn shortcuts?
Lai, Y., Zhang, C., Feng, Y., Huang, Q., and Zhao, D · 2021
Earlier work this paper cites.
Language models enable zero-shot prediction of the effects of mutations on protein function
Meier, J., Rao, R., Verkuil, R., Liu, J., Sercu, T., and Rives, A · 2021
Earlier work this paper cites.
Gradient starvation: A learning proclivity in neural networks
Pezeshki, M., Kaba, S., Bengio, Y., Courville, A. C., Precup, D., and Lajoie, G · 2021
Earlier work this paper cites.
Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences
Rives, A., Meier, J., Sercu, T., Goyal, S., Lin, Z., Liu, J., Guo, D., Ott, M., Zitnick, C. L., Ma, J., and Fergus, R · 2021
Earlier work this paper cites.
Alchemy: A benchmark and analysis toolkit for meta-reinforcement learning agents
Wang, J. X., King, M., Porcel, N., Kurth-Nelson, Z., Zhu, T., Deck, C., Choy, P., Cassin, M., Reynolds, M., Song, F., Buttimore, G., Reichert, D. P., Rabinowitz, N. C., Matthey, L., Hassabis, D., Lerchner, A., and Botvinick, M. M · 2021
Earlier work this paper cites.
Efficient training of language models to fill in the middle
Bavarian, M., Jun, H., Tezak, N. A., Schulman, J., McLeavey, C., Tworek, J., and Chen, M · 2022
Earlier work this paper cites.
Incoder: A generative model for code infilling and synthesis
Fried, D., Aghajanyan, A., Lin, J., Wang, S. I., Wallace, E., Shi, F., Zhong, R., tau Yih, W., Zettlemoyer, L., and Lewis, M · 2022
Earlier work this paper cites.
Fine-tuning pre-trained language models with noise stability regularization
Hua, H., Li, X., Dou, D., Xu, C., and Luo, J · 2022
Earlier work this paper cites.
Competition-level code generation with alphacode
Li, Y., Choi, D. H., Chung, J., Kushman, N., Schrittwieser, J., Leblond, R., Eccles, T., Keeling, J., Gimeno, F., Lago, A. D., Hubert, T., Choy, P., de Masson d’Autume, C., Babuschkin, I., Chen, X., Huang, P., Welbl, J., Gowal, S., Cherepanov, A., Molloy, J., Mankowitz, D. J., Robson, E. S., Kohli, P., de Freitas, N., Kavukcuoglu, K., and Vinyals, O · 2022
Earlier work this paper cites.
Language models are better than humans at next-token prediction
Shlegeris, B., Roger, F., Chan, L., and McLean, E · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E. H., Le, Q. V., and Zhou, D · 2022
Earlier work this paper cites.
Interactive visual reasoning under uncertainty
Xu, M., Jiang, G., Zhang, C., Zhu, S.-C., and Zhu, Y · 2022
Earlier work this paper cites.
On the inconsistencies of conditionals learned by masked language models
Young, T. and You, Y · 2022
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al · 2023
Cited alongside, same era.
Quantifying memorization across neural language models
Carlini, N., Ippolito, D., Jagielski, M., Lee, K., Tramèr, F., and Zhang, C · 2023
Cited alongside, same era.
Probing the "creativity" of large language models: Can models produce divergent semantic association?
Chen, H. and Ding, N · 2023
Cited alongside, same era.
Increasing diversity while maintaining accuracy: Text data generation with large language models and human interventions
Chung, J. J. Y., Kamar, E., and Amershi, S · 2023
Discoveryworld: A virtual environment for developing and evaluating automated scientific discovery agents
Jansen, P. A., Cot’e, M.-A., Khot, T., Bransom, E., Dalvi, B., Majumder, B. P., Tafjord, O., and Clark, P · 2024
Later among the works it cites.
Calibrated language models must hallucinate
Kalai, A. T. and Vempala, S. S · 2024
Later among the works it cites.
On the limits of language generation: Trade-offs between hallucination and mode collapse
Kalavasis, A., Mehrotra, A., and Velegkas, G · 2024
Later among the works it cites.
An analytic theory of creativity in convolutional diffusion models, 2024
Kamb, M. and Ganguli, S · 2024
Later among the works it cites.
Towards an understanding of stepwise inference in transformers: A synthetic graph navigation model
Khona, M., Okawa, M., Hula, J., Ramesh, R., Nishi, K., Dick, R. P., Lubana, E. S., and Tanaka, H · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Introduction to latent variable energy-based models: A path towards autonomous machine intelligence
Dawid, A. and LeCun, Y · 2023
Cited alongside, same era.
Autoregressive modeling with lookahead attention
Du, L., Mei, H., and Eisner, J · 2023
Cited alongside, same era.
Towards revealing the mystery behind chain of thought: a theoretical perspective
Feng, G., Zhang, B., Gu, Y., Ye, H., He, D., and Wang, L · 2023
Cited alongside, same era.
On the creativity of large language models
Franceschelli, G. and Musolesi, M · 2023
Cited alongside, same era.
Diffuseq: Sequence to sequence text generation with diffusion models
Gong, S., Li, M., Feng, J., Wu, Z., and Kong, L · 2023
Cited alongside, same era.
Protein design with guided discrete diffusion
Gruver, N., Stanton, S., Frey, N. C., Rudner, T. G. J., Hotzel, I., Lafrance-Vanasse, J., Rajpal, A., Cho, K., and Wilson, A. G · 2023
Cited alongside, same era.
Can LLMs generate random numbers? evaluating LLM sampling in controlled domains
Hopkins, A. K., Renda, A., and Carbin, M · 2023
Cited alongside, same era.
Later among the works it cites.
The factorization curse: Which tokens you predict underlie the reversal curse and more
Kitouni, O., Nolte, N., Bouchacourt, D., Williams, A., Rabbat, M., and Ibrahim, M · 2024
Later among the works it cites.
Language generation in the limit
Kleinberg, J. M. and Mullainathan, S · 2024
Later among the works it cites.
Dipper: Diversity in prompts for producing large language model ensembles in reasoning tasks, 2024
Lau, G. K. R., Hu, W., Liu, D., Chen, J., Ng, S.-K., and Low, B. K. H · 2024
Later among the works it cites.
Do large language models need sensory grounding for meaning and understanding?
LeCun, Y · 2024
Later among the works it cites.
Teaching arithmetic to small transformers
Lee, N., Sreenivasan, K., Lee, J. D., Lee, K., and Papailiopoulos, D · 2024
Later among the works it cites.
Aidanbench: Stress-testing language model creativity on open-ended questions
McLaughlin, A., Campbell, J., Uppuluri, A., and Yang, Y · 2024
Later among the works it cites.
The expressive power of transformers with chain of thought
Merrill, W. and Sabharwal, A · 2024
Later among the works it cites.
Mirowski, P. W., Love, J., Mathewson, K. W., and Mohamed, S · 2024
Later among the works it cites.
Diversity of thought improves reasoning abilities of llms, 2024
Naik, R., Chandrasekaran, V., Yuksekgonul, M., Palangi, H., and Nushi, B · 2024
Later among the works it cites.
Step-by-step diffusion: An elementary tutorial, 2024
Nakkiran, P., Bradley, A., Zhou, H., and Advani, M · 2024
Later among the works it cites.
Sequence modeling and design from molecular to genome scale with evo
Nguyen, E., Poli, M., Durrant, M. G., Kang, B., Katrekar, D., Li, D. B., Bartie, L. J., Thomas, A. W., King, S. H., Brixi, G., Sullivan, J., Ng, M. Y., Lewis, A., Lou, A., Ermon, S., Baccus, S. A., Hernandez-Boussard, T., Ré, C., Hsu, P. D., and Hie, B. L · 2024
Later among the works it cites.
Transformers can navigate mazes with multi-step prediction
Nolte, N., Kitouni, O., Williams, A., Rabbat, M., and Ibrahim, M · 2024
Later among the works it cites.
Openai o1 system card
OpenAI · 2024
Later among the works it cites.
Does writing with language models reduce content diversity?
Padmakumar, V. and He, H · 2024
Later among the works it cites.
σ \sigma -gpts: A new approach to autoregressive models
Pannatier, A., Courdier, E., and Fleuret, F · 2024
Later among the works it cites.
Is temperature the creativity parameter of large language models?
Peeperkorn, M., Kouwenhoven, T., Brown, D., and Jordanous, A · 2024
Later among the works it cites.
Mathematical discoveries from program search with large language models
Romera-Paredes, B., Barekatain, M., Novikov, A., Balog, M., Kumar, M. P., Dupont, E., Ruiz, F. J. R., Ellenberg, J. S., Wang, P., Fawzi, O., Kohli, P., and Fawzi, A · 2024
Later among the works it cites.
Understanding transformer reasoning capabilities via graph algorithms
Sanford, C., Fatemi, B., Hall, E., Tsitsulin, A., Kazemi, S. M., Halcrow, J., Perozzi, B., and Mirrokni, V · 2024
Later among the works it cites.
Transformers struggle to learn to search, 2024
Saparov, A., Pawar, S., Pimpalgaonkar, S., Joshi, N., Pang, R. Y., Padmakumar, V., Kazemi, S. M., Kim, N., and He, H · 2024
Later among the works it cites.
Can llms generate novel research ideas? A large-scale human study with 100+ NLP researchers
Si, C., Yang, D., and Hashimoto, T · 2024
Later among the works it cites.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
Snell, C., Lee, J., Xu, K., and Kumar, A · 2024
Later among the works it cites.
Does chatgpt have a poetic style?
Walsh, M., Preus, A., and Gronski, E · 2024
Later among the works it cites.
Make some noise: Unlocking language model parallel inference capability through noisy training
Wang, Y., Luo, X., Wei, F., Liu, Y., Zhu, Q., Zhang, X., Yang, Q., Xu, D., and Che, W · 2024
Later among the works it cites.
Inference scaling laws: An empirical analysis of compute-optimal inference for problem-solving with language models
Wu, Y., Sun, Z., Li, S., Welleck, S., and Yang, Y · 2024
Later among the works it cites.
Do large language models latently perform multi-hop reasoning?
Yang, S., Gribovskaya, E., Kassner, N., Geva, M., and Riedel, S · 2024
Later among the works it cites.
Beyond autoregression: Discrete diffusion for complex reasoning and planning
Ye, J., Gao, J., Gong, S., Zheng, L., Jiang, X., Li, Z., and Kong, L · 2024
Later among the works it cites.
Metamath: Bootstrap your own mathematical questions for large language models
Yu, L., Jiang, W., Shi, H., YU, J., Liu, Z., Zhang, Y., Kwok, J., Li, Z., Weller, A., and Liu, W · 2024
Later among the works it cites.
Assessing and understanding creativity in large language models
Zhao, Y., Zhang, R., Li, W., Huang, D., Guo, J., Peng, S., Hao, Y., Wen, Y., Hu, X., Du, Z., Guo, Q., Li, L., and Chen, Y · 2024
Later among the works it cites.
Let’s think outside the box: Exploring leap-of-thought in large language models with creative humor generation
Zhong, S., Huang, Z., Gao, S., Wen, W., Lin, L., Zitnik, M., and Zhou, P · 2024
Later among the works it cites.
Evaluating sakana’s ai scientist for autonomous research: Wishful thinking or an emerging reality towards ’artificial research intelligence’ (ari)?, 2025
Beel, J., Kan, M.-Y., and Baumgart, M · 2025
Closest in time.
Weight ensembling improves reasoning in language models
Dang, X., Baek, C., Wen, K., Kolter, Z., and Raghunathan, A · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
DeepSeek-AI · 2025
Closest in time.
All that glitters is not novel: Plagiarism in ai generated research, 2025
Gupta, T. and Pruthi, D · 2025
Closest in time.
Simulating 500 million years of evolution with a language model
Hayes, T., Rao, R., Akin, H., Sofroniew, N. J., Oktay, D., Lin, Z., Verkuil, R., Tran, V. Q., Deaton, J., Wiggert, M., Badkundri, R., Shafkat, I., Gong, J., Derry, A., Molina, R. S., Thomas, N., Khan, Y. A., Mishra, C., Kim, C., Bartie, L. J., Nemeth, M., Hsu, P. D., Sercu, T., Candido, S., and Rives, A · 2025
Closest in time.
The belief state transformer
Hu, E. S., Ahn, K., Liu, Q., Xu, H., Tomar, M., Langford, A., Jayaraman, D., Lamb, A., and Langford, J · 2025
Closest in time.
Why llms cannot think and how to fix it, 2025
Jahrens, M. and Martinetz, T · 2025
Closest in time.
Muennighoff, N., Yang, Z., Shi, W., Li, X. L., Fei-Fei, L., Hajishirzi, H., Zettlemoyer, L., Liang, P., Candès, E. J., and Hashimoto, T · 2025
Closest in time.
Looking beyond the next token, 2025
Thankaraj, A., Jiang, Y., Kolter, J. Z., and Bisk, Y · 2025
Closest in time.
Noveltybench: Evaluating language models for humanlike diversity
Zhang, Y., Diddee, H., Holm, S., Liu, H., Liu, X., Samuel, V., Wang, B., and Ippolito, D · 2025
Closest in time.