Fetching the paper…
Reading the bibliography…
This paper investigates whether large language models (LLMs) show agreement in assessing creativity in responses to the Alternative Uses Test (AUT).
E. P. Torrance, “Torrance tests of creative thinking,” Educational and Psychological Measurement , 1966. [Online]. Available: https://doi.org/10.1037/t05532-000
1966
Earlier work this paper cites.
J. P. Guilford, “Creativity: Yesterday, today and tomorrow,” The Journal of Creative Behavior , vol. 1, no. 1, pp. 3–14, 1967. [Online]. Available: https://onlinelibrary.wiley.com/doi/abs/10.1002/j.2162-6057.1967.tb00002.x
1967
Earlier work this paper cites.
A. Jordanous, “A standardised procedure for evaluating creative systems: Computational creativity evaluation based on what it is to be creative,” Cognitive Computation , vol. 4, no. 3, pp. 246–279, 2012. [Online]. Available: https://kar.kent.ac.uk/42379/
2012
Earlier work this paper cites.
C. França, L. F. W. Góes, A. Amorim, R. Rocha, and A. R. Da Silva, “Regent-dependent creativity: A domain independent metric for the assessment of creative artifacts,” in Proceedings of the Seventh International Conference on Computational Creativity . Citeseer, 2016, pp. 68–75
2016
Earlier work this paper cites.
R. E. Beaty and D. R. Johnson, “Automating creativity assessment with semdis: An open platform for computing semantic distance,” Behavior Research Methods , vol. 53, pp. 757–780, 2021. [Online]. Available: https://doi.org/10.3758/s13428-020-01453-w
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
T. Yang, Q. Zhang, Z. Sun, and Y. Hou, “Automatic assessment of divergent thinking in chinese language with transdis: A transformer-based language model approach,” Behavior Research Methods , vol. 56, no. 6, p. 5798–5819, Dec. 2023. [Online]. Available: http://dx.doi.org/10.3758/s13428-023-02313-z
2023
Earlier work this paper cites.
F. Góes, P. Sawicki, M. Grzes, M. Volpe, and J. Watson, “Pushing gpt’s creativity to its limits: Alternative uses and torrance tests,” in Proceedings of the 14th International Conference on Computational Creativity, Ontario, Canada, June 19-23, 2023 , A. Pease, J. M. Cunha, M. Ackerman, and D. G. Brown, Eds. Association for Computational Creativity (ACC), 2023, pp. 342–346. [Online]. Available: https://computationalcreativity.net/iccc23/papers/ICCC-2023_paper_90.pdf
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
F. Góes, P. Sawicki, M. Grzes, M. Volpe, and D. Brown, “Is GPT-4 good enough to evaluate jokes?” in Proceedings of the 14th International Conference on Computational Creativity, Ontario, Canada, June 19-23, 2023 , A. Pease, J. M. Cunha, M. Ackerman, and D. G. Brown, Eds. Association for Computational Creativity (ACC), 2023, pp. 367–371. [Online]. Available: https://computationalcreativity.net/iccc23/papers/ICCC-2023_paper_89.pdf
2023
Earlier work this paper cites.
P. Sawicki, M. Grzes, F. Góes, A. Jordanous, D. Brown, S. Paraskevopoulou, M. Peeperkorn, and A. Khatun, “On the power of special-purpose GPT models to create and evaluate new poetry in old styles,” in Proceedings of the 14th International Conference on Computational Creativity, Ontario, Canada, June 19-23, 2023 , A. Pease, J. M. Cunha, M. Ackerman, and D. G. Brown, Eds. Association for Computational Creativity (ACC), 2023, pp. 10–19. [Online]. Available: https://computationalcreativity.net/iccc23/papers/ICCC-2023_paper_18.pdf
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Closest in time.
M. Desmond, Z. Ashktorab, Q. Pan, C. Dugan, and J. M. Johnson, “Evalullm: Llm assisted evaluation of generative outputs,” in Companion Proceedings of the 29th International Conference on Intelligent User Interfaces , ser. IUI ’24 Companion. New York, NY, USA: Association for Computing Machinery, 2024, p. 30–32. [Online]. Available: https://doi.org/10.1145/3640544.3645216
2024
Closest in time.
B. Goecke, P. V. DiStefano, W. Aschauer, K. Haim, R. Beaty, and B. Forthmann, “Automated scoring of scientific creativity in german,” The Journal of Creative Behavior , 2024. [Online]. Available: https://doi.org/10.1002/jocb.658
2024
Closest in time.
T. Raz, S. Luchini, R. Beaty, and Y. Kenett, “Automated scoring of open-ended question complexity: A large language model approach,” 2024. [Online]. Available: https://doi.org/10.21203/rs.3.rs-3890828/v1
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Gómez-Rodríguez and P. Williams, “A confederacy of models: a comprehensive evaluation of LLMs on creative writing,” in Findings of the Association for Computational Linguistics: EMNLP 2023 , H. Bouamor, J. Pino, and K. Bali, Eds. Singapore: Association for Computational Linguistics, Dec. 2023, pp. 14 504–14 528. [Online]. Available: https://aclanthology.org/2023.findings-emnlp.966
2023
Cited alongside, same era.
2023
Cited alongside, same era.
P. Organisciak, S. Acar, D. Dumas, and K. Berthiaume, “Beyond semantic distance: Automated scoring of divergent thinking greatly improves with large language models,” Thinking Skills and Creativity , vol. 49, p. 101356, 2023. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S1871187123001256
2023
Cited alongside, same era.
E. Hadas and A. Hershkovitz, “Using large language models to evaluate alternative uses task flexibility score,” Thinking Skills and Creativity , vol. 52, p. 101549, 2024. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S1871187124000877
2024
Cited alongside, same era.
2024
Cited alongside, same era.
P. V. DiStefano, J. D. Patterson, and R. E. Beaty, “Automatic scoring of metaphor creativity with large language models,” Creativity Research Journal , pp. 1–15, 2024. [Online]. Available: https://doi.org/10.1080/10400419.2024.2326343
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
G. Franceschelli and M. Musolesi, “Creativity and machine learning: A survey,” ACM Comput. Surv. , vol. 56, no. 11, Jun. 2024. [Online]. Available: https://doi.org/10.1145/3664595
2024
Closest in time.
Y. Chang, X. Wang, J. Wang, Y. Wu, L. Yang, K. Zhu, H. Chen, X. Yi, C. Wang, Y. Wang, W. Ye, Y. Zhang, Y. Chang, P. S. Yu, Q. Yang, and X. Xie, “A survey on evaluation of large language models,” ACM Trans. Intell. Syst. Technol. , vol. 15, no. 3, Mar. 2024. [Online]. Available: https://doi.org/10.1145/3641289
2024
Closest in time.
Y. Wu, H. Iso, P. Pezeshkpour, N. Bhutani, and E. Hruschka, “Less is more for long document summary evaluation by LLMs,” in Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 2: Short Papers) , Y. Graham and M. Purver, Eds. St. Julian’s, Malta: Association for Computational Linguistics, Mar. 2024, pp. 330–343. [Online]. Available: https://aclanthology.org/2024.eacl-short.29
2024
Closest in time.
N. F. Liu, K. Lin, J. Hewitt, A. Paranjape, M. Bevilacqua, F. Petroni, and P. Liang, “Lost in the middle: How language models use long contexts,” Transactions of the Association for Computational Linguistics , vol. 12, pp. 157–173, 2024. [Online]. Available: https://aclanthology.org/2024.tacl-1.9
2024
Closest in time.