Fetching the paper…
Reading the bibliography…
Recent advancements have positioned AI, and particularly Large Language Models (LLMs), as transformative tools for scientific research, capable of addressing complex tasks that require reasoning, problem-solving, and decision-making.
A. B. Yoo, M. A. Jette, and M. Grondona, “Slurm: Simple Linux utility for resource management,” in Workshop on job scheduling strategies for parallel processing . Springer, 2003, pp. 44–60
2003
Earlier work this paper cites.
S. Solomon, D. Qin, M. Manning, Z. Chen, M. Marquis, K. Averyt, M. Tignor, and H. Miller, “IPCC fourth assessment report (AR4),” Climate change , vol. 374, 2007
2007
Earlier work this paper cites.
Y. Gal and Z. Ghahramani, “Dropout as a Bayesian approximation: Representing model uncertainty in deep learning,” in international conference on machine learning . PMLR, 2016, pp. 1050–1059
2016
Earlier work this paper cites.
2018
Earlier work this paper cites.
T. Kočiskỳ, J. Schwarz, P. Blunsom, C. Dyer, K. M. Hermann, G. Melis, and E. Grefenstette, “The narrativeqa reading comprehension challenge,” Transactions of the Association for Computational Linguistics , vol. 6, pp. 317–328, 2018
2018
Earlier work this paper cites.
Z. Yang, P. Qi, S. Zhang, Y. Bengio, W. Cohen, R. Salakhutdinov, and C. D. Manning, “HotpotQA: A dataset for diverse, explainable multi-hop question answering,” in Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, 2018
2018
Earlier work this paper cites.
E. Alvaro and A. Yanguas-Gil, “Characterizing the field of atomic layer deposition: Authors, topics, and collaborations,” PLOS ONE , vol. 13, no. 1, pp. 1–19, 01 2018. [Online]. Available: https://doi.org/10.1371/journal.pone.0189137
2018
Earlier work this paper cites.
P. Moritz, R. Nishihara, S. Wang, A. Tumanov, R. Liaw, E. Liang, M. Elibol, Z. Yang, W. Paul, M. I. Jordan et al. , “Ray: A distributed framework for emerging AI applications,” in 13th USENIX symposium on operating systems design and implementation (OSDI 18) , 2018, pp. 561–577
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
Y. Babuji, A. Woodard, Z. Li, D. S. Katz, B. Clifford, R. Kumar, L. Lacinski, R. Chard, J. M. Wozniak, I. Foster et al. , “Parsl: Pervasive parallel programming in Python,” in 28th International Symposium on High-Performance Parallel and Distributed Computing , 2019, pp. 25–36
2019
Earlier work this paper cites.
W. Chen, H. Zha, Z. Chen, W. Xiong, H. Wang, and W. Y. Wang, “HybridQA: A dataset of multi-hop question answering over tabular and textual data,” Findings of the Association for Computational Linguistics: EMNLP 2020 , 2020
2020
Earlier work this paper cites.
Z. Wei, W. Ji, X. Geng, Y. Chen, B. Chen, T. Qin, and D. Jiang, “ChemistryQA: A complex question answering dataset from chemistry,” Oct. 2020. [Online]. Available: https://openreview.net/forum?id=oeHTRAehiFF
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt, “Measuring mathematical problem solving with the MATH dataset,” NeurIPS , 2021
2021
Earlier work this paper cites.
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt, “Measuring massive multitask language understanding,” International Conference on Learning Representations , 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
E. O. Pyzer-Knapp, J. W. Pitera, P. W. J. Staar, S. Takeda, T. Laino, D. P. Sanders, J. Sexton, J. R. Smith, and A. Curioni, “Accelerating materials discovery using artificial intelligence, high performance computing and robotics,” npj Computational Materials , vol. 8, no. 1, p. 84, Apr 2022
2022
Earlier work this paper cites.
R. Vaid, K. Pant, and M. Shrivastava, “Towards fine-grained classification of climate change related social media text,” in 60th Annual Meeting of the Association for Computational Linguistics: Student Research Workshop . Dublin, Ireland: Association for Computational Linguistics, May 2022, pp. 434–443. [Online]. Available: https://aclanthology.org/2022.acl-srw.35
2022
Earlier work this paper cites.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou et al. , “Chain-of-thought prompting elicits reasoning in large language models,” Advances in neural information processing systems , vol. 35, pp. 24 824–24 837, 2022
2022
Earlier work this paper cites.
L. Varanasi, “GPT-4 can ace the bar, but it only has a decent chance of passing the CFA exams. Here’s a list of difficult exams the ChatGPT and GPT-4 have passed.” https://www.businessinsider.com/list-here-are-the-exams-chatgpt-has-passed-so-far-2023-1, Nov. 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, E. Hambro, L. Zettlemoyer, N. Cancedda, and T. Scialom, “Toolformer: Language models can teach themselves to use tools,” Advances in Neural Information Processing Systems , vol. 36, pp. 68 539–68 551, 2023
2023
Earlier work this paper cites.
T. Guo, B. Nan, Z. Liang, Z. Guo, N. Chawla, O. Wiest, X. Zhang et al. , “What can large language models do in chemistry? a comprehensive benchmark on eight tasks,” Advances in Neural Information Processing Systems , vol. 36, pp. 59 662–59 688, 2023
2023
Earlier work this paper cites.
E. Beeching, C. Fourrier, N. Habib, S. Han, N. Lambert, N. Rajani, O. Sanseviero, L. Tunstall, and T. Wolf, “Open llm leaderboard,” 2023
2023
Earlier work this paper cites.
B. Nicolae, T. Islam, R. Ross, H. V. Dam, K. Assogba, P. Shpilker, M. Titov, M. Turilli, T. Wang, O. Kilic, S. Jha, and L. Pouchard, “Building the i (interoperability) of fair for performance reproducibility of large-scale composable workflows in recup,” in REWORDS’23: The 3rd Workshop on Reproducible Workflows, Data Management, and Security (with eScience’23) , Limassol, Cyprus, 2023, pp. 1–7. [Online]. Available: https://hal.inria.fr/hal-04343665
2023
Earlier work this paper cites.
D. A. Boiko, R. MacKnight, B. Kline, and G. Gomes, “Autonomous chemical research with large language models,” Nature , vol. 624, no. 7992, pp. 570–578, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Z. Lin, S. Trivedi, and J. Sun, “Generating with confidence: Uncertainty quantification for black-box large language models,” Transactions on Machine Learning Research , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica, “Efficient Memory Management for Large Language Model Serving with PagedAttention,” in 29th Symposium on Operating Systems Principles . Koblenz Germany: ACM, Oct. 2023, pp. 611–626
2023
Cited alongside, same era.
R. Lacombe, K. Wu, and E. Dilworth, “Climatex: Do llms accurately assess human expert confidence in climate statements?” 2023
2023
Cited alongside, same era.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2023
Cited alongside, same era.
L. Kuhn, Y. Gal, and S. Farquhar, “Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation,” in The Eleventh International Conference on Learning Representations , 2023. [Online]. Available: https://openreview.net/forum?id=VD-AYtP0dve
2023
Cited alongside, same era.
R. Bommasani, P. Liang, and T. Lee, “Holistic evaluation of language models,” Annals of the New York Academy of Sciences , vol. 1525, no. 1, pp. 140–146, 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica, “Efficient memory management for large language model serving with PagedAttention,” in 29th Symposium on Operating Systems Principles , 2023, pp. 611–626
2023
Cited alongside, same era.
OpenAI Team, “GPT-4 Technical Report,” Preprint arXiv:2303.08774 , Mar. 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
B. Wang, W. Chen, H. Pei, C. Xie, M. Kang, C. Zhang, C. Xu, Z. Xiong, R. Dutta, R. Schaeffer et al. , “DecodingTrust: A comprehensive assessment of trustworthiness in GPT models,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Later among the works it cites.
Y. Huang, L. Sun, H. Wang, S. Wu, Q. Zhang, Y. Li, C. Gao, Y. Huang, W. Lyu, Y. Zhang et al. , “Position: Trustllm: Trustworthiness in large language models,” in International Conference on Machine Learning . PMLR, 2024, pp. 20 166–20 270
2024
Later among the works it cites.
Z.-W. Hong, I. Shenfeld, T.-H. Wang, Y.-S. Chuang, A. Pareja, J. R. Glass, A. Srivastava, and P. Agrawal, “Curiosity-driven red-teaming for large language models,” in The Twelfth International Conference on Learning Representations , 2024. [Online]. Available: https://openreview.net/forum?id=4KqkizXgXU
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
X. Du, C. Xiao, and Y. Li, “HaloScope: Harnessing unlabeled LLM generations for hallucination detection,” in The Thirty-eighth Annual Conference on Neural Information Processing Systems , 2024. [Online]. Available: https://openreview.net/forum?id=nfK0ZXFFSn
2024
Later among the works it cites.
J. Duan, H. Cheng, S. Wang, A. Zavalny, C. Wang, R. Xu, B. Kailkhura, and K. Xu, “Shifting attention to relevance: Towards the predictive uncertainty quantification of free-form large language models,” in 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2024, pp. 5050–5063
2024
Later among the works it cites.
2024
Later among the works it cites.
Z. Chen, P. Hong, and S. Madireddy, “Quantifying uncertainty in large language models: Applications in molecular chemistry tasks,” NeurIPS 2024 Workshop on Statistical Foundations of LLMs and Foundation Models , 2024
2024
Later among the works it cites.
D. Thulke, Y. Gao, P. Pelser, R. Brune, R. Jalota, F. Fok, M. Ramos, I. van Wyk, A. Nasir, H. Goldstein, T. Tragemann, K. Nguyen, A. Fowler, A. Stanco, J. Gabriel, J. Taylor, D. Moro, E. Tsymbalov, J. de Waal, E. Matusov, M. Yaghi, M. Shihadah, H. Ney, C. Dugast, J. Dotan, and D. Erasmus, “Climategpt: Towards ai synthesizing interdisciplinary research on climate change,” 2024
2024
Later among the works it cites.
H. Zhang, J. Da, D. Lee, V. Robinson, C. Wu, W. Song, T. Zhao, P. Raja, D. Slack, Q. Lyu, S. Hendryx, R. Kaplan, M. Lunati, and S. Yue, “A careful examination of large language model performance on grade school arithmetic,” 2024, arxiv 2405.00332
2024
Later among the works it cites.
2024
Later among the works it cites.
M. Tian, L. Gao, D. Zhang, X. Chen, C. Fan, X. Guo, R. Haas, P. Ji, K. Krongchon, Y. Li, S. Liu, D. Luo, Y. Ma, H. TONG, K. Trinh, C. Tian, Z. Wang, B. Wu, S. Yin, M. Zhu, K. Lieret, Y. Lu, G. Liu, Y. Du, T. Tao, O. Press, J. Callan, E. A. Huerta, and H. Peng, “Scicode: A research coding benchmark curated by scientists,” in The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track , 2024. [Online]. Available: https://openreview.net/forum?id=ADLaALtdoG
2024
Later among the works it cites.
2024
Later among the works it cites.
A. Maurya, R. Underwood, M. M. Rafique, F. Cappello, and B. Nicolae, “DataStates-LLM: Lazy asynchronous checkpointing for large language models,” in 33rd International Symposium on High-Performance Parallel and Distributed Computing , ser. HPDC ’24. New York, NY, USA: Association for Computing Machinery, 2024, p. 227–239. [Online]. Available: https://doi.org/10.1145/3625549.3658685
2024
Later among the works it cites.
J. Guo, V. Mohanty, J. H. Piazentin Ono, H. Hao, L. Gou, and L. Ren, “Investigating interaction modes and user agency in human-LLM collaboration for domain-specific data analysis,” in Extended Abstracts of the CHI Conference on Human Factors in Computing Systems , ser. CHI ’24. ACM, May 2024, p. 1–9. [Online]. Available: http://dx.doi.org/10.1145/3613905.3651042
2024
Later among the works it cites.
P. Hager, F. Jungmann, R. Holland, K. Bhagat, I. Hubrecht, M. Knauer, J. Vielhauer, M. Makowski, R. Braren, G. Kaissis, and D. Rueckert, “Evaluation and mitigation of the limitations of large language models in clinical decision-making,” Nature Medicine , vol. 30, pp. 2613–2622, 07 2024
2024
Later among the works it cites.
M. Xiong, Z. Hu, X. Lu, Y. Li, J. Fu, J. He, and B. Hooi, “Can LLMs express their uncertainty? An empirical evaluation of confidence elicitation in LLMs,” in The Twelfth International Conference on Learning Representations , 2024. [Online]. Available: https://openreview.net/forum?id=gjeQKFxFpZ
2024
Later among the works it cites.
L. Gao, J. Tow, B. Abbasi, S. Biderman, S. Black, A. DiPofi, C. Foster, L. Golding, J. Hsu, A. Le Noac’h, H. Li, K. McDonell, N. Muennighoff, C. Ociepa, J. Phang, L. Reynolds, H. Schoelkopf, A. Skowron, L. Sutawika, E. Tang, A. Thite, B. Wang, K. Wang, and A. Zou, “A framework for few-shot language model evaluation,” 2024, https://zenodo.org/records/12608602
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
S. Madireddy, C. Xu, F. Cappello, S. Bhattacharya, B. Kailkhura, M. Foltin, T. Kumar, and B. Li, “Comprehensive multi-stage evaluation of language models for scientific skill and safety red-teaming,” Accelerating the Development and Use of Generative AI for Science and Engineering: The Trillion Parameter Consortium (TPC) Workshop at The International Conference for High Performance Computing, Networking, Storage, and Analysis (SC24) , 2024
2024
Later among the works it cites.
G. Benegas, C. Ye, C. Albors, J. C. Li, and Y. S. Song, “Genomic language models: Opportunities and challenges,” Trends in Genetics , 2025
2025
Closest in time.
Center for AI Safety and Scale AI, “Humanity’s last exam,” 2024, accessed: 2025-01-15. [Online]. Available: https://agi.safe.ai/submit
2025
Closest in time.
Argonne, “Polaris supercomputer,” 2025, https://alcf.anl.gov/alcf-resources
2025
Closest in time.