Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) demonstrate promising capabilities in solving scientific problems but often suffer from the issue of hallucination.
The Adaptive Decision Maker
Payne, J. W., Bettman, J. R., and Johnson, E. J · 1993
Earlier work this paper cites.
Unskilled and unaware of it: how difficulties in recognizing one’s own incompetence lead to inflated self-assessments
Kruger, J. and Dunning, D · 1999
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Chemberta: large-scale self-supervised pretraining for molecular property prediction
Chithrananda, S., Grand, G., and Ramsundar, B · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., and others, · 2022
Earlier work this paper cites.
Mind’s eye: Grounded language model reasoning through simulation
Liu, R., Wei, J., Gu, S. S., Wu, T., Vosoughi, S., Cui, C., Zhou, D., and Dai, A. M · 2022
Earlier work this paper cites.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Lu, P., Mishra, S., Xia, T., Qiu, L., Chang, K., Zhu, S., Tafjord, O., Clark, P., and Kalyan, A · 2022
Earlier work this paper cites.
Biogpt: generative pre-trained transformer for biomedical text generation and mining
Luo, R., Sun, L., Xia, Y., Qin, T., Zhang, S., Poon, H., and Liu, T · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., and others, · 2022
Earlier work this paper cites.
Galactica: A large language model for science
Taylor, R., Kardas, M., Cucurull, G., Scialom, T., Hartshorn, A., Saravia, E., Poulton, A., Kerkez, V., and Stojnic, R · 2022
Earlier work this paper cites.
Webshop: Towards scalable real-world web interaction with grounded language agents
Yao, S., Chen, H., Yang, J., and Narasimhan, K · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., and others, · 2023
Earlier work this paper cites.
Chemcrow: Augmenting large-language models with chemistry tools
Bran, A. M., Cox, S., Schilter, O., Baldassari, C., White, A. D., and Schwaller, P · 2023
Earlier work this paper cites.
Large language models as tool makers
Cai, T., Wang, X., Ma, T., Chen, X., and Zhou, D · 2023
Earlier work this paper cites.
Raft: Reward ranked finetuning for generative foundation model alignment
Dong, H., Xiong, W., Goyal, D., Zhang, Y., Chow, W., Pan, R., Diao, S., Zhang, J., Shum, K., and Zhang, T · 2023
Cited alongside, same era.
Metatool benchmark for large language models: Deciding whether to use tools and which to use
Huang, Y., Shi, J., Li, Y., Fan, C., Wu, S., Zhang, Q., Liu, Y., Zhou, P., Wan, Y., Gong, N. Z., and others, · 2023
Cited alongside, same era.
Enhancing large language models with climate resources
Kraus, M., Bingler, J., Leippold, M., Schimanski, T., Senni, C. C., Stammbach, D., Vaghefi, S., and Webersinke, N · 2023
Cited alongside, same era.
Mycrunchgpt: A chatgpt assisted framework for scientific machine learning
Kumar, V. V., Gleyzer, L., Kahana, A., Shukla, K., and Karniadakis, G. E · 2023
Cited alongside, same era.
Rlaif: Scaling reinforcement learning from human feedback with ai feedback
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., and others, · 2024
Closest in time.
Detecting hallucinations in large language models using semantic entropy
Farquhar, S., Kossen, J., Kuhn, L., and Gal, Y · 2024
Closest in time.
Measuring mathematical problem solving with the math dataset, 2021
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J · 2024
Closest in time.
Llm and simulation as bilevel optimizers: A new paradigm to advance physical scientific discovery
Ma, P., Wang, T., Guo, M., Sun, Z., Tenenbaum, J. B., Rus, D., Gan, C., and Matusik, W · 2024
Closest in time.
Simpo: Simple preference optimization with a reference-free reward
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lee, H., Phatale, S., Mansoor, H., Mesnard, T., Ferret, J., Lu, K., Bishop, C., Hall, E., Carbune, V., Rastogi, A., and others, · 2023
Cited alongside, same era.
Gorilla: Large language model connected with massive apis, 2023
Patil, S. G., Zhang, T., Wang, X., and Gonzalez, J. E · 2023
Cited alongside, same era.
CREATOR: Tool creation for disentangling abstract and concrete reasoning of large language models
Qian, C., Han, C., Fung, Y., Qin, Y., Liu, Z., and Ji, H · 2023
Cited alongside, same era.
Toolllm: Facilitating large language models to master 16000+ real-world apis
Qin, Y., Liang, S., Ye, Y., Zhu, K., Yan, L., Lu, Y., Lin, Y., Cong, X., Tang, X., Qian, B., and others, · 2023
Cited alongside, same era.
Training language models with language feedback at scale
Scheurer, J., Campos, J. A., Korbak, T., Chan, J. S., Chen, A., Cho, K., and Perez, E · 2023
Cited alongside, same era.
Toolformer: Language models can teach themselves to use tools
Schick, T., Dwivedi-Yu, J., Dessi, R., Raileanu, R., Lomeli, M., Hambro, E., Zettlemoyer, L., Cancedda, N., and Scialom, T · 2023
Cited alongside, same era.
Toolalpaca: Generalized tool learning for language models with 3000 simulated cases, 2023
Tang, Q., Deng, Z., Lin, H., Han, X., Liang, Q., and Sun, L · 2023
Cited alongside, same era.
Deep bayesian active learning for accelerating stochastic simulation
Wu, D., Niu, R., Chinazzi, M., Vespignani, A., Ma, Y., and Yu, R · 2023
Cited alongside, same era.
Meng, Y., Xia, M., and Chen, D · 2024
Closest in time.
Multi-fidelity residual neural processes for scalable surrogate modeling
Niu, R., Wu, D., Kim, K., Ma, Y., Watson-Parris, D., and Yu, R · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2024
Closest in time.
GPQA: A graduate-level google-proof q&a benchmark
Rein, D., Hou, B. L., Stickland, A. C., Petty, J., Pang, R. Y., Dirani, J., Michael, J., and Bowman, S. R · 2024
Closest in time.
Preference ranking optimization for human alignment
Song, F., Yu, B., Li, M., Yu, H., Huang, F., Li, Y., and Wang, H · 2024
Closest in time.
Climategpt: Towards ai synthesizing interdisciplinary research on climate change
Thulke, D., Gao, Y., Pelser, P., Brune, R., Jalota, R., Fok, F., Ramos, M., Wyk, I., Nasir, A., Goldstein, H., and others, · 2024
Closest in time.
Yang, A., Yang, B., Hui, B., Zheng, B., Yu, B., Zhou, C., Li, C., Li, C., Liu, D., Huang, F., and others, · 2024
Closest in time.
Tooling or not tooling? the impact of tools on language agents for chemistry problem solving
Yu, B., Baker, F. N., Chen, Z., Herb, G., Gou, B., Adu-Ampratwum, D., Ning, X., and Sun, H · 2024
Closest in time.
AgentTuning: Enabling generalized agent abilities for LLMs
Zeng, A., Liu, M., Lu, R., Wang, B., Liu, X., Dong, Y., and Tang, J · 2024
Closest in time.
Llamafactory: Unified efficient fine-tuning of 100+ language models
Zheng, Y., Zhang, R., Zhang, J., Ye, Y., Luo, Z., Feng, Z., and Ma, Y · 2024
Closest in time.
The claude 3 model family: Opus, sonnet, haiku
Anthropic, · 2025
Closest in time.