Fetching the paper…
Reading the bibliography…
Standard benchmarks fixate on how well large language model (LLM) agents perform in finance, yet say little about whether they are safe to deploy.
The development of utility theory. i
George J Stigler · 1950
Earlier work this paper cites.
Probabilistic risk analysis: foundations and methods
Tim Bedford and Roger Cooke · 2001
Earlier work this paper cites.
Is the 2007 us sub-prime financial crisis so different? an international historical comparison
Carmen M Reinhart and Kenneth S Rogoff · 2008
Earlier work this paper cites.
Modern portfolio theory and investment analysis
Edwin J Elton, Martin J Gruber, Stephen J Brown, and William N Goetzmann · 2009
Earlier work this paper cites.
Modeling and hazard analysis using stpa
Takuto Ishimatsu, Nancy G Leveson, John Thomas, Masa Katahira, Yuko Miyamoto, and Haruka Nakao · 2010
Earlier work this paper cites.
Basel iii: an overview
Peter King and Heath Tarbert · 2011
Earlier work this paper cites.
Engineering a safer world: Systems thinking applied to safety
Nancy G Leveson · 2016
Earlier work this paper cites.
Dynamic portfolio optimization across hidden market regimes
Peter Nystrup, Henrik Madsen, and Erik Lindström · 2018
Earlier work this paper cites.
Price discovery and the accuracy of consolidated data feeds in the us equity markets
Brian F Tivnan, David Slater, James R Thompson, Tobin A Bergen-Hill, Carl D Burke, Shaun M Brady, Matthew TK Koehler, Matthew T McMahon, Brendan F Tivnan, and Jason G Veneman · 2018
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 2019
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Earlier work this paper cites.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed H. Chi, F. Xia, Quoc Le, and Denny Zhou · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Earlier work this paper cites.
Artificial intelligence risk management framework (ai rmf 1.0)
NIST AI · 2023
Earlier work this paper cites.
Unleashing the potential of prompt engineering in large language models: a comprehensive review
Banghao Chen, Zhaofeng Zhang, Nicolas Langrené, and Shengxin Zhu · 2023
Earlier work this paper cites.
Language models, agent models, and world models: The law for machine reasoning and planning
Zhiting Hu and Tianmin Shu · 2023
Earlier work this paper cites.
Walking a tightrope–evaluating large language models in high-risk domains
Chia-Chien Hung, Wiem Ben Rim, Lindsay Frost, Lars Bruckner, and Carolin Lawrence · 2023
Earlier work this paper cites.
Financebench: A new benchmark for financial question answering
Pranab Islam, Anand Kannappan, Douwe Kiela, Rebecca Qian, Nino Scherrer, and Bertie Vidgen · 2023
Earlier work this paper cites.
Deficiency of large language models in finance: An empirical examination of hallucination
Haoqiang Kang and Xiao-Yang Liu · 2023
Earlier work this paper cites.
Deductive verification of chain-of-thought reasoning
Zhan Ling, Yunhao Fang, Xuanlin Li, Zhiao Huang, Mingu Lee, Roland Memisevic, and Hao Su · 2023
Earlier work this paper cites.
A comprehensive overview of large language models
Humza Naveed, Asad Ullah Khan, Shi Qiu, Muhammad Saqib, Saeed Anwar, Muhammad Usman, Naveed Akhtar, Nick Barnes, and Ajmal Mian · 2023
Earlier work this paper cites.
Hallucination-minimized data-to-answer framework for financial decision-makers
Sohini Roychowdhury, Andres Alvarez, Brian Moore, Marko Krema, Maria Paz Gelpi, Punit Agrawal, Federico Martín Rodríguez, Ángel Rodríguez, José Ramón Cabrejas, Pablo Martínez Serrano, et al · 2023
Earlier work this paper cites.
Sander Schulhoff, Jeremy Pinto, Anaum Khan, Louis-François Bouchard, Chenglei Si, Svetlina Anati, Valen Tagliabue, Anson Liu Kost, Christopher Carnahan, and Jordan Boyd-Graber · 2023
Earlier work this paper cites.
Corex: Pushing the boundaries of complex reasoning through multi-model collaboration
Qiushi Sun, Zhangyue Yin, Xiang Li, Zhiyong Wu, Xipeng Qiu, and Lingpeng Kong · 2023
Earlier work this paper cites.
Benchmarking large language model volatility
Boyang Yu · 2023
Earlier work this paper cites.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al · 2023
Earlier work this paper cites.
Large language model in financial regulatory interpretation
Zhiyu Cao and Zachary Feinstein · 2024
Earlier work this paper cites.
A causal explainable guardrails for large language models
Zhixuan Chu, Yan Wang, Longfei Li, Zhibo Wang, Zhan Qin, and Kui Ren · 2024
Earlier work this paper cites.
Mind2web: Towards a generalist agent for the web
Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Sam Stevens, Boshi Wang, Huan Sun, and Yu Su · 2024
Cited alongside, same era.
Opportunities and challenges of generative-ai in finance
Akshar Prabhu Desai, Tejasvi Ravi, Mohammad Luqman, Ganesh Mallya, Nithya Kota, and Pranjul Yadav · 2024
Cited alongside, same era.
Large language model agent in financial trading: A survey, 2024
Han Ding, Yinheng Li, Junhao Wang, and Hang Chen · 2024
Cited alongside, same era.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Cited alongside, same era.
Agent ai: Surveying the horizons of multimodal interaction
Zane Durante, Qiuyuan Huang, Naoki Wake, Ran Gong, Jae Sung Park, Bidipta Sarkar, Rohan Taori, Yusuke Noda, Demetri Terzopoulos, Yejin Choi, et al · 2024
Risk management in the artificial intelligence act
Jonas Schuett · 2024
Later among the works it cites.
Scaling core earnings measurement with large language models
Matthew Shaffer and Charles CY Wang · 2024
Later among the works it cites.
" do anything now": Characterizing and evaluating in-the-wild jailbreak prompts on large language models
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang · 2024
Later among the works it cites.
Large language model safety: A holistic survey
Dan Shi, Yidi Chen, Yujia Zhao, Huai Su, Jiaming Zhang, Han Sun, Yeyun Shen, Zeyu Chen, Guohua Wang, Wenbo Qian, Jianfeng Gao, Lili Dong, Furu Wang, Haoyu Wang, Hongning Yang, Lifeng Wang, Qisheng Wu, Michael Zeng, Lifu Zhang, Shiyu Chang, Ee-Peng Lim, Kevin Chen-Chuan Chang, Min Zhang, Yi Chang, and Min Yang · 2024
Later among the works it cites.
Quantifying uncertainty in natural language explanations of large language models
Sree Harsha Tanneru, Chirag Agarwal, and Himabindu Lakkaraju · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A framework for measures of risk under uncertainty
Tolulope Fadina, Yang Liu, and Ruodu Wang · 2024
Cited alongside, same era.
Enhancing financial question answering with a multi-agent reflection framework
Sorouralsadat Fatemi and Yuheng Hu · 2024
Cited alongside, same era.
Tao Feng, Chuanyang Jin, Jingyu Liu, Kunlun Zhu, Haoqin Tu, Zirui Cheng, Guanyu Lin, and Jiaxuan You · 2024
Cited alongside, same era.
Fine-tuning large language models for stock return prediction using newsflow
Tian Guo and Emmanuel Hauptmann · 2024
Cited alongside, same era.
Construction of a Japanese financial benchmark for large language models
Masanori Hirano · 2024
Cited alongside, same era.
Metagpt: Meta programming for a multi-agent collaborative framework
Sirui Hong, Mingchen Zhuge, Jonathan Chen, Xiawu Zheng, Yuheng Cheng, Jinlin Wang, Ceyao Zhang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, et al · 2024
Cited alongside, same era.
Multi-modal and multi-agent systems meet rationality: A survey
Bowen Jiang, Yangxinyu Xie, Xiaomeng Wang, Weijie J Su, Camillo Jose Taylor, and Tanwi Mallick · 2024
Cited alongside, same era.
Language models don’t always say what they think: unfaithful explanations in chain-of-thought prompting
Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman · 2024
Later among the works it cites.
Mobile-agent: Autonomous multi-modal mobile device agent with visual perception
Junyang Wang, Haiyang Xu, Jiabo Ye, Ming Yan, Weizhou Shen, Ji Zhang, Fei Huang, and Jitao Sang · 2024
Later among the works it cites.
Efficient adversarial training in llms with continuous attacks
Sophie Xhonneux, Alessandro Sordoni, Stephan Günnemann, Gauthier Gidel, and Leo Schwinn · 2024
Later among the works it cites.
Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms
Miao Xiong, Zhiyuan Hu, Xinyang Lu, YIFEI LI, Jie Fu, Junxian He, and Bryan Hooi · 2024
Later among the works it cites.
An llm can fool itself: A prompt-based adversarial attack
Xilie Xu, Keyi Kong, Ning Liu, Lizhen Cui, Di Wang, Jingfeng Zhang, and Mohan Kankanhalli · 2024
Later among the works it cites.
Yahoo Finance, 2024
Yahoo Finance · 2024
Later among the works it cites.
Hongyang Yang, Boyu Zhang, Neng Wang, Cheng Guo, Xiaoli Zhang, Likun Lin, Junlin Wang, Tianyu Zhou, Mao Guan, Runjia Zhang, and Christina Dan Wang · 2024
Later among the works it cites.
Yangyang Yu, Zhiyuan Yao, Haohang Li, Zhiyang Deng, Yupeng Cao, Zhi Chen, Jordan W. Suchow, Rong Liu, Zhenyu Cui, Zhaozhuo Xu, Denghui Zhang, Koduvayur Subbalakshmi, Guojun Xiong, Yueru He, Jimin Huang, Dong Li, and Qianqian Xie · 2024
Later among the works it cites.
R-judge: Benchmarking safety risk awareness for LLM agents
Tongxin Yuan, Zhiwei He, Lingzhong Dong, Yiming Wang, Ruijie Zhao, Tian Xia, Lizhen Xu, Binglin Zhou, Fangqi Li, Zhuosheng Zhang, Rui Wang, and Gongshen Liu · 2024
Later among the works it cites.
Attacks on third-party apis of large language models
Wanru Zhao, Vidit Khazanchi, Haodi Xing, Xuanli He, Qiongkai Xu, and Nicholas Donald Lane · 2024
Later among the works it cites.
Optimizing llm based retrieval augmented generation pipelines in the financial domain
Yiyun Zhao, Prateek Singh, Hanoz Bhathena, Bernardo Ramos, Aviral Joshi, Swaroop Gadiyaram, and Saket Sharma · 2024
Later among the works it cites.
Engaging with ai: How interface design shapes human-ai collaboration in high-stakes decision-making
Zichen Chen, Yunhao Luo, and Misha Sra · 2025
Closest in time.
Ai agents under threat: A survey of key security challenges and future pathways
Zehang Deng, Yongjian Guo, Changzhou Han, Wanlun Ma, Junwu Xiong, Sheng Wen, and Yang Xiang · 2025
Closest in time.
Agent2Agent (A2A) Protocol Documentation
Google · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al · 2025
Closest in time.
Model context protocol (mcp): Landscape, security threats, and future research directions
Xinyi Hou, Yanjie Zhao, Shenao Wang, and Haoyu Wang · 2025
Closest in time.
Omniact: A dataset and benchmark for enabling multimodal generalist autonomous agents for desktop and web
Raghav Kapoor, Yash Parag Butala, Melisa Russak, Jing Yu Koh, Kiran Kamble, Waseem AlShikh, and Ruslan Salakhutdinov · 2025
Closest in time.
Demystifying domain-adaptive post-training for financial llms
Zixuan Ke, Yifei Ming, Xuan-Phi Nguyen, Caiming Xiong, and Shafiq Joty · 2025
Closest in time.
Safety at scale: A comprehensive survey of large model safety
Xingjun Ma, Zhaicheng Wang, Zehui Zhang, Dan Shi, Lichao Yang, Wenbo Qian, Shihan Chen, Zefeng Du, Xin Zheng, Ruixiang Zhu, Michael Zeng, Jianfeng Gao, Lili Dong, Haoyu Wang, Hongning Yang, Lifeng Wang, Qisheng Wu, Yujia Zhao, Lifu Zhang, Shiyu Chang, Ee-Peng Lim, Kevin Chen-Chuan Chang, Min Zhang, Yi Chang, and Min Yang · 2025
Closest in time.
Shortcutsbench: A large-scale real-world benchmark for API-based agents
Haiyang Shen, Yue Li, Desong Meng, Dongqi Cai, Sheng Qi, Li Zhang, Mengwei Xu, and Yun Ma · 2025
Closest in time.
A comprehensive survey in llm(-agent) full stack safety: Data, training and deployment
Kun Wang, Dan Shi, Yidi Chen, Weikai Li, Wenbo Qian, Zhiyong Xu, Jiahuan Chen, Hongning Yang, Haoyu Wang, Qisheng Wu, Zhipeng Liu, Baoyuan Zhang, Xin Zheng, Ruixiang Zhu, Jinyan Li, Jian Zhao, Michael Zeng, Jianfeng Gao, Lili Dong, Lifeng Wang, Lifu Zhang, Shiyu Chang, Ee-Peng Lim, Kevin Chen-Chuan Chang, Min Zhang, Yi Chang, and Min Yang · 2025
Closest in time.
Tradingagents: Multi-agents llm financial trading framework
Yijia Xiao, Edward Sun, Di Luo, and Wei Wang · 2025
Closest in time.
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al · 2025
Closest in time.