Fetching the paper…
Reading the bibliography…
Real-world financial analysis involves information across multiple languages and modalities, from reports and news to scanned filings and meeting recordings.
Comparison of the predicted and observed secondary structure of t4 phage lysozyme
Brian W Matthews. 1975 · 1975
Earlier work this paper cites.
Accuracy measures: theoretical and practical concerns
Spyros Makridakis. 1993 · 1993
Earlier work this paper cites.
The sharpe ratio
William F Sharpe. 1998 · 1998
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
A systematic analysis of performance measures for classification tasks
Marina Sokolova and Guy Lapalme. 2009 · 2009
Earlier work this paper cites.
Layoutlmv2: Multi-modal pre-training for visually-rich document understanding
Yiheng Xu, Tengchao Xu, Lei Cui, Guoxin Wang, Shaohan Huang, Furu Wei, and Ming Zhou. 2021 · 2012
Earlier work this paper cites.
Handbook of inter-rater reliability: The definitive guide to measuring the extent of agreement among raters
Kilem L Gwet. 2014 · 2014
Earlier work this paper cites.
Complementarity, F-score, and NLP evaluation
Leon Derczynski. 2016 · 2016
Earlier work this paper cites.
Semeval-2017 task 5: Fine-grained sentiment analysis on financial microblogs and news
Keith Cortis, André Freitas, Tobias Daudert, Manuela Huerlimann, Manel Zarrouk, Siegfried Handschuh, and Brian Davis. 2017 · 2017
Earlier work this paper cites.
chabsa: Aspect based sentiment analysis dataset in japanese
Takahiro Kubo, Hiroki Nakayama, and Junya Kamura. 2018 · 2018
Earlier work this paper cites.
SAMSum corpus: A human-annotated dialogue dataset for abstractive summarization
Bogdan Gliwa, Iwona Mochol, Maciej Biesek, and Aleksander Wawer. 2019 · 2019
Earlier work this paper cites.
Funsd: A dataset for form understanding in noisy scanned documents
Guillaume Jaume, Hazim Kemal Ekenel, and Jean-Philippe Thiran. 2019 · 2019
Earlier work this paper cites.
Cord: A consolidated receipt dataset for post-ocr parsing
Seungryong Park, Seung Shin, Byoungjip Lee, Sangdoo Lee, Junhwa Lee, and In So Kweon. 2019 · 2019
Earlier work this paper cites.
What you say and how you say it matters: Predicting stock volatility using verbal and vocal cues
Yu Qin and Yi Yang. 2019 · 2019
Earlier work this paper cites.
The financial document causality detection shared task (FinCausal 2020)
Dominique Mariko, Hanna Abi-Akl, Estelle Labidurie, Stephane Durfort, Hugues De Mazancourt, and Mahmoud El-Haj. 2020 · 2020
Earlier work this paper cites.
Finqa: A dataset of numerical reasoning over financial data
Zhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan R Routledge, and 1 others. 2021 · 2021
Earlier work this paper cites.
Impact of news on the commodity market: Dataset and results
Ankur Sinha and Tanmay Khandait. 2021 · 2021
Earlier work this paper cites.
Tat-qa: A question answering benchmark on a hybrid of tabular and textual content in finance
Fengbin Zhu, Wenqiang Lei, Youcheng Huang, Chao Wang, Shuo Zhang, Jiancheng Lv, Fuli Feng, and Tat-Seng Chua. 2021 · 2021
Earlier work this paper cites.
Convfinqa: Exploring the chain of numerical reasoning in conversational finance question answering
Zhiyu Chen, Shiyang Li, Charese Smiley, Zhiqiang Ma, Sameena Shah, and William Yang Wang. 2022 · 2022
Earlier work this paper cites.
ECTSum: A new benchmark dataset for bullet point summarization of long earnings call transcripts
Rajdeep Mukherjee, Abhinav Bohra, Akash Banerjee, Soumya Sharma, Manjunath Hegde, Afreen Shaikh, Shivani Shrivastava, Koustuv Dasgupta, Niloy Ganguly, Saptarshi Ghosh, and 1 others. 2022 · 2022
Earlier work this paper cites.
Finred: A dataset for relation extraction in financial domain
Soumya Sharma, Tapas Nayak, Arusarka Bose, Ajay Kumar Meena, Koustuv Dasgupta, Niloy Ganguly, and Pawan Goyal. 2022 · 2022
Earlier work this paper cites.
Accurate stock movement prediction with self-supervised learning from sparse noisy tweets
Yejun Soun, Jaemin Yoo, Minyong Cho, Jihyeong Jeon, and U Kang. 2022 · 2022
Earlier work this paper cites.
Cheng-Yu Hsieh, Chun-Liang Li, Chih-Kuan Yeh, Hootan Nakhost, Yasuhisa Fujii, Alexander Ratner, Ranjay Krishna, Chen-Yu Lee, and Tomas Pfister. 2023 · 2023
Earlier work this paper cites.
Financebench: A new benchmark for financial question answering
Pranab Islam, Anand Kannappan, Douwe Kiela, Rebecca Qian, Nino Scherrer, and Bertie Vidgen. 2023 · 2023
Earlier work this paper cites.
Multifin: A dataset for multilingual financial nlp
Rasmus Jørgensen, Oliver Brandt, Mareike Hartmann, Xiang Dai, Christian Igel, and Desmond Elliott. 2023a · 2023
Earlier work this paper cites.
MultiFin: A dataset for multilingual financial NLP
Rasmus Jørgensen, Oliver Brandt, Mareike Hartmann, Xiang Dai, Christian Igel, and Desmond Elliott. 2023b · 2023
Cited alongside, same era.
Bizbench: A quantitative reasoning benchmark for business and finance
Rik Koncel-Kedziorski, Michael Krumdick, Viet Lai, Varshini Reddy, Charles Lovering, and Chris Tanner. 2023 · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. 2023 · 2023
Cited alongside, same era.
Cfbenchmark: Chinese financial assistant benchmark for large language model
Yang Lei, Jiangtong Li, Dawei Cheng, Zhijun Ding, and Changjun Jiang. 2023 · 2023
Cited alongside, same era.
Evaluation of transformer models for financial targeted sentiment analysis in spanish
Ronghao Pan, José Antonio García-Díaz, Francisco Garcia-Sanchez, and Rafael Valencia-García. 2023 · 2023
Deepseek-vl: Towards real-world vision-language understanding
Haoyu Lu, Wen Liu, Bo Zhang, Bingxuan Wang, Kai Dong, Bo Liu, Jingxiang Sun, Tongzheng Ren, Zhuoshu Li, Yaofeng Sun, Chengqi Deng, Hanwei Xu, Zhenda Xie, and Chong Ruan. 2024 · 2024
Later among the works it cites.
CFinBench: A comprehensive chinese financial benchmark for large language models
Ying Nie, Binwei Yan, Tianyu Guo, Hao Liu, Haoyu Wang, Wei He, Binfan Zheng, Weihao Wang, Qiang Li, Weijian Sun, Yunhe Wang, and Dacheng Tao. 2024 · 2024
Later among the works it cites.
Evaluation of generative ai q&a chatbot chained to optical character recognition models for financial documents
Yu Qiu, Venkata C Duvvuri, Pratibha Yadavalli, and Neal Prasad. 2024 · 2024
Later among the works it cites.
Betterbench: Assessing ai benchmarks, uncovering issues, and establishing best practices
Anka Reuel, Amelia Hardy, Chandler Smith, Max Lamparth, Malcolm Hardy, and Mykel J Kochenderfer. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Robust speech recognition via large-scale weak supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine Mcleavey, and Ilya Sutskever. 2023 · 2023
Cited alongside, same era.
AudioPALM: A large language model that can speak and listen
Paul K Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen, Ankur Bapna, Zalán Borsos, Félix de Chaumont Quitry, Peter Chen, Dalia El Badawy, Wei Han, Eugene Kharitonov, and 1 others. 2023 · 2023
Cited alongside, same era.
Finer: Financial named entity recognition dataset and weak-supervision model
Agam Shah, Ruchit Vithani, Abhinav Gullapalli, and Sudheer Chava. 2023 · 2023
Cited alongside, same era.
BloombergGPT: A large language model for finance
Shijie Wu, Ozan Irsoy, Steven Lu, Vadim Dabravolski, Mark Dredze, Sebastian Gehrmann, Prabhanjan Kambadur, David Rosenberg, and Gideon Mann. 2023 · 2023
Cited alongside, same era.
PIXIU: A large language model, instruction data and evaluation benchmark for finance
Qianqian Xie, Weiguang Han, Xiao Zhang, Yanzhao Lai, Min Peng, Alejandro Lopez-Lira, and Jimin Huang. 2023 · 2023
Cited alongside, same era.
The financial narrative summarisation shared task (fns 2023)
Elias Zavitsanos, Aris Kosmopoulos, George Giannakopoulos, Marina Litvak, Blanca Carbajo-Coronado, Antonio Moreno-Sandoval, and Mo El-Haj. 2023 · 2023
Cited alongside, same era.
Xuanyuan 2.0: A large chinese financial chat model with hundreds of billions parameters
Xuanyu Zhang and Qing Yang. 2023 · 2023
Cited alongside, same era.
ACLSum: A new dataset for aspect-based summarization of scientific publications
Sotaro Takeshita, Tommaso Green, Ines Reinig, Kai Eckert, and Simone Ponzetto. 2024 · 2024
Later among the works it cites.
Salmonn: Towards generic hearing abilities for large language models
Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun Ma, and Chao Zhang. 2024 · 2024
Later among the works it cites.
Mmlu-pro: A more robust and challenging multi-task language understanding benchmark
Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, and 1 others. 2024 · 2024
Later among the works it cites.
Yangyang Yu, Zhiyuan Yao, Haohang Li, Zhiyang Deng, Yupeng Cao, Zhi Chen, Jordan W. Suchow, Rong Liu, Zhenyu Cui, Zhaozhuo Xu, Denghui Zhang, Koduvayur Subbalakshmi, Guojun Xiong, Yueru He, Jimin Huang, Dong Li, and Qianqian Xie. 2024 · 2024
Later among the works it cites.
The llama 4 herd: The beginning of a new era of natively multimodal ai innovation
Meta AI. 2025 · 2025
Closest in time.
Benchmarking large language models on answering and explaining challenging medical questions
Hanjie Chen, Zhouxiang Fang, Yash Singla, and Mark Dredze. 2025a · 2025
Closest in time.
Deepseek-r1-distill-qwen-32b-japanese
Ryosuke Ishigami. 2025 · 2025
Closest in time.
Resurrecting saturated llm benchmarks with adversarial encoding
Igor Ivanov and Dmitrii Volkov. 2025 · 2025
Closest in time.
FinMME: Benchmark dataset for financial multi-modal reasoning evaluation
Junyu Luo, Zhizhuo Kou, Liming Yang, Xiao Luo, Jinsheng Huang, Zhiping Xiao, Jingshu Peng, Chengzhong Liu, Jiaming Ji, Xuanzhe Liu, Sirui Han, Ming Zhang, and Yike Guo. 2025 · 2025
Closest in time.
Dolfin–document-level financial test set for machine translation
Mariam Nakhlé, Marco Dinarelli, Raheel Qader, Emmanuelle Esperança-Rodier, and Hervé Blanchon. 2025 · 2025
Closest in time.
Dolfin – document-level financial test set for machine translation
Mariam Nakhlé, Marco Dinarelli, Raheel Qader, Emmanuelle Esperança-Rodier, and Hervé Blanchon. 2025 · 2025
Closest in time.
nvingest: High-performance multimodal table extraction in the wild
NVIDIA Research. 2024 · 2025
Closest in time.
Openai o3-mini: Pushing the frontier of cost-effective reasoning
OpenAI. 2025 · 2025
Closest in time.
Plutus: Benchmarking large language models in low-resource greek finance
Xueqing Peng, Triantafillos Papadopoulos, Efstathia Soufleri, Polydoros Giannouris, Ruoyu Xiang, Yan Wang, Lingfei Qian, Jimin Huang, Qianqian Xie, and Sophia Ananiadou. 2025 · 2025
Closest in time.
Du xiaoman-xuanyuan
Duxiaoman DI Team. 2024 · 2025
Closest in time.
Gemma Team. 2025 · 2025
Closest in time.
Flag-trader: Fusion llm-agent with gradient-based reinforcement learning for financial trading
Guojun Xiong, Zhiyang Deng, Keyi Wang, Yupeng Cao, Haohang Li, Yangyang Yu, Xueqing Peng, Mingquan Lin, Kaleb E Smith, Xiao-Yang Liu, Jimin Huang, Sophia Ananiadou, and Qianqian Xie. 2025 · 2025
Closest in time.
Jin Xu, Zhifang Guo, Jinzheng He, Hangrui Hu, Ting He, Shuai Bai, Keqin Chen, Jialin Wang, Yang Fan, Kai Dang, Bin Zhang, Xiong Wang, Yunfei Chu, and Junyang Lin. 2025 · 2025
Closest in time.
An Yang, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoyan Huang, Jiandong Jiang, Jianhong Tu, Jianwei Zhang, Jingren Zhou, Junyang Lin, Kai Dang, Kexin Yang, Le Yu, Mei Li, Minmin Sun, Qin Zhu, Rui Men, Tao He, and 9 others. 2025 · 2025
Closest in time.
Mmvu: Measuring expert-level multi-discipline video understanding
Yilun Zhao, Lujing Xie, Haowei Zhang, Guo Gan, Yitao Long, Zhiyuan Hu, Tongyan Hu, Weiyuan Chen, Chuhan Li, Junyang Song, Zhijian Xu, Chengye Wang, Weifeng Pan, Ziyao Shangguan, Xiangru Tang, Zhenwen Liang, Yixin Liu, Chen Zhao, and Arman Cohan. 2025 · 2025
Closest in time.
An empirical analysis of word error rate and keyword error rate
Youngja Park, Siddharth Patwardhan, Karthik Visweswariah, and Stephen C Gates. 2008 · 2073
Closest in time.