Fetching the paper…
Reading the bibliography…
Embedding models play a crucial role in representing and retrieving information across various NLP applications.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu. 2019 · 1907
Earlier work this paper cites.
Finbert: A pretrained language model for financial communications
Yi Yang, Mark Christopher Siy Uy, and Allen Huang. 2020 · 2006
Earlier work this paper cites.
V-measure: A conditional entropy-based external cluster evaluation measure
Andrew Rosenberg and Julia Hirschberg. 2007 · 2007
Earlier work this paper cites.
Long text and multi-table summarization: Dataset and method
Shuaiqi Liu, Jiannong Cao, Ruosong Yang, and Zhiyuan Wen. 2022 · 2010
Earlier work this paper cites.
Syntactic dependency distance as sentence complexity measure
Masanori Oya. 2011 · 2011
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013 · 2013
Earlier work this paper cites.
Good debt or bad debt: Detecting semantic orientations in economic texts
Pekka Malo, Ankur Sinha, Pekka Korhonen, Jyrki Wallenius, and Pyry Takala. 2014 · 2014
Earlier work this paper cites.
GloVe: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014 · 2014
Earlier work this paper cites.
Thuctc: An efficient chinese text classifier
Maosong Sun, Jingyang Li, Zhipeng Guo, Yu Zhao, Yabin Zheng, Xiance Si, and Zhiyuan Liu. 2016 · 2016
Earlier work this paper cites.
SemEval-2017 task 5: Fine-grained sentiment analysis on financial microblogs and news
Keith Cortis, André Freitas, Tobias Daudert, Manuela Huerlimann, Manel Zarrouk, Siegfried Handschuh, and Brian Davis. 2017 · 2017
Earlier work this paper cites.
The BQ corpus: A large-scale domain-specific Chinese corpus for sentence semantic equivalence identification
Jing Chen, Qingcai Chen, Xin Liu, Haijun Yang, Daohe Lu, and Buzhou Tang. 2018 · 2018
Earlier work this paper cites.
Financial question answering
FiQA. 2018 · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. 2018 · 2018
Earlier work this paper cites.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Earlier work this paper cites.
Biosentvec: creating sentence embeddings for biomedical texts
Qingyu Chen, Yifan Peng, and Zhiyong Lu. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych. 2019 · 2019
Earlier work this paper cites.
Biowordvec, improving biomedical word embeddings with subword information and mesh
Yijia Zhang, Qingyu Chen, Zhihao Yang, Hongfei Lin, and Zhiyong Lu. 2019 · 2019
Earlier work this paper cites.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Earlier work this paper cites.
FinQA: A dataset of numerical reasoning over financial data
Zhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan Routledge, and William Yang Wang. 2021 · 2021
Earlier work this paper cites.
Multilingual and cross-lingual intent detection from spoken data
Daniela Gerz, Pei-Hao Su, Razvan Kusztos, Avishek Mondal, Michał Lis, Eshan Singhal, Nikola Mrkšić, Tsung-Hsien Wen, and Ivan Vulić. 2021a · 2021
Earlier work this paper cites.
Mdfend: Multi-domain fake news detection
Qiong Nan, Juan Cao, Yongchun Zhu, Yanyan Wang, and Jintao Li. 2021 · 2021
Earlier work this paper cites.
Impact of news on the commodity market: Dataset and results
Ankur Sinha and Tanmay Khandait. 2021 · 2021
Cited alongside, same era.
BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models
Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2021 · 2021
Cited alongside, same era.
Trade the event: Corporate events detection for news-based event-driven trading
Zhihan Zhou, Liqian Ma, and Han Liu. 2021 · 2021
Cited alongside, same era.
Tat-qa: A question answering benchmark on a hybrid of tabular and textual content in finance
Fengbin Zhu, Wenqiang Lei, Youcheng Huang, Chao Wang, Shuo Zhang, Jiancheng Lv, Fuli Feng, and Tat-Seng Chua. 2021 · 2021
Cited alongside, same era.
Ccks2022: Few-shot event extraction for the financial sector
CCKS. 2022 · 2022
Cited alongside, same era.
Bloomberggpt: A large language model for finance
Shijie Wu, Ozan Irsoy, Steven Lu, Vadim Dabravolski, Mark Dredze, Sebastian Gehrmann, Prabhanjan Kambadur, David Rosenberg, and Gideon Mann. 2023 · 2023
Later among the works it cites.
C-pack: Packaged resources to advance general chinese embedding
Shitao Xiao, Zheng Liu, Peitian Zhang, and Niklas Muennighoff. 2023 · 2023
Later among the works it cites.
Fineval: A chinese financial domain knowledge evaluation benchmark for large language models
Liwen Zhang, Weige Cai, Zhaowei Liu, Zhi Yang, Wei Dai, Yujie Liao, Qianru Qin, Yifei Li, Xingyu Liu, Zhiqiang Liu, et al. 2023 · 2023
Later among the works it cites.
Greenback bears and fiscal hawks: Finance is a jungle and text embeddings must adapt
Peter Anderson, Mano Vikash Janardhanan, Jason He, Wei Cheng, and Charlie Flanagan. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The financial narrative summarisation shared task (FNS 2022)
Mahmoud El-Haj, Nadhem Zmandar, Paul Rayson, Ahmed AbuRa’ed, Marina Litvak, Nikiforos Pittaras, George Giannakopoulos, Aris Kosmopoulos, Blanca Carbajo-Coronado, and Antonio Moreno-Sandoval. 2022 · 2022
Cited alongside, same era.
Mteb: Massive text embedding benchmark
Niklas Muennighoff, Nouamane Tazi, Loïc Magne, and Nils Reimers. 2022 · 2022
Cited alongside, same era.
Ectsum: A new benchmark dataset for bullet point summarization of long earnings call transcripts
Rajdeep Mukherjee, Abhinav Bohra, Akash Banerjee, Soumya Sharma, Manjunath Hegde, Afreen Shaikh, Shivani Shrivastava, Koustuv Dasgupta, Niloy Ganguly, Saptarshi Ghosh, et al. 2022 · 2022
Cited alongside, same era.
One embedder, any task: Instruction-finetuned text embeddings
Hongjin Su, Weijia Shi, Jungo Kasai, Yizhong Wang, Yushi Hu, Mari Ostendorf, Wen-tau Yih, Noah A Smith, Luke Zettlemoyer, and Tao Yu. 2022 · 2022
Cited alongside, same era.
Disc-finllm: A chinese financial large language model based on multiple experts fine-tuning
Wei Chen, Qiushi Wang, Zefei Long, Xianyin Zhang, Zhongtian Lu, Bingxuan Li, Siyuan Wang, Jiarong Xu, Xiang Bai, Xuanjing Huang, and Zhongyu Wei. 2023 · 2023
Cited alongside, same era.
Lawbench: Benchmarking legal knowledge of large language models
Zhiwei Fei, Xiaoyu Shen, Dawei Zhu, Fengzhe Zhou, Zhuo Han, Songyang Zhang, Kai Chen, Zongwen Shen, and Jidong Ge. 2023 · 2023
Cited alongside, same era.
How close is chatgpt to human experts? comparison corpus, evaluation, and detection
Biyang Guo, Xin Zhang, Ziyuan Wang, Minqi Jiang, Jinran Nie, Yuxuan Ding, Jianwei Yue, and Yupeng Wu. 2023 · 2023
Cited alongside, same era.
Elliot Bolton, Abhinav Venigalla, Michihiro Yasunaga, David Hall, Betty Xiong, Tony Lee, Roxana Daneshjou, Jonathan Frankle, Percy Liang, Michael Carbin, et al. 2024 · 2024
Later among the works it cites.
Consumer finance complaints
CFPB. 2024 · 2024
Later among the works it cites.
Saullm-7b: A pioneering large language model for law
Pierre Colombo, Telmo Pessoa Pires, Malik Boudiaf, Dominic Culver, Rui Melo, Caio Corro, Andre FT Martins, Fabrizio Esposito, Vera Lúcia Raposo, Sofia Morgado, et al. 2024 · 2024
Later among the works it cites.
Scaling synthetic data creation with 1,000,000,000 personas
Tao Ge, Xin Chan, Xiaoyang Wang, Dian Yu, Haitao Mi, and Dong Yu. 2024 · 2024
Later among the works it cites.
Drbenchmark: A large language understanding evaluation benchmark for french biomedical domain
Yanis Labrak, Adrien Bazoge, Oumaima El Khettari, Mickaël Rouvier, Natalia Grabar, Beatrice Daille, Solen Quiniou, Emmanuel Morin, Pierre-Antoine Gourraud, Richard Dufour, et al. 2024 · 2024
Later among the works it cites.
Nv-embed: Improved techniques for training llms as generalist embedding models
Chankyu Lee, Rajarshi Roy, Mengyao Xu, Jonathan Raiman, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping. 2024 · 2024
Later among the works it cites.
Alphafin: Benchmarking financial analysis with retrieval-augmented stock-chain framework
Xiang Li, Zhenyu Li, Chen Shi, Yong Xu, Qing Du, Mingkui Tan, Jun Huang, and Wei Lin. 2024 · 2024
Later among the works it cites.
Beyond surface similarity: Detecting subtle semantic shifts in financial narratives
Jiaxin Liu, Yi Yang, and Kar Yan Tam. 2024a · 2024
Later among the works it cites.
Sfr-embedding-2: Advanced text embedding with multi-stage training
Rui Meng, Ye Liu, Shafiq Rayhan Joty, Caiming Xiong, Yingbo Zhou, and Semih Yavuz. 2024 · 2024
Later among the works it cites.
Openai (august 24 version)
OpenAI. 2024 · 2024
Later among the works it cites.
Repetition improves language model embeddings
Jacob Mitchell Springer, Suhas Kotha, Daniel Fried, Graham Neubig, and Aditi Raghunathan. 2024 · 2024
Later among the works it cites.
Qwen2.5: A party of foundation models
Qwen Team. 2024 · 2024
Later among the works it cites.
Synthetic-PII-Financial-Documents-North-America: A synthetic dataset for training language models to label and detect pii in domain specific formats
Alex Watson, Yev Meyer, Maarten Van Segbroeck, Matthew Grossman, Sami Torbey, Piotr Mlocek, and Johnny Greco. 2024 · 2024
Later among the works it cites.
Fintruthqa: A benchmark dataset for evaluating the quality of financial information disclosure
Ziyue Xu, Peilin Zhou, Xinyu Shi, Jiageng Wu, Yikang Jiang, Bin Ke, and Jie Yang. 2024 · 2024
Later among the works it cites.
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, et al. 2024 · 2024
Later among the works it cites.
Jie Zhu, Junhui Li, Yalong Wen, and Lifan Guo. 2024 · 2024
Later among the works it cites.
RE-FIN: Retrieval-based enrichment for financial data
Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica, and Filippo Pallucchini. 2025 · 2025
Closest in time.
Voyageai (jan 25 version)
VoyageAI. 2025 · 2025
Closest in time.