Fetching the paper…
Reading the bibliography…
Embedding models play a crucial role in representing and retrieving information across various NLP applications.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych · 1908
Earlier work this paper cites.
The technique of clear writing, 1952
Robert Gunning · 1952
Earlier work this paper cites.
V-measure: A conditional entropy-based external cluster evaluation measure
Andrew Rosenberg and Julia Hirschberg · 2007
Earlier work this paper cites.
Syntactic dependency distance as sentence complexity measure
Masanori Oya · 2011
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
Good debt or bad debt: Detecting semantic orientations in economic texts
Pekka Malo, Ankur Sinha, Pekka Korhonen, Jyrki Wallenius, and Pyry Takala · 2014
Earlier work this paper cites.
GloVe: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning · 2014
Earlier work this paper cites.
Thuctc: An efficient chinese text classifier, 2016
Maosong Sun, Jingyang Li, Zhipeng Guo, Yu Zhao, Yabin Zheng, Xiance Si, and Zhiyuan Liu · 2016
Earlier work this paper cites.
The BQ corpus: A large-scale domain-specific Chinese corpus for sentence semantic equivalence identification
Jing Chen, Qingcai Chen, Xin Liu, Haijun Yang, Daohe Lu, and Buzhou Tang · 2018
Earlier work this paper cites.
Financial question answering., 2018
FiQA · 2018
Earlier work this paper cites.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Earlier work this paper cites.
Generation of company descriptions using concept-to-text and text-to-text deep models: dataset collection and systems evaluation
Raheel Qader, Khoder Jneid, François Portet, and Cyril Labbé · 2018
Earlier work this paper cites.
Biosentvec: creating sentence embeddings for biomedical texts
Qingyu Chen, Yifan Peng, and Zhiyong Lu · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu · 2019
Earlier work this paper cites.
Biowordvec, improving biomedical word embeddings with subword information and mesh
Yijia Zhang, Qingyu Chen, Zhihao Yang, Hongfei Lin, and Zhiyong Lu · 2019
Earlier work this paper cites.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Earlier work this paper cites.
Finbert: A pretrained language model for financial communications
Yi Yang, Mark Christopher Siy Uy, and Allen Huang · 2020
Earlier work this paper cites.
FinQA: A dataset of numerical reasoning over financial data
Zhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan Routledge, and William Yang Wang · 2021
Earlier work this paper cites.
Multilingual and cross-lingual intent detection from spoken data
Daniela Gerz, Pei-Hao Su, Razvan Kusztos, Avishek Mondal, Michał Lis, Eshan Singhal, Nikola Mrkšić, Tsung-Hsien Wen, and Ivan Vulić · 2021
Earlier work this paper cites.
Mdfend: Multi-domain fake news detection
Qiong Nan, Juan Cao, Yongchun Zhu, Yanyan Wang, and Jintao Li · 2021
Earlier work this paper cites.
Impact of news on the commodity market: Dataset and results
Ankur Sinha and Tanmay Khandait · 2021
Cited alongside, same era.
BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models
Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych · 2021
Cited alongside, same era.
Trade the event: Corporate events detection for news-based event-driven trading
Zhihan Zhou, Liqian Ma, and Han Liu · 2021
Cited alongside, same era.
Tat-qa: A question answering benchmark on a hybrid of tabular and textual content in finance
Fengbin Zhu, Wenqiang Lei, Youcheng Huang, Chao Wang, Shuo Zhang, Jiancheng Lv, Fuli Feng, and Tat-Seng Chua · 2021
Cited alongside, same era.
Ccks2022: Few-shot event extraction for the financial sector, 2022
CCKS · 2022
Cited alongside, same era.
Dakuan Lu, Hengkui Wu, Jiaqing Liang, Yipei Xu, Qianyu He, Yipeng Geng, Mengkun Han, Yingsi Xin, and Yanghua Xiao · 2023
Later among the works it cites.
Trillion dollar words: A new financial dataset, task & market analysis
Agam Shah, Suvan Paturi, and Sudheer Chava · 2023
Later among the works it cites.
Improving text embeddings with large language models
Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei · 2023
Later among the works it cites.
Bloomberggpt: A large language model for finance
Shijie Wu, Ozan Irsoy, Steven Lu, Vadim Dabravolski, Mark Dredze, Sebastian Gehrmann, Prabhanjan Kambadur, David Rosenberg, and Gideon Mann · 2023
Later among the works it cites.
C-pack: Packaged resources to advance general chinese embedding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The financial narrative summarisation shared task (FNS 2022)
Mahmoud El-Haj, Nadhem Zmandar, Paul Rayson, Ahmed AbuRa’ed, Marina Litvak, Nikiforos Pittaras, George Giannakopoulos, Aris Kosmopoulos, Blanca Carbajo-Coronado, and Antonio Moreno-Sandoval · 2022
Cited alongside, same era.
Long text and multi-table summarization: Dataset and method
Shuaiqi Liu, Jiannong Cao, Ruosong Yang, and Zhiyuan Wen · 2022
Cited alongside, same era.
Mteb: Massive text embedding benchmark
Niklas Muennighoff, Nouamane Tazi, Loïc Magne, and Nils Reimers · 2022
Cited alongside, same era.
Ectsum: A new benchmark dataset for bullet point summarization of long earnings call transcripts
Rajdeep Mukherjee, Abhinav Bohra, Akash Banerjee, Soumya Sharma, Manjunath Hegde, Afreen Shaikh, Shivani Shrivastava, Koustuv Dasgupta, Niloy Ganguly, Saptarshi Ghosh, et al · 2022
Cited alongside, same era.
One embedder, any task: Instruction-finetuned text embeddings
Hongjin Su, Weijia Shi, Jungo Kasai, Yizhong Wang, Yushi Hu, Mari Ostendorf, Wen-tau Yih, Noah A Smith, Luke Zettlemoyer, and Tao Yu · 2022
Cited alongside, same era.
Self-instruct: Aligning language models with self-generated instructions
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi · 2022
Cited alongside, same era.
Disc-finllm: A chinese financial large language model based on multiple experts fine-tuning
Wei Chen, Qiushi Wang, Zefei Long, Xianyin Zhang, Zhongtian Lu, Bingxuan Li, Siyuan Wang, Jiarong Xu, Xiang Bai, Xuanjing Huang, and Zhongyu Wei · 2023
Cited alongside, same era.
Shitao Xiao, Zheng Liu, Peitian Zhang, and Niklas Muennighoff · 2023
Later among the works it cites.
Fineval: A chinese financial domain knowledge evaluation benchmark for large language models
Liwen Zhang, Weige Cai, Zhaowei Liu, Zhi Yang, Wei Dai, Yujie Liao, Qianru Qin, Yifei Li, Xingyu Liu, Zhiqiang Liu, et al · 2023
Later among the works it cites.
Biomedlm: A 2.7 b parameter language model trained on biomedical text
Elliot Bolton, Abhinav Venigalla, Michihiro Yasunaga, David Hall, Betty Xiong, Tony Lee, Roxana Daneshjou, Jonathan Frankle, Percy Liang, Michael Carbin, et al · 2024
Closest in time.
Consumer finance complaints, 2024
CFPB · 2024
Closest in time.
Saullm-7b: A pioneering large language model for law
Pierre Colombo, Telmo Pessoa Pires, Malik Boudiaf, Dominic Culver, Rui Melo, Caio Corro, Andre FT Martins, Fabrizio Esposito, Vera Lúcia Raposo, Sofia Morgado, et al · 2024
Closest in time.
Drbenchmark: A large language understanding evaluation benchmark for french biomedical domain
Yanis Labrak, Adrien Bazoge, Oumaima El Khettari, Mickaël Rouvier, Natalia Grabar, Beatrice Daille, Solen Quiniou, Emmanuel Morin, Pierre-Antoine Gourraud, Richard Dufour, et al · 2024
Closest in time.
Nv-embed: Improved techniques for training llms as generalist embedding models
Chankyu Lee, Rajarshi Roy, Mengyao Xu, Jonathan Raiman, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping · 2024
Closest in time.
Alphafin: Benchmarking financial analysis with retrieval-augmented stock-chain framework
Xiang Li, Zhenyu Li, Chen Shi, Yong Xu, Qing Du, Mingkui Tan, Jun Huang, and Wei Lin · 2024
Closest in time.
Bellm: Backward dependency enhanced large language model for sentence embeddings
Xianming Li and Jing Li · 2024
Closest in time.
Beyond surface similarity: Detecting subtle semantic shifts in financial narratives
Jiaxin Liu, Yi Yang, and Kar Yan Tam · 2024
Closest in time.
Sfr-embedding-2: Advanced text embedding with multi-stage training, 2024
Rui Meng, Ye Liu, Shafiq Rayhan Joty, Caiming Xiong, Yingbo Zhou, and Semih Yavuz · 2024
Closest in time.
Repetition improves language model embeddings
Jacob Mitchell Springer, Suhas Kotha, Daniel Fried, Graham Neubig, and Aditi Raghunathan · 2024
Closest in time.
Synthetic-PII-Financial-Documents-North-America: A synthetic dataset for training language models to label and detect pii in domain specific formats, June 2024
Alex Watson, Yev Meyer, Maarten Van Segbroeck, Matthew Grossman, Sami Torbey, Piotr Mlocek, and Johnny Greco · 2024
Closest in time.
Fintruthqa: A benchmark dataset for evaluating the quality of financial information disclosure
Ziyue Xu, Peilin Zhou, Xinyu Shi, Jiageng Wu, Yikang Jiang, Bin Ke, and Jie Yang · 2024
Closest in time.
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, et al · 2024
Closest in time.
Docmath-eval: Evaluating math reasoning capabilities of llms in understanding financial documents
Yilun Zhao, Yitao Long, Hongjun Liu, Ryo Kamoi, Linyong Nan, Lyuhao Chen, Yixin Liu, Xiangru Tang, Rui Zhang, and Arman Cohan · 2024
Closest in time.
SemEval-2017 task 5: Fine-grained sentiment analysis on financial microblogs and news
Keith Cortis, André Freitas, Tobias Daudert, Manuela Huerlimann, Manel Zarrouk, Siegfried Handschuh, and Brian Davis · 2089
Closest in time.