Fetching the paper…
Reading the bibliography…
Text embeddings from PLM-based models enable a wide range of applications, yet their performance often degrades on longer texts.
Www’18 open challenge: financial opinion mining and question answering
Macedo Maia, Siegfried Handschuh, André Freitas, Brian Davis, Ross McDermott, Manel Zarrouk, and Alexandra Balahur. 2018 · 1942
Earlier work this paper cites.
The sum of log-normal probability distributions in scatter transmission systems
Lawrence Fenton. 1960 · 1960
Earlier work this paper cites.
On the exact variance of products
Leo A Goodman. 1960 · 1960
Earlier work this paper cites.
Matrix analysis and applied linear algebra , volume 71
Carl D Meyer. 2000 · 2000
Earlier work this paper cites.
Efficient intent detection with dual sentence encoders
Iñigo Casanueva, Tadas Temčinas, Daniela Gerz, Matthew Henderson, and Ivan Vulić. 2020 · 2003
Earlier work this paper cites.
Specter: Document-level representation learning using citation-informed transformers
Arman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey, and Daniel S Weld. 2020a · 2004
Earlier work this paper cites.
Fact or fiction: Verifying scientific claims
David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi. 2020 · 2004
Earlier work this paper cites.
A note on over-smoothing for graph neural networks
Chen Cai and Yusu Wang. 2020 · 2006
Earlier work this paper cites.
V-measure: A conditional entropy-based external cluster evaluation measure
Andrew Rosenberg and Julia Hirschberg. 2007 · 2007
Earlier work this paper cites.
Top2vec: Distributed representations of topics
Dimo Angelov. 2020 · 2008
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew Maas, Raymond E Daly, Peter T Pham, Dan Huang, Andrew Y Ng, and Christopher Potts. 2011 · 2011
Earlier work this paper cites.
A survey of text clustering algorithms
Charu C Aggarwal and ChengXiang Zhai. 2012 · 2012
Earlier work this paper cites.
Semeval-2012 task 6: A pilot on semantic textual similarity
Eneko Agirre, Daniel Cer, Mona Diab, and Aitor Gonzalez-Agirre. 2012 · 2012
Earlier work this paper cites.
* sem 2013 shared task: Semantic textual similarity
Eneko Agirre, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, and Weiwei Guo. 2013 · 2013
Earlier work this paper cites.
Generating a word-emotion lexicon from# emotional tweets
Anil Bandhakavi, Nirmalie Wiratunga, P Deepak, and Stewart Massie. 2014 · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning. 2014 · 2014
Earlier work this paper cites.
Rtm-dcu: Predicting semantic similarity with referential translation machines
Ergun Biçici. 2015 · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015 · 2015
Earlier work this paper cites.
A full-text learning to rank dataset for medical information retrieval
Vera Boteva, Demian Gholipour, Artem Sokolov, and Stefan Riezler. 2016 · 2016
Earlier work this paper cites.
SemEval-2016 task 4: Sentiment analysis in Twitter
Preslav Nakov, Alan Ritter, Sara Rosenthal, Fabrizio Sebastiani, and Veselin Stoyanov. 2016 · 2016
Earlier work this paper cites.
Task-oriented intrinsic evaluation of semantic textual similarity
Nils Reimers, Philip Beyer, and Iryna Gurevych. 2016 · 2016
Earlier work this paper cites.
Biosses: a semantic sentence similarity estimation system for the biomedical domain
Gizem Soğancıoğlu, Hakime Öztürk, and Arzucan Özgür. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
A Waswani, N Shazeer, N Parmar, J Uszkoreit, L Jones, A Gomez, L Kaiser, and I Polosukhin. 2017 · 2017
Earlier work this paper cites.
The narrativeqa reading comprehension challenge
Tomáš Kočiskỳ, Jonathan Schwarz, Phil Blunsom, Chris Dyer, Karl Moritz Hermann, Gábor Melis, and Edward Grefenstette. 2018 · 2018
Cited alongside, same era.
Linkso: a dataset for learning to retrieve similar question answer pairs on software development forums
Xueqing Liu, Chi Wang, Yue Leng, and ChengXiang Zhai. 2018 · 2018
Cited alongside, same era.
Carer: Contextualized affect representations for emotion recognition
Elvis Saravia, Hsien-Chi Toby Liu, Yen-Hao Huang, Junlin Wu, and Yi-Shin Chen. 2018 · 2018
Cited alongside, same era.
How contextual are contextualized word representations? comparing the geometry of bert, elmo, and gpt-2 embeddings
Kawin Ethayarajh. 2019 · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Overcoming a theoretical limitation of self-attention
David Chiang and Peter Cholak. 2022 · 2022
Later among the works it cites.
Jack FitzGerald, Christopher Hench, Charith Peris, Scott Mackie, Kay Rottmann, Ana Sanchez, Aaron Nash, Liam Urbach, Vishesh Kakarala, Richa Singh, et al. 2022 · 2022
Later among the works it cites.
Large dual encoders are generalizable retrievers
Jianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernandez Abrego, Ji Ma, Vincent Zhao, Yi Luan, Keith Hall, Ming-Wei Chang, et al. 2022 · 2022
Later among the works it cites.
Retrieval-augmented generation for large language models: A survey
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, and Haofen Wang. 2023 · 2023
Later among the works it cites.
Jina embeddings 2: 8192-token general-purpose text embeddings for long documents
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Cited alongside, same era.
Overview of touché 2020: argument retrieval
Alexander Bondarenko, Maik Fröbe, Meriem Beloucif, Lukas Gienapp, Yamen Ajjour, Alexander Panchenko, Chris Biemann, Benno Stein, Henning Wachsmuth, Martin Potthast, et al. 2020 · 2020
Cited alongside, same era.
Evaluation of sentence representations in Polish
Slawomir Dadas, Michał Perełkiewicz, and Rafał Poświata. 2020 · 2020
Cited alongside, same era.
Constructing a multi-hop QA dataset for comprehensive evaluation of reasoning steps
Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa. 2020 · 2020
Cited alongside, same era.
On the sentence embeddings from pre-trained language models
Bohan Li, Hao Zhou, Junxian He, Mingxuan Wang, Yiming Yang, and Lei Li. 2020 · 2020
Cited alongside, same era.
Graph neural networks exponentially lose expressive power for node classification
Kenta Oono and Taiji Suzuki. 2020 · 2020
Cited alongside, same era.
Understanding contrastive representation learning through alignment and uniformity on the hypersphere
Tongzhou Wang and Phillip Isola. 2020 · 2020
Cited alongside, same era.
Michael Günther, Jackmin Ong, Isabelle Mohr, Alaeddine Abdessalem, Tanguy Abel, Mohammad Kalim Akram, Susana Guzman, Georgios Mastrapas, Saba Sturua, Bo Wang, et al. 2023 · 2023
Later among the works it cites.
Evaluation of chatgpt as a question answering system for answering complex questions
Yiming Tan, Dehai Min, Yu Li, Wenbo Li, Nan Hu, Yongrui Chen, and Guilin Qi. 2023 · 2023
Later among the works it cites.
C-pack: Packaged resources to advance general chinese embedding
Shitao Xiao, Zheng Liu, Peitian Zhang, Niklas Muennighoff, Defu Lian, and Jian-Yun Nie. 2023 · 2023
Later among the works it cites.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2024 · 2024
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024 · 2024
Closest in time.
A survey on rag meeting llms: Towards retrieval-augmented large language models
Wenqi Fan, Yujuan Ding, Liangbo Ning, Shijie Wang, Hengyun Li, Dawei Yin, Tat-Seng Chua, and Qing Li. 2024 · 2024
Closest in time.
Longrag: Enhancing retrieval-augmented generation with long-context llms
Ziyan Jiang, Xueguang Ma, and Wenhu Chen. 2024 · 2024
Closest in time.
Nv-embed: Improved techniques for training llms as generalist embedding models
Chankyu Lee, Rajarshi Roy, Mengyao Xu, Jonathan Raiman, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping. 2024 · 2024
Closest in time.
Generative representational instruction tuning
Niklas Muennighoff, SU Hongjin, Liang Wang, Nan Yang, Furu Wei, Tao Yu, Amanpreet Singh, and Douwe Kiela · 2024
Closest in time.
Linear log-normal attention with unbiased concentration
Yury Nahshan, Joseph Kampeas, and Emir Haleva. 2024 · 2024
Closest in time.
Nomic embed: Training a reproducible long context text embedder
Zach Nussbaum, John X Morris, Brandon Duderstadt, and Andriy Mulyar. 2024 · 2024
Closest in time.
Benchmarking and building long-context retrieval models with loco and m2-bert
Jon Saad-Falcon, Daniel Y Fu, Simran Arora, Neel Guha, and Christopher Re. 2024 · 2024
Closest in time.
Gistembed: Guided in-sample selection of training negatives for text embedding fine-tuning
Aivin V Solatorio. 2024 · 2024
Closest in time.
Improving text embeddings with large language models
Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei. 2024 · 2024
Closest in time.
Search-in-the-chain: Interactively enhancing large language models with search for knowledge-intensive tasks
Shicheng Xu, Liang Pang, Huawei Shen, Xueqi Cheng, and Tat-Seng Chua. 2024 · 2024
Closest in time.
Dense text retrieval based on pretrained language models: A survey
Wayne Xin Zhao, Jing Liu, Ruiyang Ren, and Ji-Rong Wen. 2024 · 2024
Closest in time.
Longembed: Extending embedding models for long context retrieval
Dawei Zhu, Liang Wang, Nan Yang, Yifan Song, Wenhao Wu, Furu Wei, and Sujian Li. 2024 · 2024
Closest in time.
Mteb: Massive text embedding benchmark
Niklas Muennighoff, Nouamane Tazi, Loic Magne, and Nils Reimers. 2023 · 2037
Closest in time.