Fetching the paper…
Reading the bibliography…
Text embeddings are commonly evaluated on a small set of datasets from a single task not covering their possible applications to other tasks.
Scibert: A pretrained language model for scientific text
Iz Beltagy, Kyle Lo, and Arman Cohan. 2019 · 1903
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych. 2019 · 1908
Earlier work this paper cites.
Lxmert: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal. 2019 · 1908
Earlier work this paper cites.
Codesearchnet challenge: Evaluating the state of semantic code search
Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt. 2019 · 1909
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Specter: Document-level representation learning using citation-informed transformers
Arman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey, and Daniel S Weld. 2020a · 2004
Earlier work this paper cites.
Language-agnostic bert sentence embedding
Fangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan, and Wei Wang. 2020 · 2007
Earlier work this paper cites.
V-measure: A conditional entropy-based external cluster evaluation measure
Andrew Rosenberg and Julia Hirschberg. 2007 · 2007
Earlier work this paper cites.
Top2vec: Distributed representations of topics
Dimo Angelov. 2020 · 2008
Earlier work this paper cites.
Xor qa: Cross-lingual open-retrieval question answering
Akari Asai, Jungo Kasai, Jonathan H Clark, Kenton Lee, Eunsol Choi, and Hannaneh Hajishirzi. 2020 · 2010
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011 · 2011
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. 2011 · 2011
Earlier work this paper cites.
A survey of text clustering algorithms
Charu C Aggarwal and ChengXiang Zhai. 2012 · 2012
Earlier work this paper cites.
Semeval-2012 task 6: A pilot on semantic textual similarity
Eneko Agirre, Daniel Cer, Mona Diab, and Aitor Gonzalez-Agirre. 2012 · 2012
Earlier work this paper cites.
Vilio: State-of-the-art visio-linguistic models applied to hateful memes
Niklas Muennighoff. 2020 · 2012
Earlier work this paper cites.
* sem 2013 shared task: Semantic textual similarity
Eneko Agirre, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, and Weiwei Guo. 2013 · 2013
Earlier work this paper cites.
Hidden factors and hidden topics: Understanding rating dimensions with review text
Julian McAuley and Jure Leskovec. 2013 · 2013
Earlier work this paper cites.
Semeval-2014 task 10: Multilingual semantic textual similarity
Eneko Agirre, Carmen Banea, Claire Cardie, Daniel M Cer, Mona T Diab, Aitor Gonzalez-Agirre, Weiwei Guo, Rada Mihalcea, German Rigau, and Janyce Wiebe. 2014 · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning. 2014 · 2014
Earlier work this paper cites.
Semeval-2015 task 2: Semantic textual similarity, english, spanish and pilot on interpretability
Eneko Agirre, Carmen Banea, Claire Cardie, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, Weiwei Guo, Inigo Lopez-Gazpio, Montse Maritxalar, Rada Mihalcea, et al. 2015 · 2015
Earlier work this paper cites.
Semeval-2015 task 1: Paraphrase and semantic similarity in twitter (pit)
Wei Xu, Chris Callison-Burch, and William B Dolan. 2015 · 2015
Earlier work this paper cites.
Semeval-2016 task 1: Semantic textual similarity, monolingual and cross-lingual evaluation
Eneko Agirre, Carmen Banea, Daniel Cer, Mona Diab, Aitor Gonzalez Agirre, Rada Mihalcea, German Rigau Claramunt, and Janyce Wiebe. 2016 · 2016
Earlier work this paper cites.
Dependency based embeddings for sentence classification tasks
Alexandros Komninos and Suresh Manandhar. 2016 · 2016
Earlier work this paper cites.
Task-oriented intrinsic evaluation of semantic textual similarity
Nils Reimers, Philip Beyer, and Iryna Gurevych. 2016 · 2016
Earlier work this paper cites.
Towards preparation of the second bucc shared task: Detecting parallel sentences in comparable corpora
Pierre Zweigenbaum, Serge Sharoff, and Reinhard Rapp. 2016 · 2016
Earlier work this paper cites.
Supervised learning of universal sentence representations from natural language inference data
Alexis Conneau, Douwe Kiela, Holger Schwenk, Loic Barrault, and Antoine Bordes. 2017 · 2017
Cited alongside, same era.
A continuously growing dataset of sentential paraphrases
Wuwei Lan, Siyu Qiu, Hua He, and Wei Xu. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Overview of the second bucc shared task: Spotting parallel sentences in comparable corpora
Pierre Zweigenbaum, Serge Sharoff, and Reinhard Rapp. 2017 · 2017
Cited alongside, same era.
Cross-modal retrieval in the cooking context: Learning semantic text-image embeddings
Micael Carvalho, Rémi Cadène, David Picard, Laure Soulier, Nicolas Thome, and Matthieu Cord. 2018 · 2018
Cited alongside, same era.
GPT-NeoX: Large scale autoregressive language modeling in pytorch
Alex Andonian, Quentin Anthony, Stella Biderman, Sid Black, Preetham Gali, Leo Gao, Eric Hallahan, Josh Levy-Kramer, Connor Leahy, Lucas Nestler, Kip Parker, Michael Pieler, Shivanshu Purohit, Tri Songz, Phil Wang, and Samuel Weinbach. 2021 · 2021
Later among the works it cites.
Unsupervised corpus aware language model pre-training for dense passage retrieval
Luyu Gao and Jamie Callan. 2021 · 2021
Later among the works it cites.
Tweac: Transformer with extendable qa agent classifiers
Gregor Geigle, Nils Reimers, Andreas Rücklé, and Iryna Gurevych. 2021 · 2021
Later among the works it cites.
Towards unsupervised dense information retrieval with contrastive learning
Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. 2021 · 2021
Later among the works it cites.
Datasets: A community library for natural language processing
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alexis Conneau and Douwe Kiela. 2018 · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
Linkso: a dataset for learning to retrieve similar question answer pairs on software development forums
Xueqing Liu, Chi Wang, Yue Leng, and ChengXiang Zhai. 2018 · 2018
Cited alongside, same era.
CARER: Contextualized affect representations for emotion recognition
Elvis Saravia, Hsien-Chi Toby Liu, Yen-Hao Huang, Junlin Wu, and Yi-Shin Chen. 2018 · 2018
Cited alongside, same era.
Adversarial domain adaptation for duplicate question detection
Darsh Shah, Tao Lei, Alessandro Moschitti, Salvatore Romeo, and Preslav Nakov. 2018 · 2018
Cited alongside, same era.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2018 · 2018
Cited alongside, same era.
Overview of the third bucc shared task: Spotting parallel sentences in comparable corpora
Pierre Zweigenbaum, Serge Sharoff, and Reinhard Rapp. 2018 · 2018
Cited alongside, same era.
Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, et al. 2021 · 2021
Later among the works it cites.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. 2021 · 2021
Later among the works it cites.
I wish i would have loved this one, but i didn’t – a multilingual dataset for counterfactual detection in product reviews
James O’Neill, Polina Rozenshtein, Ryuichi Kiryo, Motoko Kubota, and Danushka Bollegala. 2021 · 2021
Later among the works it cites.
Beir: A heterogenous benchmark for zero-shot evaluation of information retrieval models
Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2021 · 2021
Later among the works it cites.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Ben Wang and Aran Komatsuzaki. 2021 · 2021
Later among the works it cites.
Kexin Wang, Nils Reimers, and Iryna Gurevych. 2021 · 2021
Later among the works it cites.
Bing delivers more contextualized search using quantized transformer inference on nvidia gpus in azure
Jeffrey Zhu, Mingqin Li, Jason Li, and Cassandra Oduola. 2021 · 2021
Later among the works it cites.
Improving language models by retrieving from trillions of tokens
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al. 2022 · 2022
Closest in time.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022 · 2022
Closest in time.
Massive: A 1m-example multilingual natural language understanding dataset with 51 typologically-diverse languages
Jack FitzGerald, Christopher Hench, Charith Peris, Scott Mackie, Kay Rottmann, Ana Sanchez, Aaron Nash, Liam Urbach, Vishesh Kakarala, Richa Singh, Swetha Ranganath, Laurie Crist, Misha Britan, Wouter Leeuwis, Gokhan Tur, and Prem Natarajan. 2022 · 2022
Closest in time.
Bitext mining using distilled sentence representations for low-resource languages
Kevin Heffernan, Onur Çelebi, and Holger Schwenk. 2022 · 2022
Closest in time.
Sgpt: Gpt sentence embeddings for semantic search
Niklas Muennighoff. 2022 · 2022
Closest in time.
Crosslingual generalization through multitask finetuning
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng-Xin Yong, Hailey Schoelkopf, et al. 2022 · 2022
Closest in time.
Text and code embeddings by contrastive pre-training
Arvind Neelakantan, Tao Xu, Raul Puri, Alec Radford, Jesse Michael Han, Jerry Tworek, Qiming Yuan, Nikolas Tezak, Jong Wook Kim, Chris Hallacy, et al. 2022 · 2022
Closest in time.
Photorealistic text-to-image diffusion models with deep language understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S Sara Mahdavi, Rapha Gontijo Lopes, et al. 2022 · 2022
Closest in time.
Bloom: A 176b-parameter open-access multilingual language model
Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2022 · 2022
Closest in time.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al. 2022 · 2022
Closest in time.
M-vader: A model for diffusion with multimodal context
Samuel Weinbach, Marco Bellagente, Constantin Eichenberg, Andrew Dai, Robert Baldock, Souradeep Nanda, Björn Deiseroth, Koen Oostermeijer, Hannah Teufel, and Andres Felipe Cruz-Salinas. 2022 · 2022
Closest in time.
Making a miracl: Multilingual information retrieval across a continuum of languages
Xinyu Zhang, Nandan Thakur, Odunayo Ogundepo, Ehsan Kamalloo, David Alfonso-Hermelo, Xiaoguang Li, Qun Liu, Mehdi Rezagholizadeh, and Jimmy Lin. 2022 · 2022
Closest in time.
Santacoder: don’t reach for the stars!
Loubna Ben Allal, Raymond Li, Denis Kocetkov, Chenghao Mou, Christopher Akiki, Carlos Munoz Ferrandis, Niklas Muennighoff, Mayank Mishra, Alex Gu, Manan Dey, et al. 2023 · 2023
Closest in time.