Fetching the paper…
Reading the bibliography…
Exploring and quantifying semantic relatedness is central to representing language and holds significant implications across various NLP tasks.
The theory of the estimation of test reliability
G Frederic Kuder and Marion W Richardson. 1937 · 1937
Earlier work this paper cites.
Coefficient alpha and the internal structure of tests
Lee J Cronbach. 1951 · 1951
Earlier work this paper cites.
Diglossia
Charles A Ferguson. 1959 · 1959
Earlier work this paper cites.
Cohesion in English
Ruqaiya Hasan and Michael AK Halliday. 1976 · 1976
Earlier work this paper cites.
Best-worst scaling: A model for the largest difference judgments
Jordan J Louviere and George G Woodworth. 1991 · 1991
Earlier work this paper cites.
Contextual correlates of semantic similarity
George A Miller and Walter G Charles. 1991 · 1991
Earlier work this paper cites.
Lexical cohesion computed by thesaural relations as an indicator of the structure of text
Jane Morris and Graeme Hirst. 1991 · 1991
Earlier work this paper cites.
WordNet: A lexical database for English
George A. Miller. 1994 · 1994
Earlier work this paper cites.
BRUJA: Question classification for Spanish. using machine translationand an English classifier
Miguel Á. García Cumbreras, L. Alfonso Ureña López, and Fernando Martínez Santiago. 2006 · 2006
Earlier work this paper cites.
Language-agnostic BERT sentence embedding
Fangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan, and Wei Wang. 2020 · 2007
Earlier work this paper cites.
Measuring Semantic Distance Using Distributional Profiles of Concepts
Saif Mohammad. 2008 · 2008
Earlier work this paper cites.
Maxdiff analysis : Simple counting , individual-level logit , and hb
Bryan K. Orme. 2009 · 2009
Earlier work this paper cites.
SemEval-2012 Task 6: A pilot on semantic textual similarity
Eneko Agirre, Daniel Cer, Mona Diab, and Aitor Gonzalez-Agirre. 2012 · 2012
Earlier work this paper cites.
Distributional Measures of Semantic Distance: A Survey
Saif M Mohammad and Graeme Hirst. 2012 · 2012
Earlier work this paper cites.
*SEM 2013 shared task: Semantic textual similarity
Eneko Agirre, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, and Weiwei Guo. 2013a · 2013
Earlier work this paper cites.
*SEM 2013 shared task: Semantic textual similarity
Eneko Agirre, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, and Weiwei Guo. 2013b · 2013
Earlier work this paper cites.
SemEval-2014 task 10: Multilingual semantic textual similarity
Eneko Agirre, Carmen Banea, Claire Cardie, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, Weiwei Guo, Rada Mihalcea, German Rigau, and Janyce Wiebe. 2014 · 2014
Earlier work this paper cites.
Best-worst scaling: Theory and methods
Terry N Flynn and Anthony AJ Marley. 2014 · 2014
Earlier work this paper cites.
A SICK cure for the evaluation of compositional distributional semantic models
Marco Marelli, Stefano Menini, Marco Baroni, Luisa Bentivogli, Raffaella Bernardi, Roberto Zamparelli, et al. 2014 · 2014
Earlier work this paper cites.
SemEval-2015 task 2: Semantic textual similarity, English, Spanish and pilot on interpretability
Eneko Agirre, Carmen Banea, Claire Cardie, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, Weiwei Guo, Inigo Lopez-Gazpio, Montse Maritxalar, Rada Mihalcea, et al. 2015 · 2015
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Machine translation experiments on PADIC: A parallel Arabic dialect corpus
Karima Meftouh, Salima Harrat, Salma Jamoussi, Mourad Abbas, and Kamel Smaili. 2015 · 2015
Earlier work this paper cites.
Improving SMT by model filtering and phrase embedding
Chengqing Zong. 2015 · 2015
Earlier work this paper cites.
SemEval-2016 Task 1: Semantic textual similarity, monolingual and cross-lingual evaluation
Eneko Agirre, Carmen Banea, Daniel Cer, Mona Diab, Aitor Gonzalez Agirre, Rada Mihalcea, German Rigau Claramunt, and Janyce Wiebe. 2016 · 2016
Earlier work this paper cites.
Capturing reliable fine-grained sentiment associations by crowdsourcing and best–worst scaling
Svetlana Kiritchenko and Saif M. Mohammad. 2016 · 2016
Cited alongside, same era.
CALYOU: A comparable spoken Algerian corpus harvested from YouTube
Karima Abidi, Mohamed Amine Menacer, and Kamel Smaili. 2017 · 2017
Cited alongside, same era.
Semeval-2017 task 1: Semantic textual similarity-multilingual and cross-lingual focused evaluation
Daniel Cer, Mona Diab, Eneko Agirre, Inigo Lopez-Gazpio, and Lucia Specia. 2017a · 2017
Cited alongside, same era.
SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation
Daniel Cer, Mona Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia. 2017b · 2017
Cited alongside, same era.
Best-worst scaling more reliable than rating scales: A case study on sentiment intensity annotation
Svetlana Kiritchenko and Saif M Mohammad. 2017 · 2017
A monolingual approach to contextualized word embeddings for mid-resource languages
Pedro Javier Ortiz Suárez, Laurent Romary, and Benoît Sagot. 2020 · 2020
Later among the works it cites.
Multi-simlex: A large-scale evaluation of multilingual and crosslingual lexical semantic similarity
Ivan Vulić, Simon Baker, Edoardo Maria Ponti, Ulla Petti, Ira Leviant, Kelly Wing, Olga Majewska, Eden Bar, Matt Malone, Thierry Poibeau, et al. 2020 · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Later among the works it cites.
Bertscore: Evaluating text generation with BERT
Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Later among the works it cites.
Dziribert: A pre-trained language model for the Algerian dialect
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Stance and sentiment in tweets
Saif M Mohammad, Parinaz Sobhani, and Svetlana Kiritchenko. 2017 · 2017
Cited alongside, same era.
John Wieting and Kevin Gimpel. 2017 · 2017
Cited alongside, same era.
A resource-light method for cross-lingual semantic textual similarity
Goran Glavaš, Marc Franco-Salvador, Simone P Ponzetto, and Paolo Rosso. 2018 · 2018
Cited alongside, same era.
Indosum: A new benchmark dataset for Indonesian text summarization
Kemal Kurniawan and Samuel Louvan. 2018 · 2018
Cited alongside, same era.
Dear sir or madam, may I introduce the GYAFC dataset: Corpus, benchmarks and metrics for formality style transfer
Sudha Rao and Joel Tetreault. 2018 · 2018
Cited alongside, same era.
Semi-supervised textual entailment on Indonesian wikipedia data
Ken Nabila Setya and Rahmad Mahendra. 2018 · 2018
Cited alongside, same era.
Xin Tang, Shanbo Cheng, Loc Do, Zhiyu Min, Feng Ji, Heng Yu, Ji Zhang, and Haiqin Chen. 2018 · 2018
Cited alongside, same era.
Amine Abdaoui, Mohamed Berrimi, Mourad Oussalah, and Abdelouahab Moussaoui. 2021 · 2021
Later among the works it cites.
ARBERT & MARBERT: Deep bidirectional transformers for Arabic
Muhammad Abdul-Mageed, AbdelRahim Elmadany, and El Moatez Billah Nagoudi. 2021 · 2021
Later among the works it cites.
Xl-sum: Large-scale multilingual abstractive summarization for 44 languages
Tahmid Hasan, Abhik Bhattacharjee, Md Saiful Islam, Kazi Samin, Yuan-Fang Li, Yong-Bin Kang, M Sohel Rahman, and Rifat Shahriyar. 2021 · 2021
Later among the works it cites.
Silt: Efficient transformer training for inter-lingual inference
Javier Huertas-Tato, Alejandro Martín, and David Camacho. 2021 · 2021
Later among the works it cites.
Introducing various semantic models for Amharic: Experimentation and evaluation with multiple tasks and datasets
Seid Muhie Yimam, Abinew Ali Ayele, Gopalakrishnan Venkatesh, Ibrahim Gashaw, and Chris Biemann. 2021 · 2021
Later among the works it cites.
A word pair dataset for semantic similarity and relatedness in Korean medical vocabulary: Reference development and validation
Yunjin Yum, Jeong Moon Lee, Moon Joung Jang, Yoojoong Kim, Jong-Ho Kim, Seongtae Kim, Unsub Shin, Sanghoun Song, and Hyung Joon Joo. 2021 · 2021
Later among the works it cites.
MasakhaNER 2.0: Africa-centric transfer learning for named entity recognition
David Adelani, Graham Neubig, Sebastian Ruder, Shruti Rijhwani, Michael Beukman, Chester Palen-Michel, Constantine Lignos, Jesujoba Alabi, Shamsuddeen Muhammad, Peter Nabende, Cheikh M. Bamba Dione, Andiswa Bukula, Rooweither Mabuya, Bonaventure F. P. Dossou, Blessing Sibanda, Happy Buzaaba, Jonathan Mukiibi, Godson Kalipe, Derguene Mbaye, Amelia Taylor, Fatoumata Kabore, Chris Chinenye Emezue, Anuoluwapo Aremu, Perez Ogayo, Catherine Gitau, Edwin Munkoh-Buabeng, Victoire Memdjokam Koagne, Allahsera Auguste Tapo, Tebogo Macucwa, Vukosi Marivate, Mboning Tchiaze Elvis, Tajuddeen Gwadabe, Tosin Adewumi, Orevaoghene Ahia, Joyce Nakatumba-Nabende, Neo Lerato Mokono, Ignatius Ezeani, Chiamaka Chukwuneke, Mofetoluwa Oluwaseun Adeyemi, Gilles Quentin Hacheme, Idris Abdulmumin, Odunayo Ogundepo, Oreen Yousuf, Tatiana Moteu, and Dietrich Klakow. 2022 · 2022
Later among the works it cites.
Ultimate Arabic News Dataset
Ahmed Hashim Al-Dulaimi. 2022 · 2022
Later among the works it cites.
Adapting pre-trained language models to African languages via multilingual adaptive fine-tuning
Jesujoba O. Alabi, David Ifeoluwa Adelani, Marius Mosbach, and Dietrich Klakow. 2022 · 2022
Later among the works it cites.
Evaluation benchmarks for Spanish sentence representations
Vladimir Araujo, Andrés Carvallo, Souvik Kundu, José Cañete, Marcelo Mendoza, Robert E. Mercer, Felipe Bravo-Marquez, Marie-Francine Moens, and Alvaro Soto. 2022 · 2022
Later among the works it cites.
ALBETO and DistilBETO: Lightweight Spanish language models
José Cañete, Sebastian Donoso, Felipe Bravo-Marquez, Andrés Carvallo, and Vladimir Araujo. 2022 · 2022
Later among the works it cites.
Maria: Spanish language models
Asier Gutiérrez Fandiño, Jordi Armengol Estapé, Marc Pàmies, Joan Llop Palao, Joaquin Silveira Ocampo, Casimiro Pio Carrino, Carme Armentano Oller, Carlos Rodriguez Penagos, Aitor Gonzalez Agirre, and Marta Villegas. 2022 · 2022
Later among the works it cites.
Goud.ma: a news article dataset for summarization in Moroccan Darija
Abderrahmane Issam and Khalil Mrini. 2022 · 2022
Later among the works it cites.
The bigscience roots corpus: A 1.6 tb composite multilingual dataset
Hugo Laurençon, Lucile Saulnier, Thomas Wang, Christopher Akiki, Albert Villanova del Moral, Teven Le Scao, Leandro Von Werra, Chenghao Mou, Eduardo González Ponferrada, Huu Nguyen, et al. 2022 · 2022
Later among the works it cites.
Just rank: Rethinking evaluation with word and sentence similarities
Bin Wang, C.-C. Jay Kuo, and Haizhou Li. 2022 · 2022
Later among the works it cites.
Relational sentence embedding for flexible semantic matching
Bin Wang and Haizhou Li. 2022 · 2022
Later among the works it cites.
What makes sentences semantically related? a textual relatedness dataset and empirical study
Mohamed Abdalla, Krishnapriya Vishnubhotla, and Saif Mohammad. 2023 · 2023
Later among the works it cites.
Mukhyansh: A headline generation dataset for Indic languages
Lokesh Madasu, Gopichand Kanumolu, Nirmal Surange, and Manish Shrivastava. 2023 · 2023
Later among the works it cites.
KUISAIL at SemEval-2020 task 12: BERT-CNN for offensive speech identification in social media
Ali Safaya, Moutasem Abdullatif, and Deniz Yuret. 2020 · 2059
Closest in time.