Fetching the paper…
Reading the bibliography…
We present a model and methodology for learning paraphrastic sentence embeddings directly from bitext, removing the time-consuming intermediate step of creating paraphrase corpora.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Unsupervised construction of large paraphrase corpora: Exploiting massively parallel news sources
Bill Dolan, Chris Quirk, and Chris Brockett. 2004 · 2004
Earlier work this paper cites.
Improved statistical machine translation using paraphrases
Chris Callison-Burch, Philipp Koehn, and Miles Osborne. 2006 · 2006
Earlier work this paper cites.
Simple English Wikipedia: a new text simplification task
William Coster and David Kauchak. 2011 · 2011
Earlier work this paper cites.
SemEval-2012 task 6: A pilot on semantic textual similarity
Eneko Agirre, Mona Diab, Daniel Cer, and Aitor Gonzalez-Agirre. 2012 · 2012
Earlier work this paper cites.
* sem 2013 shared task: Semantic textual similarity
Eneko Agirre, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, and Weiwei Guo. 2013 · 2013
Earlier work this paper cites.
PPDB: The Paraphrase Database
Juri Ganitkevitch, Benjamin Van Durme, and Chris Callison-Burch. 2013 · 2013
Earlier work this paper cites.
SemEval-2014 task 10: Multilingual semantic textual similarity
Eneko Agirre, Carmen Banea, Claire Cardie, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, Weiwei Guo, Rada Mihalcea, German Rigau, and Janyce Wiebe. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014 · 2014
Earlier work this paper cites.
SemEval-2015 task 2: Semantic textual similarity, English, Spanish and pilot on interpretability
Eneko Agirre, Carmen Banea, Claire Cardie, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, Weiwei Guo, Inigo Lopez-Gazpio, Montse Maritxalar, Rada Mihalcea, German Rigau, Larraitz Uria, and Janyce Wiebe. 2015 · 2015
Earlier work this paper cites.
Jointly optimizing word representations for lexical and sentential tasks with the c-phrase model
Nghia The Pham, Germán Kruszewski, Angeliki Lazaridou, and Marco Baroni. 2015 · 2015
Earlier work this paper cites.
SemEval-2016 task 1: Semantic textual similarity, monolingual and cross-lingual evaluation
Eneko Agirre, Carmen Banea, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, Rada Mihalcea, German Rigau, and Janyce Wiebe. 2016 · 2016
Earlier work this paper cites.
CzEng 1.6: Enlarged Czech-English Parallel Corpus with Processing Tools Dockered
Ondřej Bojar, Ondřej Dušek, Tom Kocmi, Jindřich Libovický, Michal Novák, Martin Popel, Roman Sudarikov, and Dušan Variš. 2016 · 2016
Cited alongside, same era.
Learning distributed representations of sentences from unlabelled data
Felix Hill, Kyunghyun Cho, and Anna Korhonen. 2016 · 2016
Cited alongside, same era.
Opensubtitles2016: Extracting large parallel corpora from movie and tv subtitles
Pierre Lison and Jörg Tiedemann. 2016 · 2016
Cited alongside, same era.
Charagram: Embedding words and sentences via character n n -grams
John Wieting, Mohit Bansal, Kevin Gimpel, and Karen Livescu. 2016a · 2016
Cited alongside, same era.
SemEval-2017 Task 1: Semantic textual similarity multilingual and crosslingual focused evaluation
Daniel Cer, Mona Diab, Eneko Agirre, Inigo Lopez-Gazpio, and Lucia Specia. 2017 · 2017
Cited alongside, same era.
Learning cross-lingual sentence representations via a multi-task dual-encoder model
Muthuraman Chidambaram, Yinfei Yang, Daniel Cer, Steve Yuan, Yun-Hsuan Sung, Brian Strope, and Ray Kurzweil. 2018 · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Later among the works it cites.
Francis Grégoire and Philippe Langlais. 2018 · 2018
Later among the works it cites.
Effective parallel corpus mining using bilingual sentence embeddings
Mandy Guo, Qinlan Shen, Yinfei Yang, Heming Ge, Daniel Cer, Gustavo Hernandez Abrego, Keith Stevens, Yun-Hsuan Sung, Brian Strope, and Ray Kurzweil. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Supervised learning of universal sentence representations from natural language inference data
Alexis Conneau, Douwe Kiela, Holger Schwenk, Loïc Barrault, and Antoine Bordes. 2017 · 2017
Cited alongside, same era.
An empirical analysis of nmt-derived interlingual embeddings and their use in parallel sentence identification
Cristina Espana-Bonet, Adám Csaba Varga, Alberto Barrón-Cedeño, and Josef van Genabith. 2017 · 2017
Cited alongside, same era.
A continuously growing dataset of sentential paraphrases
Wuwei Lan, Siyu Qiu, Hua He, and Wei Xu. 2017 · 2017
Cited alongside, same era.
Uriel and lang2vec: Representing languages as typological, geographical, and phylogenetic vectors
Patrick Littell, David R. Mortensen, Ke Lin, Katherine Kairis, Carlisle Turner, and Lori Levin. 2017 · 2017
Cited alongside, same era.
Learned in translation: Contextualized word vectors
Bryan McCann, James Bradbury, Caiming Xiong, and Richard Socher. 2017 · 2017
Cited alongside, same era.
Unsupervised learning of sentence embeddings using compositional n-gram features
Matteo Pagliardini, Prakhar Gupta, and Martin Jaggi. 2017 · 2017
Cited alongside, same era.
Learning joint multilingual sentence representations with neural machine translation
Holger Schwenk and Matthijs Douze. 2017 · 2017
Cited alongside, same era.
Taku Kudo and John Richardson. 2018 · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
Filtering and mining parallel data in a joint multilingual space
Holger Schwenk. 2018 · 2018
Later among the works it cites.
A multi-task approach to learning multilingual representations
Karan Singla, Dogan Can, and Shrikanth Narayanan. 2018 · 2018
Later among the works it cites.
ParaNMT-50M: Pushing the limits of paraphrastic sentence embeddings with millions of machine translations
John Wieting and Kevin Gimpel. 2018 · 2018
Later among the works it cites.
Overview of the third bucc shared task: Spotting parallel sentences in comparable corpora
Pierre Zweigenbaum, Serge Sharoff, and Reinhard Rapp. 2018 · 2018
Later among the works it cites.
ParaBank: Monolingual bitext generation and sentential paraphrasing via lexically-constrained neural machine translation
J Edward Hu, Rachel Rudinger, Matt Post, and Benjamin Van Durme. 2019 · 2019
Closest in time.
Revisiting recurrent networks for paraphrastic sentence embeddings
John Wieting and Kevin Gimpel. 2017 · 2088
Closest in time.