Fetching the paper…
Reading the bibliography…
Continuous-output neural machine translation (CoNMT) replaces the discrete next-word prediction problem with an embedding prediction.
compare-mt: A tool for holistic comparison of language generation systems
Graham Neubig, Zi-Yi Dou, Junjie Hu, Paul Michel, Danish Pruthi, Xinyi Wang, and John Wieting. 2019 · 1903
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Word representations: A simple and general method for semi-supervised learning
Joseph Turian, Lev-Arie Ratinov, and Yoshua Bengio. 2010 · 2010
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomás Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013 · 2013
Earlier work this paper cites.
Improving zero-shot learning by mitigating the hubness problem
Georgiana Dinu and Marco Baroni. 2014 · 2014
Earlier work this paper cites.
Hubness and pollution: Delving into cross-space mapping for zero-shot learning
Angeliki Lazaridou, Georgiana Dinu, and Marco Baroni. 2015 · 2015
Earlier work this paper cites.
Enriching word vectors with subword information
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Subword regularization: Improving neural network translation models with multiple subword candidates
Taku Kudo. 2018 · 2018
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Earlier work this paper cites.
Advances in pre-training distributed word representations
Tomas Mikolov, Edouard Grave, Piotr Bojanowski, Christian Puhrsch, and Armand Joulin. 2018 · 2018
Cited alongside, same era.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Cited alongside, same era.
A margin-based loss with synthetic negative samples for continuous-output machine translation
Gayatri Bhat, Sachin Kumar, and Yulia Tsvetkov. 2019 · 2019
Cited alongside, same era.
Hubless nearest neighbor search for bilingual lexicon induction
Jiaji Huang, Qiang Qiu, and Kenneth Church. 2019 · 2019
Cited alongside, same era.
Von Mises-Fisher loss for training sequence to sequence models with continuous outputs
Sachin Kumar and Yulia Tsvetkov. 2019 · 2019
Cited alongside, same era.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
BERTScore: Evaluating text generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Later among the works it cites.
Nearest neighbor machine translation
Urvashi Khandelwal, Angela Fan, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2021 · 2021
Later among the works it cites.
von Mises–Fisher loss: An exploration of embedding geometries for supervised learning
Tyler R. Scott, Andrew C. Gallagher, and Michael C. Mozer. 2021 · 2021
Later among the works it cites.
Adaptive semiparametric language models
Dani Yogatama, Cyprien de Masson d’Autume, and Lingpeng Kong. 2021 · 2021
Later among the works it cites.
Language-agnostic BERT sentence embedding
Fangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan, and Wei Wang. 2022 · 2022
Later among the works it cites.
Low-rank softmax can have unargmaxable classes in theory but rarely in practice
Andreas Grivas, Nikolay Bogoychev, and Adam Lopez. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Sentence-BERT: Sentence embeddings using Siamese BERT-networks
Nils Reimers and Iryna Gurevych. 2019 · 2019
Cited alongside, same era.
Contextual embeddings: When are they worth it?
Simran Arora, Avner May, Jian Zhang, and Christopher Ré. 2020 · 2020
Cited alongside, same era.
Multilingual denoising pre-training for neural machine translation
Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020 · 2020
Cited alongside, same era.
Jiahuan Li, Shanbo Cheng, Zewei Sun, Mingxuan Wang, and Shujian Huang. 2022a
Cited in the paper.
Diffusion-LM improves controllable text generation
Xiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang, and Tatsunori Hashimoto. 2022b
Cited in the paper.
Later among the works it cites.
On target representation in continuous-output neural machine translation
Evgeniia Tokarchuk and Vlad Niculae. 2022 · 2022
Later among the works it cites.
Problems with cosine as a measure of embedding similarity for high frequency words
Kaitlyn Zhou, Kawin Ethayarajh, Dallas Card, and Dan Jurafsky. 2022 · 2022
Later among the works it cites.
Multilingual k k -nearest-neighbor machine translation
David Stap and Christof Monz. 2023 · 2023
Closest in time.