Fetching the paper…
Reading the bibliography…
Though nearest neighbor Machine Translation ($k$NN-MT) \citep{khandelwal2020nearest} has proved to introduce significant performance boosts over standard neural MT systems, it is prohibitively slow since it uses the entire reference corpus as the datastore for the nearest neighbor search.
Massively multilingual neural machine translation
Roee Aharoni, Melvin Johnson, and Orhan Firat. 2019 · 1903
Earlier work this paper cites.
Non-parametric adaptation for neural machine translation
Ankur Bapna and Orhan Firat. 2019 · 1903
Earlier work this paper cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 1904
Earlier work this paper cites.
Learning deep transformer models for machine translation
Qiang Wang, Bei Li, Tong Xiao, Jingbo Zhu, Changliang Li, Derek F Wong, and Lidia S Chao. 2019 · 1906
Earlier work this paper cites.
Massively multilingual neural machine translation in the wild: Findings and challenges
Naveen Arivazhagan, Ankur Bapna, Orhan Firat, Dmitry Lepikhin, Melvin Johnson, Maxim Krikun, Mia Xu Chen, Yuan Cao, George Foster, Colin Cherry, et al. 2019 · 1907
Earlier work this paper cites.
Facebook fair’s wmt19 news translation task submission
Nathan Ng, Kyra Yee, Alexei Baevski, Myle Ott, Michael Auli, and Sergey Edunov. 2019 · 1907
Earlier work this paper cites.
Large-scale pretraining for neural machine translation with tens of billions of sentence pairs
Yuxian Meng, Xiangyuan Ren, Zijun Sun, Xiaoya Li, Arianna Yuan, Fei Wu, and Jiwei Li. 2019 · 1909
Earlier work this paper cites.
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2019 · 1910
Earlier work this paper cites.
Transformers without tears: Improving the normalization of self-attention
Toan Q Nguyen and Julian Salazar. 2019 · 1910
Earlier work this paper cites.
Generalization through memorization: Nearest neighbor language models
Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2019 · 1911
Earlier work this paper cites.
The mathematics of statistical machine translation: Parameter estimation
Peter F Brown, Stephen A Della Pietra, Vincent J Della Pietra, and Robert L Mercer. 1993 · 1993
Earlier work this paper cites.
Realm: Retrieval-augmented language model pre-training
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020 · 2002
Earlier work this paper cites.
Incorporating bert into neural machine translation
Jinhua Zhu, Yingce Xia, Lijun Wu, Di He, Tao Qin, Wengang Zhou, Houqiang Li, and Tie-Yan Liu. 2020 · 2002
Earlier work this paper cites.
Sac: Accelerating and structuring self-attention via sparse adaptive connection
Xiaoya Li, Yuxian Meng, Mingxin Zhou, Qinghong Han, Fei Wu, and Jiwei Li. 2020 · 2003
Earlier work this paper cites.
A systematic comparison of various statistical alignment models
Franz Josef Och and Hermann Ney. 2003 · 2003
Earlier work this paper cites.
Unsupervised domain clusters in pretrained language models
Roee Aharoni and Yoav Goldberg. 2020 · 2004
Earlier work this paper cites.
Augmenting transformers with knn-based composite memory for dialogue
Angela Fan, Claire Gardent, Chloe Braud, and Antoine Bordes. 2020 · 2004
Earlier work this paper cites.
Understanding the difficulty of training transformers
Liyuan Liu, Xiaodong Liu, Jianfeng Gao, Weizhu Chen, and Jiawei Han. 2020a · 2004
Cited alongside, same era.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandara Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020b · 2005
Cited alongside, same era.
Synthesizer: Rethinking self-attention in transformer models
Yi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan, Zhe Zhao, and Che Zheng. 2020 · 2005
Cited alongside, same era.
Deep encoder, shallow decoder: Reevaluating the speed-quality tradeoff in machine translation
Jungo Kasai, Nikolaos Pappas, Hao Peng, James Cross, and Noah A Smith. 2020 · 2006
Cited alongside, same era.
Six challenges for neural machine translation
Philipp Koehn and Rebecca Knowles. 2017 · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Later among the works it cites.
Encoding gated translation memory into neural machine translation
Qian Cao and Deyi Xiong. 2018 · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Later among the works it cites.
Search engine guided neural machine translation
Jiatao Gu, Yong Wang, Kyunghyun Cho, and Victor OK Li. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mike Lewis, Marjan Ghazvininejad, Gargi Ghosh, Armen Aghajanyan, Sida Wang, and Luke Zettlemoyer. 2020a · 2006
Cited alongside, same era.
Approximate nearest neighbor negative contrastive learning for dense text retrieval
Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul Bennett, Junaid Ahmed, and Arnold Overwijk. 2020a · 2007
Cited alongside, same era.
Incorporating bert into parallel sequence decoding with adapters
Junliang Guo, Zhirui Zhang, Linli Xu, Hao-Ran Wei, Boxing Chen, and Enhong Chen. 2020 · 2010
Cited alongside, same era.
Product quantization for nearest neighbor search
Herve Jegou, Matthijs Douze, and Cordelia Schmid. 2010 · 2010
Cited alongside, same era.
Nearest neighbor machine translation
Urvashi Khandelwal, Angela Fan, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2020 · 2010
Cited alongside, same era.
A simple, fast, and effective reparameterization of ibm model 2
Chris Dyer, Victor Chahuneau, and Noah A Smith. 2013 · 2013
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014 · 2014
Cited alongside, same era.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Later among the works it cites.
Learning to remember translation history with a continuous cache
Zhaopeng Tu, Yang Liu, Shuming Shi, and Tong Zhang. 2018 · 2018
Later among the works it cites.
Retrieve and refine: Improved sequence generation models for dialogue
Jason Weston, Emily Dinan, and Alexander H Miller. 2018 · 2018
Later among the works it cites.
Guiding neural machine translation with retrieved translation pieces
Jingyi Zhang, Masao Utiyama, Eiichro Sumita, Graham Neubig, and Satoshi Nakamura. 2018 · 2018
Later among the works it cites.
Neural fuzzy repair: Integrating fuzzy matches into neural machine translation
Bram Bulté and Arda Tezcan. 2019 · 2019
Later among the works it cites.
Billion-scale similarity search with gpus
Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2019 · 2019
Later among the works it cites.
Boosting neural machine translation with similar translations
XU Jitao, Josep M Crego, and Jean Senellart. 2020 · 2020
Later among the works it cites.
Time-aware large kernel convolutions
Vasileios Lioutas and Yuhong Guo. 2020 · 2020
Later among the works it cites.
Finetuning pretrained transformers into rnns
Jungo Kasai, Hao Peng, Yizhe Zhang, Dani Yogatama, Gabriel Ilharco, Nikolaos Pappas, Yi Mao, Weizhu Chen, and Noah A Smith. 2021 · 2021
Closest in time.
Hao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz, Noah A Smith, and Lingpeng Kong. 2021 · 2021
Closest in time.
Efficient retrieval augmented generation from unstructured knowledge for task-oriented dialog
David Thulke, Nico Daheim, Christian Dugast, and Hermann Ney. 2021 · 2021
Closest in time.