Fetching the paper…
Reading the bibliography…
The scarcity of large parallel corpora is an important obstacle for neural machine translation.
Distilling the knowledge of bert for text generation
Yen-Chun Chen, Zhe Gan, Yu Cheng, Jingzhou Liu, and Jingjing Liu. 2019 · 1911
Earlier work this paper cites.
The mathematical theory of communication
Claude E Shannon and Warren Weaver. 1949 · 1949
Earlier work this paper cites.
The mathematics of statistical machine translation: Parameter estimation
Peter F. Brown, Stephen A. Della Pietra, Vincent J. Della Pietra, and Robert L. Mercer. 1993 · 1993
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Model compression
Cristian Bucila, Rich Caruana, and Alexandru Niculescu-Mizil. 2006 · 2006
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio. 2010 · 2010
Earlier work this paper cites.
Statistical Machine Translation , 1st edition
Philipp Koehn. 2010 · 2010
Earlier work this paper cites.
Rnnlm-recurrent neural network language modeling toolkit
Tomas Mikolov, Stefan Kombrink, Anoop Deoras, Lukar Burget, and Jan Cernocky. 2011 · 2011
Earlier work this paper cites.
Do deep nets really need to be deep?
Jimmy Ba and Rich Caruana. 2014 · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014 · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
On using monolingual corpora in neural machine translation
Caglar Gulcehre, Orhan Firat, Kelvin Xu, Kyunghyun Cho, Loic Barrault, Huei-Chi Lin, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Thang Luong, Hieu Pham, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Lei Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016 · 2016
Earlier work this paper cites.
Findings of the 2016 conference on machine translation
Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Aurélie Névéol, Mariana Neves, Martin Popel, Matt Post, Raphael Rubino, Carolina Scarton, Lucia Specia, Marco Turchi, Karin Verspoor, and Marcos Zampieri. 2016 · 2016
Earlier work this paper cites.
Sequence-level knowledge distillation
Yoon Kim and Alexander M. Rush. 2016 · 2016
Earlier work this paper cites.
Language as a Latent Variable: Discrete Generative Models for Sentence Compression
Yishu Miao and Phil Blunsom. 2016 · 2016
Earlier work this paper cites.
Improving neural machine translation models with monolingual data
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. 2016 · 2016
Cited alongside, same era.
Using target-side monolingual data for neural machine translation through multi-task learning
Tobias Domhan and Felix Hieber. 2017 · 2017
Cited alongside, same era.
Emergence of language with multi-agent games: Learning to communicate with sequences of symbols
Serhii Havrylov and Ivan Titov. 2017 · 2017
Cited alongside, same era.
Tying word vectors and word classifiers: A loss framework for language modeling
Hakan Inan, Khashayar Khosravi, and Richard Socher. 2017 · 2017
Cited alongside, same era.
Six challenges for neural machine translation
Philipp Koehn and Rebecca Knowles. 2017 · 2017
Cited alongside, same era.
Using the output embedding to improve language models
On the use of BERT for neural machine translation
Stephane Clinchant, Kweon Woo Jung, and Vassilina Nikoulina. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Recycling a pre-trained BERT encoder for neural machine translation
Kenji Imamura and Eiichiro Sumita. 2019 · 2019
Later among the works it cites.
Joey NMT: A minimalist NMT toolkit for novices
Julia Kreutzer, Joost Bastings, and Stefan Riezler. 2019 · 2019
Later among the works it cites.
When does label smoothing help?
Rafael Müller, Simon Kornblith, and Geoffrey E Hinton. 2019 · 2019
Later among the works it cites.
Transformers without tears: Improving the normalization of self-attention
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ofir Press and Lior Wolf. 2017 · 2017
Cited alongside, same era.
Unsupervised pretraining for sequence to sequence learning
Prajit Ramachandran, Peter Liu, and Quoc Le. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
The neural noisy channel
Lei Yu, Phil Blunsom, Chris Dyer, Edward Grefenstette, and Tomás Kociský. 2017 · 2017
Cited alongside, same era.
Prior knowledge integration for neural machine translation using posterior regularization
Jiacheng Zhang, Yang Liu, Huanbo Luan, Jingfang Xu, and Maosong Sun. 2017 · 2017
Cited alongside, same era.
Findings of the conference on machine translation (WMT)
Ondřej Bojar, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Philipp Koehn, and Christof Monz. 2018 · 2018
Cited alongside, same era.
Iterative back-translation for neural machine translation
Vu Cong Duy Hoang, Philipp Koehn, Gholamreza Haffari, and Trevor Cohn. 2018 · 2018
Cited alongside, same era.
Toan Q. Nguyen and Julian Salazar. 2019 · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
Unsupervised neural machine translation with smt as posterior regularization
Shuo Ren, Zhirui Zhang, Shujie Liu, Ming Zhou, and Shuai Ma. 2019 · 2019
Later among the works it cites.
Revisiting low-resource neural machine translation: A case study
Rico Sennrich and Biao Zhang. 2019 · 2019
Later among the works it cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019 · 2019
Later among the works it cites.
Towards making the most of BERT in neural machine translation
Jiacheng Yang, Mingxuan Wang, Hao Zhou, Chengqi Zhao, Yong Yu, Weinan Zhang, and Lei Li. 2019 · 2019
Later among the works it cites.
Simple and effective noisy channel modeling for neural machine translation
Kyra Yee, Yann Dauphin, and Michael Auli. 2019 · 2019
Later among the works it cites.
Putting machine translation in context with the noisy channel model
Lei Yu, Laurent Sartran, Wojciech Stokowiec, Wang Ling, Lingpeng Kong, Phil Blunsom, and Chris Dyer. 2019 · 2019
Later among the works it cites.
Extreme language model compression with optimal subwords and shared projections
Sanqiang Zhao, Raghav Gupta, Yang Song, and Denny Zhou. 2019 · 2019
Later among the works it cites.
On layer normalization in the transformer architecture
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tie-Yan Liu. 2020 · 2020
Closest in time.
Understanding knowledge distillation in non-autoregressive machine translation
Chunting Zhou, Jiatao Gu, and Graham Neubig. 2020 · 2020
Closest in time.
Incorporating bert into neural machine translation
Jinhua Zhu, Yingce Xia, Lijun Wu, Di He, Tao Qin, Wengang Zhou, Houqiang Li, and Tieyan Liu. 2020 · 2020
Closest in time.
Posterior regularization for structured latent variable models
Kuzman Ganchev, Jennifer Gillenwater, Ben Taskar, et al. 2010 · 2049
Closest in time.