Fetching the paper…
Reading the bibliography…
In this paper, we investigate the problem of training neural machine translation (NMT) systems with a dataset of more than 40 billion bilingual sentence pairs, which is larger than the largest dataset to date by orders of magnitude.
Transfer learning for low-resource neural machine translation
Barret Zoph, Deniz Yuret, Jonathan May, and Kevin Knight · 1909
Earlier work this paper cites.
The part-of-speech tagging guidelines for the penn chinese treebank (3.0)
Fei Xia · 2000
Earlier work this paper cites.
Processing norms of modern chinese corpus
Shiwen Yu, Jianming Lu, Xuefeng Zhu, Huiming Duan, Shiyong Kang, Honglin Sun, Hui Wang, Qiang Zhao, and Weidong Zhan · 2001
Earlier work this paper cites.
Latent dirichlet allocation
David M Blei, Andrew Y Ng, and Michael I Jordan · 2003
Earlier work this paper cites.
Adaptation of the translation model for statistical machine translation based on information retrieval
Almut Silja Hildebrand, Matthias Eck, Stephan Vogel, and Alex Waibel · 2005
Earlier work this paper cites.
Bitam: Bilingual topic admixture models for word alignment
Bing Zhao and Eric P Xing · 2006
Earlier work this paper cites.
Domain adaptation in statistical machine translation with mixture modelling
Jorge Civera and Alfons Juan · 2007
Earlier work this paper cites.
Mixture-model adaptation for smt
George Foster and Roland Kuhn · 2007
Earlier work this paper cites.
Experiments in domain adaptation for statistical machine translation
Philipp Koehn and Josh Schroeder · 2007
Earlier work this paper cites.
Hm-bitam: Bilingual topic exploration, word alignment, and translation
Bing Zhao and Eric P Xing · 2008
Earlier work this paper cites.
Discriminative instance weighting for domain adaptation in statistical machine translation
George Foster, Cyril Goutte, and Roland Kuhn · 2010
Earlier work this paper cites.
Domain adaptation in statistical machine translation using factored translation models
Jan Niehues and Alex Waibel · 2010
Earlier work this paper cites.
Large scale parallel document mining for machine translation
Jakob Uszkoreit, Jay M Ponte, Ashok C Popat, and Moshe Dubiner · 2010
Earlier work this paper cites.
Domain adaptation via pseudo in-domain data selection
Amittai Axelrod, Xiaodong He, and Jianfeng Gao · 2011
Earlier work this paper cites.
Emergence of hierarchical structure mirroring linguistic composition in a recurrent neural network
Wataru Hinoshita, Hiroaki Arie, Jun Tani, Hiroshi G Okuno, and Tetsuya Ogata · 2011
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Andrew Senior, Paul Tucker, Ke Yang, Quoc V Le, et al · 2012
Earlier work this paper cites.
Topic models for dynamic translation model adaptation
Vladimir Eidelman, Jordan Boyd-Graber, and Philip Resnik · 2012
Earlier work this paper cites.
Perplexity minimization for translation model domain adaptation in statistical machine translation
Rico Sennrich · 2012
Earlier work this paper cites.
A topic similarity model for hierarchical phrase-based translation
Xinyan Xiao, Deyi Xiong, Min Zhang, Qun Liu, and Shouxun Lin · 2012
Earlier work this paper cites.
Learning factored representations in a deep mixture of experts
David Eigen, Marc’Aurelio Ranzato, and Ilya Sutskever · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean · 2013
Earlier work this paper cites.
Multi-gpu training of convnets
Omry Yadan, Keith Adams, Yaniv Taigman, and Marc’Aurelio Ranzato · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
One weird trick for parallelizing convolutional neural networks
Alex Krizhevsky · 2014
Cited alongside, same era.
Scaling distributed machine learning with the parameter server
Mu Li, David G Andersen, Jun Woo Park, Alexander J Smola, Amr Ahmed, Vanja Josifovski, James Long, Eugene J Shekita, and Bor-Yiing Su · 2014
Cited alongside, same era.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning · 2014
Cited alongside, same era.
Opennmt: Open-source toolkit for neural machine translation
Guillaume Klein, Yoon Kim, Yuntian Deng, Jean Senellart, and Alexander M Rush · 2017
Later among the works it cites.
Modeling source syntax for neural machine translation
Junhui Li, Deyi Xiong, Zhaopeng Tu, Muhua Zhu, Min Zhang, and Guodong Zhou · 2017
Later among the works it cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Later among the works it cites.
Don’t decay the learning rate, increase the batch size
Samuel L Smith, Pieter-Jan Kindermans, Chris Ying, and Quoc V Le · 2017
Later among the works it cites.
Attention is all you need
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le · 2014
Cited alongside, same era.
Generating sentences from a continuous space
Samuel R Bowman, Luke Vilnis, Oriol Vinyals, Andrew M Dai, Rafal Jozefowicz, and Samy Bengio · 2015
Cited alongside, same era.
Semi-supervised sequence learning
Andrew M Dai and Quoc V Le · 2015
Cited alongside, same era.
A diversity-promoting objective function for neural conversation models
Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan · 2015
Cited alongside, same era.
Effective approaches to attention-based neural machine translation
Minh-Thang Luong, Hieu Pham, and Christopher D Manning · 2015
Cited alongside, same era.
Sequence level training with recurrent neural networks
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba · 2015
Cited alongside, same era.
Guided alignment training for topic-aware neural machine translation
Wenhu Chen, Evgeny Matusov, Shahram Khadivi, and Jan-Thorsten Peter · 2016
Cited alongside, same era.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Later among the works it cites.
Sogou neural machine translation systems for wmt17
Yuguang Wang, Shanbo Cheng, Liyang Jiang, Jiajun Yang, Wei Chen, Muze Li, Lin Shi, Yanfeng Wang, and Hongtao Yang · 2017
Later among the works it cites.
Generating more interesting responses in neural conversation models with distributional constraints
Ashutosh Baheti, Alan Ritter, Jiwei Li, and Bill Dolan · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Later among the works it cites.
Understanding back-translation at scale
Sergey Edunov, Myle Ott, Michael Auli, and David Grangier · 2018
Later among the works it cites.
Achieving human parity on automatic chinese to english news translation
Hany Hassan, Anthony Aue, Chang Chen, Vishal Chowdhary, Jonathan Clark, Christian Federmann, Xuedong Huang, Marcin Junczys-Dowmunt, William Lewis, Mu Li, et al · 2018
Later among the works it cites.
Sequence to sequence mixture model for diverse machine translation
Xuanli He, Gholamreza Haffari, and Mohammad Norouzi · 2018
Later among the works it cites.
A tutorial on deep latent variable models of natural language
Yoon Kim, Sam Wiseman, and Alexander M Rush · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Later among the works it cites.
A stochastic decoder for neural machine translation
Philip Schulz, Wilker Aziz, and Trevor Cohn · 2018
Later among the works it cites.
Diverse beam search for improved description of complex scenes
Ashwin K Vijayakumar, Michael Cogswell, Ramprasaath R Selvaraju, Qing Sun, Stefan Lee, David Crandall, and Dhruv Batra · 2018
Later among the works it cites.
Do latent tree learning models identify meaningful structure in sentences?
Adina Williams, Andrew Drozdov*, and Samuel R Bowman · 2018
Later among the works it cites.
Massively multilingual neural machine translation
Roee Aharoni, Melvin Johnson, and Orhan Firat · 2019
Closest in time.
Hard but robust, easy but sensitive: How encoder and decoder perform in neural machine translation
Tianyu He, Xu Tan, and Tao Qin · 2019
Closest in time.
Cross-lingual language model pretraining
Guillaume Lample and Alexis Conneau · 2019
Closest in time.
Is word segmentation necessary for deep learning of chinese representations?
Yuxian Meng, Xiaoya Li, Xiaofei Sun, Qinghong Han, Arianna Yuan, and Jiwei Li · 2019
Closest in time.
Facebook fair’s wmt19 news translation task submission
Nathan Ng, Kyra Yee, Alexei Baevski, Myle Ott, Michael Auli, and Sergey Edunov · 2019
Closest in time.
Mixture models for diverse machine translation: Tricks of the trade
Tianxiao Shen, Myle Ott, Michael Auli, and Marc’Aurelio Ranzato · 2019
Closest in time.
Mass: Masked sequence to sequence pre-training for language generation
Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu · 2019
Closest in time.