Fetching the paper…
Reading the bibliography…
We show that Bayes' rule provides an effective mechanism for creating document translation models that can be learned from only parallel sentences and monolingual documents---a compelling benefit as parallel documents are not always available.
Cross-lingual language model pretraining
Guillaume Lample and Alexis Conneau. 2019 · 1901
Earlier work this paper cites.
Pretraining-based natural language generation for text summarization
Haoyu Zhang, Yeyun Gong, Yu Yan, Nan Duan, Jianjun Xu, Ji Wang, Ming Gong, and Ming Zhou. 2019 · 1902
Earlier work this paper cites.
Unified language model pre-training for natural language understanding and generation
Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. 2019 · 1905
Earlier work this paper cites.
Defending against neural fake news
Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi. 2019 · 1905
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019 · 1906
Earlier work this paper cites.
Repurposing decoder-transformer language models for abstractive summarization
Luke de Oliveira and Alfredo Láinez Rodrigo. 2019 · 1909
Earlier work this paper cites.
Encoder-agnostic adaptation for conditional language generation
Zachary M. Ziegler, Luke Melas-Kyriazi, Sebastian Gehrmann, and Alexander M. Rush. 2019 · 1910
Earlier work this paper cites.
The mathematics of statistical machine translation: Parameter estimation
Peter F. Brown, Stephen Della Pietra, Vincent J. Della Pietra, and Robert L. Mercer. 1993 · 1993
Earlier work this paper cites.
A program for aligning sentences in bilingual corpora
William A. Gale and Kenneth W. Church. 1993 · 1993
Earlier work this paper cites.
Bayes-ball: The rational pastime (for determining irrelevance and requisite information in belief networks and influence diagrams)
Ross D. Shachter. 1998 · 1998
Earlier work this paper cites.
Finding consensus in speech recognition: word error minimization and other applications of confusion networks
Lidia Mangu, Eric Brill, and Andreas Stolcke. 2000 · 2000
Earlier work this paper cites.
A study of translation edit rate with targeted human annotation
Matthew Snover, Bonnie Dorr, Richard Schwarz, Linnea Micciulla, and John Makhoul. 2006 · 2006
Earlier work this paper cites.
Moses: Open source toolkit for statistical machine translation
Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, Chris Dyer, Ondrej Bojar, Alexandra Constantin, and Evan Herbst. 2007 · 2007
Earlier work this paper cites.
On using monolingual corpora in neural machine translation
Çaglar Gülçehre, Orhan Firat, Kelvin Xu, Kyunghyun Cho, Loïc Barrault, Huei-Chi Lin, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Semi-supervised learning for neural machine translation
Yong Cheng, Wei Xu, Zhongjun He, Wei He, Hua Wu, Maosong Sun, and Yang Liu. 2016 · 2016
Earlier work this paper cites.
Sequence-level knowledge distillation
Yoon Kim and Alexander M. Rush. 2016 · 2016
Earlier work this paper cites.
Online segment to segment neural transduction
Lei Yu, Jan Buys, and Phil Blunsom. 2016 · 2016
Cited alongside, same era.
Does neural machine translation benefit from larger context?
Sébastien Jean, Stanislas Lauly, Orhan Firat, and Kyunghyun Cho. 2017 · 2017
Cited alongside, same era.
Cache-based document-level neural machine translation
Shaohui Kuang, Deyi Xiong, Weihua Luo, and Guodong Zhou. 2017 · 2017
Cited alongside, same era.
Learned in translation: Contextualized word vectors
Bryan McCann, James Bradbury, Caiming Xiong, and Richard Socher. 2017 · 2017
Cited alongside, same era.
Neural machine translation with extended context
Jörg Tiedemann and Yves Scherrer. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Document-level neural machine translation with hierarchical attention networks
Lesly Miculicich Werlen, Dhananjay Ram, Nikolaos Pappas, and James Henderson. 2018 · 2018
Later among the works it cites.
Improving the transformer translation model with document-level context
Jiacheng Zhang, Huanbo Luan, Maosong Sun, Feifei Zhai, Jingfang Xu, Min Zhang, and Yang Liu. 2018 · 2018
Later among the works it cites.
An effective approach to unsupervised machine translation
Mikel Artetxe, Gorka Labaka, and Eneko Agirre. 2019 · 2019
Closest in time.
An embarrassingly simple approach for transfer learning from pretrained language models
Alexandra Chronopoulou, Christos Baziotis, and Alexandros Potamianos. 2019 · 2019
Closest in time.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G. Carbonell, Quoc Viet Le, and Ruslan Salakhutdinov. 2019 · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Exploiting cross-sentence context for neural machine translation
Longyue Wang, Zhaopeng Tu, Andy Way, and Qun Liu. 2017 · 2017
Cited alongside, same era.
The neural noisy channel
Lei Yu, Phil Blunsom, Chris Dyer, Edward Grefenstette, and Tomás Kociský. 2017 · 2017
Cited alongside, same era.
Unsupervised neural machine translation
Mikel Artetxe, Gorka Labaka, Eneko Agirre, and Kyunghyun Cho. 2018 · 2018
Cited alongside, same era.
Evaluating discourse phenomena in neural machine translation
Rachel Bawden, Rico Sennrich, Alexandra Birch, and Barry Haddow. 2018 · 2018
Cited alongside, same era.
Understanding back-translation at scale
Sergey Edunov, Myle Ott, Michael Auli, and David Grangier. 2018 · 2018
Cited alongside, same era.
Document context neural machine translation with memory networks
Gholamreza Haffari and Sameen Maruf. 2018 · 2018
Cited alongside, same era.
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Closest in time.
Pre-trained language model representations for language generation
Sergey Edunov, Alexei Baevski, and Michael Auli. 2019 · 2019
Closest in time.
Microsoft translator at WMT 2019: Towards large-scale document-level neural machine translation
Marcin Junczys-Dowmunt. 2019 · 2019
Closest in time.
Selective attention for context-aware neural machine translation
Sameen Maruf, André F. T. Martins, and Gholamreza Haffari. 2019 · 2019
Closest in time.
Facebook fair’s WMT19 news translation task submission
Nathan Ng, Kyra Yee, Alexei Baevski, Myle Ott, Michael Auli, and Sergey Edunov. 2019 · 2019
Closest in time.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Closest in time.
Mixture models for diverse machine translation: Tricks of the trade
Tianxiao Shen, Myle Ott, Michael Auli, and Marc’Aurelio Ranzato. 2019 · 2019
Closest in time.
MASS: masked sequence to sequence pre-training for language generation
Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu. 2019 · 2019
Closest in time.
Baidu neural machine translation systems for wmt19
Meng Sun, Bojian Jiang, Hao Xiong, Zhongjun He, Hua Wu, and Haifeng Wang. 2019 · 2019
Closest in time.
Microsoft research asia’s systems for wmt19
Yingce Xia, Xu Tan, Fei Tian, Fei Gao, Di He, Weicong Chen, Yang Fan, Linyuan Gong, Yichong Leng, Renqian Luo, et al. 2019 · 2019
Closest in time.
Modeling coherence for discourse neural machine translation
Hao Xiong, Zhongjun He, Hua Wu, and Haifeng Wang. 2019 · 2019
Closest in time.
Simple and effective noisy channel modeling for neural machine translation
Kyra Yee, Yann Dauphin, and Michael Auli. 2019 · 2019
Closest in time.