V. Karpukhin, O. Levy, J. Eisenstein, and M. Ghazvininejad, “Training on synthetic noise improves robustness to natural noise in machine translation,” CoRR , vol. abs/1902.01509, 2019. [Online]. Available: http://arxiv.org/abs/1902.01509
Original
1902
Earlier work this paper cites.
Vaibhav, S. Singh, C. Stewart, and G. Neubig, “Improving robustness of machine translation with synthetic noise,” CoRR , vol. abs/1902.09508, 2019. [Online]. Available: http://arxiv.org/abs/1902.09508
Original
1902
Earlier work this paper cites.
S. Ding, A. Renduchintala, and K. Duh, “A call for prudent choice of subword merge operations,” CoRR , vol. abs/1905.10453, 2019. [Online]. Available: http://arxiv.org/abs/1905.10453
Original
1905
Earlier work this paper cites.
Q. Wang, B. Li, T. Xiao, J. Zhu, C. Li, D. F. Wong, and L. S. Chao, “Learning deep transformer models for machine translation,” CoRR , vol. abs/1906.01787, 2019. [Online]. Available: http://arxiv.org/abs/1906.01787
Original
1906
Earlier work this paper cites.
H. Ney, U. Essen, and R. Kneser, “On structuring probabilistic dependencies in stochastic language modelling,” Computer Speech and Language , vol. 8, pp. 1–38, 1994
1994
Earlier work this paper cites.
P. Koehn, “Europarl: A parallel corpus for statistical machine translation,” vol. 5, 11 2004
2004
Earlier work this paper cites.
D. Vilar, J.-T. Peter, and H. Ney, “Can we translate letters?” in Proceedings of the Second Workshop on Statistical Machine Translation , ser. StatMT ’07. Stroudsburg, PA, USA: Association for Computational Linguistics, 2007, pp. 33–39. [Online]. Available: http://dl.acm.org/citation.cfm?id=1626355.1626360
2007
Earlier work this paper cites.
J. Quionero-Candela, M. Sugiyama, A. Schwaighofer, and N. D. Lawrence, Dataset Shift in Machine Learning . The MIT Press, 2009
2009
Earlier work this paper cites.
X. Glorot and Y. Bengio, “Understanding the difficulty of training deep feedforward neural networks,” in Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics , ser. Proceedings of Machine Learning Research, Y. W. Teh and M. Titterington, Eds., vol. 9. Chia Laguna Resort, Sardinia, Italy: PMLR, 13–15 May 2010, pp. 249–256. [Online]. Available: http://proceedings.mlr.press/v9/glorot10a.html
2010
Earlier work this paper cites.
V. Nair and G. E. Hinton, “Rectified linear units improve restricted boltzmann machines,” in Proceedings of the 27th International Conference on International Conference on Machine Learning , ser. ICML’10. USA: Omnipress, 2010, pp. 807–814. [Online]. Available: http://dl.acm.org/citation.cfm?id=3104322.3104425
2010
Earlier work this paper cites.
M. Li, Y. Zhao, D. Zhang, and M. Zhou, “Adaptive development data selection for log-linear model in statistical machine translation,” in Proceedings of the 23rd International Conference on Computational Linguistics (Coling 2010) . Beijing, China: Coling 2010 Organizing Committee, Aug. 2010, pp. 662–670. [Online]. Available: https://www.aclweb.org/anthology/C10-1075
2010
Earlier work this paper cites.
T. Mikolov, I. Sutskever, A. Deoras, H.-S. Le, S. Kombrink, and J. Cernocký, “Subword language modeling with neural networks,” 2011
2011
Earlier work this paper cites.
A. Axelrod, X. He, and J. Gao, “Domain adaptation via pseudo in-domain data selection,” in Proceedings of the Conference on Empirical Methods in Natural Language Processing , ser. EMNLP ’11. Stroudsburg, PA, USA: Association for Computational Linguistics, 2011, pp. 355–362. [Online]. Available: http://dl.acm.org/citation.cfm?id=2145432.2145474
2011
Earlier work this paper cites.
T. Zesch, “Measuring contextual fitness using error contexts extracted from the Wikipedia revision history,” in Proceedings of the 13th Conference of the European Chapter of the Association for Computational Linguistics . Avignon, France: Association for Computational Linguistics, Apr. 2012, pp. 529–538. [Online]. Available: https://www.aclweb.org/anthology/E12-1054
2012
Earlier work this paper cites.
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean, “Distributed representations of words and phrases and their compositionality,” in Advances in Neural Information Processing Systems 26 , C. J. C. Burges, L. Bottou, M. Welling, Z. Ghahramani, and K. Q. Weinberger, Eds. Curran Associates, Inc., 2013, pp. 3111–3119
2013
Earlier work this paper cites.
G. Neubig, T. Watanabe, S. Mori, and T. Kawahara, “Substring-based machine translation,” Machine Translation , vol. 27, no. 2, pp. 139–166, June 2013. [Online]. Available: http://dx.doi.org/10.1007/s10590-013-9136-6
2013
Earlier work this paper cites.
J. R. Smith, H. Saint-Amand, M. Plamada, P. Koehn, C. Callison-Burch, and A. Lopez, “Dirt cheap web-scale parallel text from the common crawl,” vol. 1, 07 2013
2013
Earlier work this paper cites.
K. Wisniewski, K. Schöne, L. Nicolas, C. Vettori, A. Boyd, D. Meurers, A. Abel, and J. Hana, “Merlin: An online trilingual learner corpus empirically grounding the european reference levels in authentic learner data,” 10 2013
2013
Earlier work this paper cites.