Fetching the paper…
Reading the bibliography…
In sequence prediction tasks like neural machine translation, training with cross-entropy loss often leads to models that overgeneralize and plunge into local optima.
S. Kullback and R. A. Leibler, “On information and sufficiency,” Ann. Math. Statist. , vol. 22, no. 1, pp. 79–86, 03 1951
1951
Earlier work this paper cites.
S. F. Chen and J. Goodman, “An empirical study of smoothing techniques for language modeling,” in 34th Annual Meeting of the Association for Computational Linguistics . Santa Cruz, California, USA: Association for Computational Linguistics, Jun. 1996, pp. 310–318. [Online]. Available: https://www.aclweb.org/anthology/P96-1041
1996
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Comput. , vol. 9, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
L. Lee, “Measures of distributional similarity,” in Proceedings of the 37th Annual Meeting of the Association for Computational Linguistics , College Park, Maryland, USA, June 1999, pp. 25–32
1999
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Proceedings of 40th Annual Meeting of the Association for Computational Linguistics , Philadelphia, Pennsylvania, USA, July 2002, pp. 311–318
2002
Earlier work this paper cites.
A. Stolcke, “Srilm — an extensible language modeling toolkit,” in International Conference on Spoken Language Processing , 2002, pp. 901–904
2002
Earlier work this paper cites.
P. Koehn, F. J. Och, and D. Marcu, “Statistical phrase-based translation,” in Proceedings of the 2003 Conference of the North American Chapter of the Association for Computational Linguistics on Human Language Technology - Volume 1 , Stroudsburg, PA, USA, 2003, pp. 48–54
2003
Earlier work this paper cites.
F. J. Och, “Minimum error rate training in statistical machine translation,” in Proceedings of the 41st Annual Meeting of the Association for Computational Linguistics , Sapporo, Japan, July 2003, pp. 160–167
2003
Earlier work this paper cites.
M. Collins, P. Koehn, and I. Kučerová, “Clause restructuring for statistical machine translation,” in Proceedings of the 43rd annual meeting on association for computational linguistics . Association for Computational Linguistics, 2005, pp. 531–540
2005
Earlier work this paper cites.
K. J. Åström, T. Hägglund, and K. J. Astrom, Advanced PID control . ISA-The Instrumentation, Systems, and Automation Society Research Triangle Park, 2006, vol. 461
2006
Earlier work this paper cites.
Koehn, Philipp, Hoang, Hieu, Alexandra, CallisonBurch, Chris, Federico, and Marcello, “Moses: open source toolkit for statistical machine translation,” in Proceedings of the 45th Annual Meeting of the ACL on Interactive Poster and Demonstration Sessions , 2007, pp. 177–180
2007
Earlier work this paper cites.
N. Kalchbrenner and P. Blunsom, “Recurrent continuous translation models,” in Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing , Seattle, Washington, USA, October 2013, pp. 1700–1709
2013
Earlier work this paper cites.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” in Advances in Neural Information Processing Systems , 2014, pp. 3104–3112
2014
Cited alongside, same era.
K. Cho, B. van Merrienboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase representations using rnn encoder–decoder for statistical machine translation,” in Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing , Doha, Qatar, October 2014, pp. 1724–1734
2014
Cited alongside, same era.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” Advances in Neural Information Processing Systems , vol. 27, pp. 2672–2680, 2014
2014
Cited alongside, same era.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in Proceedings of 3rd International Conference on Learning Representations , 2015
2015
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems , 2017, pp. 6000–6010
2017
Later among the works it cites.
D. Bahdanau, P. Brakel, K. Xu, A. Goyal, R. Lowe, J. Pineau, A. C. Courville, and Y. Bengio, “An actor-critic algorithm for sequence prediction,” in Proceedings of 5rd International Conference on Learning Representations , 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
P. Koehn and R. Knowles, “Six challenges for neural machine translation,” in Proceedings of the First Workshop on Neural Machine Translation . Vancouver: Association for Computational Linguistics, Aug. 2017, pp. 28–39. [Online]. Available: https://www.aclweb.org/anthology/W17-3204
2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
T. Luong, H. Pham, and C. D. Manning, “Effective approaches to attention-based neural machine translation,” in Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing , Lisbon, Portugal, September 2015, pp. 1412–1421
2015
Cited alongside, same era.
S. Bengio, O. Vinyals, N. Jaitly, and N. Shazeer, “Scheduled sampling for sequence prediction with recurrent neural networks,” in Advances in Neural Information Processing Systems , 2015, pp. 1171–1179
2015
Cited alongside, same era.
2015
Cited alongside, same era.
S. Jean, K. Cho, R. Memisevic, and Y. Bengio, “On using very large target vocabulary for neural machine translation,” in Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing , Beijing, China, July 2015, pp. 1–10
2015
Cited alongside, same era.
M. Ranzato, S. Chopra, M. Auli, and W. Zaremba, “Sequence level training with recurrent neural networks,” in Proceedings of 4rd International Conference on Learning Representations , 2016
2016
Cited alongside, same era.
S. Wiseman and A. M. Rush, “Sequence-to-sequence learning as beam-search optimization,” in Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , Austin, Texas, November 2016, pp. 1296–1306. [Online]. Available: https://aclweb.org/anthology/D16-1137
2016
Cited alongside, same era.
S. Shen, Y. Cheng, Z. He, W. He, H. Wu, M. Sun, and Y. Liu, “Minimum risk training for neural machine translation,” in Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics , Berlin, Germany, August 2016, pp. 1683–1692. [Online]. Available: http://www.aclweb.org/anthology/P16-1159
2016
Cited alongside, same era.
R. Sennrich, B. Haddow, and A. Birch, “Neural machine translation of rare words with subword units,” in Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics , Berlin, Germany, August 2016, pp. 1715–1725
2016
Cited alongside, same era.
Later among the works it cites.
M. Ott, M. Auli, D. Grangier, and M. Ranzato, “Analyzing uncertainty in neural machine translation,” in International Conference on Machine Learning . PMLR, 2018, pp. 3956–3965
2018
Later among the works it cites.
S. Edunov, M. Ott, M. Auli, D. Grangier, and M. Ranzato, “Classical structured prediction losses for sequence to sequence learning,” in Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers) . Association for Computational Linguistics, 2018, pp. 355–364. [Online]. Available: http://aclweb.org/anthology/N18-1033
2018
Later among the works it cites.
E. Cohen and C. Beck, “Empirical analysis of beam search performance degradation in neural sequence models,” in International Conference on Machine Learning . PMLR, 2019, pp. 1290–1299
2019
Closest in time.
Z. Zhang, S. Wu, S. Liu, M. Li, M. Zhou, and T. Xu, “Regularizing neural machine translation by target-bidirectional agreement,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 33, no. 01, 2019, pp. 443–450
2019
Closest in time.
Z. Li, R. Wang, K. Chen, M. Utiyama, E. Sumita, Z. Zhang, and H. Zhao, “Data-dependent gaussian prior objective for language generation,” in International Conference on Learning Representations , 2020. [Online]. Available: https://openreview.net/forum?id=S1efxTVYDr
2020
Closest in time.
H. Shao, S. Yao, D. Sun, A. Zhang, S. Liu, D. Liu, J. Wang, and T. Abdelzaher, “Controlvae: Controllable variational autoencoder,” in International Conference on Machine Learning . PMLR, 2020, pp. 8655–8664
2020
Closest in time.
H. Jiang, P. He, W. Chen, X. Liu, J. Gao, and T. Zhao, “SMART: Robust and efficient fine-tuning for pre-trained natural language models through principled regularized optimization,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . Online: Association for Computational Linguistics, Jul. 2020, pp. 2177–2190. [Online]. Available: https://www.aclweb.org/anthology/2020.acl-main.197
2020
Closest in time.
A. Aghajanyan, A. Shrivastava, A. Gupta, N. Goyal, L. Zettlemoyer, and S. Gupta, “Better fine-tuning by reducing representational collapse,” in International Conference on Learning Representations , 2021. [Online]. Available: https://openreview.net/forum?id=OQ08SN70M1V
2021
Closest in time.