Fetching the paper…
Reading the bibliography…
Beam search is the go-to method for decoding auto-regressive machine translation models.
Statistical theory of extreme values and some practical applications: a series of lectures, 1954
Gumbel, E. J · 1954
Earlier work this paper cites.
Speech understanding systems: A summary of results of the five-year research effort
Reddy, R · 1977
Earlier work this paper cites.
A learning algorithm for boltzmann machines
Ackley, D. H., Hinton, G. E., and Sejnowski, T. J · 1985
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
Efficient selectivity and backup operators in monte-carlo tree search
Coulom, R · 2006
Earlier work this paper cites.
Bandit based monte-carlo planning
Kocsis, L. and Szepesvári, C · 2006
Earlier work this paper cites.
A survey of monte carlo tree search methods
Browne, C., Powley, E., Whitehouse, D., Lucas, S., Cowling, P., Rohlfshagen, P., Tavener, S., Perez Liebana, D., Samothrakis, S., and Colton, S · 2012
Earlier work this paper cites.
Meteor universal: Language specific translation evaluation for any target language
Denkowski, M. and Lavie, A · 2014
Earlier work this paper cites.
Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks
Bengio, S., Vinyals, O., Jaitly, N., and Shazeer, N · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Reward Augmented Maximum Likelihood for Neural Structured Prediction
Norouzi, M., Bengio, S., Chen, Z., Jaitly, N., Schuster, M., Wu, Y., and Schuurmans, D · 2016
Earlier work this paper cites.
Sequence Level Training with Recurrent Neural Networks
Ranzato, M., Chopra, S., Auli, M., and Zaremba, W · 2016
Earlier work this paper cites.
Minimum Risk Training for Neural Machine Translation
Shen, S., Cheng, Y., He, Z., He, W., Wu, H., Sun, M., and Liu, Y · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Wu, Y., Schuster, M., Chen, Z., Le, Q. V., Norouzi, M., Macherey, W., Krikun, M., Cao, Y., Gao, Q., Macherey, K., Klingner, J., Shah, A., Johnson, M., Liu, X., Kaiser, L., Gouws, S., Kato, Y., Kudo, T., Kazawa, H., Stevens, K., Kurian, G., Patil, N., Wang, W., Young, C., Smith, J., Riesa, J., Rudnick, A., Vinyals, O., Corrado, G., Hughes, M., and Dean, J · 2016
Earlier work this paper cites.
An Actor-Critic Algorithm for Sequence Prediction
Bahdanau, D., Brakel, P., Xu, K., Goyal, A., Lowe, R., Pineau, J., Courville, A., and Bengio, Y · 2017
Cited alongside, same era.
Decoding with value networks for neural machine translation
He, D., Lu, H., Xia, Y., Qin, T., Wang, L., and Liu, T.-Y · 2017
Cited alongside, same era.
Pushing the limits of translation quality estimation
Martins, A. F. T., Junczys-Dowmunt, M., Kepler, F. N., Astudillo, R., Hokamp, C., and Grundkiewicz, R · 2017
Cited alongside, same era.
Deep reinforcement learning-based image captioning with embedding reward
Ren, Z., Wang, X., Zhang, N., Lv, X., and Li, L · 2017
Cited alongside, same era.
Mastering the game of Go without human knowledge
Silver, D., Schrittwieser Julian, Simonyan Karen, Antonoglou Ioannis, Huang Aja, Guez Arthur, Hubert Thomas, Baker Lucas, Lai Matthew, Bolton Adrian, Chen Yutian, Lillicrap Timothy, Hui Fan, Sifre Laurent, van den Driessche George, Graepel Thore, and Hassabis Demis · 2017
Cited alongside, same era.
The curious case of neural text degeneration
Holtzman, A., Buys, J., Du, L., Forbes, M., and Choi, Y · 2019
Later among the works it cites.
Stochastic beams and where to find them: The Gumbel-top-k trick for sampling sequences without replacement
Kool, W., Van Hoof, H., and Welling, M · 2019
Later among the works it cites.
Fast transformer decoding: One write-head is all you need
Shazeer, N · 2019
Later among the works it cites.
On NMT search errors and model errors: Cat got your tongue?
Stahlberg, F. and Byrne, B · 2019
Later among the works it cites.
Non-monotonic sequential text generation
Welleck, S., Brantley, K., Daumé, III, H., and Cho, K · 2019
Later among the works it cites.
BERTScore: Evaluating text generation with BERT
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., and Artzi, Y · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Cited alongside, same era.
Classical structured prediction losses for sequence to sequence learning
Edunov, S., Ott, M., Auli, M., Grangier, D., and Ranzato, M · 2018
Cited alongside, same era.
Hierarchical neural story generation
Fan, A., Lewis, M., and Dauphin, Y · 2018
Cited alongside, same era.
Non-autoregressive neural machine translation
Gu, J., Bradbury, J., Xiong, C., Li, V. O., , and Socher, R · 2018
Cited alongside, same era.
Quality estimation for machine translation
Specia, L., Scarton, C., and Paetzold, G. H · 2018
Cited alongside, same era.
On the weaknesses of reinforcement learning for neural machine translation
Choshen, L., Fox, L., Aizenbud, Z., and Abend, O · 2019
Cited alongside, same era.
Empirical analysis of beam search performance degradation in neural sequence models
Cohen, E. and Beck, J. C · 2019
Cited alongside, same era.
Later among the works it cites.
Is map decoding all you need? the inadequacy of the mode in neural machine translation
Eikema, B. and Aziz, W · 2020
Later among the works it cites.
Array programming with NumPy
Harris, C. R., Millman, K. J., van der Walt, S. J., Gommers, R., Virtanen, P., Cournapeau, D., Wieser, E., Taylor, J., Berg, S., Smith, N. J., Kern, R., Picus, M., Hoyer, S., van Kerkwijk, M. H., Brett, M., Haldane, A., del R’ıo, J. F., Wiebe, M., Peterson, P., G’erard-Marchant, P., Sheppard, K., Reddy, T., Weckesser, W., Abbasi, H., Gohlke, C., and Oliphant, T. E · 2020
Later among the works it cites.
If beam search is the answer, what was the question?
Meister, C., Cotterell, R., and Vieira, T · 2020
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., Lillicrap, T., and Silver, D · 2020
Later among the works it cites.
Investigating the decoders of maximum likelihood sequence models: A look-ahead approach
Wang, Y.-S., Kuo, Y.-L., and Katz, B · 2020
Later among the works it cites.
Consistency of a recurrent language model with respect to incomplete decoding
Welleck, S., Kulikov, I., Kim, J., Pang, R. Y., and Cho, K · 2020
Later among the works it cites.
The DeepMind Chinese–English document translation system at WMT2020
Yu, L., Sartran, L., Huang, P.-S., Stokowiec, W., Donato, D., Srinivasan, S., Andreev, A., Ling, W., Mokra, S., Dal Lago, A., Doron, Y., Young, S., Blunsom, P., and Dyer, C · 2020
Later among the works it cites.
Energy-based reranking: Improving neural machine translation using energy-based models
Bhattacharyya, S., Rooshenas, A., Naskar, S., Sun, S., Iyyer, M., and McCallum, A · 2021
Closest in time.