Fetching the paper…
Reading the bibliography…
Prior work has explored directly regularizing the output distributions of probabilistic models to alleviate peaky (i.e.
Dual skew divergence loss for neural machine translation
Fengshun Xiao, Yingting Wu, Hai Zhao, Rui Wang, and Shu Jiang. 2019 · 1908
Earlier work this paper cites.
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2019 · 1910
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu. 2016 · 1937
Earlier work this paper cites.
Optimization by simulated annealing
Scott Kirkpatrick, C. Daniel Gelatt, and Mario P. Vecchi. 1983 · 1983
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
Ronald Williams and Jing Peng. 1991 · 1991
Earlier work this paper cites.
Nonparametric entropy estimation: An overview
J. Beirlant, E. Dudewicz, L. Gyor, and E. C. Meulen. 1997 · 1997
Earlier work this paper cites.
Comparison of permutation methods for the partial correlation and partial mantel tests
Pierre Legendre. 2000 · 2000
Earlier work this paper cites.
The Elements of Statistical Learning
Trevor Hastie, Robert Tibshirani, and Jerome Friedman. 2001 · 2001
Earlier work this paper cites.
BLEU: A method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Semi-supervised learning by entropy minimization
Yves Grandvalet and Yoshua Bengio. 2005 · 2005
Earlier work this paper cites.
Bootstrapping feature-rich dependency parsers with entropic priors
David A. Smith and Jason Eisner. 2007 · 2007
Earlier work this paper cites.
Structured sparsity in structured prediction
André F. T. Martins, Noah A. Smith, Pedro M. Q. Aguiar, and Mário A. T. Figueiredo. 2011 · 2011
Earlier work this paper cites.
The Burbea-Rao and Bhattacharyya centroids
Frank Nielsen and Sylvain Boltz. 2011 · 2011
Earlier work this paper cites.
WIT 3 : Web inventory of transcribed and translated talks
Mauro Cettolo, Christian Girardi, and Marcello Federico. 2012 · 2012
Cited alongside, same era.
Practical Bayesian optimization of machine learning algorithms
Jasper Snoek, Hugo Larochelle, and Ryan P. Adams. 2012 · 2012
Cited alongside, same era.
Comparison of co-expression measures: mutual information, correlation, and model based indices
Lin Song, Peter Langfelder, and Steve Horvath. 2012 · 2012
Cited alongside, same era.
Findings of the 2014 workshop on statistical machine translation
Ondrej Bojar, Christian Buck, Christian Federmann, Barry Haddow, Philipp Koehn, Johannes Leveling, Christof Monz, Pavel Pecina, Matt Post, Herve Saint-Amand, Radu Soricut, Lucia Specia, and Aleš Tamchyna. 2014 · 2014
Cited alongside, same era.
Scheduled sampling for sequence prediction with recurrent neural networks
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer. 2015 · 2015
Cited alongside, same era.
The best of both worlds: Combining recent advances in neural machine translation
Mia Xu Chen, Orhan Firat, Ankur Bapna, Melvin Johnson, Wolfgang Macherey, George Foster, Llion Jones, Mike Schuster, Noam Shazeer, Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Zhifeng Chen, Yonghui Wu, and Macduff Hughes. 2018 · 2018
Later among the works it cites.
The multitarget TED talks task
Kevin Duh. 2018 · 2018
Later among the works it cites.
Understanding back-translation at scale
Sergey Edunov, Myle Ott, Michael Auli, and David Grangier. 2018 · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. 2018 · 2018
Later among the works it cites.
SparseMAP: Differentiable sparse structured inference
Vlad Niculae, André Martins, Mathieu Blondel, and Claire Cardie. 2018 · 2018
Later among the works it cites.
Analyzing uncertainty in neural machine translation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015 · 2015
Cited alongside, same era.
Taming the noise in reinforcement learning via soft updates
Roy Fox, Ari Pakman, and Naftali Tishby. 2016 · 2016
Cited alongside, same era.
From softmax to sparsemax: A sparse model of attention and multi-label classification
André Martins and Ramon Astudillo. 2016 · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. 2015 · 2016
Cited alongside, same era.
Towards better decoding and language model integration in sequence to sequence models
Jan Chorowski and Navdeep Jaitly. 2017 · 2017
Cited alongside, same era.
Convolutional sequence to sequence learning
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N. Dauphin. 2017 · 2017
Cited alongside, same era.
Myle Ott, Michael Auli, David Grangier, and Marc’Aurelio Ranzato. 2018 · 2018
Later among the works it cites.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Later among the works it cites.
Tensor2Tensor for neural machine translation
Ashish Vaswani, Samy Bengio, Eugene Brevdo, Francois Chollet, Aidan N. Gomez, Stephan Gouws, Llion Jones, Łukasz Kaiser, Nal Kalchbrenner, Niki Parmar, Ryan Sepassi, Noam Shazeer, and Jakob Uszkoreit. 2018 · 2018
Later among the works it cites.
Von Mises–Fisher loss for training sequence to sequence models with continuous outputs
Sachin Kumar and Yulia Tsvetkov. 2019 · 2019
Later among the works it cites.
When does label smoothing help?
Rafael Müller, Simon Kornblith, and Geoffrey E. Hinton. 2019 · 2019
Later among the works it cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Later among the works it cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Maxwell Forbes, and Yejin Choi. 2020 · 2020
Closest in time.
Data-dependent Gaussian prior objective for language generation
Zuchao Li, Rui Wang, Kehai Chen, Masso Utiyama, Eiichiro Sumita, Zhuosheng Zhang, and Hai Zhao. 2020 · 2020
Closest in time.