Fetching the paper…
Reading the bibliography…
Beam search is the default decoding strategy for many sequence generation tasks in NLP.
A generalization of sampling without replacement from a finite universe
D. G. Horvitz and D. J. Thompson. 1952 · 1952
Earlier work this paper cites.
Asymptotic theory of rejective sampling with varying probabilities from a finite population
Jaroslav Hájek. 1964 · 1964
Earlier work this paper cites.
Speech understanding systems: A summary of results of the five-year research effort at Carnegie Mellon University
Raj Reddy. 1977 · 1977
Earlier work this paper cites.
Sampling from a Finite Population
J. Hájek. 1981 · 1981
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams. 1992 · 1992
Earlier work this paper cites.
Algorithms to find exact inclusion probabilities for conditional Poisson sampling and Pareto π \pi ps sampling designs
Nibia Aires. 1999 · 1999
Earlier work this paper cites.
BLEU: A method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Minimum Bayes-risk decoding for statistical machine translation
Shankar Kumar and William Byrne. 2004 · 2004
Earlier work this paper cites.
Pareto sampling versus Sampford and conditional Poisson sampling
Lennart Bondesson, Imbi Traat, and Anders Lundqvist. 2006 · 2006
Earlier work this paper cites.
Automatic Differentiation: Applications, Theory, and Implementations (Lecture Notes in Computational Science and Engineering)
Martin Bücker, George Corliss, Paul Hovland, Uwe Naumann, and Boyana Norris. 2006 · 2006
Earlier work this paper cites.
Sampling Algorithms
Yves Tillé. 2006 · 2006
Earlier work this paper cites.
Non-rejective implementations of the Sampford sampling design
Anton Grafström. 2009 · 2009
Earlier work this paper cites.
Determinantal point processes for machine learning
Alex Kulesza and Ben Taskar. 2012 · 2012
Cited alongside, same era.
Monte Carlo theory, methods and examples
Art B. Owen. 2013 · 2013
Cited alongside, same era.
Findings of the 2014 workshop on statistical machine translation
Ondřej Bojar, Christian Buck, Christian Federmann, Barry Haddow, Philipp Koehn, Johannes Leveling, Christof Monz, Pavel Pecina, Matt Post, Herve Saint-Amand, Radu Soricut, Lucia Specia, and Aleš Tamchyna. 2014 · 2014
Cited alongside, same era.
Gumbel-max trick and weighted reservoir sampling
Tim Vieira. 2014 · 2014
Cited alongside, same era.
Mathematical Statistics: Basic ideas and Selected Topics
Peter J. Bickel and Kjell A. Docksum. 2015 · 2015
Cited alongside, same era.
Diverse beam search for improved description of complex scenes
Ashwin Vijayakumar, Michael Cogswell, Ramprasaath Selvaraju, Qing Sun, Stefan Lee, David Crandall, and Dhruv Batra. 2018 · 2018
Later among the works it cites.
Stochastic beams and where to find them: The Gumbel-top- k k trick for sampling sequences without replacement
Wouter Kool, Herke Van Hoof, and Max Welling. 2019 · 2019
Later among the works it cites.
Rao-Blackwellized stochastic gradients for discrete distributions
Runjing Liu, Jeffrey Regier, Nilesh Tripuraneni, Michael Jordan, and Jon Mcauliffe. 2019 · 2019
Later among the works it cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Later among the works it cites.
XLNet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba. 2016 · 2016
Cited alongside, same era.
Simulation and the Monte Carlo Method , 3rd edition
Reuven Y. Rubinstein and Dirk P. Kroese. 2016 · 2016
Cited alongside, same era.
Minimum risk training for neural machine translation
Shiqi Shen, Yong Cheng, Zhongjun He, Wei He, Hua Wu, Maosong Sun, and Yang Liu. 2016 · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna. 2016 · 2016
Cited alongside, same era.
Complementary sum sampling for likelihood approximation in large scale classification
Aleksandar Botev, Bowen Zheng, and David Barber. 2017 · 2017
Cited alongside, same era.
Multiresolution recurrent neural networks: An application to dialogue response generation
Iulian Vlad Serban, Tim Klinger, Gerald Tesauro, Kartik Talamadupula, Bowen Zhou, Yoshua Bengio, and Aaron Courville. 2017 · 2017
Cited alongside, same era.
Understanding back-translation at scale
Sergey Edunov, Myle Ott, Michael Auli, and David Grangier. 2018 · 2018
Cited alongside, same era.
Is MAP decoding all you need? The inadequacy of the mode in neural machine translation
Bryan Eikema and Wilker Aziz. 2020 · 2020
Later among the works it cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020 · 2020
Later among the works it cites.
Estimating gradients for discrete random variables by sampling without replacement
Wouter Kool, Herke van Hoof, and Max Welling. 2020 · 2020
Later among the works it cites.
If beam search is the answer, what was the question?
Clara Meister, Ryan Cotterell, and Tim Vieira. 2020a · 2020
Later among the works it cites.
Incremental sampling without replacement for sequence models
Kensen Shi, David Bieber, and Charles Sutton. 2020 · 2020
Later among the works it cites.
Determinantal beam search
Clara Meister, Martina Forster, and Ryan Cotterell. 2021 · 2021
Closest in time.