Fetching the paper…
Reading the bibliography…
The well-known Gumbel-Max trick for sampling from a categorical distribution can be extended to sample $k$ elements without replacement.
A law of comparative judgement
Thurstone, L. L · 1927
Earlier work this paper cites.
Statistical theory of extreme values and some practical applications: a series of lectures
Gumbel, E. J · 1954
Earlier work this paper cites.
Individual choice behavior
Luce, R. D · 1959
Earlier work this paper cites.
The analysis of permutations
Plackett, R. L · 1975
Earlier work this paper cites.
The relationship between luce’s choice axiom, thurstone’s theory of comparative judgment, and the double exponential distribution
Yellott, J. I · 1977
Earlier work this paper cites.
Advances in importance sampling
Hesterberg, T. C · 1988
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
Weighted random sampling with a reservoir
Efraimidis, P. S. and Spirakis, P. G · 2006
Earlier work this paper cites.
Priority sampling for estimation of arbitrary subset sums
Duffield, N., Lund, C., and Thorup, M · 2007
Earlier work this paper cites.
Search-based structured prediction
Daumé, H., Langford, J., and Marcu, D · 2009
Earlier work this paper cites.
Perturb-and-map random fields: Using discrete optimization to learn and sample from energy models
Papandreou, G. and Yuille, A. L · 2011
Earlier work this paper cites.
On the partition function and random maximum a-posteriori perturbations
Hazan, T. and Jaakkola, T · 2012
Earlier work this paper cites.
Accurately computing log ( 1 − exp ( − | a | ) ) \log(1-\exp(-|a|)) assessed by the Rmpfr package, 2012
Mächler, M · 2012
Earlier work this paper cites.
Randomized optimum models for structured prediction
Tarlow, D., Adams, R., and Zemel, R · 2012
Earlier work this paper cites.
Embed and project: Discrete sampling with universal hashing
Ermon, S., Gomes, C. P., Sabharwal, A., and Selman, B · 2013
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
Graves, A., Mohamed, A.-r., and Hinton, G · 2013
Cited alongside, same era.
A* sampling
Maddison, C. J., Tarlow, D., and Minka, T · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., and Le, Q. V · 2014
Cited alongside, same era.
Gumbel-max trick and weighted reservoir sampling, 2014
Vieira, T · 2014
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2015
Cited alongside, same era.
Structured training for neural network transition-based parsing
Weiss, D., Alberti, C., Collins, M., and Petrov, S · 2015
Cited alongside, same era.
Lost relatives of the gumbel trick
Balog, M., Tripuraneni, N., Ghahramani, Z., and Weller, A · 2017
Later among the works it cites.
Convolutional sequence to sequence learning
Gehring, J., Auli, M., Grangier, D., Yarats, D., and Dauphin, Y. N · 2017
Later among the works it cites.
Automatic differentiation in PyTorch
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A · 2017
Later among the works it cites.
Generating high-quality and informative conversation responses with sequence-to-sequence models
Shao, Y., Gouws, S., Britz, D., Goldie, A., Strope, B., and Kurzweil, R · 2017
Later among the works it cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Later among the works it cites.
Estimating means in a finite universe, 2017
Vieira, T · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scalable discrete sampling as a multi-armed bandit problem
Chen, Y. and Ghahramani, Z · 2016
Cited alongside, same era.
Categorical reparameterization with gumbel-softmax
Jang, E., Gu, S., and Poole, B · 2016
Cited alongside, same era.
A simple, fast diverse decoding algorithm for neural generation
Li, J., Monroe, W., and Jurafsky, D · 2016
Cited alongside, same era.
The concrete distribution: A continuous relaxation of discrete random variables
Maddison, C. J., Mnih, A., and Teh, Y. W · 2016
Cited alongside, same era.
Sequence level training with recurrent neural networks
Ranzato, M., Chopra, S., Auli, M., and Zaremba, W · 2016
Cited alongside, same era.
Sequence-to-sequence learning as beam-search optimization
Wiseman, S. and Rush, A. M · 2016
Cited alongside, same era.
Understanding back-translation at scale
Edunov, S., Ott, M., Auli, M., and Grangier, D · 2018
Later among the works it cites.
Classical structured prediction losses for sequence to sequence learning
Edunov, S., Ott, M., Auli, M., Grangier, D., et al · 2018
Later among the works it cites.
A continuous relaxation of beam search for end-to-end training of neural sequence models
Goyal, K., Neubig, G., Dyer, C., and Berg-Kirkpatrick, T · 2018
Later among the works it cites.
Neural machine translation with Gumbel-greedy decoding
Gu, J., Im, D. J., and Li, V. O · 2018
Later among the works it cites.
Learning beam search policies via imitation learning
Negrinho, R., Gormley, M., and Gordon, G. J · 2018
Later among the works it cites.
Diverse beam search for improved description of complex scenes
Vijayakumar, A. K., Cogswell, M., Selvaraju, R. R., Sun, Q., Lee, S., Crandall, D. J., and Batra, D · 2018
Later among the works it cites.
Stochastic optimization of sorting networks via continuous relaxations
Grover, A., Wang, E., Zweig, A., and Ermon, S · 2019
Closest in time.
fairseq: A fast, extensible toolkit for sequence modeling
Ott, M., Edunov, S., Baevski, A., Fan, A., Gross, S., Ng, N., Grangier, D., and Auli, M · 2019
Closest in time.