Fetching the paper…
Reading the bibliography…
The Gumbel-Max trick is the basis of many relaxed gradient estimators.
A law of comparative judgment
Louis L Thurstone · 1927
Earlier work this paper cites.
The hungarian method for the assignment problem
Harold W Kuhn · 1955
Earlier work this paper cites.
On the shortest spanning subtree of a graph and the traveling salesman problem
Joseph B Kruskal · 1956
Earlier work this paper cites.
Individual Choice Behavior: A Theoretical Analysis
R Duncan Luce · 1959
Earlier work this paper cites.
On the shortest arborescence of a directed graph
Y.J. Chu and T. H. Liu · 1965
Earlier work this paper cites.
Optimum branchings”
Jack Edmonds · 1967
Earlier work this paper cites.
Concerning nonnegative matrices and doubly stochastic matrices
Richard Sinkhorn and Paul Knopp · 1967
Earlier work this paper cites.
Convex Analysis
R. Tyrrell Rockafellar · 1970
Earlier work this paper cites.
The analysis of permutations
Robin L Plackett · 1975
Earlier work this paper cites.
Finding the nearest point in a polytope
Philip Wolfe · 1976
Earlier work this paper cites.
Graph Theory
William T. Tutte · 1984
Earlier work this paper cites.
Likelihood ratio gradient estimation for stochastic systems
Peter W Glynn · 1990
Earlier work this paper cites.
Graph drawing by force-directed placement
Thomas MJ Fruchterman and Edward M Reingold · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Introduction to linear optimization
Dimitris Bertsimas and John N Tsitsiklis · 1997
Earlier work this paper cites.
Second-order convex analysis
R Tyrrell Rockafellar · 1999
Earlier work this paper cites.
Combinatorial optimization: polyhedra and efficiency
Alexander Schrijver · 2003
Earlier work this paper cites.
Algorithm Design
Jon Kleinberg and Éva Tardos · 2006
Earlier work this paper cites.
Convergent tree-reweighted message passing for energy minimization
Vladimir Kolmogorov · 2006
Earlier work this paper cites.
Stochastic simulation: algorithms and analysis
Søren Asmussen and Peter W Glynn · 2007
Earlier work this paper cites.
Structured prediction models via the matrix-tree theorem
Terry Koo, Amir Globerson, Xavier Carreras, and Michael Collins · 2007
Earlier work this paper cites.
Efficient projections onto the l 1-ball for learning in high dimensions
John Duchi, Shai Shalev-Shwartz, Yoram Singer, and Tushar Chandra · 2008
Earlier work this paper cites.
Graphical models, exponential families, and variational inference
Martin J Wainwright and Michael I Jordan · 2008
Earlier work this paper cites.
Probabilistic graphical models: principles and techniques
Daphne Koller and Nir Friedman · 2009
Earlier work this paper cites.
Efficient euclidean projections in linear time
Jun Liu and Jieping Ye · 2009
Earlier work this paper cites.
Implicit differentiation by perturbation
Justin Domke · 2010
Earlier work this paper cites.
Ranking via sinkhorn propagation
Ryan Prescott Adams and Richard S Zemel · 2011
Earlier work this paper cites.
Perturb-and-MAP Random Fields: Using Discrete Optimization to Learn and Sample from Energy Models
G. Papandreou and A. Yuille · 2011
Earlier work this paper cites.
Sum-product networks: A new deep architecture
Hoifung Poon and Pedro Domingos · 2011
Earlier work this paper cites.
Learning message-passing inference machines for structured prediction
Stephane Ross, Daniel Munoz, Martial Hebert, and J. Andrew Bagnell · 2011
Cited alongside, same era.
On the partition function and random maximum a-posteriori perturbations
Tamir Hazan and Tommi Jaakkola · 2012
Cited alongside, same era.
Learning attitudes and attributes from multi-aspect reviews
Julian McAuley, Jure Leskovec, and Dan Jurafsky · 2012
Cited alongside, same era.
Cardinality restricted boltzmann machines
Kevin Swersky, Ilya Sutskever, Daniel Tarlow, Richard S Zemel, Russ R Salakhutdinov, and Ryan P Adams · 2012
Cited alongside, same era.
Randomized optimum models for structured prediction
Daniel Tarlow, Ryan Adams, and Richard Zemel · 2012
Cited alongside, same era.
Fast exact inference for recursive cardinality models
Daniel Tarlow, Kevin Swersky, Richard S Zemel, Ryan P Adams, and Brendan J Frey · 2012
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Later among the works it cites.
Rebar: Low-variance, unbiased gradient estimates for discrete latent variable models
George Tucker, Andriy Mnih, Chris J Maddison, John Lawson, and Jascha Sohl-Dickstein · 2017
Later among the works it cites.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, and Skye Wanderman-Milne · 2018
Later among the works it cites.
Learning to explain: An information-theoretic perspective on model interpretation
Jianbo Chen, Le Song, Martin Wainwright, and Michael Jordan · 2018
Later among the works it cites.
Backpropagation through the void: Optimizing control variates for black-box gradient estimation
Will Grathwohl, Dami Choi, Yuhuai Wu, Geoff Roeder, and David Duvenaud · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning graphical model parameters with approximate marginal inference
Justin Domke · 2013
Cited alongside, same era.
On Sampling from the Gibbs Distribution with Random Maximum A-Posteriori Perturbations
Tamir Hazan, Subhransu Maji, and Tommi Jaakkola · 2013
Cited alongside, same era.
Tighter linear program relaxations for high order graphical models
Elad Mezuman, Daniel Tarlow, Amir Globerson, and Yair Weiss · 2013
Cited alongside, same era.
Learning with maximum a-posteriori perturbation models
Andreea Gane, Tamir Hazan, and Tommi Jaakkola · 2014
Cited alongside, same era.
Alex Graves, Greg Wayne, and Ivo Danihelka · 2014
Cited alongside, same era.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2014
Cited alongside, same era.
Neural relational inference for interacting systems
Thomas Kipf, Ethan Fetaya, Kuan-Chieh Wang, Max Welling, and Richard Zemel · 2018
Later among the works it cites.
Reparameterization gradient for non-differentiable models
Wonyeol Lee, Hangyeol Yu, and Hongseok Yang · 2018
Later among the works it cites.
Learning latent permutations with gumbel-sinkhorn networks
Gonzalo Mena, David Belanger, Scott Linderman, and Jasper Snoek · 2018
Later among the works it cites.
Listops: A diagnostic dataset for latent tree learning
Nikita Nangia and Samuel R Bowman · 2018
Later among the works it cites.
Sparsemap: Differentiable sparse structured inference
Vlad Niculae, André FT Martins, Mathieu Blondel, and Claire Cardie · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Differentiable convex optimization layers
A. Agrawal, B. Amos, S. Barratt, S. Boyd, S. Diamond, and Z. Kolter · 2019
Later among the works it cites.
Differentiating through a conic program
Akshay Agrawal, Shane Barratt, Stephen Boyd, Enzo Busseti, and Walaa M Moursi · 2019
Later among the works it cites.
Differentiable optimization-based modeling for machine learning
Brandon Amos · 2019
Later among the works it cites.
The Limited Multi-Label Projection Layer
Brandon Amos, Vladlen Koltun, and J. Zico Kolter · 2019
Later among the works it cites.
Structured prediction with projection oracles
Mathieu Blondel · 2019
Later among the works it cites.
Differentiable perturb-and-parse: Semi-supervised parsing with a structured variational autoencoder
Caio Corro and Ivan Titov · 2019
Later among the works it cites.
Stochastic optimization of sorting networks via continuous relaxations
Aditya Grover, Eric Wang, Aaron Zweig, and Stefano Ermon · 2019
Later among the works it cites.
Buy 4 reinforce samples, get a baseline for free!
Wouter Kool, Herke van Hoof, and Max Welling · 2019
Later among the works it cites.
Direct optimization through argmax for discrete variational auto-encoder
Guy Lorberbom, Andreea Gane, Tommi Jaakkola, and Tamir Hazan · 2019
Later among the works it cites.
Monte Carlo Gradient Estimation in Machine Learning
Shakir Mohamed, Mihaela Rosca, Michael Figurnov, and Andriy Mnih · 2019
Later among the works it cites.
Reparameterizable subset sampling via continuous relaxations
Sang Michael Xie and Stefano Ermon · 2019
Later among the works it cites.
ARM: Augment-REINFORCE-merge gradient for stochastic binary networks
Mingzhang Yin and Mingyuan Zhou · 2019
Later among the works it cites.
Learning with Differentiable Perturbed Optimizers
Quentin Berthet, Mathieu Blondel, Olivier Teboul, Marco Cuturi, Jean-Philippe Vert, and Francis Bach · 2020
Closest in time.
Learning with fenchel-young losses
Mathieu Blondel, André FT Martins, and Vlad Niculae · 2020
Closest in time.
Fast differentiable sorting and ranking
Mathieu Blondel, Olivier Teboul, Quentin Berthet, and Josip Djolonga · 2020
Closest in time.
Ancestral gumbel-top-k sampling for sampling without replacement
Wouter Kool, Herke van Hoof, and Max Welling · 2020
Closest in time.
Estimating gradients for discrete random variables by sampling without replacement
Wouter Kool, Herke van Hoof, and Max Welling · 2020
Closest in time.
Torch-struct: Deep structured prediction library
Alexander M Rush · 2020
Closest in time.