Fetching the paper…
Reading the bibliography…
The Gumbel-max trick is a method to draw a sample from a categorical distribution, given by its unnormalized (log-)probabilities.
L. L. Thurstone, “A law of comparative judgement,” Psychological Review , vol. 34, pp. 278–286, 1927
1927
Earlier work this paper cites.
E. J. Gumbel, “Les valeurs extrêmes des distributions statistiques,” Ann. l’institut Henri Poincaré , vol. 5, no. 2, pp. 115–158, 1935
1935
Earlier work this paper cites.
R. v. Mises, “La distribution de la plus grande de n valeurs,” Rev. Math. Union Interbalcanique , vol. 1, pp. 141–160, 1936
1936
Earlier work this paper cites.
E. J. Gumbel, Statistical Theory of Extreme Values and Some Practical Applications: A series of lectures . US Department of Commerce, 1954, vol. 33
1954
Earlier work this paper cites.
D. Raj, “Some estimators in sampling with varying probabilities without replacement,” Journal of the American Statistical Association , vol. 51, no. 274, pp. 269–284, 1956
1956
Earlier work this paper cites.
M. Murthy, “Ordered and unordered estimators in sampling without replacement,” Sankhyā: The Indian Journal of Statistics (1933-1960) , vol. 18, no. 3/4, pp. 379–390, 1957
1957
Earlier work this paper cites.
R. D. Luce, “Individual choice behavior.” 1959
1959
Earlier work this paper cites.
K. T. Wallenius, “Biased sampling; the noncentral hypergeometric probability distribution,” Stanford Univ CA Applied Mathematics and Statistics Labs, Tech. Rep., 1963
1963
Earlier work this paper cites.
W. K. Hastings, “Monte carlo sampling methods using markov chains and their applications,” 1970
1970
Earlier work this paper cites.
R. Plackett, “The Analysis of Permutations,” Appl. Stat. , vol. 24, no. 2, pp. 193–202, 1975
1975
Earlier work this paper cites.
J. Chesson, “A non-central multivariate hypergeometric distribution arising from biased sampling with application to selective predation,” Journal of Applied Probability , pp. 795–797, 1976
1976
Earlier work this paper cites.
J. I. Yellott, “The relationship between Luce’s Choice Axiom, Thurstone’s Theory of Comparative Judgment, and the double exponential distribution,” J. Math. Psychol. , vol. 15, no. 2, pp. 109–144, 1977
1977
Earlier work this paper cites.
R. S. Sutton, “Integrated architectures for learning, planning, and reacting based on approximating dynamic programming,” in Machine learning proceedings 1990 . Elsevier, 1990, pp. 216–224
1990
Earlier work this paper cites.
P. W. Glynn, “Likelihood ratio gradient estimation for stochastic systems,” Communications of the ACM , vol. 33, no. 10, pp. 75–84, 1990
1990
Earlier work this paper cites.
R. J. Williams, “Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning,” Mach. Learn. , vol. 8, no. 3, pp. 229–256, 1992
1992
Earlier work this paper cites.
J. F. C. Kingman, Poisson processes . Clarendon Press, 1992, vol. 3
1992
Earlier work this paper cites.
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine learning , vol. 8, no. 3, pp. 229–256, 1992
1992
Earlier work this paper cites.
2002
Earlier work this paper cites.
P. S. Efraimidis and P. G. Spirakis, “Weighted random sampling with a reservoir,” Inf. Process. Lett. , vol. 97, no. 5, pp. 181–185, 2006
2006
Earlier work this paper cites.
N. Duffield, C. Lund, and M. Thorup, “Priority sampling for estimation of arbitrary subset sums,” Journal of the ACM (JACM) , vol. 54, no. 6, p. 32, 2007
2007
Earlier work this paper cites.
A. Fog, “Calculation methods for wallenius’ noncentral hypergeometric distribution,” Communications in Statistics—Simulation and Computation® , vol. 37, no. 2, pp. 258–273, 2008
2008
Earlier work this paper cites.
G. Papandreou and A. L. Yuille, “Perturb-and-MAP random fields: Using discrete optimization to learn and sample from energy models,” in Proceedings of the Intern. Conf. on Comp. Vis. (ICCV) , 2011, pp. 193–200
2011
Earlier work this paper cites.
Y. C. Eldar and G. Kutyniok, Compressed sensing: theory and applications . Cambridge university press, 2012
2012
Earlier work this paper cites.
D. Tarlow, R. P. Adams, and R. S. Zemel, “Randomized optimum models for structured prediction,” in Proceedings of the Intern. Conf. on Artif. Intell. and Stats. (AISTATS) , vol. 22, 2012, pp. 1221–1229
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
T. Hazan, S. Maji, and T. Jaakkola, “On sampling from the Gibbs distribution with random maximum a-posteriori perturbations,” in Proceedings of the Conf. on Neur. Inf. Process. Syst. (NIPS) , 2013, pp. 1–9
2013
Earlier work this paper cites.
T. Salimans and D. A. Knowles, “Fixed-form variational posterior approximation through stochastic linear regression,” Bayesian Analysis , vol. 8, no. 4, Dec 2013. [Online]. Available: http://dx.doi.org/10.1214/13-BA858
2013
Earlier work this paper cites.
I. J. Goodfellow, J. Pouget-Abadie et al. , “Generative adversarial nets,” in NIPS , 2014
2014
Earlier work this paper cites.
C. J. Maddison, D. Tarlow, and T. Minka, “A* Sampling,” in Proceedings of the Conf. on Neur. Inf. Process. Syst. (NIPS) , vol. 27, 2014, pp. 3086–3094
2014
Earlier work this paper cites.
T. Vieira, “Gumbel-max trick and weighted reservoir sampling,” 2014, accessed on Feb. 3, 2021. [Online]. Available: http://timvieira.github.io/blog/post/2014/08/01/gumbel-max-trick-and-weighted-reservoir-sampling/
2014
Earlier work this paper cites.
D. J. Rezende, S. Mohamed, and D. Wierstra, “Stochastic backpropagation and approximate inference in deep generative models,” in ”Proceedings of the Intern. Conf. on Mach. Learn. (ICML)” . PMLR, 2014, pp. 1278–1286
2014
Earlier work this paper cites.
A. Mnih and K. Gregor, “Neural variational inference and learning in belief networks,” 2014
2014
Earlier work this paper cites.
K. S. Tai, R. Socher, and C. D. Manning, “Improved semantic representations from tree-structured long short-Term memory networks,” ACL , vol. 1, pp. 1556–1566, 2015
2015
Earlier work this paper cites.
M. Abadi, A. Agarwal et al. , “TensorFlow: Large-scale machine learning on heterogeneous systems,” 2015, software available from tensorflow.org. [Online]. Available: https://www.tensorflow.org/
2015
Earlier work this paper cites.
A. v. d. Oord, N. Kalchbrenner, and K. Kavukcuoglu, “Pixel recurrent neural networks,” in Proceedings of the Intern. Conf. on Mach. Learn. (ICML) , vol. 48, 2016, pp. 1747–1756. [Online]. Available: https://proceedings.mlr.press/v48/oord16.html
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Y. Chen and Z. Ghahramani, “Scalable discrete sampling as a multi-armed bandit problem,” in Proceedings of the Intern. Conf. on Mach. Learn. (ICML) , vol. 5, 2016, pp. 3691–3707
2016
Earlier work this paper cites.
C. Kim, A. Sabharwal, and S. Ermon, “Exact sampling with integer linear programs and random perturbations,” in Proceedings of the Conf. on Artif. Intell. (AAAI) , 2016, pp. 3248–3254
2016
Earlier work this paper cites.
S. Mussmann and S. Ermon, “Learning and inference via maximum inner product search,” in Proceedings of the Intern. Conf. on Mach. Learn. (ICML) , vol. 6, 2016, pp. 3814–3826
2016
Earlier work this paper cites.
A. Mnih and D. J. Rezende, “Variational inference for monte carlo objectives,” 2016
2016
Earlier work this paper cites.
E. Jang, S. Gu, and B. Poole, “Categorical reparameterization with Gumbel-softmax,” in Proceedings of the Intern. Conf. on Learn. Repr. (ICLR) , 2017
2017
Earlier work this paper cites.
C. J. Maddison, A. Mnih, and Y. W. Teh, “The concrete distribution: A continuous relaxation of discrete random variables,” in Proceedings of the Intern. Conf. on Learn. Repr. (ICLR) , 2017
2017
Earlier work this paper cites.
N. Cesa-Bianchi, C. Gentile et al. , “Boltzmann Exploration Done Right,” in Proceedings of the Conf. on Neur. Inf. Process. Syst. (NIPS) , 2017, pp. 5094–5101
2017
Earlier work this paper cites.
A. Van Den Oord, O. Vinyals et al. , “Neural discrete representation learning,” Proceedings of the Conf. on Neur. Inf. Process. Syst. (NIPS) , vol. 30, 2017
2017
Earlier work this paper cites.
J. A. Figueroa and A. R. Rivera, “Is Simple Better?: Revisiting Simple Generative Models for Unsupervised Clustering,” in Proceedings of the Conf. on Neur. Inf. Process. Syst. (NIPS) , no. Nips, 2017, pp. 1–6
2017
Earlier work this paper cites.
S. Yan, J. S. Smith et al. , “Hierarchical Multi-scale Attention Networks for action recognition,” Signal Process. Image Commun. , vol. 61, no. August 2017, pp. 73–84, 2018. [Online]. Available: https://doi.org/10.1016/j.image.2017.11.005
2017
Earlier work this paper cites.
S. Havrylov and I. Titov, “Emergence of Language with Multi-agent Games: Learning to Communicate with Sequences of Symbols,” in Proceedings of the Conf. on Neur. Inf. Process. Syst. (NIPS) , 2017
2017
Earlier work this paper cites.
I. Mordatch and P. Abbeel, “Emergence of grounded compositional language in multi-agent populations,” in Proceedings of the Conf. on Artif. Intell. (AAAI) , vol. 32, no. 1, 2017, pp. 1495–1502
2017
Earlier work this paper cites.
B. Xu, W. Kong, and J. Chen, “Semi-supervised Image Captioning via Reconstruction,” in Proceedings of the Intern. Conf. on Comp. Vis. (ICCV) , 2017, pp. 4135–4144
2017
Earlier work this paper cites.
R. Shetty, M. Rohrbach et al. , “Speaking the Same Language: Matching Machine to Human Captions by Adversarial Training,” in Proceedings of the Intern. Conf. on Comp. Vis. (ICCV) , 2017. [Online]. Available: https://goo.gl/3yRVnq
2017
Earlier work this paper cites.
J. Lu, A. Kannan et al. , “Best of Both Worlds: Transferring Knowledge from Discriminative Learning to a Generative Visual Dialog Model,” in Proceedings of the Conf. on Neur. Inf. Process. Syst. (NIPS) , 2017
2017
Earlier work this paper cites.
N. Baram, O. Ansehel et al. , “End-to-end differentiable adversarial imitation Learning,” in Proceedings of the Intern. Conf. on Mach. Learn. (ICML) , vol. 1, 2017, pp. 622–631
2017
Earlier work this paper cites.
Y. Gal, J. Hron, and A. Kendall, “Concrete dropout,” in Proceedings of the Conf. on Neur. Inf. Process. Syst. (NIPS) , 2017
2017
Earlier work this paper cites.
G. Tucker, A. Mnih et al. , “REBAR: Low-variance, unbiased gradient estimates for discrete latent variable models,” in Proceedings of the Conf. on Neur. Inf. Process. Syst. (NIPS) , vol. 30, 2017, pp. 2627–2636
2017
Earlier work this paper cites.
T. Vieira, “Estimating means in a finite universe,” 2017, accessed on April 8, 2021. [Online]. Available: https://timvieira.github.io/blog/post/2017/07/03/estimating-means-in-a-finite-universe/
2017
Earlier work this paper cites.
S. Mussmann, D. Levy, and S. Ermon, “Fast Amortized Inference and Learning in Log-linear Models with Randomly Perturbed Nearest Neighbor Search,” in Conf. Uncertain. Artif. Intell. , 2017
2017
Earlier work this paper cites.
S. Tokui and I. Sato, “Evaluating the variance of likelihood-ratio gradient estimators,” in Proceedings of the Intern. Conf. on Mach. Learn. (ICML) , vol. 7, 2017, pp. 5244–5257
2017
Cited alongside, same era.
A. Paszke, S. Gross et al. , “Automatic differentiation in pytorch,” 2017
2017
Cited alongside, same era.
B. Amos and J. Z. Kolter, “OptNet: Differentiable optimization as a layer in neural networks,” ICML , vol. 1, pp. 179–191, 2017
2017
Cited alongside, same era.
J. Djolonga and A. Krause, “Differentiable learning of submodular models,” in Proceedings of the Conf. on Neur. Inf. Process. Syst. (NIPS) , vol. December, no. 1, 2017, pp. 1014–1024
2017
Cited alongside, same era.
D. Silveira, A. Carvalho et al. , “Topic Modeling using Variational Auto-Encoders with Gumbel-Softmax and Logistic-Normal Mixture Distributions,” in Proceedings of the Intern. Joint Conf. on Neur. Netw. (IJCNN) , vol. 2018-July. IEEE, 2018, pp. 1–8
A. Biswas, T. T. Pham et al. , “Seeker: Real-Time Interactive Search,” in Int. Conf. Knowl. Discov. Data Min. , 2019, pp. 2867–2875. [Online]. Available: https://doi.org/10.1145/3292500.3330733
2019
Later among the works it cites.
M. Oberst and D. Sontag, “Counterfactual off-policy evaluation with gumbel-max structural causal models,” in Proceedings of the Intern. Conf. on Mach. Learn. (ICML) . PMLR, May 2019, pp. 4881–4890. [Online]. Available: http://proceedings.mlr.press/v97/oberst19a.html
2019
Later among the works it cites.
W. Kool, H. v. Hoof, and M. Welling, “Buy 4 reinforce samples, get a baseline for free!” Deep Reinf. Learn. Meets Struct. Predict. Deep. 2019 Work. , pp. 1–14, 2019
2019
Later among the works it cites.
M. Yin, Y. Yue, and M. Zhou, “ARSM: Augment-REINFORCE-Swap-Merge Estimator for Gradient Backpropagation Through Categorical Variables,” in Proceedings of the Intern. Conf. on Mach. Learn. (ICML) . PMLR, May 2019, pp. 7095–7104. [Online]. Available: https://github.com/ARM-gradient/ARSM
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
H. Liu, L. He et al. , “Structured inference for recurrent hidden semi-Markov model,” in Proceedings of the Intern. Joint Conf. on Artif. Intell. (IJCAI) , 2018, pp. 2447–2453
2018
Cited alongside, same era.
E. Dupont, “Learning Disentangled Joint Continuous and Discrete Representations,” in Proceedings of the Conf. on Neur. Inf. Process. Syst. (NIPS) , 2018
2018
Cited alongside, same era.
P. H. Chen, S. Si et al. , “Learning to screen for fast softmax inference on large vocabulary neural networks,” in Proceedings of the Intern. Conf. on Learn. Repr. (ICLR) , 2018
2018
Cited alongside, same era.
M. Asai and A. Fukunaga, “Classical planning in deep latent space: Bridging the subsymbolic-symbolic boundary,” in Proceedings of the Conf. on Artif. Intell. (AAAI) , vol. 32, no. 1, 2018
2018
Cited alongside, same era.
J. Chen, L. Song et al. , “Learning to explain: An information-theoretic perspective on model interpretation,” in Proceedings of the Intern. Conf. on Mach. Learn. (ICML) , vol. 2, 2018, pp. 1386–1418
2018
Cited alongside, same era.
Y. Tay, L. A. Tuan, and S. C. Hui, “Multi-pointer co-attention networks for recommendation,” ACM SIGKDD Int. Conf. Knowl. Discov. Data Min. , pp. 2309–2318, 2018
2018
Cited alongside, same era.
J. Gu, D. J. Im, and V. O. K. Li, “Neural Machine Translation with Gumbel-Greedy Decoding,” AAAI , vol. 32, no. 1, 2018
2018
Cited alongside, same era.
2019
Later among the works it cites.
A. Potapczynski, G. Loaiza-Ganem, and J. P. Cunningham, “Invertible Gaussian reparameterization: Revisiting the Gumbel-Softmax,” in Proceedings of the Conf. on Neur. Inf. Process. Syst. (NIPS) , 2019. [Online]. Available: https://github.com/cunningham-lab/
2019
Later among the works it cites.
A. Grover, E. Wang et al. , “Stochastic optimization of sorting networks via continuous relaxations,” in Proceedings of the Intern. Conf. on Learn. Repr. (ICLR) , Mar 2019
2019
Later among the works it cites.
G. Lorberbom, A. Gane et al. , “Direct optimization through arg max for discrete variational auto-encoder,” in Proceedings of the Conf. on Neur. Inf. Process. Syst. (NIPS) , 2019, pp. 6203–6214
2019
Later among the works it cites.
L. Ma, B. Ding et al. , “Active Learning for ML Enhanced Database Systems,” in ACM SIGMOD Int. Conf. Manag. Data , vol. 17, no. 20. ACM, 2020, pp. 175–191
2020
Later among the works it cites.
M. Firdaus, A. P. Shandeelya, and A. Ekbal, “More to diverse: Generating diversified responses in a task oriented multimodal dialog system,” PLoS One , vol. 15, no. 11, 2020. [Online]. Available: https://doi.org/10.1371/journal.pone.0241271.g001
2020
Later among the works it cites.
W. Kool, H. v. Hoof, and M. Welling, “Ancestral gumbel-top-k sampling for sampling without replacement,” Journal of Mach. Learn. Research , vol. 21, pp. 1–36, 2020. [Online]. Available: http://jmlr.org/papers/v21/19-985.html
2020
Later among the works it cites.
C. Wang, C. Deng, and V. Ivanov, “SAG-VAE: End-to-end joint inference of data representations and feature relations,” in Proceedings of the Intern. Joint Conf. on Neur. Netw. (IJCNN) , 2020, pp. 1–9
2020
Later among the works it cites.
B. Gao, Y. Yang et al. , “Deep clustering with concrete K-means,” in Proceedings of the Conf. on Neur. Inf. Process. Syst. (NIPS) , 2020, pp. 4252–4256
2020
Later among the works it cites.
A. Baevski, S. Schneider, and M. Auli, “VQ-WAV2VEC: Self-supervised learning of discrete speech representations,” in Proceedings of the Intern. Conf. on Learn. Repr. (ICLR) , 2020
2020
Later among the works it cites.
Y. Dai, C. Guo et al. , “Drug–drug interaction prediction with Wasserstein Adversarial Autoencoder-based knowledge graph embeddings,” Brief. Bioinform. , pp. 1–16, 2020. [Online]. Available: https://academic.oup.com/bib/advance-article/doi/10.1093/bib/bbaa256/5943784
2020
Later among the works it cites.
D. B. Acharya and H. Zhang, “Community Detection Clustering via Gumbel Softmax,” SN Comput. Sci. , vol. 1, no. 5, pp. 1–11, 2020. [Online]. Available: https://doi.org/10.1007/s42979-020-00264-2
2020
Later among the works it cites.
N. T. Ngo, T. N. Nguyen, and T. H. Nguyen, “Learning to Select Important Context Words for Event Detection,” in Proceedings of the Pacific Asia Conf. on Knowl. Disc. and Data Mining (PAKDD) , vol. 12085 LNAI. Springer International Publishing, 2020, pp. 756–768
2020
Later among the works it cites.
H. Y. Tseng, H. Y. Lee et al. , “RetrieveGAN: Image Synthesis via Differentiable Patch Retrieval,” in Proceedings of the Europ. Conf. on Comp. Vis. (ECCV) , vol. 12353, 2020, pp. 242–257
2020
Later among the works it cites.
P. Yang, J. Chen et al. , “Greedy attack and gumbel attack: Generating adversarial examples for discrete data,” J. Mach. Learn. Res. , vol. 21, pp. 1–36, 2020
2020
Later among the works it cites.
D. Bhaskar Acharya and H. Zhang, “Feature Selection and Extraction for Graph Neural Networks,” in ACM Southeast Conf. , 2020. [Online]. Available: https://doi.org/10.1145/3374135.3385309
2020
Later among the works it cites.
P. Chonwiharnphan, P. Thienprapasith, and E. Chuangsuwanich, “Generating Realistic Users Using Generative Adversarial Network with Recommendation-Based Embedding,” IEEE Access , vol. 8, pp. 41 384–41 393, 2020
2020
Later among the works it cites.
H. Zhao and R. P. Wildes, “On Diverse Asynchronous Activity Anticipation,” in Proceedings of the Europ. Conf. on Comp. Vis. (ECCV) , vol. 12374. Springer Science and Business Media Deutschland GmbH, Aug 2020, pp. 781–799. [Online]. Available: https://doi.org/10.1007/978-3-030-58526-6{\\_}46
2020
Later among the works it cites.
W. C. Wei, C. J. Jhang et al. , “A Relaxed Quantization Training Method for Hardware Limitations of Resistive Random Access Memory (ReRAM)-Based Computing-in-Memory,” IEEE J. Explor. Solid-State Comput. Devices Circuits , vol. 6, no. 1, pp. 45–52, 2020
2020
Later among the works it cites.
Y. Kim, K. Kim, and S. Lee, “Adaptive Compression of Word Embeddings,” in Proceedings of the Annual Meeting of the Assoc. for Computational Linguistics (ACL) , 2020, pp. 3950–3959
2020
Later among the works it cites.
Y. Yang, R. Bamler, and S. Mandt, “Improving Inference for Neural Image Compression,” in Proceedings of the Conf. on Neur. Inf. Process. Syst. (NIPS) , 2020
2020
Later among the works it cites.
I. A. M. Huijben, B. S. Veeling, and R. J. G. v. Sloun, “Deep probabilistic subsampling for task-adaptive compressed sensing,” in Proceedings of the Intern. Conf. on Learn. Repr. (ICLR) , 2020, pp. 1–16
2020
Later among the works it cites.
I. A. M. Huijben, B. S. Veeling et al. , “Learning sub-sampling and signal recovery with applications in ultrasound imaging,” IEEE Trans. Med. Imaging , pp. 1–13, 2020
2020
Later among the works it cites.
I. A. M. Huijben, B. S. Veeling, and R. J. G. v. Sloun, “Learning Sampling and Model-Based Signal Recovery for Compressed Sensing MRI,” in Proceedings of the IEEE Intern. Conf. on Acoustics, Speech, and Signal Process. (ICASSP) , vol. 2020-May, Apr 2020, pp. 8906–8910
2020
Later among the works it cites.
Y. Yang, S. Zhang et al. , “Deep Learning Based Antenna Selection for Channel Extrapolation in FDD Massive MIMO,” in WCSP , 2020, pp. 182–187
2020
Later among the works it cites.
C. Herrmann, R. S. Bowen, and R. Zabih, “Channel Selection Using Gumbel Softmax,” in Proceedings of the Europ. Conf. on Comp. Vis. (ECCV) , 2020, pp. 241–257
2020
Later among the works it cites.
A. Wan, X. Dai et al. , “FBNetV2: Differentiable Neural Architecture Search for Spatial and Channel Dimensions,” in Proceedings of the Conf. on Comp. Vis. and Pattern Recogn. (CVPR) , 2020, pp. 12 962–12 971
2020
Later among the works it cites.
M. Kang and B. Han, “Operation-Aware Soft Channel Pruning using Differentiable Masks,” in Proceedings of the Intern. Conf. on Mach. Learn. (ICML) , 2020, pp. 5122–5131
2020
Later among the works it cites.
Y. Wang, X. Zhang et al. , “Dynamic Network Pruning with Interpretable Layerwise Channel Selection,” AAAI , vol. 34, no. 04, pp. 6299–6306, 2020
2020
Later among the works it cites.
T. Verelst and T. Tuytelaars, “Dynamic convolutions: Exploiting spatial sparsity for faster inference,” in Proceedings of the Conf. on Comp. Vis. and Pattern Recogn. (CVPR) , 2020, pp. 2317–2326. [Online]. Available: https://github.com/thomasverelst/dynconv
2020
Later among the works it cites.
P. Guo, C. Lee, and D. Ulbricht, “Learning to Branch for Multi-Task Learning,” in Proceedings of the Intern. Conf. on Mach. Learn. (ICML) , 2020
2020
Later among the works it cites.
C. Su, H. Huang et al. , “Neural machine translation with Gumbel Tree-LSTM based encoder,” J. Vis. Commun. Image Represent. , vol. 71, p. 102811, Aug 2020
2020
Later among the works it cites.
A. Holtzman, J. Buys et al. , “The curious case of neural text degeneration,” in Proceedings of the Intern. Conf. on Learn. Repr. (ICLR) , 2020
2020
Later among the works it cites.
K. Shi, D. Bieber, and C. Sutton, “Incremental sampling without replacement for sequence models,” in Proceedings of the Intern. Conf. on Mach. Learn. (ICML) , 2020, pp. 8785–8795
2020
Later among the works it cites.
A. Wiggers and E. Hoogeboom, “Predictive Sampling with Forecasting Autoregressive Models,” in Proceedings of the Intern. Conf. on Mach. Learn. (ICML) , 2020, pp. 10 260–10 269
2020
Later among the works it cites.
Y. Qi, P. Wang et al. , “Fast Generating A Large Number of Gumbel-Max Variables,” in Proceedings of the World Wide Web Conf. (WWW) . ACM, 2020, pp. 796–807
2020
Later among the works it cites.
E. Beck, C. Bockelmann, and A. Dekorsy, “Concrete MAP detection: A machine learning inspired relaxation,” in Proceedings of the IEEE Intern. ITG Workshop on Smart Antennas (WSA) , 2020, pp. 1–5
2020
Later among the works it cites.
Y. Li, J. Liu et al. , “Gumbel-softmax-based Optimization: A simple general framework for optimization problems on graphs,” in Int. Conf. Complex Networks Their Appl. Springer International Publishing, 2020, pp. 879–890. [Online]. Available: http://dx.doi.org/10.1007/978-3-030-36687-2{\\_}73
2020
Later among the works it cites.
S. Mohamed, M. Rosca et al. , “Monte carlo gradient estimation in machine learning,” 2020
2020
Later among the works it cites.
M. B. Paulus, E. Zürich et al. , “Gradient Estimation with Stochastic Softmax Tricks,” in Proceedings of the Conf. on Neur. Inf. Process. Syst. (NIPS) , 2020
2020
Later among the works it cites.
F. Guo, J. Boyd-Graber, and L. Findlater, “Which Evaluations Uncover Sense Representations that Actually Make Sense?” in Lang. Resour. Eval. Conf. , 2020, pp. 1727–1738
2020
Later among the works it cites.
O. Wieder, S. Kohlbacher et al. , “A compact review of molecular property prediction with graph neural networks,” Drug Discovery Today: Technologies , 2020. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S1740674920300305
2020
Later among the works it cites.
Q. Berthet, M. Blondel et al. , “Learning with differentiable perturbed optimizers,” in Proceedings of the Conf. on Neur. Inf. Process. Syst. (NIPS) , 2020, pp. 9508—-9519
2020
Later among the works it cites.
A. Ramesh, M. Pavlov et al. , “Zero-shot text-to-image generation,” in Proceedings of the Intern. Conf. on Mach. Learn. (ICML) . PMLR, 2021, pp. 8821–8831
2021
Closest in time.
M. Rezaee and F. Ferraro, “Event Representation with Sequential, Semi-Supervised Discrete Variables,” NAACL , vol. June, pp. 4701–4716, 2021
2021
Closest in time.
T. Kojima, Y. Iwasawa, and Y. Matsuo, “Making Use of Latent Space in Language {GAN}s for Generating Diverse Text without Pre-training,” Proc. 16th Conf. Eur. Chapter Assoc. Comput. Linguist. Student Res. Work. , pp. 175–182, 2021. [Online]. Available: https://www.aclweb.org/anthology/2021.eacl-srw.23
2021
Closest in time.
H. v. Gorp, I. A. M. Huijben et al. , “Active deep probabilistic subsampling,” in Proceedings of the Intern. Conf. on Mach. Learn. (ICML) , 2021
2021
Closest in time.
2021
Closest in time.
S. Cai, Y. Shu, and W. Wang, “Dynamic Routing Networks,” in Proceedings of the Winter Conf. on Applications of Comp. Vis. (WACV) , 2021, pp. 3588–3597
2021
Closest in time.
M. S. Schlichtkrull, N. De Cao, and I. Titov, “Interpreting graph neural networks for NLP with differentiable edge masking,” in Proceedings of the Intern. Conf. on Learn. Repr. (ICLR) , 2021
2021
Closest in time.
M. B. Paulus, C. J. Maddison, and A. Krause, “Rao-Blackwellizing the Straight-Through Gumbel-Softmax Gradient Estimator,” in Proceedings of the Intern. Conf. on Learn. Repr. (ICLR) , Oct 2021
2021
Closest in time.
E. Hoogeboom, D. Nielsen et al. , “Argmax Flows and Multinomial Diffusion: Towards Non-Autoregressive Language Models,” arxiv , 2021
2021
Closest in time.
C. J. Maddison and D. Tarlow, “Gumbel machinery,” 2017, accessed on Feb. 10, 2021. [Online]. Available: https://cmaddis.github.io/gumbel-machinery
2021
Closest in time.
S. Gould, R. Hartley, and D. J. Campbell, “Deep Declarative Networks,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 8828, 2021
2021
Closest in time.