Fetching the paper…
Reading the bibliography…
We consider the problem of sampling from a discrete and structured distribution as a sequential decision problem, where the objective is to find a stochastic policy such that objects are sampled at the end of this sequential process proportionally to some predefined reward.
Monte Carlo sampling methods using Markov chains and their applications
W. Keith Hastings · 1970
Earlier work this paper cites.
Sampling-based approaches to calculating marginal densities
Alan E. Gelfand and Adrian FM. Smith · 1990
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 1992
Earlier work this paper cites.
Learning Gaussian Networks
Dan Geiger and David Heckerman · 1994
Earlier work this paper cites.
Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping
Andrew Y. Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
Being Bayesian About Network Structure. A Bayesian Approach to Structure Discovery in Bayesian Networks
Nir Friedman and Daphne Koller · 2003
Earlier work this paper cites.
Linearly-solvable Markov decision problems
Emanuel Todorov · 2006
Earlier work this paper cites.
Maximum Entropy Inverse Reinforcement Learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, Anind K Dey, et al · 2008
Earlier work this paper cites.
Modeling Purposeful Adaptive Behavior with the Principle of Maximum Causal Entropy
Brian D Ziebart · 2010
Earlier work this paper cites.
Optimal control as a graphical model inference problem
Hilbert J Kappen, Vicenç Gómez, and Manfred Opper · 2012
Earlier work this paper cites.
Auto-Encoding Variational Bayes
Diederik P Kingma and Max Welling · 2014
Earlier work this paper cites.
Stochastic Backpropagation and Approximate Inference in Deep Generative Models
Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra · 2014
Earlier work this paper cites.
Data Generation as Sequential Decision Making
Philip Bachman and Doina Precup · 2015
Earlier work this paper cites.
Variational Inference with Normalizing Flows
Danilo Rezende and Shakir Mohamed · 2015
Earlier work this paper cites.
Reinforced variational inference
Theophane Weber, Nicolas Heess, Ali Eslami, John Schulman, David Wingate, and David Silver · 2015
Earlier work this paper cites.
Unifying Count-Based Exploration and Intrinsic Motivation
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Earlier work this paper cites.
Taming the Noise in Reinforcement Learning via Soft Updates
Roy Fox, Ari Pakman, and Naftali Tishby · 2016
Earlier work this paper cites.
Dueling Network Architectures for Deep Reinforcement Learning
Ziyu Wang, Tom Schaul, Matteo Hessel, Hado Hasselt, Marc Lanctot, and Nando Freitas · 2016
Earlier work this paper cites.
Reinforcement Learning with Deep Energy-Based Policies
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine · 2017
Earlier work this paper cites.
Categorical Reparameterization with Gumbel-Softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2017
Earlier work this paper cites.
The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables
Chris J. Maddison, Andriy Mnih, and Yee Whye Teh · 2017
Earlier work this paper cites.
Bridging the Gap Between Value and Policy Based Reinforcement Learning
Ofir Nachum, Mohammad Norouzi, Kelvin Xu, and Dale Schuurmans · 2017
Earlier work this paper cites.
Curiosity-driven Exploration by Self-supervised Prediction
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell · 2017
Cited alongside, same era.
Equivalence Between Policy Gradients and Soft Q-Learning
John Schulman, Xi Chen, and Pieter Abbeel · 2017
Cited alongside, same era.
Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review
Sergey Levine · 2018
Cited alongside, same era.
Learning Latent Permutations with Gumbel-Sinkhorn Networks
Gonzalo Mena, David Belanger, Scott Linderman, and Jasper Snoek · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Graph Convolutional Policy Network for Goal-Directed Molecular Graph Generation
Jiaxuan You, Bowen Liu, Zhitao Ying, Vijay Pande, and Jure Leskovec · 2018
Optimizing ddpm sampling with shortcut fine-tuning
Ying Fan and Kangwook Lee · 2023
Later among the works it cites.
Multi-Fidelity Active Learning with GFlowNets
Alex Hernandez-Garcia, Nikita Saxena, Moksh Jain, Cheng-Hao Liu, and Yoshua Bengio · 2023
Later among the works it cites.
GFlowNet-EM for Learning Compositional Latent Variable Models
Edward J Hu, Nikolay Malkin, Moksh Jain, Katie E Everett, Alexandros Graikos, and Yoshua Bengio · 2023
Later among the works it cites.
A Theory of Continuous Generative Flow Networks
Salem Lahlou, Tristan Deleu, Pablo Lemos, Dinghuai Zhang, Alexandra Volokhova, Alex Hernández-García, Léna Néhale Ezzine, Yoshua Bengio, and Nikolay Malkin · 2023
Later among the works it cites.
CFlowNets: Continuous control with Generative Flow Networks
Yinchuan Li, Shuang Luo, Haozhi Wang, and Jianye Hao · 2023
Later among the works it cites.
Learning GFlowNets from partial episodes for improved convergence and stability
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Soft Actor-Critic for Discrete Action Settings
Petros Christodoulou · 2019
Cited alongside, same era.
A Theory of Regularized Markov Decision Processes
Matthieu Geist, Bruno Scherrer, and Olivier Pietquin · 2019
Cited alongside, same era.
Model-based reinforcement learning for biological sequence design
Christof Angermueller, David Dohan, David Belanger, Ramya Deshpande, Kevin Murphy, and Lucy Colwell · 2020
Cited alongside, same era.
Approximate Inference in Discrete Distributions with Monte Carlo Tree Search and Value Functions
Lars Buesing, Nicolas Heess, and Theophane Weber · 2020
Cited alongside, same era.
Monte Carlo Gradient Estimation in Machine Learning
Shakir Mohamed, Mihaela Rosca, Michael Figurnov, and Andriy Mnih · 2020
Cited alongside, same era.
Flow Network based Generative Models for Non-Iterative Diverse Candidate Generation
Emmanuel Bengio, Moksh Jain, Maksym Korablyov, Doina Precup, and Yoshua Bengio · 2021
Cited alongside, same era.
Kanika Madan, Jarrid Rector-Brooks, Maksym Korablyov, Emmanuel Bengio, Moksh Jain, Andrei Nica, Tom Bosc, Yoshua Bengio, and Nikolay Malkin · 2023
Later among the works it cites.
GFlowNets and variational inference
Nikolay Malkin, Salem Lahlou, Tristan Deleu, Xu Ji, Edward Hu, Katie Everett, Dinghuai Zhang, and Yoshua Bengio · 2023
Later among the works it cites.
Crystal-GFN: sampling crystals with desirable properties and constraints
Mila AI4Science, Alex Hernandez-Garcia, Alexandre Duval, Alexandra Volokhova, Yoshua Bengio, Divya Sharma, Pierre Luc Carrier, Michał Koziarski, and Victor Schmidt · 2023
Later among the works it cites.
Bayesian learning of Causal Structure and Mechanisms with GFlowNets and Variational Bayes
Mizu Nishikawa-Toomey, Tristan Deleu, Jithendaraa Subramanian, Yoshua Bengio, and Laurent Charlin · 2023
Later among the works it cites.
Thompson sampling for improved exploration in GFlowNets
Jarrid Rector-Brooks, Kanika Madan, Moksh Jain, Maksym Korablyov, Cheng-Hao Liu, Sarath Chandar, Nikolay Malkin, and Yoshua Bengio · 2023
Later among the works it cites.
Towards Understanding and Improving GFlowNet Training
Max W Shen, Emmanuel Bengio, Ehsan Hajiramezanali, Andreas Loukas, Kyunghyun Cho, and Tommaso Biancalani · 2023
Later among the works it cites.
An Empirical Study of the Effectiveness of Using a Replay Buffer on Mode Discovery in GFlowNets
Nikhil Vemgal, Elaine Lau, and Doina Precup · 2023
Later among the works it cites.
Training Diffusion Models with Reinforcement Learning
Kevin Black, Michael Janner, Yilun Du, Ilya Kostrikov, and Sergey Levine · 2024
Closest in time.
Delta-AI: Local objectives for amortized inference in sparse graphical models
Jean-Pierre Falet, Hae Beom Lee, Nikolay Malkin, Chen Sun, Dragos Secrieru, Dinghuai Zhang, Guillaume Lajoie, and Yoshua Bengio · 2024
Closest in time.
Amortizing intractable inference in large language models
Edward J Hu, Moksh Jain, Eric Elmoznino, Younesse Kaddar, Guillaume Lajoie, Yoshua Bengio, and Nikolay Malkin · 2024
Closest in time.
Expected flow networks in stochastic environments and two-player zero-sum games
Marco Jiralerspong, Bilun Sun, Danilo Vucetic, Tianyu Zhang, Yoshua Bengio, Gauthier Gidel, and Nikolay Malkin · 2024
Closest in time.
Maximum entropy GFlowNets with soft Q-learning
Sobhan Mohammadpour, Emmanuel Bengio, Emma Frejinger, and Pierre-Luc Bacon · 2024
Closest in time.
GFlowNet Training by Policy Gradients
Puhua Niu, Shili Wu, Mingzhou Fan, and Xiaoning Qian · 2024
Closest in time.
On diffusion models for amortized inference: Benchmarking and improving stochastic control and sampling
Marcin Sendera, Minsu Kim, Sarthak Mittal, Pablo Lemos, Luca Scimeca, Jarrid Rector-Brooks, Alexandre Adam, Yoshua Bengio, and Nikolay Malkin · 2024
Closest in time.
Generative Flow Networks as Entropy-Regularized RL
Daniil Tiapkin, Nikita Morozov, Alexey Naumov, and Dmitry Vetrov · 2024
Closest in time.
PhyloGFN: Phylogenetic inference with generative flow networks
Mingyang Zhou, Zichao Yan, Elliot Layne, Nikolay Malkin, Dinghuai Zhang, Moksh Jain, Mathieu Blanchette, and Yoshua Bengio · 2024
Closest in time.