Fetching the paper…
Reading the bibliography…
Generative flow networks (GFlowNets) are amortized variational inference algorithms that treat sampling from a distribution over compositional objects as a sequential decision-making problem with a learnable action policy.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, W. R · 1933
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Ziebart, B. D · 2010
Earlier work this paper cites.
Analysis of thompson sampling for the multi-armed bandit problem
Agrawal, S. and Goyal, N · 2012
Earlier work this paper cites.
(more) efficient reinforcement learning via posterior sampling
Osband, I., Russo, D., and Van Roy, B · 2013
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Near-optimal regret bounds for thompson sampling
Agrawal, S. and Goyal, N · 2017
Earlier work this paper cites.
Optimistic posterior sampling for reinforcement learning: worst-case regret bounds
Agrawal, S. and Jia, R · 2017
Earlier work this paper cites.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Earlier work this paper cites.
Bridging the gap between value and policy based reinforcement learning
Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Exploration by random network distillation
Burda, Y., Edwards, H., Storkey, A., and Klimov, O · 2018
Earlier work this paper cites.
Diversity is all you need: Learning skills without a reward function
Eysenbach, B., Gupta, A., Ibarz, J., and Levine, S · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
Ray: A distributed framework for emerging { \{ AI } \} applications
Moritz, P., Nishihara, R., Wang, S., Tumanov, A., Liaw, R., Liang, E., Elibol, M., Yang, Z., Paul, W., Jordan, M. I., et al · 2018
Cited alongside, same era.
Information-directed exploration for deep reinforcement learning
Nikolov, N., Kirschner, J., Berkenkamp, F., and Krause, A · 2018
Cited alongside, same era.
Randomized prior functions for deep reinforcement learning
Osband, I., Aslanides, J., and Cassirer, A · 2018
Cited alongside, same era.
Flow network based generative models for non-iterative diverse candidate generation
Bengio, E., Jain, M., Korablyov, M., Precup, D., and Bengio, Y · 2021
Later among the works it cites.
Sunrise: A simple unified framework for ensemble learning in deep reinforcement learning
Lee, K., Laskin, M., Srinivas, A., and Abbeel, P · 2021
Later among the works it cites.
Exploration by maximizing rényi entropy for reward-free rl framework
Zhang, C., Cai, Y., Huang, L., and Li, J · 2021
Later among the works it cites.
Bayesian structure learning with generative flow networks
Deleu, T., Góis, A., Emezue, C., Rankawat, M., Lacoste-Julien, S., Bauer, S., and Bengio, Y · 2022
Later among the works it cites.
Trajectory balance: Improved credit assignment in GFlowNets
Malkin, N., Jain, M., Bengio, E., Sun, C., and Bengio, Y · 2022
Later among the works it cites.
Generative augmented flow networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The uncertainty bellman equation and exploration
O’Donoghue, B., Osband, I., Munos, R., and Mnih, V · 2018
Cited alongside, same era.
A tutorial on thompson sampling
Russo, D. J., Van Roy, B., Kazerouni, A., Osband, I., Wen, Z., et al · 2018
Cited alongside, same era.
Optuna: A next-generation hyperparameter optimization framework
Akiba, T., Sano, S., Yanase, T., Ohta, T., and Koyama, M · 2019
Cited alongside, same era.
Better exploration with optimistic actor critic
Ciosek, K., Vuong, Q., Loftin, R., and Hofmann, K · 2019
Cited alongside, same era.
Provably efficient maximum entropy exploration
Hazan, E., Kakade, S., Singh, K., and Van Soest, A · 2019
Cited alongside, same era.
Marginalized state distribution entropy regularization in policy optimization
Islam, R., Ahmed, Z., and Precup, D · 2019
Cited alongside, same era.
Deep exploration via randomized value functions
Osband, I., Van Roy, B., Russo, D. J., Wen, Z., et al · 2019
Cited alongside, same era.
Pan, L., Zhang, D., Courville, A., Huang, L., and Bengio, Y · 2022
Later among the works it cites.
Generative flow networks for discrete probabilistic modeling
Zhang, D., Malkin, N., Liu, Z., Volokhova, A., Courville, A., and Bengio, Y · 2022
Later among the works it cites.
GFlowNet foundations
Bengio, Y., Lahlou, S., Deleu, T., Hu, E., Tiwari, M., and Bengio, E · 2023
Closest in time.
GFlowNet-EM for learning compositional latent variable models
Hu, E. J., Malkin, N., Jain, M., Everett, K., Graikos, A., and Bengio, Y · 2023
Closest in time.
Learning GFlowNets from partial episodes for improved convergence and stability
Madan, K., Rector-Brooks, J., Korablyov, M., Bengio, E., Jain, M., Nica, A., Bosc, T., Bengio, Y., and Malkin, N · 2023
Closest in time.
GFlowNets and variational inference
Malkin, N., Lahlou, S., Deleu, T., Ji, X., Hu, E., Everett, K., Zhang, D., and Bengio, Y · 2023
Closest in time.
Better training of gflownets with local credit and incomplete trajectories
Pan, L., Malkin, N., Zhang, D., and Bengio, Y · 2023
Closest in time.
Robust scheduling with gflownets
Zhang, D. W., Rainone, C., Peschl, M., and Bondesan, R · 2023
Closest in time.