Fetching the paper…
Reading the bibliography…
Generative Flow Networks (GFNs) have emerged as a powerful tool for sampling discrete objects from unnormalized distributions, offering a scalable alternative to Markov Chain Monte Carlo (MCMC) methods.
The theory of dynamic programming
Bellman, R. (1954) · 1954
Earlier work this paper cites.
Optimal control and nonlinear filtering for nondegenerate diffusion processes
Fleming, W. H. and Mitter, S. K. (1981) · 1981
Earlier work this paper cites.
Optimal replacement of gmc bus engines: An empirical model of harold zurcher
Rust, J. (1987) · 1987
Earlier work this paper cites.
Robust estimation of a location parameter
Huber, P. J. (1992) · 1992
Earlier work this paper cites.
Death and discounting
Shwartz, A. (2001) · 2001
Earlier work this paper cites.
Fundamentals of convex analysis
Hiriart-Urruty, J.-B. and Lemaréchal, C. (2004) · 2004
Earlier work this paper cites.
Linearly-solvable markov decision problems
Todorov, E. (2006) · 2006
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Ziebart, B. D., Maas, A., Bagnell, J. A., and Dey, A. K. (2008) · 2008
Earlier work this paper cites.
Smoothing Techniques for Computing Nash Equilibria of Sequential Games
Hoda, S., Gilpin, A., Peña, J., and Sandholm, T. (2010) · 2010
Earlier work this paper cites.
Relative entropy policy search
Peters, J., Mulling, K., and Altun, Y. (2010) · 2010
Earlier work this paper cites.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Sutton, R. S., Modayil, J., Delp, M., Degris, T., Pilarski, P. M., White, A., and Precup, D. (2011) · 2011
Earlier work this paper cites.
On stochastic optimal control and reinforcement learning by approximate inference
Rawlik, K., Toussaint, M., and Vijayakumar, S. (2012) · 2012
Earlier work this paper cites.
A link based network route choice model with unrestricted choice set
Fosgerau, M., Frejinger, E., and Karlstrom, A. (2013) · 2013
Earlier work this paper cites.
Quantum chemistry structures and properties of 134 kilo molecules
Ramakrishnan, R., Dral, P. O., Rupp, M., and Von Lilienfeld, O. A. (2014) · 2014
Earlier work this paper cites.
Learning of non-parametric control policies with high-dimensional state features
Van Hoof, H., Peters, J., and Neumann, G. (2015) · 2015
Cited alongside, same era.
Developing a predictive approach to knowledge
White, A. (2015) · 2015
Cited alongside, same era.
Taming the noise in reinforcement learning via soft updates
Fox, R., Pakman, A., and Tishby, N. (2016) · 2016
Cited alongside, same era.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S. (2017) · 2017
Cited alongside, same era.
Bridging the gap between value and policy based reinforcement learning
Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D. (2017) · 2017
Cited alongside, same era.
Junction tree variational autoencoder for molecular graph generation
Jin, W., Barzilay, R., and Jaakkola, T. (2018) · 2018
Cited alongside, same era.
Bayesian structure learning with generative flow networks
Deleu, T., Góis, A., Emezue, C., Rankawat, M., Lacoste-Julien, S., Bauer, S., and Bengio, Y. (2022) · 2022
Later among the works it cites.
Undiscounted recursive path choice models: Convergence properties and algorithms
Mai, T. and Frejinger, E. (2022) · 2022
Later among the works it cites.
Trajectory balance: Improved credit assignment in gflownets
Malkin, N., Jain, M., Bengio, E., Sun, C., and Bengio, Y. (2022) · 2022
Later among the works it cites.
Generative flow networks for discrete probabilistic modeling
Zhang, D., Malkin, N., Liu, Z., Volokhova, A., Courville, A., and Bengio, Y. (2022) · 2022
Later among the works it cites.
Gflownet foundations
Bengio, Y., Lahlou, S., Deleu, T., Hu, E. J., Tiwari, M., and Bengio, E. (2023) · 2023
Closest in time.
Extreme q-learning: Maxent rl without entropy
Garg, D., Hejna, J., Geist, M., and Ermon, S. (2023) · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Differentiable dynamic programming for structured prediction and attention
Mensch, A. and Blondel, M. (2018) · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Cited alongside, same era.
A theory of regularized markov decision processes
Geist, M., Scherrer, B., and Pietquin, O. (2019) · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al. (2019) · 2019
Cited alongside, same era.
Approximate inference in discrete distributions with monte carlo tree search and value functions
Buesing, L., Heess, N., and Weber, T. (2020) · 2020
Cited alongside, same era.
Munchausen reinforcement learning
Vieillard, N., Pietquin, O., and Geist, M. (2020) · 2020
Cited alongside, same era.
Closest in time.
Multi-objective gflownets
Jain, M., Raparthy, S. C., Hernández-García, A., Rector-Brooks, J., Bengio, Y., Miret, S., and Bengio, E. (2023) · 2023
Closest in time.
Expected flow networks in stochastic environments and two-player zero-sum games
Jiralerspong, M., Sun, B., Vucetic, D., Zhang, T., Bengio, Y., Gidel, G., and Malkin, N. (2023) · 2023
Closest in time.
Kim, M., Yun, T., Bengio, E., Zhang, D., Bengio, Y., Ahn, S., and Park, J. (2023) · 2023
Closest in time.
Learning gflownets from partial episodes for improved convergence and stability
Madan, K., Rector-Brooks, J., Korablyov, M., Bengio, E., Jain, M., Nica, A. C., Bosc, T., Bengio, Y., and Malkin, N. (2023) · 2023
Closest in time.
Stochastic generative flow networks
Pan, L., Zhang, D., Jain, M., Huang, L., and Bengio, Y. (2023) · 2023
Closest in time.
Towards understanding and improving gflownet training
Shen, M. W., Bengio, E., Hajiramezanali, E., Loukas, A., Cho, K., and Biancalani, T. (2023) · 2023
Closest in time.
Generative flow networks as entropy-regularized rl
Tiapkin, D., Morozov, N., Naumov, A., and Vetrov, D. (2023) · 2023
Closest in time.
Discrete probabilistic inference as control in multi-path environments
Deleu, T., Nouri, P., Malkin, N., Precup, D., and Bengio, Y. (2024) · 2024
Closest in time.