Fetching the paper…
Reading the bibliography…
The recently proposed generative flow networks (GFlowNets) are a method of training a policy to sample compositional discrete objects with probabilities proportional to a given reward via a sequence of actions.
Soft actor-critic for discrete action settings
Christodoulou, P. (2019) · 1910
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S. (1988) · 1988
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, L.-J. (1992) · 1992
Earlier work this paper cites.
Elements of information theory
Cover, T. M. (1999) · 1999
Earlier work this paper cites.
Dueling network architectures for deep reinforcement learning
Wang, Z., Schaul, T., Hessel, M., Hasselt, H., Lanctot, M., and Freitas, N. (2016) · 2003
Earlier work this paper cites.
Autodock vina: improving the speed and accuracy of docking with a new scoring function, efficient optimization, and multithreading
Trott, O. and Olson, A. J. (2010) · 2010
Earlier work this paper cites.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M. (2014) · 2014
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (2015) · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015) · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R. (2016) · 2016
Earlier work this paper cites.
Taming the noise in reinforcement learning via soft updates
Fox, R., Pakman, A., and Tishby, N. (2016) · 2016
Earlier work this paper cites.
Deep exploration via bootstrapped dqn
Osband, I., Blundell, C., Pritzel, A., and Van Roy, B. (2016) · 2016
Earlier work this paper cites.
Prioritized experience replay
Schaul, T., Quan, J., Antonoglou, I., and Silver, D. (2016) · 2016
Earlier work this paper cites.
Neural message passing for quantum chemistry
Gilmer, J., Schoenholz, S. S., Riley, P. F., Vinyals, O., and Dahl, G. E. (2017) · 2017
Earlier work this paper cites.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S. (2017) · 2017
Earlier work this paper cites.
A unified view of entropy-regularized markov decision processes
Neu, G., Jonsson, A., and Gómez, V. (2017) · 2017
Earlier work this paper cites.
Count-based exploration with neural density models
Ostrovski, G., Bellemare, M. G., Oord, A., and Munos, R. (2017) · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017) · 2017
Cited alongside, same era.
Efficient reinforcement learning in deterministic systems with value function generalization
Wen, Z. and Van Roy, B. (2017) · 2017
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S. (2018) · 2018
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., Lillicrap, T., Simonyan, K., and Hassabis, D. (2018) · 2018
Cited alongside, same era.
A variational perspective on generative flow networks
Zimmermann, H., Lindsten, F., van de Meent, J.-W., and Naesseth, C. A. (2022) · 2022
Later among the works it cites.
Distributional Reinforcement Learning
Bellemare, M. G., Dabney, W., and Rowland, M. (2023) · 2023
Closest in time.
Gflownet foundations
Bengio, Y., Lahlou, S., Deleu, T., Hu, E. J., Tiwari, M., and Bengio, E. (2023) · 2023
Closest in time.
Torchrl: A data-driven decision-making library for pytorch
Bou, A., Bettini, M., Dittert, S., Kumar, V., Sodhani, S., Yang, X., De Fabritiis, G., and Moens, V. (2023) · 2023
Closest in time.
Generative flow networks: a markov chain perspective
Deleu, T. and Bengio, Y. (2023) · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sutton, R. S. and Barto, A. G. (2018) · 2018
Cited alongside, same era.
Exploration by random network distillation
Burda, Y., Edwards, H., Storkey, A., and Klimov, O. (2019) · 2019
Cited alongside, same era.
A theory of regularized markov decision processes
Geist, M., Scherrer, B., and Pietquin, O. (2019) · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al. (2019) · 2019
Cited alongside, same era.
Junction tree variational autoencoder for molecular graph generation
Jin, W., Barzilay, R., and Jaakkola, T. (2020) · 2020
Cited alongside, same era.
Flow network based generative models for non-iterative diverse candidate generation
Bengio, E., Jain, M., Korablyov, M., Precup, D., and Bengio, Y. (2021) · 2021
Cited alongside, same era.
Hpc resources of the higher school of economics
Kostenetskiy, P., Chulkevich, R., and Kozyrev, V. (2021) · 2021
Cited alongside, same era.
Learning gflownets from partial episodes for improved convergence and stability
Madan, K., Rector-Brooks, J., Korablyov, M., Bengio, E., Jain, M., Nica, A. C., Bosc, T., Bengio, Y., and Malkin, N. (2023) · 2023
Closest in time.
GFlownets and variational inference
Malkin, N., Lahlou, S., Deleu, T., Ji, X., Hu, E. J., Everett, K. E., Zhang, D., and Bengio, Y. (2023) · 2023
Closest in time.
Generative augmented flow networks
Pan, L., Zhang, D., Courville, A., Huang, L., and Bengio, Y. (2023) · 2023
Closest in time.
Thompson sampling for improved exploration in gflownets
Rector-Brooks, J., Madan, K., Jain, M., Korablyov, M., Liu, C.-H., Chandar, S., Malkin, N., and Bengio, Y. (2023) · 2023
Closest in time.
Unlocking the power of representations in long-term novelty-based exploration
Saade, A., Kapturowski, S., Calandriello, D., Blundell, C., Sprechmann, P., Sarra, L., Groth, O., Valko, M., and Piot, B. (2023) · 2023
Closest in time.
Towards understanding and improving gflownet training
Shen, M. W., Bengio, E., Hajiramezanali, E., Loukas, A., Cho, K., and Biancalani, T. (2023) · 2023
Closest in time.
VA-learning as a more efficient alternative to q-learning
Tang, Y., Munos, R., Rowland, M., and Valko, M. (2023) · 2023
Closest in time.
Fast rates for maximum entropy exploration
Tiapkin, D., Belomestny, D., Calandriello, D., Moulines, E., Munos, R., Naumov, A., Perrault, P., Tang, Y., Valko, M., and Ménard, P. (2023) · 2023
Closest in time.
An empirical study of the effectiveness of using a replay buffer on mode discovery in gflownets
Vemgal, N., Lau, E., and Precup, D. (2023) · 2023
Closest in time.
Distributional gflownets with quantile flows
Zhang, D., Pan, L., Chen, R. T., Courville, A., and Bengio, Y. (2023) · 2023
Closest in time.