Fetching the paper…
Reading the bibliography…
Generative flow networks (GFlowNets) are a family of algorithms for training a sequential sampler of discrete objects under an unnormalized target density and have been successfully used for various probabilistic modeling tasks.
Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
Williams, R. and Peng, J · 1991
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Baird, L · 1995
Earlier work this paper cites.
Bias-variance error bounds for temporal difference updates
Kearns, M. J. and Singh, S. P · 2000
Earlier work this paper cites.
Protein structure prediction using Rosetta
Rohl, C. A., Strauss, C. E., Misura, K. M., and Baker, D · 2004
Earlier work this paper cites.
PyRosetta: a script-based interface for implementing molecular modeling algorithms using Rosetta
Chaudhury, S., Lyskov, S., and Gray, J. J · 2010
Earlier work this paper cites.
AutoDock Vina: improving the speed and accuracy of docking with a new scoring function, efficient optimization, and multithreading
Trott, O. and Olson, A. J · 2010
Earlier work this paper cites.
Fragment based drug design: from experimental to computational approaches
Kumar, A., Voet, A., and Zhang, K. Y · 2012
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Local fitness landscape of the green fluorescent protein
Sarkisyan, K. S., Bolotin, D. A., Meer, M. V., Usmanova, D. R., Mishin, A. S., Sharonov, G. V., Ivankov, D. N., Bozhanova, N. G., Baranov, M. S., Soylemez, O., et al · 2016
Earlier work this paper cites.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Earlier work this paper cites.
Bridging the gap between value and policy based reinforcement learning
Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Diversity is all you need: Learning skills without a reward function
Eysenbach, B., Gupta, A., Ibarz, J., and Levine, S · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Cited alongside, same era.
Deep reinforcement learning and the deadly triad
van Hasselt, H., Doron, Y., Strub, F., Hessel, M., Sonnerat, N., and Modayil, J · 2018
Cited alongside, same era.
Conditioning by adaptive sampling for robust design
Brookes, D., Park, H., and Listgarten, J · 2019
Cited alongside, same era.
Soft actor-critic for discrete action settings
Christodoulou, P · 2019
Cited alongside, same era.
Dbaasp v3: Database of antimicrobial/cytotoxic activity and structure of peptides as a resource for development of new therapeutics
Pirtskhalava, M., Amstrong, A. A., Grigolava, M., Chubinidze, M., Alimbarashvili, E., Vishnepolsky, B., Gabrielian, A., Rosenthal, A., Hurt, D. E., and Tartakovsky, M · 2021
Later among the works it cites.
MARS: Markov molecular sampling for multi-objective drug discovery
Xie, Y., Shi, C., Zhou, H., Yang, Y., Zhang, W., Yu, Y., and Li, L · 2021
Later among the works it cites.
Exploration by maximizing Rényi entropy for reward-free RL framework
Zhang, C., Cai, Y., Huang, L., and Li, J · 2021
Later among the works it cites.
Bayesian structure learning with generative flow networks
Deleu, T., Góis, A., Emezue, C., Rankawat, M., Lacoste-Julien, S., Bauer, S., and Bengio, Y · 2022
Closest in time.
Biological sequence design with GFlowNets
Jain, M., Bengio, E., Hernandez-Garcia, A., Rector-Brooks, J., Dossou, B. F., Ekbote, C., Fu, J., Zhang, T., Kilgour, M., Zhang, D., Simine, L., Das, P., and Bengio, Y · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Provably efficient maximum entropy exploration
Hazan, E., Kakade, S., Singh, K., and Van Soest, A · 2019
Cited alongside, same era.
Marginalized state distribution entropy regularization in policy optimization
Islam, R., Ahmed, Z., and Precup, D · 2019
Cited alongside, same era.
Interference and generalization in temporal difference learning
Bengio, E., Pineau, J., and Precup, D · 2020
Cited alongside, same era.
A closer look at deep policy gradients
Ilyas, A., Engstrom, L., Santurkar, S., Tsipras, D., Janoos, F., Rudolph, L., and Madry, A · 2020
Cited alongside, same era.
AdaLead: A simple and robust adaptive greedy search algorithm for sequence design
Sinai, S., Wang, R., Whatley, A., Slocum, S., Locane, E., and Kelsic, E · 2020
Cited alongside, same era.
Deep residual reinforcement learning
Zhang, S., Boehmer, W., and Whiteson, S · 2020
Cited alongside, same era.
Trajectory balance: Improved credit assignment in GFlowNets
Malkin, N., Jain, M., Bengio, E., Sun, C., and Bengio, Y · 2022
Closest in time.
Design-bench: Benchmarks for data-driven offline model-based optimization
Trabucco, B., Geng, X., Kumar, A., and Levine, S · 2022
Closest in time.
Generative flow networks for discrete probabilistic modeling
Zhang, D., Malkin, N., Liu, Z., Volokhova, A., Courville, A., and Bengio, Y · 2022
Closest in time.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Ziebart, B. D · 2022
Closest in time.
GFlowNet-EM for learning compositional latent variable models
Hu, E. J., Malkin, N., Jain, M., Everett, K., Graikos, A., and Bengio, Y · 2023
Closest in time.
GFlowNets and variational inference
Malkin, N., Lahlou, S., Deleu, T., Ji, X., Hu, E., Everett, K., Zhang, D., and Bengio, Y · 2023
Closest in time.
Better training of GFlowNets with local credit and incomplete trajectories
Pan, L., Malkin, N., Zhang, D., and Bengio, Y · 2023
Closest in time.
Chapter 11. junction tree variational autoencoder for molecular graph generation
Jin, W., Barzilay, R., and Jaakkola, T · 2041
Closest in time.