Fetching the paper…
Reading the bibliography…
Reinforcement Learning (RL) algorithms aim to learn an optimal policy by iteratively sampling actions to learn how to maximize the total expected return, $R(x)$.
Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, L.-J · 1992
Earlier work this paper cites.
Rdkit: Open-source cheminformatics. 2006
Landrum, G · 2006
Earlier work this paper cites.
Autodock vina: improving the speed and accuracy of docking with a new scoring function, efficient optimization, and multithreading
Trott, O. and Olson, A. J · 2010
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
Schaul, T., Quan, J., Antonoglou, I., and Silver, D · 2015
Cited alongside, same era.
Neural message passing for quantum chemistry
Gilmer, J., Schoenholz, S. S., Riley, P. F., Vinyals, O., and Dahl, G. E · 2017
Cited alongside, same era.
Chapter 11: Junction tree variational autoencoder for molecular graph generation
Jin, W., Barzilay, R., and Jaakkola, T · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al · 2019
Cited alongside, same era.
Flow network based generative models for non-iterative diverse candidate generation
Bengio, E., Jain, M., Korablyov, M., Precup, D., and Bengio, Y
Cited in the paper.
Bengio, Y., Lahlou, S., Deleu, T., Hu, E. J., Tiwari, M., and Bengio, E
Cited in the paper.
Python 3 Reference Manual
Van Rossum, G. and Drake, F. L · 2019
Later among the works it cites.
Learning gflownets from partial episodes for improved convergence and stability
Madan, K., Rector-Brooks, J., Korablyov, M., Bengio, E., Jain, M., Nica, A., Bosc, T., Bengio, Y., and Malkin, N · 2022
Later among the works it cites.
Trajectory balance: Improved credit assignment in gflownets
Malkin, N., Jain, M., Bengio, E., Sun, C., and Bengio, Y · 2022
Later among the works it cites.
Towards understanding and improving gflownet training
Shen, M. W., Bengio, E., Hajiramezanali, E., Loukas, A., Cho, K., and Biancalani, T · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…