Fetching the paper…
Reading the bibliography…
Generative Flow Networks (GFlowNets), a new family of probabilistic samplers, have demonstrated remarkable capabilities to generate diverse sets of high-reward candidates, in contrast to standard return maximization approaches (e.g., reinforcement learning) which often converge to a single optimal solution.
Equation of state calculations by fast computing machines
Nicholas Metropolis, Arianna W Rosenbluth, Marshall N Rosenbluth, Augusta H Teller, and Edward Teller · 1953
Earlier work this paper cites.
Monte carlo sampling methods using markov chains and their applications
W Keith Hastings · 1970
Earlier work this paper cites.
Bayesian estimates of equation system parameters, an application of integration by monte carlo
Teun Kloek and Herman K. van Dijk · 1976
Earlier work this paper cites.
An introduction to mcmc for machine learning
Christophe Andrieu, Nando De Freitas, Arnaud Doucet, and Michael I Jordan · 2003
Earlier work this paper cites.
Visualizing data using t-sne
Laurens Van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
Learning in markov random fields using tempered transitions
Russ R Salakhutdinov · 2009
Earlier work this paper cites.
Viennarna package 2.0
Ronny Lorenz, Stephan H. Bernhart, Christian Höner zu Siederdissen, Hakim Tafer, Christoph Flamm, Peter F. Stadler, and Ivo L. Hofacker · 2011
Earlier work this paper cites.
Better mixing via deep representations
Yoshua Bengio, Grégoire Mesnil, Yann Dauphin, and Salah Rifai · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Universal value function approximators
Tom Schaul, Daniel Horgan, Karol Gregor, and David Silver · 2015
Earlier work this paper cites.
Empirical evaluation of rectified activations in convolutional network
Bing Xu, Naiyan Wang, Tianqi Chen, and Mu Li · 2015
Earlier work this paper cites.
Prioritized experience replay
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2016
Earlier work this paper cites.
Hindsight experience replay
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, OpenAI Pieter Abbeel, and Wojciech Zaremba · 2017
Earlier work this paper cites.
Forward-backward reinforcement learning
Ashley D Edwards, Laura Downs, and James C Davidson · 2018
Earlier work this paper cites.
Dher: Hindsight experience replay for dynamic goals
Meng Fang, Cheng Zhou, Bei Shi, Boqing Gong, Jia Xu, and Tong Zhang · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
Marlos C Machado, Marc G Bellemare, Erik Talvitie, Joel Veness, Matthew Hausknecht, and Michael Bowling · 2018
Earlier work this paper cites.
Many-goals reinforcement learning
Vivek Veeriah, Junhyuk Oh, and Satinder Singh · 2018
Earlier work this paper cites.
Goal-conditioned imitation learning
Yiming Ding, Carlos Florensa, Pieter Abbeel, and Mariano Phielipp · 2019
Earlier work this paper cites.
Curriculum-guided hindsight experience replay
Meng Fang, Tianyi Zhou, Yali Du, Lei Han, and Zhengyou Zhang · 2019
Earlier work this paper cites.
Recall traces: Backtracking models for efficient reinforcement learning
Anirudh Goyal, Philemon Brakel, William Fedus, Soumye Singhal, Timothy Lillicrap, Sergey Levine, Hugo Larochelle, and Yoshua Bengio · 2019
Cited alongside, same era.
Q-learning algorithms: A comprehensive classification and applications
Beakcheol Jang, Myeonghwi Kim, Gaspard Harerimana, and Jong Wook Kim · 2019
Cited alongside, same era.
Planning with goal-conditioned policies
Soroush Nasiriany, Vitchyr Pong, Steven Lin, and Sergey Levine · 2019
Cited alongside, same era.
Goal reasoning in the clips executive for integrated planning and execution
Tim Niemueller, Till Hofmann, and Gerhard Lakemeyer · 2019
Cited alongside, same era.
Plangan: Model-based planning with sparse rewards and multiple goals
Henry Charlesworth and Giovanni Montana · 2020
Cited alongside, same era.
C-learning: Learning to achieve goals via recursive classification
Goal-conditioned reinforcement learning: Problems and solutions
Minghuan Liu, Menghui Zhu, and Weinan Zhang · 2022
Later among the works it cites.
Trajectory balance: Improved credit assignment in GFlownets
Nikolay Malkin, Moksh Jain, Emmanuel Bengio, Chen Sun, and Yoshua Bengio · 2022
Later among the works it cites.
Rethinking goal-conditioned supervised learning and its connection to offline rl
Rui Yang, Yiming Lu, Wenzhe Li, Hao Sun, Meng Fang, Yali Du, Xiu Li, Lei Han, and Chongjie Zhang · 2022
Later among the works it cites.
Gflownet foundations
Yoshua Bengio, Salem Lahlou, Tristan Deleu, Edward J Hu, Mo Tiwari, and Emmanuel Bengio · 2023
Later among the works it cites.
Mastering diverse domains through world models
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Benjamin Eysenbach, Ruslan Salakhutdinov, and Sergey Levine · 2020
Cited alongside, same era.
One solution is not all you need: Few-shot extrapolation via structured maxent rl
Saurabh Kumar, Aviral Kumar, Sergey Levine, and Chelsea Finn · 2020
Cited alongside, same era.
Bidirectional model-based policy optimization
Hang Lai, Jian Shen, Weinan Zhang, and Yong Yu · 2020
Cited alongside, same era.
Flow network based generative models for non-iterative diverse candidate generation
Emmanuel Bengio, Moksh Jain, Maksym Korablyov, Doina Precup, and Yoshua Bengio · 2021
Cited alongside, same era.
Goal-conditioned reinforcement learning with imagined subgoals
Elliot Chane-Sane, Cordelia Schmid, and Ivan Laptev · 2021
Cited alongside, same era.
Adversarial intrinsic motivation for reinforcement learning
Ishan Durugkar, Mauricio Tec, Scott Niekum, and Peter Stone · 2021
Cited alongside, same era.
Landmark-guided subgoal generation in hierarchical reinforcement learning
Junsu Kim, Younggyo Seo, and Jinwoo Shin · 2021
Cited alongside, same era.
Backward learning for goal-conditioned policies
Marc Höftmann, Jan Robine, and Stefan Harmeling · 2023
Later among the works it cites.
Gflownets for ai-driven scientific discovery
Moksh Jain, Tristan Deleu, Jason Hartford, Cheng-Hao Liu, Alex Hernandez-Garcia, and Yoshua Bengio · 2023
Later among the works it cites.
Goal-conditioned gflownets for controllable multi-objective molecular design
Julien Roy, Pierre-Luc Bacon, Christopher Pal, and Emmanuel Bengio · 2023
Later among the works it cites.
Towards understanding and improving gflownet training
Max W Shen, Emmanuel Bengio, Ehsan Hajiramezanali, Andreas Loukas, Kyunghyun Cho, and Tommaso Biancalani · 2023
Later among the works it cites.
Prioritizing samples in reinforcement learning with reducible loss
Shivakanth Sujit, Somjit Nath, Pedro Braga, and Samira Ebrahimi Kahou · 2023
Later among the works it cites.
GOPlan: Goal-conditioned offline reinforcement learning by planning with learned models
Mianchu Wang, Rui Yang, Xi Chen, and Meng Fang · 2023
Later among the works it cites.
Human-inspired goal reasoning implementations: A survey
Ursula Addison · 2024
Closest in time.
Dyngfn: Towards bayesian inference of gene regulatory networks with gflownets
Lazar Atanackovic, Alexander Tong, Bo Wang, Leo J Lee, Yoshua Bengio, and Jason S Hartford · 2024
Closest in time.
Rectifying reinforcement learning for reward matching
Haoran He, Emmanuel Bengio, Qingpeng Cai, and Ling Pan · 2024
Closest in time.
Stitching sub-trajectories with conditional diffusion model for goal-conditioned offline rl
Sungyoon Kim, Yunseon Choi, Daiki E Matsunaga, and Kee-Eung Kim · 2024
Closest in time.
Qgfn: Controllable greediness with action values
Elaine Lau, Stephen Zhewen Lu, Ling Pan, Doina Precup, and Emmanuel Bengio · 2024
Closest in time.
Hiql: Offline goal-conditioned rl with latent states as actions
Seohong Park, Dibya Ghosh, Benjamin Eysenbach, and Sergey Levine · 2024
Closest in time.
Generative flow networks as entropy-regularized rl
Daniil Tiapkin, Nikita Morozov, Alexey Naumov, and Dmitry P Vetrov · 2024
Closest in time.
Guided cooperation in hierarchical reinforcement learning via model-based rollout
Haoran Wang, Zeshen Tang, Yaoru Sun, Fang Wang, Siyu Zhang, and Yeming Chen · 2024
Closest in time.
Let the flows tell: Solving graph combinatorial problems with gflownets
Dinghuai Zhang, Hanjun Dai, Nikolay Malkin, Aaron C Courville, Yoshua Bengio, and Ling Pan · 2024
Closest in time.