Fetching the paper…
Reading the bibliography…
We present flow Q-learning (FQL), a simple and performant offline reinforcement learning (RL) method that leverages an expressive flow-matching policy to model arbitrarily complex action distributions in data.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2005
Earlier work this paper cites.
Reinforcement learning by reward-weighted regression for operational space control
Peters, J. and Schaal, S · 2007
Earlier work this paper cites.
Batch reinforcement learning
Lange, S., Gabel, T., and Riedmiller, M · 2012
Earlier work this paper cites.
Introduction to Smooth Manifolds
Lee, J. M · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. A · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Ba, J., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Hendrycks, D. and Gimpel, K · 2016
Earlier work this paper cites.
Wasserstein generative adversarial networks
Arjovsky, M., Chintala, S., and Bottou, L · 2017
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q · 2018
Earlier work this paper cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., Legg, S., and Kavukcuoglu, K · 2018
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Fujimoto, S., van Hoof, H., and Meger, D · 2018
Earlier work this paper cites.
Leveraging exploration in off-policy algorithms via normalizing flows
Mazoure, B., Doan, T., Durand, A., Pineau, J., and Hjelm, R. D · 2019
Earlier work this paper cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Peng, X. B., Kumar, A., Zhang, G., and Levine, S · 2019
Earlier work this paper cites.
Behavior regularized offline reinforcement learning
Wu, Y., Tucker, G., and Nachum, O · 2019
Earlier work this paper cites.
D4rl: Datasets for deep data-driven reinforcement learning
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Morel : Model-based offline reinforcement learning
Kidambi, R., Rajeswaran, A., Netrapalli, P., and Joachims, T · 2020
Earlier work this paper cites.
Conservative q-learning for offline reinforcement learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S · 2020
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J · 2020
Earlier work this paper cites.
Accelerating online reinforcement learning with offline datasets
Nair, A., Dalal, M., Gupta, A., and Levine, S · 2020
Earlier work this paper cites.
Mopo: Model-based offline policy optimization
Yu, T., Thomas, G., Yu, L., Ermon, S., Zou, J. Y., Levine, S., Finn, C., and Ma, T · 2020
Earlier work this paper cites.
Uncertainty-based offline reinforcement learning with diversified q-ensemble
An, G., Moon, S., Kim, J.-H., and Song, H. O · 2021
Earlier work this paper cites.
Decision transformer: Reinforcement learning via sequence modeling
Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I · 2021
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
Dhariwal, P. and Nichol, A · 2021
Earlier work this paper cites.
A minimalist approach to offline reinforcement learning
Fujimoto, S. and Gu, S. S · 2021
Earlier work this paper cites.
Reinforcement learning as one big sequence modeling problem
Janner, M., Li, Q., and Levine, S · 2021
Cited alongside, same era.
What matters in learning from offline human demonstrations for robot manipulation
Mandlekar, A., Xu, D., Wong, J., Nasiriany, S., Wang, C., Kulkarni, R., Fei-Fei, L., Savarese, S., Zhu, Y., and Mart’in-Mart’in, R · 2021
Cited alongside, same era.
Score-based generative modeling through stochastic differential equations
Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B · 2021
Cited alongside, same era.
A closer look at offline rl agents
Fu, Y., Wu, D., and Boulet, B · 2022
Cited alongside, same era.
Planning with diffusion for flexible behavior synthesis
Janner, M., Du, Y., Tenenbaum, J. B., and Levine, S · 2022
Cited alongside, same era.
Offline reinforcement learning with implicit q-learning
Kostrikov, I., Nair, A., and Levine, S · 2022
Guided flows for generative modeling and decision making
Zheng, Q., Le, M., Shaul, N., Lipman, Y., Grover, A., and Chen, R. T · 2023
Later among the works it cites.
Diffusion policies for out-of-distribution generalization in offline reinforcement learning
Ada, S. E., Oztop, E., and Ugur, E · 2024
Later among the works it cites.
Diffusion for world modeling: Visual details matter in atari
Alonso, E., Jelley, A., Micheli, V., Kanervisto, A., Storkey, A., Pearce, T., and Fleuret, F · 2024
Later among the works it cites.
π 0 \pi_{0} : A vision-language-action flow model for general robot control
Black, K., Brown, N., Driess, D., Esmail, A., Equi, M., Finn, C., Fusai, N., Groom, L., Hausman, K., Ichter, B., et al · 2024
Later among the works it cites.
Open x-embodiment: Robotic learning datasets and rt-x models
Collaboration, O. X.-E., O’Neill, A., Rehman, A., Maddukuri, A., Gupta, A., Padalkar, A., Lee, A., Pooley, A., Gupta, A., Mandlekar, A., Jain, A., et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Is conditional generative modeling all you need for decision-making?
Ajay, A., Du, Y., Gupta, A., Tenenbaum, J., Jaakkola, T., and Agrawal, P · 2023
Cited alongside, same era.
Building normalizing flows with stochastic interpolants
Albergo, M. S. and Vanden-Eijnden, E · 2023
Cited alongside, same era.
Efficient online reinforcement learning with offline data
Ball, P. J., Smith, L., Kostrikov, I., and Levine, S · 2023
Cited alongside, same era.
Offline reinforcement learning via high-fidelity generative behavior modeling
Chen, H., Lu, C., Ying, C., Su, H., and Zhu, J · 2023
Cited alongside, same era.
Extreme q-learning: Maxent rl without entropy
Garg, D., Hejna, J., Geist, M., and Ermon, S · 2023
Cited alongside, same era.
Idql: Implicit q-learning as an actor-critic method with diffusion policies
Hansen-Estruch, P., Kostrikov, I., Janner, M., Kuba, J. G., and Levine, S · 2023
Cited alongside, same era.
Later among the works it cites.
Consistency models as a rich and efficient policy class for reinforcement learning
Ding, Z. and Jin, C · 2024
Later among the works it cites.
Scaling rectified flow transformers for high-resolution image synthesis
Esser, P., Kulal, S., Blattmann, A., Entezari, R., Müller, J., Saini, H., Levi, Y., Lorenz, D., Sauer, A., Boesel, F., et al · 2024
Later among the works it cites.
Diffusion meets flow matching: Two sides of the same coin, 2024
Gao, R., Hoogeboom, E., Heek, J., Bortoli, V. D., Murphy, K. P., and Salimans, T · 2024
Later among the works it cites.
Aligniql: Policy alignment in implicit q-learning through constrained optimization
He, L., Shen, L., Tan, J., and Wang, X · 2024
Later among the works it cites.
Policy-guided diffusion
Jackson, M. T., Matthews, M. T., Lu, C., Ellis, B., Whiteson, S., and Foerster, J · 2024
Later among the works it cites.
Lipman, Y., Havasi, M., Holderrieth, P., Shaul, N., Le, M., Karrer, B., Chen, R. T. Q., Lopez-Paz, D., Ben-Hamu, H., and Gat, I · 2024
Later among the works it cites.
Instaflow: One step is enough for high-quality diffusion-based text-to-image generation
Liu, X., Zhang, X., Ma, J., Peng, J., et al · 2024
Later among the works it cites.
Diffusion-dice: In-sample diffusion guidance for offline reinforcement learning
Mao, L., Xu, H., Zhan, X., Zhang, W., and Zhang, A · 2024
Later among the works it cites.
Policy agnostic rl: Offline rl and online rl fine-tuning of any class and backbone
Mark, M. S., Gao, T., Sampaio, G. G., Srirama, M. K., Sharma, A., Finn, C., and Kumar, A · 2024
Later among the works it cites.
Bigger, regularized, optimistic: scaling for compute and sample-efficient continuous control
Nauman, M., Ostaszewski, M., Jankowski, K., Miłoś, P., and Cygan, M · 2024
Later among the works it cites.
Learning a diffusion model policy from rewards via q-score matching
Psenka, M., Escontrela, A., Abbeel, P., and Ma, Y · 2024
Later among the works it cites.
D5rl: Diverse datasets for data-driven deep reinforcement learning
Rafailov, R., Hatch, K. B., Singh, A., Kumar, A., Smith, L., Kostrikov, I., Hansen-Estruch, P., Kolev, V., Ball, P. J., Wu, J., et al · 2024
Later among the works it cites.
Dual rl: Unification and new methods for reinforcement and imitation learning
Sikchi, H. S., Zheng, Q., Zhang, A., and Niekum, S · 2024
Later among the works it cites.
Reasoning with latent diffusion in offline reinforcement learning
Venkatraman, S., Khaitan, S., Akella, R. T., Dolan, J., Schneider, J., and Berseth, G · 2024
Later among the works it cites.
Entropy-regularized diffusion policy with q-ensembles for offline reinforcement learning
Zhang, R., Luo, Z., Sjölund, J., Schön, T. B., and Mattsson, P · 2024
Later among the works it cites.
Diffusion actor-critic: Formulating constrained policy iteration as diffusion noise regression for offline reinforcement learning
Fang, L., Liu, R., Zhang, J., Wang, W., and Jing, B · 2025
Closest in time.
One step diffusion via shortcut models
Frans, K., Hafner, D., Levine, S., and Abbeel, P · 2025
Closest in time.
Simba: Simplicity bias for scaling up parameters in deep reinforcement learning
Lee, H., Hwang, D., Kim, D., Kim, H., Tai, J. J., Subramanian, K., Wurman, P. R., Choo, J., Stone, P., and Seno, T · 2025
Closest in time.
Ogbench: Benchmarking offline goal-conditioned rl
Park, S., Frans, K., Eysenbach, B., and Levine, S · 2025
Closest in time.
Diffusion policy policy optimization
Ren, A. Z., Lidard, J., Ankile, L. L., Simeonov, A., Agrawal, P., Majumdar, A., Burchfiel, B., Dai, H., and Simchowitz, M · 2025
Closest in time.
Energy-weighted flow matching for offline reinforcement learning
Zhang, S., Zhang, W., and Gu, Q · 2025
Closest in time.