Fetching the paper…
Reading the bibliography…
Offline reinforcement learning (RL) enables policy learning from pre-collected offline datasets, relaxing the need to interact directly with the environment.
Computer methods for ordinary differential equations and differential-algebraic equations
Uri M Ascher and Linda R Petzold · 1998
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2007
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
Aravind Rajeswaran, Vikash Kumar, Abhishek Gupta, Giulia Vezzani, John Schulman, Emanuel Todorov, and Sergey Levine · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Sigmoid-weighted linear units for neural network function approximation in reinforcement learning
Stefan Elfwing, Eiji Uchibe, and Kenji Doya · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2019
Earlier work this paper cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
Aviral Kumar, Justin Fu, Matthew Soh, George Tucker, and Sergey Levine · 2019
Earlier work this paper cites.
Cityflow: A multi-agent reinforcement learning environment for large scale city traffic scenario
Huichu Zhang, Siyuan Feng, Chang Liu, Yaoyao Ding, Yichen Zhu, Zihan Zhou, Weinan Zhang, Yong Yu, Haiming Jin, and Zhenhui Li · 2019
Earlier work this paper cites.
D4rl: Datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
Conservative q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Earlier work this paper cites.
Counterfactual data augmentation using locally factored dynamics
Silviu Pitis, Elliot Creager, and Animesh Garg · 2020
Cited alongside, same era.
Score-based generative modeling through stochastic differential equations
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole · 2020
Cited alongside, same era.
Decision transformer: Reinforcement learning via sequence modeling
Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Misha Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch · 2021
Cited alongside, same era.
A minimalist approach to offline reinforcement learning
Scott Fujimoto and Shixiang Shane Gu · 2021
Cited alongside, same era.
Offline reinforcement learning with fisher divergence critic regularization
Ilya Kostrikov, Rob Fergus, Jonathan Tompson, and Ofir Nachum · 2021
Cited alongside, same era.
Data might be enough: Bridge real-world traffic signal control using offline reinforcement learning
Liang Zhang and Jianming Deng · 2023
Later among the works it cites.
Robotic offline rl from internet videos via value-function learning
Chethan Bhateja, Derek Guo, Dibya Ghosh, Anikait Singh, Manan Tomar, Quan Vuong, Yevgen Chebotar, Sergey Levine, and Aviral Kumar · 2024
Closest in time.
Simple hierarchical planning with diffusion
Chang Chen, Fei Deng, Kenji Kawaguchi, Caglar Gulcehre, and Sungjin Ahn · 2024
Closest in time.
Guided data augmentation for offline reinforcement learning and imitation learning
Nicholas E Corrado, Yuxiao Qu, John U Balis, Adam Labiosa, and Josiah P Hanna · 2024
Closest in time.
Felight: Fairness-aware traffic signal control via sample-efficient reinforcement learning
Xinqi Du, Ziyue Li, Cheng Long, Yongheng Xing, S Yu Philip, and Hechang Chen · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Combo: Conservative offline model-based policy optimization
Tianhe Yu, Aviral Kumar, Rafael Rafailov, Aravind Rajeswaran, Sergey Levine, and Chelsea Finn · 2021
Cited alongside, same era.
Planning with diffusion for flexible behavior synthesis
Michael Janner, Yilun Du, Joshua B Tenenbaum, and Sergey Levine · 2022
Cited alongside, same era.
Elucidating the design space of diffusion-based generative models
Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine · 2022
Cited alongside, same era.
Mildly conservative q-learning for offline reinforcement learning
Jiafei Lyu, Xiaoteng Ma, Xiu Li, and Zongqing Lu · 2022
Cited alongside, same era.
A policy-guided imitation approach for offline reinforcement learning
Haoran Xu, Li Jiang, Li Jianxiong, and Xianyuan Zhan · 2022
Cited alongside, same era.
Expression might be enough: representing pressure and demand for reinforcement learning based traffic signal control
Liang Zhang, Qiang Wu, Jun Shen, Linyuan Lü, Bo Du, and Jianqing Wu · 2022
Cited alongside, same era.
Prompt-tuning decision transformer with preference ranking
Shengchao Hu, Li Shen, Ya Zhang, and Dacheng Tao · 2023
Cited alongside, same era.
Act: Empowering decision transformer with dynamic programming via advantage conditioning
Chen-Xiao Gao, Chenyang Wu, Mingjun Cao, Rui Kong, Zongzhang Zhang, and Yang Yu · 2024
Closest in time.
Diffstitch: Boosting offline reinforcement learning with diffusion-based trajectory stitching
Guanghe Li, Yixiang Shan, Zhengbang Zhu, Ting Long, and Weinan Zhang · 2024
Closest in time.
Synthetic experience replay
Cong Lu, Philip Ball, Yee Whye Teh, and Jack Parker-Holder · 2024
Closest in time.
Doubly mild generalization for offline reinforcement learning
Yixiu Mao, Qi Wang, Yun Qu, Yuhang Jiang, and Xiangyang Ji · 2024
Closest in time.
Corl: Research-oriented deep offline reinforcement learning library
Denis Tarasov, Alexander Nikulin, Dmitry Akimov, Vladislav Kurenkov, and Sergey Kolesnikov · 2024
Closest in time.
Efficient exploration in continuous-time model-based reinforcement learning
Lenart Treven, Jonas Hübotter, Florian Dorfler, and Andreas Krause · 2024
Closest in time.
Elastic decision transformer
Yueh-Hua Wu, Xiaolong Wang, and Masashi Hamaya · 2024
Closest in time.
Dual diffusion for unified image generation and understanding
Zijie Li, Henry Li, Yichun Shi, Amir Barati Farimani, Yuval Kluger, Linjie Yang, and Peng Wang · 2025
Closest in time.
Behaviorally-aware multi-agent rl with dynamic optimization for autonomous driving
Hamid Taghavifar, Chuan Hu, Chongfeng Wei, Ardashir Mohammadzadeh, and Chunwei Zhang · 2025
Closest in time.
Deep reinforcement learning for robotics: A survey of real-world successes
Chen Tang, Ben Abbatematteo, Jiaheng Hu, Rohan Chandra, Roberto Martín-Martín, and Peter Stone · 2025
Closest in time.
Hgat and multi-agent rl-based method for multi-intersection traffic signal control
Ziyang Zhai, Ruru Hao, Boyang Cui, and Siyi Wang · 2025
Closest in time.