Fetching the paper…
Reading the bibliography…
Despite recent advances in improving the sample-efficiency of reinforcement learning (RL) algorithms, designing an RL algorithm that can be practically deployed in real-world environments remains a challenge.
Admissibility and measurable utility functions
J. P. Quirk and R. Saposnik · 1962
Earlier work this paper cites.
Rules for ordering uncertain prospects
J. Hadar and W. R. Russell · 1969
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
B. T. Polyak and A. B. Juditsky · 1992
Earlier work this paper cites.
Feudal reinforcement learning
P. Dayan and G. E. Hinton · 1992
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
L. P. Kaelbling, M. L. Littman, and A. R. Cassandra · 1998
Earlier work this paper cites.
Actor-critic algorithms
V. Konda and J. Tsitsiklis · 1999
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
R. S. Sutton, D. Precup, and S. Singh · 1999
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
M. Deisenroth and C. E. Rasmussen · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. Gordon, and D. Bagnell · 2011
Earlier work this paper cites.
V-rep: A versatile and scalable robot simulation framework
E. Rohmer, S. P. Singh, and M. Freese · 2013
Earlier work this paper cites.
Deterministic policy gradient algorithms
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller · 2014
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
K. Cho, B. Van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Earlier work this paper cites.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Earlier work this paper cites.
T. Schaul, J. Quan, I. Antonoglou, and D. Silver · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2016
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Earlier work this paper cites.
Benchmarking deep reinforcement learning for continuous control
Y. Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel · 2016
Earlier work this paper cites.
Dueling network architectures for deep reinforcement learning
Z. Wang, T. Schaul, M. Hessel, H. Hasselt, M. Lanctot, and N. Freitas · 2016
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
D. Hendrycks and K. Gimpel · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Mastering the game of go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al · 2017
Earlier work this paper cites.
Discrete sequential prediction of continuous actions for deep rl
L. Metz, J. Ibarz, N. Jaitly, and J. Davidson · 2017
Earlier work this paper cites.
Hybrid reward architecture for reinforcement learning
H. Van Seijen, M. Fatemi, J. Romoff, R. Laroche, T. Barnes, and J. Tsang · 2017
Earlier work this paper cites.
A distributional perspective on reinforcement learning
M. G. Bellemare, W. Dabney, and R. Munos · 2017
Earlier work this paper cites.
Feudal networks for hierarchical reinforcement learning
A. S. Vezhnevets, S. Osindero, T. Schaul, N. Heess, M. Jaderberg, D. Silver, and K. Kavukcuoglu · 2017
Earlier work this paper cites.
Learning multi-level hierarchies with hindsight
A. Levy, G. Konidaris, R. Platt, and K. Saenko · 2017
Earlier work this paper cites.
Stochastic neural networks for hierarchical reinforcement learning
C. Florensa, Y. Duan, and P. Abbeel · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Scalable deep reinforcement learning for vision-based robotic manipulation
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke, et al · 2018
Earlier work this paper cites.
Soft actor-critic algorithms and applications
T. Haarnoja, A. Zhou, K. Hartikainen, G. Tucker, S. Ha, J. Tan, V. Kumar, H. Zhu, A. Gupta, P. Abbeel, et al · 2018
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
S. Fujimoto, H. Hoof, and D. Meger · 2018
Earlier work this paper cites.
Deep reinforcement learning that matters
P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, and D. Meger · 2018
Earlier work this paper cites.
Action branching architectures for deep reinforcement learning
A. Tavakoli, F. Pardo, and P. Kormushev · 2018
Earlier work this paper cites.
Learn what not to learn: Action elimination with deep reinforcement learning
T. Zahavy, M. Haroush, N. Merlis, D. J. Mankowitz, and S. Mannor · 2018
Earlier work this paper cites.
Soft actor-critic algorithms and applications
T. Haarnoja, A. Zhou, K. Hartikainen, G. Tucker, S. Ha, J. Tan, V. Kumar, H. Zhu, A. Gupta, P. Abbeel, et al · 2018
Earlier work this paper cites.
Sim-to-real reinforcement learning for deformable object manipulation
J. Matas, S. James, and A. J. Davison · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Earlier work this paper cites.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
A. Rajeswaran, V. Kumar, A. Gupta, G. Vezzani, J. Schulman, E. Todorov, and S. Levine · 2018
Earlier work this paper cites.
Deep q-learning from demonstrations
T. Hester, M. Vecerik, O. Pietquin, M. Lanctot, T. Schaul, B. Piot, D. Horgan, J. Quan, A. Sendonaris, I. Osband, et al · 2018
Earlier work this paper cites.
Self-imitation learning
J. Oh, Y. Guo, S. Singh, and H. Lee · 2018
Earlier work this paper cites.
Noisy networks for exploration
M. Fortunato, M. G. Azar, B. Piot, J. Menick, M. Hessel, I. Osband, A. Graves, V. Mnih, R. Munos, D. Hassabis, O. Pietquin, C. Blundell, and S. Legg · 2018
Earlier work this paper cites.
Parameter space noise for exploration
M. Plappert, R. Houthooft, P. Dhariwal, S. Sidor, R. Y. Chen, X. Chen, T. Asfour, P. Abbeel, and M. Andrychowicz · 2018
Earlier work this paper cites.
Averaging weights leads to wider optima and better generalization
P. Izmailov, D. Podoprikhin, T. Garipov, D. Vetrov, and A. G. Wilson · 2018
Earlier work this paper cites.
Time-contrastive networks: Self-supervised learning from video
P. Sermanet, C. Lynch, Y. Chebotar, J. Hsu, E. Jang, S. Schaal, S. Levine, and G. Brain · 2018
Earlier work this paper cites.
Sim-to-real: Learning agile locomotion for quadruped robots
J. Tan, T. Zhang, E. Coumans, A. Iscen, Y. Bai, D. Hafner, S. Bohez, and V. Vanhoucke · 2018
Earlier work this paper cites.
Learning to walk via deep reinforcement learning
T. Haarnoja, S. Ha, A. Zhou, J. Tan, G. Tucker, and S. Levine · 2018
Cited alongside, same era.
Data-efficient hierarchical reinforcement learning
O. Nachum, S. S. Gu, H. Lee, and S. Levine · 2018
Cited alongside, same era.
Diversity is all you need: Learning skills without a reward function
B. Eysenbach, A. Gupta, J. Ibarz, and S. Levine · 2018
Cited alongside, same era.
Learning by playing solving sparse reward tasks from scratch
M. Riedmiller, R. Hafner, T. Lampe, M. Neunert, J. Degrave, T. Wiele, V. Mnih, N. Heess, and J. T. Springenberg · 2018
Cited alongside, same era.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2019
Cited alongside, same era.
Diffusion policy: Visuomotor policy learning via action diffusion
C. Chi, S. Feng, Y. Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song · 2023
Later among the works it cites.
Perceiver-actor: A multi-task transformer for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2023
Later among the works it cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn, et al · 2023
Later among the works it cites.
Solving continuous control via q-learning
T. Seyde, P. Werner, W. Schwarting, I. Gilitschenski, M. Riedmiller, D. Rus, and M. Wulfmeier · 2023
Later among the works it cites.
Daydreamer: World models for physical robot learning
P. Wu, A. Escontrela, D. Hafner, P. Abbeel, and K. Goldberg · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. James, M. Freese, and A. J. Davison · 2019
Cited alongside, same era.
Learning agile and dynamic motor skills for legged robots
J. Hwangbo, J. Lee, A. Dosovitskiy, D. Bellicoso, V. Tsounis, V. Koltun, and M. Hutter · 2019
Cited alongside, same era.
Reinforcement learning on variable impedance controller for high-precision robotic assembly
J. Luo, E. Solowjow, C. Wen, J. A. Ojea, A. M. Agogino, A. Tamar, and P. Abbeel · 2019
Cited alongside, same era.
Residual reinforcement learning for robot control
T. Johannink, S. Bahl, A. Nair, J. Luo, A. Kumar, M. Loskyll, J. A. Ojea, E. Solowjow, and S. Levine · 2019
Cited alongside, same era.
Rlbench: The robot learning benchmark & learning environment
S. James, Z. Ma, D. R. Arrojo, and A. J. Davison · 2020
Cited alongside, same era.
Autonomous navigation of stratospheric balloons using reinforcement learning
M. G. Bellemare, S. Candido, P. S. Castro, J. Gong, M. C. Machado, S. Moitra, S. S. Ponda, and Z. Wang · 2020
Cited alongside, same era.
Mastering atari, go, chess and shogi by planning with a learned model
J. Schrittwieser, I. Antonoglou, T. Hubert, K. Simonyan, L. Sifre, S. Schmitt, A. Guez, E. Lockhart, D. Hassabis, T. Graepel, et al · 2020
Cited alongside, same era.
Action-quantized offline reinforcement learning for robotic skill learning
J. Luo, P. Dong, J. Wu, A. Kumar, X. Geng, and S. Levine · 2023
Later among the works it cites.
Mastering diverse domains through world models
D. Hafner, J. Pasukonis, J. Ba, and T. Lillicrap · 2023
Later among the works it cites.
Bigger, better, faster: Human-level atari with human-level efficiency
M. Schwarzer, J. S. O. Ceron, A. Courville, M. G. Bellemare, R. Agarwal, and P. S. Castro · 2023
Later among the works it cites.
Modem: Accelerating visual model-based reinforcement learning with demonstrations
N. Hansen, Y. Lin, H. Su, X. Wang, V. Kumar, and A. Rajeswaran · 2023
Later among the works it cites.
Efficient online reinforcement learning with offline data
P. J. Ball, L. Smith, I. Kostrikov, and S. Levine · 2023
Later among the works it cites.
Multi-view masked world models for visual robotic manipulation
Y. Seo, J. Kim, S. James, K. Lee, J. Shin, and P. Abbeel · 2023
Later among the works it cites.
Sample-efficient reinforcement learning by breaking the replay ratio barrier
P. D’Oro, M. Schwarzer, E. Nikishin, P.-L. Bacon, M. G. Bellemare, and A. Courville · 2023
Later among the works it cites.
Imitation bootstrapped reinforcement learning
H. Hu, S. Mirchandani, and D. Sadigh · 2023
Later among the works it cites.
Gnfactor: Multi-task real robot learning with generalizable neural feature fields
Y. Ze, G. Yan, Y.-H. Wu, A. Macaluso, Y. Ge, J. Ye, N. Hansen, L. E. Li, and X. Wang · 2023
Later among the works it cites.
Visual reinforcement learning with self-supervised 3d representations
Y. Ze, N. Hansen, Y. Chen, M. Jain, and X. Wang · 2023
Later among the works it cites.
Act3d: Infinite resolution action detection transformer for robotic manipulation
T. Gervet, Z. Xian, N. Gkanatsios, and K. Fragkiadaki · 2023
Later among the works it cites.
Pirlnav: Pretraining with imitation and rl finetuning for objectnav
R. Ramrakhya, D. Batra, E. Wijmans, and A. Das · 2023
Later among the works it cites.
Real-world robot learning with masked visual pre-training
I. Radosavovic, T. Xiao, S. James, P. Abbeel, J. Malik, and T. Darrell · 2023
Later among the works it cites.
Legged locomotion in challenging terrains using egocentric vision
A. Agarwal, A. Kumar, J. Malik, and D. Pathak · 2023
Later among the works it cites.
A walk in the park: Learning to walk in 20 minutes with model-free reinforcement learning
L. Smith, I. Kostrikov, and S. Levine · 2023
Later among the works it cites.
Reboot: Reuse data for bootstrapping efficient real-world dexterous manipulation
Z. Hu, A. Rovinsky, J. Luo, V. Kumar, A. Gupta, and S. Levine · 2023
Later among the works it cites.
Furniturebench: Reproducible real-world benchmark for long-horizon complex manipulation
M. Heo, Y. Lee, D. Lee, and J. J. Lim · 2023
Later among the works it cites.
Maniskill2: A unified benchmark for generalizable manipulation skills
J. Gu, F. Xiang, X. Li, Z. Ling, X. Liu, T. Mu, Y. Tang, S. Tao, X. Wei, Y. Yao, et al · 2023
Later among the works it cites.
Decomposing the generalization gap in imitation learning for visual robotic manipulation
A. Xie, L. Lee, T. Xiao, and C. Finn · 2023
Later among the works it cites.
Scaling robot learning with semantically imagined experience
T. Yu, T. Xiao, A. Stone, J. Tompson, A. Brohan, S. Wang, J. Singh, C. Tan, J. Peralta, B. Ichter, et al · 2023
Later among the works it cites.
Lrm: Large reconstruction model for single image to 3d
Y. Hong, K. Zhang, J. Gu, S. Bi, Y. Zhou, D. Liu, F. Liu, K. Sunkavalli, T. Bui, and H. Tan · 2023
Later among the works it cites.
Dmv3d: Denoising multi-view diffusion using 3d large reconstruction model
Y. Xu, H. Tan, F. Luan, S. Bi, P. Wang, J. Li, Z. Shi, K. Sunkavalli, G. Wetzstein, Z. Xu, et al · 2023
Later among the works it cites.
Shap-e: Generating conditional 3d implicit functions
H. Jun and A. Nichol · 2023
Later among the works it cites.
Gta: A geometry-aware attention mechanism for multi-view transformers
T. Miyato, B. Jaeger, M. Welling, and A. Geiger · 2023
Later among the works it cites.
Objaverse: A universe of annotated 3d objects
M. Deitke, D. Schwenk, J. Salvador, L. Weihs, O. Michel, E. VanderBilt, L. Schmidt, K. Ehsani, A. Kembhavi, and A. Farhadi · 2023
Later among the works it cites.
Mvimgnet: A large-scale dataset of multi-view images
X. Yu, M. Xu, Y. Zhang, H. Liu, C. Ye, Y. Wu, Z. Yan, C. Zhu, Z. Xiong, T. Liang, et al · 2023
Later among the works it cites.
Td-mpc2: Scalable, robust world models for continuous control
N. Hansen, H. Su, and X. Wang · 2023
Later among the works it cites.
Mamba: Linear-time sequence modeling with selective state spaces
A. Gu and T. Dao · 2023
Later among the works it cites.
Bridgedata v2: A dataset for robot learning at scale
H. R. Walke, K. Black, T. Z. Zhao, Q. Vuong, C. Zheng, P. Hansen-Estruch, A. W. He, V. Myers, M. J. Kim, M. Du, et al · 2023
Later among the works it cites.
Open x-embodiment: Robotic learning datasets and rt-x models
A. Padalkar, A. Pooley, A. Jain, A. Bewley, A. Herzog, A. Irpan, A. Khazatsky, A. Rai, A. Singh, A. Brohan, et al · 2023
Later among the works it cites.
Preference transformer: Modeling human preferences using transformers for rl
C. Kim, J. Park, J. Shin, H. Lee, P. Abbeel, and K. Lee · 2023
Later among the works it cites.
Small batch deep reinforcement learning
J. Obando Ceron, M. Bellemare, and P. S. Castro · 2023
Later among the works it cites.
Growing q-networks: Solving continuous control tasks with adaptive control resolution
T. Seyde, P. Werner, W. Schwarting, M. Wulfmeier, and D. Rus · 2024
Closest in time.
Reverse forward curriculum learning for extreme sample and demonstration efficiency in reinforcement learning
S. Tao, A. Shukla, T.-k. Chan, and H. Su · 2024
Closest in time.
3d diffuser actor: Policy diffusion with 3d scene representations
T.-W. Ke, N. Gkanatsios, and K. Fragkiadaki · 2024
Closest in time.
Y. Ze, G. Zhang, K. Zhang, C. Hu, M. Wang, and H. Xu · 2024
Closest in time.
A recipe for unbounded data augmentation in visual reinforcement learning
A. Almuzairee, N. Hansen, and H. I. Christensen · 2024
Closest in time.
Rapid locomotion via reinforcement learning
G. B. Margolis, G. Yang, K. Paigwar, T. Chen, and P. Agrawal · 2024
Closest in time.
Serl: A software suite for sample-efficient robotic reinforcement learning
J. Luo, Z. Hu, C. Xu, Y. L. Tan, J. Berg, A. Sharma, S. Schaal, C. Finn, A. Gupta, and S. Levine · 2024
Closest in time.
One-2-3-45: Any single image to 3d mesh in 45 seconds without per-shape optimization
M. Liu, C. Xu, H. Jin, L. Chen, M. Varma T, Z. Xu, and H. Su · 2024
Closest in time.
Stop regressing: Training value functions via classification for scalable deep rl
J. Farebrother, J. Orbay, Q. Vuong, A. A. Taïga, Y. Chebotar, T. Xiao, A. Irpan, S. Levine, P. S. Castro, A. Faust, et al · 2024
Closest in time.
Dissecting deep rl with high update ratios: Combatting value overestimation and divergence
M. Hussing, C. Voelcker, I. Gilitschenski, A.-m. Farahmand, and E. Eaton · 2024
Closest in time.