Fetching the paper…
Reading the bibliography…
Model-Free Reinforcement Learning (MFRL), leveraging the policy gradient theorem, has demonstrated considerable success in continuous control tasks.
Difftaichi: Differentiable programming for physical simulation
Hu, Y., Anderson, L., Li, T.-M., Sun, Q., Carr, N., Ragan-Kelley, J., and Durand, F · 1910
Earlier work this paper cites.
Dream to control: Learning behaviors by latent imagination
Hafner, D., Lillicrap, T., Ba, J., and Norouzi, M · 1912
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. and Tsitsiklis, J · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D., Singh, S., and Mansour, Y · 1999
Earlier work this paper cites.
Randomized smoothing for stochastic optimization
Duchi, J. C., Bartlett, P. L., and Wainwright, M. J · 2012
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Earlier work this paper cites.
Gradient estimation using stochastic computation graphs
Schulman, J., Heess, N., Weber, T., and Abbeel, P · 2015
Earlier work this paper cites.
Anymal-a highly mobile and dynamic quadrupedal robot
Hutter, M., Gehring, C., Jud, D., Lauber, A., Bellicoso, C. D., Tsounis, V., Hwangbo, J., Bodie, K., Fankhauser, P., Bloesch, M., et al · 2016
Earlier work this paper cites.
Control of a quadrotor with reinforcement learning
Hwangbo, J., Sa, I., Siegwart, R., and Hutter, M · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Model-based value expansion for efficient model-free reinforcement learning
Feinberg, V., Wan, A., Stoica, I., Jordan, M. I., Gonzalez, J. E., and Levine, S · 2018
Earlier work this paper cites.
Soft actor-critic algorithms and applications
Haarnoja, T., Zhou, A., Hartikainen, K., Tucker, G., Ha, S., Tan, J., Kumar, V., Zhu, H., Gupta, A., Abbeel, P., et al · 2018
Earlier work this paper cites.
Pipps: Flexible model-based policy search robust to the curse of chaos
Parmas, P., Rasmussen, C. E., Peters, J., and Doya, K · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Earlier work this paper cites.
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., Casas, D. d. L., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., et al · 2018
Cited alongside, same era.
Solving rubik’s cube with a robot hand
Akkaya, I., Andrychowicz, M., Chociej, M., Litwin, M., McGrew, B., Petron, A., Paino, A., Plappert, M., Powell, G., Ribas, R., et al · 2019
Cited alongside, same era.
Chainqueen: A real-time differentiable physical simulator for soft robotics
Hu, Y., Liu, J., Spielberg, A., Tenenbaum, J. B., Freeman, W. T., Wu, J., Rus, D., and Matusik, W · 2019
Cited alongside, same era.
Learning agile and dynamic motor skills for legged robots
Hwangbo, J., Lee, J., Dosovitskiy, A., Bellicoso, D., Tsounis, V., Koltun, V., and Hutter, M · 2019
Cited alongside, same era.
When to trust your model: Model-based policy optimization
Janner, M., Fu, J., Zhang, M., and Levine, S · 2019
Cited alongside, same era.
DiSECt: A Differentiable Simulation Engine for Autonomous Robotic Cutting
Heiden, E., Macklin, M., Narang, Y. S., Fox, D., Garg, A., and Ramos, F · 2021
Later among the works it cites.
Plasticinelab: A soft-body manipulation benchmark with differentiable physics
Huang, Z., Hu, Y., Du, T., Zhou, S., Su, H., Tenenbaum, J. B., and Gan, C · 2021
Later among the works it cites.
An end-to-end differentiable framework for contact-aware robot design
Xu, J., Chen, T., Zlokapa, L., Foshey, M., Matusik, W., Sueda, S., and Agrawal, P · 2021
Later among the works it cites.
A theoretical and empirical comparison of gradient approximations in derivative-free optimization
Berahas, A. S., Cao, L., Choromanski, K., and Scheinberg, K · 2022
Later among the works it cites.
Warp: A high-performance python framework for gpu simulation and graphics
Macklin, M · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning-based model predictive control for autonomous racing
Kabzan, J., Hewing, L., Liniger, A., and Zeilinger, M. N · 2019
Cited alongside, same era.
Differentiable cloth simulation for inverse problems
Liang, J., Lin, M., and Koltun, V · 2019
Cited alongside, same era.
Imagined value gradients: Model-based policy optimization with tranferable latent dynamics models
Byravan, A., Springenberg, J. T., Abdolmaleki, A., Hafner, R., Neunert, M., Lampe, T., Siegel, N., Heess, N., and Riedmiller, M · 2020
Cited alongside, same era.
Kaufmann, E., Loquercio, A., Ranftl, R., Müller, M., Koltun, V., and Scaramuzza, D · 2020
Cited alongside, same era.
Monte carlo gradient estimation in machine learning
Mohamed, S., Rosca, M., Figurnov, M., and Mnih, A · 2020
Cited alongside, same era.
Deep reinforcement learning at the edge of the statistical precipice
Agarwal, R., Schwarzer, M., Castro, P. S., Courville, A. C., and Bellemare, M · 2021
Cited alongside, same era.
On the model-based stochastic value gradient for continuous reinforcement learning
Amos, B., Stanton, S., Yarats, D., and Wilson, A. G · 2021
Cited alongside, same era.
Learning to walk in minutes using massively parallel deep reinforcement learning
Rudin, N., Hoeller, D., Reist, P., and Hutter, M · 2022
Later among the works it cites.
Do differentiable simulators give better policy gradients?
Suh, H. J., Simchowitz, M., Zhang, K., and Tedrake, R · 2022
Later among the works it cites.
Accelerated policy learning with parallel differentiable simulation
Xu, J., Makoviychuk, V., Narang, Y., Ramos, F., Matusik, W., Garg, A., and Macklin, M · 2022
Later among the works it cites.
Mastering diverse domains through world models
Hafner, D., Pasukonis, J., Ba, J., and Lillicrap, T · 2023
Later among the works it cites.
Td-mpc2: Scalable, robust world models for continuous control
Hansen, N., Su, H., and Wang, X · 2023
Later among the works it cites.
Differentiable dynamics simulation using invariant contact mapping and damped contact force
Lee, M., Lee, J., and Lee, D · 2023
Later among the works it cites.
Model-based reinforcement learning with scalable composite policy gradient estimators
Parmas, P., Seno, T., and Aoki, Y · 2023
Later among the works it cites.
Improving gradient computation for differentiable physics simulation with contacts
Zhong, Y. D., Han, J., Dey, B., and Brikis, G. O · 2023
Later among the works it cites.
Facing off world model backbones: Rnns, transformers, and s4
Deng, F., Park, J., and Ahn, S · 2024
Closest in time.