Fetching the paper…
Reading the bibliography…
Reinforcement learning often requires extensive training data.
1901
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Comput. , 1997. [Online]. Available: https://doi.org/10.1162/neco.1997.9.8.1735
1997
Earlier work this paper cites.
2007
Earlier work this paper cites.
2010
Earlier work this paper cites.
E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control,” in 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems , 2012, pp. 5026–5033
2012
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal policy optimization algorithms,” 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2018
Cited alongside, same era.
A. O. Onol, P. Long, and T. Padlr, “A comparative analysis of contact models in trajectory optimization for manipulation,” in 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , 2018, pp. 1–9
2018
Cited alongside, same era.
2019
Cited alongside, same era.
M. A. Z. Mora, M. Peychev, S. Ha, M. Vechev, and S. Coros, “Pods: Policy optimization via differentiable simulation,” in Proceedings of the 38th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, M. Meila and T. Zhang, Eds., vol. 139. PMLR, 18–24 Jul 2021, pp. 7805–7817. [Online]. Available: https://proceedings.mlr.press/v139/mora21a.html
H. J. Suh, M. Simchowitz, K. Zhang, and R. Tedrake, “Do differentiable simulators give better policy gradients?” in Proceedings of the 39th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, K. Chaudhuri, S. Jegelka, L. Song, C. Szepesvari, G. Niu, and S. Sabato, Eds., vol. 162. PMLR, 17–23 Jul 2022, pp. 20 668–20 696. [Online]. Available: https://proceedings.mlr.press/v162/suh22b.html
2022
Later among the works it cites.
J. Xu, V. Makoviychuk, Y. Narang, F. Ramos, W. Matusik, A. Garg, and M. Macklin, “Accelerated policy learning with parallel differentiable simulation,” 2022
2022
Later among the works it cites.
J. Du, H. Yan, J. Feng, J. T. Zhou, L. Zhen, R. S. M. Goh, and V. Tan, “Efficient sharpness-aware minimization for improved training of neural networks,” in International Conference on Learning Representations , 2022. [Online]. Available: https://openreview.net/forum?id=n0OeTdNRG0Q
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
F. Muratore, C. Eilers, M. Gienger et al. , “Data-efficient domain randomization with bayesian optimization,” IEEE Robotics and Automation Letters , 2021
2021
Cited alongside, same era.
J. Kwon, J. Kim, H. Park, and I. K. Choi, “Asam: Adaptive sharpness-aware minimization for scale-invariant learning of deep neural networks,” 2021
2021
Cited alongside, same era.
2022
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
2024
Closest in time.
2024
Closest in time.
J. Du, D. Zhou, J. Feng, V. Y. F. Tan, and J. T. Zhou, “Sharpness-aware training for free,” in Proceedings of the 36th International Conference on Neural Information Processing Systems , ser. NIPS ’22. Red Hook, NY, USA: Curran Associates Inc., 2024
2024
Closest in time.
D. Samuel, “sam: Sharpness-aware minimization for efficiently improving generalization,” https://github.com/davda54/sam , 2020, accessed: 2024-09-14
2024
Closest in time.