Fetching the paper…
Reading the bibliography…
When learning policies for robot control, the required real-world data is typically prohibitively expensive to acquire, so learning in simulation is a popular strategy.
B. F. Hobbs and A. Hepenstal, “Is optimization optimistically biased?” Water Resources Research , vol. 25, no. 2, pp. 152–160, 1989
1989
Earlier work this paper cites.
C. E. Rasmussen and C. K. I. Williams, Gaussian processes for machine learning , ser. Adaptive computation and machine learning. MIT Press, 2006
2006
Earlier work this paper cites.
M. Toussaint, “Robot trajectory optimization using approximate inference,” in ICML Montreal, Quebec, Canada, June 14-18 , vol. 382, 2009, pp. 1049–1056
2009
Earlier work this paper cites.
J. Peters, K. Mülling, and Y. Altun, “Relative entropy policy search,” in AAAI, Atlanta, Georgia, USA, July 11-15 , 2010
2010
Earlier work this paper cites.
J. Kober and J. Peters, “Policy search for motor primitives in robotics,” Machine Learning , vol. 84, no. 1-2, pp. 171–203, 2011
2011
Earlier work this paper cites.
J. Snoek, H. Larochelle, and R. P. Adams, “Practical bayesian optimization of machine learning algorithms,” in NIPS, Lake Tahoe, Nevada, United States, December 3-6 , 2012, pp. 2960–2968
2012
Earlier work this paper cites.
K. Rawlik, M. Toussaint, and S. Vijayakumar, “On stochastic optimal control and reinforcement learning by approximate inference,” in IJCAI, Beijing, China, August 3-9 , 2013, pp. 3052–3056
2013
Earlier work this paper cites.
I. Mordatch, K. Lowrey, and E. Todorov, “Ensemble-cio: Full-body dynamic motion planning that transfers to physical humanoids,” in IROS, Hamburg, Germany, September 28 - October 2 , 2015, pp. 5307–5314
2015
Earlier work this paper cites.
A. Cully, J. Clune, D. Tarapore, and J.-B. Mouret, “Robots that can adapt like animals,” Nature , vol. 521, no. 7553, pp. 503–507, 2015
2015
Earlier work this paper cites.
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal policy optimization algorithms,” ArXiv e-prints , 2017
2017
Earlier work this paper cites.
A. Rajeswaran, S. Ghotra, B. Ravindran, and S. Levine, “Epopt: Learning robust neural network policies using model ensembles,” in ICLR, Toulon, France, April 24-26 , 2017
2017
Cited alongside, same era.
J. Tobin et al
2017
Cited alongside, same era.
F. Sadeghi and S. Levine, “CAD2RL: real single-image flight without a single real image,” in RSS, Cambridge, Massachusetts, USA, July 12-16 , 2017
2017
Cited alongside, same era.
OpenAI et al
2018
Cited alongside, same era.
K. Lowrey, S. Kolev, J. Dao, A. Rajeswaran, and E. Todorov, “Reinforcement learning for non-prehensile manipulation: Transfer from simulation to physical system,” in SIMPAR 2018, Brisbane, Australia, May 16-19 , 2018, pp. 35–42
2018
Cited alongside, same era.
F. Muratore, M. Gienger, and J. Peters, “Assessing transferability from simulation to reality for reinforcement learning,” PAMI , vol. PP, pp. 1–1, 11 2019
2019
Later among the works it cites.
Y. Chebotar et al
2019
Later among the works it cites.
W. Yu, V. C. V. Kumar, G. Turk, and C. K. Liu, “Sim-to-real transfer for biped locomotion,” in IROS, Macau, SAR, China, November 3-8 . IEEE, 2019, pp. 3503–3510
2019
Later among the works it cites.
M. Balandat, B. Karrer, D. R. Jiang, S. Daulton, B. Letham, A. G. Wilson, and E. Bakshy, “BoTorch: Programmable bayesian optimization in pytorch,” ArXiv e-prints , 2019
2019
Later among the works it cites.
P. Klink, H. Abdulsamad, B. Belousov, and J. Peters, “Self-paced contextual reinforcement learning,” ArXiv e-prints , vol. 1910.02826, 2019
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
X. B. Peng, M. Andrychowicz, W. Zaremba, and P. Abbeel, “Sim-to-real transfer of robotic control with dynamics randomization,” in ICRA, Brisbane, Australia, May 21-25 , 2018, pp. 1–8
2018
Cited alongside, same era.
J. Tobin et al
2018
Cited alongside, same era.
N. Ruiz, S. Schulter, and M. Chandraker, “Learning to simulate,” ArXiv e-prints , vol. 1810.02513, 2018
2018
Cited alongside, same era.
S. Paul, M. A. Osborne, and S. Whiteson, “Fingerprint policy optimisation for robust reinforcement learning,” ArXiv e-prints , vol. 1805.10662, 2018
2018
Cited alongside, same era.
“mujoco-py,” Online. [Online]. Available: https://github.com/openai/mujoco-py
Cited in the paper.
B. Mehta, M. Diaz, F. Golemo, C. J. Pal, and L. Paull, “Active domain randomization,” in CoRL, Osaka, Japan, October 30 - November 1 , vol. 100. PMLR, 2019, pp. 1162–1176
2019
Later among the works it cites.
F. Ramos, R. Possas, and D. Fox, “Bayessim: Adaptive domain randomization via probabilistic inference for robotics simulators,” in RSS, University of Freiburg, Freiburg im Breisgau, Germany, June 22-26 , 2019
2019
Later among the works it cites.
F. Muratore, “SimuRLacra - a framework for reinforcement learning from randomized simulations,” https://github.com/famura/SimuRLacra , 2020
2020
Closest in time.
R. Moriconi, K. S. S. Kumar, and M. P. Deisenroth, “High-dimensional bayesian optimization with projections using quantile gaussian processes,” Optim. Lett. , vol. 14, no. 1, pp. 51–64, 2020
2020
Closest in time.