Fetching the paper…
Reading the bibliography…
World models learn behaviors in a latent imagination space to enhance the sample-efficiency of deep reinforcement learning (RL) algorithms.
R. S. Sutton, “Dyna, an integrated architecture for learning, planning, and reacting,” SIGART Bull. , vol. 2, no. 4, p. 160–163, Jul. 1991. [Online]. Available: https://doi.org/10.1145/122344.122377
1991
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
S. Schaal, “Is imitation learning the route to humanoid robots?” Trends in cognitive sciences , vol. 3, no. 6, pp. 233–242, 1999
1999
Earlier work this paper cites.
A. Y. Ng, S. J. Russell et al. , “Algorithms for inverse reinforcement learning.” in Icml , vol. 1, 2000, p. 2
2000
Earlier work this paper cites.
F. Lamiraux and J. . Lammond, “Smooth motion planning for car-like vehicles,” IEEE Transactions on Robotics and Automation , vol. 17, no. 4, pp. 498–501, 2001
2001
Earlier work this paper cites.
E. Velenis and P. Tsiotras, “ Minimum Time vs Maximum Exit Velocity Path Optimization During Cornering ,” in Proceedings of the IEEE International Symposium on Industrial Electronics, 2005. ISIE 2005. , vol. 1, 2005, pp. 355–360
2005
Earlier work this paper cites.
M. Riedmiller, M. Montemerlo, and H. Dahlkamp, “Learning to drive a real car in 20 minutes,” in 2007 Frontiers in the Convergence of Bioscience and Information Technologies , 2007, pp. 645–650
2007
Earlier work this paper cites.
2008
Earlier work this paper cites.
F. Braghin, F. Cheli, S. Melzi, and E. Sabbioni, “Race driver model,” Computers & Structures , vol. 86, no. 13, pp. 1503–1516, 2008, structural Optimization. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0045794908000163
2008
Earlier work this paper cites.
M. P. Deisenroth and C. E. Rasmussen, “ PILCO: A Model-Based and Data-Efficient Approach to Policy Search ,” in Proceedings of the 28th International Conference on Machine Learning , ser. ICML’11. Madison, WI, USA: Omnipress, 2011, p. 465–472
2011
Earlier work this paper cites.
S. Ross, G. Gordon, and D. Bagnell, “A reduction of imitation learning and structured prediction to no-regret online learning,” in Proceedings of the fourteenth international conference on artificial intelligence and statistics . JMLR Workshop and Conference Proceedings, 2011, pp. 627–635
2011
Earlier work this paper cites.
D. Liberzon, Calculus of Variations and Optimal Control Theory: A Concise Introduction . USA: Princeton University Press, 2011
2011
Earlier work this paper cites.
V. Sezer and M. Gokasan, “A novel obstacle avoidance algorithm: “follow the gap method”,” Robotics and Autonomous Systems , vol. 60, no. 9, pp. 1123–1134, 2012. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0921889012000838
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
J. P. Timings and D. J. Cole, “ Minimum maneuver time calculation using convex optimization ,” Journal of Dynamic Systems, Measurement, and Control , vol. 135, no. 3, 2013
2013
Earlier work this paper cites.
A. Rucco, G. Notarstefano, and J. Hauser, “ An Efficient Minimum-Time Trajectory Generation Strategy for Two-Track Car Vehicles ,” IEEE Trans on Control Systems Tech , vol. 23, no. 4, pp. 1505–1519, 2015
2015
Earlier work this paper cites.
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal policy optimization algorithms,” 2017
2017
Earlier work this paper cites.
G. Williams, N. Wagener, B. Goldfain, P. Drews, J. M. Rehg, B. Boots, and E. A. Theodorou, “Information theoretic mpc for model-based reinforcement learning,” in 2017 IEEE International Conference on Robotics and Automation (ICRA) , 2017, pp. 1714–1721
2017
Earlier work this paper cites.
R. S. Sutton and A. G. Barto, RL: An introduction . MIT press, 2018
2018
Earlier work this paper cites.
D. Ha and J. Schmidhuber, “Recurrent world models facilitate policy evolution,” in Advances in Neural Information Processing Systems , S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, Eds., vol. 31. Curran Associates, Inc., 2018. [Online]. Available: https://proceedings.neurips.cc/paper/2018/file/2de5d16682c3c35007e4e92982f1a2ba-Paper.pdf
2018
Earlier work this paper cites.
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,” in Int. Conf. on Machine Learning , 2018, pp. 1861–1870
2018
Earlier work this paper cites.
G. Barth-Maron, M. W. Hoffman, D. Budden, W. Dabney, D. Horgan, D. TB, A. Muldal, N. Heess, and T. Lillicrap, “ Distributed Distributional Deterministic Policy Gradients ,” in International Conference on Learning Representations , 2018
2018
Cited alongside, same era.
A. Abdolmaleki, J. T. Springenberg, Y. Tassa, R. Munos, N. Heess, and M. Riedmiller, “Maximum a posteriori policy optimisation,” in International Conference on Learning Representations , 2018. [Online]. Available: https://openreview.net/forum?id=S1ANxQW0b
2018
Cited alongside, same era.
Y. Yu, “Towards sample efficient reinforcement learning,” in Proceedings of the 27th International Joint Conference on Artificial Intelligence , ser. IJCAI’18. AAAI Press, 2018, p. 5739–5743
2018
Cited alongside, same era.
X. B. Peng, M. Andrychowicz, W. Zaremba, and P. Abbeel, “Sim-to-real transfer of robotic control with dynamics randomization,” 2018 IEEE International Conference on Robotics and Automation (ICRA) , May 2018. [Online]. Available: http://dx.doi.org/10.1109/ICRA.2018.8460528
H. Zhu, J. Yu, A. Gupta, D. Shah, K. Hartikainen, A. Singh, V. Kumar, and S. Levine, “The ingredients of real world robotic reinforcement learning,” in International Conference on Learning Representations , 2020. [Online]. Available: https://openreview.net/forum?id=rJe2syrtvS
2020
Later among the works it cites.
W. Zhao, J. P. Queralta, and T. Westerlund, “ Sim-to-Real Transfer in Deep RL for Robotics: a Survey ,” in IEEE Symposium Series on Computational Intelligence (SSCI) , 2020, pp. 737–744
2020
Later among the works it cites.
M. Kaspar, J. D. Muñoz Osorio, and J. Bock, “ Sim2Real Transfer for Reinforcement Learning without Dynamics Randomization ,” in 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , 2020, pp. 4383–4388
2020
Later among the works it cites.
L. Andresen, A. Brandemuehl, A. Honger, B. Kuan, N. Vödisch, H. Blum, V. Reijgwart, L. Bernreiter, L. Schaupp, J. J. Chung, M. Burki, M. R. Oswald, R. Siegwart, and A. Gawel, “ Accurate Mapping and Planning for Autonomous Racing ,” in 2020 IEEE/RSJ Int. Conference on Intelligent Robots and Systems (IROS) , 2020, pp. 4743–4749
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
J. Tan, T. Zhang, E. Coumans, A. Iscen, Y. Bai, D. Hafner, S. Bohez, and V. Vanhoucke, “ Sim-to-Real: Learning Agile Locomotion For Quadruped Robots ,” in Proceedings of Robotics: Science and Systems , Pittsburgh, Pennsylvania, June 2018
2018
Cited alongside, same era.
2019
Cited alongside, same era.
D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson, “Learning latent dynamics for planning from pixels,” in Int. Conf. on Machine Learning . PMLR, 2019, pp. 2555–2565
2019
Cited alongside, same era.
T. Haarnoja, S. Ha, A. Zhou, J. Tan, G. Tucker, and S. Levine, “Learning to walk via deep reinforcement learning,” in Proceedings of Robotics: Science and Systems , FreiburgimBreisgau, Germany, June 2019
2019
Cited alongside, same era.
A. Singh, L. Yang, C. Finn, and S. Levine, “ End-To-End Robotic Reinforcement Learning without Reward Engineering ,” in Proceedings of Robotics: Science and Systems , 2019
2019
Cited alongside, same era.
J. Kabzan, M. de la Iglesia Valls, V. Reijgwart, H. F. C. Hendrikx, C. Ehmke, M. Prajapat, A. Bühler, N. Gosala, M. Gupta, R. Sivanesan, A. Dhall, E. Chisari, N. Karnchanachari, S. Brits, M. Dangel, I. Sa, R. Dubé, A. Gawel, M. Pfeiffer, A. Liniger, J. Lygeros, and R. Siegwart, “Amz driverless: The full autonomous racing system,” 2019
2019
Cited alongside, same era.
J. Kabzan, L. Hewing, A. Liniger, and M. N. Zeilinger, “Learning-based model predictive control for autonomous racing,” IEEE Robotics and Automation Letters , vol. 4, no. 4, pp. 3363 – 3370, 2019-10
2019
Cited alongside, same era.
A. Kendall, J. Hawke, D. Janz, P. Mazur, D. Reda, J. Allen, V. Lam, A. Bewley, and A. Shah, “Learning to drive in a day,” in 2019 International Conference on Robotics and Automation (ICRA) , 2019, pp. 8248–8254
2019
Cited alongside, same era.
2020
Later among the works it cites.
J. L. Vázquez, M. Brühlmeier, A. Liniger, A. Rupenyan, and J. Lygeros, “Optimization-based hierarchical motion planning for autonomous racing,” in 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , 2020, pp. 2397–2403
2020
Later among the works it cites.
G. Bellegarda and K. Byl, “ An Online Training Method for Augmenting MPC with Deep Reinforcement Learning ,” in 2020 IEEE/RSJ Int. Conf. on Intelligent Robots and Systems (IROS) , 2020, pp. 5453–5459
2020
Later among the works it cites.
M. Lechner, R. Hasani, D. Rus, and R. Grosu, “Gershgorin loss stabilizes the recurrent neural network compartment of an end-to-end robot learning scheme,” in 2020 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2020, pp. 5446–5452
2020
Later among the works it cites.
M. Lechner, R. Hasani, A. Amini, T. A. Henzinger, D. Rus, and R. Grosu, “Neural circuit policies enabling auditable autonomy,” Nature Machine Intelligence , vol. 2, no. 10, pp. 642–652, 2020
2020
Later among the works it cites.
A. Amini, I. Gilitschenski, J. Phillips, J. Moseyko, R. Banerjee, S. Karaman, and D. Rus, “Learning robust control policies for end-to-end autonomous driving from data-driven simulation,” IEEE Robotics and Automation Letters , vol. 5, no. 2, pp. 1143–1150, 2020
2020
Later among the works it cites.
V. S. Babu and M. Behl, “f1tenth. dev-an open-source ros based f1/10 autonomous racing simulator,” in 16th Int. Conf. on Automation Science and Engineering (CASE) . IEEE, 2020, pp. 1614–1620
2020
Later among the works it cites.
M. O’Kelly, H. Zheng, A. Jain, J. Auckley, K. Luong, and R. Mangharam, “Tunercar: A superoptimization toolchain for autonomous racing,” in 2020 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2020, pp. 5356–5362
2020
Later among the works it cites.
2020
Later among the works it cites.
2021
Closest in time.
C. Vorbach, R. Hasani, A. Amini, M. Lechner, and D. Rus, “Causal navigation by continuous-time neural networks,” Advances in Neural Information Processing Systems , vol. 34, 2021
2021
Closest in time.
M. Lechner, R. Hasani, R. Grosu, D. Rus, and T. A. Henzinger, “Adversarial training is not ready for robot learning,” in 2021 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2021, pp. 4140–4147
2021
Closest in time.
R. Hasani, M. Lechner, A. Amini, D. Rus, and R. Grosu, “Liquid time-constant networks,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 35, no. 9, 2021, pp. 7657–7666
2021
Closest in time.
2021
Closest in time.
S. Mysore, B. Mabsout, R. Mancuso, and K. Saenko, “Regularizing action policies for smooth control with reinforcement learning,” 2021
2021
Closest in time.
T. Seyde, I. Gilitschenski, W. Schwarting, B. Stellato, M. Riedmiller, M. Wulfmeier, and D. Rus, “Is bang-bang control all you need? solving continuous control with bernoulli policies,” Advances in Neural Information Processing Systems , vol. 34, 2021
2021
Closest in time.
M. Jaritz, R. de Charette, M. Toromanoff, E. Perot, and F. Nashashibi, “End-to-end race driving with deep reinforcement learning,” in 2018 IEEE International Conference on Robotics and Automation (ICRA) , 2018, pp. 2070–2075
2075
Closest in time.