Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) is a promising approach for robotic navigation, allowing robots to learn through trial and error.
J. Canny, The complexity of robot motion planning . MIT press, 1988
1988
Earlier work this paper cites.
M. J. Mataric, “Reward functions for accelerated learning,” in Machine Learning Proceedings 1994 , San Francisco (CA), 1994, pp. 181–189
1994
Earlier work this paper cites.
K. Schilling, “Autonomous navigation of rovers for planetary exploration,” IFAC Proceedings Volumes , vol. 31, no. 21, pp. 83–87, 1998, 14th IFAC Symposium on Automatic Control in Aerospace 1998, Seoul, Korea, 24-28 August 1998
1998
Earlier work this paper cites.
R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour, “Policy gradient methods for reinforcement learning with function approximation,” Advances in neural information processing systems , vol. 12, 1999
1999
Earlier work this paper cites.
S. LaValle, “Planning algorithms,” Cambridge University Press google schola , vol. 2, pp. 3671–3678, 2006
2006
Earlier work this paper cites.
R. Montenegro, P. Tetali, et al. , “Mathematical aspects of mixing times in markov chains,” Foundations and Trends® in Theoretical Computer Science , vol. 1, no. 3, pp. 237–354, 2006
2006
Earlier work this paper cites.
M. Rickert, O. Brock, and A. Knoll, “Balancing exploration and exploitation in motion planning,” in 2008 IEEE International Conference on Robotics and Automation . IEEE, 2008, pp. 2812–2817
2008
Earlier work this paper cites.
R. Martinez-Cantin, N. De Freitas, E. Brochu, J. Castellanos, and A. Doucet, “A bayesian exploration-exploitation approach for optimal online sensing and planning with a visually guided mobile robot,” Autonomous Robots , vol. 27, pp. 93–103, 2009
2009
Earlier work this paper cites.
E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control.” in IROS . IEEE, 2012, pp. 5026–5033
2012
Earlier work this paper cites.
R. R. Murphy, Disaster robotics . MIT press, 2014
2014
Earlier work this paper cites.
B. Nakisa, M. N. Rastgoo, and M. J. Norodin, “Balancing exploration and exploitation in particle swarm optimization on search tasking,” Research Journal of Applied Science, Engineering and Technology , vol. 8, no. 12, pp. 1429–1434, 2014
2014
Earlier work this paper cites.
D. J. Hsu, A. Kontorovich, and C. Szepesvári, “Mixing time estimation in reversible markov chains from a single sample path,” CoRR , 2015
2015
Earlier work this paper cites.
R. Houthooft, X. Chen, Y. Duan, J. Schulman, F. De Turck, and P. Abbeel, “Vime: Variational information maximizing exploration,” Advances in neural information processing systems , vol. 29, 2016
2016
Earlier work this paper cites.
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell, “Curiosity-driven exploration by self-supervised prediction,” in International conference on machine learning . PMLR, 2017, pp. 2778–2787
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
D. A. Levin and Y. Peres, Markov chains and mixing times . American Mathematical Soc., 2017, vol. 107
2017
Earlier work this paper cites.
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction . MIT press, 2018
2018
Cited alongside, same era.
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,” in International conference on machine learning , 2018
2018
Cited alongside, same era.
A. Nair, B. McGrew, M. Andrychowicz, W. Zaremba, and P. Abbeel, “Overcoming exploration in reinforcement learning with demonstrations,” in 2018 IEEE international conference on robotics and automation (ICRA) . IEEE, 2018, pp. 6292–6299
2018
Cited alongside, same era.
G. Brunner, O. Richter, Y. Wang, and R. Wattenhofer, “Teaching a machine to read maps with deep reinforcement learning,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 32, 2018
2018
Cited alongside, same era.
K. Zhu and T. Zhang, “Deep reinforcement learning based mobile robot navigation: A review,” Tsinghua Science and Technology , vol. 26, no. 5, pp. 674–691, 2021
2021
Later among the works it cites.
D. Rengarajan, G. Vaidya, A. Sarvesh, D. Kalathil, and S. Shakkottai, “Reinforcement learning with sparse rewards using guidance from offline demonstration,” in International Conference on Learning Representations , 2021
2021
Later among the works it cites.
S. Ao, T. Zhou, G. Long, Q. Lu, L. Zhu, and J. Jiang, “Co-pilot: Collaborative planning and reinforcement learning on sub-task curriculum,” in Advances in Neural Information Processing Systems , M. Ranzato, A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan, Eds., vol. 34. Curran Associates, Inc., 2021, pp. 10 444–10 456
2021
Later among the works it cites.
Y. Yao, L. Xiao, Z. An, W. Zhang, and D. Luo, “Sample efficient reinforcement learning via model-ensemble exploration and exploitation,” in 2021 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2021, pp. 4202–4208
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
H.-T. L. Chiang, A. Faust, M. Fiser, and A. Francis, “Learning navigation behaviors end-to-end with autorl,” IEEE Robotics and Automation Letters , vol. 4, no. 2, pp. 2007–2014, 2019
2019
Cited alongside, same era.
P. Ghassemi and S. Chowdhury, “Decentralized informative path planning with balanced exploration-exploitation for swarm robotic search,” in International design engineering technical conferences and computers and information in engineering conference , vol. 59179. American Society of Mechanical Engineers, 2019, p. V001T02A058
2019
Cited alongside, same era.
J. Tarbouriech and A. Lazaric, “Active exploration in markov decision processes,” in The 22nd International Conference on Artificial Intelligence and Statistics . PMLR, 2019, pp. 974–982
2019
Cited alongside, same era.
C. Wang, J. Wang, J. Wang, and X. Zhang, “Deep-reinforcement-learning-based autonomous uav navigation with sparse rewards,” IEEE Internet of Things Journal , 02 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
M. Mutti and M. Restelli, “An intrinsically-motivated approach for learning highly exploring and fast mixing policies,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 34, no. 04, 2020
2020
Cited alongside, same era.
A. Staroverov, D. A. Yudin, I. Belkin, V. Adeshkin, Y. K. Solomentsev, and A. I. Panov, “Real-time object navigation with deep neural networks and hierarchical reinforcement learning,” IEEE Access , vol. 8, pp. 195 608–195 621, 2020
2020
Cited alongside, same era.
2021
Later among the works it cites.
K. M. B. Lee, F. Kong, R. Cannizzaro, J. L. Palmer, D. Johnson, C. Yoo, and R. Fitch, “An upper confidence bound for simultaneous exploration and exploitation in heterogeneous multi-robot systems,” in 2021 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2021, pp. 8685–8691
2021
Later among the works it cites.
X. Xiao, B. Liu, G. Warnell, and P. Stone, “Motion planning and control for mobile robot navigation using machine learning: a survey,” Autonomous Robots , pp. 1–29, 2022
2022
Later among the works it cites.
K. Weerakoon, A. J. Sathyamoorthy, U. Patel, and D. Manocha, “Terp: Reliable planning in uneven outdoor environments using deep reinforcement learning,” in 2022 International Conference on Robotics and Automation (ICRA) , 2022, pp. 9447–9453
2022
Later among the works it cites.
K. Weerakoon, S. Chakraborty, N. Karapetyan, A. J. Sathyamoorthy, A. Bedi, and D. Manocha, “HTRON: Efficient outdoor navigation with sparse rewards via heavy tailed adaptive reinforce algorithm,” in 6th Annual Conference on Robot Learning , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
R. Dorfman and K. Y. Levy, “Adapting to mixing time in stochastic optimization with Markovian data,” in Proceedings of the 39th International Conference on Machine Learning , 2022
2022
Later among the works it cites.
M. Riemer, S. C. Raparthy, I. Cases, G. Subbaraj, M. Puelma Touzel, and I. Rish, “Continual learning in environments with polynomial mixing times,” Advances in Neural Information Processing Systems , 2022
2022
Later among the works it cites.
B. Eysenbach and S. Levine, “Maximum entropy rl (provably) solves some robust rl problems,” in International Conference on Learning Representations , 2022
2022
Later among the works it cites.
S. Chakraborty, A. S. Bedi, K. Weerakoon, P. Poddar, A. Koppel, P. Tokekar, and D. Manocha, “Dealing with sparse rewards in continuous control robotics via heavy-tailed policy optimization,” in IEEE International Conference on Robotics and Automation (ICRA) , 2023
2023
Closest in time.
W. A. Suttle, A. Bedi, B. Patel, B. M. Sadler, A. Koppel, and D. Manocha, “Beyond exponentially fast mixing in average-reward reinforcement learning via multi-level monte carlo actor-critic,” in International Conference on Machine Learning , 2023
2023
Closest in time.
A. S. Bedi, A. Parayil, J. Zhang, M. Wang, and A. Koppel, “On the sample complexity and metastability of heavy-tailed policy search in continuous control,” Journal of Machine Learning Research , vol. 25, no. 39, pp. 1–58, 2024
2024
Closest in time.