Fetching the paper…
Reading the bibliography…
Reinforcement Learning has revolutionized decision-making processes in dynamic environments, yet it often struggles with autonomously detecting and achieving goals without clear feedback signals.
Equation of state calculations by fast computing machines
Metropolis, N.; Rosenbluth, A. W.; Rosenbluth, M. N.; Teller, A. H.; and Teller, E. 1953 · 1953
Earlier work this paper cites.
A new approach to linear filtering and prediction problems
Kalman, R. E. 1960 · 1960
Earlier work this paper cites.
Outline of a new approach to the analysis of complex systems and decision processes
Zadeh, L. A. 1973 · 1973
Earlier work this paper cites.
Feedback and optimal sensitivity: Model reference transformations, multiplicative seminorms, and approximate inverses
Zames, G. 1981 · 1981
Earlier work this paper cites.
Stochastic relaxation, Gibbs distributions, and the Bayesian restoration of images
Geman, S.; and Geman, D. 1984 · 1984
Earlier work this paper cites.
Bayesian netwcrks: A model cf self-activated memory for evidential reasoning
Pearl, J. 1985 · 1985
Earlier work this paper cites.
Novel approach to nonlinear/non-Gaussian Bayesian state estimation
Gordon, N. J.; Salmond, D. J.; and Smith, A. F. 1993 · 1993
Earlier work this paper cites.
New extension of the Kalman filter to nonlinear systems
Julier, S. J.; and Uhlmann, J. K. 1997 · 1997
Earlier work this paper cites.
Reinforcement learning with hierarchies of machines
Parr, R.; and Russell, S. 1997 · 1997
Earlier work this paper cites.
An introduction to variational methods for graphical models
Jordan, M. I.; Ghahramani, Z.; Jaakkola, T. S.; and Saul, L. K. 1999 · 1999
Earlier work this paper cites.
Sequential Monte Carlo methods in practice , volume 1
Doucet, A.; De Freitas, N.; Gordon, N. J.; et al. 2001 · 2001
Earlier work this paper cites.
The semantic web
Lassila, O.; Hendler, J.; and Berners-Lee, T. 2001 · 2001
Earlier work this paper cites.
Large eddy simulation of a turbulent reacting jet with conditional source-term estimation
Steiner, H.; and Bushe, W. 2001 · 2001
Earlier work this paper cites.
Beyond the Kalman Filter: Particle Filters for Tracking Applications
Ristic, B.; Arulampalam, S.; and Gordon, N. J. 2004 · 2004
Earlier work this paper cites.
Optimal filtering
Anderson, B. D.; and Moore, J. B. 2005 · 2005
Earlier work this paper cites.
‘Infotaxis’ as a strategy for searching without gradients
Vergassola, M.; Villermaux, E.; and Shraiman, B. I. 2007 · 2007
Earlier work this paper cites.
Freebase: a collaboratively created graph database for structuring human knowledge
Bollacker, K.; Evans, C.; Paritosh, P.; Sturge, T.; and Taylor, J. 2008 · 2008
Earlier work this paper cites.
Importance sampling: a review
Tokdar, S. T.; and Kass, R. E. 2010 · 2010
Earlier work this paper cites.
Odor source localization using a mobile robot in outdoor airflow environments with a particle filter algorithm
Li, J.-G.; Meng, Q.-H.; Wang, Y.; and Zeng, M. 2011 · 2011
Cited alongside, same era.
Long short-term memory
Graves, A.; and Graves, A. 2012 · 2012
Cited alongside, same era.
Gas source localization with a micro-drone using bio-inspired and particle filter-based algorithms
Neumann, P. P.; Hernandez Bennetts, V.; Lilienthal, A. J.; Bartholmai, M.; and Schiller, J. H. 2013 · 2013
Cited alongside, same era.
Probabilistic reasoning in intelligent systems: networks of plausible inference
Pearl, J. 2014 · 2014
Cited alongside, same era.
Kalman filter and its application
Li, Q.; Li, R.; Ji, K.; and Dai, W. 2015 · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A. A.; Veness, J.; Bellemare, M. G.; Graves, A.; Riedmiller, M.; Fidjeland, A. K.; Ostrovski, G.; et al. 2015 · 2015
Magnetic control of tokamak plasmas through deep reinforcement learning
Degrave, J.; Felici, F.; Buchli, J.; Neunert, M.; Tracey, B.; Carpanese, F.; Ewalds, T.; Hafner, R.; Abdolmaleki, A.; de Las Casas, D.; et al. 2022 · 2022
Later among the works it cites.
Meta-Reward-Net: Implicitly Differentiable Reward Learning for Preference-based Reinforcement Learning
Liu, R.; Bai, F.; Du, Y.; and Yang, Y. 2022 · 2022
Later among the works it cites.
Searching for a source without gradients: how good is infotaxis and how to beat it
Loisy, A.; and Eloy, C. 2022 · 2022
Later among the works it cites.
Autonomous source term estimation in unknown environments: From a dual control concept to UAV deployment
Rhodes, C.; Liu, C.; and Chen, W.-H. 2022 · 2022
Later among the works it cites.
Model-based multi-agent reinforcement learning: Recent progress and prospects
Wang, X.; Zhang, Z.; and Zhang, W. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Real-time wind estimation on a micro unmanned aerial vehicle using its inertial measurement unit
Neumann, P. P.; and Bartholmai, M. 2015 · 2015
Cited alongside, same era.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Kulkarni, T. D.; Narasimhan, K.; Saeedi, A.; and Tenenbaum, J. 2016 · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Levine, S.; Finn, C.; Darrell, T.; and Abbeel, P. 2016 · 2016
Cited alongside, same era.
Kalman filtering
Chui, C. K.; Chen, G.; et al. 2017 · 2017
Cited alongside, same era.
Motion planning for industrial robots using reinforcement learning
Meyes, R.; Tercan, H.; Roggendorf, S.; Thiele, T.; Büscher, C.; Obdenbusch, M.; Brecher, C.; Jeschke, S.; and Meisen, T. 2017 · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
Silver, D.; Schrittwieser, J.; Simonyan, K.; Antonoglou, I.; Huang, A.; Guez, A.; Hubert, T.; Baker, L.; Lai, M.; Bolton, A.; et al. 2017 · 2017
Cited alongside, same era.
Mingling foresight with imagination: Model-based cooperative multi-agent reinforcement learning
Xu, Z.; Zhang, B.; Zhan, Y.; Baiia, Y.; Fan, G.; et al. 2022 · 2022
Later among the works it cites.
A deep reinforcement learning based searching method for source localization
Zhao, Y.; Chen, B.; Wang, X.; Zhu, Z.; Wang, Y.; Cheng, G.; Wang, R.; Wang, R.; He, M.; and Liu, Y. 2022 · 2022
Later among the works it cites.
PiCor: Multi-Task Deep Reinforcement Learning with Policy Correction
Bai, F.; Zhang, H.; Tao, T.; Wu, Z.; Wang, Y.; and Xu, B. 2023 · 2023
Later among the works it cites.
Dense reinforcement learning for safety validation of autonomous vehicles
Feng, S.; Sun, H.; Yan, X.; Zhu, H.; Zou, Z.; Shen, S.; and Liu, H. X. 2023 · 2023
Later among the works it cites.
Passivity-based formation control for second-order multi-agent systems with linear or nonlinear coupling
Li, R.; Wang, J.-L.; and Shi, Y.-W. 2023 · 2023
Later among the works it cites.
Cooperative open-ended learning framework for zero-shot coordination
Li, Y.; Zhang, S.; Sun, J.; Du, Y.; Wen, Y.; Wang, X.; and Pan, W. 2023 · 2023
Later among the works it cites.
Order Matters: Agent-by-agent Policy Optimization
Wang, X.; Tian, Z.; Wan, Z.; Wen, Y.; Wang, J.; and Zhang, W. 2023 · 2023
Later among the works it cites.
Efficient Preference-based Reinforcement Learning via Aligned Experience Estimation
Bai, F.; Zhao, R.; Zhang, H.; Cui, S.; Wen, Y.; Yang, Y.; Xu, B.; and Han, L. 2024 · 2024
Closest in time.
Tackling cooperative incompatibility for zero-shot human-ai coordination
Li, Y.; Zhang, S.; Sun, J.; Zhang, W.; Du, Y.; Wen, Y.; Wang, X.; and Pan, W. 2024 · 2024
Closest in time.
Reinforcement Learning for Source Location Estimation: A Multi-Step Approach
Shi, Y.; McAreavey, K.; Liu, C.; and Liu, W. 2024 · 2024
Closest in time.
Zsc-eval: An evaluation toolkit and benchmark for multi-agent zero-shot coordination
Wang, X.; Zhang, S.; Zhang, W.; Dong, W.; Chen, J.; Wen, Y.; and Zhang, W. 2024 · 2024
Closest in time.
Xu, Z.; Mao, H.; Zhang, N.; Xin, X.; Ren, P.; Li, D.; Zhang, B.; Fan, G.; Chen, Z.; Wang, C.; et al. 2024 · 2024
Closest in time.