Fetching the paper…
Reading the bibliography…
We introduce CriticSMC, a new algorithm for planning as inference built from a composition of sequential Monte Carlo with learned Soft-Q function heuristic factors.
CAQL: Continuous Action Q-Learning
Moonkyung Ryu, Yinlam Chow, Ross Anderson, Christian Tjandraatmadja, and Craig Boutilier · 1909
Earlier work this paper cites.
Urban Driving with Conditional Imitation Learning
Jeffrey Hawke, Richard Shen, Corina Gurau, Siddharth Sharma, Daniele Reda, Nikolay Nikolov, Przemyslaw Mazur, Sean Micklethwaite, Nicolas Griffiths, Amar Shah, and Alex Kendall · 1912
Earlier work this paper cites.
Two-filter formulae for discrete-time non-linear bayesian smoothing
Yoram Bresler · 1986
Earlier work this paper cites.
Q-learning
Christopher J. C. H. Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Novel approach to nonlinear/non-gaussian bayesian state estimation
N.J. Gordon, D.J. Salmond, and A.F.M. Smith · 1993
Earlier work this paper cites.
The two-filter formula for smoothing and an implementation of the Gaussian-sum smoother
Genshiro Kitagawa · 1994
Earlier work this paper cites.
Monte carlo filter and smoother for non-gaussian nonlinear state space models
Genshiro Kitagawa · 1996
Earlier work this paper cites.
Sequential monte carlo methods for dynamic systems
Jun S. Liu and Rong Chen · 1998
Earlier work this paper cites.
The Unscented Particle Filter
Rudolph van der Merwe, Arnaud Doucet, Nando de Freitas, and Eric Wan · 2000
Earlier work this paper cites.
Following a moving target—Monte Carlo inference for dynamic Bayesian models
Walter R. Gilks and Carlo Berzuini · 2001
Earlier work this paper cites.
A tutorial on particle filters for online nonlinear/non-gaussian bayesian tracking
M. Sanjeev Arulampalam, Simon Maskell, Neil Gordon, and Tim Clapp · 2002
Earlier work this paper cites.
Particle methods for change detection, system identification, and control
Christophe Andrieu, A. Doucet, Sumeetpal S. Singh, and Vladislav Z. B. Tadić · 2003
Earlier work this paper cites.
On-line inference for hidden markov models via particle filters
Paul Fearnhead and Peter Clifford · 2003
Earlier work this paper cites.
Particle filters for mixture models with an unknown number of components
Paul Fearnhead · 2004
Earlier work this paper cites.
Monte carlo smoothing for nonlinear time series
Simon J. Godsill, Arnaud Doucet, and Mike West · 2004
Earlier work this paper cites.
Feynman-Kac Formulae: Genealogical and Interacting Particle Systems with Applications
Pierre Del Moral · 2004
Earlier work this paper cites.
Comparison of resampling schemes for particle filtering
R. Douc and O. Cappe · 2005
Earlier work this paper cites.
An overview of existing methods and recent advances in sequential monte carlo
Olivier Cappe, Simon J. Godsill, and Eric Moulines · 2007
Earlier work this paper cites.
Reinforcement learning in continuous action spaces through sequential monte carlo methods
Alessandro Lazaric, Marcello Restelli, and Andrea Bonarini · 2007
Earlier work this paper cites.
Particle Markov chain Monte Carlo methods
Christophe Andrieu, Arnaud Doucet, and Roman Holenstein · 2009
Earlier work this paper cites.
Forward smoothing using sequential Monte Carlo
P.D. Moral, A. Doucet, and S.S. Singh · 2009
Earlier work this paper cites.
Modeling interaction via the principle of maximum causal entropy
Brian D Ziebart, J Andrew Bagnell, and Anind K Dey · 2010
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Marc Deisenroth and Carl E Rasmussen · 2011
Earlier work this paper cites.
Sequential Monte Carlo smoothing for general state space hidden Markov models
Randal Douc, Aurélien Garivier, Eric Moulines, and Jimmy Olsson · 2011
Earlier work this paper cites.
Variational inference for policy search in changing situations
Gerhard Neumann · 2011
Earlier work this paper cites.
Smc2: an efficient algorithm for sequential analysis of state space models
Nicolas Chopin, Pierre E. Jacob, and Omiros Papaspiliopoulos · 2012
Cited alongside, same era.
Optimal control as a graphical model inference problem
Hilbert J Kappen, Vicenç Gómez, and Manfred Opper · 2012
Cited alongside, same era.
On some properties of markov chain monte carlo simulation methods based on the particle filter
Michael K. Pitt, Ralph dos Santos Silva, Paolo Giordani, and Robert Kohn · 2012
Cited alongside, same era.
On stochastic optimal control and reinforcement learning by approximate inference
Konrad Rawlik, Marc Toussaint, and Sethu Vijayakumar · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Cited alongside, same era.
Backward Simulation Methods for Monte Carlo Statistical Inference
Overcoming exploration in reinforcement learning with demonstrations
Ashvin Nair, Bob McGrew, Marcin Andrychowicz, Wojciech Zaremba, and Pieter Abbeel · 2018
Later among the works it cites.
Learning by playing solving sparse reward tasks from scratch
Martin Riedmiller, Roland Hafner, Thomas Lampe, Michael Neunert, Jonas Degrave, Tom van de Wiele, Vlad Mnih, Nicolas Heess, and Jost Tobias Springenberg · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Reinforcement learning and optimal control
Dimitri Bertsekas · 2019
Later among the works it cites.
Probabilistic planning with sequential monte carlo methods
Alexandre Piché, Valentin Thomas, Cyril Ibrahim, Yoshua Bengio, and Chris Pal · 2019
Later among the works it cites.
V-mpo: On-policy maximum a posteriori policy optimization for discrete and continuous control
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fredrik Lindsten and Thomans B. Schön · 2013
Cited alongside, same era.
Auto-encoding variational bayes
Diederik P. Kingma and Max Welling · 2014
Cited alongside, same era.
Sequential monte carlo for graphical models
Christian A. Naesseth, Fredrik Lindsten, and Thomas B. Schön · 2014
Cited alongside, same era.
Neural Adaptive Sequential Monte Carlo
Shixiang Gu, Zoubin Ghahramani, and Richard E. Turner · 2015
Cited alongside, same era.
Coarse-to-Fine Sequential Monte Carlo for Probabilistic Programs
Andreas Stuhlmüller, Robert X. D. Hawkins, N. Siddharth, and Noah D. Goodman · 2015
Cited alongside, same era.
End to End Learning for Self-Driving Cars
Mariusz Bojarski, Davide Del Testa, Daniel Dworakowski, Bernhard Firner, Beat Flepp, Prasoon Goyal, Lawrence D. Jackel, Mathew Monfort, Urs Muller, Jiakai Zhang, Xin Zhang, Jake Zhao, and Karol Zieba · 2016
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
H Francis Song, Abbas Abdolmaleki, Jost Tobias Springenberg, Aidan Clark, Hubert Soyer, Jack W Rae, Seb Noury, Arun Ahuja, Siqi Liu, Dhruva Tirumala, et al · 2019
Later among the works it cites.
Wei Zhan, Liting Sun, Di Wang, Haojie Shi, Aubrey Clausse, Maximilian Naumann, Julius Kümmerle, Hendrik Königshof, Christoph Stiller, Arnaud de La Fortelle, and Masayoshi Tomizuka · 2019
Later among the works it cites.
Variational inference for sequential data with future likelihood estimates
Geon-Hyeong Kim, Youngsoo Jang, Hongseok Yang, and Kee-Eung Kim · 2020
Later among the works it cites.
Implicit under-parameterization inhibits data-efficient deep reinforcement learning
Aviral Kumar, Rishabh Agarwal, Dibya Ghosh, and Sergey Levine · 2020
Later among the works it cites.
Deep dynamics models for learning dexterous manipulation
Anusha Nagabandi, Kurt Konolige, Sergey Levine, and Vikash Kumar · 2020
Later among the works it cites.
Learning to Locomote: Understanding How Environment Design Matters for Deep Reinforcement Learning
Daniele Reda, Tianxin Tao, and Michiel van de Panne · 2020
Later among the works it cites.
SimNet: Learning Reactive Self-driving Simulations from Real-world Observations
Luca Bergamini, Yawei Ye, Oliver Scheel, Long Chen, Chih Hu, Luca Del Pero, Blazej Osinski, Hugo Grimmett, and Peter Ondruska · 2021
Later among the works it cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, Erik Brynjolfsson, S. Buch, Dallas Card, Rodrigo Castellon, Niladri S. Chatterji, Annie S. Chen, Kathleen A. Creel, Jared Davis, Dora Demszky, Chris Donahue, Moussa Doumbouya, Esin Durmus, Stefano Ermon, John Etchemendy, Kawin Ethayarajh, Li Fei-Fei, Chelsea Finn, Trevor Gale, Lauren E. Gillespie, Karan Goel, Noah D. Goodman, Shelby Grossman, Neel Guha, Tatsunori Hashimoto, Peter Henderson, John Hewitt, Daniel E. Ho, Jenny Hong, Kyle Hsu, Jing Huang, Thomas F. Icard, Saahil Jain, Dan Jurafsky, Pratyusha Kalluri, Siddharth Karamcheti, Geoff Keeling, Fereshte Khani, O. Khattab, Pang Wei Koh, Mark S. Krass, Ranjay Krishna, Rohith Kuditipudi, Ananya Kumar, Faisal Ladhak, Mina Lee, Tony Lee, Jure Leskovec, Isabelle Levent, Xiang Lisa Li, Xuechen Li, Tengyu Ma, Ali Malik, Christopher D. Manning, Suvir P. Mirchandani, Eric Mitchell, Zanele Munyikwa, Suraj Nair, Avanika Narayan, Deepak Narayanan, Benjamin Newman, Allen Nie, Juan Carlos Niebles, Hamed Nilforoshan, J. F. Nyarko, Giray Ogut, Laurel Orr, Isabel Papadimitriou, Joon Sung Park, Chris Piech, Eva Portelance, Christopher Potts, Aditi Raghunathan, Robert Reich, Hongyu Ren, Frieda Rong, Yusuf H. Roohani, Camilo Ruiz, Jack Ryan, Christopher R’e, Dorsa Sadigh, Shiori Sagawa, Keshav Santhanam, Andy Shih, Krishna Parasuram Srinivasan, Alex Tamkin, Rohan Taori, Armin W. Thomas, Florian Tramèr, Rose E. Wang, William Wang, Bohan Wu, Jiajun Wu, Yuhuai Wu, Sang Michael Xie, Michihiro Yasunaga, Jiaxuan You, Matei A. Zaharia, Michael Zhang, Tianyi Zhang, Xikun Zhang, Yuhui Zhang, Lucia Zheng, Kaitlyn Zhou, and Percy Liang · 2021
Later among the works it cites.
Greedification operators for policy optimization: Investigating forward and reverse kl divergences
Alan Chan, Hugo Silva, Sungsu Lim, Tadashi Kozuno, A Rupam Mahmood, and Martha White · 2021
Later among the works it cites.
Challenges of real-world reinforcement learning: definitions, benchmarks and analysis
Gabriel Dulac-Arnold, Nir Levine, Daniel J Mankowitz, Jerry Li, Cosmin Paduraru, Sven Gowal, and Todd Hester · 2021
Later among the works it cites.
Reimagining an autonomous vehicle
Jeffrey Hawke, E Haibo, Vijay Badrinarayanan, and Alex Kendall · 2021
Later among the works it cites.
Autonomy 2.0: Why is self-driving always 5 years away?
Ashesh Jain, Luca Del Pero, Hugo Grimmett, and Peter Ondruska · 2021
Later among the works it cites.
A closer look at gradient estimators with reinforcement learning as inference
Jonathan Wilder Lavington, Michael Teng, Mark Schmidt, and Frank Wood · 2021
Later among the works it cites.
Sunrise: A simple unified framework for ensemble learning in deep reinforcement learning
Kimin Lee, Michael Laskin, Aravind Srinivas, and Pieter Abbeel · 2021
Later among the works it cites.
Imagining The Road Ahead: Multi-Agent Trajectory Prediction via Differentiable Simulation
Adam Ścibior, Vasileios Lioutas, Daniele Reda, Peyman Bateni, and Frank Wood · 2021
Later among the works it cites.
TrafficSim: Learning to Simulate Realistic Multi-Agent Behaviors
Simon Suo, Sebastian Regalado, Sergio Casas, and Raquel Urtasun · 2021
Later among the works it cites.
Symphony: Learning Realistic and Diverse Agents for Autonomous Driving Simulation
Maximilian Igl, Daewoo Kim, Alex Kuefler, Paul Mougin, Punit Shah, Kyriacos Shiarlis, Dragomir Anguelov, Mark Palatucci, Brandyn White, and Shimon Whiteson · 2022
Closest in time.
Unbiasedness of the Sequential Monte Carlo Based Normalizing Constant Estimator
Tuan Anh Le · 2022
Closest in time.
TITRATED: Learned human driving behavior without infractions via amortized inference
Vasileios Lioutas, Adam Scibior, and Frank Wood · 2022
Closest in time.