Fetching the paper…
Reading the bibliography…
Consider the problem of exploration in sparse-reward or reward-free environments, such as in Montezuma's Revenge.
A possibility for implementing curiosity and boredom in model-building neural controllers
Jürgen Schmidhuber · 1991
Earlier work this paper cites.
Curious model-building control systems
Jürgen Schmidhuber · 1991
Earlier work this paper cites.
A possibility for implementing curiosity and boredom in model-building neural controllers
Jürgen Schmidhuber · 1991
Earlier work this paper cites.
Curious model-building control systems
Jürgen Schmidhuber · 1991
Earlier work this paper cites.
Exploration in active learning
Sebastian Thrun · 1995
Earlier work this paper cites.
Exploration in active learning
Sebastian Thrun · 1995
Earlier work this paper cites.
Active learning with statistical models
David A Cohn, Zoubin Ghahramani, and Michael I Jordan · 1996
Earlier work this paper cites.
Active learning with statistical models
David A Cohn, Zoubin Ghahramani, and Michael I Jordan · 1996
Earlier work this paper cites.
Intrinsically motivated learning of hierarchical collections of skills
Andrew G Barto, Satinder Singh, Nuttapong Chentanez, et al · 2004
Earlier work this paper cites.
The im algorithm: a variational approach to information maximization
David Barber Felix Agakov · 2004
Earlier work this paper cites.
Intrinsically motivated learning of hierarchical collections of skills
Andrew G Barto, Satinder Singh, Nuttapong Chentanez, et al · 2004
Earlier work this paper cites.
The im algorithm: a variational approach to information maximization
David Barber Felix Agakov · 2004
Earlier work this paper cites.
Intrinsic motivation systems for autonomous mental development
Pierre-Yves Oudeyer, Frdric Kaplan, and Verena V Hafner · 2007
Earlier work this paper cites.
Intrinsic motivation systems for autonomous mental development
Pierre-Yves Oudeyer, Frdric Kaplan, and Verena V Hafner · 2007
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Alexander L Strehl and Michael L Littman · 2008
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Alexander L Strehl and Michael L Littman · 2008
Earlier work this paper cites.
Bayesian surprise attracts human attention
Laurent Itti and Pierre Baldi · 2009
Earlier work this paper cites.
Causality
Judea Pearl · 2009
Earlier work this paper cites.
Bayesian surprise attracts human attention
Laurent Itti and Pierre Baldi · 2009
Earlier work this paper cites.
Causality
Judea Pearl · 2009
Earlier work this paper cites.
A pomdp extension with belief-dependent rewards
Mauricio Araya, Olivier Buffet, Vincent Thomas, and Françcois Charpillet · 2010
Earlier work this paper cites.
A pomdp extension with belief-dependent rewards
Mauricio Araya, Olivier Buffet, Vincent Thomas, and Françcois Charpillet · 2010
Earlier work this paper cites.
Planning to be surprised: Optimal bayesian exploration in dynamic environments
Yi Sun, Faustino Gomez, and Jürgen Schmidhuber · 2011
Earlier work this paper cites.
Planning to be surprised: Optimal bayesian exploration in dynamic environments
Yi Sun, Faustino Gomez, and Jürgen Schmidhuber · 2011
Earlier work this paper cites.
An information-theoretic approach to curiosity-driven reinforcement learning
Susanne Still and Doina Precup · 2012
Earlier work this paper cites.
An information-theoretic approach to curiosity-driven reinforcement learning
Susanne Still and Doina Precup · 2012
Earlier work this paper cites.
Universal knowledge-seeking agents for stochastic environments
Laurent Orseau, Tor Lattimore, and Marcus Hutter · 2013
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Universal knowledge-seeking agents for stochastic environments
Laurent Orseau, Tor Lattimore, and Marcus Hutter · 2013
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
On the properties of neural machine translation: Encoder-decoder approaches
Kyunghyun Cho, Bart Van Merriënboer, Dzmitry Bahdanau, and Yoshua Bengio · 2014
Earlier work this paper cites.
On the properties of neural machine translation: Encoder-decoder approaches
Kyunghyun Cho, Bart Van Merriënboer, Dzmitry Bahdanau, and Yoshua Bengio · 2014
Earlier work this paper cites.
Action-conditional video prediction using deep networks in atari games
Junhyuk Oh, Xiaoxiao Guo, Honglak Lee, Richard L Lewis, and Satinder Singh · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Action-conditional video prediction using deep networks in atari games
Junhyuk Oh, Xiaoxiao Guo, Honglak Lee, Richard L Lewis, and Satinder Singh · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Unsupervised learning for physical interaction through video prediction
Chelsea Finn, Ian Goodfellow, and Sergey Levine · 2016
Earlier work this paper cites.
Incentivizing exploration in reinforcement learning with deep predictive models
Bradly C Stadie, Sergey Levine, and Pieter Abbeel · 2016
Earlier work this paper cites.
Deep exploration via bootstrapped dqn
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy · 2016
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Earlier work this paper cites.
Vime: Variational information maximizing exploration
Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2016
Earlier work this paper cites.
The variational fair autoencoder
Christos Louizos, Kevin Swersky, Yujia Li, Max Welling, and Richard Zemel · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
Hado Van Hasselt, Arthur Guez, and David Silver · 2016
Earlier work this paper cites.
Unsupervised learning for physical interaction through video prediction
Chelsea Finn, Ian Goodfellow, and Sergey Levine · 2016
Earlier work this paper cites.
Incentivizing exploration in reinforcement learning with deep predictive models
Bradly C Stadie, Sergey Levine, and Pieter Abbeel · 2016
Earlier work this paper cites.
Deep exploration via bootstrapped dqn
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy · 2016
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Earlier work this paper cites.
Vime: Variational information maximizing exploration
Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2016
Earlier work this paper cites.
The variational fair autoencoder
Christos Louizos, Kevin Swersky, Yujia Li, Max Welling, and Richard Zemel · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
Hado Van Hasselt, Arthur Guez, and David Silver · 2016
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell · 2017
Earlier work this paper cites.
# exploration: A study of count-based exploration for deep reinforcement learning
Haoran Tang, Rein Houthooft, Davis Foote, Adam Stooke, OpenAI Xi Chen, Yan Duan, John Schulman, Filip DeTurck, and Pieter Abbeel · 2017
Earlier work this paper cites.
Count-based exploration with neural density models
Georg Ostrovski, Marc G Bellemare, Aäron Oord, and Rémi Munos · 2017
Earlier work this paper cites.
Ex2: Exploration with exemplar models for deep reinforcement learning
Justin Fu, John Co-Reyes, and Sergey Levine · 2017
Earlier work this paper cites.
Variational intrinsic control
Karol Gregor, Danilo Jimenez Rezende, and Daan Wierstra · 2017
Earlier work this paper cites.
Hindsight experience replay
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, OpenAI Pieter Abbeel, and Wojciech Zaremba · 2017
Earlier work this paper cites.
The pycolab game engine, 2017
Thomas Stepleton · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell · 2017
Earlier work this paper cites.
# exploration: A study of count-based exploration for deep reinforcement learning
Haoran Tang, Rein Houthooft, Davis Foote, Adam Stooke, OpenAI Xi Chen, Yan Duan, John Schulman, Filip DeTurck, and Pieter Abbeel · 2017
Earlier work this paper cites.
Count-based exploration with neural density models
Georg Ostrovski, Marc G Bellemare, Aäron Oord, and Rémi Munos · 2017
Earlier work this paper cites.
Ex2: Exploration with exemplar models for deep reinforcement learning
Justin Fu, John Co-Reyes, and Sergey Levine · 2017
Earlier work this paper cites.
Variational intrinsic control
Karol Gregor, Danilo Jimenez Rezende, and Daan Wierstra · 2017
Earlier work this paper cites.
Hindsight experience replay
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, OpenAI Pieter Abbeel, and Wojciech Zaremba · 2017
Earlier work this paper cites.
The pycolab game engine, 2017
Thomas Stepleton · 2017
Earlier work this paper cites.
Cognitive computational neuroscience
Nikolaus Kriegeskorte and Pamela K Douglas · 2018
Earlier work this paper cites.
Dora the explorer: Directed outreaching reinforcement action-selection
Leshem Choshen, Lior Fox, and Yonatan Loewenstein · 2018
Earlier work this paper cites.
Randomized prior functions for deep reinforcement learning
Ian Osband, John Aslanides, and Albin Cassirer · 2018
Earlier work this paper cites.
Curiosity-driven experience prioritization via density estimation
Rui Zhao and Volker Tresp · 2018
Earlier work this paper cites.
Count-based exploration with the successor representation
Marlos C Machado, Marc G Bellemare, and Michael Bowling · 2018
Earlier work this paper cites.
Directed exploration in pac model-free reinforcement learning
Min-hwan Oh and Garud Iyengar · 2018
Earlier work this paper cites.
Variational option discovery algorithms
Joshua Achiam, Harrison Edwards, Dario Amodei, and Pieter Abbeel · 2018
Cited alongside, same era.
Automatic goal generation for reinforcement learning agents
Carlos Florensa, David Held, Xinyang Geng, and Pieter Abbeel · 2018
Cited alongside, same era.
Visual reinforcement learning with imagined goals
Ashvin V Nair, Vitchyr Pong, Murtaza Dalal, Shikhar Bahl, Steven Lin, and Sergey Levine · 2018
Cited alongside, same era.
Woulda, coulda, shoulda: Counterfactually-guided policy search
Lars Buesing, Theophane Weber, Yori Zwols, Sebastien Racaniere, Arthur Guez, Jean-Baptiste Lespiau, and Nicolas Heess · 2018
Cited alongside, same era.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
Marlos C Machado, Marc G Bellemare, Erik Talvitie, Joel Veness, Matthew Hausknecht, and Michael Bowling · 2018
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Later among the works it cites.
Learning invariant representations for reinforcement learning without reconstruction
Amy Zhang, Rowan McAllister, Roberto Calandra, Yarin Gal, and Sergey Levine · 2020
Later among the works it cites.
Adversarial active exploration for inverse dynamics model learning
Zhang-Wei Hong, Tsu-Jui Fu, Tzu-Yun Shann, and Chun-Yi Lee · 2020
Later among the works it cites.
Active world model learning with progress curiosity
Kuno Kim, Megumi Sano, Julian De Freitas, Nick Haber, and Daniel Yamins · 2020
Later among the works it cites.
Never give up: Learning directed exploration strategies
Adrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, Daniel Guo, Bilal Piot, Steven Kapturowski, Olivier Tieleman, Martín Arjovsky, Alexander Pritzel, Andew Bolt, et al · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Cited alongside, same era.
Invariant representations without adversarial training
Daniel Moyer, Shuyang Gao, Rob Brekelmans, Aram Galstyan, and Greg Ver Steeg · 2018
Cited alongside, same era.
Group normalization
Yuxin Wu and Kaiming He · 2018
Cited alongside, same era.
Cognitive computational neuroscience
Nikolaus Kriegeskorte and Pamela K Douglas · 2018
Cited alongside, same era.
Dora the explorer: Directed outreaching reinforcement action-selection
Leshem Choshen, Lior Fox, and Yonatan Loewenstein · 2018
Cited alongside, same era.
Randomized prior functions for deep reinforcement learning
Ian Osband, John Aslanides, and Albin Cassirer · 2018
Cited alongside, same era.
Curiosity-driven experience prioritization via density estimation
Rui Zhao and Volker Tresp · 2018
Cited alongside, same era.
Agent57: Outperforming the atari human benchmark
Adrià Puigdomènech Badia, Bilal Piot, Steven Kapturowski, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Daniel Guo, and Charles Blundell · 2020
Later among the works it cites.
Count-based exploration with the successor representation
Marlos C Machado, Marc G Bellemare, and Michael Bowling · 2020
Later among the works it cites.
Planning to explore via self-supervised world models
Ramanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel, Danijar Hafner, and Deepak Pathak · 2020
Later among the works it cites.
Explore, discover and learn: Unsupervised discovery of state-covering skills
Víctor Campos, Alexander Trott, Caiming Xiong, Richard Socher, Xavier Giró-i Nieto, and Jordi Torres · 2020
Later among the works it cites.
Dynamics-aware unsupervised discovery of skills
Archit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar, and Karol Hausman · 2020
Later among the works it cites.
Automatic curriculum learning through value disagreement
Yunzhi Zhang, Pieter Abbeel, and Lerrel Pinto · 2020
Later among the works it cites.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2020
Later among the works it cites.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Later among the works it cites.
Learning invariant representations for reinforcement learning without reconstruction
Amy Zhang, Rowan McAllister, Roberto Calandra, Yarin Gal, and Sergey Levine · 2020
Later among the works it cites.
Exploration in deep reinforcement learning: a comprehensive survey
Tianpei Yang, Hongyao Tang, Chenjia Bai, Jinyi Liu, Jianye Hao, Zhaopeng Meng, and Peng Liu · 2021
Later among the works it cites.
Density-based bonuses on learned representations for reward-free exploration in deep reinforcement learning
Omar Darwiche Domingues, Corentin Tallec, Remi Munos, and Michal Valko · 2021
Later among the works it cites.
Adversarially guided actor-critic
Yannis Flet-Berliac, Johan Ferret, Olivier Pietquin, Philippe Preux, and Matthieu Geist · 2021
Later among the works it cites.
Discovering and achieving goals via world models
Russell Mendonca, Oleh Rybkin, Kostas Daniilidis, Danijar Hafner, and Deepak Pathak · 2021
Later among the works it cites.
Behavior from the void: Unsupervised active pre-training
Hao Liu and Pieter Abbeel · 2021
Later among the works it cites.
Geometric entropic exploration
Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, Alaa Saade, Shantanu Thakoor, Bilal Piot, Bernardo Avila Pires, Michal Valko, Thomas Mesnard, Tor Lattimore, and Rémi Munos · 2021
Later among the works it cites.
Reinforcement learning with prototypical representations
Denis Yarats, Rob Fergus, Alessandro Lazaric, and Lerrel Pinto · 2021
Later among the works it cites.
Relative variational intrinsic control
Kate Baumli, David Warde-Farley, Steven Hansen, and Volodymyr Mnih · 2021
Later among the works it cites.
Is curiosity all you need? on the utility of emergent behaviours from curious exploration
Oliver Groth, Markus Wulfmeier, Giulia Vezzani, Vibhavari Dasagi, Tim Hertweck, Roland Hafner, Nicolas Heess, and Martin Riedmiller · 2021
Later among the works it cites.
Aps: Active pretraining with successor features
Hao Liu and Pieter Abbeel · 2021
Later among the works it cites.
Learning generalized gumbel-max causal mechanisms
Guy Lorberbom, Daniel D Johnson, Chris J Maddison, Daniel Tarlow, and Tamir Hazan · 2021
Later among the works it cites.
Understanding the behaviour of contrastive loss
Feng Wang and Huaping Liu · 2021
Later among the works it cites.
Posterior value functions: Hindsight baselines for policy gradient methods
Chris Nota, Philip Thomas, and Bruno C Da Silva · 2021
Later among the works it cites.
Counterfactual credit assignment in model-free reinforcement learning
Thomas Mesnard, Théophane Weber, Fabio Viola, Shantanu Thakoor, Alaa Saade, Anna Harutyunyan, Will Dabney, Tom Stepleton, Nicolas Heess, Arthur Guez, et al · 2021
Later among the works it cites.
Invariant causal imitation learning for generalizable policies
Ioana Bica, Daniel Jarrett, and Mihaela van der Schaar · 2021
Later among the works it cites.
Time-series generation by contrastive imitation
Daniel Jarrett, Ioana Bica, and Mihaela van der Schaar · 2021
Later among the works it cites.
Benchmarking the spectrum of agent capabilities
Danijar Hafner · 2021
Later among the works it cites.
Exploration in deep reinforcement learning: a comprehensive survey
Tianpei Yang, Hongyao Tang, Chenjia Bai, Jinyi Liu, Jianye Hao, Zhaopeng Meng, and Peng Liu · 2021
Later among the works it cites.
Density-based bonuses on learned representations for reward-free exploration in deep reinforcement learning
Omar Darwiche Domingues, Corentin Tallec, Remi Munos, and Michal Valko · 2021
Later among the works it cites.
Adversarially guided actor-critic
Yannis Flet-Berliac, Johan Ferret, Olivier Pietquin, Philippe Preux, and Matthieu Geist · 2021
Later among the works it cites.
Discovering and achieving goals via world models
Russell Mendonca, Oleh Rybkin, Kostas Daniilidis, Danijar Hafner, and Deepak Pathak · 2021
Later among the works it cites.
Behavior from the void: Unsupervised active pre-training
Hao Liu and Pieter Abbeel · 2021
Later among the works it cites.
Geometric entropic exploration
Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, Alaa Saade, Shantanu Thakoor, Bilal Piot, Bernardo Avila Pires, Michal Valko, Thomas Mesnard, Tor Lattimore, and Rémi Munos · 2021
Later among the works it cites.
Reinforcement learning with prototypical representations
Denis Yarats, Rob Fergus, Alessandro Lazaric, and Lerrel Pinto · 2021
Later among the works it cites.
Relative variational intrinsic control
Kate Baumli, David Warde-Farley, Steven Hansen, and Volodymyr Mnih · 2021
Later among the works it cites.
Is curiosity all you need? on the utility of emergent behaviours from curious exploration
Oliver Groth, Markus Wulfmeier, Giulia Vezzani, Vibhavari Dasagi, Tim Hertweck, Roland Hafner, Nicolas Heess, and Martin Riedmiller · 2021
Later among the works it cites.
Aps: Active pretraining with successor features
Hao Liu and Pieter Abbeel · 2021
Later among the works it cites.
Learning generalized gumbel-max causal mechanisms
Guy Lorberbom, Daniel D Johnson, Chris J Maddison, Daniel Tarlow, and Tamir Hazan · 2021
Later among the works it cites.
Understanding the behaviour of contrastive loss
Feng Wang and Huaping Liu · 2021
Later among the works it cites.
Posterior value functions: Hindsight baselines for policy gradient methods
Chris Nota, Philip Thomas, and Bruno C Da Silva · 2021
Later among the works it cites.
Counterfactual credit assignment in model-free reinforcement learning
Thomas Mesnard, Théophane Weber, Fabio Viola, Shantanu Thakoor, Alaa Saade, Anna Harutyunyan, Will Dabney, Tom Stepleton, Nicolas Heess, Arthur Guez, et al · 2021
Later among the works it cites.
Invariant causal imitation learning for generalizable policies
Ioana Bica, Daniel Jarrett, and Mihaela van der Schaar · 2021
Later among the works it cites.
Time-series generation by contrastive imitation
Daniel Jarrett, Ioana Bica, and Mihaela van der Schaar · 2021
Later among the works it cites.
Benchmarking the spectrum of agent capabilities
Danijar Hafner · 2021
Later among the works it cites.
Byol-explore: Exploration by bootstrapped prediction
Zhaohan Daniel Guo, Shantanu Thakoor, Miruna Pîslar, Bernardo Avila Pires, Florent Altché, Corentin Tallec, Alaa Saade, Daniele Calandriello, Jean-Bastien Grill, Yunhao Tang, et al · 2022
Closest in time.
How to stay curious while avoiding noisy tvs using aleatoric uncertainty estimation
Augustine Mavor-Parker, Kimberly Young, Caswell Barry, and Lewis Griffin · 2022
Closest in time.
Variational intrinsic control revisited
Taehwan Kwon · 2022
Closest in time.
The information geometry of unsupervised reinforcement learning
Benjamin Eysenbach, Ruslan Salakhutdinov, and Sergey Levine · 2022
Closest in time.
Cic: Contrastive intrinsic control for unsupervised skill discovery
Michael Laskin, Hao Liu, Xue Bin Peng, Denis Yarats, Aravind Rajeswaran, and Pieter Abbeel · 2022
Closest in time.
Invariant causal representation learning for generalization in imitation and reinforcement learning
Chaochao Lu, José Miguel Hernández-Lobato, and Bernhard Schölkopf · 2022
Closest in time.
Contrastive mixture of posteriors
Adam Foster, Árpi Vezér, Craig A Glastonbury, Páidí Creed, Samer Abujudeh, and Aaron Sim · 2022
Closest in time.
Exploration via elliptical episodic bonuses
Mikael Henaff, Roberta Raileanu, Minqi Jiang, and Tim Rocktäschel · 2022
Closest in time.
Blade: Robust exploration via diffusion models
Bilal Piot, Zhaohan Daniel Guo, Shantanu Thakoor, and Mohammad Gheshlaghi Azar · 2022
Closest in time.
Understanding self-predictive learning for reinforcement learning
Yunhao Tang, Zhaohan Daniel Guo, Pierre Harvey Richemond, Bernardo Ávila Pires, Yash Chandak, Rémi Munos, Mark Rowland, Mohammad Gheshlaghi Azar, Charline Le Lan, Clare Lyle, et al · 2022
Closest in time.
Mengjiao Yang, Dale Schuurmans, Pieter Abbeel, and Ofir Nachum · 2022
Closest in time.
Byol-explore: Exploration by bootstrapped prediction
Zhaohan Daniel Guo, Shantanu Thakoor, Miruna Pîslar, Bernardo Avila Pires, Florent Altché, Corentin Tallec, Alaa Saade, Daniele Calandriello, Jean-Bastien Grill, Yunhao Tang, et al · 2022
Closest in time.
How to stay curious while avoiding noisy tvs using aleatoric uncertainty estimation
Augustine Mavor-Parker, Kimberly Young, Caswell Barry, and Lewis Griffin · 2022
Closest in time.
Variational intrinsic control revisited
Taehwan Kwon · 2022
Closest in time.
The information geometry of unsupervised reinforcement learning
Benjamin Eysenbach, Ruslan Salakhutdinov, and Sergey Levine · 2022
Closest in time.
Cic: Contrastive intrinsic control for unsupervised skill discovery
Michael Laskin, Hao Liu, Xue Bin Peng, Denis Yarats, Aravind Rajeswaran, and Pieter Abbeel · 2022
Closest in time.
Invariant causal representation learning for generalization in imitation and reinforcement learning
Chaochao Lu, José Miguel Hernández-Lobato, and Bernhard Schölkopf · 2022
Closest in time.
Contrastive mixture of posteriors
Adam Foster, Árpi Vezér, Craig A Glastonbury, Páidí Creed, Samer Abujudeh, and Aaron Sim · 2022
Closest in time.
Exploration via elliptical episodic bonuses
Mikael Henaff, Roberta Raileanu, Minqi Jiang, and Tim Rocktäschel · 2022
Closest in time.
Blade: Robust exploration via diffusion models
Bilal Piot, Zhaohan Daniel Guo, Shantanu Thakoor, and Mohammad Gheshlaghi Azar · 2022
Closest in time.
Understanding self-predictive learning for reinforcement learning
Yunhao Tang, Zhaohan Daniel Guo, Pierre Harvey Richemond, Bernardo Ávila Pires, Yash Chandak, Rémi Munos, Mark Rowland, Mohammad Gheshlaghi Azar, Charline Le Lan, Clare Lyle, et al · 2022
Closest in time.
Mengjiao Yang, Dale Schuurmans, Pieter Abbeel, and Ofir Nachum · 2022
Closest in time.
Robust exploration via clustering-based online density estimation
Alaa Saade, Steven Kapturowski, Daniele Calandriello, Charles Blundell, Michal Valko, Pablo Sprechmann, and Bilal Piot · 2023
Closest in time.
Robust exploration via clustering-based online density estimation
Alaa Saade, Steven Kapturowski, Daniele Calandriello, Charles Blundell, Michal Valko, Pablo Sprechmann, and Bilal Piot · 2023
Closest in time.