Fetching the paper…
Reading the bibliography…
The past few years have seen rapid progress in combining reinforcement learning (RL) with deep learning.
Efficient exploration via state marginal matching
Lee, L., Eysenbach, B., Parisotto, E., Xing, E., Levine, S., and Salakhutdinov, R. (2019) · 1906
Earlier work this paper cites.
Using bisimulation for policy transfer in mdps
Castro, P. S., and Precup, D. (2010) · 1907
Earlier work this paper cites.
Solving rubik’s cube with a robot hand
Akkaya, I., Andrychowicz, M., Chociej, M., Litwin, M., McGrew, B., Petron, A., Paino, A., Plappert, M., Powell, G., Ribas, R., et al. (2019) · 1910
Earlier work this paper cites.
Multi-agent reinforcement learning: A selective overview of theories and algorithms
Zhang, K., Yang, Z., and Basar, T. (2019) · 1911
Earlier work this paper cites.
Dota 2 with large scale deep reinforcement learning
Berner, C., Brockman, G., Chan, B., Cheung, V., Debiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., Józefowicz, R., Gray, S., Olsson, C., Pachocki, J., Petrov, M., de Oliveira Pinto, H. P., Raiman, J., Salimans, T., Schlatter, J., Schneider, J., Sidor, S., Sutskever, I., Tang, J., Wolski, F., and Zhang, S. (2019) · 1912
Earlier work this paper cites.
An algorithm for quadratic programming
Frank, M., and Wolfe, P. (1956) · 1956
Earlier work this paper cites.
Variational empowerment as representation learning for goal-conditioned reinforcement learning
Choi, J., Sharma, A., Lee, H., Levine, S., and Gu, S. S. (2021) · 1963
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
Pomerleau, D. A. (1988) · 1988
Earlier work this paper cites.
Bisimulation through probabilistic testing
Larsen, K. G., and Skou, A. (1991) · 1991
Earlier work this paper cites.
Curious model-building control systems
Schmidhuber, J. (1991) · 1991
Earlier work this paper cites.
Nonparametric entropy estimation for stationary processes and random fields, with applications to english text
Kontoyiannis, I., Algoet, P., Suhov, Y., and Wyner, A. (1998) · 1998
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. (2020) · 2001
Earlier work this paper cites.
Recent advances in hierarchical reinforcement learning
Barto, A. G., and Mahadevan, S. (2003) · 2003
Earlier work this paper cites.
Equivalence notions and model minimization in markov decision processes
Givan, R., Dean, T., and Greig, M. (2003) · 2003
Earlier work this paper cites.
The im algorithm: a variational approach to information maximization
Agakov, D. B. F. (2004) · 2004
Earlier work this paper cites.
Metrics for finite markov decision processes.
Ferns, N., Panangaden, P., and Precup, D. (2004) · 2004
Earlier work this paper cites.
D4RL: datasets for deep data-driven reinforcement learning
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S. (2020) · 2004
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
Singh, S. P., Barto, A. G., and Chentanez, N. (2004) · 2004
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J. (2020) · 2005
Earlier work this paper cites.
The development of embodied cognition: Six lessons from babies
Smith, L., and Gasser, M. (2005) · 2005
Earlier work this paper cites.
A tutorial on energy-based learning
LeCun, Y., Chopra, S., Hadsell, R., Ranzato, M., and Huang, F. (2006) · 2006
Earlier work this paper cites.
Towards a unified theory of state abstraction for mdps.
Li, L., Walsh, T. J., and Littman, M. L. (2006) · 2006
Earlier work this paper cites.
Accelerating online reinforcement learning with offline datasets
Nair, A., Dalal, M., Gupta, A., and Levine, S. (2020) · 2006
Earlier work this paper cites.
Intrinsic motivation systems for autonomous mental development
Oudeyer, P.-Y., Kaplan, F., and Hafner, V. V. (2007) · 2007
Earlier work this paper cites.
A survey of robot learning from demonstration
Argall, B. D., Chernova, S., Veloso, M., and Browning, B. (2009) · 2008
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Strehl, A. L., and Littman, M. L. (2008) · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L., Li, K., and Li, F. (2009) · 2009
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Gutmann, M., and Hyvärinen, A. (2010) · 2010
Earlier work this paper cites.
Toward understanding natural language directions
Kollar, T., Tellex, S., Roy, D., and Roy, N. (2010) · 2010
Earlier work this paper cites.
Deep auto-encoder neural networks in reinforcement learning
Lange, S., and Riedmiller, M. A. (2010) · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G., and Bagnell, D. (2011) · 2011
Earlier work this paper cites.
Understanding natural language commands for robotic navigation and mobile manipulation
Tellex, S., Kollar, T., Dickerson, S., Walter, M., Banerjee, A., Teller, S., and Roy, N. (2011) · 2011
Earlier work this paper cites.
Towards continual reinforcement learning: A review and perspectives
Khetarpal, K., Riemer, M., Rish, I., and Precup, D. (2020) · 2012
Earlier work this paper cites.
Batch reinforcement learning
Lange, S., Gabel, T., and Riedmiller, M. (2012) · 2012
Earlier work this paper cites.
Curiosity and Motivation
Silvia, P. J. (2012) · 2012
Earlier work this paper cites.
A framework for efficient robotic manipulation
Zhan, A., Zhao, P., Pinto, L., Abbeel, P., and Laskin, M. (2020) · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P., and Welling, M. (2014) · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y. (2015) · 2015
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
Goodfellow, I. J., and Vinyals, O. (2015) · 2015
Earlier work this paper cites.
Incentivizing exploration in reinforcement learning with deep predictive models
Stadie, B. C., Levine, S., and Abbeel, P. (2015) · 2015
Earlier work this paper cites.
Embed to control: A locally linear latent dynamics model for control from raw images
Watter, M., Springenberg, J., Boedecker, J., and Riedmiller, M. (2015) · 2015
Earlier work this paper cites.
Youtube-8m: A large-scale video classification benchmark
Abu-El-Haija, S., Kothari, N., Lee, J., Natsev, P., Toderici, G., Varadarajan, B., and Vijayanarasimhan, S. (2016) · 2016
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Bellemare, M. G., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R. (2016) · 2016
Earlier work this paper cites.
Transfer from simulation to real world through learning deep inverse dynamics model
Christiano, P. F., Shah, Z., Mordatch, I., Schneider, J., Blackwell, T., Tobin, J., Abbeel, P., and Zaremba, W. (2016) · 2016
Earlier work this paper cites.
Gregor, K., Rezende, D. J., and Wierstra, D. (2016) · 2016
Earlier work this paper cites.
VIME: variational information maximizing exploration
Houthooft, R., Chen, X., Duan, Y., Schulman, J., Turck, F. D., and Abbeel, P. (2016) · 2016
Earlier work this paper cites.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Kulkarni, T. D., Narasimhan, K., Saeedi, A., and Tenenbaum, J. (2016) · 2016
Earlier work this paper cites.
Rusu, A. A., Rabinowitz, N. C., Desjardins, G., Soyer, H., Kirkpatrick, J., Kavukcuoglu, K., Pascanu, R., and Hadsell, R. (2016) · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al. (2016) · 2016
Earlier work this paper cites.
Successor features for transfer in reinforcement learning
Barreto, A., Dabney, W., Munos, R., Hunt, J. J., Schaul, T., Silver, D., and van Hasselt, H. (2017) · 2017
Earlier work this paper cites.
Density estimation using real NVP
Dinh, L., Sohl-Dickstein, J., and Bengio, S. (2017) · 2017
Earlier work this paper cites.
Count-based exploration with neural density models
Ostrovski, G., Bellemare, M. G., van den Oord, A., and Munos, R. (2017) · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T. (2017) · 2017
Earlier work this paper cites.
#exploration: A study of count-based exploration for deep reinforcement learning
Tang, H., Houthooft, R., Foote, D., Stooke, A., Chen, X., Duan, Y., Schulman, J., Turck, F. D., and Abbeel, P. (2017) · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I. (2017) · 2017
Earlier work this paper cites.
Variational option discovery algorithms
Achiam, J., Edwards, H., Amodei, D., and Abbeel, P. (2018) · 2018
Earlier work this paper cites.
Ha, D., and Schmidhuber, J. (2018) · 2018
Earlier work this paper cites.
Learning to play with intrinsically-motivated, self-aware agents
Haber, N., Mrowca, D., Wang, S., Li, F., and Yamins, D. L. (2018) · 2018
Earlier work this paper cites.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
Rajeswaran, A., Kumar, V., Gupta, A., Vezzani, G., Schulman, J., Todorov, E., and Levine, S. (2018) · 2018
Earlier work this paper cites.
Kickstarting deep reinforcement learning
Schmitt, S., Hudson, J. J., Zídek, A., Osindero, S., Doersch, C., Czarnecki, W. M., Leibo, J. Z., Küttler, H., Zisserman, A., Simonyan, K., and Eslami, S. M. A. (2018) · 2018
Earlier work this paper cites.
Time-contrastive networks: Self-supervised learning from video
Sermanet, P., Lynch, C., Chebotar, Y., Hsu, J., Jang, E., Schaal, S., Levine, S., and Brain, G. (2018) · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S., and Barto, A. G. (2018) · 2018
Earlier work this paper cites.
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., de Las Casas, D., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., Lillicrap, T. P., and Riedmiller, M. A. (2018) · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
van den Oord, A., Li, Y., and Vinyals, O. (2018) · 2018
Earlier work this paper cites.
A dissection of overfitting and generalization in continuous reinforcement learning
Zhang, A., Ballas, N., and Pineau, J. (2018) · 2018
Earlier work this paper cites.
Deep imitation learning for complex manipulation tasks from virtual reality teleoperation
Zhang, T., McCarthy, Z., Jow, O., Lee, D., Chen, X., Goldberg, K., and Abbeel, P. (2018) · 2018
Earlier work this paper cites.
Reinforcement and imitation learning for diverse visuomotor skills
Zhu, Y., Wang, Z., Merel, J., Rusu, A., Erez, T., Cabi, S., Tunyasuvunakool, S., Kramár, J., Hadsell, R., de Freitas, N., and Heess, N. (2018) · 2018
Earlier work this paper cites.
Unsupervised state representation learning in atari
Anand, A., Racah, E., Ozair, S., Bengio, Y., Côté, M., and Hjelm, R. D. (2019) · 2019
Earlier work this paper cites.
Large-scale study of curiosity-driven learning
Burda, Y., Edwards, H., Pathak, D., Storkey, A. J., Darrell, T., and Efros, A. A. (2019a) · 2019
Cited alongside, same era.
Exploration by random network distillation
Burda, Y., Edwards, H., Storkey, A. J., and Klimov, O. (2019b) · 2019
Cited alongside, same era.
Exploring the limitations of behavior cloning for autonomous driving
Codevilla, F., Santana, E., Lopez, A. M., and Gaidon, A. (2019) · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2019) · 2019
Cited alongside, same era.
Diversity is all you need: Learning skills without a reward function
Eysenbach, B., Gupta, A., Ibarz, J., and Levine, S. (2019) · 2019
Cited alongside, same era.
Deepmdp: Learning continuous latent space models for representation learning
Gelada, C., Kumar, S., Buckman, J., Nachum, O., and Bellemare, M. G. (2019) · 2019
Aleatoric and epistemic uncertainty in machine learning: An introduction to concepts and methods
Hüllermeier, E., and Waegeman, W. (2021) · 2021
Later among the works it cites.
Offline reinforcement learning as one big sequence modeling problem
Janner, M., Li, Q., and Levine, S. (2021) · 2021
Later among the works it cites.
Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms
Jin, C., Liu, Q., and Miryoosefi, S. (2021) · 2021
Later among the works it cites.
A survey of generalisation in deep reinforcement learning
Kirk, R., Zhang, A., Grefenstette, E., and Rocktäschel, T. (2021) · 2021
Later among the works it cites.
Urlb: Unsupervised reinforcement learning benchmark
Laskin, M., Yarats, D., Liu, H., Lee, K., Zhan, A., Lu, K., Cang, C., Pinto, L., and Abbeel, P. (2021) · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Provably efficient maximum entropy exploration
Hazan, E., Kakade, S. M., Singh, K., and Soest, A. V. (2019) · 2019
Cited alongside, same era.
Parameter-efficient transfer learning for NLP
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S. (2019) · 2019
Cited alongside, same era.
Compile: Compositional imitation learning and execution
Kipf, T., Li, Y., Dai, H., Zambaldi, V. F., Sanchez-Gonzalez, A., Grefenstette, E., Kohli, P., and Battaglia, P. W. (2019) · 2019
Cited alongside, same era.
Stabilizing off-policy q-learning via bootstrapping error reduction
Kumar, A., Fu, J., Soh, M., Tucker, G., and Levine, S. (2019) · 2019
Cited alongside, same era.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Lu, J., Batra, D., Parikh, D., and Lee, S. (2019) · 2019
Cited alongside, same era.
Near-optimal representation learning for hierarchical reinforcement learning
Nachum, O., Gu, S., Lee, H., and Levine, S. (2019) · 2019
Cited alongside, same era.
The power of scale for parameter-efficient prompt tuning
Lester, B., Al-Rfou, R., and Constant, N. (2021) · 2021
Later among the works it cites.
Understanding the world through action
Levine, S. (2021) · 2021
Later among the works it cites.
Celebrating diversity in shared multi-agent reinforcement learning
Li, C., Wang, T., Wu, C., Zhao, Q., Yang, J., and Zhang, C. (2021) · 2021
Later among the works it cites.
APS: active pretraining with successor features
Liu, H., and Abbeel, P. (2021a) · 2021
Later among the works it cites.
Offline pre-trained multi-agent decision transformer: One big sequence model tackles all SMAC tasks
Meng, L., Wen, M., Yang, Y., Le, C., Li, X., Zhang, W., Wen, Y., Zhang, H., Wang, J., and Xu, B. (2021) · 2021
Later among the works it cites.
Emergent social learning via multi-agent reinforcement learning
Ndousse, K. K., Eck, D., Levine, S., and Jaques, N. (2021) · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. (2021) · 2021
Later among the works it cites.
A call to build models like we build open-source software.
Raffel, C. (2021) · 2021
Later among the works it cites.
Which mutual-information representation learning objectives are sufficient for control?
Rakelly, K., Gupta, A., Florensa, C., and Levine, S. (2021) · 2021
Later among the works it cites.
Data-efficient reinforcement learning with self-predictive representations
Schwarzer, M., Anand, A., Goel, R., Hjelm, R. D., Courville, A. C., and Bachman, P. (2021a) · 2021
Later among the works it cites.
State entropy maximization with random encoders for efficient exploration
Seo, Y., Chen, L., Shin, J., Lee, H., Abbeel, P., and Lee, K. (2021) · 2021
Later among the works it cites.
RRL: resnet as representation for reinforcement learning
Shah, R. M., and Kumar, V. (2021) · 2021
Later among the works it cites.
Parrot: Data-driven behavioral priors for reinforcement learning
Singh, A., Liu, H., Zhou, G., Yu, A., Rhinehart, N., and Levine, S. (2021) · 2021
Later among the works it cites.
Unsupervised learning for reinforcement learning.
Srinivas, A., and Abbeel, P. (2021) · 2021
Later among the works it cites.
Decoupling representation learning from reinforcement learning
Stooke, A., Lee, K., Abbeel, P., and Laskin, M. (2021) · 2021
Later among the works it cites.
Representation matters: Offline pretraining for sequential decision making
Yang, M., and Nachum, O. (2021) · 2021
Later among the works it cites.
Reinforcement learning with prototypical representations
Yarats, D., Fergus, R., Lazaric, A., and Pinto, L. (2021) · 2021
Later among the works it cites.
Learning invariant representations for reinforcement learning without reconstruction
Zhang, A., McAllister, R. T., Calandra, R., Gal, Y., and Levine, S. (2021a) · 2021
Later among the works it cites.
Hierarchical reinforcement learning by discovering intrinsic options
Zhang, J., Yu, H., and Xu, W. (2021c) · 2021
Later among the works it cites.
Provable benefits of representational transfer in reinforcement learning
Agarwal, A., Song, Y., Sun, W., Wang, K., Wang, M., and Zhang, X. (2022) · 2022
Closest in time.
Reincarnating reinforcement learning: Reusing prior computation to accelerate progress.
Agarwal, R., Schwarzer, M., Castro, P. S., Courville, A., and Bellemare, M. G. (2022) · 2022
Closest in time.
Do as i can, not as i say: Grounding language in robotic affordances
Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., et al. (2022) · 2022
Closest in time.
Video pretraining (vpt): Learning to act by watching unlabeled online videos
Baker, B., Akkaya, I., Zhokhov, P., Huizinga, J., Tang, J., Ecoffet, A., Houghton, B., Sampedro, R., and Clune, J. (2022) · 2022
Closest in time.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al. (2022) · 2022
Closest in time.
Autotelic agents with intrinsically motivated goal-conditioned reinforcement learning: A short survey
Colas, C., Karch, T., Sigaud, O., and Oudeyer, P.-Y. (2022) · 2022
Closest in time.
The information geometry of unsupervised reinforcement learning
Eysenbach, B., Salakhutdinov, R., and Levine, S. (2022) · 2022
Closest in time.
Minedojo: Building open-ended embodied agents with internet-scale knowledge
Fan, L., Wang, G., Jiang, Y., Mandlekar, A., Yang, Y., Zhu, H., Tang, A., Huang, D.-A., Zhu, Y., and Anandkumar, A. (2022) · 2022
Closest in time.
Implicit behavioral cloning
Florence, P., Lynch, C., Zeng, A., Ramirez, O. A., Wahid, A., Downs, L., Wong, A., Lee, J., Mordatch, I., and Tompson, J. (2022) · 2022
Closest in time.
Generalized decision transformer for offline hindsight information matching
Furuta, H., Matsuo, Y., and Gu, S. S. (2022) · 2022
Closest in time.
Hierarchical few-shot imitation with skill transition models
Hakhamaneshi, K., Zhao, R., Zhan, A., Abbeel, P., and Laskin, M. (2022) · 2022
Closest in time.
Towards a unified view of parameter-efficient transfer learning
He, J., Zhou, C., Ma, X., Berg-Kirkpatrick, T., and Neubig, G. (2022) · 2022
Closest in time.
Adarl: What, where, and how to adapt in transfer reinforcement learning
Huang, B., Feng, F., Lu, C., Magliacane, S., and Zhang, K. (2022a) · 2022
Closest in time.
Perceiver IO: A general architecture for structured inputs & outputs
Jaegle, A., Borgeaud, S., Alayrac, J.-B., Doersch, C., Ionescu, C., Ding, D., Koppula, S., Zoran, D., Brock, A., Shelhamer, E., Henaff, O. J., Botvinick, M., Zisserman, A., Vinyals, O., and Carreira, J. (2022) · 2022
Closest in time.
Filtered-cophy: Unsupervised learning of counterfactual physics in pixel space
Janny, S., Baradel, F., Neverova, N., Nadri, M., Mori, G., and Wolf, C. (2022) · 2022
Closest in time.
Direct then diffuse: Incremental unsupervised skill discovery for state covering and goal reaching
Kamienny, P.-A., Tarbouriech, J., Lazaric, A., and Denoyer, L. (2022) · 2022
Closest in time.
Human-level atari 200x faster.
Kapturowski, S., Campos, V., Jiang, R., Rakićević, N., van Hasselt, H., Blundell, C., and Badia, A. P. (2022) · 2022
Closest in time.
Offline reinforcement learning with implicit q-learning
Kostrikov, I., Nair, A., and Levine, S. (2022) · 2022
Closest in time.
The challenges of exploration for offline reinforcement learning
Lambert, N., Wulfmeier, M., Whitney, W., Byravan, A., Bloesch, M., Dasagi, V., Hertweck, T., and Riedmiller, M. (2022) · 2022
Closest in time.
CIC: contrastive intrinsic control for unsupervised skill discovery
Laskin, M., Liu, H., Peng, X. B., Yarats, D., Rajeswaran, A., and Abbeel, P. (2022) · 2022
Closest in time.
Multi-game decision transformers
Lee, K.-H., Nachum, O., Yang, M., Lee, L., Freeman, D., Xu, W., Guadarrama, S., Fischer, I., Jang, E., Michalewski, H., et al. (2022) · 2022
Closest in time.
Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble
Lee, S., Seo, Y., Lee, K., Abbeel, P., and Shin, J. (2022) · 2022
Closest in time.
Aw-opt: Learning robotic skills with imitation and reinforcement at scale
Lu, Y., Hausman, K., Chebotar, Y., Yan, M., Jang, E., Herzog, A., Xiao, T., Irpan, A., Khansari, M., Kalashnikov, D., and Levine, S. (2022) · 2022
Closest in time.
Zero-shot reward specification via grounded natural language
Mahmoudieh, P., Pathak, D., and Darrell, T. (2022) · 2022
Closest in time.
How to stay curious while avoiding noisy tvs using aleatoric uncertainty estimation
Mavor-Parker, A., Young, K., Barry, C., and Griffin, L. (2022) · 2022
Closest in time.
Multimodal contrastive learning with limoe: the language-image mixture of experts
Mustafa, B., Riquelme, C., Puigcerver, J., Jenatton, R., and Houlsby, N. (2022) · 2022
Closest in time.
The unsurprising effectiveness of pre-trained vision models for control
Parisi, S., Rajeswaran, A., Purushwalkam, S., and Gupta, A. (2022) · 2022
Closest in time.
Lipschitz-constrained unsupervised skill discovery
Park, S., Choi, J., Kim, J., Lee, H., and Kim, G. (2022) · 2022
Closest in time.
Reed, S., Zolna, K., Parisotto, E., Colmenarejo, S. G., Novikov, A., Barth-Maron, G., Gimenez, M., Sulsky, Y., Kay, J., Springenberg, J. T., et al. (2022) · 2022
Closest in time.
Can wikipedia help offline reinforcement learning?
Reid, M., Yamada, Y., and Gu, S. S. (2022) · 2022
Closest in time.
Reinforcement learning with action-free pre-training from videos
Seo, Y., Lee, K., James, S., and Abbeel, P. (2022) · 2022
Closest in time.
Semantic exploration from language abstractions and pretrained representations
Tam, A. C., Rabinowitz, N. C., Lampinen, A. K., Roy, N. A., Chan, S. C., Strouse, D., Wang, J. X., Banino, A., and Hill, F. (2022) · 2022
Closest in time.
The role of pretrained representations for the OOD generalization of RL agents
Träuble, F., Dittadi, A., Wuthrich, M., Widmaier, F., Gehler, P. V., Winther, O., Locatello, F., Bachem, O., Schölkopf, B., and Bauer, S. (2022) · 2022
Closest in time.
Image as a foreign language: Beit pretraining for all vision and vision-language tasks
Wang, W., Bao, H., Dong, L., Bjorck, J., Peng, Z., Liu, Q., Aggarwal, K., Mohammed, O. K., Singhal, S., Som, S., and Wei, F. (2022) · 2022
Closest in time.
Finetuned language models are zero-shot learners
Wei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V. (2022) · 2022
Closest in time.
Masked visual pre-training for motor control
Xiao, T., Radosavovic, I., Darrell, T., and Malik, J. (2022) · 2022
Closest in time.
Prompting decision transformer for few-shot policy generalization
Xu, M., Shen, Y., Zhang, S., Lu, Y., Zhao, D., Tenenbaum, B. J., and Gan, C. (2022) · 2022
Closest in time.
Task-induced representation learning
Yamada, J., Pertsch, K., Gunjal, A., and Lim, J. J. (2022) · 2022
Closest in time.
TRAIL: Near-optimal imitation learning with suboptimal data
Yang, M., Levine, S., and Nachum, O. (2022) · 2022
Closest in time.
Don’t change the algorithm, change the data: Exploratory data for offline reinforcement learning
Yarats, D., Brandfonbrener, D., Liu, H., Laskin, M., Abbeel, P., Lazaric, A., and Pinto, L. (2022) · 2022
Closest in time.
Learning to drive by watching youtube videos: Action-conditioned contrastive policy pretraining.
Zhang, Q., Peng, Z., and Zhou, B. (2022) · 2022
Closest in time.
Online decision transformer
Zheng, Q., Zhang, A., and Grover, A. (2022) · 2022
Closest in time.
Leveraging procedural generation to benchmark reinforcement learning
Cobbe, K., Hesse, C., Hilton, J., and Schulman, J. (2020) · 2056
Closest in time.